Files
parsedmarc/pyproject.toml
T
07bca1ad28 Make the output and mailbox integrations optional extras (#888)
* Make the output and mailbox integrations optional extras (#883)

Breaking change for the next major release: pip install parsedmarc now
installs the parsing core plus a working core CLI (file, IMAP, Maildir,
and mbox input; CSV/JSON, Splunk HEC, webhook, and syslog output).
Everything else moves behind an extra: elastic, opensearch, kafka, s3,
gelf, loganalytics, msgraph, and gmail, joining the existing postgresql
extra, with an umbrella [all] that deliberately excludes postgresql
(psycopg's binary wheels do not exist on every platform, so
parsedmarc[all] must never fail to install there).

cli.py imports the six SDK-dependent output modules behind the #884
TYPE_CHECKING/try-except guard; a configured section whose extra is
missing fails fast with a ConfigurationError naming the section and the
exact pip install command — including the msgraph and gmail_api mailbox
sections (detected via parsedmarc.mail's placeholder classes) and
postgresql (checked before the constructor so the startup retry loop
does not retry a missing dependency for a minute). The Azure/kiota Graph
error types fall back to never-raised sentinel classes.

The Docker image installs [all,postgresql], so container users see no
change. CI lint installs [build,all,postgresql]; the unit-test job
installs [build,all], deliberately without postgresql so
test_postgres.py's absent-psycopg arm stays exercised. The
never-imported dateparser dependency is dropped in favor of declaring
python-dateutil, which utils.py actually imports; pytz moves to the
build extra for the one test that uses it.

Verified live: a no-extras wheel install imports, parses samples, and
reports the install hint for each gated section; a [all] install
restores every integration; the Docker image builds with every SDK
importable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Patch psycopg presence in the PostgreSQL CLI wiring tests

CI's unit-test job deliberately installs [build,all] without the
postgresql extra, so parsedmarc.cli.postgres.psycopg is None there and
the new missing-extra presence check correctly made _main exit 1 before
the wiring under test ran. The tests simulate the SDK being available
(PostgreSQLClient is mocked at the SDK boundary), so the module-level
psycopg handle is now patched present in setUp. Verified against a
simulated psycopg-absent environment as well as the local full install.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address Copilot review: narrow guards to ModuleNotFoundError, fix docs

- The optional-integration and Graph error-type import guards now catch
  ModuleNotFoundError instead of ImportError, so only a genuinely absent
  package reads as a missing extra; a broken-but-present SDK fails
  loudly with its real error instead of masquerading as one. The test
  blocker raises ModuleNotFoundError accordingly — the exact exception a
  missing package produces.
- _missing_extra_hint docstring no longer calls every gated integration
  an output module (it also serves the msgraph/gmail_api mailbox
  sections).
- Fix the pre-existing passsword typo in usage.md's kafka section; the
  INI key the code reads is password (cli.py _parse_config).

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Quote extras specs in copy-paste install commands

From Copilot's second review round: zsh treats an unquoted .[build,all]
as a glob and fails with 'no matches found', so the commands shown in
AGENTS.md, CONTRIBUTING.md, dashboards/README.md, and the bootstrap
script's comment are now quoted. The CI workflows keep the unquoted
form: they run under bash, which passes unmatched globs through
literally. The suggestion to change the 'Choosing what to install'
heading level was rejected — it is a subsection of 'Installing
parsedmarc', matching the file's existing hierarchy.

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix upgrade command in the changelog

* Documentation review: accuracy, spelling, grammar, and clarity pass

A full prose review of docs/source, README, CONTRIBUTING, and the
dashboards README, with every accuracy claim verified against the code
before changing it. Highlights:

- usage.md: documented six missing [general] options (the CSV/JSON
  filename options, prettify_json, normalize_timespan_threshold_hours),
  the required kafka smtp_tls_topic, [imap] timeout/max_retries, and
  the postgresql env-var prefix; corrected the maildir_path default
  (None, not INBOX — cli.py Namespace defaults), the mailbox
  check_timeout option name, the systemd restart interval (RestartSec
  is 5m), and merged the duplicate silent entry; quoted every
  copy-paste extras spec for zsh safety.
- elasticsearch.md: fixed an invalid openssl command (rsa:4096 -nodes),
  the dashboards filename (opensearch_dashboards.ndjson, matching the
  file the link serves), and assorted grammar.
- davmail.md: the service-enable command now enables davmail.service
  (was parsedmarc.service — a copy-paste error that left DavMail
  unenabled), plus a view typo and DavMail capitalization.
- output.md: the example schema reference is RFC 7489 Appendix C
  (7480 is RDAP). kibana.md: SPF relies on the SMTP envelope, not
  session headers (RFC 7208). dmarc.md: DKM -> DKIM.
- README: the intro now also names the OpenSearch/Grafana stack,
  matching the feature list. CONTRIBUTING: pre-PR checks now include
  ruff format --check and pyright, matching CI's lint job.
- dashboards/README: the service table and seed description now include
  the PostgreSQL backend the compose stack runs.

Sample data blocks, the CLI-help mirror block, and released CHANGELOG
entries were deliberately left untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Docstring review: accuracy, spelling, grammar, and clarity pass

Every docstring in parsedmarc/, parsedmarc/mail/, the maps maintainer
scripts, and the test suite reviewed with each claim verified against
the code it documents. Text-only — no behavior changes. Highlights:

- Copy-paste errors corrected: parsed_smtp_tls_reports_to_csv and
  splunk/loganalytics save functions described aggregate or failure
  reports they do not handle; LogAnalyticsException claimed to be an
  Elasticsearch error.
- Docstring/behavior mismatches: parse_report_email's report_type
  enumeration omitted smtp_tls; parse_failure_report typed msg_date as
  str (it is datetime); strip_attachment_payloads claimed payloads are
  replaced with None (the key is deleted); kafkaclient's failure and
  SMTP TLS savers claimed per-record slicing while sending the whole
  list in one message (docstrings now describe reality — whether
  slicing was intended is flagged for follow-up); the postgres savers
  claimed to take parse_report_file's return value but receive the
  inner report dict; elastic/opensearch save functions' Raises listed
  only AlreadySaved.
- None-as-semantic-state documented where missing (get_base_domain,
  get_ip_address_country), enumeration completeness fixed
  (get_ip_address_info's 9 result keys, maps script outputs, TSV
  columns), and the stale 44-industry-types count corrected to the
  46 the authoritative README list defines.
- Test docstrings aligned with what the tests actually assert,
  including two that overstated coverage of the elastic/opensearch
  address-list tests.
- Two argparse help strings fixed: file_path now names SMTP TLS report
  files alongside aggregate and failure, mirrored into usage.md's
  CLI-help block; --offline's doubled spaces removed (rendered help
  unchanged).
- elasticsearch.md's security claim corrected against Elastic's docs:
  security is enabled and auto-configured on first startup since 8.0
  (not "8.7 secure mode"), so the settings are verified, not
  hand-written.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-08-28 17:33:18 -04:00

252 lines
9.4 KiB
TOML

[build-system]
requires = [
"hatchling>=1.27.0",
]
requires_python = ">=3.10,<3.15"
build-backend = "hatchling.build"
[project]
name = "parsedmarc"
dynamic = [
"version",
]
description = "A Python package and CLI for parsing aggregate, failure, and SMTP TLS DMARC reports"
readme = "README.md"
license = "Apache-2.0"
authors = [
{ name = "Sean Whalen", email = "whalenster@gmail.com" },
]
keywords = [
"DMARC",
"parser",
"reporting",
]
classifiers = [
"Development Status :: 5 - Production/Stable",
"Intended Audience :: Developers",
"Intended Audience :: Information Technology",
"License :: OSI Approved :: Apache Software License",
"Operating System :: OS Independent",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3 :: Only",
"Programming Language :: Python :: 3.10",
"Programming Language :: Python :: 3.11",
"Programming Language :: Python :: 3.12",
"Programming Language :: Python :: 3.13",
"Programming Language :: Python :: 3.14",
]
requires-python = ">=3.10"
# The base install is the parsing core plus a working core CLI: file,
# IMAP, Maildir, and mbox input; CSV/JSON, Splunk HEC, webhook, and
# syslog output. Every other output and mailbox integration lives in an extra
# below, so a library or small-mail-host install does not have to carry
# the Elasticsearch, OpenSearch, Kafka, AWS, Azure, Gmail, and Microsoft
# Graph SDKs (see issue #883).
dependencies = [
# The [doh] extra supplies the httpx/h2 floors DNS over HTTPS needs;
# 2.7.0 is the floor verified against the dns.nameserver and
# dns.query.https(session=...) APIs utils.py builds on.
"dnspython[doh]>=2.7.0",
"expiringdict>=1.1.4",
# The runtime HTTP library (utils.py fetches, webhook and Splunk HEC
# clients, Graph error handling in cli.py). The floor matches
# microsoft-kiota-http's own requirement.
"httpx>=0.25",
"lxml>=4.4.0",
# No extras: the base mailsuite supplies IMAP and Maildir connections
# plus mail-parser. Gmail and Microsoft Graph come from the gmail and
# msgraph extras below.
"mailsuite>=2.3.1",
"maxminddb>=2.0.0",
"publicsuffixlist>=0.10.0",
# Imported directly at utils.py:39 (dateutil.parser). It used to
# arrive transitively via dateparser, which nothing ever imported.
"python-dateutil>=2.8.0",
"tqdm>=4.31.1",
"xmltodict>=0.12.0",
"PyYAML>=6.0.3"
]
# Splunk HEC, webhook, and syslog outputs deliberately have no extra:
# they need only httpx (a base dependency) and the standard library, so
# do not add empty marker extras for them.
[project.optional-dependencies]
elastic = [
"elasticsearch>=8.18,<9",
]
opensearch = [
"opensearch-py>=2.4.2,<=4.0.0",
# boto3 supplies the SigV4 signer opensearch.py imports for AWS auth.
"boto3>=1.16.63",
]
kafka = [
"kafka-python>=2.3.2",
]
s3 = [
"boto3>=1.16.63",
]
gelf = [
"pygelf>=0.4.2",
]
loganalytics = [
"azure-identity>=1.8.0",
"azure-monitor-ingestion>=1.0.0",
]
msgraph = [
"mailsuite[msgraph]>=2.3.1",
# Imported directly in cli.py for Graph error handling; otherwise
# only a transitive dep of mailsuite[msgraph] -> msgraph-sdk.
"microsoft-kiota-abstractions>=1.8.0",
]
gmail = [
"mailsuite[gmail]>=2.3.1",
]
# `all` must stay the union of every extra above — and must keep
# excluding `postgresql`: psycopg's prebuilt binary wheels do not exist
# for every platform/arch, so folding it in here would make
# `pip install parsedmarc[all]` fail on platforms it does not cover.
# Spelled out as an explicit package list rather than a self-referential
# `parsedmarc[elastic,...]` spec, which would make the project a
# dependency of itself; keep this list in sync when an extra above
# gains, drops, or re-floors a package.
all = [
"azure-identity>=1.8.0",
"azure-monitor-ingestion>=1.0.0",
"boto3>=1.16.63",
"elasticsearch>=8.18,<9",
"kafka-python>=2.3.2",
"mailsuite[gmail,msgraph]>=2.3.1",
"microsoft-kiota-abstractions>=1.8.0",
"opensearch-py>=2.4.2,<=4.0.0",
"pygelf>=0.4.2",
]
postgresql = [
# Optional output backend. psycopg ships prebuilt binary wheels via the
# [binary] extra, but those wheels don't exist for every platform/arch,
# so PostgreSQL support is opt-in rather than a mandatory dependency.
"psycopg[binary]>=3.1.0",
]
build = [
# Used only by maintainer tooling under parsedmarc/resources/maps/ —
# `collect_domain_info.py --use-search-fallback` falls back to a
# DuckDuckGo search when the homepage fetch returns a bot-block / parked
# / empty page. Optional import; the script runs without it as long as
# the fallback flag isn't passed.
"ddgs>=9.0.0",
"hatch>=1.14.0",
"myst-parser[linkify]",
"nose",
# Pinned exactly: pyright's checks evolve between releases, so an
# unpinned version could break CI without any code change. Bump
# deliberately (and fix any new findings) rather than implicitly.
"pyright==1.1.411",
"pytest",
"pytest-cov",
# Test-only: tests/test_init.py imports pytz for a fixed-offset
# timezone. It used to arrive transitively via dateparser, which the
# base dependencies no longer declare.
"pytz",
# Used only by the out-of-wheel maintainer script
# parsedmarc/resources/maps/collect_domain_info.py, which deliberately
# stays on requests because its permissive-TLS fallback is built on
# urllib3's HTTPAdapter machinery.
"requests>=2.22.0",
# Pinned exactly for the same reason as pyright: ruff's default rule
# set evolves between releases (e.g. 0.16.0 began flagging the
# str.format() style this codebase used until then), so an unpinned
# version breaks CI without any code change. Bump deliberately and fix
# any new findings in the same PR as the bump.
"ruff==0.16.0",
"sphinx",
"sphinx_rtd_theme",
]
[project.scripts]
parsedmarc = "parsedmarc.cli:_main"
[project.urls]
Homepage = "https://domainaware.github.io/parsedmarc"
[tool.hatch.version]
path = "parsedmarc/constants.py"
[tool.hatch.build.targets.sdist]
include = [
"/parsedmarc",
]
[tool.hatch.build]
exclude = [
"base_reverse_dns.csv",
"unknown_base_reverse_dns.csv",
"README.md",
"AGENTS.md",
"CLAUDE.md",
"*.bak",
# Maintenance tooling: any Python file under parsedmarc/resources/maps/
# whose name doesn't start with `_` (i.e. everything except __init__.py,
# which must keep shipping for `importlib.resources.files()` lookups).
"parsedmarc/resources/maps/[!_]*.py",
]
[tool.ruff.lint]
# The rule set is selected explicitly rather than floating on ruff's
# defaults: ruff 0.16.0 expanded the default selection from the
# long-standing E4/E7/E9/F to many more rule families (BLE, SIM, C4, DTZ,
# I, PL, S, ...), some of which conflict with deliberate house style —
# e.g. BLE001 flags the parser's intentional broad catches that keep one
# malformed report from crashing a batch. E4/E7/E9/F is the pre-0.16
# default set; adopting any of the new families is a deliberate
# per-family decision, made here with a comment, not an upgrade side
# effect.
#
# UP006/UP007/UP035/UP045 enforce modern type-hint syntax: with
# requires-python >=3.10, PEP 585 builtins (list[int]) and PEP 604 unions
# (X | Y, X | None) are available, so keep the deprecated typing.List /
# Union / Optional spellings out of the codebase.
select = [
"E4", # import rules from pycodestyle
"E7", # statement rules from pycodestyle
"E9", # runtime/syntax error rules from pycodestyle
"F", # pyflakes
"UP006", # non-pep585-annotation: List -> list, Dict -> dict
"UP007", # non-pep604-annotation-union: Union[X, Y] -> X | Y
"UP030", # format-literals: "{0}".format(x) -> "{}".format(x)
"UP032", # f-string: "{}".format(x) -> f"{x}"
"UP035", # deprecated-import: typing.List etc. / typing -> collections.abc
"UP045", # non-pep604-annotation-optional: Optional[X] -> X | None
]
[tool.pyright]
# The whole codebase passes pyright with zero errors and warnings; CI
# enforces this (see .github/workflows/python-tests.yml). Run locally with
# `pyright` from the repo root. Requires the [all] and [postgresql] extras
# to be installed so the optional integration imports (parsedmarc/cli.py's
# guarded output modules, the psycopg import in parsedmarc/postgres.py)
# resolve.
include = ["parsedmarc", "tests", "docs"]
typeCheckingMode = "standard"
[tool.pytest.ini_options]
# Default to the per-module test layout under tests/. New tests should go
# into tests/test_<module>.py to match the file they exercise; do not
# reintroduce a monolithic tests.py.
testpaths = ["tests"]
[tool.coverage.run]
# Coverage measures shipped code only. Master's reported ≈66.9% on
# Codecov was an artefact of the old monolithic tests.py having no
# [tool.coverage.run] block, which let coverage's default behaviour
# measure every file imported during the run — including the test file
# itself at ~99% "covered". That inflated the headline by ~8 percentage
# points without any actual testing signal. Restricting to the parsedmarc
# package gives a meaningful number that tracks how much of the shipped
# library the test suite actually exercises.
source = ["parsedmarc"]
# Maintainer-only batch scripts under parsedmarc/resources/maps/ ship
# out of the wheel (see the [tool.hatch.build] exclude block above) —
# omit them so the headline number reflects only installed library code.
omit = [
"*/parsedmarc/resources/maps/*.py",
]