mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-09-05 13:38:00 +00:00
* Make the output and mailbox integrations optional extras (#883) Breaking change for the next major release: pip install parsedmarc now installs the parsing core plus a working core CLI (file, IMAP, Maildir, and mbox input; CSV/JSON, Splunk HEC, webhook, and syslog output). Everything else moves behind an extra: elastic, opensearch, kafka, s3, gelf, loganalytics, msgraph, and gmail, joining the existing postgresql extra, with an umbrella [all] that deliberately excludes postgresql (psycopg's binary wheels do not exist on every platform, so parsedmarc[all] must never fail to install there). cli.py imports the six SDK-dependent output modules behind the #884 TYPE_CHECKING/try-except guard; a configured section whose extra is missing fails fast with a ConfigurationError naming the section and the exact pip install command — including the msgraph and gmail_api mailbox sections (detected via parsedmarc.mail's placeholder classes) and postgresql (checked before the constructor so the startup retry loop does not retry a missing dependency for a minute). The Azure/kiota Graph error types fall back to never-raised sentinel classes. The Docker image installs [all,postgresql], so container users see no change. CI lint installs [build,all,postgresql]; the unit-test job installs [build,all], deliberately without postgresql so test_postgres.py's absent-psycopg arm stays exercised. The never-imported dateparser dependency is dropped in favor of declaring python-dateutil, which utils.py actually imports; pytz moves to the build extra for the one test that uses it. Verified live: a no-extras wheel install imports, parses samples, and reports the install hint for each gated section; a [all] install restores every integration; the Docker image builds with every SDK importable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Patch psycopg presence in the PostgreSQL CLI wiring tests CI's unit-test job deliberately installs [build,all] without the postgresql extra, so parsedmarc.cli.postgres.psycopg is None there and the new missing-extra presence check correctly made _main exit 1 before the wiring under test ran. The tests simulate the SDK being available (PostgreSQLClient is mocked at the SDK boundary), so the module-level psycopg handle is now patched present in setUp. Verified against a simulated psycopg-absent environment as well as the local full install. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot review: narrow guards to ModuleNotFoundError, fix docs - The optional-integration and Graph error-type import guards now catch ModuleNotFoundError instead of ImportError, so only a genuinely absent package reads as a missing extra; a broken-but-present SDK fails loudly with its real error instead of masquerading as one. The test blocker raises ModuleNotFoundError accordingly — the exact exception a missing package produces. - _missing_extra_hint docstring no longer calls every gated integration an output module (it also serves the msgraph/gmail_api mailbox sections). - Fix the pre-existing passsword typo in usage.md's kafka section; the INI key the code reads is password (cli.py _parse_config). Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Quote extras specs in copy-paste install commands From Copilot's second review round: zsh treats an unquoted .[build,all] as a glob and fails with 'no matches found', so the commands shown in AGENTS.md, CONTRIBUTING.md, dashboards/README.md, and the bootstrap script's comment are now quoted. The CI workflows keep the unquoted form: they run under bash, which passes unmatched globs through literally. The suggestion to change the 'Choosing what to install' heading level was rejected — it is a subsection of 'Installing parsedmarc', matching the file's existing hierarchy. Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix upgrade command in the changelog * Documentation review: accuracy, spelling, grammar, and clarity pass A full prose review of docs/source, README, CONTRIBUTING, and the dashboards README, with every accuracy claim verified against the code before changing it. Highlights: - usage.md: documented six missing [general] options (the CSV/JSON filename options, prettify_json, normalize_timespan_threshold_hours), the required kafka smtp_tls_topic, [imap] timeout/max_retries, and the postgresql env-var prefix; corrected the maildir_path default (None, not INBOX — cli.py Namespace defaults), the mailbox check_timeout option name, the systemd restart interval (RestartSec is 5m), and merged the duplicate silent entry; quoted every copy-paste extras spec for zsh safety. - elasticsearch.md: fixed an invalid openssl command (rsa:4096 -nodes), the dashboards filename (opensearch_dashboards.ndjson, matching the file the link serves), and assorted grammar. - davmail.md: the service-enable command now enables davmail.service (was parsedmarc.service — a copy-paste error that left DavMail unenabled), plus a view typo and DavMail capitalization. - output.md: the example schema reference is RFC 7489 Appendix C (7480 is RDAP). kibana.md: SPF relies on the SMTP envelope, not session headers (RFC 7208). dmarc.md: DKM -> DKIM. - README: the intro now also names the OpenSearch/Grafana stack, matching the feature list. CONTRIBUTING: pre-PR checks now include ruff format --check and pyright, matching CI's lint job. - dashboards/README: the service table and seed description now include the PostgreSQL backend the compose stack runs. Sample data blocks, the CLI-help mirror block, and released CHANGELOG entries were deliberately left untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Docstring review: accuracy, spelling, grammar, and clarity pass Every docstring in parsedmarc/, parsedmarc/mail/, the maps maintainer scripts, and the test suite reviewed with each claim verified against the code it documents. Text-only — no behavior changes. Highlights: - Copy-paste errors corrected: parsed_smtp_tls_reports_to_csv and splunk/loganalytics save functions described aggregate or failure reports they do not handle; LogAnalyticsException claimed to be an Elasticsearch error. - Docstring/behavior mismatches: parse_report_email's report_type enumeration omitted smtp_tls; parse_failure_report typed msg_date as str (it is datetime); strip_attachment_payloads claimed payloads are replaced with None (the key is deleted); kafkaclient's failure and SMTP TLS savers claimed per-record slicing while sending the whole list in one message (docstrings now describe reality — whether slicing was intended is flagged for follow-up); the postgres savers claimed to take parse_report_file's return value but receive the inner report dict; elastic/opensearch save functions' Raises listed only AlreadySaved. - None-as-semantic-state documented where missing (get_base_domain, get_ip_address_country), enumeration completeness fixed (get_ip_address_info's 9 result keys, maps script outputs, TSV columns), and the stale 44-industry-types count corrected to the 46 the authoritative README list defines. - Test docstrings aligned with what the tests actually assert, including two that overstated coverage of the elastic/opensearch address-list tests. - Two argparse help strings fixed: file_path now names SMTP TLS report files alongside aggregate and failure, mirrored into usage.md's CLI-help block; --offline's doubled spaces removed (rendered help unchanged). - elasticsearch.md's security claim corrected against Elastic's docs: security is enabled and auto-configured on first startup since 8.0 (not "8.7 secure mode"), so the settings are verified, not hand-written. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
252 lines
9.4 KiB
TOML
252 lines
9.4 KiB
TOML
[build-system]
|
|
requires = [
|
|
"hatchling>=1.27.0",
|
|
]
|
|
requires_python = ">=3.10,<3.15"
|
|
build-backend = "hatchling.build"
|
|
|
|
[project]
|
|
name = "parsedmarc"
|
|
dynamic = [
|
|
"version",
|
|
]
|
|
description = "A Python package and CLI for parsing aggregate, failure, and SMTP TLS DMARC reports"
|
|
readme = "README.md"
|
|
license = "Apache-2.0"
|
|
authors = [
|
|
{ name = "Sean Whalen", email = "whalenster@gmail.com" },
|
|
]
|
|
keywords = [
|
|
"DMARC",
|
|
"parser",
|
|
"reporting",
|
|
]
|
|
classifiers = [
|
|
"Development Status :: 5 - Production/Stable",
|
|
"Intended Audience :: Developers",
|
|
"Intended Audience :: Information Technology",
|
|
"License :: OSI Approved :: Apache Software License",
|
|
"Operating System :: OS Independent",
|
|
"Programming Language :: Python :: 3",
|
|
"Programming Language :: Python :: 3 :: Only",
|
|
"Programming Language :: Python :: 3.10",
|
|
"Programming Language :: Python :: 3.11",
|
|
"Programming Language :: Python :: 3.12",
|
|
"Programming Language :: Python :: 3.13",
|
|
"Programming Language :: Python :: 3.14",
|
|
]
|
|
requires-python = ">=3.10"
|
|
# The base install is the parsing core plus a working core CLI: file,
|
|
# IMAP, Maildir, and mbox input; CSV/JSON, Splunk HEC, webhook, and
|
|
# syslog output. Every other output and mailbox integration lives in an extra
|
|
# below, so a library or small-mail-host install does not have to carry
|
|
# the Elasticsearch, OpenSearch, Kafka, AWS, Azure, Gmail, and Microsoft
|
|
# Graph SDKs (see issue #883).
|
|
dependencies = [
|
|
# The [doh] extra supplies the httpx/h2 floors DNS over HTTPS needs;
|
|
# 2.7.0 is the floor verified against the dns.nameserver and
|
|
# dns.query.https(session=...) APIs utils.py builds on.
|
|
"dnspython[doh]>=2.7.0",
|
|
"expiringdict>=1.1.4",
|
|
# The runtime HTTP library (utils.py fetches, webhook and Splunk HEC
|
|
# clients, Graph error handling in cli.py). The floor matches
|
|
# microsoft-kiota-http's own requirement.
|
|
"httpx>=0.25",
|
|
"lxml>=4.4.0",
|
|
# No extras: the base mailsuite supplies IMAP and Maildir connections
|
|
# plus mail-parser. Gmail and Microsoft Graph come from the gmail and
|
|
# msgraph extras below.
|
|
"mailsuite>=2.3.1",
|
|
"maxminddb>=2.0.0",
|
|
"publicsuffixlist>=0.10.0",
|
|
# Imported directly at utils.py:39 (dateutil.parser). It used to
|
|
# arrive transitively via dateparser, which nothing ever imported.
|
|
"python-dateutil>=2.8.0",
|
|
"tqdm>=4.31.1",
|
|
"xmltodict>=0.12.0",
|
|
"PyYAML>=6.0.3"
|
|
]
|
|
|
|
# Splunk HEC, webhook, and syslog outputs deliberately have no extra:
|
|
# they need only httpx (a base dependency) and the standard library, so
|
|
# do not add empty marker extras for them.
|
|
[project.optional-dependencies]
|
|
elastic = [
|
|
"elasticsearch>=8.18,<9",
|
|
]
|
|
opensearch = [
|
|
"opensearch-py>=2.4.2,<=4.0.0",
|
|
# boto3 supplies the SigV4 signer opensearch.py imports for AWS auth.
|
|
"boto3>=1.16.63",
|
|
]
|
|
kafka = [
|
|
"kafka-python>=2.3.2",
|
|
]
|
|
s3 = [
|
|
"boto3>=1.16.63",
|
|
]
|
|
gelf = [
|
|
"pygelf>=0.4.2",
|
|
]
|
|
loganalytics = [
|
|
"azure-identity>=1.8.0",
|
|
"azure-monitor-ingestion>=1.0.0",
|
|
]
|
|
msgraph = [
|
|
"mailsuite[msgraph]>=2.3.1",
|
|
# Imported directly in cli.py for Graph error handling; otherwise
|
|
# only a transitive dep of mailsuite[msgraph] -> msgraph-sdk.
|
|
"microsoft-kiota-abstractions>=1.8.0",
|
|
]
|
|
gmail = [
|
|
"mailsuite[gmail]>=2.3.1",
|
|
]
|
|
# `all` must stay the union of every extra above — and must keep
|
|
# excluding `postgresql`: psycopg's prebuilt binary wheels do not exist
|
|
# for every platform/arch, so folding it in here would make
|
|
# `pip install parsedmarc[all]` fail on platforms it does not cover.
|
|
# Spelled out as an explicit package list rather than a self-referential
|
|
# `parsedmarc[elastic,...]` spec, which would make the project a
|
|
# dependency of itself; keep this list in sync when an extra above
|
|
# gains, drops, or re-floors a package.
|
|
all = [
|
|
"azure-identity>=1.8.0",
|
|
"azure-monitor-ingestion>=1.0.0",
|
|
"boto3>=1.16.63",
|
|
"elasticsearch>=8.18,<9",
|
|
"kafka-python>=2.3.2",
|
|
"mailsuite[gmail,msgraph]>=2.3.1",
|
|
"microsoft-kiota-abstractions>=1.8.0",
|
|
"opensearch-py>=2.4.2,<=4.0.0",
|
|
"pygelf>=0.4.2",
|
|
]
|
|
postgresql = [
|
|
# Optional output backend. psycopg ships prebuilt binary wheels via the
|
|
# [binary] extra, but those wheels don't exist for every platform/arch,
|
|
# so PostgreSQL support is opt-in rather than a mandatory dependency.
|
|
"psycopg[binary]>=3.1.0",
|
|
]
|
|
build = [
|
|
# Used only by maintainer tooling under parsedmarc/resources/maps/ —
|
|
# `collect_domain_info.py --use-search-fallback` falls back to a
|
|
# DuckDuckGo search when the homepage fetch returns a bot-block / parked
|
|
# / empty page. Optional import; the script runs without it as long as
|
|
# the fallback flag isn't passed.
|
|
"ddgs>=9.0.0",
|
|
"hatch>=1.14.0",
|
|
"myst-parser[linkify]",
|
|
"nose",
|
|
# Pinned exactly: pyright's checks evolve between releases, so an
|
|
# unpinned version could break CI without any code change. Bump
|
|
# deliberately (and fix any new findings) rather than implicitly.
|
|
"pyright==1.1.411",
|
|
"pytest",
|
|
"pytest-cov",
|
|
# Test-only: tests/test_init.py imports pytz for a fixed-offset
|
|
# timezone. It used to arrive transitively via dateparser, which the
|
|
# base dependencies no longer declare.
|
|
"pytz",
|
|
# Used only by the out-of-wheel maintainer script
|
|
# parsedmarc/resources/maps/collect_domain_info.py, which deliberately
|
|
# stays on requests because its permissive-TLS fallback is built on
|
|
# urllib3's HTTPAdapter machinery.
|
|
"requests>=2.22.0",
|
|
# Pinned exactly for the same reason as pyright: ruff's default rule
|
|
# set evolves between releases (e.g. 0.16.0 began flagging the
|
|
# str.format() style this codebase used until then), so an unpinned
|
|
# version breaks CI without any code change. Bump deliberately and fix
|
|
# any new findings in the same PR as the bump.
|
|
"ruff==0.16.0",
|
|
"sphinx",
|
|
"sphinx_rtd_theme",
|
|
]
|
|
|
|
[project.scripts]
|
|
parsedmarc = "parsedmarc.cli:_main"
|
|
|
|
[project.urls]
|
|
Homepage = "https://domainaware.github.io/parsedmarc"
|
|
|
|
[tool.hatch.version]
|
|
path = "parsedmarc/constants.py"
|
|
|
|
[tool.hatch.build.targets.sdist]
|
|
include = [
|
|
"/parsedmarc",
|
|
]
|
|
|
|
[tool.hatch.build]
|
|
exclude = [
|
|
"base_reverse_dns.csv",
|
|
"unknown_base_reverse_dns.csv",
|
|
"README.md",
|
|
"AGENTS.md",
|
|
"CLAUDE.md",
|
|
"*.bak",
|
|
# Maintenance tooling: any Python file under parsedmarc/resources/maps/
|
|
# whose name doesn't start with `_` (i.e. everything except __init__.py,
|
|
# which must keep shipping for `importlib.resources.files()` lookups).
|
|
"parsedmarc/resources/maps/[!_]*.py",
|
|
]
|
|
|
|
[tool.ruff.lint]
|
|
# The rule set is selected explicitly rather than floating on ruff's
|
|
# defaults: ruff 0.16.0 expanded the default selection from the
|
|
# long-standing E4/E7/E9/F to many more rule families (BLE, SIM, C4, DTZ,
|
|
# I, PL, S, ...), some of which conflict with deliberate house style —
|
|
# e.g. BLE001 flags the parser's intentional broad catches that keep one
|
|
# malformed report from crashing a batch. E4/E7/E9/F is the pre-0.16
|
|
# default set; adopting any of the new families is a deliberate
|
|
# per-family decision, made here with a comment, not an upgrade side
|
|
# effect.
|
|
#
|
|
# UP006/UP007/UP035/UP045 enforce modern type-hint syntax: with
|
|
# requires-python >=3.10, PEP 585 builtins (list[int]) and PEP 604 unions
|
|
# (X | Y, X | None) are available, so keep the deprecated typing.List /
|
|
# Union / Optional spellings out of the codebase.
|
|
select = [
|
|
"E4", # import rules from pycodestyle
|
|
"E7", # statement rules from pycodestyle
|
|
"E9", # runtime/syntax error rules from pycodestyle
|
|
"F", # pyflakes
|
|
"UP006", # non-pep585-annotation: List -> list, Dict -> dict
|
|
"UP007", # non-pep604-annotation-union: Union[X, Y] -> X | Y
|
|
"UP030", # format-literals: "{0}".format(x) -> "{}".format(x)
|
|
"UP032", # f-string: "{}".format(x) -> f"{x}"
|
|
"UP035", # deprecated-import: typing.List etc. / typing -> collections.abc
|
|
"UP045", # non-pep604-annotation-optional: Optional[X] -> X | None
|
|
]
|
|
|
|
[tool.pyright]
|
|
# The whole codebase passes pyright with zero errors and warnings; CI
|
|
# enforces this (see .github/workflows/python-tests.yml). Run locally with
|
|
# `pyright` from the repo root. Requires the [all] and [postgresql] extras
|
|
# to be installed so the optional integration imports (parsedmarc/cli.py's
|
|
# guarded output modules, the psycopg import in parsedmarc/postgres.py)
|
|
# resolve.
|
|
include = ["parsedmarc", "tests", "docs"]
|
|
typeCheckingMode = "standard"
|
|
|
|
[tool.pytest.ini_options]
|
|
# Default to the per-module test layout under tests/. New tests should go
|
|
# into tests/test_<module>.py to match the file they exercise; do not
|
|
# reintroduce a monolithic tests.py.
|
|
testpaths = ["tests"]
|
|
|
|
[tool.coverage.run]
|
|
# Coverage measures shipped code only. Master's reported ≈66.9% on
|
|
# Codecov was an artefact of the old monolithic tests.py having no
|
|
# [tool.coverage.run] block, which let coverage's default behaviour
|
|
# measure every file imported during the run — including the test file
|
|
# itself at ~99% "covered". That inflated the headline by ~8 percentage
|
|
# points without any actual testing signal. Restricting to the parsedmarc
|
|
# package gives a meaningful number that tracks how much of the shipped
|
|
# library the test suite actually exercises.
|
|
source = ["parsedmarc"]
|
|
# Maintainer-only batch scripts under parsedmarc/resources/maps/ ship
|
|
# out of the wheel (see the [tool.hatch.build] exclude block above) —
|
|
# omit them so the headline number reflects only installed library code.
|
|
omit = [
|
|
"*/parsedmarc/resources/maps/*.py",
|
|
]
|