* Make the output and mailbox integrations optional extras (#883)
Breaking change for the next major release: pip install parsedmarc now
installs the parsing core plus a working core CLI (file, IMAP, Maildir,
and mbox input; CSV/JSON, Splunk HEC, webhook, and syslog output).
Everything else moves behind an extra: elastic, opensearch, kafka, s3,
gelf, loganalytics, msgraph, and gmail, joining the existing postgresql
extra, with an umbrella [all] that deliberately excludes postgresql
(psycopg's binary wheels do not exist on every platform, so
parsedmarc[all] must never fail to install there).
cli.py imports the six SDK-dependent output modules behind the #884
TYPE_CHECKING/try-except guard; a configured section whose extra is
missing fails fast with a ConfigurationError naming the section and the
exact pip install command — including the msgraph and gmail_api mailbox
sections (detected via parsedmarc.mail's placeholder classes) and
postgresql (checked before the constructor so the startup retry loop
does not retry a missing dependency for a minute). The Azure/kiota Graph
error types fall back to never-raised sentinel classes.
The Docker image installs [all,postgresql], so container users see no
change. CI lint installs [build,all,postgresql]; the unit-test job
installs [build,all], deliberately without postgresql so
test_postgres.py's absent-psycopg arm stays exercised. The
never-imported dateparser dependency is dropped in favor of declaring
python-dateutil, which utils.py actually imports; pytz moves to the
build extra for the one test that uses it.
Verified live: a no-extras wheel install imports, parses samples, and
reports the install hint for each gated section; a [all] install
restores every integration; the Docker image builds with every SDK
importable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Patch psycopg presence in the PostgreSQL CLI wiring tests
CI's unit-test job deliberately installs [build,all] without the
postgresql extra, so parsedmarc.cli.postgres.psycopg is None there and
the new missing-extra presence check correctly made _main exit 1 before
the wiring under test ran. The tests simulate the SDK being available
(PostgreSQLClient is mocked at the SDK boundary), so the module-level
psycopg handle is now patched present in setUp. Verified against a
simulated psycopg-absent environment as well as the local full install.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Address Copilot review: narrow guards to ModuleNotFoundError, fix docs
- The optional-integration and Graph error-type import guards now catch
ModuleNotFoundError instead of ImportError, so only a genuinely absent
package reads as a missing extra; a broken-but-present SDK fails
loudly with its real error instead of masquerading as one. The test
blocker raises ModuleNotFoundError accordingly — the exact exception a
missing package produces.
- _missing_extra_hint docstring no longer calls every gated integration
an output module (it also serves the msgraph/gmail_api mailbox
sections).
- Fix the pre-existing passsword typo in usage.md's kafka section; the
INI key the code reads is password (cli.py _parse_config).
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Quote extras specs in copy-paste install commands
From Copilot's second review round: zsh treats an unquoted .[build,all]
as a glob and fails with 'no matches found', so the commands shown in
AGENTS.md, CONTRIBUTING.md, dashboards/README.md, and the bootstrap
script's comment are now quoted. The CI workflows keep the unquoted
form: they run under bash, which passes unmatched globs through
literally. The suggestion to change the 'Choosing what to install'
heading level was rejected — it is a subsection of 'Installing
parsedmarc', matching the file's existing hierarchy.
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Fix upgrade command in the changelog
* Documentation review: accuracy, spelling, grammar, and clarity pass
A full prose review of docs/source, README, CONTRIBUTING, and the
dashboards README, with every accuracy claim verified against the code
before changing it. Highlights:
- usage.md: documented six missing [general] options (the CSV/JSON
filename options, prettify_json, normalize_timespan_threshold_hours),
the required kafka smtp_tls_topic, [imap] timeout/max_retries, and
the postgresql env-var prefix; corrected the maildir_path default
(None, not INBOX — cli.py Namespace defaults), the mailbox
check_timeout option name, the systemd restart interval (RestartSec
is 5m), and merged the duplicate silent entry; quoted every
copy-paste extras spec for zsh safety.
- elasticsearch.md: fixed an invalid openssl command (rsa:4096 -nodes),
the dashboards filename (opensearch_dashboards.ndjson, matching the
file the link serves), and assorted grammar.
- davmail.md: the service-enable command now enables davmail.service
(was parsedmarc.service — a copy-paste error that left DavMail
unenabled), plus a view typo and DavMail capitalization.
- output.md: the example schema reference is RFC 7489 Appendix C
(7480 is RDAP). kibana.md: SPF relies on the SMTP envelope, not
session headers (RFC 7208). dmarc.md: DKM -> DKIM.
- README: the intro now also names the OpenSearch/Grafana stack,
matching the feature list. CONTRIBUTING: pre-PR checks now include
ruff format --check and pyright, matching CI's lint job.
- dashboards/README: the service table and seed description now include
the PostgreSQL backend the compose stack runs.
Sample data blocks, the CLI-help mirror block, and released CHANGELOG
entries were deliberately left untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Docstring review: accuracy, spelling, grammar, and clarity pass
Every docstring in parsedmarc/, parsedmarc/mail/, the maps maintainer
scripts, and the test suite reviewed with each claim verified against
the code it documents. Text-only — no behavior changes. Highlights:
- Copy-paste errors corrected: parsed_smtp_tls_reports_to_csv and
splunk/loganalytics save functions described aggregate or failure
reports they do not handle; LogAnalyticsException claimed to be an
Elasticsearch error.
- Docstring/behavior mismatches: parse_report_email's report_type
enumeration omitted smtp_tls; parse_failure_report typed msg_date as
str (it is datetime); strip_attachment_payloads claimed payloads are
replaced with None (the key is deleted); kafkaclient's failure and
SMTP TLS savers claimed per-record slicing while sending the whole
list in one message (docstrings now describe reality — whether
slicing was intended is flagged for follow-up); the postgres savers
claimed to take parse_report_file's return value but receive the
inner report dict; elastic/opensearch save functions' Raises listed
only AlreadySaved.
- None-as-semantic-state documented where missing (get_base_domain,
get_ip_address_country), enumeration completeness fixed
(get_ip_address_info's 9 result keys, maps script outputs, TSV
columns), and the stale 44-industry-types count corrected to the
46 the authoritative README list defines.
- Test docstrings aligned with what the tests actually assert,
including two that overstated coverage of the elastic/opensearch
address-list tests.
- Two argparse help strings fixed: file_path now names SMTP TLS report
files alongside aggregate and failure, mirrored into usage.md's
CLI-help block; --offline's doubled spaces removed (rendered help
unchanged).
- elasticsearch.md's security claim corrected against Elastic's docs:
security is enabled and auto-configured on first startup since 8.0
(not "8.7 secure mode"), so the settings are verified, not
hand-written.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
codecov/test-results-action@v1 emits two warnings on every test run:
its bundled actions/github-script pin targets deprecated Node.js 20, and
Codecov deprecated the action itself in favor of running
codecov/codecov-action@v5 with report_type: test_results. Upload the
JUnit results through the same codecov-action already used for coverage.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Port mailsuite's tag-triggered release pipeline:
- Add release.yml: pushing a version tag runs the full CI suite
(python-tests.yml via workflow_call), then builds the package (the tag
must match the version in parsedmarc/constants.py, checked with
`hatch version`), publishes to PyPI via Trusted Publishing, creates
the GitHub Release with notes from the tag's CHANGELOG.md section and
the built distributions attached, pushes the multi-arch Docker image,
and deploys the Sphinx docs
- Add docs.yml: reusable docs build/deploy to GitHub Pages, also
runnable on demand (workflow_dispatch) for documentation-only changes
between releases
- docker.yml: add a workflow_call trigger with a push_image input, since
a GitHub Release created with the workflow's own GITHUB_TOKEN emits no
`release: published` event; release.yml calls it directly instead
- Remove the legacy build.sh / publish-docs.sh manual process
- AGENTS.md: CRITICAL rule that releases require explicit maintainer
permission, plus docs for the new release flow and its one-time
repo/PyPI configuration prerequisites
- Bump the mailsuite floor to >=2.3.0 (raises the transitive mail-parser
floor to >=4.6.2 and cryptography to >=50.0.0)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Make the whole codebase pass pyright cleanly and enforce it in CI
Fix all 102 pyright (1.1.410, standard mode) errors across the library,
tests, and maps scripts, then pin and enforce the zero-errors bar:
- postgres.py: make the optional psycopg import TYPE_CHECKING-aware so
the module is properly typed while keeping the runtime install-hint
fallback; import psycopg.types.json explicitly as psycopg_json (the
old psycopg_types.json attribute access only worked because psycopg
imports the submodule eagerly); have _connect()/_ensure_connected()
return the live connection so save methods use a non-Optional local;
type the DDL list as list[LiteralString] to match psycopg's execute()
overloads.
- kafkaclient.py: resolve the kafka-python 2.x/3.x bootstrap-error
fallback statically via TYPE_CHECKING (kafka-python 3.0 removed
NoBrokersAvailable), which also fixes _BootstrapError's import
resolution in tests.
- syslog.py: go through getattr/setattr for SysLogHandler.socket
(absent from typeshed); type the save_* methods with the report
TypedDicts (single or list, matching cli.py call sites — gelf.py gets
the same signatures); raise ValueError when retry_attempts < 1
instead of falling through and registering a None handler (bug fix,
with a regression test and a CHANGELOG entry).
- elastic.py / opensearch.py: human_result params are Optional[str].
- maps scripts: sort_csv declared a return type but never returned
(now -> None); seen_sort_field_values was possibly unbound;
convert_to_utf8's src_encoding is Optional[str].
- tests: cast sample-report dict helpers to their TypedDicts; mark
deliberate wrong-type calls with targeted pyright ignores; add
narrowing asserts for Optional results; access the mocked
KafkaProducer through a cast helper; match the mailsuite
fetch_message base signature (**kwargs); patch the renamed
parsedmarc.postgres.psycopg_json in test_postgres's setUpModule.
Enforcement: [tool.pyright] in pyproject.toml (include parsedmarc,
tests, docs; standard mode), pyright==1.1.410 pinned in the [build]
extra (pinned exactly so a new pyright release can't break CI without a
code change), and a "Check types" step in the lint CI job — which now
also runs ruff format --check and installs the [postgresql] extra so
the optional psycopg import resolves. Documented in AGENTS.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Set session headers via update() instead of replacing the dict
requests 2.34 ships inline type annotations, and Session.headers is a
CaseInsensitiveDict[str] — assigning a plain dict fails pyright there
(the CI runner resolved 2.34.2; the local venv's untyped 2.32.4 hid
it). headers.update() is correctly typed against both versions, and is
the documented requests idiom: it overrides User-Agent and the
client-specific headers while keeping the session's defaults
(Accept-Encoding, Connection) instead of wiping them.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Split tests.py into per-module tests/test_<module>.py
The 5174-line tests.py monolith is split into per-module files under
tests/, mirroring the checkdmarc layout:
tests/test_init.py parsedmarc/__init__.py parsing surface
tests/test_cli.py parsedmarc/cli.py + config / env-vars / SIGHUP
tests/test_utils.py parsedmarc/utils.py (DNS, IP info, PSL, etc.)
tests/test_webhook.py parsedmarc/webhook.py
tests/test_kafkaclient.py parsedmarc/kafkaclient.py
tests/test_splunk.py parsedmarc/splunk.py
tests/test_syslog.py parsedmarc/syslog.py
tests/test_loganalytics.py parsedmarc/loganalytics.py
tests/test_gelf.py parsedmarc/gelf.py
tests/test_s3.py parsedmarc/s3.py
tests/test_maps.py parsedmarc/resources/maps/ maintainer scripts
The split is purely a redistribution — no test bodies changed, no tests
added or removed. All 276 existing tests pass under the new layout.
The current tests.py contains two kitchen-sink classes (`Test` at line 54
and `TestEnvVarConfig` at line 2360) holding tests that span many
modules. Their methods are routed to the correct per-module file by name
prefix; the wholly-thematic classes (TestExtractReport, TestUtilsXxx,
TestSighupReload, etc.) move whole. Each target file gets its own
`class Test(unittest.TestCase)` for the redistributed kitchen-sink
methods, plus the thematic classes verbatim.
Wiring updates:
- `.github/workflows/python-tests.yml`: `pytest ... tests.py` →
`python -m pytest ... tests/` (also switches to `python -m pytest` per
the checkdmarc convention so cwd lands on the project root).
- `pyproject.toml`: adds `[tool.pytest.ini_options] testpaths = ["tests"]`
and `[tool.coverage.run] source = ["parsedmarc"]` with an `omit` for
`parsedmarc/resources/maps/*.py`. The maps scripts are maintainer-only
batch tooling that ships out of the wheel; excluding them from
coverage makes the headline number reflect only installed library
code. Runtime coverage on the new layout is 59% (was 45% with maps
counted), and PR-B will push it to 90%+.
- `AGENTS.md`: documents the new layout and how to run individual files
/ tests; tells future contributors not to reintroduce a monolithic
tests.py.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Restore 66.9% coverage baseline (count tests/ + parsedmarc)
Master's headline 66.9% number on Codecov includes the tests.py file
itself (99.35% covered) being measured alongside parsedmarc/*. The
original tests.py had no `[tool.coverage.run]` block, so coverage's
default — "measure every file imported during the run" — counted the
test code as if it were product code.
The split commit added `source = ["parsedmarc"]` which suppressed
measurement of the test files (correct in principle, since test files
aren't shipped code), and that alone made the headline number drop by
~8 percentage points without any actual loss of testing. This commit
swaps `source` for an explicit `include = ["parsedmarc/*", "tests/*"]`
so both halves are measured the way they were on master. Verified:
276 tests, 66.96% line coverage (effectively unchanged from master's
66.90%).
If you want the shipped-code-only number (was the headline that this
commit overrides), run `pytest --cov=parsedmarc tests/`. That number
is currently 59% and is the focus of the upcoming coverage-expansion PR.
Also adds junit.xml to .gitignore so the CI artefact doesn't get
accidentally committed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Restrict coverage to shipped code (`source = ["parsedmarc"]`)
Reverts the prior commit's `include = ["tests/*"]`. Counting the test
files toward coverage was wrong — it conflates "shipped code exercised
by tests" with "test code that pytest auto-runs", inflates the headline
number, and rewards writing more tests rather than tests that verify
more code. Master's apparent 66.9% was an artefact of the old
monolithic tests.py having no [tool.coverage.run] block at all; coverage's
default behaviour measured every imported file, including the test file
itself at ~99% "covered", which added ~8 percentage points to the
displayed number without any real testing signal.
Restricting to `source = ["parsedmarc"]` plus the existing maps omit
gives a meaningful baseline: 59% of shipped code is exercised by the
test suite today. That's the number the next PR is targeting to lift
to 90%+ before the 10.0.0 release; the Codecov "drop" here is a
measurement correction, not a regression.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Add DMARCbis report support; rename forensic→failure project-wide
Rebased on top of master @ 2cda5bf (9.9.0), which added the ASN
source attribution work (#712, #713, #714, #715). Individual Copilot
iteration commits squashed into this single commit — the per-commit
history on the feature branch was iterative (add tests, fix lint,
move field, revert, etc.) and not worth preserving; GitHub squash-
merges PRs anyway.
New fields from the DMARCbis XSD, plumbed through types, parsing, CSV
output, and the Elasticsearch / OpenSearch mappings:
- ``np`` — non-existent subdomain policy (``none`` / ``quarantine`` /
``reject``)
- ``testing`` — testing mode flag (``n`` / ``y``), replaces RFC 7489
``pct``
- ``discovery_method`` — policy discovery method (``psl`` /
``treewalk``)
- ``generator`` — report generator software identifier (metadata)
- ``human_result`` — optional descriptive text on DKIM / SPF results
RFC 7489 reports parse with ``None`` for DMARCbis-only fields.
Forensic reports have been renamed to failure reports throughout the
project to reflect the proper naming since RFC 7489.
- Core: ``types.py``, ``__init__.py`` — ``ForensicReport`` →
``FailureReport``, ``parse_forensic_report`` →
``parse_failure_report``, report type ``"failure"``.
- Output modules: ``elastic.py``, ``opensearch.py``, ``splunk.py``,
``kafkaclient.py``, ``syslog.py``, ``gelf.py``, ``webhook.py``,
``loganalytics.py``, ``s3.py``.
- CLI: ``cli.py`` — args, config keys, index names
(``dmarc_failure``).
- Docs + dashboards: all markdown, Grafana JSON, Kibana NDJSON,
Splunk XML.
Backward compatibility preserved: old function / type names remain as
aliases (``parse_forensic_report = parse_failure_report``,
``ForensicReport = FailureReport``, etc.), CLI accepts both the old
(``save_forensic``, ``forensic_topic``) and new (``save_failure``,
``failure_topic``) config keys, and updated dashboards query both
old and new index / sourcetype names so data from before and after
the rename appears together.
Merge conflicts resolved in ``parsedmarc/constants.py`` (took bis's
10.0.0 bump), ``parsedmarc/__init__.py`` (combined bis's "failure"
wording with master's IPinfo MMDB mention), ``parsedmarc/elastic.py``
and ``parsedmarc/opensearch.py`` (kept master's ``source_asn`` /
``source_asn_name`` / ``source_asn_domain`` on the failure doc path
while renaming ``forensic_report`` → ``failure_report``), and
``CHANGELOG.md`` (10.0.0 entry now sits above the 9.9.0 entry).
All 324 tests pass; ``ruff check`` / ``ruff format --check`` clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Apply post-RFC review fixes: RFC 9990 detection, langAttrString, CFWS-aware RUF parsing
Aligns the implementation with the final RFCs (9989/9990/9991) instead of
inferring DMARCbis support from the version element or the namespace alone.
Aggregate parsing (RFC 9990):
- _text() helper unwraps langAttrString values (extra_contact_info, error,
comment, human_result, generator) — when reporters include the lang
attribute, xmltodict yields {"#text": ..., "@lang": ...} dicts instead
of strings; the parser now stores the text payload in both shapes.
- New xml_namespace field on AggregateReport records the declared XML
namespace (urn:ietf:params:xml:ns:dmarc-2.0 for RFC 9990 reports).
- RFC 9990 detection accepts namespaceless reports that follow the
RFC 9990 shape (presence of np / testing / discovery_method / generator),
so reporters that don't declare the namespace still receive RFC 9990-
aware validation.
- Warnings: missing DKIM <selector> (REQUIRED in RFC 9990); legacy
forwarded / sampled_out policy-override types (removed by RFC 9990);
unknown policy-override types per the RFC 9990 enumeration.
- xml_namespace added to Elasticsearch and OpenSearch document mappings.
Failure parsing (RFC 9991):
- Identity-Alignment and Auth-Failure are split on commas with CFWS
whitespace stripped per the RFC 9991 ABNF; previously "dkim, spf"
yielded ["dkim", " spf"] with a leading space on the second token.
- Warnings logged when either REQUIRED field is missing.
Terminology: every reference to "DMARCbis" in code, tests, sample
filenames, AGENTS.md, and CHANGELOG.md is replaced with the appropriate
RFC number (9989 for the policy spec, 9990 for aggregate reports, 9991
for failure reports). Sample contents are unchanged.
Docs: corrects the prior claim that fo was dropped from RFC 9990 (only
pct was), reframes testing as a new field (not a pct replacement, since
RFC 9989 Appendix A.6 removed pct with no per-message substitute), and
documents the policy_override_reason enum changes (added policy_test_mode;
removed forwarded / sampled_out).
Tests: 8 new tests covering xml_namespace capture, RFC 9990 detection
from field shape, missing-DKIM-selector warning, legacy-override-type
warning, langAttrString unwrapping across all four affected elements,
and CFWS-aware Identity-Alignment / Auth-Failure parsing plus their
missing-field warnings. 276 tests total, all passing; ruff clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Sean Whalen <44679+seanthegeek@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fixing ES/OS forensic report lookup and storage, extracting ES to separate CI service
* bumping CI ES version to current latest
* reshuffling CI job attributes
* removing EOL Python 3.8 from the CI pipeline