mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-09-05 21:47:58 +00:00
758d1ffe4943d27c7cca85ffb8972eb0178abcdb
* Decode failure report MIME parts per their Content-Transfer-Encoding Fixes #882. parse_report_email() read every MIME part's payload without asking the standard library to decode it, leaving any transfer encoding in place: - A quoted-printable text/rfc822-headers sample part kept its RFC 2045 §6.7 soft line breaks, which split long headers without RFC 5322 folding whitespace. The sample's From header became unparseable and, with no Reported-Domain field in the report, the whole failure report was discarded with "TypeError: 'NoneType' object is not subscriptable". - A quoted-printable message/feedback-report part parsed "successfully" with silently corrupted values (e.g. "dmarc=3Dfail"). A new _decode_mime_payload() helper decodes quoted-printable and base64 parts only, applied to the message/feedback-report and sample branches; all other branches still receive the raw payload because they do their own base64/magic-byte handling. Parts with a 7bit/8bit/absent CTE are returned as-is: the message is parsed from a str, so compat32's get_payload(decode=True) would round-trip the already-correct text through raw-unicode-escape and corrupt non-ASCII characters. For nested message/* parts (the stdlib nests every message/* subtype, so decode=True returns None), the encoding is undone by hand, including removing the header/body separator the Generator inserts when the still-encoded text stops looking like headers mid-block (MissingHeaderBodySeparatorDefect) — without that, values were truncated at the first soft line break. Also fixed in the process, per the same-PR rule for bugs found while writing tests: - feedback_report_regex captured the CR of CRLF line endings (RFC 5322 §2.1 mandates CRLF; re.MULTILINE's "$" matches before the LF, not the CR). Previously masked because the base64 branch decoded through a bytes repr and stripped literal "\r" escapes. - A report with no Reported-Domain field and no parseable sample From domain now raises InvalidFailureReport with a clear message instead of the opaque TypeError; reported_domain is a required str in the FailureReport contract (types.py) consumed unconditionally by the Elasticsearch/OpenSearch outputs, so defaulting it to None is not an option. InvalidFailureReport raised inside parse_failure_report() now propagates without the "Unexpected error:" re-wrap. CLI output over the whole sample corpus is byte-identical before and after (PYTHONHASHSEED=0, n_procs=1). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Use RFC 5322 header folding instead of implicit string concatenation The hand-built test message's long Content-Type header was split across two adjacent string literals inside a list, which reads like a missing comma (flagged by code review). Fold the header with a tab continuation line instead — truer to the wire format the builder exists to produce. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Reword test builder docstring to claim only what it does Copilot review: the builder joins lines with "\n" and embeds CRLF inside the feedback-report block, so it does not preserve on-the-wire bytes exactly. What matters for the test is only that the non-ASCII sample text stays unencoded; say that instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
fix: OSD Global-tenant import + dropped report files with glob metacharacters; validate dev stack on OpenSearch 3.x with PostgreSQL (#781)
parsedmarc
parsedmarc is a Python module and CLI utility for parsing DMARC
reports. When used with Elasticsearch and Kibana (or Splunk), it works
as a self-hosted open-source alternative to commercial DMARC report
processing services such as Agari Brand Protection, Dmarcian, OnDMARC,
ProofPoint Email Fraud Defense, and Valimail.
Note
Domain-based Message Authentication, Reporting, and Conformance (DMARC) is an email authentication protocol.
Sponsors
This project is maintained by one developer. Please consider sponsoring my work if you or your organization benefit from it.
Features
- Parses aggregate/rua DMARC reports: the legacy draft and 1.0 schemas (RFC 7489) and the new RFC 9990 schema for the final DMARC standard (RFC 9989)
- Parses failure/ruf DMARC reports (RFC 6591 and RFC 9991; formerly called forensic reports)
- Parses reports from SMTP TLS Reporting (TLS-RPT, RFC 8460)
- Can parse reports from an inbox over IMAP, Microsoft Graph, or Gmail API
- Transparently handles gzip or zip compressed reports
- Consistent data structures
- Simple JSON and/or CSV output
- Optionally email the results
- Optionally send the results to Elasticsearch, OpenSearch, Splunk, or PostgreSQL, for use with premade dashboards
- Optionally send the results to Apache Kafka, Amazon S3, Azure Log Analytics (Microsoft Sentinel), a Graylog (GELF) endpoint, a syslog server, or an HTTP webhook
Python Compatibility
This project supports the following Python versions, which are either actively maintained or are the default versions for RHEL or Debian.
| Version | Supported | Reason |
|---|---|---|
| < 3.6 | ❌ | End of Life (EOL) |
| 3.6 | ❌ | Used in RHEL 8, but not supported by project dependencies |
| 3.7 | ❌ | End of Life (EOL) |
| 3.8 | ❌ | End of Life (EOL) |
| 3.9 | ❌ | Used in Debian 11 and RHEL 9, but not supported by project dependencies |
| 3.10 | ✅ | Actively maintained |
| 3.11 | ✅ | Actively maintained; supported until June 2028 (Debian 12) |
| 3.12 | ✅ | Actively maintained; supported until May 2035 (RHEL 10) |
| 3.13 | ✅ | Actively maintained; supported until June 2030 (Debian 13) |
| 3.14 | ✅ | Supported (requires imapclient>=3.1.0) |
Languages
Python
98.5%
Shell
1.4%
