mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-09-05 13:38:00 +00:00
* Decode failure report MIME parts per their Content-Transfer-Encoding Fixes #882. parse_report_email() read every MIME part's payload without asking the standard library to decode it, leaving any transfer encoding in place: - A quoted-printable text/rfc822-headers sample part kept its RFC 2045 §6.7 soft line breaks, which split long headers without RFC 5322 folding whitespace. The sample's From header became unparseable and, with no Reported-Domain field in the report, the whole failure report was discarded with "TypeError: 'NoneType' object is not subscriptable". - A quoted-printable message/feedback-report part parsed "successfully" with silently corrupted values (e.g. "dmarc=3Dfail"). A new _decode_mime_payload() helper decodes quoted-printable and base64 parts only, applied to the message/feedback-report and sample branches; all other branches still receive the raw payload because they do their own base64/magic-byte handling. Parts with a 7bit/8bit/absent CTE are returned as-is: the message is parsed from a str, so compat32's get_payload(decode=True) would round-trip the already-correct text through raw-unicode-escape and corrupt non-ASCII characters. For nested message/* parts (the stdlib nests every message/* subtype, so decode=True returns None), the encoding is undone by hand, including removing the header/body separator the Generator inserts when the still-encoded text stops looking like headers mid-block (MissingHeaderBodySeparatorDefect) — without that, values were truncated at the first soft line break. Also fixed in the process, per the same-PR rule for bugs found while writing tests: - feedback_report_regex captured the CR of CRLF line endings (RFC 5322 §2.1 mandates CRLF; re.MULTILINE's "$" matches before the LF, not the CR). Previously masked because the base64 branch decoded through a bytes repr and stripped literal "\r" escapes. - A report with no Reported-Domain field and no parseable sample From domain now raises InvalidFailureReport with a clear message instead of the opaque TypeError; reported_domain is a required str in the FailureReport contract (types.py) consumed unconditionally by the Elasticsearch/OpenSearch outputs, so defaulting it to None is not an option. InvalidFailureReport raised inside parse_failure_report() now propagates without the "Unexpected error:" re-wrap. CLI output over the whole sample corpus is byte-identical before and after (PYTHONHASHSEED=0, n_procs=1). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Use RFC 5322 header folding instead of implicit string concatenation The hand-built test message's long Content-Type header was split across two adjacent string literals inside a list, which reads like a missing comma (flagged by code review). Fold the header with a tab continuation line instead — truer to the wire format the builder exists to produce. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Reword test builder docstring to claim only what it does Copilot review: the builder joins lines with "\n" and embeds CRLF inside the feedback-report block, so it does not preserve on-the-wire bytes exactly. What matters for the test is only that the non-ASCII sample text stays unencoded; say that instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>