Files
parsedmarc/tests
Sean WhalenandClaude Fable 5 758d1ffe49 Decode failure report MIME parts per their Content-Transfer-Encoding (#885)
* Decode failure report MIME parts per their Content-Transfer-Encoding

Fixes #882.

parse_report_email() read every MIME part's payload without asking the
standard library to decode it, leaving any transfer encoding in place:

- A quoted-printable text/rfc822-headers sample part kept its RFC 2045
  §6.7 soft line breaks, which split long headers without RFC 5322
  folding whitespace. The sample's From header became unparseable and,
  with no Reported-Domain field in the report, the whole failure report
  was discarded with "TypeError: 'NoneType' object is not subscriptable".
- A quoted-printable message/feedback-report part parsed "successfully"
  with silently corrupted values (e.g. "dmarc=3Dfail").

A new _decode_mime_payload() helper decodes quoted-printable and base64
parts only, applied to the message/feedback-report and sample branches;
all other branches still receive the raw payload because they do their
own base64/magic-byte handling. Parts with a 7bit/8bit/absent CTE are
returned as-is: the message is parsed from a str, so compat32's
get_payload(decode=True) would round-trip the already-correct text
through raw-unicode-escape and corrupt non-ASCII characters. For nested
message/* parts (the stdlib nests every message/* subtype, so
decode=True returns None), the encoding is undone by hand, including
removing the header/body separator the Generator inserts when the
still-encoded text stops looking like headers mid-block
(MissingHeaderBodySeparatorDefect) — without that, values were truncated
at the first soft line break.

Also fixed in the process, per the same-PR rule for bugs found while
writing tests:

- feedback_report_regex captured the CR of CRLF line endings (RFC 5322
  §2.1 mandates CRLF; re.MULTILINE's "$" matches before the LF, not the
  CR). Previously masked because the base64 branch decoded through a
  bytes repr and stripped literal "\r" escapes.
- A report with no Reported-Domain field and no parseable sample From
  domain now raises InvalidFailureReport with a clear message instead of
  the opaque TypeError; reported_domain is a required str in the
  FailureReport contract (types.py) consumed unconditionally by the
  Elasticsearch/OpenSearch outputs, so defaulting it to None is not an
  option. InvalidFailureReport raised inside parse_failure_report() now
  propagates without the "Unexpected error:" re-wrap.

CLI output over the whole sample corpus is byte-identical before and
after (PYTHONHASHSEED=0, n_procs=1).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Use RFC 5322 header folding instead of implicit string concatenation

The hand-built test message's long Content-Type header was split across
two adjacent string literals inside a list, which reads like a missing
comma (flagged by code review). Fold the header with a tab continuation
line instead — truer to the wire format the builder exists to produce.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Reword test builder docstring to claim only what it does

Copilot review: the builder joins lines with "\n" and embeds CRLF inside
the feedback-report block, so it does not preserve on-the-wire bytes
exactly. What matters for the test is only that the non-ASCII sample
text stays unencoded; say that instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 12:14:07 -04:00
..