Centralize configuration handling with ParserConfig (#503) (#851)

* Centralize configuration handling with ParserConfig (#503)

Add parsedmarc/config.py with ParserConfig, a frozen dataclass carrying
every parsing/enrichment option plus the three shared caches (IP address
info, seen aggregate report IDs, reverse DNS map). All eight public
parsing/mailbox functions accept a keyword-only config= argument; when
provided, the individual option keyword arguments are ignored in favor
of the config's values, and every existing per-option keyword argument
keeps working unchanged. The three hand-copied parse_kwargs dicts and
the dns_timeout<->timeout rename chain are gone; the CLI builds one
ParserConfig per run (rebuilt on SIGHUP) and passes it everywhere.

Explicitly constructed configs own fresh isolated caches;
dataclasses.replace() shares them; pickling drops cache contents and
rebinds the unpickling process's module defaults, preserving the
per-worker cache behavior of n_procs parallel parsing. The module
globals IP_ADDRESS_CACHE / SEEN_AGGREGATE_REPORT_IDS / REVERSE_DNS_MAP
remain, identity-preserved, as re-exports of the default caches.

Bug fixes that ride along, each with a regression test:

- One-shot mailbox runs now honor [general] dns_timeout/dns_retries;
  the CLI call site never forwarded dns_timeout, so the library's
  stray 6.0 default silently applied.
- Lazily-triggered reverse DNS map loads (get_ip_address_info /
  get_service_from_reverse_dns_base_domain, including in n_procs
  workers) now thread psl_overrides_path/psl_overrides_url through to
  load_reverse_dns_map instead of clobbering operator-configured PSL
  overrides with the bundled defaults.
- get_dmarc_reports_from_mailbox() and watch_inbox() dns_timeout
  defaults unified to DEFAULT_DNS_TIMEOUT (2.0s) from a stray 6.0, and
  normalize_timespan_threshold_hours to the float 24.0 used everywhere
  else.

Closes #503

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address CI and review feedback on #851

- Wrap the IMAPConnection example in usage.md so ruff format is clean
  over the docs code blocks (CI runs ruff format --check on the whole
  repo; the local runs were scoped to parsedmarc/ and tests/ and
  missed it).
- Fix the pre-existing "URL ro a reverse DNS map" docstring typo in
  get_service_from_reverse_dns_base_domain, caught by Copilot on the
  adjacent hunk.
- Import parsedmarc.config once, as an aliased plain import, in
  tests/test_config.py instead of mixing import and import-from of the
  same module (flagged by code quality scanning); the aliased module
  import also keeps pyright able to resolve the submodule attribute
  access.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address second Copilot review round on #851

- Add parse_aggregate_report_file() to the library entry-point list in
  usage.md; the following paragraph describes the config= contract for
  "each of these functions", so the list must name all eight
  config-accepting entry points.
- ParserConfig.__setstate__ now initializes every non-cache field to
  its class default before applying the pickled state, so a config
  serialized by an older parsedmarc version (whose state predates
  fields added later) unpickles with the newer fields at their
  defaults instead of unset entirely (__init__ never runs during
  unpickling, so an absent field would raise AttributeError on first
  access). Covered by a regression test that feeds __setstate__ a
  partial state dict.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Sean Whalen
2026-07-25 19:27:09 -04:00
committed by GitHub
co-authored by Claude Fable 5
parent 4f4733003c
commit 4e80047e68
13 changed files with 1301 additions and 293 deletions
+7
View File
@@ -7,6 +7,13 @@
:members:
```
## parsedmarc.config
```{eval-rst}
.. automodule:: parsedmarc.config
:members:
```
## parsedmarc.elastic
```{eval-rst}
+44
View File
@@ -944,6 +944,50 @@ For sections with underscores in the name, the full section name is used:
| `gelf` | `PARSEDMARC_GELF_` |
| `webhook` | `PARSEDMARC_WEBHOOK_` |
## Using parsedmarc as a library
`parsedmarc` is also importable as a regular Python package, not just a CLI
tool. The main entry points — `parse_report_file()`, `parse_aggregate_report_xml()`,
`parse_aggregate_report_file()`, `parse_failure_report()`, `parse_report_email()`,
`get_dmarc_reports_from_mbox()`, `get_dmarc_reports_from_mailbox()`, and
`watch_inbox()` — are all importable
directly from the `parsedmarc` package. See the [API reference](api.md) for
the full set of modules and members.
Each of these functions accepts either individual option keyword arguments
(`offline`, `nameservers`, `dns_timeout`, etc.) or a single `config=` keyword
argument carrying a `ParserConfig` instance:
```python
from parsedmarc import ParserConfig, parse_report_file, get_dmarc_reports_from_mailbox
from parsedmarc.mail import IMAPConnection
config = ParserConfig(
offline=False,
nameservers=["1.1.1.1", "1.0.0.1"],
dns_timeout=5.0,
)
report = parse_report_file("aggregate_report.xml.gz", config=config)
connection = IMAPConnection(
host="imap.example.com", user="dmarc@example.com", password="..."
)
results = get_dmarc_reports_from_mailbox(connection, config=config)
```
A few things to keep in mind:
- When `config=` is passed, the individual option keyword arguments are
ignored in favor of the values carried on the `ParserConfig` instance.
- Each explicitly constructed `ParserConfig` owns its own isolated caches
(IP address info, seen aggregate report IDs, and the reverse DNS map).
Omitting `config=` falls back to the process-wide caches shared by every
call that doesn't pass one.
- `keep_alive` and `n_procs` are not part of `ParserConfig` — they control
process/worker orchestration rather than parsing or enrichment behavior,
so they are always passed as separate keyword arguments.
## Performance tuning
For large mailbox imports or backfills, parsedmarc can consume a noticeable amount