* Extend n_procs parallel parsing to mbox and mailbox sources (#147)
n_procs previously applied only to report files passed directly as CLI
arguments; messages from mbox files and mailbox connections (IMAP,
Microsoft Graph, Gmail API, Maildir) were always parsed sequentially.
A new parsedmarc.parallel module provides a shared bounded-window
ProcessPoolExecutor helper (parallel_map) used by all three input paths.
Only parsing fans out to a reused worker pool; message fetching, report
deduplication, mailbox archiving/deletion, and output stay sequential in
the main process. The submission window keeps at most ~2*n_procs
messages in flight, so memory stays bounded even for huge mboxes, and
the mailbox path fetches messages lazily on the connection-owning main
thread. keep_alive never crosses the process boundary - the main
process sends periodic IMAP keepalives while workers parse - and with
n_procs > 1, invalid-message disposition happens after the parse phase,
mirroring the existing deferred bulk archive moves.
get_dmarc_reports_from_mbox, get_dmarc_reports_from_mailbox (including
its tail-recursive re-check), and watch_inbox gain an n_procs keyword
argument (default 1); sequential behavior at the default is unchanged.
Replacing the CLI's hand-rolled Pipe/Process batching also fixes two
defects in the direct-file path: a child process that died from a
non-ParserError exception left the parent blocked forever on
conn.recv(), and the hard batch barrier let one slow file idle every
other worker slot. Workers are now a reused pool (no fresh interpreter
per file), with worker logging reconstructed via a spawn-safe pool
initializer instead of fork inheritance.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Address Copilot and code-quality review feedback on #849
- parallel_map now validates n_procs >= 1 itself with a clear error
instead of surfacing ProcessPoolExecutor's max_workers error later.
The check raises eagerly at the call (the generator body moved into an
inner function) rather than on first iteration, with a regression test.
- Aligned parallel_map's should_stop docstring with the implementation:
queued-but-unstarted jobs are cancelled, while in-flight jobs are
waited on and their results yielded, so the stop can block briefly but
never discards completed work.
- The parallel mailbox path keeps fetched message ids in a deque popped
as each in-order result arrives, so the id queue stays bounded by the
submission window instead of growing to message_limit.
- Closed the three sample-file handles the new tests opened without a
context manager.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Address second round of Copilot feedback on #849
- configure_logging no longer stacks duplicate FileHandlers when called
again with the same log_file (compared by FileHandler.baseFilename,
which stores the absolute path): a duplicate wrote every record twice
and leaked a file descriptor per call, e.g. across SIGHUP config
reloads. Latent in the pre-extraction cli._configure_logging too.
Regression tests in the new tests/test_log.py.
- Renamed the CHANGELOG's premature "10.4.0" heading to "Unreleased",
matching the project convention where in-progress entries accumulate
under Unreleased and the release PR renames the section and bumps
parsedmarc/constants.py together (as in the 10.3.0 release).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>