Files
parsedmarc/tests/test_log.py
T
48445c639e Extend n_procs parallel parsing to mbox and mailbox sources (#147) (#849)
* Extend n_procs parallel parsing to mbox and mailbox sources (#147)

n_procs previously applied only to report files passed directly as CLI
arguments; messages from mbox files and mailbox connections (IMAP,
Microsoft Graph, Gmail API, Maildir) were always parsed sequentially.

A new parsedmarc.parallel module provides a shared bounded-window
ProcessPoolExecutor helper (parallel_map) used by all three input paths.
Only parsing fans out to a reused worker pool; message fetching, report
deduplication, mailbox archiving/deletion, and output stay sequential in
the main process. The submission window keeps at most ~2*n_procs
messages in flight, so memory stays bounded even for huge mboxes, and
the mailbox path fetches messages lazily on the connection-owning main
thread. keep_alive never crosses the process boundary - the main
process sends periodic IMAP keepalives while workers parse - and with
n_procs > 1, invalid-message disposition happens after the parse phase,
mirroring the existing deferred bulk archive moves.

get_dmarc_reports_from_mbox, get_dmarc_reports_from_mailbox (including
its tail-recursive re-check), and watch_inbox gain an n_procs keyword
argument (default 1); sequential behavior at the default is unchanged.

Replacing the CLI's hand-rolled Pipe/Process batching also fixes two
defects in the direct-file path: a child process that died from a
non-ParserError exception left the parent blocked forever on
conn.recv(), and the hard batch barrier let one slow file idle every
other worker slot. Workers are now a reused pool (no fresh interpreter
per file), with worker logging reconstructed via a spawn-safe pool
initializer instead of fork inheritance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address Copilot and code-quality review feedback on #849

- parallel_map now validates n_procs >= 1 itself with a clear error
  instead of surfacing ProcessPoolExecutor's max_workers error later.
  The check raises eagerly at the call (the generator body moved into an
  inner function) rather than on first iteration, with a regression test.
- Aligned parallel_map's should_stop docstring with the implementation:
  queued-but-unstarted jobs are cancelled, while in-flight jobs are
  waited on and their results yielded, so the stop can block briefly but
  never discards completed work.
- The parallel mailbox path keeps fetched message ids in a deque popped
  as each in-order result arrives, so the id queue stays bounded by the
  submission window instead of growing to message_limit.
- Closed the three sample-file handles the new tests opened without a
  context manager.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address second round of Copilot feedback on #849

- configure_logging no longer stacks duplicate FileHandlers when called
  again with the same log_file (compared by FileHandler.baseFilename,
  which stores the absolute path): a duplicate wrote every record twice
  and leaked a file descriptor per call, e.g. across SIGHUP config
  reloads. Latent in the pre-extraction cli._configure_logging too.
  Regression tests in the new tests/test_log.py.
- Renamed the CHANGELOG's premature "10.4.0" heading to "Unreleased",
  matching the project convention where in-progress entries accumulate
  under Unreleased and the release PR renames the section and bumps
  parsedmarc/constants.py together (as in the 10.3.0 release).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 17:14:40 -04:00

75 lines
2.4 KiB
Python

"""Tests for parsedmarc.log"""
import logging
import os
import tempfile
import unittest
from parsedmarc.log import configure_logging, logger
class TestConfigureLoggingFileHandlerDedup(unittest.TestCase):
"""Repeated configure_logging calls with the same log_file (e.g. a
SIGHUP config reload re-running the CLI's logging setup) must not
stack duplicate FileHandlers - a duplicate would write every record
twice and leak a file descriptor per call."""
def setUp(self):
self._saved_handlers = list(logger.handlers)
self._saved_level = logger.level
def tearDown(self):
for handler in list(logger.handlers):
if handler not in self._saved_handlers:
logger.removeHandler(handler)
if isinstance(handler, logging.FileHandler):
handler.close()
logger.handlers[:] = self._saved_handlers
logger.setLevel(self._saved_level)
def _temp_log_path(self):
with tempfile.NamedTemporaryFile(suffix=".log", delete=False) as tf:
path = tf.name
self.addCleanup(lambda: os.path.exists(path) and os.remove(path))
return path
def test_same_log_file_twice_attaches_one_handler_and_logs_once(self):
log_path = self._temp_log_path()
configure_logging(logging.INFO, log_path)
configure_logging(logging.INFO, log_path)
file_handlers = [
h
for h in logger.handlers
if isinstance(h, logging.FileHandler)
and h.baseFilename == os.path.abspath(log_path)
]
self.assertEqual(len(file_handlers), 1)
logger.info("dedup-check line")
for handler in file_handlers:
handler.flush()
with open(log_path) as f:
contents = f.read()
self.assertEqual(contents.count("dedup-check line"), 1)
def test_different_log_files_attach_one_handler_each(self):
first_path = self._temp_log_path()
second_path = self._temp_log_path()
configure_logging(logging.INFO, first_path)
configure_logging(logging.INFO, second_path)
file_paths = [
h.baseFilename
for h in logger.handlers
if isinstance(h, logging.FileHandler)
]
self.assertIn(os.path.abspath(first_path), file_paths)
self.assertIn(os.path.abspath(second_path), file_paths)
if __name__ == "__main__":
unittest.main(verbosity=2)