mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-07-27 19:04:54 +00:00
* Extend n_procs parallel parsing to mbox and mailbox sources (#147) n_procs previously applied only to report files passed directly as CLI arguments; messages from mbox files and mailbox connections (IMAP, Microsoft Graph, Gmail API, Maildir) were always parsed sequentially. A new parsedmarc.parallel module provides a shared bounded-window ProcessPoolExecutor helper (parallel_map) used by all three input paths. Only parsing fans out to a reused worker pool; message fetching, report deduplication, mailbox archiving/deletion, and output stay sequential in the main process. The submission window keeps at most ~2*n_procs messages in flight, so memory stays bounded even for huge mboxes, and the mailbox path fetches messages lazily on the connection-owning main thread. keep_alive never crosses the process boundary - the main process sends periodic IMAP keepalives while workers parse - and with n_procs > 1, invalid-message disposition happens after the parse phase, mirroring the existing deferred bulk archive moves. get_dmarc_reports_from_mbox, get_dmarc_reports_from_mailbox (including its tail-recursive re-check), and watch_inbox gain an n_procs keyword argument (default 1); sequential behavior at the default is unchanged. Replacing the CLI's hand-rolled Pipe/Process batching also fixes two defects in the direct-file path: a child process that died from a non-ParserError exception left the parent blocked forever on conn.recv(), and the hard batch barrier let one slow file idle every other worker slot. Workers are now a reused pool (no fresh interpreter per file), with worker logging reconstructed via a spawn-safe pool initializer instead of fork inheritance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot and code-quality review feedback on #849 - parallel_map now validates n_procs >= 1 itself with a clear error instead of surfacing ProcessPoolExecutor's max_workers error later. The check raises eagerly at the call (the generator body moved into an inner function) rather than on first iteration, with a regression test. - Aligned parallel_map's should_stop docstring with the implementation: queued-but-unstarted jobs are cancelled, while in-flight jobs are waited on and their results yielded, so the stop can block briefly but never discards completed work. - The parallel mailbox path keeps fetched message ids in a deque popped as each in-order result arrives, so the id queue stays bounded by the submission window instead of growing to message_limit. - Closed the three sample-file handles the new tests opened without a context manager. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address second round of Copilot feedback on #849 - configure_logging no longer stacks duplicate FileHandlers when called again with the same log_file (compared by FileHandler.baseFilename, which stores the absolute path): a duplicate wrote every record twice and leaked a file descriptor per call, e.g. across SIGHUP config reloads. Latent in the pre-extraction cli._configure_logging too. Regression tests in the new tests/test_log.py. - Renamed the CHANGELOG's premature "10.4.0" heading to "Unreleased", matching the project convention where in-progress entries accumulate under Unreleased and the release PR renames the section and bumps parsedmarc/constants.py together (as in the 10.3.0 release). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
75 lines
2.4 KiB
Python
75 lines
2.4 KiB
Python
"""Tests for parsedmarc.log"""
|
|
|
|
import logging
|
|
import os
|
|
import tempfile
|
|
import unittest
|
|
|
|
from parsedmarc.log import configure_logging, logger
|
|
|
|
|
|
class TestConfigureLoggingFileHandlerDedup(unittest.TestCase):
|
|
"""Repeated configure_logging calls with the same log_file (e.g. a
|
|
SIGHUP config reload re-running the CLI's logging setup) must not
|
|
stack duplicate FileHandlers - a duplicate would write every record
|
|
twice and leak a file descriptor per call."""
|
|
|
|
def setUp(self):
|
|
self._saved_handlers = list(logger.handlers)
|
|
self._saved_level = logger.level
|
|
|
|
def tearDown(self):
|
|
for handler in list(logger.handlers):
|
|
if handler not in self._saved_handlers:
|
|
logger.removeHandler(handler)
|
|
if isinstance(handler, logging.FileHandler):
|
|
handler.close()
|
|
logger.handlers[:] = self._saved_handlers
|
|
logger.setLevel(self._saved_level)
|
|
|
|
def _temp_log_path(self):
|
|
with tempfile.NamedTemporaryFile(suffix=".log", delete=False) as tf:
|
|
path = tf.name
|
|
self.addCleanup(lambda: os.path.exists(path) and os.remove(path))
|
|
return path
|
|
|
|
def test_same_log_file_twice_attaches_one_handler_and_logs_once(self):
|
|
log_path = self._temp_log_path()
|
|
|
|
configure_logging(logging.INFO, log_path)
|
|
configure_logging(logging.INFO, log_path)
|
|
|
|
file_handlers = [
|
|
h
|
|
for h in logger.handlers
|
|
if isinstance(h, logging.FileHandler)
|
|
and h.baseFilename == os.path.abspath(log_path)
|
|
]
|
|
self.assertEqual(len(file_handlers), 1)
|
|
|
|
logger.info("dedup-check line")
|
|
for handler in file_handlers:
|
|
handler.flush()
|
|
with open(log_path) as f:
|
|
contents = f.read()
|
|
self.assertEqual(contents.count("dedup-check line"), 1)
|
|
|
|
def test_different_log_files_attach_one_handler_each(self):
|
|
first_path = self._temp_log_path()
|
|
second_path = self._temp_log_path()
|
|
|
|
configure_logging(logging.INFO, first_path)
|
|
configure_logging(logging.INFO, second_path)
|
|
|
|
file_paths = [
|
|
h.baseFilename
|
|
for h in logger.handlers
|
|
if isinstance(h, logging.FileHandler)
|
|
]
|
|
self.assertIn(os.path.abspath(first_path), file_paths)
|
|
self.assertIn(os.path.abspath(second_path), file_paths)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
unittest.main(verbosity=2)
|