mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-09-06 05:57:58 +00:00
09c88ca2a36af198ebcc85413ce2fa92ff555093
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
07bca1ad28 |
Make the output and mailbox integrations optional extras (#888)
* Make the output and mailbox integrations optional extras (#883) Breaking change for the next major release: pip install parsedmarc now installs the parsing core plus a working core CLI (file, IMAP, Maildir, and mbox input; CSV/JSON, Splunk HEC, webhook, and syslog output). Everything else moves behind an extra: elastic, opensearch, kafka, s3, gelf, loganalytics, msgraph, and gmail, joining the existing postgresql extra, with an umbrella [all] that deliberately excludes postgresql (psycopg's binary wheels do not exist on every platform, so parsedmarc[all] must never fail to install there). cli.py imports the six SDK-dependent output modules behind the #884 TYPE_CHECKING/try-except guard; a configured section whose extra is missing fails fast with a ConfigurationError naming the section and the exact pip install command — including the msgraph and gmail_api mailbox sections (detected via parsedmarc.mail's placeholder classes) and postgresql (checked before the constructor so the startup retry loop does not retry a missing dependency for a minute). The Azure/kiota Graph error types fall back to never-raised sentinel classes. The Docker image installs [all,postgresql], so container users see no change. CI lint installs [build,all,postgresql]; the unit-test job installs [build,all], deliberately without postgresql so test_postgres.py's absent-psycopg arm stays exercised. The never-imported dateparser dependency is dropped in favor of declaring python-dateutil, which utils.py actually imports; pytz moves to the build extra for the one test that uses it. Verified live: a no-extras wheel install imports, parses samples, and reports the install hint for each gated section; a [all] install restores every integration; the Docker image builds with every SDK importable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Patch psycopg presence in the PostgreSQL CLI wiring tests CI's unit-test job deliberately installs [build,all] without the postgresql extra, so parsedmarc.cli.postgres.psycopg is None there and the new missing-extra presence check correctly made _main exit 1 before the wiring under test ran. The tests simulate the SDK being available (PostgreSQLClient is mocked at the SDK boundary), so the module-level psycopg handle is now patched present in setUp. Verified against a simulated psycopg-absent environment as well as the local full install. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot review: narrow guards to ModuleNotFoundError, fix docs - The optional-integration and Graph error-type import guards now catch ModuleNotFoundError instead of ImportError, so only a genuinely absent package reads as a missing extra; a broken-but-present SDK fails loudly with its real error instead of masquerading as one. The test blocker raises ModuleNotFoundError accordingly — the exact exception a missing package produces. - _missing_extra_hint docstring no longer calls every gated integration an output module (it also serves the msgraph/gmail_api mailbox sections). - Fix the pre-existing passsword typo in usage.md's kafka section; the INI key the code reads is password (cli.py _parse_config). Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Quote extras specs in copy-paste install commands From Copilot's second review round: zsh treats an unquoted .[build,all] as a glob and fails with 'no matches found', so the commands shown in AGENTS.md, CONTRIBUTING.md, dashboards/README.md, and the bootstrap script's comment are now quoted. The CI workflows keep the unquoted form: they run under bash, which passes unmatched globs through literally. The suggestion to change the 'Choosing what to install' heading level was rejected — it is a subsection of 'Installing parsedmarc', matching the file's existing hierarchy. Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix upgrade command in the changelog * Documentation review: accuracy, spelling, grammar, and clarity pass A full prose review of docs/source, README, CONTRIBUTING, and the dashboards README, with every accuracy claim verified against the code before changing it. Highlights: - usage.md: documented six missing [general] options (the CSV/JSON filename options, prettify_json, normalize_timespan_threshold_hours), the required kafka smtp_tls_topic, [imap] timeout/max_retries, and the postgresql env-var prefix; corrected the maildir_path default (None, not INBOX — cli.py Namespace defaults), the mailbox check_timeout option name, the systemd restart interval (RestartSec is 5m), and merged the duplicate silent entry; quoted every copy-paste extras spec for zsh safety. - elasticsearch.md: fixed an invalid openssl command (rsa:4096 -nodes), the dashboards filename (opensearch_dashboards.ndjson, matching the file the link serves), and assorted grammar. - davmail.md: the service-enable command now enables davmail.service (was parsedmarc.service — a copy-paste error that left DavMail unenabled), plus a view typo and DavMail capitalization. - output.md: the example schema reference is RFC 7489 Appendix C (7480 is RDAP). kibana.md: SPF relies on the SMTP envelope, not session headers (RFC 7208). dmarc.md: DKM -> DKIM. - README: the intro now also names the OpenSearch/Grafana stack, matching the feature list. CONTRIBUTING: pre-PR checks now include ruff format --check and pyright, matching CI's lint job. - dashboards/README: the service table and seed description now include the PostgreSQL backend the compose stack runs. Sample data blocks, the CLI-help mirror block, and released CHANGELOG entries were deliberately left untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Docstring review: accuracy, spelling, grammar, and clarity pass Every docstring in parsedmarc/, parsedmarc/mail/, the maps maintainer scripts, and the test suite reviewed with each claim verified against the code it documents. Text-only — no behavior changes. Highlights: - Copy-paste errors corrected: parsed_smtp_tls_reports_to_csv and splunk/loganalytics save functions described aggregate or failure reports they do not handle; LogAnalyticsException claimed to be an Elasticsearch error. - Docstring/behavior mismatches: parse_report_email's report_type enumeration omitted smtp_tls; parse_failure_report typed msg_date as str (it is datetime); strip_attachment_payloads claimed payloads are replaced with None (the key is deleted); kafkaclient's failure and SMTP TLS savers claimed per-record slicing while sending the whole list in one message (docstrings now describe reality — whether slicing was intended is flagged for follow-up); the postgres savers claimed to take parse_report_file's return value but receive the inner report dict; elastic/opensearch save functions' Raises listed only AlreadySaved. - None-as-semantic-state documented where missing (get_base_domain, get_ip_address_country), enumeration completeness fixed (get_ip_address_info's 9 result keys, maps script outputs, TSV columns), and the stale 44-industry-types count corrected to the 46 the authoritative README list defines. - Test docstrings aligned with what the tests actually assert, including two that overstated coverage of the elastic/opensearch address-list tests. - Two argparse help strings fixed: file_path now names SMTP TLS report files alongside aggregate and failure, mirrored into usage.md's CLI-help block; --offline's doubled spaces removed (rendered help unchanged). - elasticsearch.md's security claim corrected against Elastic's docs: security is enabled and auto-configured on first startup since 8.0 (not "8.7 secure mode"), so the settings are verified, not hand-written. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> |
||
|
|
5c7f4ba048 |
Clear shared IP cache in parallel test setUp for isolation (#866)
tests/test_parallel.py's parity test compares sequential parse_report_file(path, offline=True) results in the parent process against results from cold worker processes. When the full suite runs locally (GITHUB_ACTIONS unset), tests/test_init.py runs first with real DNS lookups and warms the shared module-level parsedmarc.IP_ADDRESS_CACHE; get_ip_address_info consults the cache before honoring offline, so the sequential baseline returned DNS-enriched entries (e.g. reverse_dns='smtp7.cardinal.com') while the workers correctly returned None, failing the test. CI never sees this because it runs offline from the start. Clear the cache in _ParallelTestCase.setUp so baselines and workers both start cold. Test-only change; the cache-before-offline ordering in utils.py is intentional (a cache hit makes no network queries). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4e80047e68 |
Centralize configuration handling with ParserConfig (#503) (#851)
* Centralize configuration handling with ParserConfig (#503) Add parsedmarc/config.py with ParserConfig, a frozen dataclass carrying every parsing/enrichment option plus the three shared caches (IP address info, seen aggregate report IDs, reverse DNS map). All eight public parsing/mailbox functions accept a keyword-only config= argument; when provided, the individual option keyword arguments are ignored in favor of the config's values, and every existing per-option keyword argument keeps working unchanged. The three hand-copied parse_kwargs dicts and the dns_timeout<->timeout rename chain are gone; the CLI builds one ParserConfig per run (rebuilt on SIGHUP) and passes it everywhere. Explicitly constructed configs own fresh isolated caches; dataclasses.replace() shares them; pickling drops cache contents and rebinds the unpickling process's module defaults, preserving the per-worker cache behavior of n_procs parallel parsing. The module globals IP_ADDRESS_CACHE / SEEN_AGGREGATE_REPORT_IDS / REVERSE_DNS_MAP remain, identity-preserved, as re-exports of the default caches. Bug fixes that ride along, each with a regression test: - One-shot mailbox runs now honor [general] dns_timeout/dns_retries; the CLI call site never forwarded dns_timeout, so the library's stray 6.0 default silently applied. - Lazily-triggered reverse DNS map loads (get_ip_address_info / get_service_from_reverse_dns_base_domain, including in n_procs workers) now thread psl_overrides_path/psl_overrides_url through to load_reverse_dns_map instead of clobbering operator-configured PSL overrides with the bundled defaults. - get_dmarc_reports_from_mailbox() and watch_inbox() dns_timeout defaults unified to DEFAULT_DNS_TIMEOUT (2.0s) from a stray 6.0, and normalize_timespan_threshold_hours to the float 24.0 used everywhere else. Closes #503 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address CI and review feedback on #851 - Wrap the IMAPConnection example in usage.md so ruff format is clean over the docs code blocks (CI runs ruff format --check on the whole repo; the local runs were scoped to parsedmarc/ and tests/ and missed it). - Fix the pre-existing "URL ro a reverse DNS map" docstring typo in get_service_from_reverse_dns_base_domain, caught by Copilot on the adjacent hunk. - Import parsedmarc.config once, as an aliased plain import, in tests/test_config.py instead of mixing import and import-from of the same module (flagged by code quality scanning); the aliased module import also keeps pyright able to resolve the submodule attribute access. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address second Copilot review round on #851 - Add parse_aggregate_report_file() to the library entry-point list in usage.md; the following paragraph describes the config= contract for "each of these functions", so the list must name all eight config-accepting entry points. - ParserConfig.__setstate__ now initializes every non-cache field to its class default before applying the pickled state, so a config serialized by an older parsedmarc version (whose state predates fields added later) unpickles with the newer fields at their defaults instead of unset entirely (__init__ never runs during unpickling, so an absent field would raise AttributeError on first access). Covered by a regression test that feeds __setstate__ a partial state dict. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
48445c639e |
Extend n_procs parallel parsing to mbox and mailbox sources (#147) (#849)
* Extend n_procs parallel parsing to mbox and mailbox sources (#147) n_procs previously applied only to report files passed directly as CLI arguments; messages from mbox files and mailbox connections (IMAP, Microsoft Graph, Gmail API, Maildir) were always parsed sequentially. A new parsedmarc.parallel module provides a shared bounded-window ProcessPoolExecutor helper (parallel_map) used by all three input paths. Only parsing fans out to a reused worker pool; message fetching, report deduplication, mailbox archiving/deletion, and output stay sequential in the main process. The submission window keeps at most ~2*n_procs messages in flight, so memory stays bounded even for huge mboxes, and the mailbox path fetches messages lazily on the connection-owning main thread. keep_alive never crosses the process boundary - the main process sends periodic IMAP keepalives while workers parse - and with n_procs > 1, invalid-message disposition happens after the parse phase, mirroring the existing deferred bulk archive moves. get_dmarc_reports_from_mbox, get_dmarc_reports_from_mailbox (including its tail-recursive re-check), and watch_inbox gain an n_procs keyword argument (default 1); sequential behavior at the default is unchanged. Replacing the CLI's hand-rolled Pipe/Process batching also fixes two defects in the direct-file path: a child process that died from a non-ParserError exception left the parent blocked forever on conn.recv(), and the hard batch barrier let one slow file idle every other worker slot. Workers are now a reused pool (no fresh interpreter per file), with worker logging reconstructed via a spawn-safe pool initializer instead of fork inheritance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot and code-quality review feedback on #849 - parallel_map now validates n_procs >= 1 itself with a clear error instead of surfacing ProcessPoolExecutor's max_workers error later. The check raises eagerly at the call (the generator body moved into an inner function) rather than on first iteration, with a regression test. - Aligned parallel_map's should_stop docstring with the implementation: queued-but-unstarted jobs are cancelled, while in-flight jobs are waited on and their results yielded, so the stop can block briefly but never discards completed work. - The parallel mailbox path keeps fetched message ids in a deque popped as each in-order result arrives, so the id queue stays bounded by the submission window instead of growing to message_limit. - Closed the three sample-file handles the new tests opened without a context manager. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address second round of Copilot feedback on #849 - configure_logging no longer stacks duplicate FileHandlers when called again with the same log_file (compared by FileHandler.baseFilename, which stores the absolute path): a duplicate wrote every record twice and leaked a file descriptor per call, e.g. across SIGHUP config reloads. Latent in the pre-extraction cli._configure_logging too. Regression tests in the new tests/test_log.py. - Renamed the CHANGELOG's premature "10.4.0" heading to "Unreleased", matching the project convention where in-progress entries accumulate under Unreleased and the release PR renames the section and bumps parsedmarc/constants.py together (as in the 10.3.0 release). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |