* Make the output and mailbox integrations optional extras (#883)
Breaking change for the next major release: pip install parsedmarc now
installs the parsing core plus a working core CLI (file, IMAP, Maildir,
and mbox input; CSV/JSON, Splunk HEC, webhook, and syslog output).
Everything else moves behind an extra: elastic, opensearch, kafka, s3,
gelf, loganalytics, msgraph, and gmail, joining the existing postgresql
extra, with an umbrella [all] that deliberately excludes postgresql
(psycopg's binary wheels do not exist on every platform, so
parsedmarc[all] must never fail to install there).
cli.py imports the six SDK-dependent output modules behind the #884
TYPE_CHECKING/try-except guard; a configured section whose extra is
missing fails fast with a ConfigurationError naming the section and the
exact pip install command — including the msgraph and gmail_api mailbox
sections (detected via parsedmarc.mail's placeholder classes) and
postgresql (checked before the constructor so the startup retry loop
does not retry a missing dependency for a minute). The Azure/kiota Graph
error types fall back to never-raised sentinel classes.
The Docker image installs [all,postgresql], so container users see no
change. CI lint installs [build,all,postgresql]; the unit-test job
installs [build,all], deliberately without postgresql so
test_postgres.py's absent-psycopg arm stays exercised. The
never-imported dateparser dependency is dropped in favor of declaring
python-dateutil, which utils.py actually imports; pytz moves to the
build extra for the one test that uses it.
Verified live: a no-extras wheel install imports, parses samples, and
reports the install hint for each gated section; a [all] install
restores every integration; the Docker image builds with every SDK
importable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Patch psycopg presence in the PostgreSQL CLI wiring tests
CI's unit-test job deliberately installs [build,all] without the
postgresql extra, so parsedmarc.cli.postgres.psycopg is None there and
the new missing-extra presence check correctly made _main exit 1 before
the wiring under test ran. The tests simulate the SDK being available
(PostgreSQLClient is mocked at the SDK boundary), so the module-level
psycopg handle is now patched present in setUp. Verified against a
simulated psycopg-absent environment as well as the local full install.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Address Copilot review: narrow guards to ModuleNotFoundError, fix docs
- The optional-integration and Graph error-type import guards now catch
ModuleNotFoundError instead of ImportError, so only a genuinely absent
package reads as a missing extra; a broken-but-present SDK fails
loudly with its real error instead of masquerading as one. The test
blocker raises ModuleNotFoundError accordingly — the exact exception a
missing package produces.
- _missing_extra_hint docstring no longer calls every gated integration
an output module (it also serves the msgraph/gmail_api mailbox
sections).
- Fix the pre-existing passsword typo in usage.md's kafka section; the
INI key the code reads is password (cli.py _parse_config).
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Quote extras specs in copy-paste install commands
From Copilot's second review round: zsh treats an unquoted .[build,all]
as a glob and fails with 'no matches found', so the commands shown in
AGENTS.md, CONTRIBUTING.md, dashboards/README.md, and the bootstrap
script's comment are now quoted. The CI workflows keep the unquoted
form: they run under bash, which passes unmatched globs through
literally. The suggestion to change the 'Choosing what to install'
heading level was rejected — it is a subsection of 'Installing
parsedmarc', matching the file's existing hierarchy.
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Fix upgrade command in the changelog
* Documentation review: accuracy, spelling, grammar, and clarity pass
A full prose review of docs/source, README, CONTRIBUTING, and the
dashboards README, with every accuracy claim verified against the code
before changing it. Highlights:
- usage.md: documented six missing [general] options (the CSV/JSON
filename options, prettify_json, normalize_timespan_threshold_hours),
the required kafka smtp_tls_topic, [imap] timeout/max_retries, and
the postgresql env-var prefix; corrected the maildir_path default
(None, not INBOX — cli.py Namespace defaults), the mailbox
check_timeout option name, the systemd restart interval (RestartSec
is 5m), and merged the duplicate silent entry; quoted every
copy-paste extras spec for zsh safety.
- elasticsearch.md: fixed an invalid openssl command (rsa:4096 -nodes),
the dashboards filename (opensearch_dashboards.ndjson, matching the
file the link serves), and assorted grammar.
- davmail.md: the service-enable command now enables davmail.service
(was parsedmarc.service — a copy-paste error that left DavMail
unenabled), plus a view typo and DavMail capitalization.
- output.md: the example schema reference is RFC 7489 Appendix C
(7480 is RDAP). kibana.md: SPF relies on the SMTP envelope, not
session headers (RFC 7208). dmarc.md: DKM -> DKIM.
- README: the intro now also names the OpenSearch/Grafana stack,
matching the feature list. CONTRIBUTING: pre-PR checks now include
ruff format --check and pyright, matching CI's lint job.
- dashboards/README: the service table and seed description now include
the PostgreSQL backend the compose stack runs.
Sample data blocks, the CLI-help mirror block, and released CHANGELOG
entries were deliberately left untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* Docstring review: accuracy, spelling, grammar, and clarity pass
Every docstring in parsedmarc/, parsedmarc/mail/, the maps maintainer
scripts, and the test suite reviewed with each claim verified against
the code it documents. Text-only — no behavior changes. Highlights:
- Copy-paste errors corrected: parsed_smtp_tls_reports_to_csv and
splunk/loganalytics save functions described aggregate or failure
reports they do not handle; LogAnalyticsException claimed to be an
Elasticsearch error.
- Docstring/behavior mismatches: parse_report_email's report_type
enumeration omitted smtp_tls; parse_failure_report typed msg_date as
str (it is datetime); strip_attachment_payloads claimed payloads are
replaced with None (the key is deleted); kafkaclient's failure and
SMTP TLS savers claimed per-record slicing while sending the whole
list in one message (docstrings now describe reality — whether
slicing was intended is flagged for follow-up); the postgres savers
claimed to take parse_report_file's return value but receive the
inner report dict; elastic/opensearch save functions' Raises listed
only AlreadySaved.
- None-as-semantic-state documented where missing (get_base_domain,
get_ip_address_country), enumeration completeness fixed
(get_ip_address_info's 9 result keys, maps script outputs, TSV
columns), and the stale 44-industry-types count corrected to the
46 the authoritative README list defines.
- Test docstrings aligned with what the tests actually assert,
including two that overstated coverage of the elastic/opensearch
address-list tests.
- Two argparse help strings fixed: file_path now names SMTP TLS report
files alongside aggregate and failure, mirrored into usage.md's
CLI-help block; --offline's doubled spaces removed (rendered help
unchanged).
- elasticsearch.md's security claim corrected against Elastic's docs:
security is enabled and auto-configured on first startup since 8.0
(not "8.7 secure mode"), so the settings are verified, not
hand-written.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Build multi-arch Docker images with PostgreSQL support
The prebuilt image now installs the `[postgresql]` extra, so the optional
PostgreSQL output backend (psycopg) works out of the box in the container
without a separate `pip install` (#792). The wheel path is resolved into a
variable before appending the extra so the shell doesn't treat
`*.whl[postgresql]` as a bracket glob.
The build workflow now sets up QEMU + Buildx and builds a multi-arch
manifest for `linux/amd64` and `linux/arm64`, so the image runs natively on
64-bit ARM hosts such as a Raspberry Pi (#789). Every compiled dependency
(psycopg[binary], lxml, maxminddb, cryptography) ships prebuilt aarch64
manylinux wheels, so the arm64 build adds no source-compilation step.
A `pull_request` trigger (scoped to the build inputs) and `workflow_dispatch`
are added so the multi-arch build can be validated on PRs and rebuilt on
demand; pushes are still gated on the release event, so neither pushes images.
Closes#789Closes#792
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Bump version to 10.0.4 to publish the new images
The docker workflow only pushes to the registry on a `release` event, so
shipping the multi-arch + PostgreSQL-enabled image requires cutting a
release. 10.0.3 is already tagged, so bump to 10.0.4 and document the
Docker changes in the changelog.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Don't run the docker build on pull requests
The pull_request trigger (added to validate the multi-arch build) re-ran the
full ~10-minute amd64+arm64 build on every commit pushed to a docker-touching
PR, because the pull_request `paths` filter matches against the PR's entire
diff, not just the newest commit. That is wasteful once the build has been
validated.
Drop the pull_request trigger and rely on workflow_dispatch for on-demand
validation (plus the existing master-push and release triggers). Also gate the
registry login on the release event so that no non-release run authenticates
to ghcr at all — a build can only ever be pushed from a published release.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Use multi-stage build to reduce image size
* Add ARGS to be more flexible during image builds
* Create user and use it instead of root
* Don't update pip in container. The Python image should have a recent
version
* add dockerfile and actions task to build image
* test on branch
* change to push only on release, update readme
* remove pip install requirements
* change to on release github action