mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-07-29 20:04:56 +00:00
7d87ba18bee874b315f6c204cafb6f49b4f5a548
* Accept plain-text uncategorized-sources lists in find_unknown_base_reverse_dns.py Dashboard exports of uncategorized email sources are plain-text lists of one source name per line — a mix of raw MMDB as_name strings (when the source IP had no PTR and resolved via the IPinfo Lite MMDB) and base reverse-DNS domains. The script already translates as_names to their as_domain and subtracts mapped/known-unknown entries, but only read a hardcoded source_name-headed CSV. Add -i/--input and -o/--output flags (defaults preserve current behavior) and auto-detect the input format from the first line: a source_name CSV header selects the existing DictReader path, anything else is read as plain text with each line taken verbatim (never comma-split, since as_names contain commas) and deduped case-insensitively. Fix the missing-input error message, which reported the map path instead of the input path. Document the new entry point in the maps README and AGENTS.md, and make explicit in the brand-quality triage rule that map display names must be human-friendly operator names — never raw as_name strings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Add MMDB coverage scan script with anti-poisoning guards find_unmapped_as_domains.py turns the manual "Checking ASN-domain coverage of the MMDB" recipe into a maintainer script: walk every IPv4 record in the bundled IPinfo Lite MMDB, aggregate routed footprint per as_domain, subtract mapped/known-unknown keys, apply PSL folding and the full-IP privacy filter, and emit domain,ipv4_count,as_name sorted by footprint for the collector -> classifier pipeline. Because ASN registration data is self-declared to the RIRs and as_domain derives from registrant-controlled WHOIS, bulk-categorizing the MMDB needs poisoning defenses: - An IPv4-footprint floor (--min-ips, default 4096, a /20) keeps tiny self-described ASNs out of the auto-classification queue; dropped counts are always printed. - A brand-collision guard in classify_unknown_domains.py loads the existing map (--map) and demotes any single-category candidate whose proposed display name matches an existing map name without a lexical relationship to that operator's keys into the ambiguous bucket (marked name-collision-with-existing-map-entry) for human review. HAND overrides bypass the guard. The guard protects the PTR-side flow as well as the MMDB-coverage flow. Verified: scan yields 132 candidates at the default floor (1512 dropped); collector accepts the output directly; a fixture titled as Comcast under an unrelated domain lands in ambiguous while a comcast-rooted sibling auto-promotes. Also fix the maps README links that still pointed at the root AGENTS.md for the classification workflow after its extraction to maps/AGENTS.md, and correct the classify_tsv docstring's return signature. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Enhance planning guidance in CLAUDE.md by specifying auto mode activation after user approval * Codify triage flagging for identified operators with no fitting type An operator confidently identified from two corroborating sources but matching none of the README's type values should be flagged during triage with a proposed new type for the reviewer, not force-fitted and not silently recorded as known-unknown — KU means "we couldn't identify this", which would bury completed research. Extends workflow rule 7 and the LLM low-confidence list in the maps AGENTS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
fix: OSD Global-tenant import + dropped report files with glob metacharacters; validate dev stack on OpenSearch 3.x with PostgreSQL (#781)
fix: OSD Global-tenant import + dropped report files with glob metacharacters; validate dev stack on OpenSearch 3.x with PostgreSQL (#781)
fix: OSD Global-tenant import + dropped report files with glob metacharacters; validate dev stack on OpenSearch 3.x with PostgreSQL (#781)
fix: OSD Global-tenant import + dropped report files with glob metacharacters; validate dev stack on OpenSearch 3.x with PostgreSQL (#781)
parsedmarc
parsedmarc is a Python module and CLI utility for parsing DMARC
reports. When used with Elasticsearch and Kibana (or Splunk), it works
as a self-hosted open-source alternative to commercial DMARC report
processing services such as Agari Brand Protection, Dmarcian, OnDMARC,
ProofPoint Email Fraud Defense, and Valimail.
Note
Domain-based Message Authentication, Reporting, and Conformance (DMARC) is an email authentication protocol.
Sponsors
This project is maintained by one developer. Please consider sponsoring my work if you or your organization benefit from it.
Features
- Parses aggregate/rua DMARC reports: the legacy draft and 1.0 schemas (RFC 7489) and the new RFC 9990 schema for the final DMARC standard (RFC 9989)
- Parses failure/ruf DMARC reports (RFC 6591 and RFC 9991; formerly called forensic reports)
- Parses reports from SMTP TLS Reporting (TLS-RPT, RFC 8460)
- Can parse reports from an inbox over IMAP, Microsoft Graph, or Gmail API
- Transparently handles gzip or zip compressed reports
- Consistent data structures
- Simple JSON and/or CSV output
- Optionally email the results
- Optionally send the results to Elasticsearch, OpenSearch, Splunk, or PostgreSQL, for use with premade dashboards
- Optionally send the results to Apache Kafka, Amazon S3, Azure Log Analytics (Microsoft Sentinel), a Graylog (GELF) endpoint, a syslog server, or an HTTP webhook
Python Compatibility
This project supports the following Python versions, which are either actively maintained or are the default versions for RHEL or Debian.
| Version | Supported | Reason |
|---|---|---|
| < 3.6 | ❌ | End of Life (EOL) |
| 3.6 | ❌ | Used in RHEL 8, but not supported by project dependencies |
| 3.7 | ❌ | End of Life (EOL) |
| 3.8 | ❌ | End of Life (EOL) |
| 3.9 | ❌ | Used in Debian 11 and RHEL 9, but not supported by project dependencies |
| 3.10 | ✅ | Actively maintained |
| 3.11 | ✅ | Actively maintained; supported until June 2028 (Debian 12) |
| 3.12 | ✅ | Actively maintained; supported until May 2035 (RHEL 10) |
| 3.13 | ✅ | Actively maintained; supported until June 2030 (Debian 13) |
| 3.14 | ✅ | Supported (requires imapclient>=3.1.0) |
Languages
Python
98.6%
Shell
1.3%
