Files
parsedmarc/docs/source/kibana.md
T
23c5ea9ad6 Fix DKIM/SPF and SMTP TLS detail-table cross-products in dashboards (#169) (#839)
* Fix DKIM/SPF alignment detail cross-product in dashboards (#169)

Elasticsearch and OpenSearch dynamic-map the dkim_results/spf_results
object arrays as `object` (create_indexes never registers the DSL
document mappings), so Lucene flattens each array into independent
multi-valued fields and stacked terms aggregations on
dkim_results.selector/.domain/.result return every combination of
values across a report's signatures — each phantom row repeating the
full message count.

Aggregate documents now also carry dkim_results_combined and
spf_results_combined: one "selector / domain / result"
("scope / domain / result") string per auth result, composed in
add_dkim_result/add_spf_result. The Kibana/OpenSearch Dashboards and
Grafana (Elasticsearch) alignment-detail tables aggregate those
instead, and the Splunk detail panels pair the values with
mvzip/mvexpand. A documented idempotent _update_by_query backfills
documents saved by older versions; the query matches only documents
that have auth results and lack the combined fields, because an
`exists` query cannot see an empty array.

Also corrects the dead _SPFResult.results (plural) declaration to
`result` (the save path always wrote the singular key), fixes the
result parameter annotations on add_dkim_result/add_spf_result, and
removes the Grafana dmarcian.com DKIM-checker data link, which
required the separate domain/selector columns.

The SMTP TLS visualizations have the same class of defect and are
tracked separately.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address Copilot review findings on #839

Reword the combined-field regression test docstrings: the DKIM/SPF
auth results are dynamic-mapped as plain `object`, not the `nested`
mapping type the previous wording implied — the distinction is the
crux of the fix. Also drop the inert renameByName entries Copilot
flagged on the Grafana Overview and DKIM Alignment Details panels,
which referenced fields those panels' queries no longer (or never)
produced.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Add PR #839 review lessons to AGENTS.md

Extend the "Review passes cover prose" section with two rules from
the #839 Copilot findings: docstrings/comments get the same
text-level review pass as docs and dashboard labels (with suspicion
for dual-use terms like "nested" near Elasticsearch code), and
inert config entries inside hunks a PR already rewrites should be
cleaned rather than preserved to minimize the diff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Reflow create_indexes comments flagged by Copilot

The line wrap placed "#169" directly after the comment marker, so the
raw source read "# #169;". Reword so the issue reference stays on one
line.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Extend the hunk-proofreading rule with rendered-text wraps

Fold the PR #839 second-round Copilot lesson into the existing rule:
proofread how wrapped lines render (comment markers, punctuation at
wrap points), not just the wording itself.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Rename Overview combined-result column labels (Copilot round 3)

The Overview table's "DKIM Auth Result" / "SPF Auth Result" labels
were kept when the columns switched to the combined
"selector / domain / result" values, leaving the headers misleading.
Rename them to match the detail panels' convention and retarget the
byName width overrides that matched the old labels, widening them for
the longer values.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Document per-signature row semantics in the alignment tables

A message carrying multiple DKIM signatures appears once per signature
in the details tables, so summing the messages column across rows can
exceed the total message count. State that explicitly rather than
leaving readers to infer it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address Copilot round-4 findings on dashboards

Fix three pre-existing saved-object title typos in the OpenSearch
ndjson (leading space on "Aggregate DMARC passed DMARC", trailing
space on "Aggregate DMARC reporting organizations", double space in
"map  of message sources by country"), in both the top-level title
and the embedded visState title.

Normalize the Splunk DKIM details placeholders: the base search's
fillnull renders wholly-missing DKIM fields as the literal string
"null", so unsigned mail showed "null / null / null" while the SPF
panel shows "none". Rewrite the values to "none" after the signature
split, where the fields are single-valued and the mvzip pairing
cannot be disturbed. Verified against the dev Splunk that no
truncation or mis-pairing occurs either way, since fillnull
guarantees the fields are never actually null.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Backfill combined DKIM/SPF fields automatically at startup

migrate_indexes() now backfills dkim_results_combined and
spf_results_combined on aggregate documents saved by older versions,
so ES/OS users get historical data in the reworked alignment tables
without running the documented _update_by_query by hand. The backfill
is submitted as a non-blocking background task
(wait_for_completion=false, conflicts=proceed) guarded by a cheap
count query, making repeated startups a fast no-op once an index is
backfilled; any cluster error is logged as a warning and retried at
the next startup rather than raised. The manual command remains
documented for users who upgrade dashboards without pointing the new
parsedmarc at the cluster or who want to control write-load timing.

The legacy published_policy.fo long-to-text reindex migration in the
OpenSearch module is kept ahead of the new backfill, for clusters
upgraded from very old data.

Verified end-to-end against the live dev environment: a real CLI
startup backfilled 9 stripped OpenSearch documents (logged with task
ID) while the already-backfilled Elasticsearch side stayed silent,
and a second startup was silent on both engines.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Match backfill guard on either domain or result subfield

End-to-end upgrade testing (real parsedmarc 10.2.4 ingest, then a
branch startup) surfaced that an exists query cannot see an empty
string: a text field with no tokens is invisible to exists. The
parsers we audited never store an auth result with an empty or
missing domain — they drop such entries entirely, so the previous
domain-only guard was sufficient for their data — but the storage
shape of every historical parsedmarc version can't be audited, so
the guard (and the documented manual command) now matches either
the domain or the result subfield per protocol. Matching either
costs nothing and cannot skip a document that has something to
backfill.

Verified by recomputing expected combined values from _source for
all 2,299 documents on both engines: every document with stored
auth results has exactly the recomputed pairs, and the legacy fo
migration correctly did not fire on typeless indexes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Add auth-result filter controls to the Kibana/OSD aggregate dashboard

The combined per-signature columns fixed the #169 cross-product but left
no way to click-filter by an individual selector, domain, or result. Add
an "Aggregate DMARC auth result filters" input_control_vis panel above
the SPF/DKIM details tables with six option-list dropdowns (DKIM
selector/domain/result, SPF scope/domain/result) that emit ordinary
dashboard-wide filter pills. Works on both Kibana 8.19 and OpenSearch
Dashboards 3, verified by driving the controls in both UIs against the
issue's two-signature repro report.

Documented in kibana.md, including the flat-mapping caveat: combining
two component filters matches documents where any signature satisfies
each condition individually.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Scale Grafana source-country map markers with message volume

The "Map of Message Source Countries" panel drew fixed 5 px dark-green
markers at 50% opacity — nearly invisible on the dark basemap, so the
panel read as empty even when data was flowing (verified via the query
API). Markers now scale with Sum(message_count) (min 4, max 30 px) at
0.8 opacity in a higher-contrast green. Pre-existing issue; the
identically-styled failure-dashboard map panel is intentionally left
untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Correct the nested-mapping rationale in the create_indexes comments

The comments claimed Kibana/OSD/Grafana "cannot terms-aggregate fields
inside a nested mapping" — too absolute. Fact-checked empirically and
against primary docs: Kibana/OSD visual editors (Lens and classic
Visualize) do not support nested fields, but Vega panels can run nested
aggregations (they just cannot render tables, per Elastic's docs), and
Grafana >= 9.4 has a nested bucket aggregation (grafana/grafana#62301)
but no reverse_nested, so parent-level metrics like Sum(message_count)
return 0 inside per-signature buckets (reproduced live). Conclusion
unchanged: the dynamic object mapping stays load-bearing for the
shipped dashboards. PR #839's body was updated with the same
correction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Pin ruff exactly, matching the existing pyright pin rationale

CI installs the [build] extra fresh on every run, and ruff was the one
lint tool left unpinned. ruff 0.16.0 (released this week) began
flagging this codebase's str.format() house style, so every PR started
failing lint on lines it never touched. Pin to 0.15.21 — the version
the codebase is clean under — with the same bump-deliberately comment
pyright carries. Upgrading to 0.16 and converting to f-strings can be
its own PR.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address Copilot findings: harden migrate_indexes, normalize panel titles

Three unresolved review threads, all verified against cli.py's
re-raising init handler before fixing:

- elastic.py/opensearch.py: connections.get_connection() sat outside
  migrate_indexes()'s try/except, so a connection-registration failure
  would abort startup despite the docstring's promise that migration
  errors are caught and logged. Now caught, logged, and skipped until
  the next startup.
- opensearch.py: the legacy published_policy.fo migration loop did
  unguarded network I/O (exists/get_field_mapping/reindex/delete), so a
  transient cluster error aborted startup on the OpenSearch path while
  the identical situation on the Elasticsearch path was logged and
  survived. Each index's migration attempt is now wrapped, warns, and
  moves on.
- opensearch_dashboards.ndjson: normalized two pre-existing panel
  titles in the aggregate dashboard's panelsJSON ("Reporting
  organizations " trailing space, "Map  of message sources by country"
  double space).

Regression tests assert migrate_indexes never propagates connection or
per-index cluster errors on either backend.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Clarify that Nested() on auth-result fields is in-memory shape only

Copilot flagged that _AggregateReportDoc declares dkim_results and
spf_results with Nested(...) while the create_indexes comment insists
the stored mapping must stay dynamic `object`. Both are true: the
Nested declaration only shapes the DSL's in-memory document building
and is never installed as a mapping, because create_indexes skips
Index.document() registration. Say so at both sites, in both backends,
so nobody "fixes" the mismatch by registering the mapping — which
would install real nested mappings and blank the dashboards.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Refer to the filter panel by its displayed title in docs and CHANGELOG

The dashboard convention is a short panel display title backed by a
long-form saved-object name ("SPF details" / "Aggregate DMARC SPF
details"), and the new controls panel follows it. The docs and
CHANGELOG named the panel by its saved-object title, which is not what
a user sees on the dashboard; use the displayed "Auth result filters"
instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Fix misspelled column label in the failure email samples table

The "DMARC failure email samples" visualization labeled its
authentication_results column "autentication_results". The underlying
field reference was already correct; only the user-facing customLabel
was misspelled. A sweep of every title and customLabel in the ndjson
found no other misspellings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Extend the combined-field fix to SMTP TLS documents

SMTP TLS reports have the same cross-product defect as the DKIM/SPF
alignment tables (issue #169), one level deeper: policies is an object
array and each policy's failure_details is an object array inside it,
so stacked terms aggregations on their subfields fabricate rows.

Documents now also carry policies_combined ("domain / type" per policy)
and failure_details_combined ("domain / type / result / sending mta /
receiving ip / mx" per failure detail), composed at save time with the
same "none" fallbacks as the aggregate fields. migrate_indexes() gains
smtp_tls_indexes and backfills old documents with the same guarded,
non-blocking update_by_query pattern; cli.py wires the index name in on
both backends, and the manual _update_by_query command is documented.

Also fixes two adjacent dead fields: add_failure_details stored
additional_information_uri under the wrong constructor kwarg
(additional_information), and receiving_mx_hostname had no declaration
despite always being stored.

Verified live on ES 8.19 and OpenSearch 3: a two-policy repro report
yields exactly 2 policy rows and 2 failure-detail rows via the combined
fields where the old stacked aggregations return 4 of each; the startup
backfill converted the 4 pre-existing sample documents on both engines
with zero recompute mismatches.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Rework the SMTP TLS dashboards onto the combined fields

Kibana/OSD: "SMTP TLS domains" replaces its stacked policy_domain ×
policy_type terms with one terms agg on policies_combined.keyword;
"SMTP TLS failure details" replaces six stacked terms spanning both
array levels with one on failure_details_combined.keyword; the
smtp_tls* index-pattern field cache gains the new fields. The
"reporting organizations" table only buckets on doc-level org_name and
needed no change.

Splunk: the base search now expands policies at the JSON level (spath +
mvexpand) so policy fields are scalars per event, and the failure
details panel expands the second level the same way — sums are the
detail's own failed_session_count, correctly paired. Verified via the
search REST API: a two-policy repro returns exactly one row per real
failure detail with per-detail counts.

kibana.md documents the per-policy/per-detail row semantics and the
honest caveat that session-count sums remain per report document.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Persist additional_info_uri from parsed SMTP TLS failure details

Copilot caught that the savers read additional_information_uri from
the parsed failure-detail dict, but the parser's key is
additional_info_uri (SMTPTLSFailureDetailsOptional in types.py, set in
parse_smtp_tls_report_json), so the URI was never persisted — the
read-side half of the dead-field bug whose write-side half (wrong
constructor kwarg) was fixed earlier. Read the parser's key first,
keeping the long-form key as a fallback for dicts built by other
callers. Regression test proven to fail on the unfixed savers.

Also restructured the expected combined-string test values into named
locals so no implicit string concatenation sits inside a list literal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Say "inner doc", not "nested doc", in the singular-key test docstrings

Final review sweep: in this codebase "nested" is reserved for the
Elasticsearch mapping type, and these InnerDoc-serialization
docstrings used it colloquially — the same dual-use-term trap
documented in AGENTS.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 12:29:40 -04:00

7.0 KiB

Using the Kibana dashboards

The Kibana DMARC dashboards are a human-friendly way to understand the results from incoming DMARC reports.

There is no separate Kibana export — Kibana 8.x's saved-object migration handlers accept the OpenSearch Dashboards format directly, so Kibana users import the bundled dashboards/opensearch/opensearch_dashboards.ndjson in Stack Management → Saved Objects → Import. A CI check imports the same file into a Kibana 8.x container on every change so this stays compatible.

:::{note} The default dashboard is DMARC aggregate reports. To switch between dashboards, click on the Dashboard link on the left side menu of Kibana. :::

DMARC aggregate reports

As the name suggests, this dashboard is the best place to start reviewing your aggregate DMARC data.

Across the top of the dashboard, three pie charts display the percentage of alignment pass/fail for SPF, DKIM, and DMARC. Clicking on any chart segment will filter for that value.

:::{note} Messages should not be considered malicious just because they failed to pass DMARC; especially if you have just started collecting data. It may be a legitimate service that needs SPF and DKIM configured correctly. :::

Start by filtering the results to only show failed DKIM alignment. While DMARC passes if a message passes SPF or DKIM alignment, only DKIM alignment remains valid when a message is forwarded without changing the from address, which is often caused by a mailbox forwarding rule. This is because DKIM signatures are part of the message headers, whereas SPF relies on SMTP session headers.

Underneath the pie charts. you can see graphs of DMARC passage and message disposition over time.

Under the graphs you will find the most useful data tables on the dashboard. On the left, there is a list of organizations that are sending you DMARC reports. In the center, there is a list of sending servers grouped by the base domain in their reverse DNS. On the right, there is the "Message volume and DMARC compliance by from domain" table, which lists email from domains with their message volume and a percentage of those messages that passed DMARC.

By hovering your mouse over a data table value and using the magnifying glass icons, you can filter on or filter out different values. Start by looking at the Message Sources by Reverse DNS table. Find a sender that you recognize, such as an email marketing service, hover over it, and click on the plus (+) magnifying glass icon, to add a filter that only shows results for that sender. Now, look at the Message volume and DMARC compliance by from domain table to the right. That shows you the domains that a sender is sending as, and what share of that traffic is passing DMARC, which might tell you which brand/business is using a particular service. With that information, you can contact them and have them set up DKIM.

:::{note} The "Message volume and DMARC compliance by from domain" table is a TSVB visualization, used because per-domain compliance percentages require a Filter Ratio metric that agg-based data tables can't compute. It renders correctly on Kibana 8.x as imported, but editing it requires first enabling the metrics:allowStringIndices advanced setting, since it references the dmarc_aggregate* index as a string pattern, which Elastic has deprecated. :::

:::{note} If you have a lot of B2C customers, you may see a high volume of emails as your domains coming from consumer email services, such as Google/Gmail and Yahoo! This occurs when customers have mailbox rules in place that forward emails from an old account to a new account, which is why DKIM authentication is so important, as mentioned earlier. Similar patterns may be observed with businesses who send from reverse DNS addressees of parent, subsidiary, and outdated brands. :::

Further down the dashboard, you can filter by source country or source IP address.

Tables showing SPF and DKIM alignment details are located under the IP address table. Each row of the DKIM details table is one real DKIM signature, shown as a combined selector / domain / result value; the SPF details table shows scope / domain / result the same way. Combining the values into one column keeps each signature's selector, domain, and result paired together, rather than aggregating them as separate columns. Because a message that carries multiple DKIM signatures appears once per signature, summing the messages column across rows can exceed the total number of messages.

The "Auth result filters" panel above the details tables provides dropdowns for the individual auth-result components — DKIM selector, DKIM domain, DKIM result, SPF scope, SPF domain, and SPF result — and filters the whole dashboard by them. Because components from different signatures of the same message are indexed together, combining two of these component filters matches documents where any signature satisfies each condition individually, not necessarily the same signature; the combined selector / domain / result (scope / domain / result) column remains the per-signature source of truth.

:::{note} The alignment tables (SPF details, DKIM details) and the per-IP source table live on the same dashboard, further down. To view failures only, use the pie chart at the top of the page as a filter. :::

Any other filters work the same way. You can also add your own custom temporary filters by clicking on Add Filter at the upper right of the page.

DMARC failure reports

The DMARC failure reports dashboard (formerly DMARC Forensic Samples) contains information on DMARC failure reports (also known as forensic or ruf reports). These reports contain samples of emails that have failed to pass DMARC.

:::{note} Most recipients do not send failure/ruf reports at all to avoid privacy leaks. Some recipients (notably Chinese webmail services) will only supply the headers of sample emails. Very few provide the entire email. :::

SMTP TLS reporting

The SMTP TLS reporting dashboard surfaces aggregate counts of TLS-RPT reporting organizations, the policy domains they report on, and the specific failure types — certificate expiry, STARTTLS not supported, STS policy fetch errors, validation failures, and similar — together with the sending and receiving MTA addresses involved.

Like the DKIM and SPF details tables above, the "SMTP TLS domains" and "SMTP TLS failure details" tables show one row per policy and one row per failure detail, respectively, using combined policy (domain / type) and failure detail (domain / type / result / sending mta / receiving ip / mx) columns so that each policy's or failure detail's fields stay paired together, rather than aggregating them as separate columns. The successful_sessions and failed_sessions columns are summed per report document, though, not per policy: when a single report carries multiple policies, a row's session sums include the sibling policies from that report as well as its own. Fully attributing session counts to a single policy would require restructuring the stored documents.