mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-07-28 19:34:55 +00:00
* Fix DKIM/SPF alignment detail cross-product in dashboards (#169) Elasticsearch and OpenSearch dynamic-map the dkim_results/spf_results object arrays as `object` (create_indexes never registers the DSL document mappings), so Lucene flattens each array into independent multi-valued fields and stacked terms aggregations on dkim_results.selector/.domain/.result return every combination of values across a report's signatures — each phantom row repeating the full message count. Aggregate documents now also carry dkim_results_combined and spf_results_combined: one "selector / domain / result" ("scope / domain / result") string per auth result, composed in add_dkim_result/add_spf_result. The Kibana/OpenSearch Dashboards and Grafana (Elasticsearch) alignment-detail tables aggregate those instead, and the Splunk detail panels pair the values with mvzip/mvexpand. A documented idempotent _update_by_query backfills documents saved by older versions; the query matches only documents that have auth results and lack the combined fields, because an `exists` query cannot see an empty array. Also corrects the dead _SPFResult.results (plural) declaration to `result` (the save path always wrote the singular key), fixes the result parameter annotations on add_dkim_result/add_spf_result, and removes the Grafana dmarcian.com DKIM-checker data link, which required the separate domain/selector columns. The SMTP TLS visualizations have the same class of defect and are tracked separately. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot review findings on #839 Reword the combined-field regression test docstrings: the DKIM/SPF auth results are dynamic-mapped as plain `object`, not the `nested` mapping type the previous wording implied — the distinction is the crux of the fix. Also drop the inert renameByName entries Copilot flagged on the Grafana Overview and DKIM Alignment Details panels, which referenced fields those panels' queries no longer (or never) produced. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Add PR #839 review lessons to AGENTS.md Extend the "Review passes cover prose" section with two rules from the #839 Copilot findings: docstrings/comments get the same text-level review pass as docs and dashboard labels (with suspicion for dual-use terms like "nested" near Elasticsearch code), and inert config entries inside hunks a PR already rewrites should be cleaned rather than preserved to minimize the diff. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Reflow create_indexes comments flagged by Copilot The line wrap placed "#169" directly after the comment marker, so the raw source read "# #169;". Reword so the issue reference stays on one line. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Extend the hunk-proofreading rule with rendered-text wraps Fold the PR #839 second-round Copilot lesson into the existing rule: proofread how wrapped lines render (comment markers, punctuation at wrap points), not just the wording itself. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Rename Overview combined-result column labels (Copilot round 3) The Overview table's "DKIM Auth Result" / "SPF Auth Result" labels were kept when the columns switched to the combined "selector / domain / result" values, leaving the headers misleading. Rename them to match the detail panels' convention and retarget the byName width overrides that matched the old labels, widening them for the longer values. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Document per-signature row semantics in the alignment tables A message carrying multiple DKIM signatures appears once per signature in the details tables, so summing the messages column across rows can exceed the total message count. State that explicitly rather than leaving readers to infer it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot round-4 findings on dashboards Fix three pre-existing saved-object title typos in the OpenSearch ndjson (leading space on "Aggregate DMARC passed DMARC", trailing space on "Aggregate DMARC reporting organizations", double space in "map of message sources by country"), in both the top-level title and the embedded visState title. Normalize the Splunk DKIM details placeholders: the base search's fillnull renders wholly-missing DKIM fields as the literal string "null", so unsigned mail showed "null / null / null" while the SPF panel shows "none". Rewrite the values to "none" after the signature split, where the fields are single-valued and the mvzip pairing cannot be disturbed. Verified against the dev Splunk that no truncation or mis-pairing occurs either way, since fillnull guarantees the fields are never actually null. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Backfill combined DKIM/SPF fields automatically at startup migrate_indexes() now backfills dkim_results_combined and spf_results_combined on aggregate documents saved by older versions, so ES/OS users get historical data in the reworked alignment tables without running the documented _update_by_query by hand. The backfill is submitted as a non-blocking background task (wait_for_completion=false, conflicts=proceed) guarded by a cheap count query, making repeated startups a fast no-op once an index is backfilled; any cluster error is logged as a warning and retried at the next startup rather than raised. The manual command remains documented for users who upgrade dashboards without pointing the new parsedmarc at the cluster or who want to control write-load timing. The legacy published_policy.fo long-to-text reindex migration in the OpenSearch module is kept ahead of the new backfill, for clusters upgraded from very old data. Verified end-to-end against the live dev environment: a real CLI startup backfilled 9 stripped OpenSearch documents (logged with task ID) while the already-backfilled Elasticsearch side stayed silent, and a second startup was silent on both engines. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Match backfill guard on either domain or result subfield End-to-end upgrade testing (real parsedmarc 10.2.4 ingest, then a branch startup) surfaced that an exists query cannot see an empty string: a text field with no tokens is invisible to exists. The parsers we audited never store an auth result with an empty or missing domain — they drop such entries entirely, so the previous domain-only guard was sufficient for their data — but the storage shape of every historical parsedmarc version can't be audited, so the guard (and the documented manual command) now matches either the domain or the result subfield per protocol. Matching either costs nothing and cannot skip a document that has something to backfill. Verified by recomputing expected combined values from _source for all 2,299 documents on both engines: every document with stored auth results has exactly the recomputed pairs, and the legacy fo migration correctly did not fire on typeless indexes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Add auth-result filter controls to the Kibana/OSD aggregate dashboard The combined per-signature columns fixed the #169 cross-product but left no way to click-filter by an individual selector, domain, or result. Add an "Aggregate DMARC auth result filters" input_control_vis panel above the SPF/DKIM details tables with six option-list dropdowns (DKIM selector/domain/result, SPF scope/domain/result) that emit ordinary dashboard-wide filter pills. Works on both Kibana 8.19 and OpenSearch Dashboards 3, verified by driving the controls in both UIs against the issue's two-signature repro report. Documented in kibana.md, including the flat-mapping caveat: combining two component filters matches documents where any signature satisfies each condition individually. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Scale Grafana source-country map markers with message volume The "Map of Message Source Countries" panel drew fixed 5 px dark-green markers at 50% opacity — nearly invisible on the dark basemap, so the panel read as empty even when data was flowing (verified via the query API). Markers now scale with Sum(message_count) (min 4, max 30 px) at 0.8 opacity in a higher-contrast green. Pre-existing issue; the identically-styled failure-dashboard map panel is intentionally left untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Correct the nested-mapping rationale in the create_indexes comments The comments claimed Kibana/OSD/Grafana "cannot terms-aggregate fields inside a nested mapping" — too absolute. Fact-checked empirically and against primary docs: Kibana/OSD visual editors (Lens and classic Visualize) do not support nested fields, but Vega panels can run nested aggregations (they just cannot render tables, per Elastic's docs), and Grafana >= 9.4 has a nested bucket aggregation (grafana/grafana#62301) but no reverse_nested, so parent-level metrics like Sum(message_count) return 0 inside per-signature buckets (reproduced live). Conclusion unchanged: the dynamic object mapping stays load-bearing for the shipped dashboards. PR #839's body was updated with the same correction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Pin ruff exactly, matching the existing pyright pin rationale CI installs the [build] extra fresh on every run, and ruff was the one lint tool left unpinned. ruff 0.16.0 (released this week) began flagging this codebase's str.format() house style, so every PR started failing lint on lines it never touched. Pin to 0.15.21 — the version the codebase is clean under — with the same bump-deliberately comment pyright carries. Upgrading to 0.16 and converting to f-strings can be its own PR. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot findings: harden migrate_indexes, normalize panel titles Three unresolved review threads, all verified against cli.py's re-raising init handler before fixing: - elastic.py/opensearch.py: connections.get_connection() sat outside migrate_indexes()'s try/except, so a connection-registration failure would abort startup despite the docstring's promise that migration errors are caught and logged. Now caught, logged, and skipped until the next startup. - opensearch.py: the legacy published_policy.fo migration loop did unguarded network I/O (exists/get_field_mapping/reindex/delete), so a transient cluster error aborted startup on the OpenSearch path while the identical situation on the Elasticsearch path was logged and survived. Each index's migration attempt is now wrapped, warns, and moves on. - opensearch_dashboards.ndjson: normalized two pre-existing panel titles in the aggregate dashboard's panelsJSON ("Reporting organizations " trailing space, "Map of message sources by country" double space). Regression tests assert migrate_indexes never propagates connection or per-index cluster errors on either backend. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Clarify that Nested() on auth-result fields is in-memory shape only Copilot flagged that _AggregateReportDoc declares dkim_results and spf_results with Nested(...) while the create_indexes comment insists the stored mapping must stay dynamic `object`. Both are true: the Nested declaration only shapes the DSL's in-memory document building and is never installed as a mapping, because create_indexes skips Index.document() registration. Say so at both sites, in both backends, so nobody "fixes" the mismatch by registering the mapping — which would install real nested mappings and blank the dashboards. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Refer to the filter panel by its displayed title in docs and CHANGELOG The dashboard convention is a short panel display title backed by a long-form saved-object name ("SPF details" / "Aggregate DMARC SPF details"), and the new controls panel follows it. The docs and CHANGELOG named the panel by its saved-object title, which is not what a user sees on the dashboard; use the displayed "Auth result filters" instead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix misspelled column label in the failure email samples table The "DMARC failure email samples" visualization labeled its authentication_results column "autentication_results". The underlying field reference was already correct; only the user-facing customLabel was misspelled. A sweep of every title and customLabel in the ndjson found no other misspellings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Extend the combined-field fix to SMTP TLS documents SMTP TLS reports have the same cross-product defect as the DKIM/SPF alignment tables (issue #169), one level deeper: policies is an object array and each policy's failure_details is an object array inside it, so stacked terms aggregations on their subfields fabricate rows. Documents now also carry policies_combined ("domain / type" per policy) and failure_details_combined ("domain / type / result / sending mta / receiving ip / mx" per failure detail), composed at save time with the same "none" fallbacks as the aggregate fields. migrate_indexes() gains smtp_tls_indexes and backfills old documents with the same guarded, non-blocking update_by_query pattern; cli.py wires the index name in on both backends, and the manual _update_by_query command is documented. Also fixes two adjacent dead fields: add_failure_details stored additional_information_uri under the wrong constructor kwarg (additional_information), and receiving_mx_hostname had no declaration despite always being stored. Verified live on ES 8.19 and OpenSearch 3: a two-policy repro report yields exactly 2 policy rows and 2 failure-detail rows via the combined fields where the old stacked aggregations return 4 of each; the startup backfill converted the 4 pre-existing sample documents on both engines with zero recompute mismatches. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Rework the SMTP TLS dashboards onto the combined fields Kibana/OSD: "SMTP TLS domains" replaces its stacked policy_domain × policy_type terms with one terms agg on policies_combined.keyword; "SMTP TLS failure details" replaces six stacked terms spanning both array levels with one on failure_details_combined.keyword; the smtp_tls* index-pattern field cache gains the new fields. The "reporting organizations" table only buckets on doc-level org_name and needed no change. Splunk: the base search now expands policies at the JSON level (spath + mvexpand) so policy fields are scalars per event, and the failure details panel expands the second level the same way — sums are the detail's own failed_session_count, correctly paired. Verified via the search REST API: a two-policy repro returns exactly one row per real failure detail with per-detail counts. kibana.md documents the per-policy/per-detail row semantics and the honest caveat that session-count sums remain per report document. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Persist additional_info_uri from parsed SMTP TLS failure details Copilot caught that the savers read additional_information_uri from the parsed failure-detail dict, but the parser's key is additional_info_uri (SMTPTLSFailureDetailsOptional in types.py, set in parse_smtp_tls_report_json), so the URI was never persisted — the read-side half of the dead-field bug whose write-side half (wrong constructor kwarg) was fixed earlier. Read the parser's key first, keeping the long-form key as a fallback for dicts built by other callers. Regression test proven to fail on the unfixed savers. Also restructured the expected combined-string test values into named locals so no implicit string concatenation sits inside a list literal. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Say "inner doc", not "nested doc", in the singular-key test docstrings Final review sweep: in this codebase "nested" is reserved for the Elasticsearch mapping type, and these InnerDoc-serialization docstrings used it colloquially — the same dual-use-term trap documented in AGENTS.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
1393 lines
60 KiB
Python
1393 lines
60 KiB
Python
"""Tests for parsedmarc.opensearch
|
|
|
|
Mocks at the opensearch-dsl SDK boundary (connections.create_connection,
|
|
Index, Search, Document.save) so the tests verify the parsedmarc-side
|
|
transformation logic — document construction, index naming, deduplication
|
|
queries, error wrapping — without needing a running OpenSearch cluster.
|
|
"""
|
|
|
|
import time
|
|
import unittest
|
|
from unittest.mock import MagicMock, call, patch
|
|
|
|
import parsedmarc.opensearch as opensearch_module
|
|
from parsedmarc import InvalidFailureReport
|
|
from parsedmarc.opensearch import (
|
|
AlreadySaved,
|
|
OpenSearchError,
|
|
create_indexes,
|
|
migrate_indexes,
|
|
save_aggregate_report_to_opensearch,
|
|
save_failure_report_to_opensearch,
|
|
save_smtp_tls_report_to_opensearch,
|
|
set_hosts,
|
|
)
|
|
from tests.tzutil import force_tz
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Sample report fixtures
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def _aggregate_report(**overrides):
|
|
base = {
|
|
"xml_schema": "draft",
|
|
"xml_namespace": None,
|
|
"report_metadata": {
|
|
"org_name": "TestOrg",
|
|
"org_email": "dmarc@example.com",
|
|
"org_extra_contact_info": None,
|
|
"report_id": "agg-1",
|
|
"begin_date": "2024-01-15 00:00:00",
|
|
"end_date": "2024-01-16 00:00:00",
|
|
"timespan_requires_normalization": False,
|
|
"original_timespan_seconds": 86400,
|
|
"errors": [],
|
|
"generator": "TestGen/1.0",
|
|
},
|
|
"policy_published": {
|
|
"domain": "example.com",
|
|
"adkim": "r",
|
|
"aspf": "r",
|
|
"p": "none",
|
|
"sp": "none",
|
|
"pct": None,
|
|
"fo": None,
|
|
"np": "reject",
|
|
"testing": "n",
|
|
"discovery_method": "treewalk",
|
|
},
|
|
"records": [
|
|
{
|
|
"interval_begin": "2024-01-15 00:00:00",
|
|
"interval_end": "2024-01-16 00:00:00",
|
|
"normalized_timespan": False,
|
|
"source": {
|
|
"ip_address": "192.0.2.1",
|
|
"country": "US",
|
|
"reverse_dns": None,
|
|
"base_domain": None,
|
|
"name": None,
|
|
"type": None,
|
|
"asn": 64496,
|
|
"as_name": "Example AS",
|
|
"as_domain": "example.net",
|
|
},
|
|
"count": 4,
|
|
"alignment": {"spf": True, "dkim": True, "dmarc": True},
|
|
"policy_evaluated": {
|
|
"disposition": "none",
|
|
"dkim": "pass",
|
|
"spf": "pass",
|
|
"policy_override_reasons": [
|
|
{"type": "local_policy", "comment": "approved"}
|
|
],
|
|
},
|
|
"identifiers": {
|
|
"header_from": "example.com",
|
|
"envelope_from": "example.com",
|
|
"envelope_to": "rcpt@example.com",
|
|
},
|
|
"auth_results": {
|
|
"dkim": [
|
|
{
|
|
"domain": "example.com",
|
|
"selector": "s",
|
|
"result": "pass",
|
|
"human_result": None,
|
|
}
|
|
],
|
|
"spf": [
|
|
{
|
|
"domain": "example.com",
|
|
"scope": "mfrom",
|
|
"result": "pass",
|
|
"human_result": None,
|
|
}
|
|
],
|
|
},
|
|
}
|
|
],
|
|
}
|
|
base.update(overrides)
|
|
return base
|
|
|
|
|
|
def _failure_report(**overrides):
|
|
base = {
|
|
"feedback_type": "auth-failure",
|
|
"user_agent": "test/1.0",
|
|
"version": "1",
|
|
"original_envelope_id": None,
|
|
"original_mail_from": "x@example.com",
|
|
"original_rcpt_to": None,
|
|
"arrival_date": "Thu, 1 Jan 2024 00:00:00 +0000",
|
|
"arrival_date_utc": "2024-01-01 00:00:00",
|
|
"authentication_results": None,
|
|
"delivery_result": "other",
|
|
"auth_failure": ["dmarc"],
|
|
"authentication_mechanisms": [],
|
|
"dkim_domain": None,
|
|
"reported_domain": "example.com",
|
|
"sample_headers_only": True,
|
|
"source": {
|
|
"ip_address": "192.0.2.5",
|
|
"country": "US",
|
|
"reverse_dns": None,
|
|
"base_domain": None,
|
|
"name": None,
|
|
"type": None,
|
|
"asn": 64496,
|
|
"as_name": "Example AS",
|
|
"as_domain": "example.net",
|
|
},
|
|
"sample": "raw",
|
|
"parsed_sample": {
|
|
"headers": {
|
|
# mailparser emits headers as [[display_name, address]]
|
|
# lists; an empty display becomes [["", address]].
|
|
"From": [["Sender Name", "sender@example.com"]],
|
|
"To": [["", "rcpt@example.com"]],
|
|
"Subject": "Test",
|
|
},
|
|
"subject": "Test",
|
|
"filename_safe_subject": "Test",
|
|
"body": "body",
|
|
"date": "Thu, 1 Jan 2024 00:00:00 +0000",
|
|
"to": [{"display_name": None, "address": "rcpt@example.com"}],
|
|
"reply_to": [],
|
|
"cc": [],
|
|
"bcc": [],
|
|
"attachments": [],
|
|
},
|
|
}
|
|
base.update(overrides)
|
|
return base
|
|
|
|
|
|
def _smtp_tls_report(**overrides):
|
|
base = {
|
|
"organization_name": "TestOrg",
|
|
"begin_date": "2024-02-03T00:00:00Z",
|
|
"end_date": "2024-02-04T00:00:00Z",
|
|
"contact_info": "tls@example.com",
|
|
"report_id": "tls-1",
|
|
"policies": [
|
|
{
|
|
"policy_domain": "example.com",
|
|
"policy_type": "sts",
|
|
"successful_session_count": 100,
|
|
"failed_session_count": 1,
|
|
"policy_strings": ["version: STSv1"],
|
|
"mx_host_patterns": ["*.example.com"],
|
|
"failure_details": [
|
|
{
|
|
"result_type": "certificate-expired",
|
|
"failed_session_count": 1,
|
|
"receiving_mx_hostname": "mx.example.com",
|
|
"sending_mta_ip": "10.0.0.1",
|
|
}
|
|
],
|
|
}
|
|
],
|
|
}
|
|
base.update(overrides)
|
|
return base
|
|
|
|
|
|
def _empty_search():
|
|
"""A Search() mock whose .execute() returns an empty hit list."""
|
|
search = MagicMock()
|
|
search.execute.return_value = []
|
|
return search
|
|
|
|
|
|
def _populated_search():
|
|
"""A Search() mock whose .execute() returns a non-empty hit list."""
|
|
search = MagicMock()
|
|
search.execute.return_value = [MagicMock()]
|
|
return search
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# set_hosts: connection-parameter assembly
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestSetHosts(unittest.TestCase):
|
|
"""Verify the conn_params dict handed to opensearch-dsl
|
|
matches each documented option. Each branch corresponds to a
|
|
real-world deployment shape (TLS, basic auth, API key, custom CA)."""
|
|
|
|
def test_single_host_string_normalized_to_list(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("https://es:9200")
|
|
kwargs = mock_conn.call_args.kwargs
|
|
self.assertEqual(kwargs["hosts"], ["https://es:9200"])
|
|
|
|
def test_host_list_preserved(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts(["es1:9200", "es2:9200"])
|
|
kwargs = mock_conn.call_args.kwargs
|
|
self.assertEqual(kwargs["hosts"], ["es1:9200", "es2:9200"])
|
|
|
|
def test_timeout_default_60s(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200")
|
|
self.assertEqual(mock_conn.call_args.kwargs["timeout"], 60.0)
|
|
|
|
def test_timeout_custom(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200", timeout=30.0)
|
|
self.assertEqual(mock_conn.call_args.kwargs["timeout"], 30.0)
|
|
|
|
def test_use_ssl_enables_verify_by_default(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200", use_ssl=True)
|
|
kwargs = mock_conn.call_args.kwargs
|
|
self.assertEqual(kwargs["use_ssl"], True)
|
|
self.assertEqual(kwargs["verify_certs"], True)
|
|
self.assertNotIn("ca_certs", kwargs)
|
|
|
|
def test_use_ssl_with_custom_ca(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200", use_ssl=True, ssl_cert_path="/etc/ca.pem")
|
|
kwargs = mock_conn.call_args.kwargs
|
|
self.assertEqual(kwargs["ca_certs"], "/etc/ca.pem")
|
|
|
|
def test_skip_certificate_verification_sets_verify_false(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200", use_ssl=True, skip_certificate_verification=True)
|
|
self.assertEqual(mock_conn.call_args.kwargs["verify_certs"], False)
|
|
|
|
def test_username_password_sets_http_auth(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200", username="u", password="p")
|
|
self.assertEqual(mock_conn.call_args.kwargs["http_auth"], ("u", "p"))
|
|
|
|
def test_username_without_password_not_set(self):
|
|
"""Half-configured auth is suspicious enough not to send."""
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200", username="u")
|
|
self.assertNotIn("http_auth", mock_conn.call_args.kwargs)
|
|
|
|
def test_api_key_set(self):
|
|
with patch("parsedmarc.opensearch.connections.create_connection") as mock_conn:
|
|
set_hosts("es:9200", api_key="base64key==")
|
|
self.assertEqual(mock_conn.call_args.kwargs["api_key"], "base64key==")
|
|
|
|
def test_awssigv4_requires_aws_region(self):
|
|
"""SigV4 needs an AWS region to sign requests; missing it
|
|
must fail loudly, not silently fall back to unsigned auth."""
|
|
with self.assertRaises(OpenSearchError) as ctx:
|
|
set_hosts("es.amazonaws.com:443", auth_type="awssigv4")
|
|
self.assertIn("aws_region", str(ctx.exception))
|
|
|
|
def test_awssigv4_uses_boto3_credentials_and_signer(self):
|
|
"""SigV4 path resolves AWS credentials via boto3 and wires
|
|
an AWSV4SignerAuth into the connection params, plus the
|
|
RequestsHttpConnection class required by the signer."""
|
|
with (
|
|
patch("parsedmarc.opensearch.boto3.Session") as mock_session,
|
|
patch("parsedmarc.opensearch.AWSV4SignerAuth") as mock_signer,
|
|
patch("parsedmarc.opensearch.connections.create_connection") as mock_conn,
|
|
):
|
|
mock_session.return_value.get_credentials.return_value = MagicMock()
|
|
set_hosts(
|
|
"es.amazonaws.com:443",
|
|
auth_type="awssigv4",
|
|
aws_region="us-west-2",
|
|
)
|
|
kwargs = mock_conn.call_args.kwargs
|
|
self.assertIs(kwargs["http_auth"], mock_signer.return_value)
|
|
mock_signer.assert_called_once()
|
|
# connection_class must be set so opensearch-py uses the
|
|
# requests-based transport AWSV4SignerAuth requires.
|
|
self.assertIn("connection_class", kwargs)
|
|
|
|
def test_awssigv4_no_credentials_raises(self):
|
|
"""If boto3 can't find credentials, fail with a clear error
|
|
rather than letting OpenSearch raise an opaque auth error later."""
|
|
with patch("parsedmarc.opensearch.boto3.Session") as mock_session:
|
|
mock_session.return_value.get_credentials.return_value = None
|
|
with self.assertRaises(OpenSearchError) as ctx:
|
|
set_hosts(
|
|
"es.amazonaws.com:443",
|
|
auth_type="awssigv4",
|
|
aws_region="us-west-2",
|
|
)
|
|
self.assertIn("credentials", str(ctx.exception).lower())
|
|
|
|
def test_unsupported_auth_type_raises(self):
|
|
with self.assertRaises(OpenSearchError) as ctx:
|
|
set_hosts("es:9200", auth_type="kerberos")
|
|
self.assertIn("Unsupported", str(ctx.exception))
|
|
self.assertIn("kerberos", str(ctx.exception))
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# create_indexes
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestCreateIndexes(unittest.TestCase):
|
|
def test_creates_missing_index_with_default_settings(self):
|
|
with patch("parsedmarc.opensearch.Index") as mock_index_cls:
|
|
mock_index = mock_index_cls.return_value
|
|
mock_index.exists.return_value = False
|
|
create_indexes(["dmarc_aggregate-2024-01-15"])
|
|
mock_index.settings.assert_called_once_with(
|
|
number_of_shards=1, number_of_replicas=0
|
|
)
|
|
mock_index.create.assert_called_once()
|
|
|
|
def test_creates_with_custom_settings(self):
|
|
with patch("parsedmarc.opensearch.Index") as mock_index_cls:
|
|
mock_index = mock_index_cls.return_value
|
|
mock_index.exists.return_value = False
|
|
create_indexes(
|
|
["idx"], settings={"number_of_shards": 3, "refresh_interval": "5s"}
|
|
)
|
|
mock_index.settings.assert_called_once_with(
|
|
number_of_shards=3, refresh_interval="5s"
|
|
)
|
|
|
|
def test_skips_existing_index(self):
|
|
with patch("parsedmarc.opensearch.Index") as mock_index_cls:
|
|
mock_index = mock_index_cls.return_value
|
|
mock_index.exists.return_value = True
|
|
create_indexes(["idx"])
|
|
mock_index.create.assert_not_called()
|
|
|
|
def test_wraps_sdk_error(self):
|
|
with patch("parsedmarc.opensearch.Index") as mock_index_cls:
|
|
mock_index_cls.return_value.exists.side_effect = RuntimeError(
|
|
"cluster down"
|
|
)
|
|
with self.assertRaises(OpenSearchError) as ctx:
|
|
create_indexes(["idx"])
|
|
self.assertIn("cluster down", str(ctx.exception))
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# migrate_indexes
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestMigrateIndexes(unittest.TestCase):
|
|
"""migrate_indexes backfills dkim_results_combined/spf_results_combined
|
|
(issue #169) on pre-existing aggregate documents as a non-blocking
|
|
background task. It is guarded by a cheap count() query so repeated
|
|
startups against an already-backfilled index are a no-op, and any SDK
|
|
error is caught and logged rather than raised, so it never blocks
|
|
parsedmarc startup."""
|
|
|
|
def test_backfill_submitted_when_old_docs_exist(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
# The legacy fo migration that runs first sees no base index.
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
mock_client = MagicMock()
|
|
mock_client.count.return_value = {"count": 42}
|
|
mock_get_conn.return_value = mock_client
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate"])
|
|
|
|
mock_client.update_by_query.assert_called_once()
|
|
kwargs = mock_client.update_by_query.call_args.kwargs
|
|
self.assertEqual(kwargs["index"], "dmarc_aggregate*")
|
|
self.assertEqual(kwargs["conflicts"], "proceed")
|
|
self.assertFalse(kwargs["wait_for_completion"])
|
|
self.assertEqual(
|
|
kwargs["body"]["query"], opensearch_module._COMBINED_BACKFILL_QUERY
|
|
)
|
|
script_source = kwargs["body"]["script"]["source"]
|
|
self.assertIn("ctx._source.dkim_results_combined", script_source)
|
|
self.assertIn("ctx._source.spf_results_combined", script_source)
|
|
|
|
# The count() guard query also targets the date-suffixed pattern.
|
|
count_kwargs = mock_client.count.call_args.kwargs
|
|
self.assertEqual(count_kwargs["index"], "dmarc_aggregate*")
|
|
self.assertEqual(
|
|
count_kwargs["body"]["query"], opensearch_module._COMBINED_BACKFILL_QUERY
|
|
)
|
|
|
|
def test_backfill_skipped_when_no_old_docs(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
mock_client = MagicMock()
|
|
mock_client.count.return_value = {"count": 0}
|
|
mock_get_conn.return_value = mock_client
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate"])
|
|
|
|
mock_client.update_by_query.assert_not_called()
|
|
|
|
def test_backfill_skipped_when_no_aggregate_indexes(self):
|
|
with patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn:
|
|
migrate_indexes()
|
|
migrate_indexes(aggregate_indexes=None)
|
|
|
|
mock_get_conn.assert_not_called()
|
|
|
|
def test_backfill_failure_does_not_raise(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
mock_client = MagicMock()
|
|
mock_client.count.side_effect = RuntimeError("cluster unreachable")
|
|
mock_get_conn.return_value = mock_client
|
|
with self.assertLogs("parsedmarc.log", level="WARNING") as cm:
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate"])
|
|
|
|
self.assertTrue(any("cluster unreachable" in msg for msg in cm.output))
|
|
mock_client.update_by_query.assert_not_called()
|
|
|
|
def test_get_connection_failure_does_not_raise(self):
|
|
"""connections.get_connection() itself sits outside the per-index
|
|
try/except for the combined-field backfill; if it raises (e.g. no
|
|
OpenSearch connection has been configured yet), migrate_indexes must
|
|
still not propagate the exception, per its docstring's promise that
|
|
any cluster error is caught and logged."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
# The legacy fo migration that runs first sees no base index.
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
mock_get_conn.side_effect = RuntimeError("no connection")
|
|
with self.assertLogs("parsedmarc.log", level="WARNING") as cm:
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate"])
|
|
|
|
self.assertTrue(
|
|
any("Skipping the dkim_results_combined" in msg for msg in cm.output)
|
|
)
|
|
self.assertTrue(any("no connection" in msg for msg in cm.output))
|
|
|
|
def test_smtp_tls_backfill_submitted_when_old_docs_exist(self):
|
|
"""SMTP TLS analogue of test_backfill_submitted_when_old_docs_exist:
|
|
policies_combined/failure_details_combined backfill (also issue
|
|
#169) is submitted with its own guard query and painless script."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
# The legacy fo migration that runs first sees no base index.
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
mock_client = MagicMock()
|
|
mock_client.count.return_value = {"count": 7}
|
|
mock_get_conn.return_value = mock_client
|
|
migrate_indexes(smtp_tls_indexes=["smtp_tls"])
|
|
|
|
mock_client.update_by_query.assert_called_once()
|
|
kwargs = mock_client.update_by_query.call_args.kwargs
|
|
self.assertEqual(kwargs["index"], "smtp_tls*")
|
|
self.assertEqual(kwargs["conflicts"], "proceed")
|
|
self.assertFalse(kwargs["wait_for_completion"])
|
|
self.assertEqual(
|
|
kwargs["body"]["query"], opensearch_module._SMTP_TLS_COMBINED_BACKFILL_QUERY
|
|
)
|
|
script_source = kwargs["body"]["script"]["source"]
|
|
self.assertIn("ctx._source.policies_combined", script_source)
|
|
self.assertIn("ctx._source.failure_details_combined", script_source)
|
|
|
|
count_kwargs = mock_client.count.call_args.kwargs
|
|
self.assertEqual(count_kwargs["index"], "smtp_tls*")
|
|
self.assertEqual(
|
|
count_kwargs["body"]["query"],
|
|
opensearch_module._SMTP_TLS_COMBINED_BACKFILL_QUERY,
|
|
)
|
|
|
|
def test_smtp_tls_backfill_skipped_when_no_old_docs(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
mock_client = MagicMock()
|
|
mock_client.count.return_value = {"count": 0}
|
|
mock_get_conn.return_value = mock_client
|
|
migrate_indexes(smtp_tls_indexes=["smtp_tls"])
|
|
|
|
mock_client.update_by_query.assert_not_called()
|
|
|
|
def test_smtp_tls_backfill_skipped_when_no_smtp_tls_indexes_or_aggregate(self):
|
|
with patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn:
|
|
migrate_indexes()
|
|
migrate_indexes(smtp_tls_indexes=None)
|
|
|
|
mock_get_conn.assert_not_called()
|
|
|
|
def test_smtp_tls_backfill_failure_does_not_raise(self):
|
|
"""SMTP TLS analogue of test_backfill_failure_does_not_raise: an
|
|
error from the cluster during the smtp_tls_indexes loop is caught
|
|
and logged rather than raised."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
mock_client = MagicMock()
|
|
mock_client.count.side_effect = RuntimeError("cluster unreachable")
|
|
mock_get_conn.return_value = mock_client
|
|
with self.assertLogs("parsedmarc.log", level="WARNING") as cm:
|
|
migrate_indexes(smtp_tls_indexes=["smtp_tls"])
|
|
|
|
self.assertTrue(any("cluster unreachable" in msg for msg in cm.output))
|
|
mock_client.update_by_query.assert_not_called()
|
|
|
|
|
|
class TestMigrateIndexesFoMigration(unittest.TestCase):
|
|
"""The legacy `published_policy.fo` field was mapped as `long` in
|
|
older indexes. migrate_indexes detects that and rebuilds the index
|
|
with the text/keyword shape. The branch is gnarly; a regression
|
|
would silently leave old data un-migrated. Each test stubs the
|
|
combined-field backfill that now runs afterwards in the same call
|
|
(count 0 → no-op)."""
|
|
|
|
@staticmethod
|
|
def _noop_backfill_client():
|
|
client = MagicMock()
|
|
client.count.return_value = {"count": 0}
|
|
return client
|
|
|
|
def test_skips_non_existent_index(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
mock_get_conn.return_value = self._noop_backfill_client()
|
|
mock_index_cls.return_value.exists.return_value = False
|
|
migrate_indexes(aggregate_indexes=["missing"])
|
|
# exists() returned False — no field_mapping fetch.
|
|
mock_index_cls.return_value.get_field_mapping.assert_not_called()
|
|
|
|
def test_skips_when_doc_mapping_absent(self):
|
|
"""An index that has 'fo' but not under the 'doc' type
|
|
(e.g., empty index with default mapping) is left alone."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
patch("parsedmarc.opensearch.reindex") as mock_reindex,
|
|
):
|
|
mock_get_conn.return_value = self._noop_backfill_client()
|
|
idx = mock_index_cls.return_value
|
|
idx.exists.return_value = True
|
|
idx.get_field_mapping.return_value = {"some_key": {"mappings": {}}}
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate-2023-01-01"])
|
|
mock_reindex.assert_not_called()
|
|
|
|
def test_migrates_when_fo_is_long(self):
|
|
"""The actual migration path: when fo is mapped as 'long',
|
|
a v2 index is created with the corrected mapping, data is
|
|
reindexed, and the old index is deleted."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.reindex") as mock_reindex,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
mock_client = self._noop_backfill_client()
|
|
mock_get_conn.return_value = mock_client
|
|
idx = mock_index_cls.return_value
|
|
idx.exists.return_value = True
|
|
idx.get_field_mapping.return_value = {
|
|
"dmarc_aggregate-2023-01-01": {
|
|
"mappings": {
|
|
"doc": {
|
|
"published_policy.fo": {"mapping": {"fo": {"type": "long"}}}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate-2023-01-01"])
|
|
# reindex called from old → new (v2) index, with the client from
|
|
# connections.get_connection().
|
|
mock_reindex.assert_called_once()
|
|
self.assertIs(mock_reindex.call_args.args[0], mock_client)
|
|
|
|
def test_skips_when_fo_already_text(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
patch("parsedmarc.opensearch.reindex") as mock_reindex,
|
|
):
|
|
mock_get_conn.return_value = self._noop_backfill_client()
|
|
idx = mock_index_cls.return_value
|
|
idx.exists.return_value = True
|
|
idx.get_field_mapping.return_value = {
|
|
"dmarc_aggregate-2024-01-01": {
|
|
"mappings": {
|
|
"doc": {
|
|
"published_policy.fo": {"mapping": {"fo": {"type": "text"}}}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate-2024-01-01"])
|
|
mock_reindex.assert_not_called()
|
|
|
|
def test_index_exists_failure_does_not_raise(self):
|
|
"""A cluster error inside the per-index fo-migration loop (e.g.
|
|
Index(...).exists() raising because the cluster is unreachable)
|
|
must not abort startup: it is caught, logged, and the loop moves
|
|
on to the combined-field backfill, which is exercised here with
|
|
its own connection failure so both warnings are asserted."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch.connections.get_connection") as mock_get_conn,
|
|
):
|
|
mock_index_cls.return_value.exists.side_effect = ConnectionError(
|
|
"cluster unreachable"
|
|
)
|
|
mock_get_conn.side_effect = RuntimeError("no connection")
|
|
with self.assertLogs("parsedmarc.log", level="WARNING") as cm:
|
|
migrate_indexes(aggregate_indexes=["dmarc_aggregate"])
|
|
|
|
self.assertTrue(
|
|
any(
|
|
"legacy published_policy.fo migration" in msg
|
|
and "cluster unreachable" in msg
|
|
for msg in cm.output
|
|
)
|
|
)
|
|
self.assertTrue(
|
|
any("Skipping the dkim_results_combined" in msg for msg in cm.output)
|
|
)
|
|
self.assertTrue(any("no connection" in msg for msg in cm.output))
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# save_aggregate_report_to_opensearch
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestSaveAggregateReport(unittest.TestCase):
|
|
"""The aggregate-report save fans out across multiple SDK calls:
|
|
Search (for dedup), Index.create (for the daily/monthly index),
|
|
Document.save. Each test patches the boundary it needs and
|
|
leaves the rest alone."""
|
|
|
|
def _patches(self, search_factory=_empty_search):
|
|
return [
|
|
patch("parsedmarc.opensearch.Search", return_value=search_factory()),
|
|
patch(
|
|
"parsedmarc.opensearch.Index",
|
|
return_value=MagicMock(exists=MagicMock(return_value=True)),
|
|
),
|
|
patch.object(opensearch_module._AggregateReportDoc, "save"),
|
|
]
|
|
|
|
def test_save_emits_one_document_per_record(self):
|
|
report = _aggregate_report()
|
|
report["records"].append(report["records"][0].copy())
|
|
patches = self._patches()
|
|
with patches[0], patches[1], patches[2] as mock_save:
|
|
save_aggregate_report_to_opensearch(report)
|
|
# Two records → two saves.
|
|
self.assertEqual(mock_save.call_count, 2)
|
|
|
|
def test_already_saved_raises_when_search_returns_hit(self):
|
|
"""The dedup query is the only thing preventing
|
|
double-indexing on re-run. A regression would silently
|
|
re-save reports, inflating Kibana counts."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_populated_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._AggregateReportDoc, "save") as mock_save,
|
|
):
|
|
with self.assertRaises(AlreadySaved):
|
|
save_aggregate_report_to_opensearch(_aggregate_report())
|
|
mock_save.assert_not_called()
|
|
|
|
def test_search_exception_wraps_to_opensearch_error(self):
|
|
bad_search = MagicMock()
|
|
bad_search.execute.side_effect = RuntimeError("network")
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=bad_search),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
):
|
|
with self.assertRaises(OpenSearchError) as ctx:
|
|
save_aggregate_report_to_opensearch(_aggregate_report())
|
|
self.assertIn("network", str(ctx.exception))
|
|
|
|
def test_save_exception_wraps_to_opensearch_error(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(
|
|
opensearch_module._AggregateReportDoc,
|
|
"save",
|
|
side_effect=RuntimeError("disk"),
|
|
),
|
|
):
|
|
with self.assertRaises(OpenSearchError) as ctx:
|
|
save_aggregate_report_to_opensearch(_aggregate_report())
|
|
self.assertIn("disk", str(ctx.exception))
|
|
|
|
def test_index_name_uses_daily_format_by_default(self):
|
|
"""Index naming: dmarc_aggregate-YYYY-MM-DD by default."""
|
|
index_calls = []
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch.object(opensearch_module._AggregateReportDoc, "save"),
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = True
|
|
save_aggregate_report_to_opensearch(_aggregate_report())
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("dmarc_aggregate-2024-01-15", index_calls)
|
|
|
|
def test_index_name_uses_monthly_format_when_flag_set(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch.object(opensearch_module._AggregateReportDoc, "save"),
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = True
|
|
save_aggregate_report_to_opensearch(
|
|
_aggregate_report(), monthly_indexes=True
|
|
)
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("dmarc_aggregate-2024-01", index_calls)
|
|
|
|
def test_index_name_honours_suffix_and_prefix(self):
|
|
"""Prefix/suffix support multi-tenant setups where one ES
|
|
cluster serves several DMARC owners."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch.object(opensearch_module._AggregateReportDoc, "save"),
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = True
|
|
save_aggregate_report_to_opensearch(
|
|
_aggregate_report(),
|
|
index_suffix="tenant_a",
|
|
index_prefix="customer1_",
|
|
)
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("customer1_dmarc_aggregate_tenant_a-2024-01-15", index_calls)
|
|
|
|
def test_dedup_search_pattern_uses_suffix_wildcard(self):
|
|
"""Existing-report search uses '*' so it matches both
|
|
daily and monthly index buckets."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Search") as mock_search_cls,
|
|
patch(
|
|
"parsedmarc.opensearch.Index",
|
|
return_value=MagicMock(exists=MagicMock(return_value=True)),
|
|
),
|
|
patch.object(opensearch_module._AggregateReportDoc, "save"),
|
|
):
|
|
mock_search_cls.return_value.execute.return_value = []
|
|
save_aggregate_report_to_opensearch(
|
|
_aggregate_report(), index_suffix="tenant_a", index_prefix="cust_"
|
|
)
|
|
# Search index pattern wraps prefix+name+suffix with trailing wildcard.
|
|
search_index = mock_search_cls.call_args.kwargs["index"]
|
|
self.assertIn("cust_dmarc_aggregate_tenant_a*", search_index)
|
|
|
|
@unittest.skipUnless(hasattr(time, "tzset"), "requires POSIX time.tzset()")
|
|
def test_interval_dates_are_utc_regardless_of_host_timezone(self):
|
|
"""interval_begin/interval_end are UTC wall-clock strings (already
|
|
converted to UTC at parse time in __init__.py); the index-date
|
|
bucketing and stored date_begin/date_end must use their true UTC
|
|
epoch on any host. Regression test for
|
|
https://github.com/domainaware/parsedmarc/issues/819: the naive
|
|
parse used to shift the stored epoch (and therefore the index
|
|
date) by the host's UTC offset."""
|
|
force_tz(self)
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch("parsedmarc.opensearch._AggregateReportDoc") as mock_doc_cls,
|
|
):
|
|
mock_index_cls.return_value.exists.return_value = True
|
|
save_aggregate_report_to_opensearch(_aggregate_report())
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("dmarc_aggregate-2024-01-15", index_calls)
|
|
# Fixture begin_date/interval_begin is 2024-01-15 00:00:00 UTC.
|
|
self.assertEqual(
|
|
mock_doc_cls.call_args.kwargs["date_begin"].timestamp(), 1705276800
|
|
)
|
|
|
|
def test_save_populates_combined_dkim_and_spf_fields(self):
|
|
"""Regression guard for issue #169: two DKIM signatures on one
|
|
record must yield exactly two combined entries, not a 4-way
|
|
cross-product. autospec=True is required on the save patch so
|
|
mock_save.call_args captures the doc instance as ``self``."""
|
|
report = _aggregate_report()
|
|
report["records"][0]["auth_results"] = {
|
|
"dkim": [
|
|
{
|
|
"domain": "example.net",
|
|
"selector": "net1",
|
|
"result": "fail",
|
|
"human_result": None,
|
|
},
|
|
{
|
|
"domain": "example.org",
|
|
"selector": "org1",
|
|
"result": "pass",
|
|
"human_result": None,
|
|
},
|
|
],
|
|
"spf": [
|
|
{
|
|
"domain": "example.org",
|
|
"scope": "mfrom",
|
|
"result": "pass",
|
|
"human_result": None,
|
|
},
|
|
],
|
|
}
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch(
|
|
"parsedmarc.opensearch.Index",
|
|
return_value=MagicMock(exists=MagicMock(return_value=True)),
|
|
),
|
|
patch.object(
|
|
opensearch_module._AggregateReportDoc, "save", autospec=True
|
|
) as mock_save,
|
|
):
|
|
save_aggregate_report_to_opensearch(report)
|
|
doc = mock_save.call_args[0][0]
|
|
self.assertEqual(
|
|
list(doc.dkim_results_combined),
|
|
["net1 / example.net / fail", "org1 / example.org / pass"],
|
|
)
|
|
self.assertEqual(list(doc.spf_results_combined), ["mfrom / example.org / pass"])
|
|
|
|
|
|
class TestAggregateDocPassedDmarc(unittest.TestCase):
|
|
"""The _AggregateReportDoc.save() override derives passed_dmarc — the
|
|
field dashboards filter on for DMARC pass/fail — from SPF/DKIM
|
|
alignment. The SDK parent (opensearchpy.Document.save) is mocked so
|
|
no cluster is needed."""
|
|
|
|
def test_passed_dmarc_derived_from_alignment(self):
|
|
cases = [
|
|
(True, False, True),
|
|
(False, True, True),
|
|
(True, True, True),
|
|
(False, False, False),
|
|
]
|
|
for spf_aligned, dkim_aligned, expected in cases:
|
|
with self.subTest(spf=spf_aligned, dkim=dkim_aligned):
|
|
with patch.object(
|
|
opensearch_module.Document, "save", return_value=None
|
|
) as mock_super_save:
|
|
doc = opensearch_module._AggregateReportDoc(
|
|
spf_aligned=spf_aligned, dkim_aligned=dkim_aligned
|
|
)
|
|
doc.save()
|
|
mock_super_save.assert_called_once()
|
|
self.assertEqual(bool(doc.passed_dmarc), expected)
|
|
|
|
|
|
class TestAggregateDocCombinedResults(unittest.TestCase):
|
|
"""add_dkim_result/add_spf_result never touch the network, so these
|
|
construct _AggregateReportDoc directly rather than going through the
|
|
save_* entry point."""
|
|
|
|
def test_add_dkim_result_appends_combined_string(self):
|
|
"""Regression guard for issue #169: dkim_results/spf_results are
|
|
arrays of objects that the engine dynamic-maps as plain ``object``
|
|
(not ``nested``) and flattens, so Kibana/Grafana tables cannot
|
|
terms-aggregate their subfields without producing a cross-product
|
|
of selector/domain/result values. The composed
|
|
"selector / domain / result" string preserves the per-signature
|
|
pairing that the flattened array loses."""
|
|
doc = opensearch_module._AggregateReportDoc()
|
|
doc.add_dkim_result(
|
|
domain="example.net", selector="net1", result="fail", human_result=None
|
|
)
|
|
doc.add_dkim_result(
|
|
domain="example.org", selector="org1", result="pass", human_result=None
|
|
)
|
|
expected = ["net1 / example.net / fail", "org1 / example.org / pass"]
|
|
# dkim_results_combined is declared as Text(multi=True, ...); the SDK
|
|
# stub types the class attribute as Text (no Iterable protocol),
|
|
# even though the runtime value is an AttrList once multi=True is
|
|
# set.
|
|
self.assertEqual(list(doc.dkim_results_combined), expected) # pyright: ignore[reportArgumentType]
|
|
self.assertEqual(doc.to_dict()["dkim_results_combined"], expected)
|
|
|
|
def test_add_spf_result_appends_combined_string(self):
|
|
doc = opensearch_module._AggregateReportDoc()
|
|
doc.add_spf_result(
|
|
domain="example.org", scope="mfrom", result="pass", human_result=None
|
|
)
|
|
expected = ["mfrom / example.org / pass"]
|
|
self.assertEqual(list(doc.spf_results_combined), expected) # pyright: ignore[reportArgumentType]
|
|
self.assertEqual(doc.to_dict()["spf_results_combined"], expected)
|
|
|
|
def test_spf_result_serializes_under_singular_result_key(self):
|
|
"""The _SPFResult class previously declared a dead ``results``
|
|
(plural) field while the save path wrote ``result``; verify the
|
|
serialized inner doc actually uses the singular key."""
|
|
doc = opensearch_module._AggregateReportDoc()
|
|
doc.add_spf_result(
|
|
domain="example.org", scope="mfrom", result="pass", human_result=None
|
|
)
|
|
d = doc.to_dict()["spf_results"][0]
|
|
self.assertEqual(d["result"], "pass")
|
|
self.assertNotIn("results", d)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# save_failure_report_to_opensearch
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestSaveFailureReport(unittest.TestCase):
|
|
def test_save_emits_one_document(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save") as mock_save,
|
|
):
|
|
save_failure_report_to_opensearch(_failure_report())
|
|
mock_save.assert_called_once()
|
|
|
|
def test_already_saved_raises_on_dedup_hit(self):
|
|
"""Failure-report dedup uses arrival_date + From/To/Subject
|
|
from the parsed sample. A hit means we've already indexed
|
|
this exact failure sample."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_populated_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save") as mock_save,
|
|
):
|
|
with self.assertRaises(AlreadySaved):
|
|
save_failure_report_to_opensearch(_failure_report())
|
|
mock_save.assert_not_called()
|
|
|
|
def test_save_exception_wraps_to_opensearch_error(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(
|
|
opensearch_module._FailureReportDoc,
|
|
"save",
|
|
side_effect=RuntimeError("disk"),
|
|
),
|
|
):
|
|
with self.assertRaises(OpenSearchError) as ctx:
|
|
save_failure_report_to_opensearch(_failure_report())
|
|
self.assertIn("disk", str(ctx.exception))
|
|
|
|
def test_keyerror_wraps_to_invalid_failure_report(self):
|
|
"""A malformed failure report (missing a required field) is
|
|
surfaced as InvalidFailureReport so the caller can route it
|
|
differently from infra errors."""
|
|
report = _failure_report()
|
|
del report["feedback_type"]
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save"),
|
|
):
|
|
with self.assertRaises(InvalidFailureReport):
|
|
save_failure_report_to_opensearch(report)
|
|
|
|
def test_index_dedup_pattern_searches_both_old_and_new_names(self):
|
|
"""The split-PR rename forensic→failure left existing data
|
|
in dmarc_forensic*; the dedup search must check both names
|
|
so re-runs don't double-index."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Search") as mock_search_cls,
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save"),
|
|
):
|
|
mock_search_cls.return_value.execute.return_value = []
|
|
save_failure_report_to_opensearch(_failure_report())
|
|
search_index = mock_search_cls.call_args.kwargs["index"]
|
|
self.assertIn("dmarc_failure*", search_index)
|
|
self.assertIn("dmarc_forensic*", search_index)
|
|
|
|
def test_index_name_uses_arrival_date_for_monthly_partition(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch.object(opensearch_module._FailureReportDoc, "save"),
|
|
):
|
|
save_failure_report_to_opensearch(_failure_report(), monthly_indexes=True)
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("dmarc_failure-2024-01", index_calls)
|
|
|
|
@unittest.skipUnless(hasattr(time, "tzset"), "requires POSIX time.tzset()")
|
|
def test_arrival_date_epoch_is_utc_regardless_of_host_timezone(self):
|
|
"""arrival_date_utc is a UTC wall-clock string; the epoch-ms
|
|
value stored in the document (and used in the dedup query) must
|
|
be its true UTC epoch on any host. Regression test for
|
|
https://github.com/domainaware/parsedmarc/issues/811 (bug 1):
|
|
the naive parse used to shift the stored epoch by the host's
|
|
UTC offset."""
|
|
force_tz(self)
|
|
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch("parsedmarc.opensearch._FailureReportDoc") as mock_doc_cls,
|
|
):
|
|
save_failure_report_to_opensearch(_failure_report())
|
|
# Fixture arrival_date_utc is 2024-01-01 00:00:00 UTC.
|
|
self.assertEqual(mock_doc_cls.call_args.kwargs["arrival_date"], 1704067200000)
|
|
|
|
def test_failure_search_index_with_suffix_and_prefix(self):
|
|
"""When both suffix and prefix are set, the dedup search
|
|
pattern joins them onto BOTH dmarc_failure* and
|
|
dmarc_forensic* (the rename back-compat)."""
|
|
with (
|
|
patch("parsedmarc.opensearch.Search") as mock_search_cls,
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save"),
|
|
):
|
|
mock_search_cls.return_value.execute.return_value = []
|
|
save_failure_report_to_opensearch(
|
|
_failure_report(),
|
|
index_suffix="tenant_a",
|
|
index_prefix="cust_",
|
|
)
|
|
search_index = mock_search_cls.call_args.kwargs["index"]
|
|
self.assertIn("cust_dmarc_failure_tenant_a*", search_index)
|
|
self.assertIn("cust_dmarc_forensic_tenant_a*", search_index)
|
|
|
|
def test_failure_index_honours_suffix_and_prefix(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch.object(opensearch_module._FailureReportDoc, "save"),
|
|
):
|
|
save_failure_report_to_opensearch(
|
|
_failure_report(),
|
|
index_suffix="tenant_a",
|
|
index_prefix="cust_",
|
|
)
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("cust_dmarc_failure_tenant_a-2024-01-01", index_calls)
|
|
|
|
def test_from_header_with_empty_display_name(self):
|
|
"""When the From display name is empty, the code uses the
|
|
address alone (covers the early-return branch in the
|
|
display-name handling)."""
|
|
report = _failure_report()
|
|
report["parsed_sample"]["headers"]["From"] = [["", "sender@example.com"]]
|
|
report["parsed_sample"]["headers"]["To"] = [["", "rcpt@example.com"]]
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save") as mock_save,
|
|
):
|
|
save_failure_report_to_opensearch(report)
|
|
mock_save.assert_called_once()
|
|
|
|
def test_to_header_with_non_empty_display_joins_with_brackets(self):
|
|
"""The other branch: non-empty display joins display+addr
|
|
with " <" and appends ">", e.g. 'RT <rcpt@example.com>'."""
|
|
report = _failure_report()
|
|
report["parsed_sample"]["headers"]["To"] = [["RT", "rcpt@example.com"]]
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save") as mock_save,
|
|
):
|
|
save_failure_report_to_opensearch(report)
|
|
mock_save.assert_called_once()
|
|
|
|
def test_sample_address_lists_indexed_for_reply_to_cc_bcc_attachments(self):
|
|
"""A failure report sample can carry reply_to / cc / bcc /
|
|
attachments. Each populates a nested InnerDoc on the sample —
|
|
if the add_* helpers regress, those nested docs would be
|
|
silently empty in OpenSearch."""
|
|
report = _failure_report()
|
|
report["parsed_sample"]["reply_to"] = [
|
|
{"display_name": "RT", "address": "rt@example.com"}
|
|
]
|
|
report["parsed_sample"]["cc"] = [
|
|
{"display_name": "CC", "address": "cc@example.com"}
|
|
]
|
|
report["parsed_sample"]["bcc"] = [
|
|
{"display_name": "", "address": "bcc@example.com"}
|
|
]
|
|
report["parsed_sample"]["attachments"] = [
|
|
{
|
|
"filename": "a.pdf",
|
|
"mail_content_type": "application/pdf",
|
|
"sha256": "deadbeef",
|
|
}
|
|
]
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._FailureReportDoc, "save") as mock_save,
|
|
):
|
|
save_failure_report_to_opensearch(report)
|
|
mock_save.assert_called_once()
|
|
|
|
def test_reply_to_header_flattened_and_indexed(self):
|
|
"""A Reply-To header is flattened to a display string on
|
|
``sample.headers["reply-to"]`` — so the failure dashboard's
|
|
``sample.headers.reply-to.keyword`` column resolves — and each
|
|
Reply-To address also populates the nested ``sample.reply_to``
|
|
docs. Asserts on the document handed to .save(), not merely
|
|
that save ran."""
|
|
report = _failure_report()
|
|
report["parsed_sample"]["headers"]["Reply-To"] = [
|
|
["Real One", "real@phish.example"]
|
|
]
|
|
report["parsed_sample"]["reply_to"] = [
|
|
{"display_name": "Real One", "address": "real@phish.example"}
|
|
]
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(
|
|
opensearch_module._FailureReportDoc, "save", autospec=True
|
|
) as mock_save,
|
|
):
|
|
save_failure_report_to_opensearch(report)
|
|
doc = mock_save.call_args.args[0]
|
|
self.assertEqual(
|
|
doc.sample.headers["reply-to"], "Real One <real@phish.example>"
|
|
)
|
|
self.assertEqual(
|
|
[a.address for a in doc.sample.reply_to], ["real@phish.example"]
|
|
)
|
|
|
|
def test_reply_to_header_without_display_name_flattens_to_address(self):
|
|
"""A Reply-To header with no display name flattens to the bare
|
|
address — the empty-display branch of the header flattening,
|
|
matching the From/To handling."""
|
|
report = _failure_report()
|
|
report["parsed_sample"]["headers"]["Reply-To"] = [["", "noname@phish.example"]]
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(
|
|
opensearch_module._FailureReportDoc, "save", autospec=True
|
|
) as mock_save,
|
|
):
|
|
save_failure_report_to_opensearch(report)
|
|
doc = mock_save.call_args.args[0]
|
|
self.assertEqual(doc.sample.headers["reply-to"], "noname@phish.example")
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# save_smtp_tls_report_to_opensearch
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
class TestSaveSmtpTlsReport(unittest.TestCase):
|
|
def test_save_emits_one_document(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._SMTPTLSReportDoc, "save") as mock_save,
|
|
):
|
|
save_smtp_tls_report_to_opensearch(_smtp_tls_report())
|
|
mock_save.assert_called_once()
|
|
|
|
def test_already_saved_raises_on_dedup_hit(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_populated_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._SMTPTLSReportDoc, "save") as mock_save,
|
|
):
|
|
with self.assertRaises(AlreadySaved):
|
|
save_smtp_tls_report_to_opensearch(_smtp_tls_report())
|
|
mock_save.assert_not_called()
|
|
|
|
def test_search_exception_wraps_to_opensearch_error(self):
|
|
bad = MagicMock()
|
|
bad.execute.side_effect = RuntimeError("network")
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=bad),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
):
|
|
with self.assertRaises(OpenSearchError):
|
|
save_smtp_tls_report_to_opensearch(_smtp_tls_report())
|
|
|
|
def test_save_exception_wraps_to_opensearch_error(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(
|
|
opensearch_module._SMTPTLSReportDoc,
|
|
"save",
|
|
side_effect=RuntimeError("disk"),
|
|
),
|
|
):
|
|
with self.assertRaises(OpenSearchError):
|
|
save_smtp_tls_report_to_opensearch(_smtp_tls_report())
|
|
|
|
def test_index_name_uses_begin_date_for_monthly_partition(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch.object(opensearch_module._SMTPTLSReportDoc, "save"),
|
|
):
|
|
save_smtp_tls_report_to_opensearch(_smtp_tls_report(), monthly_indexes=True)
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("smtp_tls-2024-02", index_calls)
|
|
|
|
def test_index_name_honours_suffix_and_prefix(self):
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index") as mock_index_cls,
|
|
patch.object(opensearch_module._SMTPTLSReportDoc, "save"),
|
|
):
|
|
save_smtp_tls_report_to_opensearch(
|
|
_smtp_tls_report(), index_suffix="t1", index_prefix="cust_"
|
|
)
|
|
index_calls = [c.args[0] for c in mock_index_cls.call_args_list]
|
|
self.assertIn("cust_smtp_tls_t1-2024-02-03", index_calls)
|
|
|
|
def test_policy_without_strings_or_mx_patterns(self):
|
|
"""policy_strings / mx_host_patterns are optional in the
|
|
report shape — verify the branch where they're absent."""
|
|
report = _smtp_tls_report()
|
|
for policy in report["policies"]:
|
|
policy.pop("policy_strings", None)
|
|
policy.pop("mx_host_patterns", None)
|
|
policy.pop("failure_details", None)
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._SMTPTLSReportDoc, "save") as mock_save,
|
|
):
|
|
save_smtp_tls_report_to_opensearch(report)
|
|
mock_save.assert_called_once()
|
|
|
|
def test_failure_details_all_optional_fields_populated(self):
|
|
"""Exercise every optional field in failure_details so the
|
|
full set of `if "x" in failure_detail` branches runs."""
|
|
report = _smtp_tls_report()
|
|
report["policies"][0]["failure_details"] = [
|
|
{
|
|
"result_type": "certificate-expired",
|
|
"failed_session_count": 1,
|
|
"receiving_mx_hostname": "mx.example.com",
|
|
"additional_information_uri": "https://example.com/why",
|
|
"failure_reason_code": "ERR_CERT",
|
|
"ip_address": "10.0.0.5",
|
|
"receiving_ip": "10.0.0.2",
|
|
"receiving_mx_helo": "mx.helo.example.com",
|
|
"sending_mta_ip": "10.0.0.1",
|
|
}
|
|
]
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(opensearch_module._SMTPTLSReportDoc, "save") as mock_save,
|
|
):
|
|
save_smtp_tls_report_to_opensearch(report)
|
|
mock_save.assert_called_once()
|
|
|
|
def test_save_populates_combined_policy_and_failure_detail_fields(self):
|
|
"""Regression guard for the SMTP TLS analogue of issue #169:
|
|
policies and their failure_details are object arrays, so stacked
|
|
terms aggregations on their subfields cross-product just like
|
|
dkim_results/spf_results did. Two policies (one with two failure
|
|
details, one with none) must yield exactly two policies_combined
|
|
entries and two failure_details_combined entries, not a
|
|
cross-product. autospec=True is required on the save patch so
|
|
mock_save.call_args captures the doc instance as ``self``."""
|
|
report = _smtp_tls_report(
|
|
policies=[
|
|
{
|
|
"policy_domain": "example.com",
|
|
"policy_type": "sts",
|
|
"successful_session_count": 100,
|
|
"failed_session_count": 2,
|
|
"failure_details": [
|
|
{
|
|
"result_type": "certificate-expired",
|
|
"failed_session_count": 1,
|
|
"sending_mta_ip": "192.0.2.1",
|
|
"receiving_ip": "203.0.113.1",
|
|
"receiving_mx_hostname": "mx1.example.com",
|
|
"additional_info_uri": (
|
|
"https://reports.example.com/tls-help"
|
|
),
|
|
},
|
|
{
|
|
"result_type": "starttls-not-supported",
|
|
"failed_session_count": 1,
|
|
"sending_mta_ip": "192.0.2.2",
|
|
"receiving_ip": "203.0.113.2",
|
|
"receiving_mx_hostname": "mx2.example.com",
|
|
},
|
|
],
|
|
},
|
|
{
|
|
"policy_domain": "example.net",
|
|
"policy_type": "tlsa",
|
|
"successful_session_count": 50,
|
|
"failed_session_count": 0,
|
|
},
|
|
]
|
|
)
|
|
with (
|
|
patch("parsedmarc.opensearch.Search", return_value=_empty_search()),
|
|
patch("parsedmarc.opensearch.Index"),
|
|
patch.object(
|
|
opensearch_module._SMTPTLSReportDoc, "save", autospec=True
|
|
) as mock_save,
|
|
):
|
|
save_smtp_tls_report_to_opensearch(report)
|
|
doc = mock_save.call_args[0][0]
|
|
self.assertEqual(
|
|
list(doc.policies_combined), ["example.com / sts", "example.net / tlsa"]
|
|
)
|
|
expected_detail_expired = (
|
|
"example.com / sts / certificate-expired / 192.0.2.1 / "
|
|
"203.0.113.1 / mx1.example.com"
|
|
)
|
|
expected_detail_starttls = (
|
|
"example.com / sts / starttls-not-supported / 192.0.2.2 / "
|
|
"203.0.113.2 / mx2.example.com"
|
|
)
|
|
self.assertEqual(
|
|
list(doc.failure_details_combined),
|
|
[expected_detail_expired, expected_detail_starttls],
|
|
)
|
|
# The parser emits additional_info_uri (SMTPTLSFailureDetailsOptional
|
|
# in types.py); the saver must persist it on the declared
|
|
# additional_information_uri field rather than dropping it.
|
|
self.assertEqual(
|
|
doc.policies[0].failure_details[0].additional_information_uri,
|
|
"https://reports.example.com/tls-help",
|
|
)
|
|
|
|
|
|
class TestBackwardCompatAlias(unittest.TestCase):
|
|
def test_save_forensic_alias_points_to_save_failure(self):
|
|
self.assertIs(
|
|
opensearch_module.save_forensic_report_to_opensearch,
|
|
opensearch_module.save_failure_report_to_opensearch,
|
|
)
|
|
|
|
def test_forensic_doc_alias_points_to_failure_doc(self):
|
|
self.assertIs(
|
|
opensearch_module._ForensicReportDoc, opensearch_module._FailureReportDoc
|
|
)
|
|
self.assertIs(
|
|
opensearch_module._ForensicSampleDoc, opensearch_module._FailureSampleDoc
|
|
)
|
|
|
|
|
|
# Silence unused-import lint in the test module preamble.
|
|
_ = call
|
|
|
|
|
|
if __name__ == "__main__":
|
|
unittest.main(verbosity=2)
|