mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-09-05 13:38:00 +00:00
* Make the output and mailbox integrations optional extras (#883) Breaking change for the next major release: pip install parsedmarc now installs the parsing core plus a working core CLI (file, IMAP, Maildir, and mbox input; CSV/JSON, Splunk HEC, webhook, and syslog output). Everything else moves behind an extra: elastic, opensearch, kafka, s3, gelf, loganalytics, msgraph, and gmail, joining the existing postgresql extra, with an umbrella [all] that deliberately excludes postgresql (psycopg's binary wheels do not exist on every platform, so parsedmarc[all] must never fail to install there). cli.py imports the six SDK-dependent output modules behind the #884 TYPE_CHECKING/try-except guard; a configured section whose extra is missing fails fast with a ConfigurationError naming the section and the exact pip install command — including the msgraph and gmail_api mailbox sections (detected via parsedmarc.mail's placeholder classes) and postgresql (checked before the constructor so the startup retry loop does not retry a missing dependency for a minute). The Azure/kiota Graph error types fall back to never-raised sentinel classes. The Docker image installs [all,postgresql], so container users see no change. CI lint installs [build,all,postgresql]; the unit-test job installs [build,all], deliberately without postgresql so test_postgres.py's absent-psycopg arm stays exercised. The never-imported dateparser dependency is dropped in favor of declaring python-dateutil, which utils.py actually imports; pytz moves to the build extra for the one test that uses it. Verified live: a no-extras wheel install imports, parses samples, and reports the install hint for each gated section; a [all] install restores every integration; the Docker image builds with every SDK importable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Patch psycopg presence in the PostgreSQL CLI wiring tests CI's unit-test job deliberately installs [build,all] without the postgresql extra, so parsedmarc.cli.postgres.psycopg is None there and the new missing-extra presence check correctly made _main exit 1 before the wiring under test ran. The tests simulate the SDK being available (PostgreSQLClient is mocked at the SDK boundary), so the module-level psycopg handle is now patched present in setUp. Verified against a simulated psycopg-absent environment as well as the local full install. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address Copilot review: narrow guards to ModuleNotFoundError, fix docs - The optional-integration and Graph error-type import guards now catch ModuleNotFoundError instead of ImportError, so only a genuinely absent package reads as a missing extra; a broken-but-present SDK fails loudly with its real error instead of masquerading as one. The test blocker raises ModuleNotFoundError accordingly — the exact exception a missing package produces. - _missing_extra_hint docstring no longer calls every gated integration an output module (it also serves the msgraph/gmail_api mailbox sections). - Fix the pre-existing passsword typo in usage.md's kafka section; the INI key the code reads is password (cli.py _parse_config). Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Quote extras specs in copy-paste install commands From Copilot's second review round: zsh treats an unquoted .[build,all] as a glob and fails with 'no matches found', so the commands shown in AGENTS.md, CONTRIBUTING.md, dashboards/README.md, and the bootstrap script's comment are now quoted. The CI workflows keep the unquoted form: they run under bash, which passes unmatched globs through literally. The suggestion to change the 'Choosing what to install' heading level was rejected — it is a subsection of 'Installing parsedmarc', matching the file's existing hierarchy. Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Fix upgrade command in the changelog * Documentation review: accuracy, spelling, grammar, and clarity pass A full prose review of docs/source, README, CONTRIBUTING, and the dashboards README, with every accuracy claim verified against the code before changing it. Highlights: - usage.md: documented six missing [general] options (the CSV/JSON filename options, prettify_json, normalize_timespan_threshold_hours), the required kafka smtp_tls_topic, [imap] timeout/max_retries, and the postgresql env-var prefix; corrected the maildir_path default (None, not INBOX — cli.py Namespace defaults), the mailbox check_timeout option name, the systemd restart interval (RestartSec is 5m), and merged the duplicate silent entry; quoted every copy-paste extras spec for zsh safety. - elasticsearch.md: fixed an invalid openssl command (rsa:4096 -nodes), the dashboards filename (opensearch_dashboards.ndjson, matching the file the link serves), and assorted grammar. - davmail.md: the service-enable command now enables davmail.service (was parsedmarc.service — a copy-paste error that left DavMail unenabled), plus a view typo and DavMail capitalization. - output.md: the example schema reference is RFC 7489 Appendix C (7480 is RDAP). kibana.md: SPF relies on the SMTP envelope, not session headers (RFC 7208). dmarc.md: DKM -> DKIM. - README: the intro now also names the OpenSearch/Grafana stack, matching the feature list. CONTRIBUTING: pre-PR checks now include ruff format --check and pyright, matching CI's lint job. - dashboards/README: the service table and seed description now include the PostgreSQL backend the compose stack runs. Sample data blocks, the CLI-help mirror block, and released CHANGELOG entries were deliberately left untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Docstring review: accuracy, spelling, grammar, and clarity pass Every docstring in parsedmarc/, parsedmarc/mail/, the maps maintainer scripts, and the test suite reviewed with each claim verified against the code it documents. Text-only — no behavior changes. Highlights: - Copy-paste errors corrected: parsed_smtp_tls_reports_to_csv and splunk/loganalytics save functions described aggregate or failure reports they do not handle; LogAnalyticsException claimed to be an Elasticsearch error. - Docstring/behavior mismatches: parse_report_email's report_type enumeration omitted smtp_tls; parse_failure_report typed msg_date as str (it is datetime); strip_attachment_payloads claimed payloads are replaced with None (the key is deleted); kafkaclient's failure and SMTP TLS savers claimed per-record slicing while sending the whole list in one message (docstrings now describe reality — whether slicing was intended is flagged for follow-up); the postgres savers claimed to take parse_report_file's return value but receive the inner report dict; elastic/opensearch save functions' Raises listed only AlreadySaved. - None-as-semantic-state documented where missing (get_base_domain, get_ip_address_country), enumeration completeness fixed (get_ip_address_info's 9 result keys, maps script outputs, TSV columns), and the stale 44-industry-types count corrected to the 46 the authoritative README list defines. - Test docstrings aligned with what the tests actually assert, including two that overstated coverage of the elastic/opensearch address-list tests. - Two argparse help strings fixed: file_path now names SMTP TLS report files alongside aggregate and failure, mirrored into usage.md's CLI-help block; --offline's doubled spaces removed (rendered help unchanged). - elasticsearch.md's security claim corrected against Elastic's docs: security is enabled and auto-configured on first startup since 8.0 (not "8.7 secure mode"), so the settings are verified, not hand-written. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
289 lines
12 KiB
Python
289 lines
12 KiB
Python
"""Tests for parsedmarc.kafkaclient"""
|
|
|
|
import json
|
|
import unittest
|
|
from typing import cast
|
|
from unittest.mock import MagicMock, patch
|
|
|
|
from kafka.errors import UnknownTopicOrPartitionError
|
|
|
|
from parsedmarc.kafkaclient import KafkaClient, KafkaError, _BootstrapError
|
|
|
|
|
|
def _producer(client: KafkaClient) -> MagicMock:
|
|
"""The patched KafkaProducer as a MagicMock, for assertion access."""
|
|
return cast(MagicMock, client.producer)
|
|
|
|
|
|
def _aggregate_report():
|
|
return {
|
|
"report_metadata": {
|
|
"org_name": "TestOrg",
|
|
"org_email": "test@example.com",
|
|
"report_id": "r-123",
|
|
"begin_date": "2024-01-01 00:00:00",
|
|
"end_date": "2024-01-02 00:00:00",
|
|
},
|
|
"policy_published": {"domain": "example.com", "p": "none"},
|
|
"records": [
|
|
{"source": {"ip_address": "192.0.2.1"}, "count": 1},
|
|
{"source": {"ip_address": "192.0.2.2"}, "count": 2},
|
|
],
|
|
}
|
|
|
|
|
|
class TestKafkaClientInit(unittest.TestCase):
|
|
"""KafkaProducer config wiring: SSL, SASL, plain — each branch has
|
|
user-facing security consequences if it's wrong."""
|
|
|
|
def test_init_plain_no_ssl(self):
|
|
"""No SSL, no auth: just bootstrap_servers and serializer."""
|
|
with patch("parsedmarc.kafkaclient.KafkaProducer") as mock_producer:
|
|
KafkaClient(kafka_hosts=["broker:9092"])
|
|
kwargs = mock_producer.call_args.kwargs
|
|
self.assertEqual(kwargs["bootstrap_servers"], ["broker:9092"])
|
|
self.assertNotIn("security_protocol", kwargs)
|
|
self.assertNotIn("sasl_plain_username", kwargs)
|
|
|
|
def test_init_ssl_enables_ssl_security_protocol(self):
|
|
with (
|
|
patch("parsedmarc.kafkaclient.KafkaProducer") as mock_producer,
|
|
patch("parsedmarc.kafkaclient.create_default_context") as mock_ctx,
|
|
):
|
|
KafkaClient(kafka_hosts=["broker:9093"], ssl=True)
|
|
kwargs = mock_producer.call_args.kwargs
|
|
self.assertEqual(kwargs["security_protocol"], "SSL")
|
|
self.assertIs(kwargs["ssl_context"], mock_ctx.return_value)
|
|
|
|
def test_init_username_implies_ssl(self):
|
|
"""Doc says ssl=True is implied when username/password supplied."""
|
|
with (
|
|
patch("parsedmarc.kafkaclient.KafkaProducer") as mock_producer,
|
|
patch("parsedmarc.kafkaclient.create_default_context"),
|
|
):
|
|
KafkaClient(kafka_hosts=["broker:9093"], username="user", password="pass")
|
|
kwargs = mock_producer.call_args.kwargs
|
|
self.assertEqual(kwargs["security_protocol"], "SSL")
|
|
self.assertEqual(kwargs["sasl_plain_username"], "user")
|
|
self.assertEqual(kwargs["sasl_plain_password"], "pass")
|
|
|
|
def test_init_uses_provided_ssl_context(self):
|
|
"""A caller-supplied SSLContext takes precedence over the
|
|
default context — this lets ops pin to a private CA."""
|
|
custom_ctx = MagicMock()
|
|
with (
|
|
patch("parsedmarc.kafkaclient.KafkaProducer") as mock_producer,
|
|
patch("parsedmarc.kafkaclient.create_default_context") as mock_default,
|
|
):
|
|
KafkaClient(kafka_hosts=["b:9093"], ssl=True, ssl_context=custom_ctx)
|
|
self.assertIs(mock_producer.call_args.kwargs["ssl_context"], custom_ctx)
|
|
mock_default.assert_not_called()
|
|
|
|
def test_init_value_serializer_emits_utf8_json(self):
|
|
"""The value_serializer turns Python objects into UTF-8 JSON
|
|
bytes. A regression here would corrupt every event sent."""
|
|
with patch("parsedmarc.kafkaclient.KafkaProducer") as mock_producer:
|
|
KafkaClient(kafka_hosts=["b"])
|
|
serializer = mock_producer.call_args.kwargs["value_serializer"]
|
|
result = serializer({"hello": "world", "n": 1})
|
|
self.assertEqual(json.loads(result.decode("utf-8")), {"hello": "world", "n": 1})
|
|
|
|
def test_init_no_brokers_available_raises_kafka_error(self):
|
|
with patch(
|
|
"parsedmarc.kafkaclient.KafkaProducer",
|
|
side_effect=_BootstrapError(),
|
|
):
|
|
with self.assertRaises(KafkaError) as ctx:
|
|
KafkaClient(kafka_hosts=["unreachable:9092"])
|
|
self.assertIn("No Kafka brokers", str(ctx.exception))
|
|
|
|
|
|
class TestKafkaClientHelpers(unittest.TestCase):
|
|
"""Static helpers used by save_aggregate."""
|
|
|
|
def test_strip_metadata_lifts_keys_to_root_and_drops_metadata(self):
|
|
report = _aggregate_report()
|
|
result = KafkaClient.strip_metadata(report)
|
|
self.assertEqual(result["org_name"], "TestOrg")
|
|
self.assertEqual(result["org_email"], "test@example.com")
|
|
self.assertEqual(result["report_id"], "r-123")
|
|
self.assertNotIn("report_metadata", result)
|
|
|
|
def test_generate_date_range_iso_format(self):
|
|
report = _aggregate_report()
|
|
date_range = KafkaClient.generate_date_range(report)
|
|
self.assertEqual(date_range, ["2024-01-01T00:00:00", "2024-01-02T00:00:00"])
|
|
|
|
|
|
class TestSaveAggregateReportsToKafka(unittest.TestCase):
|
|
"""save_aggregate sends one Kafka message per record (slice), with
|
|
the metadata + policy duplicated onto each slice for Kibana parity."""
|
|
|
|
def _client(self):
|
|
with patch("parsedmarc.kafkaclient.KafkaProducer"):
|
|
return KafkaClient(kafka_hosts=["b:9092"])
|
|
|
|
def test_sends_one_message_per_record(self):
|
|
client = self._client()
|
|
client.save_aggregate_reports_to_kafka(_aggregate_report(), "dmarc-aggregate")
|
|
# 2 records in the sample report → 2 producer.send calls.
|
|
self.assertEqual(_producer(client).send.call_count, 2)
|
|
# Topic is forwarded verbatim.
|
|
for call in _producer(client).send.call_args_list:
|
|
self.assertEqual(call.args[0], "dmarc-aggregate")
|
|
|
|
def test_each_slice_carries_metadata(self):
|
|
client = self._client()
|
|
client.save_aggregate_reports_to_kafka(_aggregate_report(), "topic")
|
|
sent = [call.args[1] for call in _producer(client).send.call_args_list]
|
|
for slice_ in sent:
|
|
self.assertEqual(slice_["org_name"], "TestOrg")
|
|
self.assertEqual(slice_["org_email"], "test@example.com")
|
|
self.assertEqual(slice_["report_id"], "r-123")
|
|
self.assertEqual(
|
|
slice_["date_range"], ["2024-01-01T00:00:00", "2024-01-02T00:00:00"]
|
|
)
|
|
self.assertEqual(
|
|
slice_["policy_published"], {"domain": "example.com", "p": "none"}
|
|
)
|
|
|
|
def test_empty_list_is_a_noop(self):
|
|
client = self._client()
|
|
client.save_aggregate_reports_to_kafka([], "topic")
|
|
_producer(client).send.assert_not_called()
|
|
|
|
def test_dict_input_normalized_to_list(self):
|
|
"""Single-report dict input is wrapped to a list."""
|
|
client = self._client()
|
|
client.save_aggregate_reports_to_kafka(_aggregate_report(), "topic")
|
|
# 2 records still sent (one report with 2 records, not multiple reports).
|
|
self.assertEqual(_producer(client).send.call_count, 2)
|
|
|
|
def test_unknown_topic_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).send.side_effect = UnknownTopicOrPartitionError()
|
|
with self.assertRaises(KafkaError) as ctx:
|
|
client.save_aggregate_reports_to_kafka(_aggregate_report(), "missing")
|
|
self.assertIn("Unknown topic or partition", str(ctx.exception))
|
|
|
|
def test_generic_send_exception_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).send.side_effect = RuntimeError("transport failure")
|
|
with self.assertRaises(KafkaError) as ctx:
|
|
client.save_aggregate_reports_to_kafka(_aggregate_report(), "topic")
|
|
self.assertIn("transport failure", str(ctx.exception))
|
|
|
|
def test_flush_exception_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).flush.side_effect = RuntimeError("flush failure")
|
|
with self.assertRaises(KafkaError) as ctx:
|
|
client.save_aggregate_reports_to_kafka(_aggregate_report(), "topic")
|
|
self.assertIn("flush failure", str(ctx.exception))
|
|
|
|
|
|
class TestSaveFailureReportsToKafka(unittest.TestCase):
|
|
def _client(self):
|
|
with patch("parsedmarc.kafkaclient.KafkaProducer"):
|
|
return KafkaClient(kafka_hosts=["b:9092"])
|
|
|
|
def test_sends_full_list_in_one_message(self):
|
|
"""Failure reports are sent as one Kafka message carrying the
|
|
whole list — unlike aggregate records, which are sent as
|
|
individual slices to stay under Kafka's default 1MB message
|
|
cap."""
|
|
client = self._client()
|
|
reports = [{"id": "f1"}, {"id": "f2"}]
|
|
client.save_failure_reports_to_kafka(reports, "dmarc-failure")
|
|
_producer(client).send.assert_called_once_with("dmarc-failure", reports)
|
|
|
|
def test_dict_input_normalized_to_list(self):
|
|
client = self._client()
|
|
client.save_failure_reports_to_kafka({"id": "single"}, "topic")
|
|
# The send payload is wrapped to a single-element list.
|
|
args = _producer(client).send.call_args.args
|
|
self.assertEqual(args[1], [{"id": "single"}])
|
|
|
|
def test_empty_list_is_a_noop(self):
|
|
client = self._client()
|
|
client.save_failure_reports_to_kafka([], "topic")
|
|
_producer(client).send.assert_not_called()
|
|
|
|
def test_unknown_topic_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).send.side_effect = UnknownTopicOrPartitionError()
|
|
with self.assertRaises(KafkaError):
|
|
client.save_failure_reports_to_kafka([{"a": 1}], "missing")
|
|
|
|
def test_generic_send_error_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).send.side_effect = OSError("net")
|
|
with self.assertRaises(KafkaError):
|
|
client.save_failure_reports_to_kafka([{"a": 1}], "topic")
|
|
|
|
def test_flush_error_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).flush.side_effect = OSError("flush")
|
|
with self.assertRaises(KafkaError):
|
|
client.save_failure_reports_to_kafka([{"a": 1}], "topic")
|
|
|
|
|
|
class TestSaveSmtpTlsReportsToKafka(unittest.TestCase):
|
|
def _client(self):
|
|
with patch("parsedmarc.kafkaclient.KafkaProducer"):
|
|
return KafkaClient(kafka_hosts=["b:9092"])
|
|
|
|
def test_sends_full_list_in_one_message(self):
|
|
client = self._client()
|
|
reports = [{"organization_name": "x"}]
|
|
client.save_smtp_tls_reports_to_kafka(reports, "smtp-tls")
|
|
_producer(client).send.assert_called_once_with("smtp-tls", reports)
|
|
|
|
def test_dict_input_normalized_to_list(self):
|
|
client = self._client()
|
|
client.save_smtp_tls_reports_to_kafka({"organization_name": "x"}, "topic")
|
|
args = _producer(client).send.call_args.args
|
|
self.assertEqual(args[1], [{"organization_name": "x"}])
|
|
|
|
def test_empty_list_is_a_noop(self):
|
|
client = self._client()
|
|
client.save_smtp_tls_reports_to_kafka([], "topic")
|
|
_producer(client).send.assert_not_called()
|
|
|
|
def test_unknown_topic_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).send.side_effect = UnknownTopicOrPartitionError()
|
|
with self.assertRaises(KafkaError):
|
|
client.save_smtp_tls_reports_to_kafka([{"a": 1}], "missing")
|
|
|
|
def test_generic_send_error_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).send.side_effect = RuntimeError("oops")
|
|
with self.assertRaises(KafkaError):
|
|
client.save_smtp_tls_reports_to_kafka([{"a": 1}], "topic")
|
|
|
|
def test_flush_error_translates_to_kafka_error(self):
|
|
client = self._client()
|
|
_producer(client).flush.side_effect = RuntimeError("flush")
|
|
with self.assertRaises(KafkaError):
|
|
client.save_smtp_tls_reports_to_kafka([{"a": 1}], "topic")
|
|
|
|
|
|
class TestKafkaClientClose(unittest.TestCase):
|
|
def test_close_calls_underlying_producer_close(self):
|
|
with patch("parsedmarc.kafkaclient.KafkaProducer"):
|
|
client = KafkaClient(kafka_hosts=["b"])
|
|
client.close()
|
|
_producer(client).close.assert_called_once()
|
|
|
|
|
|
class TestKafkaBackwardCompatAlias(unittest.TestCase):
|
|
def test_forensic_alias_points_to_failure_method(self):
|
|
self.assertIs(
|
|
KafkaClient.save_forensic_reports_to_kafka,
|
|
KafkaClient.save_failure_reports_to_kafka,
|
|
)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
unittest.main(verbosity=2)
|