mirror of
https://github.com/domainaware/parsedmarc.git
synced 2026-08-01 21:22:18 +00:00
Backfill combined DKIM/SPF fields automatically at startup
migrate_indexes() now backfills dkim_results_combined and spf_results_combined on aggregate documents saved by older versions, so ES/OS users get historical data in the reworked alignment tables without running the documented _update_by_query by hand. The backfill is submitted as a non-blocking background task (wait_for_completion=false, conflicts=proceed) guarded by a cheap count query, making repeated startups a fast no-op once an index is backfilled; any cluster error is logged as a warning and retried at the next startup rather than raised. The manual command remains documented for users who upgrade dashboards without pointing the new parsedmarc at the cluster or who want to control write-load timing. The legacy published_policy.fo long-to-text reindex migration in the OpenSearch module is kept ahead of the new backfill, for clusters upgraded from very old data. Verified end-to-end against the live dev environment: a real CLI startup backfilled 9 stripped OpenSearch documents (logged with task ID) while the already-backfilled Elasticsearch side stayed silent, and a second startup was silent on both engines. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5
parent
e892de794b
commit
bbee148d2a
@@ -235,14 +235,28 @@ result paired, which the dashboards' alignment-detail tables aggregate on.
|
||||
Reports saved by older versions lack these fields and will not appear in
|
||||
those tables.
|
||||
|
||||
Running the following once per cluster backfills the fields on existing
|
||||
documents. It is idempotent (documents that already have the fields are
|
||||
skipped), so it is safe to re-run. It works identically on OpenSearch;
|
||||
just adjust the URL and credentials. The query matches only documents
|
||||
that have at least one DKIM or SPF auth result and lack the corresponding
|
||||
combined field; documents with no auth results are skipped, because an
|
||||
`exists` query cannot see an empty array, and for search purposes an
|
||||
empty `dkim_results_combined` is identical to an absent one.
|
||||
parsedmarc now backfills this automatically. On startup, it runs a cheap
|
||||
count query against each configured aggregate index pattern to check for
|
||||
documents that have DKIM or SPF results but are missing the corresponding
|
||||
combined field. If any are found, it submits the backfill as a background
|
||||
`_update_by_query` task (`wait_for_completion=false`), so startup is never
|
||||
blocked on it; progress is logged, including the task ID. The check itself
|
||||
is idempotent — once an index is fully backfilled, later startups see a
|
||||
count of 0 and log nothing further — and it works the same way on
|
||||
OpenSearch. Any error talking to the cluster (for example, no indexes yet
|
||||
on a fresh install) is logged as a warning and retried on the next startup,
|
||||
rather than aborting parsedmarc.
|
||||
|
||||
If you upgrade the dashboards without pointing the new parsedmarc version
|
||||
at the cluster, or you'd rather control when the write load happens, you
|
||||
can still run the backfill manually. It is idempotent (documents that
|
||||
already have the fields are skipped), so it is safe to re-run. It works
|
||||
identically on OpenSearch; just adjust the URL and credentials. The query
|
||||
matches only documents that have at least one DKIM or SPF auth result and
|
||||
lack the corresponding combined field; documents with no auth results are
|
||||
skipped, because an `exists` query cannot see an empty array, and for
|
||||
search purposes an empty `dkim_results_combined` is identical to an
|
||||
absent one.
|
||||
|
||||
```bash
|
||||
curl -X POST "http://localhost:9200/dmarc_aggregate*/_update_by_query?conflicts=proceed&wait_for_completion=false" \
|
||||
|
||||
Reference in New Issue
Block a user