20 KiB
Workflow Runner Refactor — Design
Date: 2026-05-19
Updated: 2026-09-24 — see "Status update (2026-09-24)" below; line references
throughout refreshed against current dev.
Branch base: dev
Status: Approved design, pending implementation plan. Motivating bugs are now
independently fixed (see status update) — this is a structural cleanup, not an
active-bug fix.
Status update (2026-09-24)
The acute bugs this design was written against have since been closed by two
targeted, already-merged fixes that do not use the WorkflowRunContext /
ContextVar approach below:
- #12386 (cross-process "file already moved" race, cited in cause 3) was
fixed by #12389 ("Fix: avoid moving files if already moved", merged in
v2.20.12):
validate_moveinupdate_filename_and_move_filesnow detects, via a checksum comparison (_path_matches_checksum), when the target file already exists because a concurrent save already moved it, and recovers the DB pointer instead of raisingCannotMoveFilesException. This is a different mechanism from anything proposed here and stays as-is; nothing in this refactor should touch_path_matches_checksumor its call sites. - The intra-workflow tag/
m2m_changedrace described in cause 3 and design §4 (add_nested_tagsfiringm2m_changed→refresh_from_db()→ wiping an earlier-ordered action's unsavedcorrespondent/storage_path) was fixed by #13178 ("Fix: prevent tag assignment from reverting other pending workflow assignments", merged 2026-07-20 — after this design was written).apply_assignment_to_document(mutations.py) now applies tag changes to a freshly-fetchedDocumentinstance rather than the shared in-memory one, matching the patternapply_removal_to_documentalready used for tag removal. Because the mid-run rename this triggers now reads the old, still-current DB state (nothing else has been saved yet), it recomputes the same path and is a no-op; the real rename happens once, correctly, at the finaldocument.save(). Regression coverage:test_document_updated_workflow_assignment_storage_path_persists_with_tag_assignmentintest_workflows.py(the only test #13178 itself added). A related, earlier fix (#12664) covers a similar but distinct case — tag removal alongside a title assignment — viatest_document_updated_workflow_assignment_persists_when_removing_trigger_tag. - The
filename/archive_filenameexclusion fromupdate_fields(cause 3, design §4 point 2) was already in place before this design was written (added by #12390-adjacent work) and is unaffected — still correctly attributed here as a load-bearing cross-process guard, not duct tape. - #13178's own PR description is the origin of this idea: "Once the beta is out, I do have some thoughts about using ContextVar to delete this class entirely. We've run into it plenty of times." So the guard concept predates this design; the acute pain it was meant to address has since been patched piecemeal instead.
What this means for scope: causes 1 and 2 below (dual-mode branching,
staged-file parameter plumbing) are unchanged and fully present in current code
— this refactor is still worth doing for them. Cause 3's race is no longer an
active bug; the ContextVar guard is now defense-in-depth / simplification
(it also collapses two independent workarounds — the fresh-instance tag fetch
and the per-workflow update_fields restriction — into one clearer mechanism)
rather than a fix for a reproducible failure. Treat any "fixes a bug" framing
below as historical motivation, not a current defect claim.
Problem
Workflow execution and the Django signal layer have repeatedly produced fragile, hard-to-fix bugs (see the revert/refix history around password removal: #12803, #12814, #12716, and the filename race #12386, now closed — see status update above). Three structural causes:
-
run_workflowsis dual-mode. A single function handles both consumption (mutating aDocumentMetadataOverrides) and post-save (mutating a realDocument), branching on ause_overridesflag. The branching is concentrated in two places — the action dispatch insiderun_workflows(handlers.py:938-1014,use_overridesfirst set athandlers.py:881) andbuild_workflow_action_context(actions.py:33-83), each with two full code paths. Theapply_*helpers inworkflows/mutations.pyare already split by target type (apply_assignment_to_documentvsapply_assignment_to_overrides, etc.); the refactor unifies their callers, not the helpers themselves. Still fully present in current code — unchanged by any fix since this design was written. -
File location is an implicit, timing-dependent side channel. The
DOCUMENT_ADDEDworkflow fires fromrun_workflows_added(handlers.py:811-824), which runs while the consumer is still inside its transaction — before the consumed file is copied todocument.source_path(document_consumption_finishedis sent fromconsumer.py, the file copy happens after). The staged path is therefore threaded through asoriginal_file/caller_supplied_original_fileparameters (handlers.py:868,890-898,976-979). Actions that read the file (password removal, email attachments) depend on this plumbing being correct. Still fully present in current code. -
The workflow run could race the filename rename — now mitigated by two independent, narrower fixes (see "Status update" above), not eliminated structurally.
update_filename_and_move_files(handlers.py:437-439decorators, body440-673) is a rawpost_save/m2m_changedreceiver. When a workflow persists its changes viadocument.save(update_fields=[...]), or whenapply_assignment_to_documentmutates tags via a freshly-fetched instance (mutations.py:26-31), that write fires the receiver while the workflow is still executing. As of #13178 this no longer corrupts in-memory state (tags are applied to a separate instance, so the mid-run rename it triggers reads only already-committed data and is a no-op or self-consistent move); as of #12389 a genuine cross-process "someone else already moved this file" race recovers via checksum comparison instead of erroring. The comment documenting thefilename/archive_filenameexclusion from the workflow's ownupdate_fields(handlers.py:1019-1027) remains a load-bearing guard against a different, still-real cross-process hazard (an in-memorydocument.filenamegoing stale while another process moves the file) and is unaffected by either fix above.
Note: run_workflows_added / run_workflows_updated are connected to the
custom signals document_consumption_finished / document_updated, fired
explicitly by paperless code in a handful of known sites — not to raw Django
post_save. Only update_filename_and_move_files is a raw post_save receiver.
This refactor does not change where workflows are triggered from.
Scope
In scope:
- Refactor
run_workflowsand its action helpers around an execution-context abstraction. - Delete the
original_fileside-channel plumbing. - Make the workflow-execution → persist → rename sequence explicit and deterministic.
Out of scope:
- Changing where/when workflows are triggered (custom signal call sites unchanged).
- Reworking the matching logic (
matching.document_matches_workflow). - Any change to workflow models, serializers, or the REST API.
Design
1. WorkflowRunContext protocol
New module documents/workflows/context.py defining a typing.Protocol:
WorkflowRunContext (Protocol)
source_file: Path # where the file actually is, now
build_placeholder_context() -> dict
apply_assignment(action) -> None
apply_removal(action) -> None
persist() -> None # commit accumulated mutations
record_run(workflow, trigger_type) -> None
Two concrete implementations (which need not import the Protocol — structural typing):
ConsumptionContext— wrapsConsumableDocument+DocumentMetadataOverrides.source_filereturns the staged file path. Mutations land on the overrides.persist()is a no-op (the overrides object is returned to the caller).PersistedContext— wraps a realDocument. Mutations land on the in-memoryDocument.persist()performs a single save.
Context selection — run_workflows picks the context from the call shape:
- CONSUMPTION trigger (
ConsumableDocument+ non-Noneoverrides) →ConsumptionContext. - DOCUMENT_ADDED / DOCUMENT_UPDATED / SCHEDULED (a real
Document,overrides=None) →PersistedContext.
source_file for PersistedContext. It cannot unconditionally return
document.source_path: for the DOCUMENT_ADDED trigger the file has not yet
been moved there. The staged path is therefore passed into the PersistedContext
at construction time by run_workflows_added (which still receives it from the
document_consumption_finished signal). source_file returns that staged path
when supplied, otherwise document.source_path. This relocates the staged-path
information from a chain of function parameters into a single piece of
construction state — the original_file / caller_supplied_original_file
parameter plumbing through run_workflows and the action helpers is what gets
deleted, not the staged path itself.
WorkflowRunContext is a plain Protocol, not @runtime_checkable — the runner
constructs the context itself, so no isinstance check is needed. Genuinely
shared logic goes into module-level helper functions, not a base class.
2. run_workflows becomes branch-free
run_workflows keeps its current public signature so all call sites are
unchanged. Its body:
- Construct the appropriate context once, from the argument types.
- Run a single flat match-and-dispatch loop over matching workflows/actions, delegating every action to context methods.
No use_overrides flag anywhere. The branching currently scattered across
run_workflows, build_workflow_action_context, and the apply_* helpers
collapses into the two context classes.
3. File staging via source_file
source_file is a property of the context, fixed at construction. The
original_file and caller_supplied_original_file parameters threaded through
run_workflows and the execute_* helpers are deleted; each context resolves
the path itself (see "Context selection" above).
Deferred password removal. execute_password_removal_action, when given a
ConsumableDocument, currently installs a one-shot handler on
document_consumption_finished that picks up original_file from kwargs
later (actions.py:295-308). This deferred hook lives outside the context
abstraction. The refactor must explicitly decide its fate: either keep it as-is
(the context still constructs correctly around it) or fold the deferral into
ConsumptionContext. This is called out as an open implementation decision, not
silently absorbed.
4. Explicit workflow → persist → rename sequencing
What must be deferred is the file rename, not the DB save. run_workflows
keeps its per-workflow document.refresh_from_db() at the top of each iteration
— that is deliberate concurrency protection against bulk_update_documents
running simultaneously. Deferring all saves to a single final persist() would
let one workflow's refresh wipe a prior workflow's in-memory changes. So:
run_workflowsrefreshes and applies actions per workflow, andPersistedContext.persist()saves after each matching workflow, as today.- The save deliberately continues to exclude
filename/archive_filenamefromupdate_fields. This is not duct tape: it guards a cross-process hazard — another Celery task may have moved the file and writtenfilenameto the DB, and a stale in-memoryfilenamein our save would revert it. TheContextVarguard (below) only addresses intra-process ordering, so this exclusion stays. - The rename is suppressed for the whole run and invoked exactly once, afterward, against final committed state.
Historical race, already fixed by other means (see "Status update"):
apply_assignment_to_document used to assign tags via
document.add_nested_tags(...) directly on the shared in-memory document,
which fired m2m_changed on Document.tags.through before the workflow's
document.save(); the m2m_changed receiver update_filename_and_move_files
then called refresh_from_db() on that shared instance, wiping the workflow's
in-memory correspondent/type, and moved the file to a path computed from stale
metadata. #13178 fixed this by having apply_assignment_to_document (and
apply_removal_to_document, which already did this) mutate tags on a
freshly-fetched Document.objects.get(pk=document.pk) instead of the
shared one (mutations.py:26-31), so the refresh_from_db() triggered by
m2m_changed no longer touches the workflow's unsaved in-memory fields. The
guard below is not needed to fix that specific corruption anymore — it instead
gives a single, general mechanism that supersedes the fresh-instance-fetch
workaround (and the equivalent one already in apply_removal_to_document),
rather than requiring every future mutation path to remember to fetch a
separate instance.
To stop the rename from firing mid-workflow, a ContextVar guard is
introduced (e.g. documents/workflows/context.py module-level
_workflow_in_progress: ContextVar[bool]). update_filename_and_move_files
checks the guard and early-returns when set. run_workflows wraps its entire
persisted-path execution — not just the persist() call — in a context manager
that sets the guard via set()/reset(token). Token-based reset is
reentrancy-safe for nested saves or nested workflow runs.
The guard must span the whole execution, not just persist(), because
update_filename_and_move_files is also registered to m2m_changed on
Document.tags.through and to post_save on CustomFieldInstance
(handlers.py:437-438). A workflow action that assigns tags or custom fields
would otherwise trigger a rename mid-workflow through those signals. If the
ContextVar guard lands, the fresh-instance-fetch workaround in
apply_assignment_to_document/apply_removal_to_document becomes redundant
but is not itself incorrect — decide during implementation whether to simplify
those two call sites back to mutating document directly now that the guard
covers the hazard, or leave them as extra defense-in-depth. Note either way:
document passed to add_nested_tags/tags.clear/tags.remove must still be
re-fetched or refresh_from_db()'d for the tags relation to reflect the
change on the in-memory instance used later in the same action.
After execution completes, run_workflows calls persist() once and then
explicitly invokes the move logic once. The ContextVar is set/reset in the
same thread that runs these receivers synchronously, so they always observe the
value. (Celery prefork workers run each task in its own process; greenlet
pools are also contextvars-aware — non-issues, noted for completeness.)
The move body of update_filename_and_move_files is extracted into a plain
callable that the runner invokes directly. The function is already invoked
directly today for version documents — currently as
update_filename_and_move_files(Document, version_doc)
(handlers.py:670-673), i.e. the whole receiver is called with a synthetic
sender positional arg, bypassing only the @receiver decorator/dispatch, not
the guard-check-then-body split this refactor introduces. Once the move body is
extracted into its own callable (move_files_for_document(instance) per the
implementation plan), this recursive call site must be updated to call that
extracted function directly (move_files_for_document(version_doc)) rather
than the thin wrapper — otherwise recursing into version documents would
re-enter the guard-check wrapper unnecessarily (harmless, since the guard
should be set during a workflow run and unset otherwise, but pointless
indirection). The thin post_save/m2m_changed receivers remain as a
guard-checking wrapper around the extracted callable.
The two post_save receivers on Document are update_filename_and_move_files
(handlers.py:439) and update_llm_suggestions_cache (handlers.py:746-747).
The ContextVar guard suppresses only the former —
update_llm_suggestions_cache keeps running normally, as do
document_consumption_finished receivers such as
add_or_update_document_in_llm_index (which is not a post_save receiver).
This is why the guard is preferred over persisting with .update(), which would
silently suppress all post_save receivers including
update_llm_suggestions_cache.
WorkflowRun.objects.create(...) is created per matching workflow as today
(handlers.py:1039-1043); it is a separate model and is not deferred.
The comment at handlers.py:1019-1027 (added by the pre-existing update_fields
fix, predating this design) is updated to describe the new flow (per-workflow
save under the guard; single explicit rename afterward) but the filename /
archive_filename exclusion it documents is kept — see point 2 above.
Testing
- Runner loop — exercised against a fake context implementing the
WorkflowRunContextsurface that recordsapply_assignment/apply_removal/persistcalls. No DB document, no staged files, no signals. - Concrete contexts —
ConsumptionContextandPersistedContexteach get focused tests: given an action, assert the mutation lands on the overrides vs. the document, and thatsource_fileresolves to the staged vs. final path. - ContextVar guard — assert
update_filename_and_move_filesearly-returns while the guard is set, and that the rename runs exactly once afterpersist(). - Regression: the racy case is already covered, not newly needed. The
scenario this design originally asked for a new test for — a workflow that
reassigns metadata (tags + correspondent/storage path) while the document is
subject to a filename template, asserting final DB filename and file location
stay consistent — is already exercised by
test_document_updated_workflow_assignment_storage_path_persists_with_tag_assignmentintest_workflows.py(added by #13178). No new regression test is required for this; the refactor's job is to keep it passing unchanged. - Regression safety net — the existing
test_workflows.pysuite (~100+ tests; manydocument_consumption_finished.sendsites plus many directrun_workflows(...)calls for theDOCUMENT_UPDATEDpath) must stay green unchanged. A test that needs editing signals a behavior change to flag explicitly, not a silent refactor outcome.
Per project conventions: tests grouped under classes, fixtures and test signatures fully type-annotated.
Implementation sequence
Each step is independently reviewable and keeps the test suite green:
- Introduce the
Protocol+ the two contexts;run_workflowsdelegates to them. Pure refactor, no behavior change. - Move the staged path into
PersistedContextconstruction (passed byrun_workflows_added); delete theoriginal_file/caller_supplied_original_fileparameter plumbing throughrun_workflowsand theexecute_*helpers. - Extract the move body from
update_filename_and_move_filesinto a callable (updating the version-document recursive call site athandlers.py:670-673to call it directly); add theContextVarguard;run_workflowsinvokes the move once after the run completes. Thefilename/archive_filenameexclusion in the per-workflow save is kept; only the comment athandlers.py:1019-1027is updated to describe the new flow.
Pain points addressed
- Dual-mode → eliminated by the
Protocol+ two contexts; nouse_overrides. Still an open, unfixed problem in current code — this is the refactor's main remaining justification. - File staging →
source_fileis a context property; side-channel args deleted. Still an open, unfixed problem in current code. - Rename race → per-workflow save under a
ContextVarguard that suppresses the mid-workflow rename; a single explicit rename runs once at the end against final state. No longer an active bug — #12389 and #13178 independently closed the two concrete failure modes (cross-process already-moved file; intra-workflow tag/m2m clobbering unsaved fields) by narrower means. The guard is now valuable as a single general mechanism replacing two independent point-fixes, not as a bug fix.