Updates to these ideas

This commit is contained in:
stumpylog
2026-09-29 13:09:58 -07:00
parent cfe05a7e03
commit c8a00a05a8
3 changed files with 333 additions and 145 deletions
@@ -1,42 +1,110 @@
# Workflow Runner Refactor — Design
**Date:** 2026-05-19
**Updated:** 2026-09-24 — see "Status update (2026-09-24)" below; line references
throughout refreshed against current `dev`.
**Branch base:** `dev`
**Status:** Approved design, pending implementation plan
**Status:** Approved design, pending implementation plan. Motivating bugs are now
independently fixed (see status update) — this is a structural cleanup, not an
active-bug fix.
## Status update (2026-09-24)
The acute bugs this design was written against have since been closed by two
targeted, already-merged fixes that do **not** use the `WorkflowRunContext` /
`ContextVar` approach below:
- **#12386** (cross-process "file already moved" race, cited in cause 3) was
fixed by **#12389** ("Fix: avoid moving files if already moved", merged in
v2.20.12): `validate_move` in `update_filename_and_move_files` now detects,
via a checksum comparison (`_path_matches_checksum`), when the target file
already exists because a concurrent save already moved it, and recovers the
DB pointer instead of raising `CannotMoveFilesException`. This is a different
mechanism from anything proposed here and stays as-is; nothing in this
refactor should touch `_path_matches_checksum` or its call sites.
- **The intra-workflow tag/`m2m_changed` race** described in cause 3 and design
§4 (`add_nested_tags` firing `m2m_changed` → `refresh_from_db()` → wiping an
earlier-ordered action's unsaved `correspondent`/`storage_path`) was fixed by
**#13178** ("Fix: prevent tag assignment from reverting other pending
workflow assignments", merged 2026-07-20 — _after_ this design was written).
`apply_assignment_to_document` (`mutations.py`) now applies tag changes to a
**freshly-fetched** `Document` instance rather than the shared in-memory one,
matching the pattern `apply_removal_to_document` already used for tag
removal. Because the mid-run rename this triggers now reads the _old_,
still-current DB state (nothing else has been saved yet), it recomputes the
same path and is a no-op; the real rename happens once, correctly, at the
final `document.save()`. Regression coverage:
`test_document_updated_workflow_assignment_storage_path_persists_with_tag_assignment`
in `test_workflows.py` (the only test #13178 itself added). A related,
earlier fix (#12664) covers a similar but distinct case — tag removal
alongside a title assignment — via
`test_document_updated_workflow_assignment_persists_when_removing_trigger_tag`.
- The `filename`/`archive_filename` exclusion from `update_fields` (cause 3,
design §4 point 2) was already in place before this design was written (added
by #12390-adjacent work) and is unaffected — still correctly attributed here
as a load-bearing cross-process guard, not duct tape.
- **#13178's own PR description is the origin of this idea**: "Once the beta
is out, I do have some thoughts about using ContextVar to delete this class
entirely. We've run into it plenty of times." So the guard concept predates
this design; the acute pain it was meant to address has since been patched
piecemeal instead.
**What this means for scope:** causes 1 and 2 below (dual-mode branching,
staged-file parameter plumbing) are unchanged and fully present in current code
— this refactor is still worth doing for them. Cause 3's race is no longer an
_active_ bug; the `ContextVar` guard is now defense-in-depth / simplification
(it also collapses two independent workarounds — the fresh-instance tag fetch
and the per-workflow `update_fields` restriction — into one clearer mechanism)
rather than a fix for a reproducible failure. Treat any "fixes a bug" framing
below as historical motivation, not a current defect claim.
## Problem
Workflow execution and the Django signal layer have repeatedly produced fragile,
hard-to-fix bugs (see the revert/refix history around password removal: #12803,
#12814, #12716, and the filename race #12386). Three structural causes:
#12814, #12716, and the filename race #12386, now closed — see status update
above). Three structural causes:
1. **`run_workflows` is dual-mode.** A single function handles both consumption
(mutating a `DocumentMetadataOverrides`) and post-save (mutating a real
`Document`), branching on a `use_overrides` flag. The branching is
concentrated in two places — the action dispatch inside `run_workflows`
(`handlers.py:931-1001`) and `build_workflow_action_context`
(`actions.py:33-83`), each with two full code paths. The `apply_*` helpers in
`workflows/mutations.py` are _already_ split by target type
(`apply_assignment_to_document` vs `apply_assignment_to_overrides`, etc.); the
refactor unifies their callers, not the helpers themselves.
(`handlers.py:938-1014`, `use_overrides` first set at `handlers.py:881`) and
`build_workflow_action_context` (`actions.py:33-83`), each with two full code
paths. The `apply_*` helpers in `workflows/mutations.py` are _already_ split
by target type (`apply_assignment_to_document` vs
`apply_assignment_to_overrides`, etc.); the refactor unifies their callers,
not the helpers themselves. **Still fully present in current code — unchanged
by any fix since this design was written.**
2. **File location is an implicit, timing-dependent side channel.** The
`DOCUMENT_ADDED` workflow fires from `run_workflows_added`, which runs while
the consumer is still inside its transaction — _before_ the consumed file is
copied to `document.source_path` (`document_consumption_finished` is sent at
`consumer.py:654`; the file copy happens after, at `consumer.py:666+`). The
staged path is therefore threaded through as `original_file` /
`caller_supplied_original_file` parameters. Actions that read the file
(password removal, email attachments) depend on this plumbing being correct.
`DOCUMENT_ADDED` workflow fires from `run_workflows_added`
(`handlers.py:811-824`), which runs while the consumer is still inside its
transaction — _before_ the consumed file is copied to `document.source_path`
(`document_consumption_finished` is sent from `consumer.py`, the file copy
happens after). The staged path is therefore threaded through as
`original_file` / `caller_supplied_original_file` parameters
(`handlers.py:868,890-898,976-979`). Actions that read the file (password
removal, email attachments) depend on this plumbing being correct. **Still
fully present in current code.**
3. **The workflow run races the filename rename.** `update_filename_and_move_files`
is a raw `post_save` receiver on `Document`. When a workflow persists its
changes via `document.save(update_fields=[...])`, that save fires `post_save`
and runs the rename _while the workflow is still executing_. Under concurrent
Celery/UI updates the interleaved `refresh_from_db()` calls corrupt state. The
comment at `handlers.py:980-984` — deliberately excluding `filename` /
`archive_filename` from the workflow save — is a load-bearing workaround for
exactly this.
3. **The workflow run could race the filename rename — now mitigated by two
independent, narrower fixes (see "Status update" above), not eliminated
structurally.** `update_filename_and_move_files` (`handlers.py:437-439`
decorators, body `440-673`) is a raw `post_save`/`m2m_changed` receiver. When
a workflow persists its changes via `document.save(update_fields=[...])`, or
when `apply_assignment_to_document` mutates tags via a freshly-fetched
instance (`mutations.py:26-31`), that write fires the receiver _while the
workflow is still executing_. As of #13178 this no longer corrupts in-memory
state (tags are applied to a separate instance, so the mid-run rename it
triggers reads only already-committed data and is a no-op or self-consistent
move); as of #12389 a genuine cross-process "someone else already moved this
file" race recovers via checksum comparison instead of erroring. The comment
documenting the `filename`/`archive_filename` exclusion from the workflow's
own `update_fields` (`handlers.py:1019-1027`) remains a load-bearing guard
against a _different_, still-real cross-process hazard (an in-memory
`document.filename` going stale while another process moves the file) and is
unaffected by either fix above.
Note: `run_workflows_added` / `run_workflows_updated` are connected to the
_custom_ signals `document_consumption_finished` / `document_updated`, fired
@@ -155,12 +223,23 @@ let one workflow's refresh wipe a prior workflow's in-memory changes. So:
3. The rename is suppressed for the whole run and invoked **exactly once,
afterward**, against final committed state.
The actual race being fixed: `apply_assignment_to_document` assigns tags via
`document.add_nested_tags(...)`, which fires `m2m_changed` on
`Document.tags.through` _before_ the workflow's `document.save()`. The
`m2m_changed` receiver `update_filename_and_move_files` then calls
`refresh_from_db()`, wiping the workflow's in-memory correspondent/type, and
moves the file to a path computed from stale metadata. The guard prevents this.
**Historical race, already fixed by other means (see "Status update"):**
`apply_assignment_to_document` used to assign tags via
`document.add_nested_tags(...)` directly on the shared in-memory `document`,
which fired `m2m_changed` on `Document.tags.through` _before_ the workflow's
`document.save()`; the `m2m_changed` receiver `update_filename_and_move_files`
then called `refresh_from_db()` on that shared instance, wiping the workflow's
in-memory correspondent/type, and moved the file to a path computed from stale
metadata. #13178 fixed this by having `apply_assignment_to_document` (and
`apply_removal_to_document`, which already did this) mutate tags on a
**freshly-fetched** `Document.objects.get(pk=document.pk)` instead of the
shared one (`mutations.py:26-31`), so the `refresh_from_db()` triggered by
`m2m_changed` no longer touches the workflow's unsaved in-memory fields. The
guard below is not needed to fix that specific corruption anymore — it instead
gives a single, general mechanism that supersedes the fresh-instance-fetch
workaround (and the equivalent one already in `apply_removal_to_document`),
rather than requiring every future mutation path to remember to fetch a
separate instance.
To stop the rename from firing mid-workflow, a **`ContextVar` guard** is
introduced (e.g. `documents/workflows/context.py` module-level
@@ -173,8 +252,16 @@ reentrancy-safe for nested saves or nested workflow runs.
The guard must span the whole execution, not just `persist()`, because
`update_filename_and_move_files` is _also_ registered to `m2m_changed` on
`Document.tags.through` and to `post_save` on `CustomFieldInstance`
(`handlers.py:431-432`). A workflow action that assigns tags or custom fields
would otherwise trigger a rename mid-workflow through those signals.
(`handlers.py:437-438`). A workflow action that assigns tags or custom fields
would otherwise trigger a rename mid-workflow through those signals. If the
`ContextVar` guard lands, the fresh-instance-fetch workaround in
`apply_assignment_to_document`/`apply_removal_to_document` becomes redundant
but is not itself incorrect — decide during implementation whether to simplify
those two call sites back to mutating `document` directly now that the guard
covers the hazard, or leave them as extra defense-in-depth. Note either way:
`document` passed to `add_nested_tags`/`tags.clear`/`tags.remove` must still be
re-fetched or `refresh_from_db()`'d for the tags relation to reflect the
change on the in-memory instance used later in the same action.
After execution completes, `run_workflows` calls `persist()` once and then
explicitly invokes the move logic once. The `ContextVar` is set/reset in the
@@ -184,26 +271,37 @@ pools are also `contextvars`-aware — non-issues, noted for completeness.)
The move body of `update_filename_and_move_files` is extracted into a plain
callable that the runner invokes directly. The function is already invoked
directly (as a plain call, bypassing the decorator) for version documents at
`handlers.py:664-667`, so this extraction has precedent. The thin `post_save`
receiver remains as a guard-checking wrapper.
directly today for version documents — currently as
`update_filename_and_move_files(Document, version_doc)`
(`handlers.py:670-673`), i.e. the _whole receiver_ is called with a synthetic
`sender` positional arg, bypassing only the `@receiver` decorator/dispatch, not
the guard-check-then-body split this refactor introduces. Once the move body is
extracted into its own callable (`move_files_for_document(instance)` per the
implementation plan), this recursive call site must be updated to call that
extracted function directly (`move_files_for_document(version_doc)`) rather
than the thin wrapper — otherwise recursing into version documents would
re-enter the guard-check wrapper unnecessarily (harmless, since the guard
should be set during a workflow run and unset otherwise, but pointless
indirection). The thin `post_save`/`m2m_changed` receivers remain as a
guard-checking wrapper around the extracted callable.
The two `post_save` receivers on `Document` are `update_filename_and_move_files`
(`handlers.py:433`) and `update_llm_suggestions_cache` (`handlers.py:740`). The
`ContextVar` guard suppresses **only** the former — `update_llm_suggestions_cache`
keeps running normally, as do `document_consumption_finished` receivers such as
(`handlers.py:439`) and `update_llm_suggestions_cache` (`handlers.py:746-747`).
The `ContextVar` guard suppresses **only** the former —
`update_llm_suggestions_cache` keeps running normally, as do
`document_consumption_finished` receivers such as
`add_or_update_document_in_llm_index` (which is _not_ a `post_save` receiver).
This is why the guard is preferred over persisting with `.update()`, which would
silently suppress _all_ `post_save` receivers including
`update_llm_suggestions_cache`.
`WorkflowRun.objects.create(...)` is created per matching workflow as today
(`handlers.py:998-1002`); it is a separate model and is not deferred.
(`handlers.py:1039-1043`); it is a separate model and is not deferred.
The comment at `handlers.py:980-984` is updated to describe the new flow
(per-workflow save under the guard; single explicit rename afterward) but the
`filename` / `archive_filename` exclusion it documents is kept — see point 2
above.
The comment at `handlers.py:1019-1027` (added by the pre-existing `update_fields`
fix, predating this design) is updated to describe the new flow (per-workflow
save under the guard; single explicit rename afterward) but the `filename` /
`archive_filename` exclusion it documents is kept — see point 2 above.
## Testing
@@ -216,11 +314,16 @@ above.
- **ContextVar guard** — assert `update_filename_and_move_files` early-returns
while the guard is set, and that the rename runs exactly once after
`persist()`.
- **Regression: the racy case** — a workflow that reassigns metadata while the
document is subject to a filename template; assert final DB filename and file
location are consistent (the #12386 scenario).
- **Regression safety net** — the existing `test_workflows.py` suite (~100
tests; ~19 `document_consumption_finished.send` sites plus many direct
- **Regression: the racy case is already covered, not newly needed.** The
scenario this design originally asked for a new test for — a workflow that
reassigns metadata (tags + correspondent/storage path) while the document is
subject to a filename template, asserting final DB filename and file location
stay consistent — is already exercised by
`test_document_updated_workflow_assignment_storage_path_persists_with_tag_assignment`
in `test_workflows.py` (added by #13178). No new regression test is required
for this; the refactor's job is to keep it passing **unchanged**.
- **Regression safety net** — the existing `test_workflows.py` suite (~100+
tests; many `document_consumption_finished.send` sites plus many direct
`run_workflows(...)` calls for the `DOCUMENT_UPDATED` path) must stay green
**unchanged**. A test that needs editing signals a behavior change to flag
explicitly, not a silent refactor outcome.
@@ -238,16 +341,24 @@ Each step is independently reviewable and keeps the test suite green:
`run_workflows_added`); delete the `original_file` /
`caller_supplied_original_file` parameter plumbing through `run_workflows`
and the `execute_*` helpers.
3. Extract the move body from `update_filename_and_move_files` into a callable;
add the `ContextVar` guard; `run_workflows` invokes the move once after the
run completes. The `filename` / `archive_filename` exclusion in the
per-workflow save is kept; only the comment at `handlers.py:980-984` is
updated to describe the new flow.
3. Extract the move body from `update_filename_and_move_files` into a callable
(updating the version-document recursive call site at `handlers.py:670-673`
to call it directly); add the `ContextVar` guard; `run_workflows` invokes the
move once after the run completes. The `filename` / `archive_filename`
exclusion in the per-workflow save is kept; only the comment at
`handlers.py:1019-1027` is updated to describe the new flow.
## Pain points addressed
- **Dual-mode** → eliminated by the `Protocol` + two contexts; no `use_overrides`.
- **File staging** → `source_file` is a context property; side-channel args deleted.
Still an open, unfixed problem in current code — this is the refactor's main
remaining justification.
- **File staging** → `source_file` is a context property; side-channel args
deleted. Still an open, unfixed problem in current code.
- **Rename race** → per-workflow save under a `ContextVar` guard that suppresses
the mid-workflow rename; a single explicit rename runs once at the end against
final state.
final state. **No longer an active bug** — #12389 and #13178 independently
closed the two concrete failure modes (cross-process already-moved file;
intra-workflow tag/m2m clobbering unsaved fields) by narrower means. The
guard is now valuable as a single general mechanism replacing two
independent point-fixes, not as a bug fix.