Minor updates from line drifts

This commit is contained in:
stumpylog committed 2026-08-18 09:35:43 -07:00
1 parent 04a703029c
commit decaaf2f1b
10 files changed
+53 -51

No files matched your search

@@ -133,7 +133,7 @@ Individual failures are logged and counted but do not abort the run. Bidirection
| `src/documents/signals/handlers.py` | `shutil.move()` → `storage.move()`; remove `create_source_path_directory` / `delete_empty_directories` callsites |
| `src/documents/tasks.py` | Same as signals |
| `src/documents/file_handling.py` | `exists()` checks and directory references use storage API |
| `src/documents/views/` | File-serving views use `storage.open()` within context; wrap for `FileResponse` lifecycle |
| `src/documents/views.py` | File-serving views use `storage.open()` within context; wrap for `FileResponse` lifecycle |
| `src/documents/management/commands/document_importer.py` | Replace `Path.glob()` and direct copies with storage API |
| `src/documents/management/commands/document_exporter.py` | Replace direct file copies and `FileLock`-guarded writes with storage API |
@@ -24,7 +24,7 @@ hard-to-fix bugs (see the revert/refix history around password removal: #12803,
`DOCUMENT_ADDED` workflow fires from `run_workflows_added`, which runs while
the consumer is still inside its transaction — _before_ the consumed file is
copied to `document.source_path` (`document_consumption_finished` is sent at
`consumer.py:658`; the file copy happens after, at `consumer.py:670+`). The
`consumer.py:654`; the file copy happens after, at `consumer.py:666+`). The
staged path is therefore threaded through as `original_file` /
`caller_supplied_original_file` parameters. Actions that read the file
(password removal, email attachments) depend on this plumbing being correct.
@@ -6,7 +6,7 @@
`docs/superpowers/done/specs/2026-06-14-search-query-translation-design.md`.
**Builds on:** the `SearchQueryError(ValueError)` base in
`documents/search/_translate.py` and the single `except SearchQueryError` handler
in `UnifiedSearchViewSet.list` (`documents/views.py:2477`), which re-raises as DRF
in `UnifiedSearchViewSet.list` (`documents/views.py:2612`), which re-raises as DRF
`ValidationError({"query": [msg]})`. Any new subclass surfaces through that one
handler automatically, so this work is purely additive.
@@ -15,7 +15,7 @@ handler automatically, so this work is purely additive.
Every advanced-search failure other than the now-handled invalid date lands in
the view's generic `except Exception` and returns
`HttpResponseBadRequest("Error listing search results, check logs for more
detail.")` (`views.py:2479-2482`). `index.parse_query(...)` runs _outside_ the
detail.")` (`views.py:2617-2621`). `index.parse_query(...)` runs _outside_ the
`translate_query` try/except in `parse_user_query` (`_query.py:220-235`), so
anything Tantivy rejects bypasses `SearchQueryError` entirely and gets the
unhelpful generic 400. Some Tantivy errors also leak Rust internals (e.g.
@@ -131,7 +131,9 @@ Both postdate the `0.26.0` wheel.
- Tantivy side (does a translated string parse?): build a real index via
`documents.search._schema.build_schema` + `register_tokenizers`, then
`index.parse_query(translate_query(q, tz), DEFAULT_SEARCH_FIELDS, field_boosts=…)`.
- Whoosh side (what did v2 do?): the old `get_schema()` + `MultifieldParser([...]) +
DateParserPlugin(...)` still exists on `main` (`src/documents/index.py`); run a query
through it to get the ground-truth `Query`.
- Whoosh side (what did v2 do?): `src/documents/index.py` (the old `get_schema()` +
`MultifieldParser([...]) + DateParserPlugin(...)`) was deleted from `main` in
`aed9abe48` (#12471, 2026-04-02); check it out at `git show aed9abe48^:src/documents/index.py`
(or `git checkout aed9abe48^ -- src/documents/index.py`) and run a query through it
to get the ground-truth `Query`.
- A fuller empirical gap matrix lives in `SEARCH_TANTIVY_WHOOSH_COMPAT.md`.
@@ -11,11 +11,11 @@ Every document that enters paperless converges on one operation: build a
and dispatch the `consume_file` Celery task with a `trigger_source` header. That
operation is hand-rolled at **five** sites today, plus a sixth internal one:
- consume-folder watcher — `document_consumer.py:342`
- API upload + Web UI — `views.py:3181` (one endpoint, two `DocumentSource` values)
- document-version upload — `views.py:1964`
- mail attachment — `mail.py:899`
- mail `.eml` whole-message — `mail.py:987`
- consume-folder watcher — `document_consumer.py:346`
- API upload + Web UI — `views.py:3327` (one endpoint, two `DocumentSource` values)
- document-version upload — `views.py:2086`
- mail attachment — `mail.py:993`
- mail `.eml` whole-message — `mail.py:1084`
- barcode split children (internal re-enqueue) — `barcodes.py:190`/`227`
The duplication causes three concrete problems:
@@ -29,10 +29,10 @@ The duplication causes three concrete problems:
2. **A scratch leak from split staging/cleanup ownership.** Staged sources create
scratch input under `SCRATCH_DIR` that nothing ever fully removes:
`ConsumerPlugin` unlinks only the input **file**, and only on the success path
(`consumer.py:742`). The exact leak shape varies by site — mail attachments and
(`consumer.py:738`). The exact leak shape varies by site — mail attachments and
API/version use `mkdtemp` + a file inside, so the **directory** is orphaned
(empty after success, dir-with-file on failure); the mail `.eml` path uses
`mkstemp` (`mail.py:~955`), so it leaks a **file** directly in `SCRATCH_DIR` on
`mkstemp` (`mail.py:~1034`), so it leaks a **file** directly in `SCRATCH_DIR` on
failure. Either way there is no owner that removes the staged input on every
terminal path.
@@ -46,7 +46,7 @@ The duplication causes three concrete problems:
Separately, the consumption task already has **two** working temp directories that
duplicate each other: `consume_file` opens one `TemporaryDirectory` and passes it
to every plugin (`tasks.py:220`), but `ConsumerPlugin` ignores that and opens its
_own_ second `TemporaryDirectory` (`consumer.py:417`).
_own_ second `TemporaryDirectory` (`consumer.py:408`).
## Goal
@@ -194,7 +194,7 @@ with a derived work_root:
The per-task working directory passed to plugins becomes a **subfolder of
work_root**, and `ConsumerPlugin` uses that handed-in directory for its working
copy instead of opening its own second `TemporaryDirectory` (`consumer.py:417`).
copy instead of opening its own second `TemporaryDirectory` (`consumer.py:408`).
One tree per document; one cleanup.
### Barcode split children (`barcodes.py`)
@@ -216,7 +216,7 @@ independently cleanable when the parent stops.
Mail is the one source that does **not** dispatch per file: `_handle_message`
collects N attachment signatures (and optionally the `.eml` signature), then
`queue_consumption_tasks` wraps them in a single `chord(...).delay()` _after_ the
loop (`mail.py:919`). A per-file `release()` is therefore wrong — if `release()`
loop (`mail.py:1014`). A per-file `release()` is therefore wrong — if `release()`
ran per attachment and the later chord dispatch threw, every staged file would be
orphaned, reopening the leak. **The ownership boundary is the whole message:**
@@ -237,7 +237,7 @@ def _handle_message(...):
`queue_consumption_tasks` itself is unchanged. `build_consume_signature` **must
pass `input_doc`/`overrides` as keyword args** (`consume_file.s(input_doc=...,
overrides=...)`) so the resulting `Signature.kwargs` keeps the shape mail tests
assert on (`test_mail.py:365-366`).
assert on (`test_mail.py:389-390`).
### Call-site refactor (the external sites)
@@ -298,7 +298,7 @@ consume_file task (async, later)
extend to it; (b) the plan must verify the move-precedes-stop ordering, since it
is load-bearing for the cleanup rule.
- **`ConsumerPlugin`'s own cleanup becomes partly redundant.** On success it
unlinks `original_file` and `working_copy` (`consumer.py:742/744`), both of
unlinks `original_file` and `working_copy` (`consumer.py:738/740`), both of
which now live inside work_root that the task `finally` `rmtree`s. The redundant
unlinks are harmless but the plan should remove them for clarity, while keeping
the qpdf `--replace-input` recovery (`unmodified_original`, `consumer.py:452+`)
@@ -11,7 +11,7 @@ The archive-wide AI chat (`ChatStreamingView` with no `document_id`, backed by
purely via dense-vector similarity search: `VectorIndexRetriever` embeds the
user's question and does cosine-similarity nearest-neighbor search over chunk
embeddings, with a hardcoded `similarity_top_k` of 5 (`CHAT_RETRIEVER_TOP_K`,
`chat.py:20`) and no similarity cutoff.
`chat.py:22`) and no similarity cutoff.
Dense embeddings are known to perform poorly on exact keyword, rare/foreign
word, and numeric-string matching (e.g. a compound German word like
@@ -60,9 +60,9 @@ already computed for vector search wherever possible.
caller's permitted document IDs.
2. Call `documents.search.get_backend().search_ids(query_str, user=user,
search_mode=SearchMode.TEXT, limit=CHAT_LEXICAL_TOP_K)` — the same idiom
already used in `views.py:3522` — to get lexical document-ID matches.
already used in `views.py:3588` — to get lexical document-ID matches.
`user` is `None` for superusers and `request.user` otherwise, matching the
existing permission pattern (`views.py:3521`). Intersect the returned IDs
existing permission pattern (`views.py:3587`). Intersect the returned IDs
with the caller-provided `documents` set so results never exceed what the
caller already permission-scoped (this is what makes the retriever safe
to use for both the archive-wide and single-document cases).
@@ -72,7 +72,7 @@ search_mode=SearchMode.TEXT, limit=CHAT_LEXICAL_TOP_K)` — the same idiom
and a metadata filter restricted to that one document ID, reusing the
same query embedding. This is deliberate: `PaperlessSqliteVecVectorStore
.query()` runs a single global `vec0` KNN search over the WHERE-filtered
rows (`vector_store.py:409-434`) — it does not partition top-k per
rows (`vector_store.py:491-527`) — it does not partition top-k per
document — so a single batched call across N lexical-hit documents with
`top_k=N` could return several chunks from one document and none from
another. Per-document calls are the only way to guarantee each lexical
@@ -122,7 +122,7 @@ search_mode=SearchMode.TEXT, limit=CHAT_LEXICAL_TOP_K)` — the same idiom
internally exactly as today and used as step 1 of the hybrid flow.
- **Changed:** `stream_chat_with_documents()` / `_stream_chat_with_documents()`
gain a `user` parameter, threaded from `ChatStreamingView.post`
(`views.py:2268`) using the same `None`-for-superuser convention already
(`views.py:2303`) using the same `None`-for-superuser convention already
used elsewhere in `views.py`.
### Scope: applies to both chat modes
@@ -140,7 +140,7 @@ that vector similarity alone might miss.
reason, the lexical step should degrade gracefully to vector-only results
(log and continue) rather than failing the whole chat response — chat
already wraps everything in a try/except at the `stream_chat_with_documents`
level (`chat.py:82-87`), but the lexical addition should not, by itself,
level (`chat.py:138-146`), but the lexical addition should not, by itself,
turn a previously-working vector-only answer into an error.
### Testing