mirror of
https://github.com/paperless-ngx/paperless-ngx.git
synced 2026-10-08 17:17:14 +00:00
Minor updates from line drifts
This commit is contained in:
1 parent
04a703029c
commit
decaaf2f1b
10 files changed
+53
-51
No files matched your search
@@ -133,7 +133,7 @@ Individual failures are logged and counted but do not abort the run. Bidirection
|
||||
| `src/documents/signals/handlers.py` | `shutil.move()` → `storage.move()`; remove `create_source_path_directory` / `delete_empty_directories` callsites |
|
||||
| `src/documents/tasks.py` | Same as signals |
|
||||
| `src/documents/file_handling.py` | `exists()` checks and directory references use storage API |
|
||||
| `src/documents/views/` | File-serving views use `storage.open()` within context; wrap for `FileResponse` lifecycle |
|
||||
| `src/documents/views.py` | File-serving views use `storage.open()` within context; wrap for `FileResponse` lifecycle |
|
||||
| `src/documents/management/commands/document_importer.py` | Replace `Path.glob()` and direct copies with storage API |
|
||||
| `src/documents/management/commands/document_exporter.py` | Replace direct file copies and `FileLock`-guarded writes with storage API |
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ hard-to-fix bugs (see the revert/refix history around password removal: #12803,
|
||||
`DOCUMENT_ADDED` workflow fires from `run_workflows_added`, which runs while
|
||||
the consumer is still inside its transaction — _before_ the consumed file is
|
||||
copied to `document.source_path` (`document_consumption_finished` is sent at
|
||||
`consumer.py:658`; the file copy happens after, at `consumer.py:670+`). The
|
||||
`consumer.py:654`; the file copy happens after, at `consumer.py:666+`). The
|
||||
staged path is therefore threaded through as `original_file` /
|
||||
`caller_supplied_original_file` parameters. Actions that read the file
|
||||
(password removal, email attachments) depend on this plumbing being correct.
|
||||
|
||||
@@ -6,7 +6,7 @@
|
||||
`docs/superpowers/done/specs/2026-06-14-search-query-translation-design.md`.
|
||||
**Builds on:** the `SearchQueryError(ValueError)` base in
|
||||
`documents/search/_translate.py` and the single `except SearchQueryError` handler
|
||||
in `UnifiedSearchViewSet.list` (`documents/views.py:2477`), which re-raises as DRF
|
||||
in `UnifiedSearchViewSet.list` (`documents/views.py:2612`), which re-raises as DRF
|
||||
`ValidationError({"query": [msg]})`. Any new subclass surfaces through that one
|
||||
handler automatically, so this work is purely additive.
|
||||
|
||||
@@ -15,7 +15,7 @@ handler automatically, so this work is purely additive.
|
||||
Every advanced-search failure other than the now-handled invalid date lands in
|
||||
the view's generic `except Exception` and returns
|
||||
`HttpResponseBadRequest("Error listing search results, check logs for more
|
||||
detail.")` (`views.py:2479-2482`). `index.parse_query(...)` runs _outside_ the
|
||||
detail.")` (`views.py:2617-2621`). `index.parse_query(...)` runs _outside_ the
|
||||
`translate_query` try/except in `parse_user_query` (`_query.py:220-235`), so
|
||||
anything Tantivy rejects bypasses `SearchQueryError` entirely and gets the
|
||||
unhelpful generic 400. Some Tantivy errors also leak Rust internals (e.g.
|
||||
|
||||
@@ -131,7 +131,9 @@ Both postdate the `0.26.0` wheel.
|
||||
- Tantivy side (does a translated string parse?): build a real index via
|
||||
`documents.search._schema.build_schema` + `register_tokenizers`, then
|
||||
`index.parse_query(translate_query(q, tz), DEFAULT_SEARCH_FIELDS, field_boosts=…)`.
|
||||
- Whoosh side (what did v2 do?): the old `get_schema()` + `MultifieldParser([...]) +
|
||||
DateParserPlugin(...)` still exists on `main` (`src/documents/index.py`); run a query
|
||||
through it to get the ground-truth `Query`.
|
||||
- Whoosh side (what did v2 do?): `src/documents/index.py` (the old `get_schema()` +
|
||||
`MultifieldParser([...]) + DateParserPlugin(...)`) was deleted from `main` in
|
||||
`aed9abe48` (#12471, 2026-04-02); check it out at `git show aed9abe48^:src/documents/index.py`
|
||||
(or `git checkout aed9abe48^ -- src/documents/index.py`) and run a query through it
|
||||
to get the ground-truth `Query`.
|
||||
- A fuller empirical gap matrix lives in `SEARCH_TANTIVY_WHOOSH_COMPAT.md`.
|
||||
@@ -11,11 +11,11 @@ Every document that enters paperless converges on one operation: build a
|
||||
and dispatch the `consume_file` Celery task with a `trigger_source` header. That
|
||||
operation is hand-rolled at **five** sites today, plus a sixth internal one:
|
||||
|
||||
- consume-folder watcher — `document_consumer.py:342`
|
||||
- API upload + Web UI — `views.py:3181` (one endpoint, two `DocumentSource` values)
|
||||
- document-version upload — `views.py:1964`
|
||||
- mail attachment — `mail.py:899`
|
||||
- mail `.eml` whole-message — `mail.py:987`
|
||||
- consume-folder watcher — `document_consumer.py:346`
|
||||
- API upload + Web UI — `views.py:3327` (one endpoint, two `DocumentSource` values)
|
||||
- document-version upload — `views.py:2086`
|
||||
- mail attachment — `mail.py:993`
|
||||
- mail `.eml` whole-message — `mail.py:1084`
|
||||
- barcode split children (internal re-enqueue) — `barcodes.py:190`/`227`
|
||||
|
||||
The duplication causes three concrete problems:
|
||||
@@ -29,10 +29,10 @@ The duplication causes three concrete problems:
|
||||
2. **A scratch leak from split staging/cleanup ownership.** Staged sources create
|
||||
scratch input under `SCRATCH_DIR` that nothing ever fully removes:
|
||||
`ConsumerPlugin` unlinks only the input **file**, and only on the success path
|
||||
(`consumer.py:742`). The exact leak shape varies by site — mail attachments and
|
||||
(`consumer.py:738`). The exact leak shape varies by site — mail attachments and
|
||||
API/version use `mkdtemp` + a file inside, so the **directory** is orphaned
|
||||
(empty after success, dir-with-file on failure); the mail `.eml` path uses
|
||||
`mkstemp` (`mail.py:~955`), so it leaks a **file** directly in `SCRATCH_DIR` on
|
||||
`mkstemp` (`mail.py:~1034`), so it leaks a **file** directly in `SCRATCH_DIR` on
|
||||
failure. Either way there is no owner that removes the staged input on every
|
||||
terminal path.
|
||||
|
||||
@@ -46,7 +46,7 @@ The duplication causes three concrete problems:
|
||||
Separately, the consumption task already has **two** working temp directories that
|
||||
duplicate each other: `consume_file` opens one `TemporaryDirectory` and passes it
|
||||
to every plugin (`tasks.py:220`), but `ConsumerPlugin` ignores that and opens its
|
||||
_own_ second `TemporaryDirectory` (`consumer.py:417`).
|
||||
_own_ second `TemporaryDirectory` (`consumer.py:408`).
|
||||
|
||||
## Goal
|
||||
|
||||
@@ -194,7 +194,7 @@ with a derived work_root:
|
||||
|
||||
The per-task working directory passed to plugins becomes a **subfolder of
|
||||
work_root**, and `ConsumerPlugin` uses that handed-in directory for its working
|
||||
copy instead of opening its own second `TemporaryDirectory` (`consumer.py:417`).
|
||||
copy instead of opening its own second `TemporaryDirectory` (`consumer.py:408`).
|
||||
One tree per document; one cleanup.
|
||||
|
||||
### Barcode split children (`barcodes.py`)
|
||||
@@ -216,7 +216,7 @@ independently cleanable when the parent stops.
|
||||
Mail is the one source that does **not** dispatch per file: `_handle_message`
|
||||
collects N attachment signatures (and optionally the `.eml` signature), then
|
||||
`queue_consumption_tasks` wraps them in a single `chord(...).delay()` _after_ the
|
||||
loop (`mail.py:919`). A per-file `release()` is therefore wrong — if `release()`
|
||||
loop (`mail.py:1014`). A per-file `release()` is therefore wrong — if `release()`
|
||||
ran per attachment and the later chord dispatch threw, every staged file would be
|
||||
orphaned, reopening the leak. **The ownership boundary is the whole message:**
|
||||
|
||||
@@ -237,7 +237,7 @@ def _handle_message(...):
|
||||
`queue_consumption_tasks` itself is unchanged. `build_consume_signature` **must
|
||||
pass `input_doc`/`overrides` as keyword args** (`consume_file.s(input_doc=...,
|
||||
overrides=...)`) so the resulting `Signature.kwargs` keeps the shape mail tests
|
||||
assert on (`test_mail.py:365-366`).
|
||||
assert on (`test_mail.py:389-390`).
|
||||
|
||||
### Call-site refactor (the external sites)
|
||||
|
||||
@@ -298,7 +298,7 @@ consume_file task (async, later)
|
||||
extend to it; (b) the plan must verify the move-precedes-stop ordering, since it
|
||||
is load-bearing for the cleanup rule.
|
||||
- **`ConsumerPlugin`'s own cleanup becomes partly redundant.** On success it
|
||||
unlinks `original_file` and `working_copy` (`consumer.py:742/744`), both of
|
||||
unlinks `original_file` and `working_copy` (`consumer.py:738/740`), both of
|
||||
which now live inside work_root that the task `finally` `rmtree`s. The redundant
|
||||
unlinks are harmless but the plan should remove them for clarity, while keeping
|
||||
the qpdf `--replace-input` recovery (`unmodified_original`, `consumer.py:452+`)
|
||||
|
||||
@@ -11,7 +11,7 @@ The archive-wide AI chat (`ChatStreamingView` with no `document_id`, backed by
|
||||
purely via dense-vector similarity search: `VectorIndexRetriever` embeds the
|
||||
user's question and does cosine-similarity nearest-neighbor search over chunk
|
||||
embeddings, with a hardcoded `similarity_top_k` of 5 (`CHAT_RETRIEVER_TOP_K`,
|
||||
`chat.py:20`) and no similarity cutoff.
|
||||
`chat.py:22`) and no similarity cutoff.
|
||||
|
||||
Dense embeddings are known to perform poorly on exact keyword, rare/foreign
|
||||
word, and numeric-string matching (e.g. a compound German word like
|
||||
@@ -60,9 +60,9 @@ already computed for vector search wherever possible.
|
||||
caller's permitted document IDs.
|
||||
2. Call `documents.search.get_backend().search_ids(query_str, user=user,
|
||||
search_mode=SearchMode.TEXT, limit=CHAT_LEXICAL_TOP_K)` — the same idiom
|
||||
already used in `views.py:3522` — to get lexical document-ID matches.
|
||||
already used in `views.py:3588` — to get lexical document-ID matches.
|
||||
`user` is `None` for superusers and `request.user` otherwise, matching the
|
||||
existing permission pattern (`views.py:3521`). Intersect the returned IDs
|
||||
existing permission pattern (`views.py:3587`). Intersect the returned IDs
|
||||
with the caller-provided `documents` set so results never exceed what the
|
||||
caller already permission-scoped (this is what makes the retriever safe
|
||||
to use for both the archive-wide and single-document cases).
|
||||
@@ -72,7 +72,7 @@ search_mode=SearchMode.TEXT, limit=CHAT_LEXICAL_TOP_K)` — the same idiom
|
||||
and a metadata filter restricted to that one document ID, reusing the
|
||||
same query embedding. This is deliberate: `PaperlessSqliteVecVectorStore
|
||||
.query()` runs a single global `vec0` KNN search over the WHERE-filtered
|
||||
rows (`vector_store.py:409-434`) — it does not partition top-k per
|
||||
rows (`vector_store.py:491-527`) — it does not partition top-k per
|
||||
document — so a single batched call across N lexical-hit documents with
|
||||
`top_k=N` could return several chunks from one document and none from
|
||||
another. Per-document calls are the only way to guarantee each lexical
|
||||
@@ -122,7 +122,7 @@ search_mode=SearchMode.TEXT, limit=CHAT_LEXICAL_TOP_K)` — the same idiom
|
||||
internally exactly as today and used as step 1 of the hybrid flow.
|
||||
- **Changed:** `stream_chat_with_documents()` / `_stream_chat_with_documents()`
|
||||
gain a `user` parameter, threaded from `ChatStreamingView.post`
|
||||
(`views.py:2268`) using the same `None`-for-superuser convention already
|
||||
(`views.py:2303`) using the same `None`-for-superuser convention already
|
||||
used elsewhere in `views.py`.
|
||||
|
||||
### Scope: applies to both chat modes
|
||||
@@ -140,7 +140,7 @@ that vector similarity alone might miss.
|
||||
reason, the lexical step should degrade gracefully to vector-only results
|
||||
(log and continue) rather than failing the whole chat response — chat
|
||||
already wraps everything in a try/except at the `stream_chat_with_documents`
|
||||
level (`chat.py:82-87`), but the lexical addition should not, by itself,
|
||||
level (`chat.py:138-146`), but the lexical addition should not, by itself,
|
||||
turn a previously-working vector-only answer into an error.
|
||||
|
||||
### Testing
|
||||
|
||||
Reference in new issue
Block a user