Compare commits

..
Author SHA1 Message Date
stumpylogandClaude Sonnet 5 3838706194 Fix: freeze m0001's vec0 table shape instead of tracking current schema
_rebuild_into()/_create_vec_table() always reflect whatever the *current*
schema is -- correct for compact() (which only ever runs on an
already-current-schema store), but wrong for a migration that isn't the
latest one anymore. m0001 delegated to them, which worked fine as long as
v2 was current, but silently breaks the moment a v2 -> v3 migration is
added: m0001 would then produce a v3-shaped table (whatever columns
happen to be current) instead of its actual v2 shape, and the next
migration in the chain would find the columns it expects to migrate from
already gone.

Spell out m0001's own v2-shaped CREATE VIRTUAL TABLE and row copy instead,
so it keeps producing the same output forever, independent of later
schema changes -- the same reasoning already documented on the test
helper _copying_apply(), just not, until now, applied to the real
migration.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 13:34:25 -07:00
stumpylogandClaude Sonnet 5 15938945d7 Simplify: dedupe file-swap rebuild, migration-check locking, and test mocks
Prompted by an automated simplification review of the branch. Applies the
highest-value findings, all behavior-preserving (full suite still green):

- vector_store.py: extract _rebuild_file() (temp-file lifecycle: open,
  populate, swap-in-or-discard-on-failure) and _rebuild_into() (create
  table, copy meta, stream rows) out of compact() and the v1->v2
  migration, which had duplicated both almost verbatim. The migration
  module shrinks to a single _rebuild_into() call instead of reaching
  into four PaperlessSqliteVecVectorStore privates.
- vector_store.py: extract _stored_schema_version() out of
  has_pending_migration() and check_and_run_migrations(), which
  duplicated the same schema_version read and its "missing key means
  current" default.
- indexing.py: extract _with_exclusive_access(), collapsing three
  identical _exclude_readers()/Timeout/log-and-skip blocks (compaction
  in update_llm_index() and llm_index_compact(), the migration check in
  _check_and_run_migrations()) to one line each. Also trims
  _check_and_run_migrations()'s docstring, which had grown to restate
  content already documented on has_pending_migration() and
  check_and_run_migrations(), including a claim that was now stale
  (llm_index_migrate() is a second consumer of the re-embed signal).
- test_ai_indexing.py: add a mock_store fixture for the
  write_store()-yields-a-MagicMock pattern that 8 tests were hand-rolling
  identically.
- migrations/__init__.py: consolidate the "how to add a migration"
  instructions to one place instead of three.
- docs/administration.md: fix two now-inaccurate claims -- the bare-metal
  migrate step doesn't run "automatically" (it's the manual step being
  documented), and the self-contradictory "if enabled... no-op if
  disabled" phrasing on the same step.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 12:32:59 -07:00
stumpylogandClaude Sonnet 5 470bcc1751 Refactor: address PR review comments
- Move schema migrations into their own paperless_ai/migrations package,
  one module per migration (mNNNN_description.py -- a leading digit isn't
  a valid Python identifier, unlike Django's own numbered migrations,
  which load via importlib.import_module() rather than a static import
  statement), establishing the pattern before more migrations accumulate
  in vector_store.py.
- Drop the redundant "if the LLM index is enabled..." clause from the
  Docker upgrade note in docs/administration.md -- it's already covered
  by the LLM index section below, and calling out just one of several
  auto-applied migrations there was incomplete.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 11:45:18 -07:00
stumpylogandClaude Sonnet 5 cd7038a7a4 Fix: don't gate document_ids-scoped LLM index updates on Document.modified
bulk_edit.py's tag/correspondent/document_type/storage_path/custom_field
helpers write via queryset.update() or direct M2M/through-model bulk
operations, none of which call Document.save() -- so Document.modified's
auto_now never fires. Comparing against get_modified_times() for a
document_ids-scoped update_llm_index() call was silently skipping reindex
of documents whose embedded metadata had just changed via a bulk edit.
The modified-time comparison is now only used for the unscoped,
full-library incremental scan, where it is still needed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 11:13:05 -07:00
stumpylogandClaude Sonnet 5 df8a6cc715 Feature: run pending LLM index migrations automatically on startup
Add document_llmindex migrate, a cheap check-only path (no reindex) safe
to run unconditionally: has_pending_migration() short-circuits to a
metadata-only read once the store is current, so a healthy install pays
almost nothing. If a pending migration would require re-embedding, it
only logs a warning and leaves the index as-is -- re-embedding can be
slow and, for a metered embedding backend, cost money, so it stays a
deliberate manual action (document_llmindex rebuild), never automatic.

Wire it into the Docker image as a new init-llmindex-migrate oneshot
(modeled on init-search-index, gated on init-migrations), and document
the equivalent manual step for bare-metal upgrades.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 10:48:24 -07:00
stumpylogandClaude Sonnet 5 b21d50645d Fix: run pending structural migrations before delete()/upsert_document()
Only update_llm_index()'s nightly rebuild ran check_and_run_migrations(),
but llm_index_remove_document()/llm_index_add_or_update_document() call
delete()/upsert_document() directly and immediately (on every document
delete/edit). On upgrade, every pre-existing document is "pre-migration"
until the next scheduled rebuild -- up to 24h by default -- so a delete or
edit in that window would find zero document_chunks rows and silently
leave stale/orphaned vec0 rows behind.

Add has_pending_migration(), a cheap metadata-only read that needs no
exclusive access, so normal writes only pay for check_and_run_migrations()'s
exclusive locking when a migration is actually pending -- otherwise they'd
be gated by the same lock compaction uses, which
test_normal_write_is_not_gated_by_the_compaction_lock exists to forbid.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 10:30:04 -07:00
stumpylog 24b86770df Perf: store document_chunks.document_id as INTEGER
document_chunks is a plain SQLite table, not a vec0 virtual table, so
(unlike vec0's own TEXT metadata column) it gets normal type-affinity
coercion between the TEXT document ids used elsewhere in this module and
an INTEGER column -- verified empirically. INTEGER keys make the
per-document delete lookup this table exists for cheaper: smaller index
entries and integer rather than text B-tree comparisons.
2026-07-28 09:39:02 -07:00
stumpylog 6a2972f313 Test: add missing document_chunks coverage and GIVEN/WHEN/THEN docstrings
Cover compact()'s handling of a churned document's final chunk generation,
delete() on a never-indexed document, upsert_document() clearing to an empty
node list, and the v1->v2 migration's embed_model preservation and
idempotency, per repo convention.
2026-07-28 09:36:58 -07:00
Trenton HolmesandClaude Sonnet 5 fa9741b555 Simplify: dedupe row-copy/migration-test helpers, type fixture params
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 08:56:16 -07:00
Trenton HolmesandClaude Sonnet 5 63c9430eec Perf: point-delete vector store chunks by id instead of a document_id scan
vec0 only gets an efficient lookup on a metadata column inside a KNN
query; a plain document_id-filtered DELETE is a full table scan
regardless of index size. Add a document_chunks side table (plain
SQLite, real index) to look up chunk ids for a document, then delete
each by its id primary key instead. Includes a v1 -> v2 migration to
backfill document_chunks for indexes created before this existed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-28 08:56:16 -07:00
39 changed files with 1860 additions and 886 deletions
@@ -0,0 +1,12 @@
#!/command/with-contenv /usr/bin/bash
# shellcheck shell=bash
declare -r log_prefix="[init-llmindex-migrate]"
echo "${log_prefix} Checking LLM index schema..."
cd "${PAPERLESS_SRC_DIR}"
if [[ -n "${USER_IS_NON_ROOT}" ]]; then
python3 manage.py document_llmindex migrate
else
s6-setuidgid paperless python3 manage.py document_llmindex migrate
fi
@@ -0,0 +1 @@
oneshot
@@ -0,0 +1 @@
/etc/s6-overlay/s6-rc.d/init-llmindex-migrate/run
+20 -1
View File
@@ -212,6 +212,16 @@ following:
This is a no-op if the index is already up to date, so it is safe to This is a no-op if the index is already up to date, so it is safe to
run on every upgrade. run on every upgrade.
5. Apply any pending LLM index schema migrations.
```shell-session
cd src
python3 manage.py document_llmindex migrate
```
This is a no-op if the index is already up to date, or if the LLM index
is disabled, so it is safe to run on every upgrade.
### Database Upgrades ### Database Upgrades
Paperless-ngx is compatible with Django-supported versions of PostgreSQL and MariaDB and it is generally Paperless-ngx is compatible with Django-supported versions of PostgreSQL and MariaDB and it is generally
@@ -532,7 +542,7 @@ index is updated automatically on the schedule set by
can manage it manually: can manage it manually:
``` ```
document_llmindex {rebuild,update,compact} document_llmindex {rebuild,update,compact,migrate}
``` ```
Specify `rebuild` to build the index from scratch from all documents in the database. Use Specify `rebuild` to build the index from scratch from all documents in the database. Use
@@ -544,6 +554,15 @@ scheduled task runs.
Specify `compact` to reclaim space and optimize the on-disk vector store. Specify `compact` to reclaim space and optimize the on-disk vector store.
Specify `migrate` to apply any pending index schema migrations without a full reindex.
This is a no-op if the index is already up to date, so it is safe to run on every
startup or upgrade; the container's startup sequence runs it automatically, and the
[bare-metal upgrade steps](#bare-metal-updating) include it as a manual step. If a
pending migration would require re-embedding every document, `migrate` only logs a
warning and leaves the index as-is -- re-embedding can be slow and, for a metered
embedding backend, cost money, so it is never triggered automatically. Run `rebuild`
yourself when you are ready.
!!! note !!! note
These commands have no effect unless AI is enabled and an embedding backend is These commands have no effect unless AI is enabled and an embedding backend is
-28
View File
@@ -620,34 +620,6 @@ no other workflow will be executed on the document.
If a "Move to Trash" action is executed in a consume pipeline, the consumption If a "Move to Trash" action is executed in a consume pipeline, the consumption
will be aborted and the file will be deleted. will be aborted and the file will be deleted.
##### Password Removal {#workflow-action-password-removal}
"Password Removal" actions attempt to remove password protection from encrypted PDF documents. You can specify:
- One or more passwords to try, separated by commas or new lines
- Each password is tried in order until one successfully unlocks the document
Password removal never modifies a file in place. Instead, once a working password is found, the
decrypted content is consumed as a new [document version](#document-file-versions), leaving the
original (still encrypted) version in the document's version history.
**Consumption Started**: because this trigger fires before the document exists yet, the password
removal itself is deferred until after the initial consumption of the encrypted file has completed.
OCR engines cannot process an encrypted PDF, so this first version is typically stored with no
extracted text (unless the file already contained extractable text outside of OCR). Immediately
afterwards, the password is removed and the decrypted file is automatically re-consumed as a second,
new version of the same document, this time with normal OCR/text extraction applied. In other words,
a password-protected file added with this trigger will briefly exist as an un-OCR'd version before
the properly processed version is created.
**Document Added**, **Document Updated**, **Scheduled**: these triggers run against a document that
already exists, so password removal happens immediately: the decrypted content is queued for
consumption as a new version right away. Note that if the document's initial consumption also
happened while it was still encrypted, that original version will likewise be missing OCR text.
**Current limitation**: Passwords are stored as a simple list without descriptions. To handle
multiple PDF types with different passwords, create separate workflows for each use case.
#### Workflow placeholders #### Workflow placeholders
Titles and webhook payloads can be generated by workflows using [Jinja templates](https://jinja.palletsprojects.com/en/3.1.x/templates/). Titles and webhook payloads can be generated by workflows using [Jinja templates](https://jinja.palletsprojects.com/en/3.1.x/templates/).
+35 -35
View File
@@ -659,7 +659,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">449</context> <context context-type="linenumber">445</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-list/bulk-editor/custom-fields-bulk-edit-dialog/custom-fields-bulk-edit-dialog.component.html</context> <context context-type="sourcefile">src/app/components/document-list/bulk-editor/custom-fields-bulk-edit-dialog/custom-fields-bulk-edit-dialog.component.html</context>
@@ -831,7 +831,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">468</context> <context context-type="linenumber">464</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-list/document-list.component.html</context> <context context-type="sourcefile">src/app/components/document-list/document-list.component.html</context>
@@ -1355,7 +1355,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">394</context> <context context-type="linenumber">390</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context> <context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context>
@@ -1596,7 +1596,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">169</context> <context context-type="linenumber">165</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="2691296884221415710" datatype="html"> <trans-unit id="2691296884221415710" datatype="html">
@@ -1607,7 +1607,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">174</context> <context context-type="linenumber">170</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context> <context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context>
@@ -1638,7 +1638,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">178</context> <context context-type="linenumber">174</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context> <context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context>
@@ -1669,7 +1669,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">182</context> <context context-type="linenumber">178</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context> <context context-type="sourcefile">src/app/components/document-list/bulk-editor/bulk-editor.component.html</context>
@@ -4914,7 +4914,7 @@
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">360</context> <context context-type="linenumber">356</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="8057014866157903311" datatype="html"> <trans-unit id="8057014866157903311" datatype="html">
@@ -7760,14 +7760,14 @@
<source>Details</source> <source>Details</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">164</context> <context context-type="linenumber">160</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="5701618810648052610" datatype="html"> <trans-unit id="5701618810648052610" datatype="html">
<source>Title</source> <source>Title</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">167</context> <context context-type="linenumber">163</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-list/document-list.component.html</context> <context context-type="sourcefile">src/app/components/document-list/document-list.component.html</context>
@@ -7790,14 +7790,14 @@
<source>Date created</source> <source>Date created</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">171</context> <context context-type="linenumber">167</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="5607669932062416162" datatype="html"> <trans-unit id="5607669932062416162" datatype="html">
<source>Default</source> <source>Default</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">183</context> <context context-type="linenumber">179</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/manage/saved-views/saved-views.component.html</context> <context context-type="sourcefile">src/app/components/manage/saved-views/saved-views.component.html</context>
@@ -7808,14 +7808,14 @@
<source>Content</source> <source>Content</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">290</context> <context context-type="linenumber">286</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="218403386307979629" datatype="html"> <trans-unit id="218403386307979629" datatype="html">
<source>Metadata</source> <source>Metadata</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">299</context> <context context-type="linenumber">295</context>
</context-group> </context-group>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/metadata-collapse/metadata-collapse.component.ts</context> <context context-type="sourcefile">src/app/components/document-detail/metadata-collapse/metadata-collapse.component.ts</context>
@@ -7826,147 +7826,147 @@
<source>Date modified</source> <source>Date modified</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">306</context> <context context-type="linenumber">302</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="6392918669949841614" datatype="html"> <trans-unit id="6392918669949841614" datatype="html">
<source>Date added</source> <source>Date added</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">310</context> <context context-type="linenumber">306</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="146828917013192897" datatype="html"> <trans-unit id="146828917013192897" datatype="html">
<source>Media filename</source> <source>Media filename</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">314</context> <context context-type="linenumber">310</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="4500855521601039868" datatype="html"> <trans-unit id="4500855521601039868" datatype="html">
<source>Original filename</source> <source>Original filename</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">318</context> <context context-type="linenumber">314</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="2659735245739197634" datatype="html"> <trans-unit id="2659735245739197634" datatype="html">
<source>Original SHA256 checksum</source> <source>Original SHA256 checksum</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">322</context> <context context-type="linenumber">318</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="5888243105821763422" datatype="html"> <trans-unit id="5888243105821763422" datatype="html">
<source>Original file size</source> <source>Original file size</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">326</context> <context context-type="linenumber">322</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="2696647325713149563" datatype="html"> <trans-unit id="2696647325713149563" datatype="html">
<source>Original mime type</source> <source>Original mime type</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">330</context> <context context-type="linenumber">326</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="6714358112223607756" datatype="html"> <trans-unit id="6714358112223607756" datatype="html">
<source>Archive SHA256 checksum</source> <source>Archive SHA256 checksum</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">335</context> <context context-type="linenumber">331</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="6033581412811562084" datatype="html"> <trans-unit id="6033581412811562084" datatype="html">
<source>Archive file size</source> <source>Archive file size</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">341</context> <context context-type="linenumber">337</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="6992781481378431874" datatype="html"> <trans-unit id="6992781481378431874" datatype="html">
<source>Original document metadata</source> <source>Original document metadata</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">350</context> <context context-type="linenumber">346</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="2846565152091361585" datatype="html"> <trans-unit id="2846565152091361585" datatype="html">
<source>Archived document metadata</source> <source>Archived document metadata</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">353</context> <context context-type="linenumber">349</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="7206723502037428235" datatype="html"> <trans-unit id="7206723502037428235" datatype="html">
<source>Notes <x id="START_BLOCK_IF" equiv-text="@if (document()?.notes.length) {"/><x id="START_TAG_SPAN" ctype="x-span" equiv-text="&lt;span class=&quot;badge text-bg-secondary ms-1&quot;&gt;"/><x id="INTERPOLATION" equiv-text="length}}"/><x id="CLOSE_TAG_SPAN" ctype="x-span"/><x id="CLOSE_BLOCK_IF" equiv-text="}"/></source> <source>Notes <x id="START_BLOCK_IF" equiv-text="@if (document()?.notes.length) {"/><x id="START_TAG_SPAN" ctype="x-span" equiv-text="&lt;span class=&quot;badge text-bg-secondary ms-1&quot;&gt;"/><x id="INTERPOLATION" equiv-text="length}}"/><x id="CLOSE_TAG_SPAN" ctype="x-span"/><x id="CLOSE_BLOCK_IF" equiv-text="}"/></source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">372,375</context> <context context-type="linenumber">368,371</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="186236568870281953" datatype="html"> <trans-unit id="186236568870281953" datatype="html">
<source>History</source> <source>History</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">383</context> <context context-type="linenumber">379</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="8236092845697214347" datatype="html"> <trans-unit id="8236092845697214347" datatype="html">
<source> Duplicates <x id="START_TAG_SPAN" ctype="x-span" equiv-text="&lt;span class=&quot;badge text-bg-secondary ms-1&quot;&gt;"/><x id="INTERPOLATION" equiv-text="cate_documents.length }}"/><x id="CLOSE_TAG_SPAN" ctype="x-span"/></source> <source> Duplicates <x id="START_TAG_SPAN" ctype="x-span" equiv-text="&lt;span class=&quot;badge text-bg-secondary ms-1&quot;&gt;"/><x id="INTERPOLATION" equiv-text="cate_documents.length }}"/><x id="CLOSE_TAG_SPAN" ctype="x-span"/></source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">405,409</context> <context context-type="linenumber">401,405</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="6449374629822973702" datatype="html"> <trans-unit id="6449374629822973702" datatype="html">
<source>Duplicate documents detected:</source> <source>Duplicate documents detected:</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">411</context> <context context-type="linenumber">407</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="14058600336670816" datatype="html"> <trans-unit id="14058600336670816" datatype="html">
<source>In trash</source> <source>In trash</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">422</context> <context context-type="linenumber">418</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="5129524307369213584" datatype="html"> <trans-unit id="5129524307369213584" datatype="html">
<source>Save &amp; next</source> <source>Save &amp; next</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">451</context> <context context-type="linenumber">447</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="4910102545766233758" datatype="html"> <trans-unit id="4910102545766233758" datatype="html">
<source>Save &amp; close</source> <source>Save &amp; close</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">453</context> <context context-type="linenumber">449</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="3823219296477075982" datatype="html"> <trans-unit id="3823219296477075982" datatype="html">
<source>Discard</source> <source>Discard</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">455</context> <context context-type="linenumber">451</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="1309556917227148591" datatype="html"> <trans-unit id="1309556917227148591" datatype="html">
<source>Document loading...</source> <source>Document loading...</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">463</context> <context context-type="linenumber">459</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="8191371354890763172" datatype="html"> <trans-unit id="8191371354890763172" datatype="html">
<source>Enter Password</source> <source>Enter Password</source>
<context-group purpose="location"> <context-group purpose="location">
<context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context> <context context-type="sourcefile">src/app/components/document-detail/document-detail.component.html</context>
<context context-type="linenumber">517</context> <context context-type="linenumber">513</context>
</context-group> </context-group>
</trans-unit> </trans-unit>
<trans-unit id="5758784066858623886" datatype="html"> <trans-unit id="5758784066858623886" datatype="html">
@@ -129,25 +129,13 @@ describe('PngxPdfViewerComponent', () => {
;(component as any).applyScale() ;(component as any).applyScale()
expect(viewer.currentScaleValue).toBe(PdfZoomScale.PageFit) expect(viewer.currentScaleValue).toBe(PdfZoomScale.PageFit)
expect(viewer.currentScale).toBe(2) expect(viewer.currentScale).toBe(2)
})
it('does not reapply scale for page-only changes', async () => {
await initComponent()
const pdf = (component as any).pdf as { numPages: number }
pdf.numPages = 3
const viewer = (component as any).pdfViewer as PDFViewer
viewer.setDocument(pdf)
const applyScaleSpy = jest.spyOn(component as any, 'applyScale') const applyScaleSpy = jest.spyOn(component as any, 'applyScale')
component.page = 2 component.page = 2
;(component as any).lastViewerPage = 2
component.ngOnChanges({ ;(component as any).applyViewerState()
page: new SimpleChange(1, 2, false),
})
expect(viewer.currentPageNumber).toBe(2)
expect((component as any).lastViewerPage).toBeUndefined() expect((component as any).lastViewerPage).toBeUndefined()
expect(applyScaleSpy).not.toHaveBeenCalled() expect(applyScaleSpy).toHaveBeenCalled()
}) })
it('does not reset the viewer when it is already on the requested page', async () => { it('does not reset the viewer when it is already on the requested page', async () => {
@@ -116,10 +116,7 @@ export class PngxPdfViewerComponent
changes['zoomScale'] || changes['zoomScale'] ||
changes['rotation'] changes['rotation']
) { ) {
// Prevent loop with page / scale application see https://github.com/paperless-ngx/paperless-ngx/issues/13404 this.applyViewerState()
this.applyViewerState(
!!(changes['zoom'] || changes['zoomScale'] || changes['rotation'])
)
} }
if (changes['searchQuery']) { if (changes['searchQuery']) {
@@ -243,7 +240,7 @@ export class PngxPdfViewerComponent
} }
} }
private applyViewerState(applyScale = true): void { private applyViewerState(): void {
if (!this.pdfViewer) { if (!this.pdfViewer) {
return return
} }
@@ -267,7 +264,7 @@ export class PngxPdfViewerComponent
if (this.page === this.lastViewerPage) { if (this.page === this.lastViewerPage) {
this.lastViewerPage = undefined this.lastViewerPage = undefined
} }
if (hasPages && applyScale) { if (hasPages) {
this.applyScale() this.applyScale()
} }
this.dispatchFindIfReady() this.dispatchFindIfReady()
@@ -113,8 +113,8 @@
<form [formGroup]='documentForm' (ngSubmit)="save()"> <form [formGroup]='documentForm' (ngSubmit)="save()">
<div class="btn-toolbar justify-content-end mb-1 pb-3 gap-2 row-gap-2 border-bottom"> <div class="btn-toolbar mb-1 border-bottom">
<div class="btn-group me-auto"> <div class="btn-group pb-3">
<button type="button" class="btn btn-sm btn-outline-secondary" i18n-title title="Close" (click)="close()"> <button type="button" class="btn btn-sm btn-outline-secondary" i18n-title title="Close" (click)="close()">
<i-bs width="1.2em" height="1.2em" name="x"></i-bs> <i-bs width="1.2em" height="1.2em" name="x"></i-bs>
</button> </button>
@@ -127,36 +127,32 @@
</div> </div>
<ng-container *pngxIfPermissions="{ action: PermissionAction.Change, type: PermissionType.Document }"> <ng-container *pngxIfPermissions="{ action: PermissionAction.Change, type: PermissionType.Document }">
<div class="d-flex gap-2"> <div class="btn-group pb-3 ms-auto">
<div class="btn-group"> <pngx-suggestions-dropdown *pngxIfPermissions="{ action: PermissionAction.Change, type: PermissionType.Document }"
<pngx-suggestions-dropdown *pngxIfPermissions="{ action: PermissionAction.Change, type: PermissionType.Document }" [disabled]="!userCanEdit || suggestionsLoading()"
[disabled]="!userCanEdit || suggestionsLoading()" [loading]="suggestionsLoading()"
[loading]="suggestionsLoading()" [suggestions]="suggestions()"
[suggestions]="suggestions()" [aiEnabled]="aiEnabled"
[aiEnabled]="aiEnabled" (getSuggestions)="getSuggestions()"
(getSuggestions)="getSuggestions()" (addTag)="createTag($event)"
(addTag)="createTag($event)" (addDocumentType)="createDocumentType($event)"
(addDocumentType)="createDocumentType($event)" (addCorrespondent)="createCorrespondent($event)">
(addCorrespondent)="createCorrespondent($event)"> </pngx-suggestions-dropdown>
</pngx-suggestions-dropdown>
</div>
<div class="btn-group">
<pngx-custom-fields-dropdown
*pngxIfPermissions="{ action: PermissionAction.View, type: PermissionType.CustomField }"
[documentId]="documentId()"
[disabled]="!userCanEdit"
[existingFields]="document()?.custom_fields"
(created)="refreshCustomFields()"
(added)="addField($event)">
</pngx-custom-fields-dropdown>
</div>
</div> </div>
<div class="ps-3"> <div class="btn-group pb-3 ms-2">
<ng-container *ngTemplateOutlet="saveButtons"></ng-container> <pngx-custom-fields-dropdown
*pngxIfPermissions="{ action: PermissionAction.View, type: PermissionType.CustomField }"
[documentId]="documentId()"
[disabled]="!userCanEdit"
[existingFields]="document()?.custom_fields"
(created)="refreshCustomFields()"
(added)="addField($event)">
</pngx-custom-fields-dropdown>
</div> </div>
</ng-container> </ng-container>
<ng-container *ngTemplateOutlet="saveButtons"></ng-container>
</div> </div>
<ul ngbNav #nav="ngbNav" class="nav-underline flex-nowrap flex-md-wrap overflow-auto" (navChange)="onNavChange($event)" [activeId]="activeNavID()" (activeIdChange)="activeNavID.set($event)"> <ul ngbNav #nav="ngbNav" class="nav-underline flex-nowrap flex-md-wrap overflow-auto" (navChange)="onNavChange($event)" [activeId]="activeNavID()" (activeIdChange)="activeNavID.set($event)">
@@ -280,7 +276,7 @@
} }
</div> </div>
<div class="d-flex justify-content-end border-top pt-3"> <div class="d-flex border-top pt-3">
<ng-container *ngTemplateOutlet="saveButtons"></ng-container> <ng-container *ngTemplateOutlet="saveButtons"></ng-container>
</div> </div>
</ng-template> </ng-template>
@@ -444,7 +440,7 @@
</div> </div>
<ng-template #saveButtons> <ng-template #saveButtons>
<div class="btn-group"> <div class="btn-group pb-3 ms-4">
<ng-container *pngxIfPermissions="{ action: PermissionAction.Change, type: PermissionType.Document }"> <ng-container *pngxIfPermissions="{ action: PermissionAction.Change, type: PermissionType.Document }">
<button type="submit" class="order-3 btn btn-sm btn-primary" i18n [disabled]="!userCanEdit || networkActive() || (isDirty$ | async) !== true">Save</button> <button type="submit" class="order-3 btn btn-sm btn-primary" i18n [disabled]="!userCanEdit || networkActive() || (isDirty$ | async) !== true">Save</button>
@if (hasNext()) { @if (hasNext()) {
+25 -15
View File
@@ -57,7 +57,9 @@ from paperless.models import ArchiveFileGenerationChoices
from paperless.parsers import ParserContext from paperless.parsers import ParserContext
from paperless.parsers import ParserProtocol from paperless.parsers import ParserProtocol
from paperless.parsers.registry import get_parser_registry from paperless.parsers.registry import get_parser_registry
from paperless.parsers.utils import pdf_born_digital_text from paperless.parsers.utils import PDF_TEXT_MIN_LENGTH
from paperless.parsers.utils import extract_pdf_text
from paperless.parsers.utils import is_tagged_pdf
LOGGING_NAME: Final[str] = "paperless.consumer" LOGGING_NAME: Final[str] = "paperless.consumer"
@@ -136,45 +138,53 @@ def should_produce_archive(
# Must produce a PDF so the frontend can display the original format at all. # Must produce a PDF so the frontend can display the original format at all.
if parser.requires_pdf_rendition: if parser.requires_pdf_rendition:
_log.debug("Archive: yes - parser requires PDF rendition for frontend display") _log.debug("Archive: yes parser requires PDF rendition for frontend display")
return True return True
# Parser cannot produce an archive (e.g. TextDocumentParser). # Parser cannot produce an archive (e.g. TextDocumentParser).
if not parser.can_produce_archive: if not parser.can_produce_archive:
_log.debug("Archive: no - parser cannot produce archives") _log.debug("Archive: no parser cannot produce archives")
return False return False
generation = OcrConfig().archive_file_generation generation = OcrConfig().archive_file_generation
if generation == ArchiveFileGenerationChoices.ALWAYS: if generation == ArchiveFileGenerationChoices.ALWAYS:
_log.debug("Archive: yes - ARCHIVE_FILE_GENERATION=always") _log.debug("Archive: yes ARCHIVE_FILE_GENERATION=always")
return True return True
if generation == ArchiveFileGenerationChoices.NEVER: if generation == ArchiveFileGenerationChoices.NEVER:
_log.debug("Archive: no - ARCHIVE_FILE_GENERATION=never") _log.debug("Archive: no ARCHIVE_FILE_GENERATION=never")
return False return False
# auto: produce archives for scanned/image documents; skip for born-digital PDFs. # auto: produce archives for scanned/image documents; skip for born-digital PDFs.
if mime_type.startswith("image/"): if mime_type.startswith("image/"):
_log.debug("Archive: yes - image document, ARCHIVE_FILE_GENERATION=auto") _log.debug("Archive: yes image document, ARCHIVE_FILE_GENERATION=auto")
return True return True
if mime_type == "application/pdf": if mime_type == "application/pdf":
text, born_digital = pdf_born_digital_text(document_path, log=_log) text = extract_pdf_text(document_path)
text_length = len(text) if text else 0 has_text = text is not None and len(text) > 0
if born_digital: if has_text and is_tagged_pdf(document_path):
_log.debug( _log.debug(
"Archive: no - born-digital PDF (text_length=%d)," "Archive: no born-digital PDF (structure tags detected),"
" ARCHIVE_FILE_GENERATION=auto", " ARCHIVE_FILE_GENERATION=auto",
text_length,
) )
return False return False
if text is None or len(text) <= PDF_TEXT_MIN_LENGTH:
_log.debug(
"Archive: yes — scanned PDF (text_length=%d%d),"
" ARCHIVE_FILE_GENERATION=auto",
len(text) if text else 0,
PDF_TEXT_MIN_LENGTH,
)
return True
_log.debug( _log.debug(
"Archive: yes - scanned/textless PDF (text_length=%d)," "Archive: no — born-digital PDF (text_length=%d > %d),"
" ARCHIVE_FILE_GENERATION=auto", " ARCHIVE_FILE_GENERATION=auto",
text_length, len(text),
PDF_TEXT_MIN_LENGTH,
) )
return True return False
_log.debug( _log.debug(
"Archive: no - MIME type %r not eligible for auto archive generation", "Archive: no MIME type %r not eligible for auto archive generation",
mime_type, mime_type,
) )
return False return False
-3
View File
@@ -36,9 +36,6 @@ def send_email(
TODO: re-evaluate this pending https://code.djangoproject.com/ticket/35581 / https://github.com/django/django/pull/18966 TODO: re-evaluate this pending https://code.djangoproject.com/ticket/35581 / https://github.com/django/django/pull/18966
""" """
if "\r" in subject or "\n" in subject:
subject = " ".join(line.strip(" \t") for line in subject.splitlines())
email = EmailMessage( email = EmailMessage(
subject=subject, subject=subject,
body=body, body=body,
@@ -386,19 +386,10 @@ class Command(CryptMixin, PaperlessCommand):
raise DeserializationError( raise DeserializationError(
f"{model.__name__} has no updatable fields; PK-only models are not supported by the importer", f"{model.__name__} has no updatable fields; PK-only models are not supported by the importer",
) )
# MySQL/MariaDB support upserts via ON DUPLICATE KEY UPDATE but,
# unlike PostgreSQL/SQLite, cannot target a specific unique field
# for the conflict -- passing unique_fields there raises
# NotSupportedError.
unique_fields = (
[model._meta.pk.attname]
if connection.features.supports_update_conflicts_with_target
else None
)
model.objects.bulk_create( # type: ignore[attr-defined] model.objects.bulk_create( # type: ignore[attr-defined]
instances, instances,
update_conflicts=True, update_conflicts=True,
unique_fields=unique_fields, unique_fields=[model._meta.pk.attname],
update_fields=update_fields, update_fields=update_fields,
) )
loaded_models.add(model) loaded_models.add(model)
@@ -3,6 +3,7 @@ from typing import Any
from documents.management.commands.base import PaperlessCommand from documents.management.commands.base import PaperlessCommand
from documents.tasks import llmindex_index from documents.tasks import llmindex_index
from paperless_ai.indexing import llm_index_compact from paperless_ai.indexing import llm_index_compact
from paperless_ai.indexing import llm_index_migrate
class Command(PaperlessCommand): class Command(PaperlessCommand):
@@ -13,12 +14,18 @@ class Command(PaperlessCommand):
def add_arguments(self, parser: Any) -> None: def add_arguments(self, parser: Any) -> None:
super().add_arguments(parser) super().add_arguments(parser)
parser.add_argument("command", choices=["rebuild", "update", "compact"]) parser.add_argument(
"command",
choices=["rebuild", "update", "compact", "migrate"],
)
def handle(self, *args: Any, **options: Any) -> None: def handle(self, *args: Any, **options: Any) -> None:
if options["command"] == "compact": if options["command"] == "compact":
llm_index_compact() llm_index_compact()
return return
if options["command"] == "migrate":
llm_index_migrate()
return
llmindex_index( llmindex_index(
rebuild=options["command"] == "rebuild", rebuild=options["command"] == "rebuild",
iter_wrapper=lambda docs: self.track( iter_wrapper=lambda docs: self.track(
+2 -10
View File
@@ -43,16 +43,8 @@ def _fmt(dt: datetime) -> str:
def _iso_range(lo: datetime, hi: datetime) -> str: def _iso_range(lo: datetime, hi: datetime) -> str:
""" """Format a [lo TO hi] range string in ISO 8601 for Tantivy query syntax."""
Format a half-open ``[lo TO hi)`` range in ISO 8601 for Tantivy query syntax. return f"[{_fmt(lo)} TO {_fmt(hi)}]"
``hi`` is always the exclusive ceiling of a computed period (the start of
the *next* day/week/month/quarter/year), so the closing bracket must be
the Tantivy exclusive-range brace ``}`` rather than ``]`` — otherwise the
first instant of the following period (e.g. the 1st of next month) is
incorrectly included in the match.
"""
return f"[{_fmt(lo)} TO {_fmt(hi)}}}"
def _quarter_start(d: date) -> date: def _quarter_start(d: date) -> date:
+2 -13
View File
@@ -566,17 +566,6 @@ def translate_range(field: str, lo: str, hi: str, tz: tzinfo) -> str:
lo_pair, hi_pair = hi_pair, lo_pair lo_pair, hi_pair = hi_pair, lo_pair
lo_iso = _fmt(lo_pair[0]) if lo_pair is not None else OPEN_LO lo_iso = _fmt(lo_pair[0]) if lo_pair is not None else OPEN_LO
hi_iso = _fmt(hi_pair[1]) if hi_pair is not None else OPEN_HI
# A bound resolves to (floor, ceil) where floor == ceil for an exact instant return f"{field}:[{lo_iso} TO {hi_iso}]"
# (a full ISO datetime, "now", or a "+/-N unit" offset) and floor != ceil for
# a coarser period token (year/month/day precision). Only the latter needs a
# half-open close: its ceil is the start of the *next* period and must be
# excluded, or that instant (e.g. the 1st of next month) wrongly matches.
if hi_pair is not None:
hi_iso = _fmt(hi_pair[1])
hi_close = "]" if hi_pair[0] == hi_pair[1] else "}"
else:
hi_iso = OPEN_HI
hi_close = "]"
return f"{field}:[{lo_iso} TO {hi_iso}{hi_close}"
-5
View File
@@ -1134,11 +1134,6 @@ def before_task_publish_handler(
return return
try: try:
# Close stale connections without disrupting a transaction publishing a task
for connection in connections.all(initialized_only=True):
if not connection.in_atomic_block:
connection.close_if_unusable_or_obsolete()
_, task_kwargs, _ = body _, task_kwargs, _ = body
task_id = headers["id"] task_id = headers["id"]
@@ -8,6 +8,7 @@ if TYPE_CHECKING:
from pytest_mock import MockerFixture from pytest_mock import MockerFixture
_COMPACT = "documents.management.commands.document_llmindex.llm_index_compact" _COMPACT = "documents.management.commands.document_llmindex.llm_index_compact"
_MIGRATE = "documents.management.commands.document_llmindex.llm_index_migrate"
_INDEX = "documents.management.commands.document_llmindex.llmindex_index" _INDEX = "documents.management.commands.document_llmindex.llmindex_index"
@@ -17,6 +18,11 @@ class TestDocumentLlmindexCommand:
call_command("document_llmindex", "compact") call_command("document_llmindex", "compact")
mock_compact.assert_called_once_with() mock_compact.assert_called_once_with()
def test_migrate_calls_llm_index_migrate(self, mocker: MockerFixture) -> None:
mock_migrate = mocker.patch(_MIGRATE)
call_command("document_llmindex", "migrate")
mock_migrate.assert_called_once_with()
def test_rebuild_calls_llmindex_index_with_rebuild_true( def test_rebuild_calls_llmindex_index_with_rebuild_true(
self, self,
mocker: MockerFixture, mocker: MockerFixture,
+1 -3
View File
@@ -32,9 +32,7 @@ AUCKLAND = ZoneInfo("Pacific/Auckland") # UTC+13 in southern-hemisphere summer
def _range(result: str, field: str) -> tuple[str, str]: def _range(result: str, field: str) -> tuple[str, str]:
# Half-open period ranges close with "}" (exclusive); exact-instant ranges m = re.search(rf"{field}:\[(.+?) TO (.+?)\]", result)
# (full ISO datetimes, "now", relative offsets) close with "]" (inclusive).
m = re.search(rf"{field}:\[(.+?) TO (.+?)[\]}}]", result)
assert m, f"No range for {field!r} in: {result!r}" assert m, f"No range for {field!r} in: {result!r}"
return m.group(1), m.group(2) return m.group(1), m.group(2)
+34 -34
View File
@@ -214,27 +214,27 @@ class TestTranslateScalar:
( (
"created", "created",
"2020", "2020",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}", "created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z]",
), ),
( (
"created", "created",
"202003", "202003",
"created:[2020-03-01T00:00:00Z TO 2020-04-01T00:00:00Z}", "created:[2020-03-01T00:00:00Z TO 2020-04-01T00:00:00Z]",
), ),
( (
"created", "created",
"20200115", "20200115",
"created:[2020-01-15T00:00:00Z TO 2020-01-16T00:00:00Z}", "created:[2020-01-15T00:00:00Z TO 2020-01-16T00:00:00Z]",
), ),
( (
"created", "created",
"2020-01-15", "2020-01-15",
"created:[2020-01-15T00:00:00Z TO 2020-01-16T00:00:00Z}", "created:[2020-01-15T00:00:00Z TO 2020-01-16T00:00:00Z]",
), ),
( (
"created", "created",
"2020-03", "2020-03",
"created:[2020-03-01T00:00:00Z TO 2020-04-01T00:00:00Z}", "created:[2020-03-01T00:00:00Z TO 2020-04-01T00:00:00Z]",
), ),
], ],
) )
@@ -248,9 +248,9 @@ class TestTranslateScalar:
assert exc_info.value.value == "202023" assert exc_info.value.value == "202023"
def test_keyword_delegates(self) -> None: def test_keyword_delegates(self) -> None:
# keyword path produces a half-open range; just assert it is a created range # keyword path produces a range; just assert it is a created range
out = translate_scalar("created", "today", UTC) out = translate_scalar("created", "today", UTC)
assert out.startswith("created:[") and out.endswith("}") assert out.startswith("created:[") and out.endswith("]")
def test_14digit_compact_datetime(self) -> None: def test_14digit_compact_datetime(self) -> None:
out = translate_scalar("created", "20240115120000", UTC) out = translate_scalar("created", "20240115120000", UTC)
@@ -279,21 +279,21 @@ class TestTranslateRange:
@pytest.mark.parametrize( @pytest.mark.parametrize(
("lo", "hi", "expected"), ("lo", "hi", "expected"),
[ [
("2005", "2009", "created:[2005-01-01T00:00:00Z TO 2010-01-01T00:00:00Z}"), ("2005", "2009", "created:[2005-01-01T00:00:00Z TO 2010-01-01T00:00:00Z]"),
( (
"202001", "202001",
"202006", "202006",
"created:[2020-01-01T00:00:00Z TO 2020-07-01T00:00:00Z}", "created:[2020-01-01T00:00:00Z TO 2020-07-01T00:00:00Z]",
), ),
( (
"20200101", "20200101",
"20201231", "20201231",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}", "created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z]",
), ),
( (
"2020-01-01", "2020-01-01",
"2020-12-31", "2020-12-31",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}", "created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z]",
), ),
], ],
) )
@@ -302,7 +302,7 @@ class TestTranslateRange:
def test_reversed_swaps(self): def test_reversed_swaps(self):
assert translate_range("created", "2009", "2005", UTC) == ( assert translate_range("created", "2009", "2005", UTC) == (
"created:[2005-01-01T00:00:00Z TO 2010-01-01T00:00:00Z}" "created:[2005-01-01T00:00:00Z TO 2010-01-01T00:00:00Z]"
) )
def test_open_upper(self): def test_open_upper(self):
@@ -311,7 +311,7 @@ class TestTranslateRange:
def test_open_lower(self): def test_open_lower(self):
out = translate_range("created", "", "2020", UTC) out = translate_range("created", "", "2020", UTC)
assert out == f"created:[{OPEN_LO} TO 2021-01-01T00:00:00Z}}" assert out == f"created:[{OPEN_LO} TO 2021-01-01T00:00:00Z]"
def test_invalid_bound_raises(self): def test_invalid_bound_raises(self):
with pytest.raises(InvalidDateQuery) as exc_info: with pytest.raises(InvalidDateQuery) as exc_info:
@@ -334,16 +334,16 @@ class TestTranslateQuery:
[ [
( (
"created:2020", "created:2020",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}", "created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z]",
), ),
("tag:foo,bar", "tag:foo AND tag:bar"), ("tag:foo,bar", "tag:foo AND tag:bar"),
# 'type' is a user-facing alias rewritten to 'document_type' (the real schema field) # 'type' is a user-facing alias rewritten to 'document_type' (the real schema field)
("tag:foo,type:bar", "tag:foo AND document_type:bar"), ("tag:foo,type:bar", "tag:foo AND document_type:bar"),
( (
"created:[2020 TO 2021],added:[2022 TO 2023]", "created:[2020 TO 2021],added:[2022 TO 2023]",
"created:[2020-01-01T00:00:00Z TO 2022-01-01T00:00:00Z}" "created:[2020-01-01T00:00:00Z TO 2022-01-01T00:00:00Z]"
" AND " " AND "
"added:[2022-01-01T00:00:00Z TO 2024-01-01T00:00:00Z}", "added:[2022-01-01T00:00:00Z TO 2024-01-01T00:00:00Z]",
), ),
# correspondent is not multi-value: comma stays literal inside the value # correspondent is not multi-value: comma stays literal inside the value
("correspondent:foo,bar", "correspondent:foo,bar"), ("correspondent:foo,bar", "correspondent:foo,bar"),
@@ -506,7 +506,7 @@ class TestOperatorNormalization:
def test_date_range_preserved(self) -> None: def test_date_range_preserved(self) -> None:
out = translate_query("created:[2020 TO 2021]", UTC) out = translate_query("created:[2020 TO 2021]", UTC)
# Must not corrupt the ISO range # Must not corrupt the ISO range
assert out == "created:[2020-01-01T00:00:00Z TO 2022-01-01T00:00:00Z}" assert out == "created:[2020-01-01T00:00:00Z TO 2022-01-01T00:00:00Z]"
def test_date_scalar_with_or(self) -> None: def test_date_scalar_with_or(self) -> None:
out = translate_query("created:2020 OR foo", UTC) out = translate_query("created:2020 OR foo", UTC)
@@ -581,42 +581,42 @@ class TestKeywordDateResolution:
[ [
pytest.param( pytest.param(
"today", "today",
"created:[2026-03-28T00:00:00Z TO 2026-03-29T00:00:00Z}", "created:[2026-03-28T00:00:00Z TO 2026-03-29T00:00:00Z]",
id="today", id="today",
), ),
pytest.param( pytest.param(
"yesterday", "yesterday",
"created:[2026-03-27T00:00:00Z TO 2026-03-28T00:00:00Z}", "created:[2026-03-27T00:00:00Z TO 2026-03-28T00:00:00Z]",
id="yesterday", id="yesterday",
), ),
pytest.param( pytest.param(
"previous week", "previous week",
"created:[2026-03-16T00:00:00Z TO 2026-03-23T00:00:00Z}", "created:[2026-03-16T00:00:00Z TO 2026-03-23T00:00:00Z]",
id="previous-week", id="previous-week",
), ),
pytest.param( pytest.param(
"this month", "this month",
"created:[2026-03-01T00:00:00Z TO 2026-04-01T00:00:00Z}", "created:[2026-03-01T00:00:00Z TO 2026-04-01T00:00:00Z]",
id="this-month", id="this-month",
), ),
pytest.param( pytest.param(
"previous month", "previous month",
"created:[2026-02-01T00:00:00Z TO 2026-03-01T00:00:00Z}", "created:[2026-02-01T00:00:00Z TO 2026-03-01T00:00:00Z]",
id="previous-month", id="previous-month",
), ),
pytest.param( pytest.param(
"this year", "this year",
"created:[2026-01-01T00:00:00Z TO 2027-01-01T00:00:00Z}", "created:[2026-01-01T00:00:00Z TO 2027-01-01T00:00:00Z]",
id="this-year", id="this-year",
), ),
pytest.param( pytest.param(
"previous year", "previous year",
"created:[2025-01-01T00:00:00Z TO 2026-01-01T00:00:00Z}", "created:[2025-01-01T00:00:00Z TO 2026-01-01T00:00:00Z]",
id="previous-year", id="previous-year",
), ),
pytest.param( pytest.param(
"previous quarter", "previous quarter",
"created:[2025-10-01T00:00:00Z TO 2026-01-01T00:00:00Z}", "created:[2025-10-01T00:00:00Z TO 2026-01-01T00:00:00Z]",
id="previous-quarter", id="previous-quarter",
), ),
], ],
@@ -637,42 +637,42 @@ class TestKeywordDateResolution:
[ [
pytest.param( pytest.param(
"today", "today",
"added:[2026-03-27T15:00:00Z TO 2026-03-28T15:00:00Z}", "added:[2026-03-27T15:00:00Z TO 2026-03-28T15:00:00Z]",
id="today", id="today",
), ),
pytest.param( pytest.param(
"yesterday", "yesterday",
"added:[2026-03-26T15:00:00Z TO 2026-03-27T15:00:00Z}", "added:[2026-03-26T15:00:00Z TO 2026-03-27T15:00:00Z]",
id="yesterday", id="yesterday",
), ),
pytest.param( pytest.param(
"previous week", "previous week",
"added:[2026-03-15T15:00:00Z TO 2026-03-22T15:00:00Z}", "added:[2026-03-15T15:00:00Z TO 2026-03-22T15:00:00Z]",
id="previous-week", id="previous-week",
), ),
pytest.param( pytest.param(
"this month", "this month",
"added:[2026-02-28T15:00:00Z TO 2026-03-31T15:00:00Z}", "added:[2026-02-28T15:00:00Z TO 2026-03-31T15:00:00Z]",
id="this-month", id="this-month",
), ),
pytest.param( pytest.param(
"previous month", "previous month",
"added:[2026-01-31T15:00:00Z TO 2026-02-28T15:00:00Z}", "added:[2026-01-31T15:00:00Z TO 2026-02-28T15:00:00Z]",
id="previous-month", id="previous-month",
), ),
pytest.param( pytest.param(
"this year", "this year",
"added:[2025-12-31T15:00:00Z TO 2026-12-31T15:00:00Z}", "added:[2025-12-31T15:00:00Z TO 2026-12-31T15:00:00Z]",
id="this-year", id="this-year",
), ),
pytest.param( pytest.param(
"previous year", "previous year",
"added:[2024-12-31T15:00:00Z TO 2025-12-31T15:00:00Z}", "added:[2024-12-31T15:00:00Z TO 2025-12-31T15:00:00Z]",
id="previous-year", id="previous-year",
), ),
pytest.param( pytest.param(
"previous quarter", "previous quarter",
"added:[2025-09-30T15:00:00Z TO 2025-12-31T15:00:00Z}", "added:[2025-09-30T15:00:00Z TO 2025-12-31T15:00:00Z]",
id="previous-quarter", id="previous-quarter",
), ),
], ],
@@ -719,7 +719,7 @@ class TestISODatetimeBounds:
def test_translate_query_text_before_comma_separated_date_clause(self) -> None: def test_translate_query_text_before_comma_separated_date_clause(self) -> None:
result = translate_query("schäfersee,created:previous year", UTC) result = translate_query("schäfersee,created:previous year", UTC)
assert result == ( assert result == (
"schäfersee AND created:[2025-01-01T00:00:00Z TO 2026-01-01T00:00:00Z}" "schäfersee AND created:[2025-01-01T00:00:00Z TO 2026-01-01T00:00:00Z]"
) )
def test_invalid_iso_datetime_raises(self) -> None: def test_invalid_iso_datetime_raises(self) -> None:
+1 -1
View File
@@ -75,7 +75,7 @@ class TestEmail(DirectoriesMixin, SampleDirMixin, APITestCase):
{ {
"documents": [self.doc1.pk, self.doc2.pk], "documents": [self.doc1.pk, self.doc2.pk],
"addresses": "hello@paperless-ngx.com,test@example.com", "addresses": "hello@paperless-ngx.com,test@example.com",
"subject": "Bulk email\n test", "subject": "Bulk email test",
"message": "Here are your documents", "message": "Here are your documents",
}, },
), ),
-42
View File
@@ -720,48 +720,6 @@ class TestDocumentSearchApi(DirectoriesMixin, APITestCase):
self.assertEqual(results[0]["id"], 3) self.assertEqual(results[0]["id"], 3)
self.assertEqual(results[0]["title"], "bank statement 3") self.assertEqual(results[0]["title"], "bank statement 3")
def test_search_added_previous_month_excludes_next_period_start(self) -> None:
"""
GIVEN:
- One document added at the last instant of last month
- One document added exactly at the first instant of this month
WHEN:
- Query for documents added in the previous month
THEN:
- Only the document from last month is returned; the document dated
exactly at the start of this month (the exclusive upper bound of
the range) is not
"""
d1 = DocumentFactory.create(
title="end of last month",
content="last instant of last month",
checksum="A",
pk=1,
added=timezone.make_aware(datetime.datetime(2024, 1, 31, 23, 59, 59)),
)
d2 = DocumentFactory.create(
title="start of this month",
content="first instant of this month",
checksum="B",
pk=2,
added=timezone.make_aware(datetime.datetime(2024, 2, 1, 0, 0, 0)),
)
backend = get_backend()
backend.add_or_update(d1)
backend.add_or_update(d2)
with time_machine.travel(
timezone.make_aware(datetime.datetime(2024, 2, 15, 12, 0, 0)),
tick=False,
):
response = self.client.get("/api/documents/?query=added:previous month")
results = response.data["results"]
self.assertEqual(len(results), 1)
self.assertEqual(results[0]["id"], 1)
self.assertEqual(results[0]["title"], "end of last month")
def test_search_added_invalid_date(self) -> None: def test_search_added_invalid_date(self) -> None:
""" """
GIVEN: GIVEN:
+2 -2
View File
@@ -1329,7 +1329,7 @@ class PreConsumeTestCase(DirectoriesMixin, GetConsumerMixin, TestCase):
with self.get_consumer(self.test_file) as c: with self.get_consumer(self.test_file) as c:
c.run() c.run()
# Verify no pre-consume script subprocess was invoked # Verify no pre-consume script subprocess was invoked
# (run_subprocess may still be called by pdf_born_digital_text via pdftotext) # (run_subprocess may still be called by _extract_text_for_archive_check)
script_calls = [ script_calls = [
call call
for call in m.call_args_list for call in m.call_args_list
@@ -1354,7 +1354,7 @@ class PreConsumeTestCase(DirectoriesMixin, GetConsumerMixin, TestCase):
self.assertTrue(m.called) self.assertTrue(m.called)
# Find the call that invoked the pre-consume script # Find the call that invoked the pre-consume script
# (run_subprocess may also be called by pdf_born_digital_text via pdftotext) # (run_subprocess may also be called by _extract_text_for_archive_check)
script_call = next( script_call = next(
call call
for call in m.call_args_list for call in m.call_args_list
+42 -14
View File
@@ -134,32 +134,60 @@ class TestShouldProduceArchive:
assert should_produce_archive(parser, mime, Path("/tmp/doc")) is expected assert should_produce_archive(parser, mime, Path("/tmp/doc")) is expected
@pytest.mark.parametrize( @pytest.mark.parametrize(
("born_digital", "expected"), ("extracted_text", "expected"),
[ [
pytest.param(True, False, id="born-digital-skips-archive"), pytest.param(
pytest.param(False, True, id="not-born-digital-produces-archive"), "This is a born-digital PDF with lots of text content. " * 10,
False,
id="born-digital-long-text-skips-archive",
),
pytest.param(None, True, id="no-text-scanned-produces-archive"),
pytest.param("tiny", True, id="short-text-treated-as-scanned"),
], ],
) )
def test_auto_pdf_archive_decision( def test_auto_pdf_archive_decision(
self, self,
mocker: MockerFixture, mocker: MockerFixture,
settings, settings,
born_digital: bool, # noqa: FBT001 extracted_text: str | None,
expected: bool, # noqa: FBT001 expected: bool, # noqa: FBT001
) -> None: ) -> None:
"""Archive decision tracks pdf_born_digital_text()'s verdict exactly.
should_produce_archive() defers entirely to pdf_born_digital_text()
for the has-real-text decision, so both callers of that predicate
(this function and RasterisedDocumentParser.parse()) always agree.
"""
settings.ARCHIVE_FILE_GENERATION = "auto" settings.ARCHIVE_FILE_GENERATION = "auto"
mocker.patch( mocker.patch("documents.consumer.is_tagged_pdf", return_value=False)
"documents.consumer.pdf_born_digital_text", mocker.patch("documents.consumer.extract_pdf_text", return_value=extracted_text)
return_value=("some text", born_digital),
)
parser = _parser_instance(can_produce=True, requires_rendition=False) parser = _parser_instance(can_produce=True, requires_rendition=False)
assert ( assert (
should_produce_archive(parser, "application/pdf", Path("/tmp/doc.pdf")) should_produce_archive(parser, "application/pdf", Path("/tmp/doc.pdf"))
is expected is expected
) )
def test_tagged_pdf_skips_archive_in_auto_mode(
self,
mocker: MockerFixture,
settings,
) -> None:
"""Tagged PDFs (e.g. Word exports) with real text are treated as born-digital, even below PDF_TEXT_MIN_LENGTH."""
settings.ARCHIVE_FILE_GENERATION = "auto"
mocker.patch("documents.consumer.is_tagged_pdf", return_value=True)
mocker.patch("documents.consumer.extract_pdf_text", return_value="tiny")
parser = _parser_instance(can_produce=True, requires_rendition=False)
assert (
should_produce_archive(parser, "application/pdf", Path("/tmp/doc.pdf"))
is False
)
def test_tagged_pdf_without_text_produces_archive(
self,
mocker: MockerFixture,
settings,
) -> None:
"""A tagged PDF with no actual extractable text (e.g. some scanner firmware) is not
trusted as born-digital the tag alone must not bypass OCR."""
settings.ARCHIVE_FILE_GENERATION = "auto"
mocker.patch("documents.consumer.is_tagged_pdf", return_value=True)
mocker.patch("documents.consumer.extract_pdf_text", return_value=None)
parser = _parser_instance(can_produce=True, requires_rendition=False)
assert (
should_produce_archive(parser, "application/pdf", Path("/tmp/doc.pdf"))
is True
)
-26
View File
@@ -56,32 +56,6 @@ def send_publish(
@pytest.mark.django_db @pytest.mark.django_db
class TestBeforeTaskPublishHandler: class TestBeforeTaskPublishHandler:
@mock.patch("documents.signals.handlers.connections.all")
def test_closes_old_connections_outside_atomic_blocks(
self,
connections_all,
) -> None:
connection = mock.Mock(in_atomic_block=False)
connections_all.return_value = [connection]
task_id = send_publish("documents.tasks.train_classifier", (), {})
connection.close_if_unusable_or_obsolete.assert_called_once_with()
assert PaperlessTask.objects.filter(task_id=task_id).exists()
@mock.patch("documents.signals.handlers.connections.all")
def test_keeps_connections_open_inside_atomic_blocks(
self,
connections_all,
) -> None:
connection = mock.Mock(in_atomic_block=True)
connections_all.return_value = [connection]
task_id = send_publish("documents.tasks.train_classifier", (), {})
connection.close_if_unusable_or_obsolete.assert_not_called()
assert PaperlessTask.objects.filter(task_id=task_id).exists()
def test_creates_task_for_consume_file( def test_creates_task_for_consume_file(
self, self,
consume_input_doc, consume_input_doc,
+21 -6
View File
@@ -3,6 +3,7 @@ from __future__ import annotations
import importlib.resources import importlib.resources
import logging import logging
import os import os
import re
import shutil import shutil
import tempfile import tempfile
from pathlib import Path from pathlib import Path
@@ -24,9 +25,9 @@ from paperless.config import OcrConfig
from paperless.models import CleanChoices from paperless.models import CleanChoices
from paperless.models import ModeChoices from paperless.models import ModeChoices
from paperless.models import OutputTypeChoices from paperless.models import OutputTypeChoices
from paperless.parsers.utils import PDF_TEXT_MIN_LENGTH
from paperless.parsers.utils import extract_pdf_text from paperless.parsers.utils import extract_pdf_text
from paperless.parsers.utils import is_born_digital_text from paperless.parsers.utils import is_tagged_pdf
from paperless.parsers.utils import post_process_text
from paperless.parsers.utils import read_file_handle_unicode_errors from paperless.parsers.utils import read_file_handle_unicode_errors
from paperless.version import __full_version_str__ from paperless.version import __full_version_str__
@@ -509,10 +510,10 @@ class RasterisedDocumentParser:
if mime_type == "application/pdf": if mime_type == "application/pdf":
text_original = self.extract_text(None, document_path) text_original = self.extract_text(None, document_path)
original_has_text = is_born_digital_text( has_text = text_original is not None and len(text_original) > 0
text_original, original_has_text = has_text and (
document_path, is_tagged_pdf(document_path, log=self.log)
log=self.log, or len(text_original) > PDF_TEXT_MIN_LENGTH
) )
else: else:
text_original = None text_original = None
@@ -657,3 +658,17 @@ class RasterisedDocumentParser:
f"No text was found in {document_path}, the content will be empty.", f"No text was found in {document_path}, the content will be empty.",
) )
self.text = "" self.text = ""
def post_process_text(text: str | None) -> str | None:
if not text:
return None
collapsed_spaces = re.sub(r"([^\S\r\n]+)", " ", text)
no_leading_whitespace = re.sub(r"([\n\r]+)([^\S\n\r]+)", "\\1", collapsed_spaces)
no_trailing_whitespace = re.sub(r"([^\S\n\r]+)$", "", no_leading_whitespace)
# TODO: this needs a rework
# replace \0 prevents issues with saving to postgres.
# text may contain \0 when this character is present in PDF files.
return no_trailing_whitespace.strip().replace("\0", " ")
-82
View File
@@ -111,88 +111,6 @@ def extract_pdf_text(
return None return None
def post_process_text(text: str | None) -> str | None:
"""Normalize extracted PDF/OCR text: collapse whitespace, strip padding.
Returns ``None`` for ``None`` or whitespace-only input, so callers can
treat "no text" and "only layout padding" the same way.
"""
if not text:
return None
collapsed_spaces = re.sub(r"([^\S\r\n]+)", " ", text)
no_leading_whitespace = re.sub(r"([\n\r]+)([^\S\n\r]+)", "\\1", collapsed_spaces)
no_trailing_whitespace = re.sub(r"([^\S\n\r]+)$", "", no_leading_whitespace)
# replace \0 prevents issues with saving to postgres.
# text may contain \0 when this character is present in PDF files.
result = no_trailing_whitespace.strip().replace("\0", " ")
return result or None
def is_born_digital_text(
text: str | None,
path: Path,
log: logging.Logger | None = None,
) -> bool:
"""Decide whether already-extracted, normalized PDF text counts as born-digital.
This is the single source of truth for "does this PDF already have real
text", used both to decide whether to produce an archive file and to
decide whether OCR can be skipped. Both decisions must agree, or a
tagged-but-textless PDF can end up with no archive AND a forced OCR pass
(see GH #13387): raw ``pdftotext -layout`` output can be non-empty
(whitespace/form-feed padding) even when there is no real content, so
*text* must already be normalized via :func:`post_process_text`, not the
raw extraction.
Parameters
----------
text:
The normalized extracted text (or ``None``) to evaluate.
path:
Absolute path to the PDF file, used for the tagged-PDF check.
log:
Logger for warnings. Falls back to the module-level logger when omitted.
Returns
-------
bool
Whether the PDF counts as born-digital (has real text, and is either
tagged or exceeds ``PDF_TEXT_MIN_LENGTH``).
"""
if not text:
return False
return is_tagged_pdf(path, log=log) or len(text) > PDF_TEXT_MIN_LENGTH
def pdf_born_digital_text(
path: Path,
log: logging.Logger | None = None,
) -> tuple[str | None, bool]:
"""Extract a PDF's text and decide whether it should be treated as born-digital.
Convenience wrapper around :func:`is_born_digital_text` for callers that
don't already have the PDF's text extracted (e.g. the archive-generation
decision, which runs before any parser has touched the file).
Parameters
----------
path:
Absolute path to the PDF file.
log:
Logger for warnings. Falls back to the module-level logger when omitted.
Returns
-------
tuple[str | None, bool]
The normalized extracted text (or ``None``), and whether the PDF
counts as born-digital.
"""
text = post_process_text(extract_pdf_text(path, log=log))
return text, is_born_digital_text(text, path, log=log)
def read_file_handle_unicode_errors( def read_file_handle_unicode_errors(
filepath: Path, filepath: Path,
log: logging.Logger | None = None, log: logging.Logger | None = None,
-17
View File
@@ -36,23 +36,6 @@ def samples_dir() -> Path:
return (Path(__file__).parent / "samples").resolve() return (Path(__file__).parent / "samples").resolve()
@pytest.fixture(scope="session")
def tagged_no_text_pdf_file(samples_dir: Path) -> Path:
"""Path to a tagged PDF whose only "text" is pdftotext layout padding.
Reproduces GH #13387: ``/MarkInfo /Marked true`` is set, but the only
extractable content is a form-feed byte, not real text. Lives here
rather than in parsers/conftest.py so both parser tests and
paperless/tests/test_parser_utils.py can use it.
Returns
-------
Path
Absolute path to ``tesseract/tagged-but-no-text.pdf``.
"""
return samples_dir / "tesseract" / "tagged-but-no-text.pdf"
@pytest.fixture(autouse=True) @pytest.fixture(autouse=True)
def clean_registry() -> Generator[None, None, None]: def clean_registry() -> Generator[None, None, None]:
"""Reset the parser registry before and after every test. """Reset the parser registry before and after every test.
@@ -21,7 +21,7 @@ from documents.parsers import run_convert
from paperless.models import ModeChoices from paperless.models import ModeChoices
from paperless.parsers import ParserProtocol from paperless.parsers import ParserProtocol
from paperless.parsers.tesseract import RasterisedDocumentParser from paperless.parsers.tesseract import RasterisedDocumentParser
from paperless.parsers.utils import is_tagged_pdf from paperless.parsers.tesseract import post_process_text
if TYPE_CHECKING: if TYPE_CHECKING:
from pathlib import Path from pathlib import Path
@@ -151,6 +151,36 @@ class TestRasterisedDocumentParserLifecycle:
assert tempdir is not None and not tempdir.exists() assert tempdir is not None and not tempdir.exists()
# ---------------------------------------------------------------------------
# post_process_text
# ---------------------------------------------------------------------------
class TestPostProcessText:
@pytest.mark.parametrize(
("source", "expected"),
[
pytest.param(
"simple string",
"simple string",
id="collapse-spaces",
),
pytest.param(
"simple newline\n testing string",
"simple newline\ntesting string",
id="preserve-newline",
),
pytest.param(
"utf-8 строка с пробелами в конце ", # noqa: RUF001
"utf-8 строка с пробелами в конце", # noqa: RUF001
id="utf8-trailing-spaces",
),
],
)
def test_post_process_text(self, source: str, expected: str) -> None:
assert post_process_text(source) == expected
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Page count # Page count
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@@ -880,25 +910,25 @@ class TestSkipArchive:
self, self,
mocker: MockerFixture, mocker: MockerFixture,
tesseract_parser: RasterisedDocumentParser, tesseract_parser: RasterisedDocumentParser,
tagged_no_text_pdf_file: Path, tesseract_samples_dir: Path,
) -> None: ) -> None:
""" """
GIVEN: GIVEN:
- A real PDF that reports itself as tagged (/MarkInfo /Marked - A PDF that reports itself as tagged (/MarkInfo /Marked true) but
true) but whose only pdftotext output is layout padding (a has no actual extractable text (some scanner firmware produces
lone form-feed byte), not real text (see GitHub issue #13387, this see GitHub issue #13349)
originally reported against #13349's tagged-PDF handling)
- Mode: auto, produce_archive=False - Mode: auto, produce_archive=False
WHEN: WHEN:
- Document is parsed - Document is parsed
THEN: THEN:
- The tag alone is not trusted as "has text"; OCRmyPDF still runs - The tag alone is not trusted as "has text"; OCRmyPDF still runs
""" """
assert is_tagged_pdf(tagged_no_text_pdf_file) is True
tesseract_parser.settings.mode = ModeChoices.AUTO tesseract_parser.settings.mode = ModeChoices.AUTO
mocker.patch("paperless.parsers.tesseract.is_tagged_pdf", return_value=True)
mocker.patch.object(tesseract_parser, "extract_text", return_value=None)
mock_ocr = mocker.patch("ocrmypdf.ocr") mock_ocr = mocker.patch("ocrmypdf.ocr")
tesseract_parser.parse( tesseract_parser.parse(
tagged_no_text_pdf_file, tesseract_samples_dir / "multi-page-images.pdf",
"application/pdf", "application/pdf",
produce_archive=False, produce_archive=False,
) )
-110
View File
@@ -4,18 +4,10 @@ from __future__ import annotations
import codecs import codecs
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING
import pytest
from paperless.parsers.utils import is_tagged_pdf from paperless.parsers.utils import is_tagged_pdf
from paperless.parsers.utils import pdf_born_digital_text
from paperless.parsers.utils import post_process_text
from paperless.parsers.utils import read_file_handle_unicode_errors from paperless.parsers.utils import read_file_handle_unicode_errors
if TYPE_CHECKING:
from pytest_mock import MockerFixture
SAMPLES = Path(__file__).parent / "samples" / "tesseract" SAMPLES = Path(__file__).parent / "samples" / "tesseract"
@@ -68,105 +60,3 @@ class TestIsTaggedPdf:
bad = tmp_path / "bad.pdf" bad = tmp_path / "bad.pdf"
bad.write_bytes(b"not a pdf") bad.write_bytes(b"not a pdf")
assert is_tagged_pdf(bad) is False assert is_tagged_pdf(bad) is False
class TestPostProcessText:
@pytest.mark.parametrize(
("source", "expected"),
[
pytest.param(
"simple string",
"simple string",
id="collapse-spaces",
),
pytest.param(
"simple newline\n testing string",
"simple newline\ntesting string",
id="preserve-newline",
),
pytest.param(
"utf-8 строка с пробелами в конце ", # noqa: RUF001
"utf-8 строка с пробелами в конце", # noqa: RUF001
id="utf8-trailing-spaces",
),
pytest.param(None, None, id="none-input"),
pytest.param("", None, id="empty-string"),
pytest.param(" \n\x0c \n ", None, id="whitespace-and-formfeed-only"),
],
)
def test_post_process_text(
self,
source: str | None,
expected: str | None,
) -> None:
assert post_process_text(source) == expected
class TestPdfBornDigitalText:
"""Regression coverage for GH #13387.
should_produce_archive() and RasterisedDocumentParser.parse() must agree
on whether a PDF has real text, so both go through this one function.
"""
@pytest.mark.parametrize(
("extracted", "tagged", "expected_text", "expected_born_digital"),
[
pytest.param("tiny", True, "tiny", True, id="tagged-with-real-text"),
pytest.param("tiny", False, "tiny", False, id="untagged-below-min-length"),
pytest.param(
"x" * 51,
False,
"x" * 51,
True,
id="untagged-above-min-length",
),
pytest.param(None, True, None, False, id="tagged-but-no-text"),
],
)
def test_born_digital_decision(
self,
mocker: MockerFixture,
tmp_path: Path,
extracted: str | None,
tagged: bool, # noqa: FBT001
expected_text: str | None,
expected_born_digital: bool, # noqa: FBT001
) -> None:
"""
GIVEN:
- A PDF whose pdftotext output and /MarkInfo tag status vary
WHEN:
- pdf_born_digital_text() is called
THEN:
- The normalized text and born-digital verdict match; the tag
alone never counts as "has text"
"""
mocker.patch(
"paperless.parsers.utils.extract_pdf_text",
return_value=extracted,
)
mocker.patch("paperless.parsers.utils.is_tagged_pdf", return_value=tagged)
text, born_digital = pdf_born_digital_text(tmp_path / "doc.pdf")
assert text == expected_text
assert born_digital is expected_born_digital
def test_tagged_but_textless_pdf_is_not_born_digital(
self,
tagged_no_text_pdf_file: Path,
) -> None:
"""
GIVEN:
- A real PDF that is tagged (/MarkInfo /Marked true) but whose
only "text" is layout padding (a stray form-feed byte)
WHEN:
- pdf_born_digital_text() is called with no mocking
THEN:
- The normalized text is None and the PDF is not treated as
born-digital. The raw, unnormalized pdftotext output is
non-empty for this file, which is exactly what caused the
archive decision to disagree with the OCR decision in #13387.
"""
text, born_digital = pdf_born_digital_text(tagged_no_text_pdf_file)
assert text is None
assert born_digital is False
+78 -28
View File
@@ -144,6 +144,24 @@ def _exclude_readers():
lock.close() lock.close()
def _with_exclusive_access(operation: str, fn):
"""Run ``fn()`` with exclusive index access (see ``_exclude_readers()``),
for compaction/migration file swaps that must not run while readers are
active. Returns ``fn()``'s result, or None (after logging) if active
readers do not drain within ``LLM_INDEX_COMPACTION_LOCK_TIMEOUT`` --
callers skip the operation this run; it retries next time.
"""
try:
with _exclude_readers():
return fn()
except Timeout:
logger.info(
"Skipping LLM index %s: index readers are active; will retry next run.",
operation,
)
return None
@contextmanager @contextmanager
def write_store(embed_model_name: str | None = None): def write_store(embed_model_name: str | None = None):
"""Acquire the write lock and yield the vector store. """Acquire the write lock and yield the vector store.
@@ -324,6 +342,21 @@ def _exclude_document_id_filter(document_id: int | str):
) )
def _check_and_run_migrations(store: "PaperlessSqliteVecVectorStore") -> bool:
"""Run any pending structural migrations, returning True if a pending
re-embed migration needs the caller to force a rebuild -- never
triggered automatically here. Safe to call before any write, including
delete()/upsert_document(): has_pending_migration() (see its docstring)
keeps this a no-op, with no exclusive access taken, once the store is
current.
"""
if not store.has_pending_migration():
return False
return bool(
_with_exclusive_access("migration check", store.check_and_run_migrations),
)
def update_llm_index( def update_llm_index(
*, *,
iter_wrapper: IterWrapper[Document] = identity, iter_wrapper: IterWrapper[Document] = identity,
@@ -339,15 +372,7 @@ def update_llm_index(
happens, since a rebuild always covers the whole library regardless. happens, since a rebuild always covers the whole library regardless.
""" """
with write_store() as store: with write_store() as store:
try: needs_reembed = _check_and_run_migrations(store)
with _exclude_readers():
needs_reembed = store.check_and_run_migrations()
except Timeout:
logger.info(
"Skipping LLM index migration check: index readers are active; "
"will retry next run.",
)
needs_reembed = False
if needs_reembed: if needs_reembed:
logger.warning( logger.warning(
"LLM index migration requires re-embedding; forcing rebuild.", "LLM index migration requires re-embedding; forcing rebuild.",
@@ -396,12 +421,24 @@ def update_llm_index(
if document_ids is not None if document_ids is not None
else documents else documents
) )
existing = store.get_modified_times() # When document_ids is given, the caller already knows exactly
# which documents to reindex (e.g. a bulk edit) -- trust it and
# skip the modified-time comparison entirely. Bulk edits (tags,
# correspondent, document type, storage path, custom fields)
# write via queryset.update()/M2M bulk operations, which bypass
# Document.modified's auto_now, so comparing against
# get_modified_times() here would silently skip reindexing
# documents whose embedded metadata just changed. The comparison
# is only meaningful for the unscoped, full-library scan, where
# it avoids re-embedding documents that have not changed.
existing = store.get_modified_times() if document_ids is None else None
changed = 0 changed = 0
for document in iter_wrapper(scoped_documents): for document in iter_wrapper(scoped_documents):
doc_id = str(document.id) doc_id = str(document.id)
if existing.get(doc_id) == document.modified.isoformat(): if existing is not None:
continue stored_modified = existing.get(doc_id)
if stored_modified == document.modified.isoformat():
continue
nodes = build_document_node(document, chunk_size=chunk_size) nodes = build_document_node(document, chunk_size=chunk_size)
_embed_nodes(nodes, embed_model) _embed_nodes(nodes, embed_model)
store.upsert_document(doc_id, nodes) store.upsert_document(doc_id, nodes)
@@ -412,14 +449,7 @@ def update_llm_index(
else "No changes detected in LLM index." else "No changes detected in LLM index."
) )
try: _with_exclusive_access("compaction", store.compact)
with _exclude_readers():
store.compact()
except Timeout:
logger.info(
"Skipping LLM index compaction: index readers are active; "
"will retry next run.",
)
return msg return msg
@@ -434,25 +464,45 @@ def llm_index_add_or_update_document(document: Document):
_embed_nodes(new_nodes, get_embedding_model(config)) _embed_nodes(new_nodes, get_embedding_model(config))
with write_store(embed_model_name=get_configured_model_name(config)) as store: with write_store(embed_model_name=get_configured_model_name(config)) as store:
_check_and_run_migrations(store)
store.upsert_document(str(document.id), new_nodes) store.upsert_document(str(document.id), new_nodes)
def llm_index_migrate() -> None:
"""Apply any pending LLM index schema migrations, with no reindex.
Intended to run unconditionally on every startup (see the
init-llmindex-migrate container step and the bare-metal upgrade docs):
has_pending_migration() short-circuits to a metadata-only read once the
store is current, so a healthy install pays almost nothing here. Only
ever applies structural migrations -- a pending re-embed migration is
left for the explicit, deliberate rebuild path (``document_llmindex
update``/``rebuild``) to resolve, since re-embedding can be slow and,
for a metered embedding backend, cost money.
"""
if not AIConfig().llm_index_enabled:
return
with write_store() as store:
needs_reembed = _check_and_run_migrations(store)
if needs_reembed:
logger.warning(
"LLM index requires re-embedding, which this automatic migration "
"check will not do on its own -- it can be slow and, for a "
"metered embedding backend, cost money. Run "
"'document_llmindex rebuild' manually when ready.",
)
def llm_index_compact() -> None: def llm_index_compact() -> None:
"""Compact the index immediately, rebuilding the table to reclaim space.""" """Compact the index immediately, rebuilding the table to reclaim space."""
with write_store() as store: with write_store() as store:
try: _with_exclusive_access("compaction", lambda: store.compact(force=True))
with _exclude_readers():
store.compact(force=True)
except Timeout:
logger.info(
"Skipping LLM index compaction: index readers are active; "
"will retry next run.",
)
def llm_index_remove_document(document: Document): def llm_index_remove_document(document: Document):
"""Remove a document's chunks from the LLM index.""" """Remove a document's chunks from the LLM index."""
with write_store() as store: with write_store() as store:
_check_and_run_migrations(store)
store.delete(str(document.id)) store.delete(str(document.id))
+58
View File
@@ -0,0 +1,58 @@
"""Schema migrations for the sqlite-vec vector store.
Each migration lives in its own module here, named ``mNNNN_description.py``
(e.g. ``m0001_add_document_chunks.py`` -- a leading digit isn't a valid
Python identifier, hence the ``m`` prefix, unlike Django's own numbered
migrations, which load via a dynamic ``importlib.import_module()`` call
rather than a static import statement), and registers itself into
``MIGRATIONS`` at import time. ``vector_store.py`` imports those modules at
the bottom of the file, purely for that registration side effect, after
``PaperlessSqliteVecVectorStore`` is fully defined -- migrations need it to
implement ``apply()`` (see ``Migration`` below).
To add a new migration: add a new ``mNNNN_description.py`` module here that
imports ``PaperlessSqliteVecVectorStore`` from ``paperless_ai.vector_store``,
defines its ``apply()`` (most likely just a call to
``PaperlessSqliteVecVectorStore._rebuild_into()``, see ``Migration`` below),
and appends a ``Migration`` to ``MIGRATIONS``; then import that module at the
bottom of ``vector_store.py`` and bump ``SCHEMA_VERSION`` there.
"""
import sqlite3
from collections.abc import Callable
from dataclasses import dataclass
from dataclasses import field
from typing import Literal
@dataclass
class Migration:
"""A schema migration for the sqlite-vec vector store.
kind="structural": rows are copied into a new-schema file with no
re-embedding needed. Supply ``apply(src_conn, dst_conn, dim)``, which
must create the vec0 table in ``dst_conn`` and copy ``src_conn``'s rows
and index_meta into it -- usually just a call to
``PaperlessSqliteVecVectorStore._rebuild_into(src_conn, dst_conn, dim)``.
``schema_version`` is written by the migration runner after ``apply``
returns, not by ``apply`` itself.
kind="re-embed": the new schema requires fresh embeddings.
``check_and_run_migrations()`` returns True when it encounters one of
these so the caller can force a full rebuild (which recreates the table
at the current SCHEMA_VERSION).
"""
from_version: int
to_version: int
kind: Literal["structural", "re-embed"]
description: str
apply: Callable[[sqlite3.Connection, sqlite3.Connection, int], None] | None = field(
default=None,
repr=False,
)
# Registry of all schema migrations in order, populated by each migration
# module's import-time registration (see the module docstring above).
MIGRATIONS: list[Migration] = []
@@ -0,0 +1,95 @@
import sqlite3
from paperless_ai.migrations import MIGRATIONS
from paperless_ai.migrations import Migration
from paperless_ai.vector_store import COMPACT_BATCH_SIZE
from paperless_ai.vector_store import DEFAULT_TABLE_NAME
from paperless_ai.vector_store import PaperlessSqliteVecVectorStore
def _migrate_v1_to_v2_add_document_chunks(
src_conn: sqlite3.Connection,
dst_conn: sqlite3.Connection,
dim: int,
) -> None:
"""v1 -> v2: backfill the document_chunks side table.
document_chunks (see PaperlessSqliteVecVectorStore._open_connection) lets
delete()/upsert_document() find a document's chunk ids without a vec0
full table scan on the document_id metadata column. Every row written
before this migration predates that table, so without backfilling,
deleting a pre-migration document would find zero chunk ids and leave its
vec0 rows permanently orphaned.
Deliberately spells out its own v2-shaped vec0 table and row copy,
rather than delegating to PaperlessSqliteVecVectorStore's
_create_vec_table()/_rebuild_into()/_copy_rows(): those always reflect
whatever the *current* schema is. If a later migration changes that
schema (bumping SCHEMA_VERSION again), this migration must keep
producing its own historical v2 shape regardless -- otherwise a user
upgrading across multiple versions in one go (e.g. v1 straight to v4)
would have this migration silently produce a v4-shaped table instead
of v2, and the v2 -> v3 migration that runs right after it would find
the columns it expects to migrate *from* already gone.
"""
dst_conn.execute( # nosemgrep: python.sqlalchemy.security.sqlalchemy-execute-raw-query.sqlalchemy-execute-raw-query
"CREATE VIRTUAL TABLE "
+ DEFAULT_TABLE_NAME
+ " USING vec0("
+ "id TEXT PRIMARY KEY,"
+ " document_id TEXT,"
+ " modified TEXT,"
+ " +node_content TEXT,"
+ " embedding float["
+ str(int(dim))
+ "] distance_metric=cosine"
+ ")",
)
PaperlessSqliteVecVectorStore._meta_set_on(dst_conn, "dim", str(dim))
embed_model = PaperlessSqliteVecVectorStore._meta_get_on(src_conn, "embed_model")
if embed_model is not None:
PaperlessSqliteVecVectorStore._meta_set_on(dst_conn, "embed_model", embed_model)
dst_conn.execute("BEGIN IMMEDIATE")
src_cursor = src_conn.execute(
"SELECT id, document_id, modified, node_content, embedding FROM "
+ DEFAULT_TABLE_NAME,
)
live = 0
while batch := src_cursor.fetchmany(COMPACT_BATCH_SIZE):
dst_conn.executemany(
"INSERT INTO "
+ DEFAULT_TABLE_NAME
+ " (id, document_id, modified, node_content, embedding) "
"VALUES (?, ?, ?, ?, ?)",
[
(
r["id"],
r["document_id"],
r["modified"],
r["node_content"],
bytes(r["embedding"]),
)
for r in batch
],
)
dst_conn.executemany(
"INSERT INTO document_chunks (chunk_id, document_id) VALUES (?, ?)",
[(r["id"], r["document_id"]) for r in batch],
)
live += len(batch)
# This migration only ever copies live rows (like compact()), so the
# cumulative counter resets to match -- the new file has no bloat yet.
PaperlessSqliteVecVectorStore._meta_set_on(dst_conn, "total_inserts", str(live))
dst_conn.execute("COMMIT")
MIGRATIONS.append(
Migration(
from_version=1,
to_version=2,
kind="structural",
description="add document_chunks side table for O(1) per-document deletes",
apply=_migrate_v1_to_v2_add_document_chunks,
),
)
+233 -36
View File
@@ -21,6 +21,7 @@ from documents.signals import document_consumption_finished
from documents.signals import document_updated from documents.signals import document_updated
from documents.tests.factories import DocumentFactory from documents.tests.factories import DocumentFactory
from documents.tests.factories import PaperlessTaskFactory from documents.tests.factories import PaperlessTaskFactory
from documents.tests.factories import TagFactory
from paperless.models import ApplicationConfiguration from paperless.models import ApplicationConfiguration
from paperless_ai import indexing from paperless_ai import indexing
from paperless_ai.tests.conftest import FakeEmbedding from paperless_ai.tests.conftest import FakeEmbedding
@@ -36,6 +37,23 @@ def real_document(db: None) -> Document:
) )
@pytest.fixture
def mock_store(mocker: pytest_mock.MockerFixture) -> MagicMock:
"""The MagicMock store yielded by every ``with write_store() as store:``
block, for tests that only care what indexing.py does with the store,
not what the store itself does.
"""
store = mocker.MagicMock()
mocker.patch(
"paperless_ai.indexing.write_store",
return_value=mocker.MagicMock(
__enter__=mocker.MagicMock(return_value=store),
__exit__=mocker.MagicMock(return_value=False),
),
)
return store
@pytest.mark.django_db @pytest.mark.django_db
def test_build_document_node(real_document: Document) -> None: def test_build_document_node(real_document: Document) -> None:
nodes = indexing.build_document_node(real_document) nodes = indexing.build_document_node(real_document)
@@ -347,6 +365,62 @@ def test_update_llm_index_partial_update(
assert after[str(doc2.pk)] == before[str(doc2.pk)] assert after[str(doc2.pk)] == before[str(doc2.pk)]
class TestUpdateLlmIndexScopedDocumentIds:
"""A document_ids-scoped update must trust the caller and reindex every scoped
document, never gating on Document.modified.
bulk_edit.py's add_tag/remove_tag/modify_tags/set_correspondent/set_document_type/
set_storage_path/modify_custom_fields all write via queryset.update() or direct
M2M/through-model bulk operations -- none of which call Document.save(), so
Document.modified's auto_now never fires. Comparing against
get_modified_times() for a document_ids-scoped call would therefore skip
reindexing documents whose embedded tags/correspondent/etc. just changed.
"""
@pytest.mark.django_db
def test_scoped_update_reindexes_despite_unchanged_modified(
self,
temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding,
) -> None:
"""A document_ids-scoped update must pick up a tag added via the M2M
manager's bulk_create path, even though that leaves modified untouched.
Steps:
1. Build an initial index for a document with no tags.
2. Add a tag via a direct through-model bulk_create, mirroring
bulk_edit.add_tag -- this does not call Document.save().
3. Call update_llm_index(rebuild=False, document_ids=[doc.pk]).
4. Assert the stored node metadata now includes the new tag.
"""
# Step 1
tag = TagFactory.create(name="Important")
doc = DocumentFactory.create(title="Test Document", added=timezone.now())
indexing.update_llm_index(rebuild=True)
modified_before = doc.modified
# Step 2: bulk-add the tag the way bulk_edit.add_tag does -- a direct
# through-model insert, no Document.save().
DocumentTagRelationship = Document.tags.through
DocumentTagRelationship.objects.bulk_create(
[DocumentTagRelationship(document_id=doc.pk, tag_id=tag.pk)],
)
doc.refresh_from_db()
assert doc.modified == modified_before, (
"Precondition failed: expected modified to be unchanged after a "
"through-model bulk tag add"
)
# Step 3
result = indexing.update_llm_index(rebuild=False, document_ids=[doc.pk])
assert result == "LLM index updated successfully."
# Step 4
with indexing.get_vector_store() as store:
nodes = store.get_nodes()
assert any(tag.name in node.metadata.get("tags", []) for node in nodes)
@pytest.mark.django_db @pytest.mark.django_db
def test_add_or_update_document_updates_existing_entry( def test_add_or_update_document_updates_existing_entry(
temp_llm_index_dir: Path, temp_llm_index_dir: Path,
@@ -705,23 +779,104 @@ class TestLlmIndexAddOrUpdateDocumentEmptyContent:
@pytest.mark.django_db @pytest.mark.django_db
def test_llm_index_compact_uses_force( def test_llm_index_compact_uses_force(
temp_llm_index_dir: Path, temp_llm_index_dir: Path,
mocker: pytest_mock.MockerFixture, mock_store: MagicMock,
) -> None: ) -> None:
"""compact must use force=True to rebuild the table and reclaim space immediately.""" """compact must use force=True to rebuild the table and reclaim space immediately."""
mock_store = mocker.MagicMock()
mocker.patch(
"paperless_ai.indexing.write_store",
return_value=mocker.MagicMock(
__enter__=mocker.MagicMock(return_value=mock_store),
__exit__=mocker.MagicMock(return_value=False),
),
)
indexing.llm_index_compact() indexing.llm_index_compact()
mock_store.compact.assert_called_once_with(force=True) mock_store.compact.assert_called_once_with(force=True)
@pytest.mark.django_db
class TestLlmIndexMigrate:
"""llm_index_migrate() is the cheap, startup-safe migration check -- see
the init-llmindex-migrate container step and the bare-metal upgrade docs.
"""
def test_skips_when_llm_index_disabled(
self,
temp_llm_index_dir: Path,
mocker: pytest_mock.MockerFixture,
) -> None:
"""
GIVEN:
- The LLM index is disabled
WHEN:
- llm_index_migrate() is called
THEN:
- The store is never opened (no stray db file for users who
never enabled AI features)
"""
mock_config = mocker.MagicMock()
mock_config.llm_index_enabled = False
mocker.patch("paperless_ai.indexing.AIConfig", return_value=mock_config)
mock_write_store = mocker.patch("paperless_ai.indexing.write_store")
indexing.llm_index_migrate()
mock_write_store.assert_not_called()
def test_runs_pending_structural_migration(
self,
temp_llm_index_dir: Path,
mocker: pytest_mock.MockerFixture,
mock_store: MagicMock,
caplog: pytest.LogCaptureFixture,
) -> None:
"""
GIVEN:
- The LLM index is enabled and a structural migration is pending
WHEN:
- llm_index_migrate() is called
THEN:
- check_and_run_migrations() runs, and no re-embed warning is
logged (the pending migration was structural, not re-embed)
"""
mock_config = mocker.MagicMock()
mock_config.llm_index_enabled = True
mocker.patch("paperless_ai.indexing.AIConfig", return_value=mock_config)
mock_store.has_pending_migration.return_value = True
mock_store.check_and_run_migrations.return_value = False
with caplog.at_level("WARNING"):
indexing.llm_index_migrate()
mock_store.check_and_run_migrations.assert_called_once()
assert "re-embedding" not in caplog.text
def test_warns_without_rebuilding_when_reembed_pending(
self,
temp_llm_index_dir: Path,
mocker: pytest_mock.MockerFixture,
mock_store: MagicMock,
caplog: pytest.LogCaptureFixture,
) -> None:
"""
GIVEN:
- The LLM index is enabled and a pending migration requires
re-embedding
WHEN:
- llm_index_migrate() is called
THEN:
- A warning is logged telling the admin to rebuild manually, but
no rebuild is triggered automatically -- re-embedding can be
slow and, for a metered embedding backend, cost money, so it
must be a deliberate user action, never an automatic one
"""
mock_config = mocker.MagicMock()
mock_config.llm_index_enabled = True
mocker.patch("paperless_ai.indexing.AIConfig", return_value=mock_config)
mock_store.has_pending_migration.return_value = True
mock_store.check_and_run_migrations.return_value = True
with caplog.at_level("WARNING"):
indexing.llm_index_migrate()
assert "re-embedding" in caplog.text
mock_store.drop_table.assert_not_called()
mock_store.add.assert_not_called()
@pytest.mark.django_db @pytest.mark.django_db
class TestLlmIndexLocking: class TestLlmIndexLocking:
"""Index mutation functions must go through write_store(), which holds the lock. """Index mutation functions must go through write_store(), which holds the lock.
@@ -734,16 +889,9 @@ class TestLlmIndexLocking:
self, self,
temp_llm_index_dir: Path, temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding, mock_embed_model: FakeEmbedding,
mock_store: MagicMock,
mocker: pytest_mock.MockerFixture, mocker: pytest_mock.MockerFixture,
) -> None: ) -> None:
mock_store = MagicMock()
mocker.patch(
"paperless_ai.indexing.write_store",
return_value=mocker.MagicMock(
__enter__=mocker.MagicMock(return_value=mock_store),
__exit__=mocker.MagicMock(return_value=False),
),
)
mock_node = MagicMock() mock_node = MagicMock()
mock_node.get_content.return_value = "fake node text" mock_node.get_content.return_value = "fake node text"
mocker.patch( mocker.patch(
@@ -757,40 +905,89 @@ class TestLlmIndexLocking:
mock_store.upsert_document.assert_called_once() mock_store.upsert_document.assert_called_once()
@pytest.mark.parametrize("has_pending", [True, False])
def test_add_or_update_document_runs_migration_check_only_when_pending(
self,
temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding,
mock_store: MagicMock,
mocker: pytest_mock.MockerFixture,
*,
has_pending: bool,
) -> None:
"""
GIVEN:
- A document to add/update, and a store reporting whether a
migration is pending
WHEN:
- llm_index_add_or_update_document() is called
THEN:
- check_and_run_migrations() runs only when has_pending_migration()
is True, so a normal upsert never pays for the exclusive access
that check_and_run_migrations() requires
- upsert_document() is called either way
"""
mock_store.has_pending_migration.return_value = has_pending
mock_node = MagicMock()
mock_node.get_content.return_value = "fake node text"
mocker.patch(
"paperless_ai.indexing.build_document_node",
return_value=[mock_node],
)
doc = DocumentFactory.build(id=1)
indexing.llm_index_add_or_update_document(doc)
assert mock_store.check_and_run_migrations.called is has_pending
mock_store.upsert_document.assert_called_once()
def test_remove_document_uses_write_store( def test_remove_document_uses_write_store(
self, self,
temp_llm_index_dir: Path, temp_llm_index_dir: Path,
mocker: pytest_mock.MockerFixture, mock_store: MagicMock,
) -> None: ) -> None:
mock_store = MagicMock()
mocker.patch(
"paperless_ai.indexing.write_store",
return_value=mocker.MagicMock(
__enter__=mocker.MagicMock(return_value=mock_store),
__exit__=mocker.MagicMock(return_value=False),
),
)
doc = MagicMock(spec=Document) doc = MagicMock(spec=Document)
doc.id = 1 doc.id = 1
indexing.llm_index_remove_document(doc) indexing.llm_index_remove_document(doc)
mock_store.delete.assert_called_once_with("1") mock_store.delete.assert_called_once_with("1")
@pytest.mark.parametrize("has_pending", [True, False])
def test_remove_document_runs_migration_check_only_when_pending(
self,
temp_llm_index_dir: Path,
mock_store: MagicMock,
*,
has_pending: bool,
) -> None:
"""
GIVEN:
- A document to remove, and a store reporting whether a
migration is pending
WHEN:
- llm_index_remove_document() is called
THEN:
- check_and_run_migrations() runs only when has_pending_migration()
is True, so a normal delete never pays for the exclusive access
that check_and_run_migrations() requires (see
test_normal_write_is_not_gated_by_the_compaction_lock)
- delete() is called either way
"""
mock_store.has_pending_migration.return_value = has_pending
doc = DocumentFactory.build(id=1)
indexing.llm_index_remove_document(doc)
assert mock_store.check_and_run_migrations.called is has_pending
mock_store.delete.assert_called_once_with("1")
def test_update_llm_index_rebuild_uses_write_store( def test_update_llm_index_rebuild_uses_write_store(
self, self,
temp_llm_index_dir: Path, temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding, mock_embed_model: FakeEmbedding,
mock_store: MagicMock,
mocker: pytest_mock.MockerFixture, mocker: pytest_mock.MockerFixture,
) -> None: ) -> None:
mock_store = MagicMock()
mocker.patch(
"paperless_ai.indexing.write_store",
return_value=mocker.MagicMock(
__enter__=mocker.MagicMock(return_value=mock_store),
__exit__=mocker.MagicMock(return_value=False),
),
)
mock_qs = MagicMock() mock_qs = MagicMock()
mock_qs.exists.return_value = True mock_qs.exists.return_value = True
mock_qs.__iter__ = MagicMock(return_value=iter([])) mock_qs.__iter__ = MagicMock(return_value=iter([]))
File diff suppressed because it is too large Load Diff
+216 -121
View File
@@ -2,16 +2,12 @@ import json
import logging import logging
import sqlite3 import sqlite3
import struct import struct
from collections.abc import Callable
from collections.abc import Iterator from collections.abc import Iterator
from collections.abc import Sequence from collections.abc import Sequence
from contextlib import contextmanager from contextlib import contextmanager
from dataclasses import dataclass
from dataclasses import field
from pathlib import Path from pathlib import Path
from types import TracebackType from types import TracebackType
from typing import Any from typing import Any
from typing import Literal
import sqlite_vec import sqlite_vec
from llama_index.core.bridge.pydantic import PrivateAttr from llama_index.core.bridge.pydantic import PrivateAttr
@@ -26,15 +22,27 @@ from llama_index.core.vector_stores.types import VectorStoreQueryResult
from llama_index.core.vector_stores.utils import metadata_dict_to_node from llama_index.core.vector_stores.utils import metadata_dict_to_node
from llama_index.core.vector_stores.utils import node_to_metadata_dict from llama_index.core.vector_stores.utils import node_to_metadata_dict
from paperless_ai.migrations import MIGRATIONS
from paperless_ai.migrations import Migration
logger = logging.getLogger("paperless_ai.vector_store") logger = logging.getLogger("paperless_ai.vector_store")
DB_FILENAME = "llmindex.db" DB_FILENAME = "llmindex.db"
DEFAULT_TABLE_NAME = "documents" DEFAULT_TABLE_NAME = "documents"
# Current schema version. Written to index_meta at table creation and bumped _INSERT = (
# whenever a Migration is added to MIGRATIONS. check_and_run_migrations() uses "INSERT INTO "
# this to decide which migrations to run on an existing store. + DEFAULT_TABLE_NAME
SCHEMA_VERSION = 1 + " (id, document_id, modified, node_content, embedding) VALUES (?, ?, ?, ?, ?)"
)
_INSERT_CHUNK_INDEX = (
"INSERT INTO document_chunks (chunk_id, document_id) VALUES (?, ?)"
)
# Current schema version. Bump when adding a migration -- see
# paperless_ai/migrations/__init__.py for the full procedure.
SCHEMA_VERSION = 2
# compact(): rebuild when the cumulative rowid count exceeds this multiple of # compact(): rebuild when the cumulative rowid count exceeds this multiple of
# the live row count. DELETEs on vec0 tables never reclaim space (upstream # the live row count. DELETEs on vec0 tables never reclaim space (upstream
@@ -53,38 +61,6 @@ COMPACT_BATCH_SIZE = 500
_FILTER_COLUMNS = frozenset({"document_id", "modified"}) _FILTER_COLUMNS = frozenset({"document_id", "modified"})
@dataclass
class Migration:
"""A schema migration for the sqlite-vec vector store.
kind="structural": rows are copied into a new-schema file with no
re-embedding needed. Supply ``apply(src_conn, dst_conn, dim)`` which
must create the vec0 table in ``dst_conn``, copy all rows from
``src_conn``, and write ``dim`` / ``embed_model`` / ``total_inserts`` to
``dst_conn``'s ``index_meta``. ``schema_version`` is written by the
migration runner after ``apply`` returns.
kind="re-embed": the new schema requires fresh embeddings.
``check_and_run_migrations()`` returns True when it encounters one of
these so the caller can force a full rebuild (which recreates the table
at the current SCHEMA_VERSION).
"""
from_version: int
to_version: int
kind: Literal["structural", "re-embed"]
description: str
apply: Callable[[sqlite3.Connection, sqlite3.Connection, int], None] | None = field(
default=None,
repr=False,
)
# Registry of all schema migrations in order. Empty at v1 -- this is the
# baseline. Add entries here (and bump SCHEMA_VERSION) when the schema changes.
MIGRATIONS: list[Migration] = []
def _pack(embedding: Sequence[float]) -> bytes: def _pack(embedding: Sequence[float]) -> bytes:
return struct.pack(f"{len(embedding)}f", *embedding) return struct.pack(f"{len(embedding)}f", *embedding)
@@ -93,6 +69,42 @@ def _unpack(blob: bytes) -> list[float]:
return list(struct.unpack(f"{len(blob) // 4}f", blob)) return list(struct.unpack(f"{len(blob) // 4}f", blob))
def _copy_rows(src_conn: sqlite3.Connection, dst_conn: sqlite3.Connection) -> int:
"""Copy every live vec0 row from ``src_conn`` into ``dst_conn``, recording
each one in ``dst_conn``'s document_chunks side table. Returns the number
of rows copied. The caller owns ``dst_conn``'s transaction.
Rows are streamed from the source cursor in batches instead of being
materialized all at once, so a large index does not cause an OOM during a
routine compaction or migration.
"""
src_cursor = src_conn.execute(
"SELECT id, document_id, modified, node_content, embedding FROM "
+ DEFAULT_TABLE_NAME,
)
copied = 0
while batch := src_cursor.fetchmany(COMPACT_BATCH_SIZE):
dst_conn.executemany(
_INSERT,
[
(
r["id"],
r["document_id"],
r["modified"],
r["node_content"],
bytes(r["embedding"]),
)
for r in batch
],
)
dst_conn.executemany(
_INSERT_CHUNK_INDEX,
[(r["id"], r["document_id"]) for r in batch],
)
copied += len(batch)
return copied
def _build_where(filters: MetadataFilters | None) -> tuple[str, list[str]]: def _build_where(filters: MetadataFilters | None) -> tuple[str, list[str]]:
"""Translate the EQ / IN / NE filters we use into a parameterized SQL clause """Translate the EQ / IN / NE filters we use into a parameterized SQL clause
on vec0 metadata columns. Returns ("", []) when there is nothing to filter. on vec0 metadata columns. Returns ("", []) when there is nothing to filter.
@@ -189,6 +201,24 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
conn.execute( conn.execute(
"CREATE TABLE IF NOT EXISTS index_meta (key TEXT PRIMARY KEY, value TEXT)", "CREATE TABLE IF NOT EXISTS index_meta (key TEXT PRIMARY KEY, value TEXT)",
) )
# vec0 metadata columns only get an efficient lookup path inside a KNN
# (MATCH) query; a plain `WHERE document_id = ?` is a full table scan
# regardless of index size. This plain, indexed table is how delete()/
# upsert_document() find a document's chunk ids without that scan.
# document_id is INTEGER here (unlike vec0's own TEXT metadata
# column): this is a normal SQLite table, so standard type affinity
# correctly coerces the TEXT document ids written/looked-up
# elsewhere in this module -- it does not share vec0's own
# metadata-column comparison code, which silently mismatches
# non-TEXT bound values instead of coercing them.
conn.execute(
"CREATE TABLE IF NOT EXISTS document_chunks "
"(chunk_id TEXT PRIMARY KEY, document_id INTEGER NOT NULL)",
)
conn.execute(
"CREATE INDEX IF NOT EXISTS idx_document_chunks_document_id "
"ON document_chunks (document_id)",
)
return conn return conn
@property @property
@@ -223,13 +253,17 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
else: else:
self._conn.execute("COMMIT") self._conn.execute("COMMIT")
def _meta_get(self, key: str) -> str | None: @staticmethod
row = self._conn.execute( def _meta_get_on(conn: sqlite3.Connection, key: str) -> str | None:
row = conn.execute(
"SELECT value FROM index_meta WHERE key = ?", "SELECT value FROM index_meta WHERE key = ?",
(key,), (key,),
).fetchone() ).fetchone()
return row["value"] if row else None return row["value"] if row else None
def _meta_get(self, key: str) -> str | None:
return self._meta_get_on(self._conn, key)
@staticmethod @staticmethod
def _meta_set_on(conn: sqlite3.Connection, key: str, value: str) -> None: def _meta_set_on(conn: sqlite3.Connection, key: str, value: str) -> None:
conn.execute( conn.execute(
@@ -259,6 +293,7 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
def drop_table(self) -> None: def drop_table(self) -> None:
self._conn.execute("DROP TABLE IF EXISTS " + DEFAULT_TABLE_NAME) self._conn.execute("DROP TABLE IF EXISTS " + DEFAULT_TABLE_NAME)
self._conn.execute("DELETE FROM index_meta") self._conn.execute("DELETE FROM index_meta")
self._conn.execute("DELETE FROM document_chunks")
def stored_model_name(self) -> str | None: def stored_model_name(self) -> str | None:
"""Return the embedding model name recorded at table creation, or None.""" """Return the embedding model name recorded at table creation, or None."""
@@ -325,11 +360,37 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
_pack(node.get_embedding()), _pack(node.get_embedding()),
) )
_INSERT = ( def _index_chunks(self, rows: list[tuple[str, str, str, str, bytes]]) -> None:
"INSERT INTO " """Record each row's (chunk_id, document_id) in the document_chunks
+ DEFAULT_TABLE_NAME side table, kept in lockstep with every insert into the vec0 table."""
+ " (id, document_id, modified, node_content, embedding) VALUES (?, ?, ?, ?, ?)" self._conn.executemany(
) _INSERT_CHUNK_INDEX,
[(chunk_id, document_id) for chunk_id, document_id, *_ in rows],
)
def _delete_chunks_by_document_id(self, document_id: str) -> None:
"""Delete all of a document's chunks via point-deletes on `id`.
vec0 has no efficient lookup on the document_id metadata column
outside a KNN query (see _open_connection), so a plain
`DELETE ... WHERE document_id = ?` is a full table scan regardless of
index size. Looking the chunk ids up in document_chunks first (a real
indexed lookup) and deleting each by its `id` primary key instead
turns that scan into a handful of O(1) point deletes.
"""
doc_id = str(document_id)
chunk_rows = self._conn.execute(
"SELECT chunk_id FROM document_chunks WHERE document_id = ?",
(doc_id,),
).fetchall()
self._conn.executemany(
"DELETE FROM " + DEFAULT_TABLE_NAME + " WHERE id = ?",
[(row["chunk_id"],) for row in chunk_rows],
)
self._conn.execute(
"DELETE FROM document_chunks WHERE document_id = ?",
(doc_id,),
)
def _increment_total_inserts(self, count: int) -> None: def _increment_total_inserts(self, count: int) -> None:
"""Increment the cumulative insert counter stored in index_meta. """Increment the cumulative insert counter stored in index_meta.
@@ -348,7 +409,8 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
rows = [self._row(node) for node in nodes] rows = [self._row(node) for node in nodes]
with self._transaction(): with self._transaction():
self._ensure_table(len(nodes[0].get_embedding())) self._ensure_table(len(nodes[0].get_embedding()))
self._conn.executemany(self._INSERT, rows) self._conn.executemany(_INSERT, rows)
self._index_chunks(rows)
self._increment_total_inserts(len(rows)) self._increment_total_inserts(len(rows))
return [node.node_id for node in nodes] return [node.node_id for node in nodes]
@@ -365,22 +427,17 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
if nodes: if nodes:
self._ensure_table(len(nodes[0].get_embedding())) self._ensure_table(len(nodes[0].get_embedding()))
if self.table_exists(): if self.table_exists():
self._conn.execute( self._delete_chunks_by_document_id(document_id)
"DELETE FROM " + DEFAULT_TABLE_NAME + " WHERE document_id = ?",
(str(document_id),),
)
if rows: if rows:
self._conn.executemany(self._INSERT, rows) self._conn.executemany(_INSERT, rows)
self._index_chunks(rows)
self._increment_total_inserts(len(rows)) self._increment_total_inserts(len(rows))
return [node.node_id for node in nodes] return [node.node_id for node in nodes]
def delete(self, ref_doc_id: str, **delete_kwargs: Any) -> None: def delete(self, ref_doc_id: str, **delete_kwargs: Any) -> None:
if self.table_exists(): if self.table_exists():
with self._transaction(): with self._transaction():
self._conn.execute( self._delete_chunks_by_document_id(ref_doc_id)
"DELETE FROM " + DEFAULT_TABLE_NAME + " WHERE document_id = ?",
(str(ref_doc_id),),
)
def _rows_to_nodes(self, rows: list[sqlite3.Row]) -> list[BaseNode]: def _rows_to_nodes(self, rows: list[sqlite3.Row]) -> list[BaseNode]:
nodes: list[BaseNode] = [] nodes: list[BaseNode] = []
@@ -464,6 +521,67 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
result[doc_id] = str(row["modified"] or "") result[doc_id] = str(row["modified"] or "")
return result return result
@property
def _db_path(self) -> str:
return str(Path(self._uri) / DB_FILENAME)
@contextmanager
def _rebuild_file(self) -> Iterator[sqlite3.Connection]:
"""Open a fresh temp database file for a file-swap rebuild (compact
or structural migration), yielding its connection for the caller to
populate.
On success, swaps the temp file in as the live database (closing
this store's current connection first -- see _swap_in_compact()).
On any exception, discards the temp file, including its -wal/-shm,
instead, and this store's own connection is left untouched.
"""
compact_path = self._db_path + ".compact"
new_conn = self._open_connection(compact_path)
try:
yield new_conn
except BaseException:
new_conn.close()
for suffix in ["", "-wal", "-shm"]:
Path(compact_path + suffix).unlink(missing_ok=True)
raise
else:
new_conn.close()
self._swap_in_compact(compact_path, self._db_path)
@staticmethod
def _rebuild_into(
src_conn: sqlite3.Connection,
dst_conn: sqlite3.Connection,
dim: int,
meta_keys: tuple[str, ...] = ("dim", "embed_model", "schema_version"),
) -> int:
"""Create the vec0 table in ``dst_conn``, copy ``meta_keys`` from
``src_conn``'s index_meta, and stream every live row across
(populating document_chunks as it goes -- see _copy_rows()).
Returns the number of rows copied.
Used by both compact() (default meta_keys: the schema is unchanged)
and structural migrations (meta_keys minus "schema_version", which
the migration sets to its own target version instead of preserving
the source's).
"""
PaperlessSqliteVecVectorStore._create_vec_table(dst_conn, dim)
for key in meta_keys:
value = PaperlessSqliteVecVectorStore._meta_get_on(src_conn, key)
if value is not None:
PaperlessSqliteVecVectorStore._meta_set_on(dst_conn, key, value)
dst_conn.execute("BEGIN IMMEDIATE")
copied = _copy_rows(src_conn, dst_conn)
# Reset the cumulative counter: after a rebuild, total_inserts == live.
PaperlessSqliteVecVectorStore._meta_set_on(
dst_conn,
"total_inserts",
str(copied),
)
dst_conn.execute("COMMIT")
return copied
def compact(self, *, force: bool = False) -> None: def compact(self, *, force: bool = False) -> None:
"""Rebuild the database file to reclaim space left behind by DELETEs. """Rebuild the database file to reclaim space left behind by DELETEs.
@@ -496,50 +614,8 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
live, live,
total, total,
) )
db_path = str(Path(self._uri) / DB_FILENAME) with self._rebuild_file() as new_conn:
compact_path = db_path + ".compact" self._rebuild_into(self._conn, new_conn, dim)
# Copy all live rows into a fresh database file.
new_conn = self._open_connection(compact_path)
try:
self._create_vec_table(new_conn, dim)
self._meta_set_on(new_conn, "dim", str(dim))
for key in ("embed_model", "schema_version"):
value = self._meta_get(key)
if value is not None:
self._meta_set_on(new_conn, key, value)
src_cursor = self._conn.execute(
"SELECT id, document_id, modified, node_content, embedding "
"FROM " + DEFAULT_TABLE_NAME,
)
new_conn.execute("BEGIN IMMEDIATE")
# Stream rows from the source cursor in batches instead of
# materializing the whole table in memory, so a large index does
# not cause an OOM during routine maintenance compactions.
while batch := src_cursor.fetchmany(COMPACT_BATCH_SIZE):
new_conn.executemany(
self._INSERT,
[
(
r["id"],
r["document_id"],
r["modified"],
r["node_content"],
bytes(r["embedding"]),
)
for r in batch
],
)
# Reset the cumulative counter: after compact, total_inserts == live.
self._meta_set_on(new_conn, "total_inserts", str(live))
new_conn.execute("COMMIT")
except BaseException:
new_conn.close()
for p in [compact_path, compact_path + "-wal", compact_path + "-shm"]:
Path(p).unlink(missing_ok=True)
raise
new_conn.close()
self._swap_in_compact(compact_path, db_path)
def _swap_in_compact(self, compact_path: str, db_path: str) -> None: def _swap_in_compact(self, compact_path: str, db_path: str) -> None:
"""Atomically replace the live database with the compacted copy.""" """Atomically replace the live database with the compacted copy."""
@@ -551,6 +627,31 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
Path(compact_path).replace(db_path) Path(compact_path).replace(db_path)
self._conn = self._open_connection(db_path) self._conn = self._open_connection(db_path)
def _stored_schema_version(self) -> int | None:
"""The schema_version recorded in index_meta, or None if no table
exists. A missing key (a store predating version tracking) is
treated as SCHEMA_VERSION -- i.e. already current -- since no
migration in MIGRATIONS targets a version before tracking began.
"""
if not self.table_exists():
return None
raw = self._meta_get("schema_version")
return int(raw) if raw is not None else SCHEMA_VERSION
def has_pending_migration(self) -> bool:
"""Cheaply check whether a migration is pending, with no exclusive
access needed -- just a metadata read under the connection callers
already hold via the write FileLock.
Callers should only pay for check_and_run_migrations()'s exclusive
access (a structural migration's file swap must not run while
readers are active) when this returns True, so that the common
case -- already at SCHEMA_VERSION -- never contends with readers
or a concurrent compaction.
"""
current = self._stored_schema_version()
return current is not None and current < SCHEMA_VERSION
def check_and_run_migrations(self) -> bool: def check_and_run_migrations(self) -> bool:
"""Apply any pending schema migrations to the store. """Apply any pending schema migrations to the store.
@@ -559,15 +660,13 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
this method returns True when one is encountered so the caller can this method returns True when one is encountered so the caller can
force a full rebuild (which recreates the table at SCHEMA_VERSION). force a full rebuild (which recreates the table at SCHEMA_VERSION).
Must be called under the write FileLock. No-op when the table does Must be called under the write FileLock, with readers excluded (see
not exist or is already at SCHEMA_VERSION. has_pending_migration() for a cheap pre-check that avoids paying for
that exclusion in the common case). No-op when the table does not
exist or is already at SCHEMA_VERSION.
""" """
if not self.table_exists(): current = self._stored_schema_version()
return False if current is None or current >= SCHEMA_VERSION:
raw = self._meta_get("schema_version")
current = int(raw) if raw is not None else SCHEMA_VERSION
if current >= SCHEMA_VERSION:
return False return False
pending = sorted( pending = sorted(
@@ -579,7 +678,7 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
if migration.kind == "re-embed": if migration.kind == "re-embed":
logger.warning( logger.warning(
"LLM index schema v%d -> v%d requires re-embedding (%s); " "LLM index schema v%d -> v%d requires re-embedding (%s); "
"forcing full rebuild.", "the caller must force a rebuild.",
migration.from_version, migration.from_version,
migration.to_version, migration.to_version,
migration.description, migration.description,
@@ -601,16 +700,12 @@ class PaperlessSqliteVecVectorStore(BasePydanticVectorStore):
dim = self.vector_dim() dim = self.vector_dim()
if dim is None: # pragma: no cover if dim is None: # pragma: no cover
raise RuntimeError("Cannot migrate: no stored vector dimension") raise RuntimeError("Cannot migrate: no stored vector dimension")
db_path = str(Path(self._uri) / DB_FILENAME) with self._rebuild_file() as new_conn:
compact_path = db_path + ".compact"
new_conn = self._open_connection(compact_path)
try:
migration.apply(self._conn, new_conn, dim) migration.apply(self._conn, new_conn, dim)
self._meta_set_on(new_conn, "schema_version", str(migration.to_version)) self._meta_set_on(new_conn, "schema_version", str(migration.to_version))
except BaseException: # pragma: no cover
new_conn.close()
for p in [compact_path, compact_path + "-wal", compact_path + "-shm"]: # Registers m0001 into MIGRATIONS; must be at the bottom (needs
Path(p).unlink(missing_ok=True) # PaperlessSqliteVecVectorStore fully defined) -- see
raise # paperless_ai/migrations/__init__.py for the full procedure.
new_conn.close() from paperless_ai.migrations import m0001_add_document_chunks # noqa: E402, F401
self._swap_in_compact(compact_path, db_path)
Generated
+3 -3
View File
@@ -3735,15 +3735,15 @@ crypto = [
[[package]] [[package]]
name = "pymdown-extensions" name = "pymdown-extensions"
version = "11.0" version = "10.21.3"
source = { registry = "https://pypi.org/simple" } source = { registry = "https://pypi.org/simple" }
dependencies = [ dependencies = [
{ name = "markdown", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" }, { name = "markdown", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
{ name = "pyyaml", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" }, { name = "pyyaml", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
] ]
sdist = { url = "https://files.pythonhosted.org/packages/47/67/f1e79672a5f91985577c7984c9709ca110e4fd37fe7fd167b60422e6ccc2/pymdown_extensions-11.0.tar.gz", hash = "sha256:8269cef0247f9e2d0a62fcea10860aba05c1cbab5470fd4b63230b96434dc589", size = 857049, upload-time = "2026-06-23T02:27:45.146Z" } sdist = { url = "https://files.pythonhosted.org/packages/9e/26/d1015444da4d952a1ca487a236b522eb979766f0295a0bd0c5fc089989a9/pymdown_extensions-10.21.3.tar.gz", hash = "sha256:72cfcf55f07aea0d4af2c4f11dd4e52466ddfb1bb819673146398e0bd3a77354", size = 854140, upload-time = "2026-05-13T12:57:32.267Z" }
wheels = [ wheels = [
{ url = "https://files.pythonhosted.org/packages/af/b6/1ae53367e28b9cffa3be7574e13fbe4589694272fd47710fbdbafd3d63c6/pymdown_extensions-11.0-py3-none-any.whl", hash = "sha256:fbc4acb641814fa9d17521bbd21a5240ef739a662f11c06330c4b78c93e954d6", size = 269415, upload-time = "2026-06-23T02:27:43.826Z" }, { url = "https://files.pythonhosted.org/packages/7e/85/545a951eecc270fcd688288c600017e2050a1aacb56c711d208586d3e470/pymdown_extensions-10.21.3-py3-none-any.whl", hash = "sha256:d7a5d08014fc571e80ca21dd6f854e31f94c489800350564d55d15b3c41e76b6", size = 269002, upload-time = "2026-05-13T12:57:30.296Z" },
] ]
[[package]] [[package]]