The trash list needs to show a version to whoever owns its root document, but
it deliberately shows only owned and unowned documents and ignores explicit
shares, so it cannot use permitted_object_ids. #14384 gave it its own
filter_queryset with a hand-built condition through the root, which describes
the same rule as the annotations in permitted_object_ids a second time.
Move those annotations into annotate_authorizing_fields. PermittedObjectsFilter
gains a parent_field attribute that its owner-only and granted paths both use,
and the trash filter shrinks to setting include_granted and parent_field.
The search index and the document views use the newest version's content as a
root document's content, but everything the LLM side reads used the root's own
content: the text embedded in the LLM index, the content in the classification
prompt, the context blocks from similar documents, and the text that is
searched for similar documents. After a document was replaced by a new
version the LLM kept answering from the old text.
Read get_effective_content() in those four places. The index update annotates
the content in the query like the search index does, and the similar-document
lookup annotates it for the documents it fetches, so no query is made per
document. Entries already in the LLM index keep the old text until the
document next changes or the index is rebuilt.
The search index and the LLM index hold root documents only, a root being
indexed with its newest version's content. Every write path therefore had to
remember to hand them the root. Several did not: adding or deleting a note on
a version, restoring a trashed document together with its versions, and
reprocessing a version all wrote the version into the index under its own id,
where it could be returned as a separate search result. The paths that walk the
whole library did the same: the document_index reindex command put every
version in the search index, a full LLM index rebuild embedded every version,
and an incremental LLM update scoped to a version id indexed it under its own
id. Asking for documents like a version looked the version's own id up in the
index and silently found nothing.
WriteBatch.add_or_update, WriteBatch.add_or_update_ids and
llm_index_add_or_update_document now resolve a version to its root themselves.
The reindex command and update_llm_index only walk root documents, and an
incremental LLM update is scoped to the roots of the given ids through
versioning.root_document_ids, which add_or_update_ids shares. more_like_id
returns the root's id for a version. The reprocess task no longer picks the root
for the indexes and only still clears the caches of both documents.
has_perms_owner_aware judged a document by its own owner and grants, so each
endpoint that fetches a document itself had to remember to map a version to
its root document first, and one that forgot, like the more-like-this search
filter, authorized by a stale version owner.
The check now maps a Document to its root before looking at the owner and the
guardian grants, matching what permitted_document_ids does for id sets. The
eight call sites that mapped the document themselves pass it straight through.
The DRF object permission class needs no change because the document viewset
only ever serves root documents.
permitted_document_ids judged a version by its own owner and grants, so a
version whose owner had drifted from its root's was visible to the wrong
people and hidden from the right ones. Callers patched this individually by
mapping each document to its root first. The query itself was also slow on
MariaDB: the guardian grants were a UNION cast to integers and tested with
IN inside an OR with the owner checks, which MariaDB cannot materialize, so it
re-scans the user's grants for every document. At 20k documents that took
seconds for a user with a couple of hundred grants.
permitted_object_ids now looks grants up as an EXISTS keyed on the row id cast
to a string, which uses guardian's unique index, and matches the user's groups
with an IN subquery. It takes an optional parent_field naming a self-referencing
foreign key whose target authorizes the row, and permitted_document_ids passes
root_document, so a version is visible exactly when its root is. The helper
that mapped documents to their roots at the call sites is no longer needed, so
the email, selection data, share link bundle, trash, bulk download and bulk
edit checks use the id set directly.
An interrupted rebuild left an empty index stamped as current, so the next
start reported it as up to date. Mark the rebuild as in progress and only
clear the marker once it completes.
* Moves runners to 26.04 and a few jobs to -slim variant
* Probably fixing the imagemagik 7 problems and maybe the frontend playwright thing?
* Compare thumbnails, but allow a little difference in the perceptual hash
* Feature: store barcode contents, list and search them
New setting PAPERLESS_CONSUMER_STORE_BARCODE_VALUES (off by default)
stores all barcodes found during consumption with the document: page,
type and content. They are listed on the metadata tab with a copy
button, returned by the documents API and searchable with barcodes:
in the advanced search. Versions keep their own barcodes, reprocessing
reads them again. Refs #9898
* Tests: cover the remaining barcode branches
Covers unchanged and failing barcode reads on reprocessing, unsupported files and DocumentBarcode.__str__, and uses toHaveLength in the barcode list spec as suggested by SonarCloud.
* Address review: keep barcodes when they can't be read, format choices
- Reprocessing and new versions keep the stored barcodes when the file
can't be scanned or the scan fails, and replace them atomically.
- Format is a TextChoices of the zxing-cpp formats, with a test.
- Shared scan code, latest_version helper, TypedDict, serializer reuse.
- Barcodes in the split manifest, export/import tests, pytest-style tests.
* Barcode tests: TIFF reprocess, format check both ways, fixtures
- Reprocessing a TIFF with TIFF support off keeps the stored barcodes.
- The format test also fails when zxing-cpp drops a format.
- Fixtures in place of the sample dir mixin, plugin-level disabled test.
- Format labels aren't translated, OpenAPI enum named BarcodeFormatEnum.
* Review: module-level zxing reader, shorter barcode docs
- read_barcodes_zxing is a module-level function used by scan_pdf.
- Drop the trivial __str__ test, mark it no cover.
- Shorten the barcode docs and remove the duplicate in configuration.md.
* Fix header
* Use utility class
* Return barcodes from the metadata endpoint only
Drop the barcodes field and its prefetches from the document serializer,
as agreed in the review. Also remove the now empty component stylesheet.
---------
Co-authored-by: shamoon <4887959+shamoon@users.noreply.github.com>
* Chore: Refresh the mypy and pyrefly baselines
* Chore: Type factory calls as the model they build for mypy
factory-boy leaves Factory() unannotated, so mypy took UserFactory() to
be a factory instance rather than a User. A shared TypedModelFactory
base declares the call as returning the model for type checking only.
* Refresh again after rebase
bleach's repository is archived, and it depends on html5lib, which has not had
a release since 2020. turbohtml provides a bleach-compatible clean and a native
linkify that runs in linear time, so the 2048-character cap on email
linkification from #13187 is no longer needed and long messages get their
email addresses linked again.
The `created` date fallback derived a naive datetime from a file's
mtime using the OS-local zone, then labeled it as the configured
TIME_ZONE without converting. When the OS-local zone and TIME_ZONE
disagree, or when the C library can't resolve zoneinfo at all (as in
some sandboxed environments, where it silently falls back to UTC),
the resulting date can land on the wrong calendar day.
Convert the timestamp directly into the target zone with `tz=` on
fromtimestamp() instead of a naive conversion plus make_aware().
* Fix: redirect SHARE_LINK_BUNDLE_DIR to the test temp layout instead of the real media directory
* Fix: include f_to in test_filters subTest labels so each of the 8 cases reports distinctly
* Fix: run the post_consume error-log assertion after the raising call and match the actual paperless_mail logger name
* Fix: rename the blank-password workflow test to match its behavior and add a real wrong-password-fails test
* Fix: use the created social account's actual pk and remove an accidental tuple wrapping the mock provider
* Fix: assert against the created documents' actual pks instead of hardcoded 1 and 2
* Fix: assert test_compression actually produces a valid LZMA-compressed zip
* Fix: clear os.environ when patching PAPERLESS_ADMIN_* vars so a host-set value can't leak into the no-user test
* Fix: restore MIDDLEWARE, AUTHENTICATION_BACKENDS and REST_FRAMEWORK auth classes after each remote-user settings test instead of leaking the mutation into later tests
* Fix: use a guaranteed-nonexistent temp path instead of hardcoded /tmp/foo/bar in test_export_target_not_exists
Snapshot only the edited field, before and after the operation, gathering tags
and custom field instances into sorted id lists per document (empty when there
are none).
The mail message and mailbox builders, the fake libmagic and the classifier preprocessor stub lived inside test_mail.py and test_classifier.py, so other test modules imported them by importing a test module.
They now live in helpers modules beside the tests that use them.
util_call_with_backoff made every call wait 20 seconds even when the first attempt succeeded
The helper now sleeps only after a failure that has another attempt left.
Test modules in paperless, paperless_mail and documents imported filesystem assertions, the migration test base, the retry helper and the streaming-response reader out of documents/tests/utils.py, which kept each app's tests coupled to another app's test package.
They now live in paperless_testing, and the progress manager fake is renamed FakeProgressManager and now subclasses the real ProgressManager, overriding only the transport, so the payload it records is built by the production code. The twenty places that patched documents.tasks.ProgressManager by hand now use a fake_progress_manager fixture.
Two fixtures created a temporary index directory and pointed INDEX_DIR at it, and paperless_dirs did the same, so a test that requested more than one got whichever assignment ran last. The search conftest no longer defines its own index_dir fixture, the tests that took it read paperless_dirs.index_dir instead, and _search_index is now a thin wrapper that requests paperless_dirs. The fixture that yields a Document is renamed from indexed_document to searchable_document so it no longer differs by one character from the index_document factory next to it.
TantivyBackend(path=None) built an in-memory index that is not used by production ever used. The backend now requires a path, the open, write-batch and rebuild branches are gone, and the shared backend fixture and the fulltext similar-documents fixture build a real index under the per-test directory layout.
* Chore: Speed up test setup by hashing passwords with MD5 and batching index writes
Don't use Django's default PBKDF2, about 600 ms per use and 110 uses across the suite. Switches to MD5 instead.
Also fixes a test that didn't batch update the search index
* Chore: Stop the invalid webhook params test from waiting on a Celery broker
test_workflow_webhook_action_url_invalid_params_headers left send_webhook.apply_async unpatched, so it tried to actually enqueu and waited for the timeout.
Nineteen single-file fixtures in the parsers conftest had no consumers anywhere in the test tree.
Two fixtures were both called samples_dir and resolved one directory apart They are now document_samples_dir and parser_samples_dir
Enables pytest-randomly, which has sat commented out in pyproject.toml
since the Pytest 9 upgrade. Tests now run in a different order every
session, so a test cannot quietly depend on another having run first.