Compare commits

..
Author SHA1 Message Date
stumpylog 642511c7b9 Fix: use a root document's effective content for the LLM index and prompts
The search index and the document views use the newest version's content as a
root document's content, but everything the LLM side reads used the root's own
content: the text embedded in the LLM index, the content in the classification
prompt, the context blocks from similar documents, and the text that is
searched for similar documents. After a document was replaced by a new
version the LLM kept answering from the old text.

Read get_effective_content() in those four places. The index update annotates
the content in the query like the search index does, and the similar-document
lookup annotates it for the documents it fetches, so no query is made per
document. Entries already in the LLM index keep the old text until the
document next changes or the index is rebuilt.
2026-10-08 15:22:03 -07:00
stumpylog 210d97a522 Fix: index document versions as their root document
The search index and the LLM index hold root documents only, a root being
indexed with its newest version's content. Every write path therefore had to
remember to hand them the root. Several did not: adding or deleting a note on
a version, restoring a trashed document together with its versions, and
reprocessing a version all wrote the version into the index under its own id,
where it could be returned as a separate search result. The paths that walk the
whole library did the same: the document_index reindex command put every
version in the search index, a full LLM index rebuild embedded every version,
and an incremental LLM update scoped to a version id indexed it under its own
id. Asking for documents like a version looked the version's own id up in the
index and silently found nothing.

WriteBatch.add_or_update, WriteBatch.add_or_update_ids and
llm_index_add_or_update_document now resolve a version to its root themselves.
The reindex command and update_llm_index only walk root documents, and an
incremental LLM update is scoped to the roots of the given ids through
versioning.root_document_ids, which add_or_update_ids shares. more_like_id
returns the root's id for a version. The reprocess task no longer picks the root
for the indexes and only still clears the caches of both documents.
2026-10-08 15:22:02 -07:00
stumpylog 8a96190359 Fix: authorize document versions by their root in the single-object permission check
has_perms_owner_aware judged a document by its own owner and grants, so each
endpoint that fetches a document itself had to remember to map a version to
its root document first, and one that forgot, like the more-like-this search
filter, authorized by a stale version owner.

The check now maps a Document to its root before looking at the owner and the
guardian grants, matching what permitted_document_ids does for id sets. The
eight call sites that mapped the document themselves pass it straight through.
The DRF object permission class needs no change because the document viewset
only ever serves root documents.
2026-10-08 15:21:45 -07:00
stumpylog ece8769f59 Fix: authorize document versions by their root and speed up permission id sets
permitted_document_ids judged a version by its own owner and grants, so a
version whose owner had drifted from its root's was visible to the wrong
people and hidden from the right ones. Callers patched this individually by
mapping each document to its root first. The query itself was also slow on
MariaDB: the guardian grants were a UNION cast to integers and tested with
IN inside an OR with the owner checks, which MariaDB cannot materialize, so it
re-scans the user's grants for every document. At 20k documents that took
seconds for a user with a couple of hundred grants.

permitted_object_ids now looks grants up as an EXISTS keyed on the row id cast
to a string, which uses guardian's unique index, and matches the user's groups
with an IN subquery. It takes an optional parent_field naming a self-referencing
foreign key whose target authorizes the row, and permitted_document_ids passes
root_document, so a version is visible exactly when its root is. The helper
that mapped documents to their roots at the call sites is no longer needed, so
the email, selection data, share link bundle, trash, bulk download and bulk
edit checks use the id set directly.
2026-10-08 15:21:28 -07:00
shamoon 138160cbda Fix: bulk reprocess latest version for root documents (#14385) 2026-10-08 18:50:44 +00:00
GitHub Actions 6d960c11d5 Auto translate strings 2026-10-08 18:40:22 +00:00
shamoon a35bd6e238 Fix: consistently use root doc for version action permissions (#14384) 2026-10-08 18:39:05 +00:00
Trenton H 474c630aa4 Fix: Don't attempt to index created dates which are not representable inside the FTS index (#14390) 2026-10-08 08:36:35 -07:00
Trenton H d7a9894400 Fix: retry a search index rebuild that was interrupted (#14379)
An interrupted rebuild left an empty index stamped as current, so the next
start reported it as up to date. Mark the rebuild as in progress and only
clear the marker once it completes.
2026-10-07 08:54:36 -07:00
GitHub Actions ee34a6598e Auto translate strings 2026-10-06 15:13:08 +00:00
Trenton H 3a3b3ef66a Chore: Upgrade runners to Ubuntu 26.04 (#14347)
* Moves runners to 26.04 and a few jobs to -slim variant

* Probably fixing the imagemagik 7 problems and maybe the frontend playwright thing?

* Compare thumbnails, but allow a little difference in the perceptual hash
2026-10-06 08:11:44 -07:00
46 changed files with 1662 additions and 271 deletions

No files matched your search

+3 -3
View File
@@ -80,7 +80,7 @@ jobs:
needs: changes
if: needs.changes.outputs.backend_changed == 'true'
name: "Python ${{ matrix.python-version }}"
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
strategy:
@@ -114,7 +114,7 @@ jobs:
packages: unpaper tesseract-ocr imagemagick ghostscript poppler-utils
- name: Configure ImageMagick
run: |
sudo cp docker/rootfs/etc/ImageMagick-6/paperless-policy.xml /etc/ImageMagick-6/policy.xml
sudo cp docker/rootfs/etc/ImageMagick-6/paperless-policy.xml /etc/ImageMagick-7/policy.xml
- name: Install Python dependencies
env:
PYTHON_VERSION: ${{ steps.setup-python.outputs.python-version }}
@@ -158,7 +158,7 @@ jobs:
needs: changes
if: needs.changes.outputs.backend_changed == 'true'
name: Check project typing
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
env:
+3 -3
View File
@@ -24,10 +24,10 @@ jobs:
fail-fast: false
matrix:
include:
- runner: ubuntu-24.04
- runner: ubuntu-26.04
arch: amd64
platform: linux/amd64
- runner: ubuntu-24.04-arm
- runner: ubuntu-26.04-arm
arch: arm64
platform: linux/arm64
runs-on: ${{ matrix.runner }}
@@ -163,7 +163,7 @@ jobs:
archive: false
merge-and-push:
name: Merge and Push Manifest
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
needs: build-arch
if: needs.build-arch.outputs.should-push == 'true'
environment: image-publishing
+2 -2
View File
@@ -65,7 +65,7 @@ jobs:
needs: changes
if: needs.changes.outputs.docs_changed == 'true'
name: Build Documentation
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
steps:
- uses: actions/configure-pages@45bfe0192ca1faeb007ade9deae92b16b8254a0d # v6.0.0
- name: Checkout
@@ -102,7 +102,7 @@ jobs:
name: Deploy Documentation
needs: [changes, build]
if: github.event_name == 'push' && github.ref == 'refs/heads/main' && needs.changes.outputs.docs_changed == 'true'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
pages: write
id-token: write
+6 -6
View File
@@ -72,7 +72,7 @@ jobs:
needs: changes
if: needs.changes.outputs.frontend_changed == 'true'
name: Install Dependencies
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
steps:
@@ -104,7 +104,7 @@ jobs:
name: Lint
needs: [changes, install-dependencies]
if: needs.changes.outputs.frontend_changed == 'true'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
steps:
@@ -137,7 +137,7 @@ jobs:
name: "Unit Tests (${{ matrix.shard-index }}/${{ matrix.shard-count }})"
needs: [changes, install-dependencies]
if: needs.changes.outputs.frontend_changed == 'true'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
strategy:
@@ -188,10 +188,10 @@ jobs:
name: E2E Tests
needs: [changes, install-dependencies]
if: needs.changes.outputs.frontend_changed == 'true'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
container: mcr.microsoft.com/playwright:v1.62.1-noble
container: mcr.microsoft.com/playwright:v1.62.1-resolute
env:
PLAYWRIGHT_BROWSERS_PATH: /ms-playwright
PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD: 1
@@ -246,7 +246,7 @@ jobs:
name: Frontend Build
needs: [changes, unit-tests, e2e-tests]
if: needs.changes.outputs.frontend_changed == 'true'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
steps:
+5 -5
View File
@@ -14,7 +14,7 @@ permissions: {}
jobs:
wait-for-docker:
name: Wait for Docker Build
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
checks: read
statuses: read
@@ -30,7 +30,7 @@ jobs:
build-release:
name: Build Release
needs: wait-for-docker
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
steps:
@@ -73,7 +73,7 @@ jobs:
timeout-minutes: 12
uses: $/.github/actions/apt-install
with:
packages: gettext liblept5
packages: gettext libleptonica6
# ---- Build Documentation ----
- name: Build documentation
env:
@@ -145,7 +145,7 @@ jobs:
publish-release:
name: Publish Release
needs: build-release
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: write
pull-requests: write
@@ -197,7 +197,7 @@ jobs:
name: Append Changelog
needs: publish-release
if: needs.publish-release.outputs.prerelease == 'false'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: write
pull-requests: write
+2 -2
View File
@@ -15,7 +15,7 @@ permissions:
jobs:
zizmor:
name: Run zizmor
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
contents: read
actions: read
@@ -29,7 +29,7 @@ jobs:
uses: zizmorcore/zizmor-action@cc914d7f3750a2d13d75c7f184a1060aa0e9d482 # v0.6.4
semgrep:
name: Semgrep CE
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
container:
image: semgrep/semgrep:1.155.0@sha256:cc869c685dcc0fe497c86258da9f205397d8108e56d21a86082ea4886e52784d
if: github.actor != 'dependabot[bot]'
+2 -2
View File
@@ -17,7 +17,7 @@ jobs:
cleanup-images:
name: Cleanup Image Tags for ${{ matrix.primary-name }}
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
environment: registry-maintenance
strategy:
fail-fast: false
@@ -42,7 +42,7 @@ jobs:
cleanup-untagged-images:
name: Cleanup Untagged Images Tags for ${{ matrix.primary-name }}
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
needs:
- cleanup-images
environment: registry-maintenance
+1 -1
View File
@@ -21,7 +21,7 @@ on:
jobs:
analyze:
name: Analyze
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
actions: read
contents: read
+1 -1
View File
@@ -13,7 +13,7 @@ jobs:
synchronize-with-crowdin:
name: Crowdin Sync
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
environment: translation-sync
steps:
- name: Checkout
+1 -1
View File
@@ -7,7 +7,7 @@ jobs:
# Note: peakoss/anti-slop does not support the `issues` event yet (all of its
# issue inputs are still commented out upstream), so the checks that the PR Bot
# workflow gets from the action are implemented manually here.
runs-on: ubuntu-latest
runs-on: ubuntu-slim
permissions:
issues: write
steps:
+2 -2
View File
@@ -4,7 +4,7 @@ on:
types: [opened]
jobs:
Anti-slop:
runs-on: ubuntu-latest
runs-on: ubuntu-slim
permissions:
contents: read
issues: read
@@ -24,7 +24,7 @@ jobs:
ASLOP-PR-VERIFY
pr-bot:
name: Automated PR Bot
runs-on: ubuntu-latest
runs-on: ubuntu-slim
# Runs after Anti-slop so the welcome comment can see whether the PR was closed
# instead of racing it. Still runs if that job fails, so labeling is not lost.
needs: Anti-slop
+1 -1
View File
@@ -12,7 +12,7 @@ permissions:
jobs:
pr_opened_or_reopened:
name: pr_opened_or_reopened
runs-on: ubuntu-24.04
runs-on: ubuntu-slim
permissions:
# write permission is required for autolabeler
pull-requests: write
+5 -5
View File
@@ -9,7 +9,7 @@ jobs:
stale:
name: 'Stale'
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
issues: write
pull-requests: write
@@ -34,7 +34,7 @@ jobs:
lock-threads:
name: 'Lock Old Threads'
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-26.04
permissions:
issues: write
pull-requests: write
@@ -58,7 +58,7 @@ jobs:
close-answered-discussions:
name: 'Close Answered Discussions'
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-slim
permissions:
discussions: write
steps:
@@ -117,7 +117,7 @@ jobs:
close-outdated-discussions:
name: 'Close Outdated Discussions'
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-slim
permissions:
discussions: write
steps:
@@ -211,7 +211,7 @@ jobs:
close-unsupported-feature-requests:
name: 'Close Unsupported Feature Requests'
if: github.repository_owner == 'paperless-ngx'
runs-on: ubuntu-24.04
runs-on: ubuntu-slim
permissions:
discussions: write
steps:
+1 -1
View File
@@ -8,7 +8,7 @@ env:
jobs:
generate-translate-strings:
name: Generate Translation Strings
runs-on: ubuntu-latest
runs-on: ubuntu-26.04
environment: translation-sync
permissions:
contents: write
+1 -3
View File
@@ -76,9 +76,7 @@ is not supported by any of the available parsers.
**A:** Not by default. As of v3, a file whose contents match an existing document is still
consumed, and the duplicate is flagged in the UI — open the document and check the
**Duplicates** tab to review documents that share the same content, or filter the document
list by **Duplicates** to find all of them (see
[Duplicate documents](usage.md#duplicate-documents)). If you prefer the old
**Duplicates** tab to review documents that share the same content. If you prefer the old
behavior of rejecting duplicates during consumption, set
[`PAPERLESS_CONSUMER_DELETE_DUPLICATES`](configuration.md#PAPERLESS_CONSUMER_DELETE_DUPLICATES)
to `true`.
+8 -7
View File
@@ -299,18 +299,19 @@ for details.
### Duplicate documents
By default, Paperless-ngx **does not reject duplicates**. If you consume a file whose
contents match an existing document (same original or archive checksum), the new copy is
still consumed and a warning is logged.
contents exactly match an existing document (same checksum), the new copy is still
consumed and a warning is logged. The task entry for the upload also flags that a
duplicate was detected and links to the existing document(s).
When a document has duplicates, a **Duplicates** tab appears on its detail page, listing
the other documents you can view that share the same content (including any in the trash).
To find all documents with duplicates, choose **Duplicates** in the document list's text
filter dropdown, or use `has_duplicates=true` in the REST API.
To review duplicates, open a document and switch to the **Duplicates** tab on the
document detail page. It lists other documents that share the same content, including any
that are in the trash (shown with a badge), and links to each so you can decide which to
keep.
If you would rather reject duplicates at consumption time (the pre-v3 behavior), set
[`PAPERLESS_CONSUMER_DELETE_DUPLICATES`](configuration.md#PAPERLESS_CONSUMER_DELETE_DUPLICATES)
to `true`. The duplicate file is then deleted instead of consumed, and the task fails with
a "Document already exists" message linking to the existing document.
a "document already exists" message.
## Document Suggestions
+1 -1
View File
@@ -1,6 +1,6 @@
[project]
name = "paperless-ngx"
version = "3.3.0"
version = "3.2.1"
description = """\
A community-supported supercharged document management system: scan, index and archive all your physical documents\
"""
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "paperless-ngx-ui",
"version": "3.3.0",
"version": "3.2.1",
"scripts": {
"preinstall": "npx only-allow pnpm",
"ng": "ng",
+1 -1
View File
@@ -8,7 +8,7 @@ export const environment = {
apiVersion: '10', // match src/paperless/settings.py
appTitle: DEFAULT_APP_TITLE,
tag: 'prod',
version: '3.3.0',
version: '3.2.1',
webSocketHost: window.location.host,
webSocketProtocol: window.location.protocol == 'https:' ? 'wss:' : 'ws:',
webSocketBaseUrl: base_url.pathname + 'ws/',
+27 -2
View File
@@ -13,8 +13,14 @@ from celery import group
from celery import shared_task
from django.conf import settings
from django.db import transaction
from django.db.models import Case
from django.db.models import F
from django.db.models import Max
from django.db.models import OuterRef
from django.db.models import Q
from django.db.models import Subquery
from django.db.models import When
from django.db.models.functions import Coalesce
from django.utils import timezone
from documents.data_models import ConsumableDocument
@@ -36,6 +42,7 @@ from documents.tasks import remove_document_from_index
from documents.tasks import update_document_content_maybe_archive_file
from documents.versioning import get_latest_version_for_root
from documents.versioning import get_root_document
from documents.versioning import versions_newest_first
if TYPE_CHECKING:
from collections.abc import Mapping
@@ -408,10 +415,28 @@ def reprocess(doc_ids: list[int], *, remote_ocr: bool = False) -> Literal["OK"]:
Consumption workflows do not run here, so ``remote_ocr`` is how the user
asks for the remote engine when it is not configured to handle everything.
A root document with versions reprocesses its latest version, which is the
file whose content, archive and thumbnail are shown for it.
"""
for document_id in doc_ids:
latest_version = versions_newest_first(
Document.objects.filter(root_document=OuterRef("pk")),
).values("id")[:1]
source_ids = (
Document.objects.filter(id__in=doc_ids)
.annotate(
source_id=Case(
When(root_document__isnull=False, then=F("id")),
default=Coalesce(Subquery(latest_version), F("id")),
),
)
.order_by()
.values_list("source_id", flat=True)
.distinct()
)
for source_id in source_ids:
update_document_content_maybe_archive_file.apply_async(
kwargs={"document_id": document_id, "remote_ocr": remote_ocr},
kwargs={"document_id": source_id, "remote_ocr": remote_ocr},
headers={"trigger_source": PaperlessTask.TriggerSource.MANUAL},
)
@@ -67,18 +67,22 @@ class Command(PaperlessCommand):
if options.get("recreate"):
wipe_index(settings.INDEX_DIR)
documents = Document.objects.select_related(
"correspondent",
"document_type",
"storage_path",
"owner",
).prefetch_related(
"tags",
"notes__user",
"custom_fields__field",
"versions",
"barcodes",
"versions__barcodes",
documents = (
Document.objects.filter(root_document__isnull=True)
.select_related(
"correspondent",
"document_type",
"storage_path",
"owner",
)
.prefetch_related(
"tags",
"notes__user",
"custom_fields__field",
"versions",
"barcodes",
"versions__barcodes",
)
)
total = documents.count()
rebuild_kwargs = {}
+66 -12
View File
@@ -6,14 +6,19 @@ from django.contrib.auth.models import Permission
from django.contrib.auth.models import User
from django.contrib.contenttypes.models import ContentType
from django.db.models import Case
from django.db.models import CharField
from django.db.models import Count
from django.db.models import Exists
from django.db.models import F
from django.db.models import IntegerField
from django.db.models import Model
from django.db.models import OuterRef
from django.db.models import Q
from django.db.models import QuerySet
from django.db.models import Value
from django.db.models import When
from django.db.models.functions import Cast
from django.db.models.functions import Coalesce
from guardian.core import ObjectPermissionChecker
from guardian.models import GroupObjectPermission
from guardian.models import UserObjectPermission
@@ -25,6 +30,7 @@ from rest_framework.permissions import BasePermission
from rest_framework.permissions import DjangoObjectPermissions
from documents.models import Document
from documents.versioning import get_root_document
class PaperlessObjectPermissions(DjangoObjectPermissions):
@@ -348,6 +354,7 @@ def permitted_object_ids(
perm: str,
*,
include_deleted: bool = False,
parent_field: str | None = None,
) -> QuerySet[int]:
"""
Generic version of ``permitted_document_ids`` for any model with an
@@ -356,6 +363,20 @@ def permitted_object_ids(
soft-delete pattern (currently only ``Document``); for every other model
it is accepted but has no effect, since those models have no soft-delete
concept.
``parent_field`` names a self-referencing foreign key whose target
authorizes the row (``Document.root_document``). A row with a parent is
visible exactly when its parent is, judged by the parent's owner and
grants, so the row's own owner and grants are ignored.
Guardian stores ``object_pk`` as a string, so each grant is an ``EXISTS``
keyed on the row's id cast to a string, which can use guardian's
(user, permission, object_pk) unique index. Casting every ``object_pk`` to
an integer for an ``id IN (...)`` is not indexable, and MariaDB cannot
materialize it inside the owner ``OR``, so it re-scans the user's grants
for every row. The user's groups are matched with an ``IN`` subquery
rather than a join through the membership table, which SQLite plans badly
once the grant tables grow.
"""
has_soft_delete = hasattr(model, "global_objects")
manager = (
@@ -363,8 +384,21 @@ def permitted_object_ids(
)
base_qs = manager.all().only("id", "owner")
owner_field, key_field = "owner", "pk"
if parent_field is not None:
owner_field, key_field = "authorizing_owner", "authorizing_id"
base_qs = base_qs.annotate(
authorizing_id=Coalesce(f"{parent_field}_id", "id"),
authorizing_owner=Case(
When(**{f"{parent_field}_id__isnull": True}, then=F("owner_id")),
default=F(f"{parent_field}__owner_id"),
output_field=IntegerField(),
),
)
unowned = Q(**{f"{owner_field}__isnull": True})
if user is None or not getattr(user, "is_authenticated", False):
return base_qs.filter(owner__isnull=True).values_list("id", flat=True)
return base_qs.filter(unowned).values_list("id", flat=True)
# Deactivated users get nothing, deactivated superusers included, so this
# has to come before the superuser shortcut. guardian's
@@ -388,20 +422,24 @@ def permitted_object_ids(
"permission__content_type": content_type,
}
user_perm_ids = (
UserObjectPermission.objects.filter(user=user, **perm_filter)
.annotate(object_pk_int=Cast("object_pk", IntegerField()))
.values_list("object_pk_int", flat=True)
key_as_text = Cast(OuterRef(key_field), CharField(max_length=64))
granted_to_user = Exists(
UserObjectPermission.objects.filter(
user=user,
object_pk=key_as_text,
**perm_filter,
),
)
group_perm_ids = (
GroupObjectPermission.objects.filter(group__user=user, **perm_filter)
.annotate(object_pk_int=Cast("object_pk", IntegerField()))
.values_list("object_pk_int", flat=True)
granted_to_group = Exists(
GroupObjectPermission.objects.filter(
group_id__in=user.groups.values("id"),
object_pk=key_as_text,
**perm_filter,
),
)
permitted_ids = user_perm_ids.union(group_perm_ids)
return base_qs.filter(
Q(owner=user) | Q(owner__isnull=True) | Q(id__in=permitted_ids),
Q(**{owner_field: user.pk}) | unowned | granted_to_user | granted_to_group,
).values_list("id", flat=True)
@@ -470,8 +508,17 @@ def permitted_document_ids(
``include_deleted=True`` for callers that need to check permission on
soft-deleted documents (e.g. trash restore). This intentionally avoids
``get_objects_for_user`` to keep the subquery small and index-friendly.
A version is authorized by its root document, so a version's own owner and
grants never matter.
"""
return permitted_object_ids(user, Document, perm, include_deleted=include_deleted)
return permitted_object_ids(
user,
Document,
perm,
include_deleted=include_deleted,
parent_field="root_document",
)
def get_document_count_filter_for_user(user, related_name: str = "documents"):
@@ -630,7 +677,14 @@ def has_perms_owner_aware(user, perms, obj):
single-object check still has many production callers. Several callers
remain across ``documents/``, ``paperless_mail/``, and ``paperless_ai/``
-- grep for this function name before removing it.
A document version is authorized by its root document, like in
``permitted_document_ids``, so a version's own owner and grants never
matter. Fetch the root with ``select_related("root_document__owner")`` to
avoid extra queries.
"""
if isinstance(obj, Document):
obj = get_root_document(obj)
checker = ObjectPermissionChecker(user)
return obj.owner is None or obj.owner == user or checker.has_perm(perms, obj)
+70 -35
View File
@@ -7,6 +7,7 @@ import threading
import time
from datetime import UTC
from datetime import datetime
from datetime import timedelta
from enum import StrEnum
from itertools import islice
from typing import TYPE_CHECKING
@@ -31,6 +32,7 @@ from documents.search._query import parse_user_query
from documents.search._schema import _write_sentinels
from documents.search._schema import build_schema
from documents.search._schema import open_or_rebuild_index
from documents.search._schema import rebuild_in_progress
from documents.search._schema import wipe_index
from documents.search._tokenizer import ascii_fold
from documents.search._tokenizer import autocomplete_tokens
@@ -54,6 +56,19 @@ if TYPE_CHECKING:
logger = logging.getLogger("paperless.search")
# tantivy stores dates as signed 64-bit nanoseconds since the Unix epoch, which
# covers 1677-09-21T00:12:43 to 2262-04-11T23:47:16 UTC
_INDEX_DATE_NANOS_MIN: Final[int] = -(2**63)
_INDEX_DATE_NANOS_MAX: Final[int] = 2**63 - 1
_UNIX_EPOCH: Final[datetime] = datetime(1970, 1, 1, tzinfo=UTC)
def _is_indexable_date(value: datetime) -> bool:
"""Whether value, at whole-second precision, fits tantivy's date range."""
nanos = ((value - _UNIX_EPOCH) // timedelta(seconds=1)) * 1_000_000_000
return _INDEX_DATE_NANOS_MIN <= nanos <= _INDEX_DATE_NANOS_MAX
_LOCK_TIMEOUT_SECONDS: Final[float] = 10.0 # per-attempt acquire timeout
_LOCK_RETRY_ATTEMPTS: Final[int] = 4 # total attempts (1 initial + 3 retries)
_LOCK_BACKOFF_BASE: Final[float] = 1.0 # seconds
@@ -269,9 +284,14 @@ class WriteBatch:
and adding the new version. This ensures stale document data (e.g., after
permission changes) doesn't persist in the index.
Only root documents are indexed, with their effective content, so a
version is indexed as its root document.
Args:
document: Django Document instance to index
"""
if document.root_document_id is not None:
document = document.root_document
self.remove(document.pk)
doc = self._backend._build_tantivy_doc(document)
self._writer.add_document(doc)
@@ -296,20 +316,22 @@ class WriteBatch:
An id with no matching document (e.g. deleted between the caller
collecting ids and the batch running) is silently skipped, matching
``add_or_update()``'s existing single-document deferred-task behavior
rather than erroring or leaving a stale index entry.
rather than erroring or leaving a stale index entry. The id of a
version stands for its root document.
Args:
ids: Primary keys of Document instances to index
"""
from documents.models import Document
from documents.versioning import annotate_effective_content
from documents.versioning import root_document_ids
ids = list(ids)
if not ids:
return
queryset = annotate_effective_content(
Document.objects.filter(pk__in=ids)
Document.objects.filter(pk__in=root_document_ids(ids))
.select_related("correspondent", "document_type", "storage_path", "owner")
.prefetch_related(
"tags",
@@ -627,7 +649,15 @@ class TantivyBackend:
document.created.day,
tzinfo=UTC,
)
doc.add_date("created", created_date)
if _is_indexable_date(created_date):
doc.add_date("created", created_date)
else:
logger.warning(
"Document %s has a created date (%s) outside the range the search "
"index can store; it will be indexed without a created date",
document.pk,
document.created,
)
doc.add_date("modified", document.modified)
doc.add_date("added", document.added)
@@ -1110,39 +1140,44 @@ class TantivyBackend:
flushing a segment, deferring merge work; they do not avoid it.
"""
wipe_index(self._path)
new_index = tantivy.Index(build_schema(), path=str(self._path))
_write_sentinels(self._path)
register_tokenizers(new_index, settings.SEARCH_LANGUAGE)
# The marker covers the window where the empty index is already stamped
# as current but not yet populated, so an interrupted rebuild is retried.
with rebuild_in_progress(self._path):
new_index = tantivy.Index(build_schema(), path=str(self._path))
_write_sentinels(self._path)
register_tokenizers(new_index, settings.SEARCH_LANGUAGE)
# Point instance at the new index so _build_tantivy_doc uses it
old_index, old_schema = self._raw_index, self._raw_schema
self._raw_index = new_index
self._raw_schema = new_index.schema
# Stream documents one-by-one (so the progress bar advances per
# document) while fetching viewer permissions one SQL query per chunk.
# The stream is Sized, so iter_wrapper can still discover the total.
documents_stream = _DocumentViewerStream(documents, chunk_size=1000)
try:
writer = new_index.writer(heap_size=writer_heap_bytes)
for document, (viewer_ids, viewer_group_ids) in iter_wrapper(
documents_stream,
):
doc = self._build_tantivy_doc(
document,
viewer_ids=viewer_ids,
viewer_group_ids=viewer_group_ids,
)
writer.add_document(doc)
writer.commit()
# Wait for background merge threads to finish so all segments are
# fully merged and persisted before the index is considered rebuilt.
writer.wait_merging_threads()
new_index.reload()
except BaseException: # pragma: no cover
# Restore old index on failure so the backend remains usable
self._raw_index = old_index
self._raw_schema = old_schema
raise
# Point instance at the new index so _build_tantivy_doc uses it
old_index, old_schema = self._raw_index, self._raw_schema
self._raw_index = new_index
self._raw_schema = new_index.schema
# Stream documents one-by-one (so the progress bar advances per
# document) while fetching viewer permissions one SQL query per
# chunk. The stream is Sized, so iter_wrapper can still discover
# the total.
documents_stream = _DocumentViewerStream(documents, chunk_size=1000)
try:
writer = new_index.writer(heap_size=writer_heap_bytes)
for document, (viewer_ids, viewer_group_ids) in iter_wrapper(
documents_stream,
):
doc = self._build_tantivy_doc(
document,
viewer_ids=viewer_ids,
viewer_group_ids=viewer_group_ids,
)
writer.add_document(doc)
writer.commit()
# Wait for background merge threads to finish so all segments
# are fully merged and persisted before the index is considered
# rebuilt.
writer.wait_merging_threads()
new_index.reload()
except BaseException: # pragma: no cover
# Restore old index on failure so the backend remains usable
self._raw_index = old_index
self._raw_schema = old_schema
raise
def chunked(iterable, size):
+45 -4
View File
@@ -4,6 +4,7 @@ import hashlib
import json
import logging
import shutil
from contextlib import contextmanager
from typing import TYPE_CHECKING
from typing import Final
from typing import NamedTuple
@@ -16,6 +17,7 @@ from whoosh_compat import FieldKind
from documents.search._fields import PUBLIC_FIELDS
if TYPE_CHECKING:
from collections.abc import Iterator
from pathlib import Path
logger = logging.getLogger("paperless.search")
@@ -28,6 +30,11 @@ logger = logging.getLogger("paperless.search")
# v3 - barcodes JSON field for stored barcode contents
SCHEMA_VERSION: Final[int] = 3
# Present in the index directory from the moment a full rebuild starts until it
# finishes. If a rebuild is interrupted it is left behind, so the half-built
# index is not mistaken for a complete one.
REBUILD_MARKER: Final[str] = ".rebuilding"
class FieldDescriptor(NamedTuple):
"""One tantivy field, in declaration order.
@@ -255,9 +262,9 @@ def needs_rebuild(index_dir: Path) -> bool:
"""
Check if the search index needs rebuilding.
Reads .index_settings.json to compare the stored schema version, search
language and schema fingerprint against the current configuration. Returns
True if the file is missing, unparsable, or any value mismatches.
True if a previous full rebuild never finished (the rebuild marker is still
present), or if the index's stamped settings no longer match the current
configuration. See _settings_mismatch().
Args:
index_dir: Path to the search index directory
@@ -265,6 +272,40 @@ def needs_rebuild(index_dir: Path) -> bool:
Returns:
True if the index needs rebuilding, False if it's up to date
"""
if (index_dir / REBUILD_MARKER).exists():
logger.warning("Previous search index rebuild did not finish - rebuilding.")
return True
return _settings_mismatch(index_dir)
@contextmanager
def rebuild_in_progress(index_dir: Path) -> Iterator[None]:
"""
Flag the index as incomplete for the duration of a full rebuild.
The marker is cleared only if the block exits cleanly. There is deliberately
no try/finally: an exception must leave the marker behind so the next
needs_rebuild() check retries the rebuild.
"""
marker = index_dir / REBUILD_MARKER
marker.touch()
yield
marker.unlink(missing_ok=True)
def _settings_mismatch(index_dir: Path) -> bool:
"""
Check the stamped settings against the current configuration.
Reads .index_settings.json to compare the stored schema version, search
language and schema fingerprint. Returns True if the file is missing,
unparsable, or any value mismatches.
This deliberately ignores the rebuild marker: open_or_rebuild_index() uses it
so that a process opening the index while another process is mid-rebuild
(or after one died) does not wipe the partial index out from under it.
Repopulating is the job of ``document_index reindex``.
"""
settings_file = index_dir / ".index_settings.json"
if not settings_file.exists():
return True
@@ -333,7 +374,7 @@ def open_or_rebuild_index(index_dir: Path | None = None) -> tantivy.Index:
index_dir = cast("Path", settings.INDEX_DIR)
if not index_dir.exists():
return tantivy.Index(build_schema())
if needs_rebuild(index_dir):
if _settings_mismatch(index_dir):
wipe_index(index_dir)
idx = tantivy.Index(build_schema(), path=str(index_dir))
_write_sentinels(index_dir)
+2
View File
@@ -502,6 +502,8 @@ def update_document_content_maybe_archive_file(
llm_index_add_or_update_document(document)
clear_document_caches(document.pk)
if document.root_document_id is not None:
clear_document_caches(document.root_document_id)
except Exception:
logger.exception(
+199
View File
@@ -1,4 +1,6 @@
import json
import logging
from datetime import date
from pathlib import Path
import pytest
@@ -16,7 +18,10 @@ from documents.search._backend import TantivyBackend
from documents.search._backend import WriteBatch
from documents.search._backend import get_backend
from documents.search._backend import reset_backend
from documents.search._schema import REBUILD_MARKER
from documents.search._schema import needs_rebuild
from documents.signals.handlers import add_to_index
from paperless_testing.dirs import PaperlessDirs
from paperless_testing.factories import CorrespondentFactory
from paperless_testing.factories import DocumentFactory
from paperless_testing.factories import DocumentTypeFactory
@@ -288,6 +293,80 @@ class TestAddOrUpdateIds:
assert backend.search_ids("updated", user=None) == [doc.pk]
class TestVersionsAreIndexedAsTheirRoot:
"""Only root documents are indexed, with their effective content, so
every write path that is handed a version indexes its root instead."""
@staticmethod
def _root_with_version() -> tuple[Document, Document]:
root = DocumentFactory(title="Statement", content="stale text")
version = DocumentFactory(
title="Statement",
content="latest text",
root_document=root,
version_index=1,
)
return root, version
def test_add_or_update_indexes_the_root_of_a_version(
self,
backend: TantivyBackend,
) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- The version is passed to add_or_update
THEN:
- The root is indexed with the version's text, and the version is not
"""
root, version = self._root_with_version()
backend.add_or_update(version)
assert backend.search_ids("latest", user=None) == [root.pk]
assert backend.search_ids("stale", user=None) == []
def test_add_or_update_ids_indexes_each_root_once(
self,
backend: TantivyBackend,
) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- Both ids are passed to add_or_update_ids
THEN:
- The root is indexed once and the version is not indexed
"""
root, version = self._root_with_version()
with backend.batch_update() as batch:
batch.add_or_update_ids([version.pk, root.pk])
assert backend.search_ids("Statement", user=None) == [root.pk]
assert backend.search_ids("latest", user=None) == [root.pk]
def test_add_or_update_ids_resolves_a_lone_version(
self,
backend: TantivyBackend,
) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- Only the version's id is passed to add_or_update_ids
THEN:
- The root is indexed
"""
root, version = self._root_with_version()
with backend.batch_update() as batch:
batch.add_or_update_ids([version.pk])
assert backend.search_ids("latest", user=None) == [root.pk]
class TestSearch:
"""Test search query parsing and matching via search_ids."""
@@ -823,6 +902,53 @@ class TestRebuild:
backend.rebuild(Document.objects.all(), iter_wrapper=wrapper)
assert 30 in seen
def test_successful_rebuild_leaves_index_up_to_date(
self,
backend: TantivyBackend,
paperless_dirs: PaperlessDirs,
) -> None:
"""
GIVEN:
- A backend and one document
WHEN:
- rebuild() completes
THEN:
- needs_rebuild() is False and no rebuild marker remains
"""
DocumentFactory.create()
backend.rebuild(Document.objects.all())
assert needs_rebuild(paperless_dirs.index_dir) is False
assert not (paperless_dirs.index_dir / REBUILD_MARKER).exists()
def test_interrupted_rebuild_is_retried(
self,
backend: TantivyBackend,
paperless_dirs: PaperlessDirs,
) -> None:
"""
GIVEN:
- A rebuild that dies while indexing documents (e.g. the database
connection is lost)
WHEN:
- needs_rebuild() is checked afterwards
THEN:
- It is True, even though the empty index was already stamped with
current settings, so the next start rebuilds instead of reporting
the index as up to date
"""
DocumentFactory.create()
def die(pairs):
raise RuntimeError("terminating connection due to administrator command")
yield # pragma: no cover
with pytest.raises(RuntimeError):
backend.rebuild(Document.objects.all(), iter_wrapper=die)
assert needs_rebuild(paperless_dirs.index_dir) is True
def test_includes_group_granted_viewers(self, backend: TantivyBackend) -> None:
"""Rebuild must index viewer ids for group-only grants, not just direct ones.
@@ -854,6 +980,79 @@ class TestRebuild:
assert ids == [doc.pk]
class TestCreatedDateOutOfRange:
"""The index stores dates as nanosecond i64 values (1677-09-22 to 2262-04-11).
A document whose created date falls outside that window must not abort
indexing: it is indexed without a created value and a warning names it.
"""
@pytest.mark.parametrize(
("created", "expected_warnings"),
[
pytest.param(date(1677, 9, 22), 0, id="first-representable-day"),
pytest.param(date(2262, 4, 11), 0, id="last-representable-day"),
pytest.param(date(1677, 9, 21), 1, id="day-before-first"),
pytest.param(date(2262, 4, 12), 1, id="day-after-last"),
pytest.param(date(16, 8, 30), 1, id="two-digit-year-read-as-year-16"),
pytest.param(date(9999, 12, 31), 1, id="max-python-date"),
],
)
def test_add_or_update_indexes_document_and_warns_when_out_of_range(
self,
backend: TantivyBackend,
caplog: pytest.LogCaptureFixture,
created: date,
expected_warnings: int,
) -> None:
"""
GIVEN:
- A document with a created date at or beyond the index date limits
WHEN:
- The document is added to the index
THEN:
- The document is indexed and searchable either way
- A warning naming the document is logged only for out-of-range dates
"""
doc = DocumentFactory(created=created, content="boundarycontent")
with caplog.at_level(logging.WARNING, logger="paperless.search"):
backend.add_or_update(doc)
assert backend.search_ids("boundarycontent", user=None) == [doc.pk]
warnings = [r for r in caplog.records if r.levelno == logging.WARNING]
assert len(warnings) == expected_warnings
if expected_warnings:
assert f"Document {doc.pk}" in warnings[0].getMessage()
def test_rebuild_continues_past_out_of_range_document(
self,
backend: TantivyBackend,
caplog: pytest.LogCaptureFixture,
) -> None:
"""
GIVEN:
- A document with an unrepresentable created date among valid ones
WHEN:
- The index is rebuilt
THEN:
- Rebuild completes and every document is searchable
- A warning names the offending document
"""
good = DocumentFactory(created=date(2016, 8, 30), content="rebuildcontent")
bad = DocumentFactory(created=date(16, 8, 30), content="rebuildcontent")
with caplog.at_level(logging.WARNING, logger="paperless.search"):
backend.rebuild(Document.objects.all())
assert sorted(backend.search_ids("rebuildcontent", user=None)) == sorted(
[good.pk, bad.pk],
)
warnings = [r for r in caplog.records if r.levelno == logging.WARNING]
assert len(warnings) == 1
assert f"Document {bad.pk}" in warnings[0].getMessage()
class TestAutocomplete:
"""Test autocomplete functionality."""
@@ -8,6 +8,7 @@ from auditlog.models import LogEntry # type: ignore[import-untyped]
from django.contrib.contenttypes.models import ContentType
from django.core.files.uploadedfile import SimpleUploadedFile
from django.test import TestCase as DjangoTestCase
from django.test import override_settings
from django.utils import timezone
from rest_framework import status
from rest_framework.test import APITestCase
@@ -16,13 +17,17 @@ from documents.data_models import DocumentSource
from documents.filters import EffectiveContentFilter
from documents.filters import TitleContentFilter
from documents.models import Document
from documents.models import Note
from documents.models import ShareLink
from documents.versioning import annotate_effective_content
from documents.views import DocumentSelectionMixin
from paperless_testing.dirs import DirectoriesMixin
from paperless_testing.factories import DocumentFactory
from paperless_testing.factories import UserFactory
from paperless_testing.http import read_streaming_response
from paperless_testing.permissions import grant_all_global
from paperless_testing.permissions import grant_global
from paperless_testing.permissions import grant_object
if TYPE_CHECKING:
from pathlib import Path
@@ -1043,3 +1048,152 @@ class TestBulkSelectionExcludesVersions(DjangoTestCase):
)
self.assertEqual(selected, [root.id])
class TestVersionActionPermissions(DirectoriesMixin, APITestCase):
def setUp(self):
super().setUp()
self.user = UserFactory()
grant_all_global(self.user)
self.client.force_authenticate(self.user)
self.root = DocumentFactory(owner=UserFactory())
self.version = DocumentFactory(root_document=self.root, owner=None)
@override_settings(AUDIT_LOG_ENABLED=True)
def test_actions_reject_stale_version_ownership(self):
note = Note.objects.create(document=self.version, note="Version note")
for owner in (None, self.user):
self.version.owner = owner
self.version.save(update_fields=["owner"])
for action in (
"notes",
"suggestions",
"ai_suggestions",
"history",
"share_links",
):
with self.subTest(owner=owner, action=action):
response = self.client.get(
f"/api/documents/{self.version.pk}/{action}/",
)
self.assertEqual(response.status_code, 403)
response = self.client.post(
f"/api/documents/{self.version.pk}/notes/",
{"note": "New note"},
)
self.assertEqual(response.status_code, 403)
response = self.client.delete(
f"/api/documents/{self.version.pk}/notes/?id={note.pk}",
)
self.assertEqual(response.status_code, 403)
response = self.client.post(
"/api/share_links/",
{"document": self.version.pk, "file_version": "original"},
)
self.assertEqual(response.status_code, 403)
response = self.client.post(
"/api/share_link_bundles/",
{"document_ids": [self.version.pk], "file_version": "original"},
format="json",
)
self.assertEqual(response.status_code, 400)
response = self.client.post(
"/api/documents/email/",
{
"documents": [self.version.pk],
"addresses": "recipient@example.com",
"subject": "Version",
"message": "Version",
},
format="json",
)
self.assertEqual(response.status_code, 403)
with (
mock.patch("documents.views.AIConfig") as ai_config,
mock.patch("documents.views.stream_chat_with_documents") as chat,
):
ai_config.return_value.ai_enabled = True
response = self.client.post(
"/api/documents/chat/",
{"q": "Version?", "document_id": self.version.pk},
format="json",
)
self.assertEqual(response.status_code, 403)
chat.assert_not_called()
self.assertTrue(Note.objects.filter(pk=note.pk).exists())
self.assertFalse(ShareLink.objects.exists())
@mock.patch("documents.views.build_share_link_bundle.apply_async")
def test_root_permissions_allow_sharing_a_private_version(self, build_mock):
self.version.owner = UserFactory()
self.version.save(update_fields=["owner"])
grant_object(self.user, self.root, "view_document", "change_document")
note = Note.objects.create(document=self.version, note="Version note")
response = self.client.get(f"/api/documents/{self.version.pk}/notes/")
self.assertEqual(response.status_code, 200)
self.assertEqual(response.data[0]["id"], note.pk)
response = self.client.post(
"/api/share_links/",
{"document": self.version.pk, "file_version": "original"},
)
self.assertEqual(response.status_code, 201)
self.assertEqual(ShareLink.objects.get().document_id, self.version.pk)
response = self.client.get(f"/api/documents/{self.version.pk}/share_links/")
self.assertEqual(response.status_code, 200)
self.assertEqual(len(response.data), 1)
response = self.client.post(
"/api/share_link_bundles/",
{"document_ids": [self.version.pk], "file_version": "original"},
format="json",
)
self.assertEqual(response.status_code, 201)
build_mock.assert_called_once()
def test_root_view_permission_does_not_allow_note_changes(self):
grant_object(self.user, self.root, "view_document")
note = Note.objects.create(document=self.version, note="Version note")
response = self.client.get(f"/api/documents/{self.version.pk}/notes/")
self.assertEqual(response.status_code, 200)
response = self.client.post(
f"/api/documents/{self.version.pk}/notes/",
{"note": "New note"},
)
self.assertEqual(response.status_code, 403)
response = self.client.delete(
f"/api/documents/{self.version.pk}/notes/?id={note.pk}",
)
self.assertEqual(response.status_code, 403)
self.assertTrue(Note.objects.filter(pk=note.pk).exists())
@override_settings(AUDIT_LOG_ENABLED=True)
def test_history_uses_root_ownership(self):
self.root.owner = self.user
self.root.save(update_fields=["owner"])
self.version.owner = UserFactory()
self.version.save(update_fields=["owner"])
response = self.client.get(f"/api/documents/{self.version.pk}/history/")
self.assertEqual(response.status_code, 200)
def test_selection_data_rejects_stale_version_ownership(self):
for owner in (None, self.user):
self.version.owner = owner
self.version.save(update_fields=["owner"])
with self.subTest(owner=owner):
response = self.client.post(
"/api/documents/selection_data/",
{"documents": [self.version.pk]},
format="json",
)
self.assertEqual(response.status_code, 403)
def test_selection_data_allows_private_version_of_permitted_root(self):
self.version.owner = UserFactory()
self.version.save(update_fields=["owner"])
grant_object(self.user, self.root, "view_document")
other = DocumentFactory(owner=self.user)
response = self.client.post(
"/api/documents/selection_data/",
{"documents": [self.version.pk, other.pk]},
format="json",
)
self.assertEqual(response.status_code, 200)
+53
View File
@@ -1196,6 +1196,59 @@ class TestDocumentSearchApi(DirectoriesMixin, APITestCase):
self.assertIn(d3.id, result_ids)
self.assertNotIn(d4.id, result_ids)
def test_search_more_like_version_uses_its_root(self) -> None:
"""
GIVEN:
- A document similar in content to a root document, and one that is not
- A version of the root document, which is never indexed
WHEN:
- API request for more like the version
THEN:
- The documents similar to the root are returned, not the version
"""
indexed = {}
for name, title, content, day in (
("root", "bank statement 1", "things i paid for in august", (2019, 3, 4)),
(
"similar",
"bank statement 3",
"things i paid for in september",
(2020, 7, 9),
),
(
"other",
"Quarterly Report",
"quarterly revenue profit margin",
(2021, 11, 30),
),
):
with time_machine.travel(
timezone.make_aware(datetime.datetime(*day)),
tick=False,
):
indexed[name] = DocumentFactory(
title=title,
content=content,
created=datetime.date(*day),
added=timezone.make_aware(datetime.datetime(*day)),
)
version = DocumentFactory(
root_document=indexed["root"],
version_index=1,
content="things i paid for in august",
)
backend = get_backend()
for document in indexed.values():
backend.add_or_update(document)
response = self.client.get(f"/api/documents/?more_like_id={version.id}")
self.assertEqual(response.status_code, status.HTTP_200_OK)
result_ids = [r["id"] for r in response.data["results"]]
self.assertIn(indexed["similar"].id, result_ids)
self.assertNotIn(indexed["other"].id, result_ids)
self.assertNotIn(version.id, result_ids)
def test_more_like_requires_id_of_existing_document(self) -> None:
"""
GIVEN:
+51
View File
@@ -6,6 +6,7 @@ from rest_framework.test import APITestCase
from documents.models import Document
from paperless_testing.dirs import DirectoriesMixin
from paperless_testing.factories import DocumentFactory
from paperless_testing.factories import UserFactory
from paperless_testing.permissions import grant_all_global
@@ -279,3 +280,53 @@ class TestTrashAPI(DirectoriesMixin, APITestCase):
Document.objects.filter(root_document=root).values_list("id", flat=True),
[version.pk for version in versions],
)
def test_api_trash_version_follows_root_owner(self) -> None:
"""
GIVEN:
- A deleted version of user2's document, owned by nobody
- A deleted version of the user's document, owned by user2
WHEN:
- The user lists the trash and tries to restore or empty the versions
THEN:
- Only the version of the user's own document is listed
- The other version can't be restored or emptied
- The version of the user's own document can be restored
"""
user2 = UserFactory(username="user2")
other_version = DocumentFactory(
root_document=DocumentFactory(owner=user2),
version_index=1,
)
other_version.delete()
own_version = DocumentFactory(
owner=user2,
root_document=DocumentFactory(owner=self.user),
version_index=1,
)
own_version.delete()
resp = self.client.get("/api/trash/")
self.assertEqual(resp.status_code, status.HTTP_200_OK)
self.assertEqual(
[doc["id"] for doc in resp.data["results"]],
[own_version.pk],
)
for action in ("restore", "empty"):
with self.subTest(action=action):
resp = self.client.post(
"/api/trash/",
{"action": action, "documents": [other_version.pk]},
)
self.assertEqual(resp.status_code, status.HTTP_403_FORBIDDEN)
self.assertTrue(
Document.deleted_objects.filter(pk=other_version.pk).exists(),
)
resp = self.client.post(
"/api/trash/",
{"action": "restore", "documents": [own_version.pk]},
)
self.assertEqual(resp.status_code, status.HTTP_200_OK)
self.assertTrue(Document.objects.filter(pk=own_version.pk).exists())
+91 -27
View File
@@ -4,6 +4,7 @@ from pathlib import Path
from unittest import mock
import pikepdf
import pytest
from django.contrib.auth.models import Group
from django.contrib.auth.models import Permission
from django.contrib.auth.models import User
@@ -12,6 +13,7 @@ from django.test import TestCase
from django.test.utils import CaptureQueriesContext
from guardian.shortcuts import get_groups_with_perms
from guardian.shortcuts import get_users_with_perms
from pytest_mock import MockerFixture
from documents import bulk_edit
from documents.models import Correspondent
@@ -23,6 +25,7 @@ from documents.models import StoragePath
from documents.models import Tag
from documents.permissions import set_permissions_for_objects
from paperless_testing.dirs import DirectoriesMixin
from paperless_testing.factories import DocumentFactory
from paperless_testing.permissions import grant_object
@@ -1970,18 +1973,22 @@ class TestPDFActions(DirectoriesMixin, TestCase):
self.assertIn("Error removing password from document", cm.output[0])
class TestBulkEditReprocess(DirectoriesMixin, TestCase):
def setUp(self) -> None:
super().setUp()
self.doc = Document.objects.create(
title="test",
checksum="A",
mime_type="application/pdf",
@pytest.mark.django_db
class TestBulkEditReprocess:
@pytest.fixture
def mock_task(self, mocker: MockerFixture) -> mock.MagicMock:
return mocker.patch(
"documents.bulk_edit.update_document_content_maybe_archive_file",
)
@mock.patch("documents.bulk_edit.update_document_content_maybe_archive_file")
def test_reprocess_defaults_to_local(self, mock_task: mock.Mock) -> None:
@staticmethod
def _queued_ids(mock_task: mock.MagicMock) -> list[int]:
return [
call.kwargs["kwargs"]["document_id"]
for call in mock_task.apply_async.call_args_list
]
def test_reprocess_defaults_to_local(self, mock_task: mock.MagicMock) -> None:
"""
GIVEN:
- A reprocess request that says nothing about remote OCR
@@ -1990,18 +1997,17 @@ class TestBulkEditReprocess(DirectoriesMixin, TestCase):
THEN:
- The task is queued without asking for the remote engine
"""
result = bulk_edit.reprocess([self.doc.id])
doc = DocumentFactory()
assert bulk_edit.reprocess([doc.id]) == "OK"
self.assertEqual(result, "OK")
mock_task.apply_async.assert_called_once()
_, kwargs = mock_task.apply_async.call_args
self.assertEqual(
kwargs["kwargs"],
{"document_id": self.doc.id, "remote_ocr": False},
)
assert mock_task.apply_async.call_args.kwargs["kwargs"] == {
"document_id": doc.id,
"remote_ocr": False,
}
@mock.patch("documents.bulk_edit.update_document_content_maybe_archive_file")
def test_reprocess_passes_remote_ocr(self, mock_task: mock.Mock) -> None:
def test_reprocess_passes_remote_ocr(self, mock_task: mock.MagicMock) -> None:
"""
GIVEN:
- A reprocess request that explicitly asks for remote OCR
@@ -2010,14 +2016,72 @@ class TestBulkEditReprocess(DirectoriesMixin, TestCase):
THEN:
- The request is forwarded to the task for every document
"""
other = Document.objects.create(
title="test2",
checksum="B",
mime_type="application/pdf",
docs = DocumentFactory.create_batch(2)
bulk_edit.reprocess([doc.id for doc in docs], remote_ocr=True)
assert mock_task.apply_async.call_count == 2
for call in mock_task.apply_async.call_args_list:
assert call.kwargs["kwargs"]["remote_ocr"]
def test_reprocess_root_uses_latest_version(
self,
mock_task: mock.MagicMock,
) -> None:
"""
GIVEN:
- A root document with two versions
WHEN:
- reprocess is called with the root document
THEN:
- The latest version is reprocessed, not the root's original file
"""
root = DocumentFactory()
DocumentFactory(root_document=root, version_index=1)
latest = DocumentFactory(root_document=root, version_index=2)
bulk_edit.reprocess([root.id])
assert self._queued_ids(mock_task) == [latest.id]
def test_reprocess_explicit_version(self, mock_task: mock.MagicMock) -> None:
"""
GIVEN:
- A root document with two versions
WHEN:
- reprocess is called with the older version
THEN:
- That version is reprocessed
"""
root = DocumentFactory()
older = DocumentFactory(root_document=root, version_index=1)
DocumentFactory(root_document=root, version_index=2)
bulk_edit.reprocess([older.id])
assert self._queued_ids(mock_task) == [older.id]
def test_reprocess_root_and_latest_version_dispatches_once(
self,
mock_task: mock.MagicMock,
) -> None:
"""
GIVEN:
- A root document with two versions, the latest created on a
different date than the root
WHEN:
- reprocess is called with both the root and its latest version
THEN:
- The latest version is reprocessed only once
"""
root = DocumentFactory(created=date(2024, 1, 1))
DocumentFactory(root_document=root, version_index=1)
latest = DocumentFactory(
root_document=root,
version_index=2,
created=date(2025, 1, 1),
)
bulk_edit.reprocess([self.doc.id, other.id], remote_ocr=True)
bulk_edit.reprocess([root.id, latest.id])
self.assertEqual(mock_task.apply_async.call_count, 2)
for call in mock_task.apply_async.call_args_list:
self.assertTrue(call.kwargs["kwargs"]["remote_ocr"])
assert self._queued_ids(mock_task) == [latest.id]
+22
View File
@@ -22,6 +22,7 @@ from documents.models import Document
from documents.tasks import update_document_content_maybe_archive_file
from paperless_testing.assertions import FileSystemAssertsMixin
from paperless_testing.dirs import DirectoriesMixin
from paperless_testing.factories import DocumentFactory
sample_file: Path = Path(__file__).parent / "samples" / "simple.pdf"
@@ -116,6 +117,27 @@ class TestMakeIndex:
call_command("document_index", "reindex", skip_checks=True)
mock_get_backend.return_value.rebuild.assert_called_once()
def test_reindex_skips_versions(self, mocker: MockerFixture) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- The reindex command runs
THEN:
- Only the root document is handed to the rebuild, since a version
is indexed as its root
"""
root = DocumentFactory()
DocumentFactory(root_document=root, version_index=1)
mock_get_backend = mocker.patch(
"documents.management.commands.document_index.get_backend",
)
call_command("document_index", "reindex", skip_checks=True)
documents = mock_get_backend.return_value.rebuild.call_args.args[0]
assert list(documents.values_list("pk", flat=True)) == [root.pk]
def test_optimize(self) -> None:
"""Optimize command must execute without error (Tantivy handles optimization automatically)."""
call_command("document_index", "optimize", skip_checks=True)
@@ -18,6 +18,7 @@ from documents.models import Correspondent
from documents.models import DocumentType
from documents.models import StoragePath
from documents.models import Tag
from documents.permissions import has_perms_owner_aware
from documents.permissions import permitted_document_ids
from documents.permissions import permitted_object_ids
from documents.permissions import restrict_queryset_to_visible
@@ -32,6 +33,8 @@ from paperless_testing.permissions import grant_global
from paperless_testing.permissions import grant_object
if TYPE_CHECKING:
from django.contrib.auth.models import User
from paperless_testing.dirs import PaperlessDirs
@@ -178,6 +181,382 @@ class TestPermittedDocumentIdsIncludeDeleted:
)
@pytest.mark.django_db
class TestPermittedDocumentIdsVersions:
"""
A version is authorized by its root document: the version's own owner and
grants never matter.
"""
@pytest.mark.parametrize(
("root_owner", "version_owner", "expected_visible"),
[
pytest.param(
"other",
"nobody",
False,
id="unowned-version-of-private-root",
),
pytest.param("other", "user", False, id="own-version-of-private-root"),
pytest.param("user", "other", True, id="foreign-version-of-own-root"),
pytest.param("user", "nobody", True, id="unowned-version-of-own-root"),
pytest.param("nobody", "other", True, id="private-version-of-unowned-root"),
],
)
def test_version_follows_root_owner(
self,
root_owner: str,
version_owner: str,
*,
expected_visible: bool,
) -> None:
"""
GIVEN:
- A root document and a version with differing owners
WHEN:
- The permitted document ids are resolved for the user
THEN:
- The version is visible exactly when its root is
"""
user = UserFactory()
owners = {"user": user, "other": UserFactory(), "nobody": None}
root = DocumentFactory(owner=owners[root_owner])
version = DocumentFactory(root_document=root, owner=owners[version_owner])
visible = set(permitted_document_ids(user))
assert (version.pk in visible) is expected_visible
assert (root.pk in visible) is expected_visible
@staticmethod
def grantee(user: User, kind: str) -> User | Group:
"""The user itself, or a new group the user belongs to."""
if kind == "user":
return user
group = Group.objects.create(name="shared")
user.groups.add(group)
return group
@pytest.mark.parametrize(
"grantee_kind",
[pytest.param("user", id="user"), pytest.param("group", id="group")],
)
def test_grant_on_root_applies_to_version(self, grantee_kind: str) -> None:
"""
GIVEN:
- A private root document shared with a user or one of their groups
- A version of it owned by someone else
WHEN:
- The permitted document ids are resolved for the user
THEN:
- Both the root and the version are visible
- A user without the grant sees neither
"""
user = UserFactory()
stranger = UserFactory()
root = DocumentFactory(owner=UserFactory())
version = DocumentFactory(root_document=root, owner=UserFactory())
grant_object(self.grantee(user, grantee_kind), root, "view_document")
assert_visible_document_ids(
permitted_document_ids(user),
expected_visible=[root.pk, version.pk],
expected_hidden=[],
)
assert_visible_document_ids(
permitted_document_ids(stranger),
expected_visible=[],
expected_hidden=[root.pk, version.pk],
)
@pytest.mark.parametrize(
"grantee_kind",
[pytest.param("user", id="user"), pytest.param("group", id="group")],
)
def test_grant_on_version_is_ignored(self, grantee_kind: str) -> None:
"""
GIVEN:
- A private root document
- A version with an explicit grant for the user or one of their groups
WHEN:
- The permitted document ids are resolved for the user
THEN:
- Neither the root nor the version is visible
"""
user = UserFactory()
root = DocumentFactory(owner=UserFactory())
version = DocumentFactory(root_document=root, owner=UserFactory())
grant_object(self.grantee(user, grantee_kind), version, "view_document")
assert_visible_document_ids(
permitted_document_ids(user),
expected_visible=[],
expected_hidden=[root.pk, version.pk],
)
def test_grant_on_one_root_does_not_reach_another_roots_version(self) -> None:
"""
GIVEN:
- Two private roots, each with a version
- The user may view only the first root
WHEN:
- The permitted document ids are resolved for the user
THEN:
- Only the first root and its version are visible
"""
user = UserFactory()
first = DocumentFactory(owner=UserFactory())
first_version = DocumentFactory(root_document=first, owner=UserFactory())
second = DocumentFactory(owner=UserFactory())
second_version = DocumentFactory(root_document=second, owner=user)
grant_object(user, first, "view_document")
assert_visible_document_ids(
permitted_document_ids(user),
expected_visible=[first.pk, first_version.pk],
expected_hidden=[second.pk, second_version.pk],
)
def test_user_in_several_groups(self) -> None:
"""
GIVEN:
- A user in two groups
- Two private roots shared with one group each, and a third shared with nobody
- A version of each root
WHEN:
- The permitted document ids are resolved for the user
THEN:
- The two shared roots and their versions are visible
- The third root and its version are not
"""
user = UserFactory()
groups = [Group.objects.create(name=f"group{i}") for i in range(2)]
user.groups.add(*groups)
shared = [DocumentFactory(owner=UserFactory()) for _ in groups]
for root, group in zip(shared, groups, strict=True):
grant_object(group, root, "view_document")
unshared = DocumentFactory(owner=UserFactory())
shared_versions = [
DocumentFactory(root_document=root, owner=UserFactory()) for root in shared
]
unshared_version = DocumentFactory(root_document=unshared, owner=None)
assert_visible_document_ids(
permitted_document_ids(user),
expected_visible=[
*(root.pk for root in shared),
*(version.pk for version in shared_versions),
],
expected_hidden=[unshared.pk, unshared_version.pk],
)
def test_permission_is_resolved_through_the_root(self) -> None:
"""
GIVEN:
- A private root document where the user may view and change
WHEN:
- The permitted ids are resolved for view, change and delete
THEN:
- The version is visible for view and change only
"""
user = UserFactory()
root = DocumentFactory(owner=UserFactory())
version = DocumentFactory(root_document=root, owner=UserFactory())
grant_object(user, root, "view_document", "change_document")
assert version.pk in set(permitted_document_ids(user))
assert version.pk in set(permitted_document_ids(user, perm="change_document"))
assert version.pk in set(
permitted_document_ids(user, perm="documents.change_document"),
)
assert version.pk not in set(
permitted_document_ids(user, perm="delete_document"),
)
def test_anonymous_sees_versions_of_unowned_roots_only(self) -> None:
"""
GIVEN:
- A version owned by nobody under a private root
- A version owned by someone under an unowned root
WHEN:
- The permitted document ids are resolved for an anonymous user
THEN:
- Only the version of the unowned root is visible
"""
private_root = DocumentFactory(owner=UserFactory())
private_version = DocumentFactory(root_document=private_root, owner=None)
open_root = DocumentFactory(owner=None)
open_version = DocumentFactory(root_document=open_root, owner=UserFactory())
assert_visible_document_ids(
permitted_document_ids(AnonymousUser()),
expected_visible=[open_root.pk, open_version.pk],
expected_hidden=[private_root.pk, private_version.pk],
)
def test_deleted_versions_follow_their_deleted_root(self) -> None:
"""
GIVEN:
- A soft-deleted root document and its version, which deleting the
root soft-deletes too; the version is owned by someone else
WHEN:
- The permitted document ids are resolved with and without deleted
documents
THEN:
- Nothing is visible by default
- With deleted documents included, the version is visible to the
root's owner and not to the version's own owner
"""
owner = UserFactory()
version_owner = UserFactory()
root = DocumentFactory(owner=owner)
version = DocumentFactory(root_document=root, owner=version_owner)
root.delete()
assert not {root.pk, version.pk} & set(permitted_document_ids(owner))
assert_visible_document_ids(
permitted_document_ids(owner, include_deleted=True),
expected_visible=[root.pk, version.pk],
expected_hidden=[],
)
assert_visible_document_ids(
permitted_document_ids(version_owner, include_deleted=True),
expected_visible=[],
expected_hidden=[root.pk, version.pk],
)
@pytest.mark.parametrize(
"is_superuser",
[
pytest.param(False, id="regular-user"),
pytest.param(True, id="superuser"),
],
)
def test_inactive_user_sees_no_versions(self, *, is_superuser: bool) -> None:
"""
GIVEN:
- An inactive user, possibly a superuser, who owns a root and its version
WHEN:
- The permitted document ids are resolved for them
THEN:
- Nothing is visible
"""
user = UserFactory(is_active=False, is_superuser=is_superuser)
root = DocumentFactory(owner=user)
version = DocumentFactory(root_document=root, owner=user)
assert_visible_document_ids(
permitted_document_ids(user),
expected_visible=[],
expected_hidden=[root.pk, version.pk],
)
def test_superuser_sees_all_versions(self) -> None:
"""
GIVEN:
- A private root owned by someone else, with a version
WHEN:
- The permitted document ids are resolved for a superuser
THEN:
- Both the root and the version are visible
"""
superuser = UserFactory(superuser=True)
root = DocumentFactory(owner=UserFactory())
version = DocumentFactory(root_document=root, owner=UserFactory())
assert_visible_document_ids(
permitted_document_ids(superuser),
expected_visible=[root.pk, version.pk],
expected_hidden=[],
)
@pytest.mark.django_db
class TestHasPermsOwnerAwareVersions:
"""
The single-object check agrees with permitted_document_ids: a version is
authorized by its root document.
"""
@pytest.mark.parametrize(
("root_owner", "version_owner", "expected"),
[
pytest.param(
"other",
"nobody",
False,
id="unowned-version-of-private-root",
),
pytest.param("other", "user", False, id="own-version-of-private-root"),
pytest.param("user", "other", True, id="foreign-version-of-own-root"),
pytest.param("nobody", "other", True, id="private-version-of-unowned-root"),
],
)
def test_version_follows_root_owner(
self,
root_owner: str,
version_owner: str,
*,
expected: bool,
) -> None:
"""
GIVEN:
- A root document and a version with differing owners
WHEN:
- The single-object check runs for the version
THEN:
- The version is allowed exactly when its root is
"""
user = UserFactory()
owners = {"user": user, "other": UserFactory(), "nobody": None}
root = DocumentFactory(owner=owners[root_owner])
version = DocumentFactory(root_document=root, owner=owners[version_owner])
assert has_perms_owner_aware(user, "view_document", version) is expected
assert has_perms_owner_aware(user, "view_document", root) is expected
def test_grant_on_root_applies_and_grant_on_version_does_not(self) -> None:
"""
GIVEN:
- A private root with a version, and a second private root with a version
- The user may change only the first root, and was granted the second
root's version directly
WHEN:
- The single-object check runs for each version
THEN:
- Only the first root's version is allowed
"""
user = UserFactory()
shared_root = DocumentFactory(owner=UserFactory())
shared_version = DocumentFactory(root_document=shared_root, owner=UserFactory())
private_root = DocumentFactory(owner=UserFactory())
private_version = DocumentFactory(
root_document=private_root,
owner=UserFactory(),
)
grant_object(user, shared_root, "change_document")
grant_object(user, private_version, "change_document")
assert has_perms_owner_aware(user, "change_document", shared_version)
assert not has_perms_owner_aware(user, "change_document", private_version)
def test_other_models_use_their_own_owner(self) -> None:
"""
GIVEN:
- A tag owned by someone else, and one owned by the user
WHEN:
- The single-object check runs for each
THEN:
- Only the user's own tag is allowed without a grant
"""
user = UserFactory()
mine = TagFactory(owner=user)
theirs = TagFactory(owner=UserFactory())
assert has_perms_owner_aware(user, "view_tag", mine)
assert not has_perms_owner_aware(user, "view_tag", theirs)
@pytest.mark.django_db
class TestAiChatAllDocumentsPermissionBoundary:
"""
@@ -90,10 +90,13 @@ class ShareLinkBundleAPITests(DirectoriesMixin, APITestCase):
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
self.assertIn("document_ids", response.data)
@mock.patch("documents.views.permitted_document_ids", return_value=set())
def test_create_bundle_rejects_insufficient_permissions(self, perms_mock) -> None:
def test_create_bundle_rejects_insufficient_permissions(self) -> None:
requester = UserFactory(username="bundle_creator")
grant_global(requester, "add_sharelinkbundle", "view_document")
self.client.force_authenticate(requester)
document = DocumentFactory(owner=UserFactory(username="document_owner"))
payload = {
"document_ids": [self.document.pk],
"document_ids": [self.document.pk, document.pk],
"file_version": ShareLink.FileVersion.ARCHIVE,
"expiration_days": 7,
}
@@ -101,8 +104,8 @@ class ShareLinkBundleAPITests(DirectoriesMixin, APITestCase):
response = self.client.post(self.ENDPOINT, payload, format="json")
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
self.assertIn("document_ids", response.data)
perms_mock.assert_called()
self.assertIn(str(document.pk), str(response.data["document_ids"]))
self.assertFalse(ShareLinkBundle.objects.exists())
@mock.patch("documents.views.build_share_link_bundle.apply_async")
def test_rebuild_bundle_resets_state(self, delay_mock) -> None:
+75
View File
@@ -20,6 +20,7 @@ from documents.sanity_checker import SanityCheckMessages
from documents.tests.helpers import dummy_preprocess
from paperless_testing.assertions import FileSystemAssertsMixin
from paperless_testing.dirs import DirectoriesMixin
from paperless_testing.factories import DocumentFactory
@pytest.mark.django_db
@@ -287,6 +288,80 @@ class TestUpdateContent(DirectoriesMixin, TestCase):
tasks.update_document_content_maybe_archive_file(doc.pk)
self.assertNotEqual(Document.objects.get(pk=doc.pk).content, "test")
def _create_root_with_version(self) -> tuple[Document, Document]:
sample1 = self.dirs.scratch_dir / "sample.pdf"
shutil.copy(
Path(__file__).parent
/ "samples"
/ "documents"
/ "originals"
/ "0000001.pdf",
sample1,
)
root = DocumentFactory(content="root content", mime_type="application/pdf")
version = DocumentFactory(
content="my document",
filename=sample1,
mime_type="application/pdf",
root_document=root,
version_index=1,
)
return root, version
@mock.patch("documents.tasks.clear_document_caches")
@mock.patch("documents.search.get_backend")
def test_update_content_version_clears_caches_for_root(
self,
mock_get_backend: mock.Mock,
mock_clear_caches: mock.Mock,
) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- Update content task is called for the version
THEN:
- The version's content is updated, not the root's
- The document is indexed
- Caches are cleared for both
"""
root, version = self._create_root_with_version()
tasks.update_document_content_maybe_archive_file(version.pk)
self.assertNotEqual(
Document.objects.get(pk=version.pk).content,
"my document",
)
self.assertEqual(Document.objects.get(pk=root.pk).content, "root content")
mock_get_backend.return_value.add_or_update.assert_called_once()
mock_clear_caches.assert_has_calls(
[mock.call(version.pk), mock.call(root.pk)],
)
@override_settings(AI_ENABLED=True, LLM_EMBEDDING_BACKEND="huggingface")
@mock.patch("documents.tasks.llm_index_add_or_update_document")
@mock.patch("documents.search.get_backend")
def test_update_content_version_updates_llm_index(
self,
mock_get_backend: mock.Mock,
mock_llm_index: mock.Mock,
) -> None:
"""
GIVEN:
- A root document with a version
- The LLM index is enabled
WHEN:
- Update content task is called for the version
THEN:
- The LLM index is updated
"""
_, version = self._create_root_with_version()
tasks.update_document_content_maybe_archive_file(version.pk)
mock_llm_index.assert_called_once()
class TestUpdateContentRemoteOCR(DirectoriesMixin, TestCase):
"""
+17
View File
@@ -17,9 +17,26 @@ from django.db.models.functions import RowNumber
from documents.models import Document
if TYPE_CHECKING:
from collections.abc import Iterable
from rest_framework.request import Request
def root_document_ids(ids: Iterable[int]) -> QuerySet[int]:
"""
The ids of the root documents of the given documents: a root stands for
itself and a version for its root. Only the indexes' bookkeeping needs
this, since they hold root documents only.
"""
return (
Document.objects.filter(pk__in=ids)
.annotate(root_id=Coalesce("root_document_id", "id"))
.order_by()
.values_list("root_id", flat=True)
.distinct()
)
def versions_newest_first(documents: QuerySet[Document]) -> QuerySet[Document]:
"""
Sorts versions so the newest one comes first using version_index and not on id,
+84 -51
View File
@@ -329,9 +329,10 @@ def _get_tantivy_query_and_mode(params):
def _get_more_like_id(query_params: dict[str, Any], user: User | None) -> int:
try:
more_like_doc_id = int(query_params["more_like_id"])
more_like_doc = Document.objects.select_related("owner").get(
pk=more_like_doc_id,
)
more_like_doc = Document.objects.select_related(
"owner",
"root_document__owner",
).get(pk=more_like_doc_id)
except (TypeError, ValueError, Document.DoesNotExist):
raise PermissionDenied(_("Invalid more_like_id"))
@@ -342,7 +343,8 @@ def _get_more_like_id(query_params: dict[str, Any], user: User | None) -> int:
):
raise PermissionDenied(_("Insufficient permissions."))
return more_like_doc_id
# Only root documents are indexed, a version stands for its root
return more_like_doc.root_document_id or more_like_doc.pk
class SearchParams(NamedTuple):
@@ -1550,7 +1552,10 @@ class DocumentViewSet(
)
def suggestions(self, request, pk=None):
doc = get_object_or_404(
Document.objects.select_related("owner").prefetch_related("versions"),
Document.objects.select_related(
"owner",
"root_document__owner",
).prefetch_related("versions"),
pk=pk,
)
if request.user is not None and not has_perms_owner_aware(
@@ -1610,7 +1615,10 @@ class DocumentViewSet(
@method_decorator(cache_control(no_cache=True))
def ai_suggestions(self, request, pk=None):
doc = get_object_or_404(
Document.objects.select_related("owner").prefetch_related("versions"),
Document.objects.select_related(
"owner",
"root_document__owner",
).prefetch_related("versions"),
pk=pk,
)
if request.user is not None and not has_perms_owner_aware(
@@ -1856,9 +1864,14 @@ class DocumentViewSet(
currentUser = request.user
try:
doc = (
Document.objects.select_related("owner")
Document.objects.select_related("owner", "root_document__owner")
.prefetch_related("notes")
.only("pk", "owner__id")
.only(
"pk",
"owner__id",
"root_document__id",
"root_document__owner__id",
)
.get(pk=pk)
)
if currentUser is not None and not has_perms_owner_aware(
@@ -1973,7 +1986,9 @@ class DocumentViewSet(
def share_links(self, request, pk=None):
currentUser = request.user
try:
doc = Document.objects.select_related("owner").get(pk=pk)
doc = Document.objects.select_related("owner", "root_document__owner").get(
pk=pk,
)
if currentUser is not None and not has_perms_owner_aware(
currentUser,
"change_document",
@@ -2008,10 +2023,11 @@ class DocumentViewSet(
if not settings.AUDIT_LOG_ENABLED:
return HttpResponseBadRequest("Audit log is disabled")
try:
doc = Document.objects.get(pk=pk)
doc = Document.objects.select_related("root_document__owner").get(pk=pk)
root_doc = get_root_document(doc)
if not request.user.has_perm("auditlog.view_logentry") or (
doc.owner is not None
and doc.owner != request.user
root_doc.owner is not None
and root_doc.owner != request.user
and not request.user.is_superuser
):
return HttpResponseForbidden(
@@ -2102,9 +2118,7 @@ class DocumentViewSet(
documents = Document.objects.filter(pk__in=document_ids)
if (
request.user is not None
and documents.exclude(
pk__in=permitted_document_ids(request.user),
).exists()
and documents.exclude(id__in=permitted_document_ids(request.user)).exists()
):
return HttpResponseForbidden("Insufficient permissions")
@@ -2430,11 +2444,17 @@ class ChatStreamingView(GenericAPIView[Any]):
if doc_id:
try:
document = Document.objects.get(id=doc_id)
document = Document.objects.select_related(
"root_document__owner",
).get(id=doc_id)
except Document.DoesNotExist:
return HttpResponseBadRequest("Document not found")
if not has_perms_owner_aware(request.user, "view_document", document):
if not has_perms_owner_aware(
request.user,
"view_document",
document,
):
return HttpResponseForbidden("Insufficient permissions")
documents = Document.objects.filter(pk=document.pk)
@@ -2990,12 +3010,8 @@ class DocumentOperationPermissionMixin(PassUserMixin, DocumentSelectionMixin):
user.has_perm(
"documents.change_document",
)
and not Document.global_objects.filter(
pk__in=[doc.pk for doc in root_docs],
)
.exclude(
pk__in=permitted_document_ids(user, perm="change_document"),
)
and not Document.global_objects.filter(pk__in=documents)
.exclude(pk__in=permitted_document_ids(user, perm="change_document"))
.exists()
)
@@ -3605,10 +3621,11 @@ class SelectionDataView(DocumentSelectionMixin, GenericAPIView[Any]):
user=request.user,
validated_data=serializer.validated_data,
)
permitted_documents = Document.objects.filter(
id__in=permitted_document_ids(request.user),
)
if permitted_documents.filter(pk__in=ids).count() != len(ids):
documents = Document.objects.filter(pk__in=ids)
if (
documents.count() != len(ids)
or documents.exclude(id__in=permitted_document_ids(request.user)).exists()
):
return HttpResponseForbidden("Insufficient permissions")
correspondents = Correspondent.objects.annotate(
@@ -4104,21 +4121,16 @@ class BulkDownloadView(DocumentSelectionMixin, GenericAPIView[Any]):
validated_data=serializer.validated_data,
)
documents = Document.objects.filter(pk__in=ids)
versioned_documents = []
compression = serializer.validated_data.get("compression")
content = serializer.validated_data.get("content")
follow_filename_format = serializer.validated_data.get("follow_formatting")
permitted_ids = set(permitted_document_ids(request.user))
for document in documents:
root_doc = get_root_document(document)
if root_doc.pk not in permitted_ids:
return HttpResponseForbidden("Insufficient permissions")
versioned_documents.append(
get_latest_version_for_root(
root_doc,
),
)
if documents.exclude(id__in=permitted_document_ids(request.user)).exists():
return HttpResponseForbidden("Insufficient permissions")
versioned_documents = [
get_latest_version_for_root(get_root_document(document))
for document in documents
]
if content == "both":
strategy_class = OriginalAndArchiveStrategy
@@ -4790,19 +4802,23 @@ class ShareLinkBundleViewSet(PassUserMixin, ModelViewSet[ShareLinkBundle]):
},
)
documents = list(documents_qs)
permitted_ids = set(permitted_document_ids(request.user))
for document in documents:
if document.pk not in permitted_ids:
raise ValidationError(
{
"document_ids": _(
"Insufficient permissions to share document %(id)s.",
)
% {"id": document.pk},
},
)
denied_id = (
documents_qs.exclude(id__in=permitted_document_ids(request.user))
.order_by("pk")
.values_list("pk", flat=True)
.first()
)
if denied_id is not None:
raise ValidationError(
{
"document_ids": _(
"Insufficient permissions to share document %(id)s.",
)
% {"id": denied_id},
},
)
documents = list(documents_qs)
document_map = {document.pk: document for document in documents}
ordered_documents = [document_map[doc_id] for doc_id in document_ids]
@@ -5587,6 +5603,23 @@ class TrashView(ListModelMixin, PassUserMixin):
class _TrashPermittedObjectsFilter(PermittedObjectsFilter):
include_granted = False
def filter_queryset(self, request, queryset, view):
if request.user.is_superuser or not request.user.is_active:
return super().filter_queryset(request, queryset, view)
# A version belongs to whoever owns its root
def owned_or_unowned(prefix: str) -> Q:
return Q(**{f"{prefix}owner": request.user}) | Q(
**{f"{prefix}owner__isnull": True},
)
return queryset.filter(
(Q(root_document__isnull=True) & owned_or_unowned(""))
| (
Q(root_document__isnull=False) & owned_or_unowned("root_document__")
),
)
filter_backends = (_TrashPermittedObjectsFilter,)
pagination_class = StandardPagination
@@ -5617,7 +5650,7 @@ class TrashView(ListModelMixin, PassUserMixin):
else self.filter_queryset(self.get_queryset()).all()
)
if docs.exclude(
pk__in=permitted_document_ids(
id__in=permitted_document_ids(
request.user,
perm="delete_document",
include_deleted=True,
+22 -39
View File
@@ -2,7 +2,7 @@ msgid ""
msgstr ""
"Project-Id-Version: paperless-ngx\n"
"Report-Msgid-Bugs-To: \n"
"POT-Creation-Date: 2026-10-05 16:26+0000\n"
"POT-Creation-Date: 2026-10-08 18:39+0000\n"
"PO-Revision-Date: 2022-02-17 04:17\n"
"Last-Translator: \n"
"Language-Team: English\n"
@@ -1652,49 +1652,49 @@ msgstr ""
msgid "workflow runs"
msgstr ""
#: documents/serialisers.py:515 documents/serialisers.py:872
#: documents/serialisers.py:2902 documents/views.py:343 documents/views.py:2732
#: documents/serialisers.py:516 documents/serialisers.py:873
#: documents/serialisers.py:2903 documents/views.py:344 documents/views.py:2751
#: paperless_mail/serialisers.py:156
msgid "Insufficient permissions."
msgstr ""
#: documents/serialisers.py:708
#: documents/serialisers.py:709
msgid "Invalid color."
msgstr ""
#: documents/serialisers.py:2369
#: documents/serialisers.py:2370
#, python-format
msgid "File type %(type)s not supported"
msgstr ""
#: documents/serialisers.py:2413
#: documents/serialisers.py:2414
#, python-format
msgid "Custom field id must be an integer: %(id)s"
msgstr ""
#: documents/serialisers.py:2420
#: documents/serialisers.py:2421
#, python-format
msgid "Custom field with id %(id)s does not exist"
msgstr ""
#: documents/serialisers.py:2437 documents/serialisers.py:2447
#: documents/serialisers.py:2438 documents/serialisers.py:2448
msgid ""
"Custom fields must be a list of integers or an object mapping ids to values."
msgstr ""
#: documents/serialisers.py:2442
#: documents/serialisers.py:2443
msgid "Some custom fields don't exist or were specified twice."
msgstr ""
#: documents/serialisers.py:2589
#: documents/serialisers.py:2590
msgid "Invalid variable detected."
msgstr ""
#: documents/serialisers.py:2958
#: documents/serialisers.py:2959
msgid "Duplicate document identifiers are not allowed."
msgstr ""
#: documents/serialisers.py:2988 documents/views.py:4787
#: documents/serialisers.py:2989 documents/views.py:4807
#, python-format
msgid "Documents not found: %(ids)s"
msgstr ""
@@ -1941,61 +1941,44 @@ msgstr ""
msgid "As a final step, please complete the following form:"
msgstr ""
#: documents/validators.py:24
#, python-brace-format
msgid "Unable to parse URI {value}, missing scheme"
msgstr ""
#: documents/validators.py:29
#, python-brace-format
msgid "Unable to parse URI {value}, missing net location or path"
msgstr ""
#: documents/validators.py:36
msgid ""
"URI scheme '{parts.scheme}' is not allowed. Allowed schemes: {', '."
"join(allowed_schemes)}"
msgid ", "
msgstr ""
#: documents/validators.py:45
#, python-brace-format
msgid "Unable to parse URI {value}"
msgstr ""
#: documents/views.py:336 documents/views.py:2729
#: documents/views.py:337 documents/views.py:2748
msgid "Invalid more_like_id"
msgstr ""
#: documents/views.py:1676
#: documents/views.py:1683
msgid "Invalid AI configuration."
msgstr ""
#: documents/views.py:1687
#: documents/views.py:1694
msgid "AI backend request timed out."
msgstr ""
#: documents/views.py:1699
#: documents/views.py:1706
msgid "AI backend rejected the request. Check logs for details."
msgstr ""
#: documents/views.py:2554 documents/views.py:2870
#: documents/views.py:2573 documents/views.py:2889
msgid "Specify only one of text, title_search, query, or more_like_id."
msgstr ""
#: documents/views.py:4800
#: documents/views.py:4823
#, python-format
msgid "Insufficient permissions to share document %(id)s."
msgstr ""
#: documents/views.py:4846
#: documents/views.py:4870
msgid "Bundle is already being processed."
msgstr ""
#: documents/views.py:4910
#: documents/views.py:4934
msgid "The share link bundle is still being prepared. Please try again later."
msgstr ""
#: documents/views.py:4924
#: documents/views.py:4948
msgid "The share link bundle is unavailable."
msgstr ""
@@ -137,9 +137,20 @@ class TestNginxService:
reason="No Gotenberg/Tika servers to test with",
)
class TestParserLive:
@staticmethod
def imagehash(file: Path, hash_size: int = 18) -> str:
return f"{average_hash(Image.open(file), hash_size)}"
# Rasterizer versions shift a few pixels, so compare perceptual hashes by
# Hamming distance (out of 18 * 18 = 324 bits) rather than for equality
MAX_HASH_DISTANCE = 8
@classmethod
def assert_thumbnails_similar(cls, generated: Path, expected: Path) -> None:
distance = average_hash(Image.open(generated), 18) - average_hash(
Image.open(expected),
18,
)
assert distance <= cls.MAX_HASH_DISTANCE, (
f"Thumbnail {generated} differs from {expected} by {distance} bits "
f"(max {cls.MAX_HASH_DISTANCE})"
)
def test_get_thumbnail(
self,
@@ -168,12 +179,7 @@ class TestParserLive:
assert thumb.exists()
assert thumb.is_file()
assert self.imagehash(thumb) == self.imagehash(
simple_txt_email_thumbnail_file,
), (
f"Created thumbnail {thumb} differs from expected file "
f"{simple_txt_email_thumbnail_file}"
)
self.assert_thumbnails_similar(thumb, simple_txt_email_thumbnail_file)
def test_tika_parse_successful(self, mail_parser: MailDocumentParser) -> None:
"""
@@ -255,7 +261,7 @@ class TestParserLive:
THEN:
- Gotenberg shall be called to generate the PDF
- The archive PDF shall contain the expected content
- The generated thumbnail shall match the expected image hash
- The generated thumbnail shall be perceptually close to the expected image
"""
util_call_with_backoff(mail_parser.parse, [html_email_file, "message/rfc822"])
@@ -272,14 +278,4 @@ class TestParserLive:
html_email_file,
"message/rfc822",
)
generated_thumbnail_hash = self.imagehash(generated_thumbnail)
# The created PDF is not reproducible, but the converted image
# should always look the same
expected_hash = self.imagehash(html_email_thumbnail_file)
assert generated_thumbnail_hash == expected_hash, (
f"PDF thumbnail differs from expected. "
f"Generated: {generated_thumbnail}, "
f"Hash: {generated_thumbnail_hash} vs {expected_hash}"
)
self.assert_thumbnails_similar(generated_thumbnail, html_email_thumbnail_file)
+1 -1
View File
@@ -1,6 +1,6 @@
from typing import Final
__version__: Final[tuple[int, int, int]] = (3, 3, 0)
__version__: Final[tuple[int, int, int]] = (3, 2, 1)
# Version string like X.Y.Z
__full_version_str__: Final[str] = ".".join(map(str, __version__))
# Version string like X.Y
+6 -3
View File
@@ -7,6 +7,7 @@ from documents.models import Document
from documents.permissions import permitted_object_ids
from documents.permissions import restrict_queryset_to_visible
from documents.permissions import user_is_unrestricted
from documents.versioning import annotate_effective_content
from paperless.config import AIConfig
from paperless_ai.base_model import ClassificationSuggestions
from paperless_ai.base_model import TaxonomyChoiceDict
@@ -113,7 +114,7 @@ def build_prompt_without_rag(
) -> str:
filename = document.filename or ""
content = truncate_content(
document.content[:4000] or "",
(document.get_effective_content() or "")[:4000],
chunk_size=config.llm_embedding_chunk_size,
context_size=config.llm_context_size,
)
@@ -227,7 +228,9 @@ def get_taxonomy_context(
# similar_documents is already ordered by descending weight; don't lose it.
similar_document_ids = [s["document_id"] for s in similar_documents]
similar_documents_by_id = Document.objects.in_bulk(similar_document_ids)
similar_documents_by_id = annotate_effective_content(
Document.objects.all(),
).in_bulk(similar_document_ids)
similar_docs = [
similar_documents_by_id[document_id]
for document_id in similar_document_ids
@@ -235,7 +238,7 @@ def get_taxonomy_context(
][:max_docs]
context_blocks = []
for similar in similar_docs:
text = similar.content[:1000] or ""
text = (similar.get_effective_content() or "")[:1000]
title = similar.title or similar.filename or "Untitled"
context_blocks.append(f"TITLE: {title}\n{text}")
except Exception:
+1 -1
View File
@@ -135,6 +135,6 @@ def build_llm_index_text(doc: Document) -> str:
lines.append(f"Custom Field - {instance.field.name}: {instance}")
lines.append("\nContent:\n")
lines.append(doc.content or "")
lines.append(doc.get_effective_content() or "")
return _normalize_llm_index_text("\n".join(lines))
+15 -8
View File
@@ -17,6 +17,8 @@ from documents.models import PaperlessTask
from documents.utils import IterWrapper
from documents.utils import QuerySetStream
from documents.utils import identity
from documents.versioning import annotate_effective_content
from documents.versioning import root_document_ids
from paperless.config import AIConfig
from paperless_ai.db import db_connection_released
from paperless_ai.embedding import build_llm_index_text
@@ -443,11 +445,11 @@ def update_llm_index(
"Skipping LLM index update: migration check deferred; "
"will retry next run."
)
documents = Document.objects.select_related(
"correspondent",
"document_type",
"storage_path",
).prefetch_related("tags", "notes", "custom_fields__field")
documents = annotate_effective_content(
Document.objects.filter(root_document__isnull=True)
.select_related("correspondent", "document_type", "storage_path")
.prefetch_related("tags", "notes", "custom_fields__field"),
)
no_documents = not documents.exists()
# Fast exit before touching config: nothing to index and no existing index.
@@ -483,7 +485,7 @@ def update_llm_index(
msg = "LLM index rebuilt successfully."
else:
scoped_documents = (
documents.filter(id__in=document_ids)
documents.filter(id__in=root_document_ids(document_ids))
if document_ids is not None
else documents
)
@@ -510,7 +512,12 @@ def update_llm_index(
def llm_index_add_or_update_document(document: Document):
"""Add or atomically replace a document's chunks in the index."""
"""
Add or atomically replace a document's chunks in the index. Only root
documents are indexed, so a version is indexed as its root document.
"""
if document.root_document_id is not None:
document = document.root_document
config = AIConfig()
new_nodes = build_document_node(
document,
@@ -688,7 +695,7 @@ def retrieve_similar_nodes(
)
query_text = truncate_embedding_query(
(document.title or "") + "\n" + (document.content or ""),
(document.title or "") + "\n" + (document.get_effective_content() or ""),
chunk_size=config.llm_embedding_chunk_size,
)
# Hold the shared read lock for the whole retrieval so the connection is
@@ -55,6 +55,7 @@ def mock_document():
doc.storage_path = None
doc.archive_serial_number = "12345"
doc.content = "This is the document content."
doc.get_effective_content.return_value = "This is the document content."
cf1 = MagicMock(__str__=lambda x: "Value1")
cf1.field = MagicMock()
@@ -434,6 +435,55 @@ def test_get_taxonomy_context_preserves_similarity_order_and_distinct_documents(
)
@pytest.mark.django_db
class TestClassifierEffectiveContent:
"""A root document's text for the LLM is its newest version's content."""
@staticmethod
def _root_with_version() -> Document:
root = DocumentFactory(title="Statement", content="stale text")
DocumentFactory(root_document=root, version_index=1, content="latest text")
return root
def test_prompt_uses_the_newest_versions_content(self) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- The classification prompt is built for the root
THEN:
- It contains the newest version's content
"""
prompt = build_prompt_without_rag(self._root_with_version(), AIConfig())
assert "latest text" in prompt
assert "stale text" not in prompt
@override_settings(LLM_EMBEDDING_BACKEND="huggingface")
def test_similar_document_context_uses_the_newest_versions_content(self) -> None:
"""
GIVEN:
- A similar root document with a version
WHEN:
- The similar-document context is built
THEN:
- It contains the newest version's content
"""
similar = self._root_with_version()
document = DocumentFactory(content="Some content")
fake_nodes = [
SimpleNamespace(metadata={"document_id": str(similar.pk)}, score=0.9),
]
with patch(
"paperless_ai.ai_classifier.retrieve_similar_nodes",
return_value=fake_nodes,
):
_candidates, context = get_taxonomy_context(document, user=None)
assert context == "TITLE: Statement\nlatest text"
@pytest.mark.django_db
@override_settings(LLM_EMBEDDING_BACKEND="huggingface")
def test_get_taxonomy_context_no_similar_docs():
+103
View File
@@ -388,6 +388,108 @@ def test_update_llm_index_partial_update(
assert after[str(doc2.pk)] == before[str(doc2.pk)]
@pytest.mark.django_db
class TestLlmIndexVersions:
"""The LLM index holds root documents only: a version is indexed as its root."""
def test_add_or_update_document_indexes_a_version_as_its_root(
self,
temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding,
) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- The version is passed to llm_index_add_or_update_document
THEN:
- Only the root document is in the index
"""
root = DocumentFactory(content="root content")
version = DocumentFactory(root_document=root, version_index=1)
indexing.llm_index_add_or_update_document(version)
with indexing.get_vector_store() as store:
indexed = store.get_modified_times()
assert set(indexed) == {str(root.pk)}
def test_rebuild_skips_versions(
self,
temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding,
) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- The LLM index is rebuilt
THEN:
- Only the root document is in the index
"""
root = DocumentFactory()
DocumentFactory(root_document=root, version_index=1)
indexing.update_llm_index(rebuild=True)
with indexing.get_vector_store() as store:
indexed = store.get_modified_times()
assert set(indexed) == {str(root.pk)}
def test_rebuild_indexes_the_newest_versions_content(
self,
temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding,
mocker: pytest_mock.MockerFixture,
) -> None:
"""
GIVEN:
- A root document with a version
WHEN:
- The LLM index is rebuilt
THEN:
- The root's text for the index is the version's content, answered
from the query rather than a query per document
"""
root = DocumentFactory(content="stale text")
DocumentFactory(root_document=root, version_index=1, content="latest text")
spy = mocker.spy(indexing, "build_document_node")
indexing.update_llm_index(rebuild=True)
indexed = spy.call_args.args[0]
assert indexed.pk == root.pk
assert indexed.effective_content == "latest text"
def test_incremental_update_by_version_id_refreshes_the_root(
self,
temp_llm_index_dir: Path,
mock_embed_model: FakeEmbedding,
) -> None:
"""
GIVEN:
- An indexed root document with a version whose root was modified since
WHEN:
- An incremental update is scoped to the version's id
THEN:
- The root's entry is refreshed and no entry exists for the version
"""
root = DocumentFactory()
version = DocumentFactory(root_document=root, version_index=1)
indexing.update_llm_index(rebuild=True)
Document.objects.filter(pk=root.pk).update(modified=timezone.now())
root.refresh_from_db()
indexing.update_llm_index(document_ids=[version.pk])
with indexing.get_vector_store() as store:
indexed = store.get_modified_times()
assert indexed == {str(root.pk): root.modified.isoformat()}
@pytest.mark.django_db
def test_add_or_update_document_updates_existing_entry(
temp_llm_index_dir: Path,
@@ -637,6 +739,7 @@ class TestLlmIndexAddOrUpdateDocumentEmptyContent:
doc = MagicMock(spec=Document)
doc.id = 42
doc.root_document_id = None
# Must not raise
indexing.llm_index_add_or_update_document(doc)
+40 -1
View File
@@ -12,6 +12,7 @@ from paperless_ai.embedding import _normalize_llm_index_text
from paperless_ai.embedding import build_llm_index_text
from paperless_ai.embedding import get_configured_model_name
from paperless_ai.embedding import get_embedding_model
from paperless_testing.factories import DocumentFactory
@pytest.fixture
@@ -46,6 +47,7 @@ def mock_document():
doc.correspondent.name = "Test Correspondent"
doc.archive_serial_number = "12345"
doc.content = "This is the document content."
doc.get_effective_content.return_value = "This is the document content."
cf1 = MagicMock(__str__=lambda x: "Value1")
cf1.field = MagicMock()
@@ -280,7 +282,7 @@ def test_build_llm_index_text(mock_document):
def test_build_llm_index_text_normalizes_ocr_punctuation_runs(mock_document):
mock_document.content = (
mock_document.get_effective_content.return_value = (
"Introduction ................................................ 7\n"
"Hardware Limitation ________________________________________ 9\n"
"Keep short punctuation like INV-100 and ellipses..."
@@ -294,6 +296,43 @@ def test_build_llm_index_text_normalizes_ocr_punctuation_runs(mock_document):
assert "ellipses..." in result
@pytest.mark.django_db
class TestBuildLlmIndexTextVersions:
"""A root document is indexed with its effective content, like in the search index."""
def test_root_uses_the_newest_versions_content(self) -> None:
"""
GIVEN:
- A root document with two versions
WHEN:
- The LLM index text is built for the root
THEN:
- It contains the newest version's content and not the others'
"""
root = DocumentFactory(content="stale text")
DocumentFactory(root_document=root, version_index=1, content="older text")
DocumentFactory(root_document=root, version_index=2, content="latest text")
text = build_llm_index_text(root)
assert "latest text" in text
assert "stale text" not in text
assert "older text" not in text
def test_root_without_versions_uses_its_own_content(self) -> None:
"""
GIVEN:
- A root document without versions
WHEN:
- The LLM index text is built for it
THEN:
- It contains the document's own content
"""
root = DocumentFactory(content="own text")
assert "own text" in build_llm_index_text(root)
def test_normalize_llm_index_text_collapses_ocr_leaders_without_joining_lines():
assert _normalize_llm_index_text("A........B\nC____D----E") == "A B\nC D E"
Generated
+1 -1
View File
@@ -2971,7 +2971,7 @@ wheels = [
[[package]]
name = "paperless-ngx"
version = "3.3.0"
version = "3.2.1"
source = { virtual = "." }
dependencies = [
{ name = "azure-ai-documentintelligence" },