Commit Graph
4251 Commits
Author SHA1 Message Date
stumpylog d9d44520bd test: add query-count regression coverage for custom_fields sync on update 2026-08-25 10:59:47 -07:00
stumpylog 500ff563cf refactor: replace drf-writable-nested's NestedUpdateMixin with explicit custom_fields sync 2026-08-25 10:38:51 -07:00
stumpylog 7f314b5149 test: characterize custom field instance delete-on-omit as a hard delete 2026-08-25 10:28:24 -07:00
stumpylog 132918821d Merge remote-tracking branch 'origin/perf/batch-custom-field-lookup' into tmp/perf-integration 2026-08-25 10:15:52 -07:00
stumpylog 512a3fe196 Merge remote-tracking branch 'origin/perf/batch-modify-custom-fields' into tmp/perf-integration 2026-08-25 10:15:48 -07:00
stumpylog b01e0368b7 Merge remote-tracking branch 'origin/perf/batch-tags-field-lookup' into tmp/perf-integration 2026-08-25 10:15:44 -07:00
stumpylog 1c82bd15c5 Merge branch 'perf/batch-set-permissions' into tmp/perf-integration 2026-08-25 10:15:39 -07:00
stumpylog cbac71c165 Perf: avoid unnecessary full-row fetches in batch permission assignment
set_permissions_for_objects now takes a model + pks instead of instances,
and identity filtering resolves straight to ids, so bulk-editing
permissions no longer materializes full Document/User/Group rows just to
read their pk/id. Row construction for bulk_create is also chunked to
bound peak memory for very large "apply to all" operations.
2026-08-25 09:27:24 -07:00
stumpylog 8780bcd5c7 Perf: batch guardian permission assignment in bulk-edit
bulk_edit.set_permissions and BulkEditObjectPermissionsView both
looped documents/objects and called set_permissions_for_object per
object, which itself calls guardian's assign_perm/remove_perm once
per (object, user) pair -- ~10-20+ queries per object, scaling with
selection size.

Added set_permissions_for_objects, a bulk equivalent that resolves
existing permission holders once across the whole batch (not once per
object) and applies changes with a small, batch-size-independent
number of queries per action instead of one per (object, user) pair.

Deliberately does not use guardian's queryset-aware assign_perm:
passing a list as the target routes to bulk_assign_perm, which skips
creating a direct permission row for anyone who already has the
permission via ANY group membership (checked via
ObjectPermissionChecker.has_perm, which is group-inheritance-aware) --
unlike the single-object assign_perm this replaces, which always
ensures a direct row via get_or_create. Losing that guarantee would
mean a later revocation of the group's grant silently strips access
an admin explicitly asked to be direct. Bulk-creates rows straight
against UserObjectPermission/GroupObjectPermission instead
(ignore_conflicts=True, relying on the existing (identity, permission,
object_pk) unique constraint), which preserves the original semantics
exactly while still batching every object and identity into one query
per action. Also raises Permission.DoesNotExist for an unrecognized
action name instead of silently no-op-ing, matching the original
per-object path -- BulkEditObjectsSerializer never actually validates
action keys against the raw client-supplied permissions dict, so this
is reachable from client input, not just internal callers.

Verified via CaptureQueriesContext: query count is now identical at 5
vs. 50 documents/objects (was 1,123 queries for 20 documents on the
Document path, 2,806 for 50 tags on the BulkEditObjectPermissionsView
path, both now flat). Full documents test suite green (2,148 passed,
1 skipped).
2026-08-24 20:04:58 -07:00
stumpylog 9d2416c435 Perf: batch id resolution for TagsField and friends
TagsField/CorrespondentField/DocumentTypeField/StoragePathField were
plain PrimaryKeyRelatedField subclasses with no batching. When used
with many=True (only tags today: DocumentSerializer.tags,
WorkflowActionSerializer.assign_tags), DRF's ManyRelatedField resolves
each submitted id with its own query -- one query per tag on every
PATCH/PUT that sets tags.

Added BatchResolvingPrimaryKeyRelatedField as the shared base for all
four field classes and overrode many_init so the many=True form
(_BatchingManyRelatedField) resolves the whole id list with one
pk__in query, falling back to the child relation's normal per-item
validation for anything not found in that batch. Only TagsField uses
many=True today, but the fix isn't tag-specific -- if a future PR puts
many=True on one of the others, it inherits the same batching instead
of reintroducing this as a new bug.

Independent review caught a real regression: Django's IntegerFieldOverflow
guard (out-of-range int -> EmptyResultSet) only covers exact/gt/gte/lt/lte
lookups, not `in`, so an absurdly large tag id reached the batched
pk__in= query as-is and raised an unhandled OverflowError (SQLite) /
DataError (Postgres) instead of the normal 400 the original per-item
`exact` lookup produced. Guarded the batch query and fall through to
per-item resolution (which goes through the protected `exact` lookup)
on failure.

Verified via CaptureQueriesContext against a real API PATCH: 20 tags
dropped from 54 to 35 queries per request (exactly the 19 saved by
collapsing 20 individual lookups into one batched query). Full
documents/workflows/bulk-edit/retagger/custom-fields suites green
(443 passed).
2026-08-24 15:15:00 -07:00
GitHub Actions 2609327e9c Auto translate strings 2026-08-24 21:44:25 +00:00
shamoonandGitHub b90ccf910f Finally, the remote ocr workflow (#13637)
* Ok! Backend stuff for the remote ocr workflow

* Frotnend workflow stuff

* And docs

* Fix dynamic action fields thing

* Actually, fix the action dropdown thing

* Fix this validation thing, and we have to check existing actions

* Fix migration
2026-08-24 14:43:05 -07:00
shamoonandGitHub c93c996edf Remote ocr reprocess (#13636)
* Backend stuff for remote ocr reprocess, add to bulk edit pass in from ui settings

* Ok, frontend reprocess remote option

* Docs
2026-08-24 14:43:05 -07:00
shamoonandGitHub 7f1609332a Allow parsers to declare uses remote, and remote ocr_mode (#13634)
* uses_remote_service + allow_remote to allow opt-in / out of remote OCR

* Add to parser dev docs

* remote_ocr_mode config setting

* Checks for remote_ocr_mode and fix import

* Update config.component.spec.ts

* More tests for remote_ocr_mode

* Docs for remote_ocr_mode

* Ok, wire up the remote_ocr_mode with allow_remote for consumer

* Update consumer.py

* Format remote OCR mode check tests

* Use get_choice_from_env
2026-08-24 14:43:04 -07:00
stumpylog a98d0669e4 Perf: batch CustomField/Document lookups in modify_custom_fields
modify_custom_fields looped documents x fields, re-.get()-ing the
CustomField queryset per iteration and Document.objects.get() per doc
for DOCUMENTLINK fields -- same shape as the earlier custom_fields
serializer N+1 (#13779), just nested one level deeper. Resolve both
into dicts once up front instead. Also pass the resolved objects
(not bare ids) to update_or_create so newly-created CustomFieldInstance
rows cache their field/document FK, avoiding a re-fetch when auditlog's
post_save receiver calls str(instance) (which touches .field.name).

docs_by_id defers `content` (the one field guaranteed both large and
unused by this function or its receivers) rather than using .only(),
since .only() would just turn the filename-generation signal's other
field access into a deferred-reload N+1.

Verified via CaptureQueriesContext: 6 docs x 4 fields dropped from 48
CustomField queries to 1; DOCUMENTLINK per-doc Document lookups dropped
from N to 0 (single batched query instead).
2026-08-24 14:21:59 -07:00
GitHub Actions a1f20c9fe7 Auto translate strings 2026-08-24 21:19:19 +00:00
shamoonandGitHub 4fd1c60731 Enhancement: support using remote OCR engines selectively (#13633)
* Backend changes and migration for remote OCR Config

* Backend tests

* Frontend stuff, with sections

* Docs

* Update test_tesseract_parser.py

* Actually we cant use this any more, in case settings are in app config

* Dont mark entire test file for db, use a mock for empty engine settings
2026-08-24 14:17:52 -07:00
stumpylog bda506968b Handles a bad client sending malformed JSON or non-int primary keys 2026-08-24 12:35:06 -07:00
a0908f6b4a Enhancement: websocket heartbeat (#13739)
---------

Co-authored-by: shamoon <4887959+shamoon@users.noreply.github.com>
2026-08-24 15:15:21 +00:00
bab9129ff8 Fix: lazy import guardian modules to fix search language setting (#13768)
Co-authored-by: Trenton H <797416+stumpylog@users.noreply.github.com>
2026-08-24 13:46:20 +00:00
Trenton Holmes 4ccb34a70b Perf: avoid per-instance CustomField reload in DocumentMetadataOverrides
send_websocket_document_updated calls document.refresh_from_db()
before building overrides, which drops the custom_fields prefetch
(and its select_related("field")) set up by the view's queryset.
DocumentMetadataOverrides.from_document() then lazily reloads field
once per custom field instance. Since from_document() can't rely on
the caller having a prefetched document, select_related explicitly at
the point of use instead.
2026-08-23 17:33:05 -07:00
Trenton Holmes 0a466c9fcf Perf: reuse resolved CustomField objects across drf-writable-nested's per-item revalidation
drf-writable-nested's update_or_create_reverse_relations rebuilds a
fresh serializer -- and fresh field instances -- per custom_fields item
while matching existing vs. new instances during save(), so the
per-instance lookup cache alone only helped the first validation pass.
It passes the same context dict (by reference) to every one of those
serializers, so stash resolved CustomField objects there instead:
later passes reuse them for free rather than re-querying.
2026-08-23 17:12:25 -07:00
shamoonandGitHub 294328f174 Fix: version indexing fixes (#13737) 2026-08-23 23:04:47 +00:00
Trenton Holmes c9cc4f427d Perf: batch CustomField lookups when validating a document's custom_fields
DocumentSerializer.custom_fields validates each item's field id via a
plain PrimaryKeyRelatedField, which issues one SELECT per custom field
per validation pass (discussion #13690). Batch-resolve all field ids in
one query and cache them on the field instance so per-item validation
is free instead of re-querying.
2026-08-23 14:53:46 -07:00
shamoonandGitHub 0458bad5f2 Fix: append charset to file response for text files (#13759) 2026-08-22 06:15:53 -07:00
shamoonandGitHub 7e4a644714 Fix: align bulk edit perms with document model (#13757) 2026-08-22 05:24:16 -07:00
shamoonandGitHub bed95ea301 Tweakhancement: add jitter to IMAP polling schedule (#13734) 2026-08-20 17:19:47 +00:00
GitHub Actions a424dace43 Auto translate strings 2026-08-19 18:20:04 +00:00
f1c8a72f26 Enhancement: sync OIDC groups to superuser and staff roles (#13060)
Co-authored-by: shamoon <4887959+shamoon@users.noreply.github.com>
Co-authored-by: SoleroTG <github-29h@solero.quietmail.eu>
Co-authored-by: stumpylog <797416+stumpylog@users.noreply.github.com>
2026-08-19 15:27:56 +00:00
GitHub Actions fd3c525f03 Auto translate strings 2026-08-19 14:24:36 +00:00
shamoonandGitHub e389298aab Enhancement: merge documents as versions (#13515) 2026-08-19 07:20:14 -07:00
b17a512539 Refactor: render paperless_ai prompts via Jinja2 templates instead of f-strings (#13698)
* Refactor: render paperless_ai prompts via Jinja2 templates instead of f-strings

* Apply suggestions from code review

Co-authored-by: shamoon <4887959+shamoon@users.noreply.github.com>
2026-08-18 18:32:21 +00:00
shamoonandGitHub 643d0205bc Fix: dont re-render path template when checking collisions (#13718) 2026-08-18 07:19:51 -07:00
shamoonandGitHub f2806179a2 Fix: DocumentClassifierSchema bounds (#13707) 2026-08-17 09:58:35 -07:00
GitHub Actions 31746371f4 Auto translate strings 2026-08-14 22:53:12 +00:00
Trenton HandGitHub 0e5fbc973a Enhancement: prefer existing tags, types, correspondents, and storage paths in AI suggestions (#13676)
AI Suggestions previously invented near-duplicate metadata because the classification
prompt had no knowledge of the installation's own taxonomy. This surfaces
a small, ranked, permission-filtered set of existing tags/document
types/correspondents/storage paths - drawn from the document's RAG
neighbors plus its own already-assigned metadata - so the model prefers
reusing what already exists.

The LLM response schema now returns existing_ids (IDs of reused
candidates) separately from new_names (genuinely new suggestions).
Only new_names goes through localization and fuzzy name-matching;
existing_ids is resolved deterministically and never touched by the
localization pass, so exact matches can no longer be silently
corrupted by translation.
2026-08-14 15:51:34 -07:00
Trenton HandGitHub fe5d09a123 Fix: reopen a fresh Tantivy index per write to prevent orphaned segment files (#13682) 2026-08-14 16:21:15 +00:00
GitHub Actions 01c12d9ea4 Auto translate strings 2026-08-13 19:48:40 +00:00
Max TruxaandGitHub f5c0d118f7 Fix: fix validation of workflow title assignment (#13659) 2026-08-13 12:46:57 -07:00
ff13847d0a Feature: Allow selection of compression type and and level during export (#13661)
* Feature: Allow configuring the compression type and compression levels during export

Building on the zip export improvements, this now allows users to further configure the
zip to fit their needs.  A simple stored zip for speed, or a high compression zstd for
the smallest archive.  Full validation of the method and levels at the command line

Co-authored-by: shamoon <4887959+shamoon@users.noreply.github.com>
2026-08-13 18:29:33 +00:00
GitHub Actions 4789fe9a52 Auto translate strings 2026-08-12 19:04:59 +00:00
shamoonandGitHub e150c8c7c0 Enhancement: customizable icons for saved views (#13388) 2026-08-12 19:03:24 +00:00
JaydenandGitHub 879cd4a30a Enhancement: Add --url argument to document_fuzzy_match to improve output (#13123)
Added a new --url argument to specify the base URL of the Paperless instance, allowing matched documents to be displayed as clickable links. Updated the logic to fetch document titles based on the presence of the base URL.
2026-08-12 08:23:05 -07:00
Trenton HandGitHub 6a02b87dde Feature: Updates remote OCR parser to respect the OCR mode setting (#13408)
* Have the remote parser respect the provided produce_archive_file setting, as already determined via the consumer checks

* Updates the documentation to be correct about the respecting now

* merge conflict fixing
2026-08-11 19:23:46 +00:00
Trenton HandGitHub 59a2651804 Fix: pass document chat queries as a QuerySet instead of a materialized list (#13638)
In tracemalloc based profiling, not materializing the whole Document list
reduced memory to approximately 20% of the baseline, with a peak memory
that scaled with the library size.  Now, the lazt queryset is used and only
the needed pk value is actually contributing to memory
2026-08-11 15:25:08 +00:00
shamoonandGitHub 855669ddf9 Fix: fixes for workflow assign custom field values (#13630) 2026-08-10 07:38:07 -07:00
GitHub Actions 62089df2d8 Auto translate strings 2026-08-10 02:26:58 +00:00
Trenton HandGitHub 5e5f6a88a3 Fix: deny deactivated users in permission filtering and auto-login (#13623)
* Fix: Hardening sweep, ensure a user is active, not just authenticated

* Missed this test
2026-08-10 02:25:06 +00:00
Trenton HandGitHub c28c532bef Fix: check bulk mail delete permissions for the whole batch up front (#13620)
ProcessedMailViewSet.bulk_delete checked permissions inside the delete loop, so an unpermitted id returned 403 only after the mails ahead of it had already been deleted. Resolve the permitted set once via permitted_object_ids and reject before deleting anything, which also drops the per-mail permission queries.
2026-08-09 13:55:49 +00:00
Trenton HandGitHub 17dc482872 Fix: Allow DRF to validate the maximum API key length (#13614) 2026-08-08 19:51:27 +00:00