mirror of
https://github.com/paperless-ngx/paperless-ngx.git
synced 2026-08-03 01:22:17 +00:00
Performance: Tantivy indexing optimization (#13053)
* Tantivy: get permissions by chunks -40% indexing time compared to previous commit * Make progress bar process one by one with chunk -15% indexing time compared to previous commit * Prefetch FK + iterate over chunk from SQL Prefetch additional needed data (note user, custom field content) -20% indexing time compared to previous commit * Reindex: increase Tantivy heap size from 128 to 512MB Gains probably vary depending on the machine, but it seems a sweet spot compatible with low-end hardware. * Reindex: optimization on permission fetching and autocomplete word set -10% indexing time compared to previous commit * Autocomplete analyzer python->rust Splits words with underscore compared to the python analyzer. E.g.: "blue_print" -> ["blue", "print"] It can still be found with the "blue_print" keyword, as the search string is also split in two words. -50% indexing time compared to previous commit (indexing is twice faster!) * Index bigram for CJK content only Inedxing time slightly longer (~3%), but since the non-CJK content is not indexed, bigram searchs will be slightly optimized. * Fix group-based view_document permissions missing from bulk rebuild _bulk_get_viewer_ids only queried UserObjectPermission, dropping the group-permission expansion that get_users_with_perms(with_group_users=True) performs for the non-batched per-document indexing path. A user who could only see a document via group membership would lose search access to it after any full reindex. Also query GroupObjectPermission and expand group membership to user ids, matching the existing single-document behavior. * Yield (document, viewer_ids) pairs from _DocumentViewerStream Previously _DocumentViewerStream.__iter__ yielded plain Document objects while the matching viewer ids were exposed through a separate mutable attribute (viewer_ids_by_pk), overwritten each time the generator crossed a chunk boundary. rebuild() read that attribute out-of-band per document. This only worked because the current iter_wrapper (a plain progress-bar passthrough) happens to consume the stream in strict lock-step with no lookahead. Any wrapper that buffers, batches, or reorders would silently pair a document with the wrong chunk's viewer ids. Yield the pair directly so the association travels with the document regardless of how iter_wrapper consumes the stream, and drop the now-unneeded viewer_ids_by_pk attribute. * Add --heap-size-mb CLI arg to document_index reindex writer_heap_bytes was hardcoded at 512MB with no way to tune it. Expose it as a manual-rebuild-only CLI arg rather than a settings/env var, per review feedback, so lower-memory hosts can reduce it without a wider config surface. Defaults to unset so TantivyBackend.rebuild's own default stays the single source of truth. --------- Co-authored-by: stumpylog <797416+stumpylog@users.noreply.github.com>
This commit is contained in:
co-authored by
stumpylog
parent
71557d7c64
commit
df1ddb15cc
@@ -45,6 +45,18 @@ class Command(PaperlessCommand):
|
||||
"Safe to run on every startup or upgrade."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--heap-size-mb",
|
||||
type=int,
|
||||
default=None,
|
||||
help=(
|
||||
"Tantivy writer memory budget in MB for a full reindex "
|
||||
"(split across writer threads). Defaults to 512MB; lower "
|
||||
"this on memory-constrained hosts. Larger values buffer "
|
||||
"more documents before flushing a segment, deferring "
|
||||
"merge work rather than avoiding it."
|
||||
),
|
||||
)
|
||||
|
||||
def handle(self, *args, **options):
|
||||
with transaction.atomic():
|
||||
@@ -60,13 +72,26 @@ class Command(PaperlessCommand):
|
||||
"document_type",
|
||||
"storage_path",
|
||||
"owner",
|
||||
).prefetch_related("tags", "notes", "custom_fields", "versions")
|
||||
).prefetch_related(
|
||||
"tags",
|
||||
"notes__user",
|
||||
"custom_fields__field",
|
||||
"versions",
|
||||
)
|
||||
total = documents.count()
|
||||
rebuild_kwargs = {}
|
||||
if options.get("heap_size_mb") is not None:
|
||||
rebuild_kwargs["writer_heap_bytes"] = (
|
||||
options["heap_size_mb"] * 1_000_000
|
||||
)
|
||||
get_backend().rebuild(
|
||||
documents,
|
||||
iter_wrapper=lambda docs: self.track(
|
||||
docs,
|
||||
iter_wrapper=lambda pairs: self.track(
|
||||
pairs,
|
||||
description="Indexing documents...",
|
||||
total=total,
|
||||
),
|
||||
**rebuild_kwargs,
|
||||
)
|
||||
reset_backend()
|
||||
|
||||
|
||||
Reference in New Issue
Block a user