Compare commits

..
Author SHA1 Message Date
stumpylogandClaude Sonnet 5 cfb9125d9e chore: remove benchmark-commands spec/plan docs now that work is complete
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-30 13:43:01 -07:00
stumpylogandClaude Sonnet 5 fa2d7f660b Fix: apply final-review fix wave (safety guard, date spread, MariaDB explain, history fields, memory footprint, minors)
Final whole-branch review before this tooling branch settles: add a
required --yes-i-know-this-wipes-the-database confirmation flag ahead of
--reset (it deletes ALL users/groups/documents in the target DB, not just
benchmark-created ones); spread seeded documents' created dates over a
3-year window instead of leaving them all on one default date; fix
capture_explain() to use MariaDB's ANALYZE syntax instead of Postgres-only
EXPLAIN ANALYZE (verified against real MariaDB 12.3 -- the old code was a
silent 1064 syntax error); record db_vendor/document_count in run/profile
history entries; shrink SeededData's memory footprint at scale by
returning counts instead of full ORM instance tuples; and a handful of
minor fixes (storage_path assignment, --explain warning instead of silent
no-op, harness.py repeat<1 guard, type annotations). All changes verified
against real Postgres and MariaDB containers on the VM, not just SQLite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-30 13:33:28 -07:00
stumpylogandClaude Sonnet 5 2712c3fe48 Fix: correct /api/tags/ endpoint URL in benchmarking skill
The skill documentation stated that run benchmarks /api/tags/ but
src/paperless_benchmark/endpoints.py actually uses
/api/tags/?page_size=100000. Update the skill's Command reference
section to match the actual endpoint definition.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-30 11:20:14 -07:00
stumpylog 684c4ef392 docs(benchmark): add paperless-benchmarking skill 2026-07-30 11:07:33 -07:00
stumpylog 410e440bd6 fix(benchmark): fully clear users/groups on reset, grant perf_target model perms 2026-07-30 10:53:02 -07:00
stumpylogandClaude Sonnet 5 ef9ee435e8 docs: add Task 10 (quirk fixes) and Task 11 (skill) to benchmark plan
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-30 10:44:34 -07:00
stumpylog 5d8287f120 fix(benchmark): reuse seeded data in profile instead of reseeding
_handle_profile unconditionally called seed_benchmark_dataset every
run, colliding with an IntegrityError against data seeded by a prior
seed/profile call. Look up the existing perf_target user instead,
matching _handle_run's already-correct pattern, and raise a
CommandError pointing at `seed` if no dataset exists.

Scenario.run/queryset_for_explain now take a User directly instead of
a SeededData, since only data.users[0] (perf_target) was ever used.

profile no longer seeds, so it drops the --tier dependency: the
printed line and history entry no longer include a tier/(tier=...)
suffix, and --tier's help text now only mentions `seed`.
2026-07-30 10:15:12 -07:00
stumpylog acfdafbff9 feat(benchmark): add manage.py benchmark command (seed/run/profile/list-scenarios) 2026-07-30 09:52:21 -07:00
stumpylog cd02cfe02c feat(benchmark): add pluggable profile-scenario registry 2026-07-30 08:31:14 -07:00
stumpylog ff9cde7f9c feat(benchmark): add timed API endpoint scenarios 2026-07-30 07:49:05 -07:00
stumpylogandClaude Sonnet 5 259731074c docs: add implementation plan for benchmark management commands
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-30 07:16:32 -07:00
stumpylog 02aaf9b3a2 feat(benchmark): add local JSONL results history log 2026-07-29 14:58:11 -07:00
stumpylog 8a34385199 feat(benchmark): add tiered dataset seeding 2026-07-29 14:17:52 -07:00
stumpylog 3e8656a3ed feat(benchmark): add per-backend reset and explain helpers 2026-07-29 13:27:39 -07:00
stumpylog 8ec322f9aa feat(benchmark): add run_profile timing harness 2026-07-29 12:29:53 -07:00
stumpylogandClaude Sonnet 5 22e4be1548 feat(benchmark): scaffold paperless_benchmark app
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 12:11:53 -07:00
stumpylogandClaude Sonnet 5 17e0b70da5 docs: add design spec for benchmark management commands
Merges the standalone run_benchmarks.py/seed_benchmark_data.py scripts
and the tools/profiling branch's Postgres-only harness into one
cross-backend `manage.py benchmark` command suite.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-29 11:15:28 -07:00
shamoonandGitHub 668fa77428 Fix: prevent scale vs page loop in pngx PDF viewer (#13406) 2026-07-29 09:25:37 -07:00
shamoonandGitHub 5bd72014a6 Fix: handle hidden line breaks in email subjects when sending email (#13402) 2026-07-29 14:35:30 +00:00
Trenton HandGitHub bbb9c86ba4 Fix: avoid NotSupportedError from document_importer on MariaDB (#13400) 2026-07-29 07:14:13 -07:00
26 changed files with 1360 additions and 226 deletions
@@ -0,0 +1,187 @@
---
name: paperless-benchmarking
description: Use when profiling paperless-ngx performance, running `manage.py benchmark`, investigating a slow query or endpoint, or deciding whether a profiling finding should become a permanent registered scenario. Covers command reference, the fork/merge-back branch workflow for perf investigations, and how to read query-plan output.
---
# Paperless-ngx Benchmarking
This repo has a built-in benchmarking/profiling tool: `manage.py benchmark`, in
the `paperless_benchmark` Django app. It replaces ad hoc standalone scripts —
use it instead of writing new one-off seed/timing scripts.
## Command reference
```
manage.py benchmark seed --tier {home,medium,large} [--reset --yes-i-know-this-wipes-the-database] [--seed N]
manage.py benchmark run --repeat 5 [--label baseline]
manage.py benchmark profile <scenario_name> [--repeat 5] [--explain]
manage.py benchmark list-scenarios
```
- **`seed`** builds a realistic dataset at one of three scales: `home` (500
documents — fast, use this for iteration), `medium` (20,000 — the default
when `--tier` is omitted; a multi-minute seed), `large` (360,000 — matches
the scale reported in real large-install bug reports; slow, only use it
when a finding needs confirming at real scale). `--reset` wipes any
previously-seeded benchmark data first — but it is destructive and
irreversible: it deletes **all** users, **all** groups, and **all**
documents/tags/correspondents/document types/storage paths in the target
database, not just benchmark-created rows. Only run it against a disposable
benchmark database, never a real install. Because of that, `--reset` also
requires passing `--yes-i-know-this-wipes-the-database` in the same
invocation, or the command raises an error and does nothing. Omit both
flags if you want to layer more data onto an existing seed instead. `seed`
creates two named users, `perf_target` (mixed owned/shared documents,
realistic guardian permission grants) and `perf_admin` (superuser), plus a
general user/group pool with realistic permission-row ratios.
- **`run`** times the 3 built-in API endpoint benchmarks
(`/api/documents/`, `/api/documents/?page_size=50`, `/api/tags/?page_size=100000`) for both
`perf_target` and `perf_admin`, reporting min/median/max wall-clock and SQL
query count. Requires `seed` to have already run — it reuses that data, it
does not seed its own.
- **`profile`** times one named scenario from the registry (see
`list-scenarios`) via best-of-N repeat timing and SQL query count, against
`perf_target`. `--explain` additionally captures and prints the query plan:
real `EXPLAIN ANALYZE` execution stats on PostgreSQL/MariaDB, or
`EXPLAIN QUERY PLAN` (plan only, no real timing/row counts — clearly labeled
as such) on SQLite. Like `run`, `profile` requires `seed` to have already
run — it does not seed its own data either.
- Every `run`/`profile` invocation appends a JSON line to
`benchmark_results/history.jsonl` at the repo root (local-only, gitignored —
never commit this file). Use it to compare a `before`/`after` pair across
two invocations without hand-copying numbers.
- Full chain example: `seed --reset` once, then `run` and `profile` as many
times as you want against that same seeded data — no need to reseed between
them.
## Adding a new scenario
A "scenario" is a named, registered query/operation that `profile` can time
and explain. To add one, edit `src/paperless_benchmark/scenarios.py`: write a
`_<name>_run(user)` function (returns whatever `run_profile` should time) and
optionally a `_<name>_queryset(user)` function (returns the `QuerySet` for
`--explain` to analyze), then `register(Scenario(name=..., describe=...,
run=..., queryset_for_explain=...))` at module level. Both functions receive
the already-seeded `perf_target` user — they should not seed their own data.
## Branch workflow
`tools/benchmark-management-commands` is a **long-lived tooling branch**, not
a feature branch that gets merged and closed:
1. It is periodically brought up to date with `dev` (merge `dev` into it) so
the tooling doesn't drift from the schema/codebase it profiles. Do this
before starting a new investigation if it's been a while since the last
sync.
2. **Every performance investigation forks its own branch from
`tools/benchmark-management-commands`** (not from `dev`). Do the
investigation there: write throwaway profiling code, try fixes, capture
before/after numbers.
3. **That investigation branch never merges into `dev` or production.** Its
only job is to produce evidence and, optionally, a reusable scenario.
4. If the investigation turns up a scenario worth keeping permanently (see
"When to graduate a scenario" below), open a PR that adds **just that
scenario** back into `tools/benchmark-management-commands` — not the rest
of the investigation branch's throwaway code.
5. Any actual production fix the investigation motivates (e.g. an ORM query
change) goes into its own normal feature branch off `dev`, following the
project's regular contribution process — profiling evidence informs that
PR's description, but the profiling code itself does not travel with it.
## Reading query-plan output
- **PostgreSQL** `EXPLAIN ANALYZE`: look for `Seq Scan` on a large table
(missing index), a large gap between `rows=N` (planner's estimate) and the
actual row count in parentheses (stale statistics or a bad cardinality
estimate), and nested-loop joins driven by an outer relation with many
rows (usually the N+1 pattern this tool exists to catch).
- **MariaDB**: verified against a real MariaDB 12.3 container that MariaDB
does NOT accept MySQL 8.0.18+'s `EXPLAIN ANALYZE` syntax (it's a 1064
syntax error) -- `capture_explain()` instead runs MariaDB's own
`ANALYZE <statement>` form (no `EXPLAIN` keyword), which returns a
tabular plan with real per-row execution columns: `rows` (estimate) vs.
`r_rows` (actual), and `filtered` vs. `r_filtered`. A large gap between
`rows` and `r_rows`, or `type: ALL` (full table scan) on a large table,
are the signals to look for -- the same underlying concerns as Postgres's
`Seq Scan`/estimate-vs-actual gap, just in MariaDB's column-based output
instead of Postgres's nested-tree text format.
- **SQLite** `EXPLAIN QUERY PLAN`: no real timing/row-count data, only the
chosen access path (`SCAN` vs `SEARCH`, which index if any). Useful for
confirming an index is even being considered, not for judging real-world
cost — corroborate any SQLite finding against Postgres/MariaDB before
trusting it, since planner behavior differs meaningfully between them.
- Compare query **count**, not just timing, between before/after: a fix that
keeps the same wall-clock time but drops query count from O(n) to O(1) is
still a real, durable improvement — timing alone is noisy and
environment-dependent, query count is not.
## Cleaning up after an interrupted run
If a `seed`/`run`/`profile` invocation gets killed mid-run (Ctrl-C, `kill -9`,
a timed-out SSH session, etc.), check whether it left anything behind before
trusting the next benchmark's numbers. This was verified for real: a
`benchmark seed --tier large --reset ...` was started against both a fresh
PostgreSQL 18 container and a fresh MariaDB 12.3 container and `kill -9`'d a
few seconds into document seeding. In both cases, the database-side
connection disappeared immediately -- no stuck backend, no lingering query,
no held lock was observed in either backend once the killed process's PID
was confirmed gone. That said, this was one interruption point (mid
bulk-seed, between chunks); a run killed mid-query, or a driver/network
hiccup that doesn't cleanly close the socket, could behave differently, so
still check before trusting a number if any run in the session was
interrupted:
- **PostgreSQL**: look for leftover connections against the benchmark
database:
```sql
SELECT pid, state, query, query_start
FROM pg_stat_activity
WHERE datname = current_database() AND pid <> pg_backend_pid();
```
If a stuck backend shows up, clear it with:
```sql
SELECT pg_terminate_backend(pid)
FROM pg_stat_activity
WHERE datname = current_database() AND pid <> pg_backend_pid();
```
- **MariaDB**: look for leftover connections/queries:
```sql
SHOW FULL PROCESSLIST;
```
If a stuck connection shows up (anything other than your current admin
session), clear it with:
```sql
KILL <id>;
```
A stray connection left running concurrently with a subsequent benchmark run
would add real, contaminating load (extra queries competing for the same
rows, possibly held locks slowing the next run's timings) -- cheap to rule
out, expensive to silently trust a number that was actually measured
alongside a zombie connection.
## When to graduate a one-off finding into a permanent scenario
Register a scenario (rather than leaving it as throwaway code on the
investigation branch) when **both** are true:
- The query pattern is one this codebase is likely to regress on again (e.g.
it involves a permission-check join, a bulk operation, or anything else
with an easy-to-reintroduce N+1) — not a one-time fluke specific to this
investigation.
- Re-running it later, against a fresh seed, would still produce a
meaningful signal (it doesn't depend on investigation-specific throwaway
data or a fix that's already permanently landed and can't regress the same
way).
If a finding doesn't meet both bars, keep it as disposable code on the
investigation branch and let the branch's evidence (captured in the PR
description of whatever production fix it motivates) be the permanent
record instead.
+3
View File
@@ -115,3 +115,6 @@ celerybeat-schedule*
# Git worktree local folder
.worktrees
# Benchmark tooling output (local only, never committed)
/benchmark_results/
+4 -5
View File
@@ -948,11 +948,10 @@ for display in the web interface.
!!! note
The **remote OCR parser** (Azure AI) also honors this setting: when
no archive is requested (`never`, or `auto` with a born-digital PDF),
the remote engine is skipped entirely and locally-extracted text is
used instead, avoiding an unnecessary API call and a duplicate text
layer.
The **remote OCR parser** (Azure AI) always produces a searchable
PDF and stores it as the archive copy, regardless of this setting.
`ARCHIVE_FILE_GENERATION=never` has no effect when the remote
parser handles a document.
#### [`PAPERLESS_OCR_CLEAN=<mode>`](#PAPERLESS_OCR_CLEAN) {#PAPERLESS_OCR_CLEAN}
+4 -5
View File
@@ -187,11 +187,10 @@ PAPERLESS_ARCHIVE_FILE_GENERATION=auto
### Remote OCR parser
If you use the **remote OCR parser** (Azure AI), `ARCHIVE_FILE_GENERATION` is
honored the same way as for the local engine: when no archive is requested
(`never`, or `auto` with a born-digital PDF), the remote engine is skipped
entirely and locally-extracted text is used instead, avoiding an unnecessary
API call and a duplicate text layer.
If you use the **remote OCR parser** (Azure AI), note that it always produces a
searchable PDF and stores it as the archive copy. `ARCHIVE_FILE_GENERATION=never`
has no effect for documents handled by the remote parser - the archive is produced
unconditionally by the remote engine.
## Search Index (Whoosh -> Tantivy)
@@ -129,13 +129,25 @@ describe('PngxPdfViewerComponent', () => {
;(component as any).applyScale()
expect(viewer.currentScaleValue).toBe(PdfZoomScale.PageFit)
expect(viewer.currentScale).toBe(2)
})
it('does not reapply scale for page-only changes', async () => {
await initComponent()
const pdf = (component as any).pdf as { numPages: number }
pdf.numPages = 3
const viewer = (component as any).pdfViewer as PDFViewer
viewer.setDocument(pdf)
const applyScaleSpy = jest.spyOn(component as any, 'applyScale')
component.page = 2
;(component as any).lastViewerPage = 2
;(component as any).applyViewerState()
component.ngOnChanges({
page: new SimpleChange(1, 2, false),
})
expect(viewer.currentPageNumber).toBe(2)
expect((component as any).lastViewerPage).toBeUndefined()
expect(applyScaleSpy).toHaveBeenCalled()
expect(applyScaleSpy).not.toHaveBeenCalled()
})
it('does not reset the viewer when it is already on the requested page', async () => {
@@ -116,7 +116,10 @@ export class PngxPdfViewerComponent
changes['zoomScale'] ||
changes['rotation']
) {
this.applyViewerState()
// Prevent loop with page / scale application see https://github.com/paperless-ngx/paperless-ngx/issues/13404
this.applyViewerState(
!!(changes['zoom'] || changes['zoomScale'] || changes['rotation'])
)
}
if (changes['searchQuery']) {
@@ -240,7 +243,7 @@ export class PngxPdfViewerComponent
}
}
private applyViewerState(): void {
private applyViewerState(applyScale = true): void {
if (!this.pdfViewer) {
return
}
@@ -264,7 +267,7 @@ export class PngxPdfViewerComponent
if (this.page === this.lastViewerPage) {
this.lastViewerPage = undefined
}
if (hasPages) {
if (hasPages && applyScale) {
this.applyScale()
}
this.dispatchFindIfReady()
+3
View File
@@ -36,6 +36,9 @@ def send_email(
TODO: re-evaluate this pending https://code.djangoproject.com/ticket/35581 / https://github.com/django/django/pull/18966
"""
if "\r" in subject or "\n" in subject:
subject = " ".join(line.strip(" \t") for line in subject.splitlines())
email = EmailMessage(
subject=subject,
body=body,
@@ -386,10 +386,19 @@ class Command(CryptMixin, PaperlessCommand):
raise DeserializationError(
f"{model.__name__} has no updatable fields; PK-only models are not supported by the importer",
)
# MySQL/MariaDB support upserts via ON DUPLICATE KEY UPDATE but,
# unlike PostgreSQL/SQLite, cannot target a specific unique field
# for the conflict -- passing unique_fields there raises
# NotSupportedError.
unique_fields = (
[model._meta.pk.attname]
if connection.features.supports_update_conflicts_with_target
else None
)
model.objects.bulk_create( # type: ignore[attr-defined]
instances,
update_conflicts=True,
unique_fields=[model._meta.pk.attname],
unique_fields=unique_fields,
update_fields=update_fields,
)
loaded_models.add(model)
+1 -1
View File
@@ -75,7 +75,7 @@ class TestEmail(DirectoriesMixin, SampleDirMixin, APITestCase):
{
"documents": [self.doc1.pk, self.doc2.pk],
"addresses": "hello@paperless-ngx.com,test@example.com",
"subject": "Bulk email test",
"subject": "Bulk email\n test",
"message": "Here are your documents",
},
),
+7 -33
View File
@@ -3,9 +3,7 @@ Built-in remote-OCR document parser.
Handles documents by sending them to a configured remote OCR engine
(currently Azure AI Vision / Document Intelligence) and retrieving both
the extracted text and a searchable PDF with an embedded text layer. For
born-digital PDFs that need no archive copy, the remote call is skipped
entirely in favor of locally-extracted text (see ``RemoteDocumentParser.parse``).
the extracted text and a searchable PDF with an embedded text layer.
When no engine is configured, ``score()`` returns ``None`` so the parser
is effectively invisible to the registry — the tesseract parser handles
@@ -23,8 +21,6 @@ from typing import Self
from django.conf import settings
from paperless.parsers.utils import extract_pdf_text
from paperless.parsers.utils import post_process_text
from paperless.version import __full_version_str__
if TYPE_CHECKING:
@@ -73,11 +69,8 @@ class RemoteDocumentParser:
"""Parse documents via a remote OCR API (currently Azure AI Vision).
This parser sends documents to a remote engine that returns both
extracted text and a searchable PDF with an embedded text layer,
except when ``parse()`` is called with ``produce_archive=False`` for
a PDF, in which case the remote call is skipped and only locally
extracted text is returned (no archive). It does not depend on
Tesseract or ocrmypdf.
extracted text and a searchable PDF with an embedded text layer.
It does not depend on Tesseract or ocrmypdf.
Class attributes
----------------
@@ -166,11 +159,8 @@ class RemoteDocumentParser:
Returns
-------
bool
Always True — the remote engine is capable of returning a PDF
with an embedded text layer to serve as the archive copy.
Whether it actually does so for a given document depends on
``produce_archive`` passed to :meth:`parse` (see there for when
the remote engine call, and thus archive generation, is skipped).
Always True — the remote engine always returns a PDF with an
embedded text layer that serves as the archive copy.
"""
return True
@@ -227,12 +217,6 @@ class RemoteDocumentParser:
) -> None:
"""Send the document to the remote engine and store results.
When *produce_archive* is False for a PDF, the caller (via
``documents.consumer.should_produce_archive``) has already determined
that the document is born-digital and needs no archive — skip the
remote engine entirely rather than re-OCRing it and creating a
duplicate text layer.
Parameters
----------
document_path:
@@ -240,8 +224,8 @@ class RemoteDocumentParser:
mime_type:
Detected MIME type of the document.
produce_archive:
Whether an archive copy is wanted. For PDFs, False skips the
remote engine and uses locally-extracted text instead.
Ignored — the remote engine always returns a searchable PDF,
which is stored as the archive copy regardless of this flag.
"""
config = RemoteEngineConfig(
engine=settings.REMOTE_OCR_ENGINE,
@@ -256,16 +240,6 @@ class RemoteDocumentParser:
self._text = ""
return
if not produce_archive and mime_type == "application/pdf":
logger.debug(
"Remote OCR: skipped — no archive requested, "
"using locally-extracted text",
)
self._text = (
post_process_text(extract_pdf_text(document_path, log=logger)) or ""
)
return
if config.engine == "azureai":
self._text = self._azure_ai_vision_parse(document_path, config)
+21 -6
View File
@@ -3,6 +3,7 @@ from __future__ import annotations
import importlib.resources
import logging
import os
import re
import shutil
import tempfile
from pathlib import Path
@@ -24,9 +25,9 @@ from paperless.config import OcrConfig
from paperless.models import CleanChoices
from paperless.models import ModeChoices
from paperless.models import OutputTypeChoices
from paperless.parsers.utils import PDF_TEXT_MIN_LENGTH
from paperless.parsers.utils import extract_pdf_text
from paperless.parsers.utils import pdf_has_digital_text
from paperless.parsers.utils import post_process_text
from paperless.parsers.utils import is_tagged_pdf
from paperless.parsers.utils import read_file_handle_unicode_errors
from paperless.version import __full_version_str__
@@ -509,10 +510,10 @@ class RasterisedDocumentParser:
if mime_type == "application/pdf":
text_original = self.extract_text(None, document_path)
original_has_text = pdf_has_digital_text(
document_path,
text_original,
log=self.log,
has_text = text_original is not None and len(text_original) > 0
original_has_text = has_text and (
is_tagged_pdf(document_path, log=self.log)
or len(text_original) > PDF_TEXT_MIN_LENGTH
)
else:
text_original = None
@@ -657,3 +658,17 @@ class RasterisedDocumentParser:
f"No text was found in {document_path}, the content will be empty.",
)
self.text = ""
def post_process_text(text: str | None) -> str | None:
if not text:
return None
collapsed_spaces = re.sub(r"([^\S\r\n]+)", " ", text)
no_leading_whitespace = re.sub(r"([\n\r]+)([^\S\n\r]+)", "\\1", collapsed_spaces)
no_trailing_whitespace = re.sub(r"([^\S\n\r]+)$", "", no_leading_whitespace)
# TODO: this needs a rework
# replace \0 prevents issues with saving to postgres.
# text may contain \0 when this character is present in PDF files.
return no_trailing_whitespace.strip().replace("\0", " ")
-57
View File
@@ -65,63 +65,6 @@ def is_tagged_pdf(
return False
def pdf_has_digital_text(
path: Path,
text: str | None,
log: logging.Logger | None = None,
) -> bool:
"""Return True if a PDF already has a usable, born-digital text layer.
Combines the tagged-PDF check with an extracted-text length check.
Shared by the tesseract and remote OCR parsers to decide whether
OCR_MODE=auto/off should skip (re-)OCRing a document.
Parameters
----------
path:
Absolute path to the PDF file.
text:
Text already extracted from the PDF (e.g. via ``extract_pdf_text``),
or ``None``.
log:
Logger for warnings. Falls back to the module-level logger when omitted.
Returns
-------
bool
``True`` when the document already contains a text layer.
"""
return is_tagged_pdf(path, log=log) or (
text is not None and len(text) > PDF_TEXT_MIN_LENGTH
)
def post_process_text(text: str | None) -> str | None:
"""Normalise whitespace in extracted OCR/PDF text and strip NUL bytes.
Parameters
----------
text:
Raw extracted text, or ``None``.
Returns
-------
str | None
Cleaned text, or ``None`` when *text* is falsy.
"""
if not text:
return None
collapsed_spaces = re.sub(r"([^\S\r\n]+)", " ", text)
no_leading_whitespace = re.sub(r"([\n\r]+)([^\S\n\r]+)", "\\1", collapsed_spaces)
no_trailing_whitespace = re.sub(r"([^\S\n\r]+)$", "", no_leading_whitespace)
# TODO: this needs a rework
# replace \0 prevents issues with saving to postgres.
# text may contain \0 when this character is present in PDF files.
return no_trailing_whitespace.strip().replace("\0", " ")
def extract_pdf_text(
path: Path,
log: logging.Logger | None = None,
+1
View File
@@ -150,6 +150,7 @@ INSTALLED_APPS = [
"drf_spectacular",
"drf_spectacular_sidecar",
"treenode",
"paperless_benchmark.apps.PaperlessBenchmarkConfig",
*env_apps,
]
@@ -336,117 +336,6 @@ class TestRemoteParserParse:
assert remote_parser.get_date() is None
# ---------------------------------------------------------------------------
# parse() — produce_archive=False skips the remote engine (PDFs only)
# ---------------------------------------------------------------------------
class TestRemoteParserSkipsWhenNoArchiveWanted:
"""When the caller has already decided no archive is needed for a PDF
(documents.consumer.should_produce_archive), the remote engine call is
skipped entirely in favor of locally-extracted text.
"""
def test_pdf_skips_azure_when_no_archive_requested(
self,
remote_parser: RemoteDocumentParser,
simple_digital_pdf_file: Path,
azure_client: Mock,
) -> None:
"""
GIVEN: produce_archive=False for a PDF
WHEN: parse() is called
THEN: Azure is never invoked, no archive is produced, and text
comes from local pdftotext extraction
"""
remote_parser.parse(
simple_digital_pdf_file,
"application/pdf",
produce_archive=False,
)
azure_client.begin_analyze_document.assert_not_called()
assert remote_parser.get_archive_path() is None
assert remote_parser.get_text() != ""
def test_pdf_no_archive_requested_text_matches_local_extraction(
self,
remote_parser: RemoteDocumentParser,
simple_digital_pdf_file: Path,
azure_client: Mock,
mocker: MockerFixture,
) -> None:
"""
GIVEN: produce_archive=False for a PDF
WHEN: parse() is called
THEN: the returned text is exactly the locally-extracted text,
not anything from the (unused) Azure mock
"""
mocker.patch(
"paperless.parsers.remote.extract_pdf_text",
return_value="Local digital text.",
)
remote_parser.parse(
simple_digital_pdf_file,
"application/pdf",
produce_archive=False,
)
assert remote_parser.get_text() == "Local digital text."
def test_pdf_no_archive_requested_closes_no_client(
self,
remote_parser: RemoteDocumentParser,
simple_digital_pdf_file: Path,
azure_client: Mock,
) -> None:
remote_parser.parse(
simple_digital_pdf_file,
"application/pdf",
produce_archive=False,
)
azure_client.close.assert_not_called()
def test_non_pdf_still_calls_azure_when_no_archive_requested(
self,
remote_parser: RemoteDocumentParser,
simple_digital_pdf_file: Path,
azure_client: Mock,
) -> None:
"""
Images have no local-text fallback, so produce_archive=False does
not skip the remote engine for non-PDF MIME types.
"""
remote_parser.parse(
simple_digital_pdf_file,
"image/png",
produce_archive=False,
)
azure_client.begin_analyze_document.assert_called_once()
assert remote_parser.get_text() == _DEFAULT_TEXT
@pytest.mark.usefixtures("no_engine_settings")
def test_unconfigured_engine_takes_precedence_over_skip(
self,
remote_parser: RemoteDocumentParser,
simple_digital_pdf_file: Path,
) -> None:
"""An unconfigured engine still short-circuits before the
produce_archive check, returning empty text as before.
"""
remote_parser.parse(
simple_digital_pdf_file,
"application/pdf",
produce_archive=False,
)
assert remote_parser.get_text() == ""
assert remote_parser.get_archive_path() is None
# ---------------------------------------------------------------------------
# parse() — Azure failure path
# ---------------------------------------------------------------------------
@@ -21,7 +21,7 @@ from documents.parsers import run_convert
from paperless.models import ModeChoices
from paperless.parsers import ParserProtocol
from paperless.parsers.tesseract import RasterisedDocumentParser
from paperless.parsers.utils import post_process_text
from paperless.parsers.tesseract import post_process_text
if TYPE_CHECKING:
from pathlib import Path
View File
+8
View File
@@ -0,0 +1,8 @@
from django.apps import AppConfig
from django.utils.translation import gettext_lazy as _
class PaperlessBenchmarkConfig(AppConfig):
name = "paperless_benchmark"
verbose_name = _("Paperless benchmark")
+136
View File
@@ -0,0 +1,136 @@
from __future__ import annotations
from typing import TYPE_CHECKING
from django.db import connection
if TYPE_CHECKING:
from django.db.models import QuerySet
def _reset_table_names() -> list[str]:
from guardian.models import GroupObjectPermission
from guardian.models import UserObjectPermission
from documents.models import Correspondent
from documents.models import Document
from documents.models import DocumentType
from documents.models import StoragePath
from documents.models import Tag
return [
Document.tags.through._meta.db_table,
Document._meta.db_table,
Tag._meta.db_table,
Correspondent._meta.db_table,
DocumentType._meta.db_table,
StoragePath._meta.db_table,
UserObjectPermission._meta.db_table,
GroupObjectPermission._meta.db_table,
]
def _delete_all_users_and_groups() -> None:
# ASSUMPTION: this tool assumes a disposable benchmark database, never
# point it at a real install. This deletes EVERY user and group in the
# database (not just benchmark-created ones) -- there is no way to
# distinguish "real" users from seeded ones, so this is only safe against
# a database that exists solely to run this benchmarking tool. The
# `benchmark seed --reset` CLI path requires an explicit
# `--yes-i-know-this-wipes-the-database` flag before reaching here; do
# not remove that guard.
from django.contrib.auth import get_user_model
from django.contrib.auth.models import Group
get_user_model().objects.all().delete()
Group.objects.all().delete()
def _reset_postgresql() -> None:
tables = _reset_table_names()
with connection.cursor() as cursor:
cursor.execute(f"TRUNCATE TABLE {', '.join(tables)} RESTART IDENTITY CASCADE;")
_delete_all_users_and_groups()
def _reset_mariadb() -> None:
# MariaDB's TRUNCATE has no CASCADE clause and refuses to truncate a
# table referenced by a foreign key while checks are enabled, so
# checks are disabled for the duration of the reset.
tables = _reset_table_names()
with connection.cursor() as cursor:
cursor.execute("SET FOREIGN_KEY_CHECKS = 0;")
try:
for table in tables:
cursor.execute(f"TRUNCATE TABLE {table};")
finally:
cursor.execute("SET FOREIGN_KEY_CHECKS = 1;")
_delete_all_users_and_groups()
def _reset_sqlite() -> None:
from documents.models import Correspondent
from documents.models import Document
from documents.models import DocumentType
from documents.models import StoragePath
from documents.models import Tag
Document.global_objects.all().delete()
Tag.objects.all().delete()
Correspondent.objects.all().delete()
DocumentType.objects.all().delete()
StoragePath.objects.all().delete()
_delete_all_users_and_groups()
def reset_benchmark_data() -> None:
"""
Remove all previously-seeded benchmark data (documents, tags,
correspondents, document types, storage paths, guardian permission
rows, users, and groups) so a fresh `benchmark seed` run starts
from an empty slate. Dispatches per-backend because TRUNCATE syntax
and cascade behavior differ across the 3 supported databases.
"""
if connection.vendor == "postgresql":
_reset_postgresql()
elif connection.vendor == "mysql":
# MariaDB also reports vendor == "mysql" under Django's mysql backend.
_reset_mariadb()
else:
_reset_sqlite()
def capture_explain(queryset: QuerySet) -> str:
"""
Return the query plan for `queryset` using the current backend's
explain facility. PostgreSQL supports `EXPLAIN ANALYZE {sql}` (real
execution stats). MariaDB does NOT accept that syntax -- verified
against a real MariaDB 12.3 container: `EXPLAIN ANALYZE {sql}` raises a
1064 syntax error, while MariaDB's own `ANALYZE {sql}` form (no
`EXPLAIN` keyword) works and returns real per-row execution stats
(`r_rows`, `r_filtered`, etc. columns) -- this is MariaDB's
EXPLAIN-ANALYZE-equivalent, distinct from MySQL 8.0.18+'s
`EXPLAIN ANALYZE` syntax, which MariaDB does not implement. SQLite only
supports EXPLAIN QUERY PLAN (the chosen plan, not real timing/row
counts) -- that output is clearly labeled rather than silently looking
equivalent to the other two backends' output.
"""
sql, params = queryset.query.sql_with_params()
with connection.cursor() as cursor:
if connection.vendor == "postgresql":
cursor.execute(f"EXPLAIN ANALYZE {sql}", params)
return "\n".join(str(row[0]) for row in cursor.fetchall())
if connection.vendor == "mysql":
# MariaDB also reports vendor == "mysql" under Django's mysql
# backend. Unlike MySQL 8.0.18+, MariaDB has no `EXPLAIN
# ANALYZE` syntax -- its equivalent is `ANALYZE <statement>`.
cursor.execute(f"ANALYZE {sql}", params)
columns = [c[0] for c in cursor.description]
header = " | ".join(columns)
rows = "\n".join(
" | ".join(str(c) for c in row) for row in cursor.fetchall()
)
return f"{header}\n{rows}"
cursor.execute(f"EXPLAIN QUERY PLAN {sql}", params)
rows = "\n".join(" | ".join(str(c) for c in row) for row in cursor.fetchall())
return f"(plan only -- no execution stats on SQLite)\n{rows}"
+81
View File
@@ -0,0 +1,81 @@
# src/paperless_benchmark/endpoints.py
from __future__ import annotations
import statistics
import time
from dataclasses import dataclass
from typing import TYPE_CHECKING
if TYPE_CHECKING:
from django.contrib.auth.models import User
from rest_framework.test import APIClient
ENDPOINTS: tuple[tuple[str, str], ...] = (
("documents_default", "/api/documents/"),
("documents_page50", "/api/documents/?page_size=50"),
("tags_all", "/api/tags/?page_size=100000"),
)
@dataclass(frozen=True, slots=True)
class EndpointTiming:
user_label: str
endpoint_name: str
query_count: int
min_ms: float
median_ms: float
max_ms: float
def _timed_requests(client: APIClient, url: str, n: int) -> list[float]:
times = []
for _ in range(n):
t0 = time.perf_counter()
resp = client.get(url)
t1 = time.perf_counter()
if resp.status_code != 200:
raise RuntimeError(
f"GET {url} -> {resp.status_code}: {resp.content[:300]!r}",
)
times.append(t1 - t0)
return times
def _query_count(client: APIClient, url: str) -> int:
from django.db import connection
from django.test.utils import CaptureQueriesContext
with CaptureQueriesContext(connection) as ctx:
resp = client.get(url)
if resp.status_code != 200:
raise RuntimeError(f"GET {url} -> {resp.status_code}: {resp.content[:300]!r}")
return len(ctx.captured_queries)
def run_endpoint_benchmarks(
*,
perf_target: User,
perf_admin: User,
repeat: int,
) -> list[EndpointTiming]:
from rest_framework.test import APIClient
results: list[EndpointTiming] = []
for user_label, user in (("target", perf_target), ("admin", perf_admin)):
client = APIClient()
client.force_authenticate(user=user)
for name, url in ENDPOINTS:
client.get(url) # warm-up request, not counted
qcount = _query_count(client, url)
times_ms = [t * 1000 for t in _timed_requests(client, url, repeat)]
results.append(
EndpointTiming(
user_label=user_label,
endpoint_name=name,
query_count=qcount,
min_ms=min(times_ms),
median_ms=statistics.median(times_ms),
max_ms=max(times_ms),
),
)
return results
+56
View File
@@ -0,0 +1,56 @@
# src/paperless_benchmark/harness.py
from __future__ import annotations
import time
from dataclasses import dataclass
from typing import TYPE_CHECKING
from typing import Generic
from typing import TypeVar
from django.db import connection
from django.test.utils import CaptureQueriesContext
if TYPE_CHECKING:
from collections.abc import Callable
T = TypeVar("T")
@dataclass(frozen=True, slots=True)
class ProfileResult(Generic[T]):
best_seconds: float
all_seconds: tuple[float, ...]
query_count: int
result: T
def run_profile(fn: Callable[[], T], *, repeat: int = 5) -> ProfileResult[T]:
"""
Call `fn` `repeat` times, capturing wall-clock time for every call and
the SQL query count for the final call. Returns the best (minimum)
time across all repeats, since the first call(s) can be skewed by
connection warm-up or cold caches.
"""
if repeat < 1:
raise ValueError("repeat must be >= 1")
all_seconds: list[float] = []
result: T | None = None
query_count = 0
for i in range(repeat):
with CaptureQueriesContext(connection) as ctx:
start = time.perf_counter()
result = fn()
all_seconds.append(time.perf_counter() - start)
if i == repeat - 1:
query_count = len(ctx.captured_queries)
# Purely a type-narrowing aid for the type checker: the `repeat < 1`
# guard above already turns the one case that could leave `result`
# unset into a clear ValueError, so this is unreachable in practice.
assert result is not None
return ProfileResult(
best_seconds=min(all_seconds),
all_seconds=tuple(all_seconds),
query_count=query_count,
result=result,
)
@@ -0,0 +1,222 @@
# src/paperless_benchmark/management/commands/benchmark.py
from __future__ import annotations
from typing import Any
from django.core.management.base import BaseCommand
from django.core.management.base import CommandError
from django.core.management.base import CommandParser
class Command(BaseCommand):
help = "Seed, run, and profile paperless-ngx performance benchmarks."
def add_arguments(self, parser: CommandParser) -> None:
parser.add_argument(
"action",
choices=["seed", "run", "profile", "list-scenarios"],
help="Which benchmark action to perform.",
)
parser.add_argument(
"scenario",
nargs="?",
default=None,
help="Scenario name (required for `profile`; see `list-scenarios`).",
)
parser.add_argument(
"--tier",
choices=["home", "medium", "large"],
default="medium",
help="Dataset scale tier for `seed` (default: medium).",
)
parser.add_argument(
"--reset",
action="store_true",
default=False,
help="For `seed`: truncate existing benchmark data first.",
)
parser.add_argument(
"--yes-i-know-this-wipes-the-database",
action="store_true",
default=False,
help=(
"Required alongside --reset: confirms you understand `seed "
"--reset` deletes ALL users, ALL groups, and ALL documents/"
"tags/correspondents/document types/storage paths in this "
"database, not just benchmark-created ones."
),
)
parser.add_argument(
"--seed",
type=int,
default=42,
help="RNG seed for reproducible datasets (default: 42).",
)
parser.add_argument(
"--repeat",
type=int,
default=5,
help="Number of timed repetitions for `run`/`profile` (default: 5).",
)
parser.add_argument(
"--label",
default="baseline",
help="Free-text tag for a `run`, printed and recorded in history only.",
)
parser.add_argument(
"--explain",
action="store_true",
default=False,
help="For `profile`: also capture and print the query plan.",
)
def handle(self, *args: Any, **options: Any) -> None:
action = options["action"]
if action == "seed":
self._handle_seed(options)
elif action == "run":
self._handle_run(options)
elif action == "profile":
self._handle_profile(options)
else:
self._handle_list_scenarios()
def _handle_seed(self, options: dict[str, Any]) -> None:
from paperless_benchmark.db import reset_benchmark_data
from paperless_benchmark.seeding import seed_benchmark_dataset
if options["reset"]:
if not options["yes_i_know_this_wipes_the_database"]:
raise CommandError(
"--reset requires --yes-i-know-this-wipes-the-database. "
"This deletes ALL users, ALL groups, and ALL documents, "
"tags, correspondents, document types, and storage paths "
"in this database -- not just benchmark-created ones. "
"Only run this against a disposable benchmark database, "
"never a real install. Re-run with "
"--reset --yes-i-know-this-wipes-the-database to proceed.",
)
self.stdout.write("Resetting existing benchmark data...")
reset_benchmark_data()
data = seed_benchmark_dataset(options["tier"], seed=options["seed"])
self.stdout.write(
self.style.SUCCESS(
f"Seeded tier={options['tier']!r}: {data.documents} documents, "
f"{len(data.users)} users, {len(data.groups)} groups.",
),
)
def _handle_run(self, options: dict[str, Any]) -> None:
from django.contrib.auth import get_user_model
from django.db import connection
from documents.models import Document
from paperless_benchmark.endpoints import run_endpoint_benchmarks
from paperless_benchmark.results import append_history
user_model = get_user_model()
try:
perf_target = user_model.objects.get(username="perf_target")
perf_admin = user_model.objects.get(username="perf_admin")
except user_model.DoesNotExist as e:
raise CommandError(
"No benchmark dataset found. Run `manage.py benchmark seed` first.",
) from e
db_vendor = connection.vendor
document_count = Document.objects.count()
results = run_endpoint_benchmarks(
perf_target=perf_target,
perf_admin=perf_admin,
repeat=options["repeat"],
)
self.stdout.write(f"# label={options['label']} repeat={options['repeat']}")
self.stdout.write(
f"{'user':7s} {'endpoint':20s} {'queries':>8s} "
f"{'min_ms':>9s} {'median_ms':>10s} {'max_ms':>9s}",
)
for r in results:
self.stdout.write(
f"{r.user_label:7s} {r.endpoint_name:20s} {r.query_count:8d} "
f"{r.min_ms:9.1f} {r.median_ms:10.1f} {r.max_ms:9.1f}",
)
append_history(
{
"mode": "run",
"label": options["label"],
"user": r.user_label,
"endpoint": r.endpoint_name,
"query_count": r.query_count,
"min_ms": r.min_ms,
"median_ms": r.median_ms,
"max_ms": r.max_ms,
"db_vendor": db_vendor,
"document_count": document_count,
},
)
def _handle_profile(self, options: dict[str, Any]) -> None:
from django.contrib.auth import get_user_model
from django.db import connection
from documents.models import Document
from paperless_benchmark.db import capture_explain
from paperless_benchmark.harness import run_profile
from paperless_benchmark.results import append_history
from paperless_benchmark.scenarios import get as get_scenario
if not options["scenario"]:
raise CommandError(
"`profile` requires a scenario name; see `list-scenarios`.",
)
scenario = get_scenario(options["scenario"])
user_model = get_user_model()
try:
perf_target = user_model.objects.get(username="perf_target")
except user_model.DoesNotExist as e:
raise CommandError(
"No benchmark dataset found. Run `manage.py benchmark seed` first.",
) from e
profile = run_profile(
lambda: scenario.run(perf_target),
repeat=options["repeat"],
)
self.stdout.write(
f"{scenario.name}: best={profile.best_seconds:.4f}s "
f"queries={profile.query_count}",
)
if options["explain"]:
if scenario.queryset_for_explain is not None:
plan = capture_explain(scenario.queryset_for_explain(perf_target))
self.stdout.write(plan)
else:
self.stdout.write(
self.style.WARNING(
f"--explain was requested but scenario {scenario.name!r} "
"does not support it (no queryset_for_explain); skipping.",
),
)
append_history(
{
"mode": "profile",
"scenario": scenario.name,
"best_seconds": profile.best_seconds,
"query_count": profile.query_count,
"db_vendor": connection.vendor,
"document_count": Document.objects.count(),
},
)
def _handle_list_scenarios(self) -> None:
from paperless_benchmark.scenarios import all_scenarios
for scenario in all_scenarios():
self.stdout.write(f"{scenario.name}: {scenario.describe}")
+39
View File
@@ -0,0 +1,39 @@
# src/paperless_benchmark/results.py
from __future__ import annotations
import json
import subprocess
from datetime import UTC
from datetime import datetime
from pathlib import Path
from typing import Any
RESULTS_DIR = Path(__file__).resolve().parent.parent.parent / "benchmark_results"
def _current_git_ref() -> str:
result = subprocess.run(
["git", "rev-parse", "--short", "HEAD"],
capture_output=True,
text=True,
check=False,
)
return result.stdout.strip() or "unknown"
def append_history(entry: dict[str, Any], *, code_ref: str | None = None) -> None:
"""
Append one line to benchmark_results/history.jsonl -- a local-only,
append-only, cross-session record of every `benchmark run`/`profile`
invocation. Unlike a single overwritten snapshot file, this survives
across sessions so a benchmarking effort picked back up days later has
a full timeline instead of only the most recent result.
"""
RESULTS_DIR.mkdir(exist_ok=True)
record = {
"timestamp": datetime.now(UTC).isoformat(),
"code_ref": code_ref or _current_git_ref(),
**entry,
}
with (RESULTS_DIR / "history.jsonl").open("a") as f:
f.write(json.dumps(record) + "\n")
+110
View File
@@ -0,0 +1,110 @@
# src/paperless_benchmark/scenarios.py
from __future__ import annotations
from dataclasses import dataclass
from typing import TYPE_CHECKING
from typing import Any
if TYPE_CHECKING:
from collections.abc import Callable
from django.contrib.auth.models import User
from django.db.models import QuerySet
@dataclass(frozen=True, slots=True)
class Scenario:
name: str
describe: str
run: Callable[[User], Any]
queryset_for_explain: Callable[[User], QuerySet] | None = None
_SCENARIOS: dict[str, Scenario] = {}
def register(scenario: Scenario) -> None:
_SCENARIOS[scenario.name] = scenario
def get(name: str) -> Scenario:
from django.core.management.base import CommandError
try:
return _SCENARIOS[name]
except KeyError:
available = ", ".join(sorted(_SCENARIOS)) or "(none registered)"
raise CommandError(
f"Unknown benchmark scenario {name!r}. Available: {available}",
) from None
def all_scenarios() -> tuple[Scenario, ...]:
return tuple(_SCENARIOS.values())
def _guardian_visibility_query_run(user: User) -> list[int]:
from documents.models import Document
from documents.permissions import get_objects_for_user_owner_aware
return list(
get_objects_for_user_owner_aware(
user,
"documents.view_document",
Document,
).values_list("id", flat=True),
)
def _guardian_visibility_query_queryset(user: User) -> QuerySet:
from documents.models import Document
from documents.permissions import get_objects_for_user_owner_aware
return get_objects_for_user_owner_aware(user, "documents.view_document", Document)
register(
Scenario(
name="guardian_visibility_query",
describe=(
"Document-visibility queryset for a user with mixed owned/shared "
"documents -- exercises documents.permissions."
"get_objects_for_user_owner_aware's guardian permission join."
),
run=_guardian_visibility_query_run,
queryset_for_explain=_guardian_visibility_query_queryset,
),
)
def _permitted_document_ids_run(user: User) -> list[int]:
from documents.models import Document
from documents.permissions import permitted_document_ids
return list(
Document.objects.filter(id__in=permitted_document_ids(user)).values_list(
"id",
flat=True,
),
)
def _permitted_document_ids_queryset(user: User) -> QuerySet:
from documents.models import Document
from documents.permissions import permitted_document_ids
return Document.objects.filter(id__in=permitted_document_ids(user))
register(
Scenario(
name="permitted_document_ids",
describe=(
"Document-visibility query built from documents.permissions."
"permitted_document_ids -- the resolved-ID-set alternative to "
"guardian_visibility_query, for side-by-side comparison."
),
run=_permitted_document_ids_run,
queryset_for_explain=_permitted_document_ids_queryset,
),
)
+445
View File
@@ -0,0 +1,445 @@
# src/paperless_benchmark/seeding.py
from __future__ import annotations
import datetime
import random
import time
from dataclasses import dataclass
from typing import TYPE_CHECKING
from typing import Literal
if TYPE_CHECKING:
from django.contrib.auth.models import Group
from django.contrib.auth.models import User
Tier = Literal["home", "medium", "large"]
CHUNK_SIZE = 5_000
MIME_TYPES = (
"application/pdf",
"image/png",
"image/jpeg",
"text/plain",
)
# Ownership split for documents, mirroring the shape used in the #11950
# perf-benchmark dataset (owned-by-target / owned-by-other / unowned).
OWNED_BY_TARGET_FRACTION = 0.60
OWNED_BY_OTHER_FRACTION = 0.30
# remainder (0.10) is unowned
# Of documents owned by "other" users, the fraction explicitly shared
# (view, or view+change) with perf_target via guardian permissions --
# this is what exercises the get_user_can_change() per-row N+1 that the
# `run` endpoint benchmarks measure.
SHARED_WITH_TARGET_FRACTION = 0.5
SHARED_WITH_CHANGE_FRACTION = 0.5
# Guardian permission-row ratios measured from a real install (discussion
# #13276): 1,414 user-perm rows / 27,232 group-perm rows over 12,000
# documents. Layered across ALL owned documents for the general user/group
# pool (not just perf_target's shares), so `profile` scenarios exercise a
# realistic permission-join shape for arbitrary users, not only perf_target.
USER_PERM_ROWS_PER_DOC = 1_414 / 12_000
GROUP_PERM_ROWS_PER_DOC = 27_232 / 12_000
@dataclass(frozen=True, slots=True)
class _TierCounts:
documents: int
tags: int
correspondents: int
document_types: int
storage_paths: int
other_users: int
groups: int
tags_per_doc: tuple[int, int]
TIERS: dict[Tier, _TierCounts] = {
"home": _TierCounts(
documents=500,
tags=20,
correspondents=10,
document_types=8,
storage_paths=5,
other_users=3,
groups=2,
tags_per_doc=(1, 3),
),
"medium": _TierCounts(
documents=20_000,
tags=100,
correspondents=300,
document_types=50,
storage_paths=20,
other_users=10,
groups=5,
tags_per_doc=(2, 6),
),
"large": _TierCounts(
documents=360_000,
tags=1_000,
correspondents=5_000,
document_types=300,
storage_paths=50,
other_users=25,
groups=10,
tags_per_doc=(3, 7),
),
}
def log(msg: str) -> None:
print(f"[{time.strftime('%H:%M:%S')}] {msg}", flush=True) # noqa: T201
@dataclass(frozen=True, slots=True)
class SeededData:
"""
Summary of a completed seed run. `documents`/`tags`/`correspondents`/
`document_types`/`storage_paths` are counts, not the seeded ORM
instances: at the `large` tier (360,000 documents) holding every
instance in memory simultaneously is a real risk for zero benefit, since
no caller consumes anything but the counts. `perf_target`/`perf_admin`/
`users`/`groups` stay as real objects -- at most ~26 users/12 groups even
at `large` tier, and small enough to be useful to a future caller.
"""
perf_target: User
perf_admin: User
users: tuple[User, ...]
groups: tuple[Group, ...]
documents: int
tags: int
correspondents: int
document_types: int
storage_paths: int
def _grant_model_level_permissions(user: User) -> None:
"""
Grant perf_target Django model-level view/add/change permissions on
Document and Tag, on top of the per-object guardian grants seeding
creates elsewhere. DRF's PaperlessObjectPermissions checks model-level
permissions before guardian's object-level ones are ever consulted, so
without this perf_target gets a blanket 403 on /api/documents/ and
/api/tags/ regardless of which documents guardian says it can see.
"""
from django.contrib.auth.models import Permission
from django.contrib.contenttypes.models import ContentType
from documents.models import Document
from documents.models import Tag
for model in (Document, Tag):
content_type = ContentType.objects.get_for_model(model)
codenames = [
f"{action}_{model._meta.model_name}" for action in ("view", "add", "change")
]
perms = Permission.objects.filter(
content_type=content_type,
codename__in=codenames,
)
user.user_permissions.add(*perms)
def _create_users_and_groups(
counts: _TierCounts,
) -> tuple[User, User, tuple[User, ...], tuple[Group, ...]]:
from django.contrib.auth.models import Group
from documents.tests.factories import UserFactory
perf_target = UserFactory.create(username="perf_target")
_grant_model_level_permissions(perf_target)
perf_admin = UserFactory.create(username="perf_admin", superuser=True)
other_users = tuple(UserFactory.create_batch(counts.other_users))
groups = tuple(
Group.objects.create(name=f"benchmark_group_{i}") for i in range(counts.groups)
)
log(
f"Created users: 1 target, 1 superuser, {len(other_users)} other, "
f"{len(groups)} groups.",
)
return perf_target, perf_admin, other_users, groups
def _create_lookup_tables(counts: _TierCounts):
from documents.models import Correspondent
from documents.models import DocumentType
from documents.models import StoragePath
from documents.models import Tag
from documents.tests.factories import CorrespondentFactory
from documents.tests.factories import DocumentTypeFactory
from documents.tests.factories import StoragePathFactory
from documents.tests.factories import TagFactory
tags = tuple(Tag.objects.bulk_create(TagFactory.build_batch(counts.tags)))
correspondents = tuple(
Correspondent.objects.bulk_create(
CorrespondentFactory.build_batch(counts.correspondents),
),
)
document_types = tuple(
DocumentType.objects.bulk_create(
DocumentTypeFactory.build_batch(counts.document_types),
),
)
storage_paths = tuple(
StoragePath.objects.bulk_create(
StoragePathFactory.build_batch(counts.storage_paths),
),
)
log(
f"Created {len(tags)} tags, {len(correspondents)} correspondents, "
f"{len(document_types)} document types, {len(storage_paths)} storage paths.",
)
return tags, correspondents, document_types, storage_paths
def _assign_owner(rng: random.Random, perf_target: User, other_users: tuple[User, ...]):
roll = rng.random()
if roll < OWNED_BY_TARGET_FRACTION:
return perf_target, "target"
if roll < OWNED_BY_TARGET_FRACTION + OWNED_BY_OTHER_FRACTION:
return rng.choice(other_users), "other"
return None, "unowned"
def _seed_documents(
rng: random.Random,
counts: _TierCounts,
tags,
correspondents,
document_types,
storage_paths,
perf_target,
other_users,
) -> tuple[int, list[int]]:
"""
Bulk-create `counts.documents` documents in chunks. Returns the total
document count plus a lightweight list of pks for documents that ended
up with an owner (target or other) -- that's all
`_grant_general_permissions` needs to sample from, so full `Document`
instances aren't accumulated across chunks (a real memory concern at the
`large` tier's 360,000 documents).
"""
from django.contrib.auth.models import Permission
from django.contrib.contenttypes.models import ContentType
from guardian.models import UserObjectPermission
from documents.models import Document
from documents.tests.factories import DocumentFactory
tag_ids = [t.pk for t in tags]
correspondent_ids = [c.pk for c in correspondents]
document_type_ids = [d.pk for d in document_types]
storage_path_ids = [s.pk for s in storage_paths]
doc_content_type = ContentType.objects.get_for_model(Document)
view_perm = Permission.objects.get(
codename="view_document",
content_type=doc_content_type,
)
change_perm = Permission.objects.get(
codename="change_document",
content_type=doc_content_type,
)
through_model = Document.tags.through
document_count = 0
owned_document_pks: list[int] = []
remaining = counts.documents
while remaining > 0:
chunk_n = min(CHUNK_SIZE, remaining)
remaining -= chunk_n
batch = []
owner_buckets = []
for _ in range(chunk_n):
doc = DocumentFactory.build(
mime_type=rng.choice(MIME_TYPES),
page_count=rng.randint(1, 30),
correspondent_id=(
rng.choice(correspondent_ids)
if correspondent_ids and rng.random() < 0.8
else None
),
document_type_id=(
rng.choice(document_type_ids)
if document_type_ids and rng.random() < 0.6
else None
),
storage_path_id=(
rng.choice(storage_path_ids)
if storage_path_ids and rng.random() < 0.8
else None
),
# Document.created is a plain DateField (default: today).
# Leaving it unset would give every seeded document the same
# date, collapsing Document's ("-created",) ordering index
# into a single-valued sort key -- spread it over a
# realistic multi-year window instead.
created=datetime.date.today()
- datetime.timedelta(days=rng.randint(0, 365 * 3)),
)
owner, bucket = _assign_owner(rng, perf_target, other_users)
doc.owner_id = owner.pk if owner else None
batch.append(doc)
owner_buckets.append(bucket)
created = Document.objects.bulk_create(batch, batch_size=CHUNK_SIZE)
through_rows = []
for doc in created:
k = rng.randint(*counts.tags_per_doc)
for tag_id in rng.sample(tag_ids, min(k, len(tag_ids))):
through_rows.append(through_model(document_id=doc.pk, tag_id=tag_id))
if through_rows:
through_model.objects.bulk_create(through_rows, batch_size=CHUNK_SIZE)
perm_rows = []
for doc, bucket in zip(created, owner_buckets, strict=True):
if bucket == "unowned":
continue
owned_document_pks.append(doc.pk)
if bucket != "other":
continue
if rng.random() >= SHARED_WITH_TARGET_FRACTION:
continue
perm_rows.append(
UserObjectPermission(
permission=view_perm,
content_type=doc_content_type,
object_pk=str(doc.pk),
user=perf_target,
),
)
if rng.random() < SHARED_WITH_CHANGE_FRACTION:
perm_rows.append(
UserObjectPermission(
permission=change_perm,
content_type=doc_content_type,
object_pk=str(doc.pk),
user=perf_target,
),
)
if perm_rows:
UserObjectPermission.objects.bulk_create(perm_rows, batch_size=CHUNK_SIZE)
document_count += len(created)
log(f" {document_count}/{counts.documents} documents seeded")
return document_count, owned_document_pks
def _grant_general_permissions(
rng: random.Random,
owned_document_pks: list[int],
users,
groups,
) -> None:
"""
Layer realistic (issue #13276-derived) guardian permission-row ratios
across owned documents for the general user/group pool, so `profile`
scenarios exercise the same permission-join shape regardless of which
user they check visibility for (not just perf_target).
"""
from django.contrib.auth.models import Permission
from django.contrib.contenttypes.models import ContentType
from guardian.models import GroupObjectPermission
from guardian.models import UserObjectPermission
from documents.models import Document
if not owned_document_pks or not users:
return
doc_content_type = ContentType.objects.get_for_model(Document)
view_perm = Permission.objects.get(
codename="view_document",
content_type=doc_content_type,
)
n_user_perms = round(len(owned_document_pks) * USER_PERM_ROWS_PER_DOC)
n_group_perms = (
round(len(owned_document_pks) * GROUP_PERM_ROWS_PER_DOC) if groups else 0
)
user_rows = [
UserObjectPermission(
permission=view_perm,
content_type=doc_content_type,
object_pk=str(rng.choice(owned_document_pks)),
user=rng.choice(users),
)
for _ in range(n_user_perms)
]
if user_rows:
UserObjectPermission.objects.bulk_create(
user_rows,
batch_size=CHUNK_SIZE,
ignore_conflicts=True,
)
group_rows = [
GroupObjectPermission(
permission=view_perm,
content_type=doc_content_type,
object_pk=str(rng.choice(owned_document_pks)),
group=rng.choice(groups),
)
for _ in range(n_group_perms)
]
if group_rows:
GroupObjectPermission.objects.bulk_create(
group_rows,
batch_size=CHUNK_SIZE,
ignore_conflicts=True,
)
log(f" Granted {len(user_rows)} user perms, {len(group_rows)} group perms.")
def seed_benchmark_dataset(tier: Tier, *, seed: int = 42) -> SeededData:
"""
Build a benchmark dataset at the given scale tier: a named perf_target
(mixed owned/shared documents) and perf_admin (superuser) for endpoint
benchmarking, plus a general user/group pool with realistic guardian
permission-row ratios for profile scenarios.
"""
counts = TIERS[tier]
rng = random.Random(seed)
log(f"Seeding tier={tier!r}")
perf_target, perf_admin, other_users, groups = _create_users_and_groups(counts)
tags, correspondents, document_types, storage_paths = _create_lookup_tables(counts)
document_count, owned_document_pks = _seed_documents(
rng,
counts,
tags,
correspondents,
document_types,
storage_paths,
perf_target,
other_users,
)
all_users = (perf_target, *other_users)
_grant_general_permissions(rng, owned_document_pks, all_users, groups)
log(f"Done. {document_count} documents seeded for tier={tier!r}.")
return SeededData(
perf_target=perf_target,
perf_admin=perf_admin,
users=all_users,
groups=groups,
documents=document_count,
tags=len(tags),
correspondents=len(correspondents),
document_types=len(document_types),
storage_paths=len(storage_paths),
)