mirror of
https://github.com/paperless-ngx/paperless-ngx.git
synced 2026-08-10 04:43:19 +00:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d8c9d22ea1 | ||
|
|
50ed8c060b | ||
|
|
00631146ff | ||
|
|
d09caf480c | ||
|
|
62089df2d8 | ||
|
|
5e5f6a88a3 | ||
|
|
02e6c49c62 | ||
|
|
3be64da4cb |
@@ -0,0 +1,55 @@
|
|||||||
|
---
|
||||||
|
name: whoosh-compat-transition
|
||||||
|
description: Use when integrating the whoosh-compat library into paperless-ngx search, replacing src/documents/search/_translate.py or _dates.py, building the search FieldRegistry, or changing user query parsing during the whoosh-to-tantivy transition
|
||||||
|
---
|
||||||
|
|
||||||
|
# whoosh-compat transition
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
whoosh-compat (github.com/stumpylog/whoosh-compat; local checkout usually at `../whoosh-compat`) replaces the hand-maintained translation layer (`src/documents/search/_translate.py`, `_dates.py`): it parses user queries with a faithful fork of whoosh's real grammar into a typed AST and emits programmatic tantivy queries. Read its README and ARCHITECTURE.md before wiring anything; its DIVERGENCES.md lists intended behavior differences and is the authority on "is this difference a bug".
|
||||||
|
|
||||||
|
## Decisions already made (do not re-derive)
|
||||||
|
|
||||||
|
- **Queries are user-typed free text.** The advanced search box passes whatever the user types straight to the parser (that is how the issue #13568 queries exist). Do NOT try to infer the supported field surface from frontend code; the frontend only generates a few date filter strings, everything else is typed by users.
|
||||||
|
- **The field surface is a policy decision, not `KNOWN_FIELDS`.** Today's `KNOWN_FIELDS` accepts internal ID fields (`tag_id`, `owner_id`, `viewer_id`, other `*_id`) that are undocumented in `docs/usage.md` and were ruled not user-searchable by the maintainer: exclude them from the `FieldRegistry` (they stay as programmatic permission/filter fields in `build_permission_filter`, which never touches user query text). The registry is built from documented syntax in `docs/usage.md` plus the v2-compat aliases (`type`, `path`, `type_id`-style aliases follow their canonical field's fate). Undocumented-but-working fields (`asn`, `page_count`, `num_notes`, `original_filename`, `checksum`) need an explicit maintainer yes/no; since users type freely, silently dropping one breaks any saved view using it, so a drop must be a visible, documented decision.
|
||||||
|
- **Analyzer seam:** `FieldSpec.analyzer` binds the live registered tantivy analyzer's `.analyze` (the same Rust analyzer used at index time; language-keyed, so rebuild the registry when `SEARCH_LANGUAGE` changes, on the same trigger as `register_tokenizers`). `pattern_normalizer` is `_tokenizer.ascii_fold`: character-level lowercase+fold only, NEVER stemming.
|
||||||
|
- **Diagnostics before emit:** `whoosh_compat.parse()` never raises on bad input. Check `ParseResult.diagnostics` and map to `SearchQueryError`/`InvalidDateQuery` (HTTP 400) BEFORE calling `emit()`; also catch the emitter's `UnsupportedQueryError` into a 400. Never carry forward the legacy raw-string fallback (`except Exception: query_str = raw_query`) into the new path; it masks integration bugs.
|
||||||
|
- **`notes` and `custom_fields` are JSON fields** with fixed subpaths (`notes.user`/`notes.note`, `custom_fields.name`/`custom_fields.value`); the registry stays a static, language-keyed singleton, never per-request.
|
||||||
|
|
||||||
|
## Mandatory before deleting old code
|
||||||
|
|
||||||
|
- Date-grammar parity audit, line by line: every keyword, relative unit, and abbreviation `_dates.py` and `_translate.py` accept today (including the whoosh-era abbreviations kept for old saved views) must have an accepted form in whoosh-compat's dateparse grammar. Silent keyword loss is the saved-view breakage class behind issue #13568.
|
||||||
|
- Acceptance corpus compared by matched-document-ID sets, not query strings: the #13568 queries verbatim, real saved-view strings, every date keyword, field aliases, comma lists, date and numeric ranges, wildcards with bracket classes, boosts, JSON subpaths.
|
||||||
|
|
||||||
|
## Tests: what goes, what comes
|
||||||
|
|
||||||
|
Removed with their modules (do not port their string-level assertions):
|
||||||
|
|
||||||
|
- `src/documents/tests/search/test_translate.py`: its subject is deleted; string-translation unit cases are whoosh-compat's own responsibility now. Cases that encode real user-visible behavior get reincarnated as result-level acceptance cases, not string assertions.
|
||||||
|
- Date-keyword unit tests tied to `_dates.py` internals: same treatment.
|
||||||
|
- `test_query.py` cases asserting `parse_user_query` internals or intermediate query strings: rewritten against the new pipeline, asserting on matched results.
|
||||||
|
|
||||||
|
Kept: `test_migration_fulltext_query_field_prefixes.py` (data migration, orthogonal), `test_schema.py`, `test_tokenizer.py`, permission-filter and simple-search tests.
|
||||||
|
|
||||||
|
Added:
|
||||||
|
|
||||||
|
- A result-level acceptance module (paperless's analogue of whoosh-compat's `test_acceptance_e2e.py`): the corpus above against a real index built from `build_schema()`, asserting document-ID sets. Use `pytest.param(..., id="...")` for every case.
|
||||||
|
- Registry unit tests: internal `*_id` names rejected, aliases resolve to canonical fields, JSON subpaths match `docs/usage.md`, construction deterministic per language.
|
||||||
|
- One `Multitoken` case nested inside a top-level `OR` (whoosh-compat DIVERGENCES entry on Multitoken.DEFAULT) to prove it does not matter for paperless's data.
|
||||||
|
- If acceptance work surfaces a new whoosh-compat divergence, that is a whoosh-compat-repo change (its `differential-triage` skill applies), not a silent paperless workaround.
|
||||||
|
|
||||||
|
## Coordination
|
||||||
|
|
||||||
|
- whoosh-compat is pre-1.0: pin an exact version or git SHA; upgrades are deliberate, reviewed changes.
|
||||||
|
- JSON subpath emission depends on the installed tantivy-py version (fallback until quickwit-oss/tantivy-py#716 ships). The whoosh-compat repo has a `carve-out-retirement` skill; coordinate tantivy pin bumps with it, in a separate PR from the parser migration.
|
||||||
|
- Rollout: settings flag defaulting to the legacy path plus shadow-compare logging (log when old and new paths return different ID sets; sample if cost matters) for one release; delete `_translate.py`/`_dates.py` only after the flag defaults to the new path with no material reports.
|
||||||
|
|
||||||
|
## Common mistakes
|
||||||
|
|
||||||
|
- Inferring the field surface from frontend code (users type queries directly).
|
||||||
|
- Copying `KNOWN_FIELDS` into the registry wholesale (resurfaces internal fields).
|
||||||
|
- Wiring stemming into `pattern_normalizer`.
|
||||||
|
- Calling `emit()` unconditionally, or porting the legacy raw-string fallback.
|
||||||
|
- Deleting `_dates.py` without the parity audit.
|
||||||
|
- Porting `test_translate.py`'s string assertions instead of writing result-level tests.
|
||||||
@@ -173,10 +173,6 @@ RUN set -eux \
|
|||||||
&& rm --force --verbose *.deb \
|
&& rm --force --verbose *.deb \
|
||||||
&& rm --recursive --force --verbose /var/lib/apt/lists/*
|
&& rm --recursive --force --verbose /var/lib/apt/lists/*
|
||||||
|
|
||||||
# Ensure interactive shells (docker exec bash) see resolved *_FILE secrets,
|
|
||||||
# mirroring what with-contenv already does for s6 services.
|
|
||||||
RUN echo '. /etc/profile.d/contenv.sh' >> /etc/bash.bashrc
|
|
||||||
|
|
||||||
WORKDIR /usr/src/paperless/src/
|
WORKDIR /usr/src/paperless/src/
|
||||||
|
|
||||||
# Python dependencies
|
# Python dependencies
|
||||||
|
|||||||
@@ -1,18 +0,0 @@
|
|||||||
#!/bin/sh
|
|
||||||
# Source s6 container environment for interactive shells.
|
|
||||||
# Ensures variables resolved from *_FILE secret injection are visible
|
|
||||||
# when using 'docker exec bash'. Does not affect s6 services (those
|
|
||||||
# use with-contenv directly). Has no effect in non-container contexts
|
|
||||||
# because the directory will not exist.
|
|
||||||
# Note: sh/dash shells opened via 'docker exec sh' are not covered;
|
|
||||||
# only bash-based sessions benefit from this file.
|
|
||||||
_pngx_contenv="/run/s6/container_environment"
|
|
||||||
if [ -d "${_pngx_contenv}" ]; then
|
|
||||||
for _pngx_f in "${_pngx_contenv}"/*; do
|
|
||||||
[ -f "${_pngx_f}" ] || continue
|
|
||||||
_pngx_name=$(basename "${_pngx_f}")
|
|
||||||
_pngx_val=$(cat "${_pngx_f}")
|
|
||||||
export "${_pngx_name}=${_pngx_val}"
|
|
||||||
done
|
|
||||||
fi
|
|
||||||
unset _pngx_contenv _pngx_f _pngx_name _pngx_val
|
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,416 @@
|
|||||||
|
# whoosh-compat transition design
|
||||||
|
|
||||||
|
Date: 2026-08-07
|
||||||
|
Status: approved, pending spec review
|
||||||
|
Related skill: `whoosh-compat-transition`
|
||||||
|
Related issue: [stumpylog/whoosh-compat#1](https://github.com/stumpylog/whoosh-compat/issues/1)
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
Replace paperless-ngx's hand-maintained query-translation layer
|
||||||
|
(`src/documents/search/_translate.py`, `src/documents/search/_dates.py`)
|
||||||
|
with [whoosh-compat](https://github.com/stumpylog/whoosh-compat): a typed
|
||||||
|
Whoosh-grammar parser that emits programmatically constructed
|
||||||
|
`tantivy.Query` objects instead of building an intermediate Tantivy query
|
||||||
|
_string_. The integration point is narrow: `parse_user_query()` in
|
||||||
|
`src/documents/search/_query.py` is the only function whose implementation
|
||||||
|
changes; `_backend.py`, `_tokenizer.py`, simple/title search, CJK handling,
|
||||||
|
and permission filtering are all unaffected.
|
||||||
|
|
||||||
|
Delivered as a stack of four paperless-ngx PRs plus one prerequisite change
|
||||||
|
in whoosh-compat itself (same maintainer, no cross-repo coordination
|
||||||
|
overhead), landed with no feature flag and no shadow-compare rollout period
|
||||||
|
— safety comes from a result-level acceptance test corpus and a
|
||||||
|
date-grammar parity audit instead.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
```
|
||||||
|
raw_query (user-typed)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
wc.parse(raw_query, registry=FIELD_REGISTRY, default_fields=DEFAULT_SEARCH_FIELDS,
|
||||||
|
field_boosts=_FIELD_BOOSTS, tz=tz)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
ParseResult(ast, diagnostics)
|
||||||
|
│
|
||||||
|
├─ diagnostics non-empty? → map ALL diagnostics to SearchQueryError
|
||||||
|
│ subclass(es) → HTTP 400 (never just the first diagnostic)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
emit(ast, index=index, schema=schema, registry=FIELD_REGISTRY)
|
||||||
|
│ (raises UnsupportedQueryError → mapped to SearchQueryError → 400,
|
||||||
|
│ for constructs that parse but can't execute against tantivy)
|
||||||
|
▼
|
||||||
|
tantivy.Query
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
existing clause assembly in parse_user_query(): Should(exact) + optional
|
||||||
|
fuzzy re-parse of raw_query + optional CJK bigram query, unchanged from today
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
_apply_permission_filter() in _backend.py wraps the result with
|
||||||
|
build_permission_filter() — entirely independent of whoosh-compat, unchanged
|
||||||
|
```
|
||||||
|
|
||||||
|
Permission filtering is explicitly out of scope for this migration:
|
||||||
|
`build_permission_filter()` builds its `tantivy.Query` directly against
|
||||||
|
`owner_id`/`viewer_id`/`viewer_group_id`, never through the parser or
|
||||||
|
registry, and those fields are exactly the internal `*_id` fields excluded
|
||||||
|
from the `FieldRegistry` (see "Field surface" below). Nothing in this
|
||||||
|
migration's diff touches it.
|
||||||
|
|
||||||
|
## PR stack
|
||||||
|
|
||||||
|
Each PR is independently buildable, reviewable, and CI-able; later PRs
|
||||||
|
rebase on earlier ones. No PR depends on whoosh-compat behavior it hasn't
|
||||||
|
already proven correct in isolation.
|
||||||
|
|
||||||
|
1. **Refactor `_schema.py` to a shared field-definition table.** Pure
|
||||||
|
refactor — `build_schema()`'s output is byte-identical before and after.
|
||||||
|
`test_schema.py` (existing) proves it.
|
||||||
|
2. **Pin whoosh-compat as a real dependency; build `FieldRegistry`.** New
|
||||||
|
`_registry.py` built from the same table PR 1 introduced. Registry unit
|
||||||
|
tests only — no wiring into search yet.
|
||||||
|
3. **Date-grammar parity audit.** A transitional, executable differential
|
||||||
|
test using the still-present `_dates.py`/`_translate.py` as the oracle.
|
||||||
|
Any gap found is fixed in whoosh-compat directly before this PR closes.
|
||||||
|
A whoosh-compat PyPI release is expected around this point (see
|
||||||
|
"Dependency pinning").
|
||||||
|
4. **Wire it in; delete the old path.** Rewrite `parse_user_query()`,
|
||||||
|
diagnostics→exception mapping, add the result-level acceptance corpus,
|
||||||
|
expand `test_api_search.py`, delete `_translate.py`/`_dates.py`/
|
||||||
|
`test_translate.py` and the internals-testing classes in `test_query.py`,
|
||||||
|
update `docs/usage.md` and changelog.
|
||||||
|
|
||||||
|
**Prerequisite, whoosh-compat repo** (tracked as
|
||||||
|
[stumpylog/whoosh-compat#1](https://github.com/stumpylog/whoosh-compat/issues/1),
|
||||||
|
lands before PR 4 starts its diagnostics-mapping work): add `field: str |
|
||||||
|
None` and `raw_value: str | None` to `Diagnostic`, threaded through at its
|
||||||
|
three construction sites (`dateparse.py`'s `_error()`, `default.py`'s
|
||||||
|
`term_query()` and `_coerce_range_bound()`), so paperless can build typed
|
||||||
|
exceptions without parsing whoosh-compat's human-readable `message` text.
|
||||||
|
|
||||||
|
## Field surface
|
||||||
|
|
||||||
|
The `FieldRegistry` covers only query-syntax-addressable fields — a subset
|
||||||
|
of the full Tantivy schema. Internal-only schema fields with no query-syntax
|
||||||
|
meaning of their own (`title_sort`/`correspondent_sort`/`type_sort` shadow
|
||||||
|
sort fields, `bigram_*` CJK fields, `simple_title`/`simple_content`,
|
||||||
|
`autocomplete_word`, `notes_text`) stay hardcoded `sb.add_*` calls in
|
||||||
|
`_schema.py`, untouched by the shared table.
|
||||||
|
|
||||||
|
**Decision: keep and document all five currently-undocumented-but-working
|
||||||
|
fields** (`asn`, `page_count`, `num_notes`, `original_filename`,
|
||||||
|
`checksum`) rather than dropping them — least risk of silently breaking an
|
||||||
|
existing saved view. `docs/usage.md`'s advanced-search section gets these
|
||||||
|
added with examples, as part of PR 4.
|
||||||
|
|
||||||
|
**Decision: `archive_checksum` stays out of scope.** Unlike `checksum`, it
|
||||||
|
isn't indexed in the Tantivy schema at all today (confirmed: `_schema.py`
|
||||||
|
only adds `checksum`; `_build_tantivy_doc` only calls
|
||||||
|
`doc.add_text("checksum", document.checksum)`). Making it searchable is a
|
||||||
|
schema-level change (new indexed field, new document population code), not
|
||||||
|
a parser-migration concern — left as a separate follow-up.
|
||||||
|
|
||||||
|
**Decision: internal `*_id` fields (`tag_id`, `correspondent_id`,
|
||||||
|
`document_type_id`, `storage_path_id`, `owner_id`, `viewer_id`,
|
||||||
|
`viewer_group_id`) are excluded from the `FieldRegistry` entirely.** They
|
||||||
|
remain Tantivy-schema-only, used exclusively by `build_permission_filter()`.
|
||||||
|
Because whoosh-compat folds any unrecognized `field:` prefix into literal
|
||||||
|
text (Whoosh-parity leniency, confirmed in `FieldsPlugin.do_fieldnames` —
|
||||||
|
not an error), a saved view typed as `tag_id:5` won't 400: it silently
|
||||||
|
becomes a text search for the literal string `tag_id:5`, most likely
|
||||||
|
returning zero results. This is a real behavior change and gets a
|
||||||
|
**changelog callout**, not just a docs update, since a docs addition alone
|
||||||
|
wouldn't surface it to someone skimming release notes.
|
||||||
|
|
||||||
|
## Shared field-definition table (`_fields.py`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
from whoosh_compat import FieldKind # reused directly — no parallel enum
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class PublicField:
|
||||||
|
name: str
|
||||||
|
kind: FieldKind
|
||||||
|
aliases: tuple[str, ...] = ()
|
||||||
|
comma_values: bool = False
|
||||||
|
date_only: bool = False
|
||||||
|
fast: bool = False
|
||||||
|
subpaths: tuple[str, ...] = () # JSON kind only
|
||||||
|
|
||||||
|
PUBLIC_FIELDS = (
|
||||||
|
PublicField("title", FieldKind.TEXT),
|
||||||
|
PublicField("content", FieldKind.TEXT),
|
||||||
|
PublicField("correspondent", FieldKind.TEXT),
|
||||||
|
PublicField("document_type", FieldKind.TEXT, aliases=("type",)),
|
||||||
|
PublicField("storage_path", FieldKind.TEXT, aliases=("path",)),
|
||||||
|
PublicField("original_filename", FieldKind.TEXT),
|
||||||
|
PublicField("tag", FieldKind.TEXT, comma_values=True),
|
||||||
|
PublicField("checksum", FieldKind.KEYWORD),
|
||||||
|
PublicField("asn", FieldKind.U64, fast=True),
|
||||||
|
PublicField("page_count", FieldKind.U64, fast=True),
|
||||||
|
PublicField("num_notes", FieldKind.U64, fast=True),
|
||||||
|
PublicField("created", FieldKind.DATE, date_only=True, fast=True),
|
||||||
|
PublicField("modified", FieldKind.DATETIME, fast=True),
|
||||||
|
PublicField("added", FieldKind.DATETIME, fast=True),
|
||||||
|
PublicField("notes", FieldKind.JSON, subpaths=("user", "note")),
|
||||||
|
PublicField("custom_fields", FieldKind.JSON, subpaths=("name", "value")),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
`build_schema()` derives its `sb.add_*` call and tokenizer from `kind`
|
||||||
|
(TEXT/KEYWORD → `add_text_field` with `paperless_text`/`raw` tokenizer
|
||||||
|
respectively; U64 → `add_unsigned_field`; DATE/DATETIME → `add_date_field`;
|
||||||
|
JSON → `add_json_field`). The `notes_text` snippet-companion field stays a
|
||||||
|
separate hardcoded line right after the `notes` entry — schema-only
|
||||||
|
plumbing with no query-syntax meaning.
|
||||||
|
|
||||||
|
`_registry.py` maps each `PublicField` to a `whoosh_compat.FieldSpec`,
|
||||||
|
kept as one flat dataclass (no kind-specific subclassing) to mirror
|
||||||
|
whoosh-compat's own `FieldSpec` design, which validates kind-conditional
|
||||||
|
attributes (e.g. JSON requires non-empty `subpaths`) at
|
||||||
|
`FieldRegistry.__init__` rather than in the type system.
|
||||||
|
|
||||||
|
Footnote for whoever writes `_registry.py`: `FieldRegistry.__init__` forces
|
||||||
|
`date_only=True` on _any_ `FieldKind.DATE` spec regardless of what's
|
||||||
|
passed, unconditionally — `PublicField.date_only` isn't an independent
|
||||||
|
knob for DATE fields the way it might look; it only matters in the sense
|
||||||
|
that `created` sets it explicitly for clarity, while `modified`/`added`
|
||||||
|
use `FieldKind.DATETIME` instead of relying on that override.
|
||||||
|
|
||||||
|
**`subpaths` stays `tuple[str, ...]`, not a nested structure.** Confirmed
|
||||||
|
against whoosh-compat's own `FieldRegistry.resolve_json()`: it splits a
|
||||||
|
dotted query term on the _first_ dot only and matches the remainder as an
|
||||||
|
exact string against `spec.subpaths` — even the docstring's own
|
||||||
|
`"metadata.author.name"` example is a single opaque string in the tuple,
|
||||||
|
not a recursive tree. A tuple of strings is exactly as expressive as the
|
||||||
|
library it feeds; inventing richer structure in `PublicField` now would
|
||||||
|
just get flattened back to strings at the registry-construction boundary.
|
||||||
|
Real recursive nesting, if ever needed, is new whoosh-compat capability
|
||||||
|
first.
|
||||||
|
|
||||||
|
**JSON document population stays separate from `subpaths`.** `subpaths` is
|
||||||
|
query-side only — it declares which dotted names are legal to type and
|
||||||
|
which JSON keys the emitter should address. It says nothing about how
|
||||||
|
`_backend.py::_build_tantivy_doc` builds the JSON documents at index-write
|
||||||
|
time, and that logic isn't uniform attribute access (`note.user.username`
|
||||||
|
needs a null guard and isn't `note.user`; `cfi.value_for_search` is a
|
||||||
|
property, not a literal `value` attribute), so a generic
|
||||||
|
`getattr(obj, subpath_name)` scheme would silently do the wrong thing for
|
||||||
|
both. That code stays hand-written, unchanged by this migration. Mitigation
|
||||||
|
instead: a coupling test (PR 2, alongside the registry unit tests) asserting
|
||||||
|
the literal JSON keys used in `_build_tantivy_doc`'s `doc.add_json(...)`
|
||||||
|
calls match `PUBLIC_FIELDS`' `notes`/`custom_fields` `subpaths` exactly, so
|
||||||
|
drift between the two is caught rather than silently becoming an
|
||||||
|
unqueryable (or silently unindexed) field.
|
||||||
|
|
||||||
|
**JSON subpath queries (`notes.*`, `custom_fields.*`) route through
|
||||||
|
`index.parse_query()`, not programmatic construction, given paperless's
|
||||||
|
pinned tantivy version.** Installed `tantivy-py`'s `Query.term_query`
|
||||||
|
cannot resolve a JSON subpath by exact field name — it raises as if the
|
||||||
|
field didn't exist. Until
|
||||||
|
[tantivy-py#716](https://github.com/quickwit-oss/tantivy-py/pull/716) lands
|
||||||
|
and ships, whoosh-compat's `TantivyEmitter._json_paths_supported()` feature-
|
||||||
|
detects this per process and falls back to a strictly escaped, single-leaf
|
||||||
|
`index.parse_query()` call for just that one leaf (whoosh-compat's README/
|
||||||
|
ARCHITECTURE.md call this out as "the JSON subpath carve-out"). Paperless
|
||||||
|
pins `tantivy~=0.26.0`, squarely inside the affected range (whoosh-compat's
|
||||||
|
`tantivy` extra only requires `tantivy>=0.24`, so nothing prevents this
|
||||||
|
combination). Nothing needs to change in this design because of it — the
|
||||||
|
carve-out is self-retiring on whoosh-compat's side once tantivy-py catches
|
||||||
|
up — but the acceptance corpus's `notes.user:`/`custom_fields.name:` cases
|
||||||
|
(PR 4) are exercising that fallback escaping path specifically, not the
|
||||||
|
programmatic path every other field goes through, and that's worth knowing
|
||||||
|
if one of those cases ever behaves oddly around quoting/escaping.
|
||||||
|
|
||||||
|
**Analyzer wiring**: `FieldSpec.analyzer` reuses the same `tantivy
|
||||||
|
.TextAnalyzer` objects `_tokenizer.py` already builds (`_paperless_text
|
||||||
|
(language)`, etc.) — standalone objects not dependent on index
|
||||||
|
registration, so `_registry.py` calls the same builder functions and binds
|
||||||
|
`.analyze` directly; `checksum` (KEYWORD, `raw` tokenizer) gets an identity
|
||||||
|
analyzer (`lambda t: [t]`). `pattern_normalizer` for every field is
|
||||||
|
`_tokenizer.ascii_fold` (character-fold only, never stemming) per the
|
||||||
|
skill's explicit instruction. The whole `FieldRegistry` is built once,
|
||||||
|
cached keyed by `settings.SEARCH_LANGUAGE`, rebuilt on the same trigger
|
||||||
|
`register_tokenizers()` already uses.
|
||||||
|
|
||||||
|
## Error handling
|
||||||
|
|
||||||
|
```python
|
||||||
|
class SearchQueryError(ValueError): ... # unchanged, base
|
||||||
|
|
||||||
|
class InvalidDateQuery(SearchQueryError): # unchanged
|
||||||
|
def __init__(self, field, value): ...
|
||||||
|
|
||||||
|
class InvalidNumberQuery(SearchQueryError): # new
|
||||||
|
def __init__(self, field: str | None, value: str | None) -> None:
|
||||||
|
self.field = field
|
||||||
|
self.value = value
|
||||||
|
super().__init__(f"Invalid numeric value {value!r} for field {field!r}.")
|
||||||
|
|
||||||
|
class MultipleSearchQueryErrors(SearchQueryError): # new
|
||||||
|
"""Aggregates every user-fixable error from one parse, not just the first."""
|
||||||
|
def __init__(self, errors: Sequence[SearchQueryError]) -> None:
|
||||||
|
self.errors = tuple(errors)
|
||||||
|
super().__init__("; ".join(str(e) for e in self.errors))
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
def parse_user_query(index, raw_query, tz):
|
||||||
|
registry = get_field_registry(settings.SEARCH_LANGUAGE)
|
||||||
|
result = wc.parse(
|
||||||
|
raw_query, registry=registry, default_fields=DEFAULT_SEARCH_FIELDS,
|
||||||
|
field_boosts=_FIELD_BOOSTS, tz=tz,
|
||||||
|
)
|
||||||
|
if result.diagnostics:
|
||||||
|
raise _diagnostics_to_error(result.diagnostics) # ALL diagnostics, not [0]
|
||||||
|
|
||||||
|
try:
|
||||||
|
exact = tantivy_emit(result.ast, index=index, schema=index.schema, registry=registry)
|
||||||
|
except UnsupportedQueryError as e:
|
||||||
|
raise SearchQueryError(str(e)) from e
|
||||||
|
|
||||||
|
# CJK: unchanged — already re-parses raw_query directly via index.parse_query,
|
||||||
|
# never went through translate_query, so nothing here changes.
|
||||||
|
cjk_query = _build_cjk_query(index, raw_query, _CJK_ALL_FIELDS) if _has_cjk(raw_query) else None
|
||||||
|
|
||||||
|
clauses = [(tantivy.Occur.Should, exact)]
|
||||||
|
threshold = settings.ADVANCED_FUZZY_SEARCH_THRESHOLD
|
||||||
|
if threshold is not None:
|
||||||
|
# Fuzzy re-parses raw_query (not the AST) — no clean AST-level fuzzy
|
||||||
|
# equivalent exists; fuzzy matching was always an approximate,
|
||||||
|
# secondary clause, so this divergence from the exact-match path is
|
||||||
|
# acceptable.
|
||||||
|
fuzzy = index.parse_query(raw_query, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS,
|
||||||
|
fuzzy_fields={f: (True, 1, True) for f in DEFAULT_SEARCH_FIELDS})
|
||||||
|
clauses.append((tantivy.Occur.Should, tantivy.Query.boost_query(fuzzy, 0.1)))
|
||||||
|
if cjk_query is not None:
|
||||||
|
clauses.append((tantivy.Occur.Should, cjk_query))
|
||||||
|
|
||||||
|
return exact if len(clauses) == 1 else tantivy.Query.boolean_query(clauses)
|
||||||
|
|
||||||
|
|
||||||
|
def _diagnostics_to_error(diagnostics: tuple[Diagnostic, ...]) -> SearchQueryError:
|
||||||
|
errors = [_single_diagnostic_to_error(d) for d in diagnostics]
|
||||||
|
return errors[0] if len(errors) == 1 else MultipleSearchQueryErrors(errors)
|
||||||
|
|
||||||
|
|
||||||
|
def _single_diagnostic_to_error(d: Diagnostic) -> SearchQueryError:
|
||||||
|
if d.kind is DiagnosticKind.BAD_DATE:
|
||||||
|
return InvalidDateQuery(d.field, d.raw_value)
|
||||||
|
if d.kind is DiagnosticKind.BAD_NUMBER:
|
||||||
|
return InvalidNumberQuery(d.field, d.raw_value)
|
||||||
|
return SearchQueryError(d.message)
|
||||||
|
```
|
||||||
|
|
||||||
|
No `except Exception: query_str = raw_query` fallback — per the skill, that
|
||||||
|
legacy defensive branch is explicitly not carried forward. A bug in the new
|
||||||
|
path must surface as a real error, not silently degrade to stale behavior.
|
||||||
|
|
||||||
|
`views.py`'s existing `except SearchQueryError as e: raise
|
||||||
|
ValidationError({"query": [str(e)]}) from e` handler gets one added branch
|
||||||
|
to surface every aggregated message instead of just one:
|
||||||
|
|
||||||
|
```python
|
||||||
|
except SearchQueryError as e:
|
||||||
|
messages = [str(sub) for sub in e.errors] if isinstance(e, MultipleSearchQueryErrors) else [str(e)]
|
||||||
|
raise ValidationError({"query": messages}) from e
|
||||||
|
```
|
||||||
|
|
||||||
|
`d.field`/`d.raw_value` depend on the whoosh-compat prerequisite change
|
||||||
|
(issue #1) landing first; until then (or if `field`/`raw_value` are `None`
|
||||||
|
for a given diagnostic kind not yet covered), `_single_diagnostic_to_error`
|
||||||
|
falls back to `SearchQueryError(d.message)`.
|
||||||
|
|
||||||
|
Deferred, explicitly out of scope for this PR stack: any frontend use of
|
||||||
|
`startchar`/`endchar` (already present on `Diagnostic` today) to highlight
|
||||||
|
the offending span in the search box. Backend-only for now, per explicit
|
||||||
|
decision.
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
Existing test inventory (`src/documents/tests/search/` and
|
||||||
|
`test_api_search.py`):
|
||||||
|
|
||||||
|
| File | Fate |
|
||||||
|
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||||
|
| `test_translate.py` | Deleted (PR 4) — subject deleted |
|
||||||
|
| `test_query.py`: `TestCreatedDateField`, `TestDateTimeFields`, `TestWhooshQueryRewriting`, `TestYearRangeRewriting`, `TestNonDateFieldsNotRewritten`, `TestPassthrough`, `TestNormalizeQuery` | Deleted (PR 4) — test `_translate.py`/`_dates.py` internals or intermediate query strings |
|
||||||
|
| `test_query.py`: `TestParseUserQuery` | Reviewed at plan time; result-level assertions folded into the new acceptance module, internals-only assertions dropped |
|
||||||
|
| `test_query.py`: `TestParseSimpleTextHighlightQuery`, `TestPermissionFilter` | Unchanged — never touched `translate_query` |
|
||||||
|
| `test_schema.py`, `test_tokenizer.py`, `test_backend.py`, `test_lock_backoff.py`, `test_migration_fulltext_query_field_prefixes.py` | Unchanged |
|
||||||
|
| `test_api_search.py` (`TestDocumentSearchApi`, 43 tests) | **Stays green across every PR in the stack** (hard gate, not just PR 4) — full HTTP+DB+index integration coverage catches wiring mistakes none of the narrower tests would |
|
||||||
|
|
||||||
|
New tests per PR:
|
||||||
|
|
||||||
|
- **PR 2**: `test_registry.py` — internal `*_id` names rejected; `type`/
|
||||||
|
`path` aliases resolve to canonical fields; JSON subpaths match
|
||||||
|
`docs/usage.md`; registry construction deterministic per language; the
|
||||||
|
`notes`/`custom_fields` dict-key coupling test described above.
|
||||||
|
- **PR 3**: `test_date_grammar_parity.py` — transitional, parametrized over
|
||||||
|
every keyword/unit `_dates.py`/`_translate.py` accept today
|
||||||
|
(`_DATE_KEYWORDS`, all of `_UNIT_ALIASES`'s Whoosh-era abbreviations —
|
||||||
|
`yrs`/`mos`/`wks`/`hrs`/`mins`/`secs` etc. — digit-precision forms, ISO
|
||||||
|
dash forms, `now-7d`/`now+1h`/`now-30m` compact offsets, open/reversed
|
||||||
|
ranges). Each case parses through
|
||||||
|
`wc.parse()` against a DATE-kind `FieldRegistry` and asserts no
|
||||||
|
diagnostics _and_ bounds matching what `_dates.py`/`_translate.py`
|
||||||
|
compute today, using the still-present legacy code as the oracle.
|
||||||
|
Deleted again in PR 4 along with that oracle, superseded by the
|
||||||
|
permanent acceptance corpus. This audit is scoped to _parity_ only —
|
||||||
|
whoosh-compat's date grammar is a strict superset of what `_dates.py`
|
||||||
|
accepts today (e.g. `tomorrow`, `now`, `midnight`, `noon`, weekday names
|
||||||
|
like `next monday`), so the migration also grants new date vocabulary for
|
||||||
|
free. That's a nice side effect, not something this PR needs to test or
|
||||||
|
document beyond noting it in the changelog alongside the other behavior
|
||||||
|
changes.
|
||||||
|
- **PR 4**:
|
||||||
|
- Result-level acceptance module (paperless's analogue of whoosh-compat's
|
||||||
|
`test_acceptance_e2e.py`): a real index built via `build_schema()`, the
|
||||||
|
issue #13568 queries verbatim, real saved-view strings, every date
|
||||||
|
keyword/unit, field aliases, comma lists, numeric/date ranges,
|
||||||
|
bracket-class wildcards, boosts, JSON subpaths — asserted by matched
|
||||||
|
document-ID set, `pytest.param(..., id=...)` per case. Plus a
|
||||||
|
multi-diagnostic case (two bad fields → `MultipleSearchQueryErrors`
|
||||||
|
with both messages present) and one `Multitoken` case nested inside a
|
||||||
|
top-level `OR` (proves DIVERGENCES entry 15 doesn't matter for
|
||||||
|
paperless's data, per the skill).
|
||||||
|
- `test_api_search.py` expanded: a multi-bad-field query (e.g.
|
||||||
|
`created:notadate AND asn:notanumber`) asserting the 400 response's
|
||||||
|
`query` list contains both messages; end-to-end searches on the five
|
||||||
|
newly-documented fields (`asn:`, `page_count:`, `num_notes:`,
|
||||||
|
`original_filename:`, `checksum:`) returning the right documents
|
||||||
|
through the real index.
|
||||||
|
|
||||||
|
## Dependency pinning
|
||||||
|
|
||||||
|
Stays `path = "../whoosh-compat"` in `[tool.uv.sources]` through the whole
|
||||||
|
PR stack — both repos are being actively co-developed. The final swap
|
||||||
|
happens at PR 4:
|
||||||
|
|
||||||
|
- **Primary plan**: whoosh-compat is released to PyPI around PR 3 (per
|
||||||
|
your stated intent), assuming the parity audit and issue #1 don't turn
|
||||||
|
up anything else needing a second round. PR 4 switches to a pinned PyPI
|
||||||
|
version (`whoosh-compat[tantivy]==X.Y.Z` in `dependencies`, the
|
||||||
|
`[tool.uv.sources]` override removed entirely).
|
||||||
|
- **Fallback**: if the PyPI release slips past PR 4's start, pin an exact
|
||||||
|
git commit SHA instead (`whoosh-compat[tantivy] @ git+https://github.com/
|
||||||
|
stumpylog/whoosh-compat@<sha>`), per the skill's "pre-1.0: pin an exact
|
||||||
|
version or git SHA, upgrades are deliberate" guidance.
|
||||||
|
|
||||||
|
The `TODO` comment already sitting in `pyproject.toml` (from the earlier
|
||||||
|
smoke-test setup) gets updated to reflect this — "release, else pinned SHA"
|
||||||
|
— rather than committing hard to one path before it's known which applies.
|
||||||
|
|
||||||
|
## Explicitly out of scope
|
||||||
|
|
||||||
|
- `archive_checksum` indexing/search (separate schema-level follow-up).
|
||||||
|
- Frontend consumption of `Diagnostic.startchar`/`endchar` for in-box error
|
||||||
|
highlighting (backend-only for this PR stack).
|
||||||
|
- A feature flag or shadow-compare rollout period — explicitly decided
|
||||||
|
against; safety comes from the acceptance corpus and parity audit instead.
|
||||||
|
- Any change to `build_permission_filter()`/`_apply_permission_filter()` —
|
||||||
|
confirmed untouched by this migration.
|
||||||
+1
-1
@@ -66,5 +66,5 @@
|
|||||||
"ts-node": "~10.9.1",
|
"ts-node": "~10.9.1",
|
||||||
"typescript": "^6.0.3"
|
"typescript": "^6.0.3"
|
||||||
},
|
},
|
||||||
"packageManager": "pnpm@10.26.0"
|
"packageManager": "pnpm@11.15.1"
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -5,6 +5,7 @@ trustPolicy: no-downgrade
|
|||||||
trustPolicyExclude:
|
trustPolicyExclude:
|
||||||
- "chokidar@4.0.3"
|
- "chokidar@4.0.3"
|
||||||
- "semver@6.3.1 || 5.7.2"
|
- "semver@6.3.1 || 5.7.2"
|
||||||
|
blockExoticSubdeps: true
|
||||||
allowBuilds:
|
allowBuilds:
|
||||||
"@parcel/watcher": true
|
"@parcel/watcher": true
|
||||||
canvas: true
|
canvas: true
|
||||||
|
|||||||
@@ -1047,6 +1047,12 @@ class PermittedObjectsFilter(BaseFilterBackend):
|
|||||||
perm_codename: str | None = None
|
perm_codename: str | None = None
|
||||||
|
|
||||||
def filter_queryset(self, request, queryset, view):
|
def filter_queryset(self, request, queryset, view):
|
||||||
|
# Before the superuser and owner-only paths, neither of which consults
|
||||||
|
# permitted_object_ids. Scoped to authenticated users so anonymous
|
||||||
|
# access (AnonymousUser.is_active is False) keeps its existing
|
||||||
|
# unowned-only behaviour.
|
||||||
|
if request.user.is_authenticated and not request.user.is_active:
|
||||||
|
return queryset.none()
|
||||||
if request.user.is_superuser:
|
if request.user.is_superuser:
|
||||||
return queryset
|
return queryset
|
||||||
if not self.include_granted:
|
if not self.include_granted:
|
||||||
|
|||||||
@@ -54,11 +54,15 @@ class PaperlessObjectPermissions(DjangoObjectPermissions):
|
|||||||
|
|
||||||
class PaperlessAdminPermissions(BasePermission):
|
class PaperlessAdminPermissions(BasePermission):
|
||||||
def has_permission(self, request, view):
|
def has_permission(self, request, view):
|
||||||
return request.user.is_staff
|
return request.user.is_active and request.user.is_staff
|
||||||
|
|
||||||
|
|
||||||
def has_global_statistics_permission(user: User | None) -> bool:
|
def has_global_statistics_permission(user: User | None) -> bool:
|
||||||
if user is None or not getattr(user, "is_authenticated", False):
|
if (
|
||||||
|
user is None
|
||||||
|
or not getattr(user, "is_active", False)
|
||||||
|
or not getattr(user, "is_authenticated", False)
|
||||||
|
):
|
||||||
return False
|
return False
|
||||||
|
|
||||||
return getattr(user, "is_superuser", False) or user.has_perm(
|
return getattr(user, "is_superuser", False) or user.has_perm(
|
||||||
@@ -67,7 +71,11 @@ def has_global_statistics_permission(user: User | None) -> bool:
|
|||||||
|
|
||||||
|
|
||||||
def has_system_status_permission(user: User | None) -> bool:
|
def has_system_status_permission(user: User | None) -> bool:
|
||||||
if user is None or not getattr(user, "is_authenticated", False):
|
if (
|
||||||
|
user is None
|
||||||
|
or not getattr(user, "is_active", False)
|
||||||
|
or not getattr(user, "is_authenticated", False)
|
||||||
|
):
|
||||||
return False
|
return False
|
||||||
|
|
||||||
return (
|
return (
|
||||||
@@ -188,6 +196,13 @@ def permitted_object_ids(
|
|||||||
if user is None or not getattr(user, "is_authenticated", False):
|
if user is None or not getattr(user, "is_authenticated", False):
|
||||||
return base_qs.filter(owner__isnull=True).values_list("id", flat=True)
|
return base_qs.filter(owner__isnull=True).values_list("id", flat=True)
|
||||||
|
|
||||||
|
# Deactivated users get nothing, deactivated superusers included, so this
|
||||||
|
# has to come before the superuser shortcut. guardian's
|
||||||
|
# ObjectPermissionChecker denies inactive users, but get_objects_for_user
|
||||||
|
# (the pattern this replaces) does not, so it would not be inherited.
|
||||||
|
if not getattr(user, "is_active", False):
|
||||||
|
return base_qs.none().values_list("id", flat=True)
|
||||||
|
|
||||||
if getattr(user, "is_superuser", False):
|
if getattr(user, "is_superuser", False):
|
||||||
return base_qs.values_list("id", flat=True)
|
return base_qs.values_list("id", flat=True)
|
||||||
|
|
||||||
|
|||||||
@@ -496,6 +496,28 @@ class TestPermittedObjectIdsGenericModels:
|
|||||||
expected_hidden=[strangers.pk],
|
expected_hidden=[strangers.pk],
|
||||||
)
|
)
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("is_superuser", [False, True])
|
||||||
|
def test_inactive_user_sees_nothing(self, model, factory, perm, is_superuser):
|
||||||
|
suffix = f"{model.__name__}_{is_superuser}"
|
||||||
|
user = User.objects.create_user(
|
||||||
|
username=f"inactive_{suffix}",
|
||||||
|
is_active=False,
|
||||||
|
is_superuser=is_superuser,
|
||||||
|
)
|
||||||
|
other = User.objects.create_user(username=f"other_{suffix}")
|
||||||
|
granted = factory(owner=other)
|
||||||
|
assign_perm(perm, user, granted)
|
||||||
|
|
||||||
|
assert_visible_document_ids(
|
||||||
|
permitted_object_ids(user, model, perm),
|
||||||
|
expected_visible=[],
|
||||||
|
expected_hidden=[
|
||||||
|
factory(owner=None).pk,
|
||||||
|
factory(owner=user).pk,
|
||||||
|
granted.pk,
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
def test_unowned_object_visible_to_everyone(self, model, factory, perm):
|
def test_unowned_object_visible_to_everyone(self, model, factory, perm):
|
||||||
user = User.objects.create_user(username=f"user_{model.__name__}")
|
user = User.objects.create_user(username=f"user_{model.__name__}")
|
||||||
unowned = factory(owner=None)
|
unowned = factory(owner=None)
|
||||||
|
|||||||
@@ -68,3 +68,44 @@ class TestPermittedObjectsFilter:
|
|||||||
visible_ids = set(result.values_list("id", flat=True))
|
visible_ids = set(result.values_list("id", flat=True))
|
||||||
assert visible_ids == {owned.pk}
|
assert visible_ids == {owned.pk}
|
||||||
assert granted.pk not in visible_ids
|
assert granted.pk not in visible_ids
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("username", "is_superuser"),
|
||||||
|
[("inactive", False), ("inactive_super", True)],
|
||||||
|
)
|
||||||
|
def test_inactive_user_sees_nothing(self, username: str, *, is_superuser: bool):
|
||||||
|
user = User.objects.create_user(
|
||||||
|
username=username,
|
||||||
|
is_active=False,
|
||||||
|
is_superuser=is_superuser,
|
||||||
|
)
|
||||||
|
TagFactory(owner=None)
|
||||||
|
TagFactory(owner=user)
|
||||||
|
granted = TagFactory(owner=User.objects.create_user(username=f"o_{username}"))
|
||||||
|
assign_perm("view_tag", user, granted)
|
||||||
|
request = APIRequestFactory().get("/")
|
||||||
|
request.user = user
|
||||||
|
|
||||||
|
result = PermittedObjectsFilter().filter_queryset(
|
||||||
|
request,
|
||||||
|
Tag.objects.all(),
|
||||||
|
_DummyView(),
|
||||||
|
)
|
||||||
|
assert result.count() == 0
|
||||||
|
|
||||||
|
def test_inactive_user_sees_nothing_with_include_granted_false(self):
|
||||||
|
user = User.objects.create_user(username="inactive_owner", is_active=False)
|
||||||
|
TagFactory(owner=user)
|
||||||
|
TagFactory(owner=None)
|
||||||
|
request = APIRequestFactory().get("/")
|
||||||
|
request.user = user
|
||||||
|
|
||||||
|
class _OwnerOnlyFilter(PermittedObjectsFilter):
|
||||||
|
include_granted = False
|
||||||
|
|
||||||
|
result = _OwnerOnlyFilter().filter_queryset(
|
||||||
|
request,
|
||||||
|
Tag.objects.all(),
|
||||||
|
_DummyView(),
|
||||||
|
)
|
||||||
|
assert result.count() == 0
|
||||||
|
|||||||
@@ -2,7 +2,7 @@ msgid ""
|
|||||||
msgstr ""
|
msgstr ""
|
||||||
"Project-Id-Version: paperless-ngx\n"
|
"Project-Id-Version: paperless-ngx\n"
|
||||||
"Report-Msgid-Bugs-To: \n"
|
"Report-Msgid-Bugs-To: \n"
|
||||||
"POT-Creation-Date: 2026-08-08 14:28+0000\n"
|
"POT-Creation-Date: 2026-08-10 02:25+0000\n"
|
||||||
"PO-Revision-Date: 2022-02-17 04:17\n"
|
"PO-Revision-Date: 2022-02-17 04:17\n"
|
||||||
"Last-Translator: \n"
|
"Last-Translator: \n"
|
||||||
"Language-Team: English\n"
|
"Language-Team: English\n"
|
||||||
@@ -53,7 +53,7 @@ msgstr ""
|
|||||||
msgid "Maximum nesting depth exceeded."
|
msgid "Maximum nesting depth exceeded."
|
||||||
msgstr ""
|
msgstr ""
|
||||||
|
|
||||||
#: documents/filters.py:1073
|
#: documents/filters.py:1079
|
||||||
msgid "Custom field not found"
|
msgid "Custom field not found"
|
||||||
msgstr ""
|
msgstr ""
|
||||||
|
|
||||||
|
|||||||
@@ -19,7 +19,10 @@ class AutoLoginMiddleware(MiddlewareMixin):
|
|||||||
if request.path.startswith("/api/token/") and request.method == "POST":
|
if request.path.startswith("/api/token/") and request.method == "POST":
|
||||||
return None
|
return None
|
||||||
try:
|
try:
|
||||||
request.user = User.objects.get(username=settings.AUTO_LOGIN_USERNAME)
|
request.user = User.objects.get(
|
||||||
|
username=settings.AUTO_LOGIN_USERNAME,
|
||||||
|
is_active=True,
|
||||||
|
)
|
||||||
auth.login(
|
auth.login(
|
||||||
request=request,
|
request=request,
|
||||||
user=request.user,
|
user=request.user,
|
||||||
|
|||||||
@@ -0,0 +1,53 @@
|
|||||||
|
from django.contrib.auth.models import AnonymousUser
|
||||||
|
from django.contrib.auth.models import User
|
||||||
|
from django.test import RequestFactory
|
||||||
|
from django.test import TestCase
|
||||||
|
from django.test import override_settings
|
||||||
|
|
||||||
|
from paperless.auth import AutoLoginMiddleware
|
||||||
|
|
||||||
|
|
||||||
|
@override_settings(AUTO_LOGIN_USERNAME="autologin")
|
||||||
|
class TestAutoLoginMiddleware(TestCase):
|
||||||
|
def setUp(self) -> None:
|
||||||
|
super().setUp()
|
||||||
|
self.factory = RequestFactory()
|
||||||
|
self.middleware = AutoLoginMiddleware(lambda request: None)
|
||||||
|
|
||||||
|
def _process(self, request):
|
||||||
|
# login() needs a session to write to
|
||||||
|
request.session = self.client.session
|
||||||
|
self.middleware.process_request(request)
|
||||||
|
return request
|
||||||
|
|
||||||
|
def test_active_user_is_logged_in(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- AUTO_LOGIN_USERNAME names an active user
|
||||||
|
WHEN:
|
||||||
|
- A request is processed by the middleware
|
||||||
|
THEN:
|
||||||
|
- That user is attached to the request
|
||||||
|
"""
|
||||||
|
user = User.objects.create_user(username="autologin")
|
||||||
|
|
||||||
|
request = self._process(self.factory.get("/"))
|
||||||
|
|
||||||
|
self.assertEqual(request.user, user)
|
||||||
|
|
||||||
|
def test_deactivated_user_is_not_logged_in(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- AUTO_LOGIN_USERNAME names a user who has been deactivated
|
||||||
|
WHEN:
|
||||||
|
- A request is processed by the middleware
|
||||||
|
THEN:
|
||||||
|
- The request is left anonymous rather than authenticated as them
|
||||||
|
"""
|
||||||
|
User.objects.create_user(username="autologin", is_active=False)
|
||||||
|
|
||||||
|
request = self.factory.get("/")
|
||||||
|
request.user = AnonymousUser()
|
||||||
|
self._process(request)
|
||||||
|
|
||||||
|
self.assertFalse(request.user.is_authenticated)
|
||||||
@@ -1298,16 +1298,16 @@ wheels = [
|
|||||||
|
|
||||||
[[package]]
|
[[package]]
|
||||||
name = "fpdf2"
|
name = "fpdf2"
|
||||||
version = "2.8.7"
|
version = "2.8.8"
|
||||||
source = { registry = "https://pypi.org/simple" }
|
source = { registry = "https://pypi.org/simple" }
|
||||||
dependencies = [
|
dependencies = [
|
||||||
{ name = "defusedxml", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "defusedxml", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "fonttools", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "fonttools", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "pillow", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "pillow", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
]
|
]
|
||||||
sdist = { url = "https://files.pythonhosted.org/packages/27/f2/72feae0b2827ed38013e4307b14f95bf0b3d124adfef4d38a7d57533f7be/fpdf2-2.8.7.tar.gz", hash = "sha256:7060ccee5a9c7ab0a271fb765a36a23639f83ef8996c34e3d46af0a17ede57f9", size = 362351, upload-time = "2026-02-28T05:39:16.456Z" }
|
sdist = { url = "https://files.pythonhosted.org/packages/1e/bc/8fd4321aed40cadadddc8f311c65b6082346b252bca048f7b476d8f35d72/fpdf2-2.8.8.tar.gz", hash = "sha256:9e94e155e85e8053329a9a1fce8b566fd7a7c5bb79e98a1a3952d379b947c5b9", size = 374689, upload-time = "2026-08-09T23:32:45.334Z" }
|
||||||
wheels = [
|
wheels = [
|
||||||
{ url = "https://files.pythonhosted.org/packages/66/0a/cf50ecffa1e3747ed9380a3adfc829259f1f86b3fdbd9e505af789003141/fpdf2-2.8.7-py3-none-any.whl", hash = "sha256:d391fc508a3ce02fc43a577c830cda4fe6f37646f2d143d489839940932fbc19", size = 327056, upload-time = "2026-02-28T05:39:14.619Z" },
|
{ url = "https://files.pythonhosted.org/packages/f5/be/af012eda9507494f28b99b077423806c43a11573eb6225dd46f19ae2d263/fpdf2-2.8.8-py3-none-any.whl", hash = "sha256:3557a478fc577a929c94aace9666aed4dcc432b5ab6764232e6a59f1ccd75f17", size = 337000, upload-time = "2026-08-09T23:32:43.728Z" },
|
||||||
]
|
]
|
||||||
|
|
||||||
[[package]]
|
[[package]]
|
||||||
@@ -2927,8 +2927,8 @@ dependencies = [
|
|||||||
{ name = "sqlite-vec", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "sqlite-vec", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "tantivy", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "tantivy", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "tika-client", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "tika-client", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "torch", version = "2.13.0", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "sys_platform == 'darwin'" },
|
{ name = "torch", version = "2.13.0", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "torch", version = "2.13.0+cpu", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "sys_platform == 'linux'" },
|
{ name = "torch", version = "2.13.0+cpu", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "watchfiles", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "watchfiles", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "whitenoise", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "whitenoise", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "zxing-cpp", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "zxing-cpp", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
@@ -4511,8 +4511,8 @@ dependencies = [
|
|||||||
{ name = "numpy", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "numpy", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "scikit-learn", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "scikit-learn", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "scipy", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "scipy", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "torch", version = "2.13.0", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "sys_platform == 'darwin'" },
|
{ name = "torch", version = "2.13.0", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "torch", version = "2.13.0+cpu", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "sys_platform == 'linux'" },
|
{ name = "torch", version = "2.13.0+cpu", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "tqdm", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "tqdm", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "transformers", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "transformers", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
{ name = "typing-extensions", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
{ name = "typing-extensions", marker = "sys_platform == 'darwin' or sys_platform == 'linux'" },
|
||||||
@@ -4957,18 +4957,17 @@ name = "torch"
|
|||||||
version = "2.13.0"
|
version = "2.13.0"
|
||||||
source = { registry = "https://download.pytorch.org/whl/cpu" }
|
source = { registry = "https://download.pytorch.org/whl/cpu" }
|
||||||
resolution-markers = [
|
resolution-markers = [
|
||||||
"python_full_version >= '3.15' and sys_platform == 'darwin'",
|
|
||||||
"python_full_version >= '3.12' and python_full_version < '3.15' and sys_platform == 'darwin'",
|
"python_full_version >= '3.12' and python_full_version < '3.15' and sys_platform == 'darwin'",
|
||||||
"python_full_version < '3.12' and sys_platform == 'darwin'",
|
"python_full_version < '3.12' and sys_platform == 'darwin'",
|
||||||
]
|
]
|
||||||
dependencies = [
|
dependencies = [
|
||||||
{ name = "filelock", marker = "sys_platform == 'darwin'" },
|
{ name = "filelock", marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "fsspec", marker = "sys_platform == 'darwin'" },
|
{ name = "fsspec", marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "jinja2", marker = "sys_platform == 'darwin'" },
|
{ name = "jinja2", marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "networkx", marker = "sys_platform == 'darwin'" },
|
{ name = "networkx", marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "setuptools", marker = "sys_platform == 'darwin'" },
|
{ name = "setuptools", marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "sympy", marker = "sys_platform == 'darwin'" },
|
{ name = "sympy", marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
{ name = "typing-extensions", marker = "sys_platform == 'darwin'" },
|
{ name = "typing-extensions", marker = "python_full_version < '3.15' and sys_platform == 'darwin'" },
|
||||||
]
|
]
|
||||||
wheels = [
|
wheels = [
|
||||||
{ url = "https://download-r2.pytorch.org/whl/cpu/torch-2.13.0-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:e76f9bcecc52b8ff711239a2f7547d5353df95878ab232f0773c1d95928b92f8", upload-time = "2026-07-08T12:26:13Z" },
|
{ url = "https://download-r2.pytorch.org/whl/cpu/torch-2.13.0-cp311-cp311-macosx_14_0_arm64.whl", hash = "sha256:e76f9bcecc52b8ff711239a2f7547d5353df95878ab232f0773c1d95928b92f8", upload-time = "2026-07-08T12:26:13Z" },
|
||||||
@@ -4983,6 +4982,7 @@ name = "torch"
|
|||||||
version = "2.13.0+cpu"
|
version = "2.13.0+cpu"
|
||||||
source = { registry = "https://download.pytorch.org/whl/cpu" }
|
source = { registry = "https://download.pytorch.org/whl/cpu" }
|
||||||
resolution-markers = [
|
resolution-markers = [
|
||||||
|
"python_full_version >= '3.15' and sys_platform == 'darwin'",
|
||||||
"python_full_version == '3.12.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
|
"python_full_version == '3.12.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
|
||||||
"python_full_version == '3.12.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
|
"python_full_version == '3.12.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
|
||||||
"python_full_version >= '3.15' and sys_platform == 'linux'",
|
"python_full_version >= '3.15' and sys_platform == 'linux'",
|
||||||
@@ -4990,13 +4990,13 @@ resolution-markers = [
|
|||||||
"python_full_version < '3.12' and sys_platform == 'linux'",
|
"python_full_version < '3.12' and sys_platform == 'linux'",
|
||||||
]
|
]
|
||||||
dependencies = [
|
dependencies = [
|
||||||
{ name = "filelock", marker = "sys_platform == 'linux'" },
|
{ name = "filelock", marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "fsspec", marker = "sys_platform == 'linux'" },
|
{ name = "fsspec", marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "jinja2", marker = "sys_platform == 'linux'" },
|
{ name = "jinja2", marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "networkx", marker = "sys_platform == 'linux'" },
|
{ name = "networkx", marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "setuptools", marker = "sys_platform == 'linux'" },
|
{ name = "setuptools", marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "sympy", marker = "sys_platform == 'linux'" },
|
{ name = "sympy", marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
{ name = "typing-extensions", marker = "sys_platform == 'linux'" },
|
{ name = "typing-extensions", marker = "(python_full_version >= '3.15' and sys_platform == 'darwin') or sys_platform == 'linux'" },
|
||||||
]
|
]
|
||||||
wheels = [
|
wheels = [
|
||||||
{ url = "https://download-r2.pytorch.org/whl/cpu/torch-2.13.0%2Bcpu-cp311-cp311-linux_s390x.whl", hash = "sha256:6e9817dbdf5ea76789babd46e457eac5bf14ff566cf85f8addbfdff2d56601ce", upload-time = "2026-07-08T19:27:52Z" },
|
{ url = "https://download-r2.pytorch.org/whl/cpu/torch-2.13.0%2Bcpu-cp311-cp311-linux_s390x.whl", hash = "sha256:6e9817dbdf5ea76789babd46e457eac5bf14ff566cf85f8addbfdff2d56601ce", upload-time = "2026-07-08T19:27:52Z" },
|
||||||
|
|||||||
Reference in New Issue
Block a user