Compare commits

...
Author SHA1 Message Date
stumpylogandClaude Opus 5 a65efd3e4c docs(search): state the wildcard rule the alternatives actually implement
The wildcard caveat still described the pre-alternatives world: "where
stemming shortens a word, a pattern that reaches past the point it was cut
off matches nothing". Measured against a real index that is false for most
of the set it ranges over. university*, companie*, happiness* and
universities* all match, because the stem of the typed run is offered as a
second branch. Only a fragment landing strictly between the stem and the
surface form fails, which is what both of the examples happened to be, so
the sentence read as convincing. A user who believed it would type a
shorter fragment, which genuinely finds nothing, instead of the full word,
which works.

The rule stated now is the measured one: a trailing * matches a stored term
when either the run as typed or its stem is a prefix of that term. Both
branches are load-bearing, verified one word per document. copy* reaches
"copies" only via the stem and "copyright" only via the surface form, so a
rule naming just the stem would predict the wrong answer for half of it.

Also corrects two more claims that were each defensible alone and wrong
together, both measured before rewriting:

- checksum patterns are lowercased even though checksum terms are stored
  verbatim, so checksum:9F86D081* does match; "matched exactly as typed and
  nothing else" told the reader otherwise. Pinned as a positive case beside
  the uppercase *term* that really does fail.
- relative offsets and long-form dates must be quoted only as a value
  standing alone; added:[-1 week to now] works unquoted and returns the same
  documents as the quoted spelling.

test_pattern_stemming.py's partial-prefix docstring was fabricated: it said
"librar" stems to the longer "librari" and so matches nothing on its own.
Measured, the stemmer leaves "librar" alone, "univers" stems to the shorter
"univ", and librari* does match this file's own fixture. Rewritten against
measured values, and renamed, since a future editor reasoning from it would
have believed these params exercise the two-alternative path. They do not,
so a regression breaking that path would have left them green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 16:37:38 -07:00
stumpylogandClaude Opus 5 8f371e39ae docs(test): state what the library now asserts about the 8-digit date form
The docstring named test_compact_numeric_datetime as covering the 8-digit
calendar-day form, but that test asserted only the lower bound, so the
day-window property this file's deleted test_eight_digits_is_a_calendar_day_window
used to pin was asserted in neither repo. whoosh-compat b672741 pins both
bounds; this says so accurately rather than implying coverage that did not
exist. "A library test covers this spelling" and "a library test asserts the
property my deleted test asserted" are different claims, and only the second
licenses a deletion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 16:06:36 -07:00
stumpylogandClaude Opus 5 caf87136f6 test(search): collapse the date survivors onto one representative
Follow-up to 98bd31f7d, applying the stricter reading of "one representative
case per integration seam". The trimmed unit-abbreviation and reversed-range
survivors reached the same seams as test_compact_date_forms.py by different
spellings: no mutation kills either without also killing the compact case,
and the compact case additionally fails when queries are resolved in a fixed
non-local timezone, which neither of the other two catches because now-5h and
now+/-1h are timezone-invariant. Verified before deleting: the compact
survivor fails under all three mutations (fixed non-local timezone, `added`
registered as TEXT instead of DATETIME, `added` indexed at day precision).

The spellings themselves are asserted by the library, per unit in
tests/test_relative_date_unit_abbreviations.py and for both range kinds in
tests/test_parser_dates.py. The reversed-range file's other purpose, to invert
as the signal that the library's bound swap landed, is spent: it inverted.

Also give TestTagCommaList a docstring saying what it is. It is a docs-contract
regression test for a published spelling, not proof of paperless's field
configuration: removing comma_values from the tag FieldSpec leaves it passing,
because the analyzer splits the literal value into the same tokens. The
registry fact is owned by test_registry.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 15:49:31 -07:00
stumpylogandClaude Opus 5 cf1cce5f9d test(search): keep one case per integration seam, not one per grammar spelling
whoosh-compat now owns the query grammar and asserts it directly, so the
parallel grammar corpus this tree accumulated re-proves the library rather
than paperless. Each surviving case was checked by mutation: break the seam
it claims (index `added` at day precision, resolve queries in a fixed
non-local timezone, tokenize on whitespace instead of separators, index the
metadata counters off by one) and the case fails.

- delete test_dash_prefix_negation.py: ported wholesale to the library, and
  the analyzer half it also covered is held by test_documented_syntax.py's
  documented leading-hyphen case, which the whitespace-tokenizer mutation
  fails.
- delete test_comma_value_lists.py: end to end the value-list and literal
  readings of a comma agree, because the analyzer splits the literal value
  on the comma anyway. Declaring comma_values on correspondent, or removing
  it from tag, left every case in that file passing, so it proved the
  grammar and not the registry. The registry fact moves to test_registry.py,
  where the same mutation does fail.
- trim the unit-abbreviation, compact-date and reversed-range files to one
  representative each.
- drop the second quoted zero-width keyword from test_documented_syntax.py;
  which keyword sits inside the quotes is library grammar.

Docstrings that justified a file by the absence of library coverage are
rewritten: that premise expired when the library absorbed the assertions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 15:43:45 -07:00
stumpylogandClaude Opus 5 c56f910cc9 docs(search): scope the relative-offset warning to a value standing alone
The warning said relative offsets such as "-1 week" "in practice mean no
documents at all", unqualified, one bullet after the range-bound rule
offered added:['-1 week' to now] as the example of quoting a bound. Both
sentences were true in their own scope, but read together they talk a
user out of a query that works: the bare value is a zero-width instant,
the same offset as a bound is a real seven-day window (verified:
lo=2026-06-08T12:00, hi=2026-06-15T12:00 from a frozen 2026-06-15T12:00).

Also widens the last sentence to say now-3days and "3 days ago" are
rejected wherever they appear, having checked they are rejected as range
bounds too, single- or double-quoted, not only as bare values.

Restores the 'added:"-1 week"' case to the standalone-forms list, which
the docs name by that exact spelling, and pins the bound reading beside
it so neither half of the warning is an unpinned claim.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 15:19:26 -07:00
stumpylogandClaude Opus 5 5ecd237a3d feat(search)!: let a wildcard match the typed run or its stem
whoosh-compat's pattern_normalizer now accepts several alternative forms
per literal run, ORed and deduplicated by the emitter, so the
shorter-of-the-two heuristic that had to pick one form is gone. The typed
run and its stem are both offered: neither is a prefix of the other once
the stemmer substitutes rather than truncates ("copy" -> "copi"), so
"copy*" now reaches "copies" and "copyright" alike instead of trading one
for the other.

checksum keeps _fold_normalizer. It is the only KEYWORD field, indexed
with the raw tokenizer, and a stemmed prefix there ("ceded" -> "cede")
returns documents whose checksum does not start with what was typed.

Also picks up two date-grammar fixes from the same library release: a
reversed relative range now swaps its bounds like the absolute case
instead of day-bumping the upper one, and a date value the grammar can
only half-consume (a bare, unquoted "added:2005-03-04T15:30:00Z") is
rejected as an InvalidDateQuery rather than silently matching nothing.

docs/usage.md: the "copy* does not find copyright" caveat is no longer
true; the bare timestamp is now an error rather than a silent non-match;
and the range-bracket quoting rule was wrong in a user-visible way. Only
double-quoted bounds are rejected, single-quoted ones parse.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 15:02:33 -07:00
stumpylogandClaude Opus 5 d17232043b docs: correct three stale pointers in the search comments and admin docs
_map_emit_error's docstring pointed at the QueryParserError arm of views.py's
generic handler, which was deleted; QueryParserError no longer appears in
production code at all. _REGEX_TIMEOUT justified itself as ReDoS protection,
but the two pre-parse rewrites it was written for are gone and its one
remaining use is a character class that cannot backtrack. The timeout stays,
its rationale is corrected. administration.md listed two triggers for
--if-needed; the schema fingerprint is now a third.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:14:31 -07:00
stumpylogandClaude Opus 5 0d609074dd fix(search): stop logging an operator ERROR for field:* on a JSON field
_map_emit_error routed every MISCONFIGURED diagnostic to an ERROR log, so the
six user-typeable spellings of a JSON existence search (notes:*, notes.note:*,
notes.user:*, custom_fields:*, custom_fields.name:*, custom_fields.value:*)
each wrote one permanent ERROR line per request, and any authenticated user
could generate them in a loop.

Nothing is misconfigured. whoosh-compat decides EXISTS_REQUIRES_FAST from the
registry's own FieldSpec (kind plus fast) without consulting the index schema,
and field_descriptors() builds the JSON fields non-fast on purpose, so no
operator action can clear the condition. It is ordinary user error and now
gets the 400 with no alert. SCHEMA_FIELD_MISSING, the other MISCONFIGURED
kind, really is a registry-versus-schema comparison and keeps the ERROR.

The 400 and its user-facing message are unchanged. No log deduplication is
introduced; the classification is what was wrong, not the logging policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:11:31 -07:00
stumpylogandClaude Opus 5 82546f13b3 fix(search): keep wildcard patterns literal on unstemmed KEYWORD fields
get_field_registry() branched the analyzer by field kind but gave every field
the same stemming pattern normalizer. checksum is indexed with the raw
tokenizer, so its terms are neither folded nor stemmed, yet its wildcard
patterns were: "checksum:ceded*" normalized to "cede*" and matched a document
whose checksum starts with "cedef00d". About 2.8% of random hex prefixes were
rewritten this way. Always over-matching rather than missing, but for a field
whose whole purpose is exact identification, returning a different checksum is
a wrong answer.

KEYWORD fields now get a fold-and-lower normalizer, which is what every field
used before pattern stemming was added; TEXT fields keep the stemming one so
"invoice*" still reaches the indexed "invoic".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 12:08:47 -07:00
stumpylogandClaude Opus 5 0b080d9415 test(search): restore the comma clause-separator and compact date forms
Two properties were asserted before the whoosh-compat migration and nowhere
after it, though both still hold.

A comma directly before another field name separates two clauses rather than
delimiting a value list (the old TestNormalizeQuery and _translate.py's
TestCommaResolution clause-separator cases). The new class pins both halves
against a corpus that separates the readings: on tag (the one comma_values
field) the value-list reading demands a tag literally named
"added:2005-03-04" and matches nothing -- measured tag:"foo,added:2005-03-04"
-> [] against the same corpus where tag:foo,added:2005-03-04 -> [both] -- and
on title the literal reading matches nothing either, title:"Alpha,tag:foo" ->
[] against title:Alpha,tag:foo -> [both, tag_only]. Both assertions are exact
sets, so neither alternative reading survives.

The compact separator-free date spellings get their own file. 8 digits is a
calendar-day window and 14 digits a single instant, so the corpus carries a
same-day/other-hour document and a next-day/same-hour one: a form degrading
into the other width, or into a non-match, fails rather than passing on the
one document that matches either way. Measured: added:20050304 -> [instant,
same_day] (identical to added:2005-03-04), added:20050304153000 -> [instant],
added:20050304090000 -> [same_day].

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:58:40 -07:00
stumpylogandClaude Opus 5 7d61c3769f test(search): assert the empty highlight query matches nothing
test_empty_query_returns_empty_query and test_all_operators_returns_empty_query
asserted isinstance(result, tantivy.Query), which parse_simple_text_highlight_query
cannot fail to satisfy: it either returns a Query or raises. Replacing its
`return tantivy.Query.empty_query()` with `all_query()` left both green, so
the contract they were named for -- highlight nothing, rather than highlight
every document -- was unpinned.

They now count hits against an index holding one document. A third test
asserts a real token matches that document, so an empty corpus cannot make
the other two pass for a match-everything query.

Under the empty_query -> all_query mutation both new tests fail; before this
change the mutation killed nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:56:42 -07:00
stumpylogandClaude Opus 5 42c4f648a4 test(search): delete the alias tests that cannot fail
TestFieldAliases in test_documented_syntax.py pinned type:/path: against a
single indexed document with no decoy, so neither test could tell alias
resolution from the demotion that happens without it. Stripping `aliases`
from every FieldSpec and clearing the registry cache left both green:

- type:invoice demotes to unfielded text, and document_type is a default
  search field, so the token still matched the typed document.
- path:archive demotes and matched via the title: the fixture title was
  "Pathed", and with SEARCH_LANGUAGE=en that stems to "path".

test_acceptance.py::TestFieldAliases covers the same syntax at the same
result level with content decoys chosen to make demotion visible, and dies
on that mutation. Deleted rather than given decoys of their own, so the
property has one home instead of two.

Under the alias-strip mutation the suite now fails 4 tests: both
test_acceptance.py::TestFieldAliases cases and both
test_registry.py resolution cases.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:54:14 -07:00
stumpylogandClaude Opus 5 6ab4c3d689 fix(search): cap the global search query too
GlobalSearchView calls the backend directly rather than through the shared
helper the cap lives in, so "every query string is length-checked" was a
claim with an exception rather than an invariant.

It hardcodes SearchMode.TEXT, which is linear rather than quadratic, so
this path was never the CPU-exhaustion vector and this is not a fix for
one. It is capped so the invariant holds without a footnote: the view
already bounds the query from below, and a later change letting it select
a search mode would otherwise reopen the hole with nothing to catch it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:37:28 -07:00
stumpylogandClaude Opus 5 885bc2fdf3 fix(search): cap query length at the shared search-param helper (F3)
whoosh-compat's fieldname tagger is O(n^2) in plain word characters,
reachable only through SearchMode.QUERY's whoosh grammar. Measured
against the real field registry: ~1s at 10k chars, ~3.7s at 20k, ~14.4s
at 40k. The POST selection-filter path (bulk edit, bulk download) has
no server-imposed length bound the way the GET path incidentally does
via header limits, making an unbounded query a single-request CPU
exhaustion vector.

Cap both entry points at their shared choke point,
_get_tantivy_query_and_mode, with a new QueryTooLongError that reuses
the existing SearchQueryError -> 400 routing both callers already
have. 4096 chars bounds the worst case to roughly 0.16s by quadratic
extrapolation, far beyond any plausible hand-typed query. TEXT and
TITLE modes route through simple_search_tokens instead and measure
linear even at 20k chars, so the same cap is hygiene for them rather
than a fix. Hardcoded rather than a PAPERLESS_* setting: this is a
security boundary, and a raisable ceiling could reintroduce the exact
hazard it exists to close.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:30:30 -07:00
stumpylogandClaude Opus 5 7cb6b32a8f chore(deps): re-lock with CI's uv, restore two unrelated downgrades
The lock carried ~490 redundant `sys_platform == 'darwin' or sys_platform
== 'linux'` markers and had quietly pinned sqlparse to 0.5.5 and
pymdown-extensions to 11.0, both older than dev and neither related to
search. Re-resolving with uv 0.12.x (CI's pinned series) drops the markers
and restores both to dev's versions, taking the lock's diff against dev
from 1428 lines to 51. The only version line that now differs from dev is
whoosh-compat's own.

Three versions move here, not two: whoosh-compat also goes 0.1.0.dev0 ->
0.1.0, picked up from the sibling repo's own bump through the local path
dependency rather than from anything this re-lock decided.

A plain `uv lock` is a no-op here, since the lock is already
self-consistent; the markers are only recomputed when the resolution
actually re-runs.

Also drops a TODO's pointer to a spec file deleted in e50421542, and
corrects a comment claiming custom fields have a companion text field for
full-text search. They do not: no such field is ever written, and their
values are reachable only through the JSON field. Records that notes_text
is absent from _DEFAULT_SEARCH_FIELDS, so it is highlight-only by design.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 11:20:52 -07:00
stumpylog 86d90539ea fix(search): guard the TEXT-mode highlight query against a 500
parse_simple_text_highlight_query re-parsed simple-search tokens through
Tantivy's query-string parser without quoting, so any token carrying
Tantivy grammar (a bare quote, a colon, brackets, a slash) raised an
unguarded ValueError once the search itself had already matched a
document. With DocumentViewSet.list's blanket exception handler narrowed
earlier on this branch, that ValueError now reaches the client as a bare
500, not the 400 it used to be -- confirmed against the real endpoint
before this change.

Quote and escape each token as its own phrase before parsing so ordinary
punctuation in a plain-text search no longer trips the grammar parser,
and keeps producing real highlight snippets instead of none. Still guard
the call with a narrow ValueError catch (matching the sibling notes_text
guard's shape, not its broader Exception catch) as defense in depth for
inputs quoting alone cannot save, falling back to the query that already
matched.
2026-08-20 10:57:48 -07:00
stumpylogandClaude Opus 5 f9398caf4b test(search): restore the generic internal-id-field guard
The prune replaced a generic "no PUBLIC_FIELDS name ends in _id" invariant
with a fixed list of the seven names that were dropped. That list catches
the seven; nothing catches the eighth.

Dropping write-only *_id fields from the query surface is the whole point
of the schema change earlier in this branch, so the generic form is what
guards the class against recurrence. Both tests coexist: the list pins
that specific names stay unregistered, this pins that no new one leaks in.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 10:44:03 -07:00
stumpylogandClaude Opus 5 dbb0b8217b test(search): prune trivial and superseded search tests (Task 10 Prune list)
- test_fields.py: reduced to test_json_fields_have_subpaths - the rest
  asserted properties of PUBLIC_FIELDS' 16-line literal tuple, already
  covered behaviourally by test_registry.py's resolve()-based tests.
- test_registry.py::TestJsonSubpathCoupling: deleted - its own docstring
  admitted it hardcodes both sides of the comparison it claims to guard.
  test_acceptance.py::TestJsonSubpaths already proves the coupling
  against a real index, and the new test_json_subpath_completeness.py
  proves it exhaustively for every declared subpath.
- test_query.py: dropped the asn/checksum isinstance-only parse checks,
  now duplicated by result-level matches in test_documented_syntax.py
  and test_api_search.py.
- conftest.py: dropped the module-scoped `index` fixture, dead since
  test_translate.py was deleted.

419 passed (documents/tests/search/ + test_api_search.py +
test_api_search_errors.py), down from 432 before this prune.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 10:32:19 -07:00
stumpylogandClaude Opus 5 2d2dad0e1a test(search): pin currently-unasserted behaviours (Task 10 Add list)
Adds dedicated result-level tests, each indexing real documents and
asserting on matched-ID sets rather than parse shape:

- Deferred `-term` negation (G1): bare `-taxes` requires the term, and
  fielded `-title:alpha` drops the negation entirely, both matching v2.
- Reversed date ranges: absolute bounds swap (whoosh parity); relative
  `now±` bounds day-bump instead, an inconsistency pinned as a
  whoosh-compat follow-up rather than "fixed" locally.
- The six whoosh unit abbreviations (yrs/mos/wks/hrs/mins/secs).
- Unterminated `[` date range brackets 400 at the API level.
- `tag:foo,bar` comma value lists are conjunctive, with a decoy proving
  `correspondent:foo,bar` is not treated as a list.
- Date keyword phrases (`today`) honour the active (non-UTC) timezone,
  not just relative ranges.
- `_DEFAULT_SEARCH_FIELDS` stays a subset of PUBLIC_FIELDS.
- Every declared JSON subpath is actually written at index time.

Also replaces test_schema.py::TestFastFlagAgreement's tantivy-py
__reduce__() pickling probe with one built on the paperless-owned,
tantivy-independent field_descriptors().

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 10:27:05 -07:00
stumpylogandClaude Opus 5 b3e3d8b23e docs(search): correct the RFC3339 timestamp claim, it works when quoted
The previous commit's warning said a timestamp carrying a time of day was
not understood at all, and told users to fall back to whole-day range
bounds. Both halves were wrong. Only the bare unquoted spelling fails:

    added:2005-01-01T00:00:00Z                            -> no match
    added:"2005-01-01T00:00:00Z"                          -> matches
    added:[2005-01-01T00:00:00Z to 2006-01-01T00:00:00Z]  -> matches

That is the ordinary quoting rule the surrounding docs already state, the
same one "-1 week" and "next monday" obey, so present the timestamp as a
working form rather than as a limitation and drop the false workaround:
range bounds carrying a time of day work fine.

A bound must be bare inside range brackets, where quoting it is rejected
outright, so document both halves of the rule rather than just "quote it".

Keep the zero-width warning distinct from the quoting rule now sitting above
it, since a reader who just learned quoting rescues "next monday" would
otherwise assume it rescues "-3 days". It does not: re-verified against
documents added at exactly those instants, the quoted offsets still match
only that one instant.

Pin the working spellings, which is the assertion that was missing: nothing
covered the quoted or range-bound forms, so a regression of a working
feature went undetected. Also pin that quoting does not rescue the
zero-width group.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 10:06:12 -07:00
stumpylogandClaude Opus 5 059759b83f docs(search): describe the query grammar paperless actually supports
usage.md linked to tantivy's QueryParser documentation, promising a grammar
paperless neither implements nor intends to. Replace the link with a
description of the surface that was verified end-to-end against a real index,
and correct the claims that did not survive that verification.

- checksum: the field is stored verbatim, so only a complete lowercase
  checksum matches. The old a1b2c3d4 example matched nothing.
- A leading - is not negation. Separators are stripped at index time, so
  "invoice -secret" requires "secret", the opposite of the intent. Document
  NOT as the way to exclude a term.
- Document the aliases type: and path:, num_notes:, numeric ranges, quoted
  phrases, tag:'s comma list (which requires all listed tags, not any), and
  the date forms that resolve to a real span: tomorrow, ISO dates, month
  names, "next monday"/"last monday".
- Warn about now/noon/midnight and offsets like "-3 days": they parse, but
  resolve to a single instant rather than a span, so they match nothing.
  Likewise a T/Z timestamp, whose time portion is split off as loose text.

Add test_documented_syntax.py, which asserts on matched document IDs rather
than on parsed queries, so the docs cannot drift from the code again. Its
negative cases pin the behaviours the warnings describe.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 09:53:19 -07:00
stumpylogandClaude Opus 5 48ff9a8218 docs(search): point the field-table comments at field_descriptors()
Both comments still described the pre-fingerprint layout: _fields.py said
the internal-only fields "stay hardcoded in build_schema()", and the
fast-flag test said build_schema() honors the flag only in its U64 and
DATE branches. Both now live in field_descriptors(), and _fields.py's
header is the one thing a future editor reads before touching the field
table, so a stale pointer there is the expensive kind.

Comments only; no executable line is touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 09:44:10 -07:00
stumpylogandClaude Opus 5 44886aabd1 feat(search): detect schema shape changes with a schema fingerprint
build_schema() was half table-driven and half hardcoded, so editing it for
parser reasons could change the on-disk field list without anyone bumping
SCHEMA_VERSION. tantivy compares schemas by ordered field list, so such an
edit leaves reads working while every write raises.

Complete the table: build_schema() now iterates an explicit list of field
descriptors covering id, the PUBLIC_FIELDS expansion, the sort shadow,
bigram, simple_* and autocomplete fields and the permission columns. The
same list is hashed into a schema_fingerprint() that is stamped into
.index_settings.json and compared by needs_rebuild() as a fourth check
alongside the existing schema version and language checks.

The fingerprint is computed from paperless' own descriptors rather than
tantivy's schema representation, so a tantivy-py option-key rename or
addition cannot silently force a global reindex.

The emitted schema is byte-identical to the previous one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 09:36:02 -07:00
stumpylogandClaude Opus 5 9554390a08 fix(search): neutralize tantivy's boolean keywords in the fuzzy words
The fuzzy clause's word string is cut to \w+ runs so no query grammar
reaches index.parse_query, but tantivy's boolean keywords are themselves
word runs. Under analyzed=True the field analyzer lowercased them into
ordinary terms before they got that far; now that the words are raw
query text, an uppercase keyword out of a quoted phrase arrives as
grammar: '"tax AND reports"' quietly made the clause a conjunction,
'"tax NOT reports"' gave it its own exclusion, and '"tax AND"' (or IN
anywhere) failed the parse and cost the query its fuzzy clause outright.

Lowercase exactly AND/OR/NOT/IN, which is what the analyzer used to do
and is the only spelling tantivy reads as grammar ("And" is a term).
Nothing else is touched: tantivy already lowercases query terms with the
field's analyzer, and doing it ourselves first is not the same operation
for every input (Python folds a final sigma differently, and turns 'İ'
into a sequence tantivy then splits in two), which would search for
terms the index does not contain.

Also pins two behaviours that were reasoned about but untested: the
fielded-CJK test now runs with the fuzzy clause on as well, where the
clause's documented unfielded contribution does bring the other document
back, and the negation tests pin the CJK over-admission for an exclusion
under an Or, which cannot be hoisted without dropping the other branch's
documents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 09:17:57 -07:00
stumpylogandClaude Opus 5 a678c6ff82 fix(search): stop analyzing the fuzzy clause's words twice
_try_parse_fuzzy_query collected free_text_tokens with the default
analyzed=True and handed the analyzer's output back to
index.parse_query, which analyzes it again. Analysis is not idempotent:
'universities' stems to 'univers', and re-stemming that yields 'univ', a
term the index does not contain. prefix=True hid the mistake as
over-broad matching rather than as no matches at all, which is why no
test caught it: searching 'universities' also returned documents whose
only relevant word was 'univalent' or 'unicycle'.

Collect the raw text instead. Raw text has not been tokenized, so the
whole-token \w+ filter that keeps tantivy query grammar out of the
re-parse would now reject ordinary input outright: 'COVID-19',
'hello@example.com' and the phrase "tax reports" each arrive as a single
token containing punctuation, and a query made only of such terms would
lose its fuzzy clause entirely. Cut each token into its word runs and
keep those, which recovers the terms and keeps the guarantee the filter
exists for: only word characters ever reach the parser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 08:52:02 -07:00
stumpylogandClaude Opus 5 352312e97d fix(search): apply the query's exclusions to the whole blended query
parse_user_query ORs three top-level Should clauses: the exact query, an
optional fuzzy blend and an optional CJK bigram clause. The latter two
are built from positive terms only and cannot express an exclusion, so
each one re-admitted precisely the documents the exact clause had
excluded: 'invoice NOT secret' returned the secret document as soon as
ADVANCED_FUZZY_SEARCH_THRESHOLD was set, and '東京 NOT secret' returned
it unconditionally, since nothing gates the CJK clause.

Building the CJK clause from the AST does not fix this: there the
excluded term is not the CJK one, so the clause legitimately contains
東京 and still matches the document.

Hoist the exclusions instead. _ConjunctiveNegations walks the parsed
tree for the subtrees that constrain every matching document, and each
is emitted as an ordinary positive query attached with MustNot above the
Must-ed blend. Or is not descended into: in 'invoice OR NOT secret' the
negation is one branch's condition, and hoisting it would drop documents
the other branch matches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 08:48:04 -07:00
stumpylogandClaude Opus 5 4e2d71513a fix(search): build the CJK clause from the parsed AST, not the raw query
_build_cjk_query scanned the raw query string for CJK runs, so a CJK term
the user negated ('invoice NOT 漢字') or restricted to one field
('title:漢字', 'notes:漢字') came straight back as a top-level Should
clause over every bigram field. The fuzzy clause already collects its
words from the parsed tree for exactly this reason; the CJK clause a few
lines below did not.

Collect the CJK runs from whoosh_compat's free_text_tokens over
result.ast instead, one default field at a time so the tokens keep their
field attribution: a bare term (already copied onto every default field
by the parser) still searches every bigram field, while title:東京
reaches bigram_title alone, and a term on a non-default field
contributes nothing. Fields sharing identical CJK text share one parse.

The raw-string builder stays for the simple TEXT/TITLE modes, whose
input is plain text with no query grammar to respect, as does
extract_cjk_text, which the indexing side calls per bigram field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 08:44:47 -07:00
stumpylogandClaude Opus 5 7fddad849a refactor(search): delete the date-keyword-phrase pre-parse rewrite
whoosh-compat's grammar already accepts the closed multi-word date
keyword vocabulary (previous month, this year, etc.) unquoted after a
date field, making _quote_date_keyword_phrases redundant. Like its
sibling rewrite removed in an earlier commit, it was not quote-aware
and could insert quotes mid-phrase inside an unrelated quoted string
(e.g. title:"see added:previous month notes"), corrupting the parse.
Deleting it removes that hazard entirely.

Docs are adjusted to scope the quoted-or-unquoted equivalence to the
documented keyword list; other date expressions the grammar accepts
(relative offsets, absolute dates) still require quoting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 08:27:07 -07:00
stumpylogandClaude Sonnet 5 0fb36dae43 fix(search): delete regex bare-JSON-prefix rewrite, use default subpaths
The regex rewrite that turned bare notes:/custom_fields: prefixes into
their subpath spelling was blind to quoting: content:"payment notes:
none" was silently rewritten mid-phrase into a notes-field search and
matched zero documents. whoosh-compat's FieldSpec now supports a
default subpath per JSON field (SubpathSpec(default=True)), which
resolves during parsing where quoting is already understood, so the
pre-parse string rewrite is no longer needed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-20 08:13:29 -07:00
stumpylogandClaude Opus 5 95b600ebdd docs(search): correct the recall claim in the pattern normalizer
The docstring said a shorter prefix "only widens recall". That holds for a
stem that truncates, not for one that substitutes: English y -> i moves the
pattern sideways, so "copy*" gains "copies" and loses "copyright". Stating
it as a general invariant is what hid that class in the first place.

Length stays the rule; only its justification is corrected. No behavior
change -- no executable line is touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 08:02:20 -07:00
stumpylogandClaude Opus 5 e48182c358 docs(search): pin the stem-substitution limit wildcards inherit
Stemming substitutes as well as truncates ("copy" and "copies" both index as
"copi" while "copyright" keeps its literal y), so a stemmed pattern reaches a
word's inflections but no longer reaches compounds that keep the surface
spelling. That trade is accepted: the same substitution is what makes company*
and library* work, and no rule over one normalized string separates them. So
usage.md stops claiming a trailing star just works, and a test pins copy* to the
base word rather than the compound. Also adds a parity test tying
stem_pattern_text to paperless_text_analyzer's own output, so a filter added to
the index analyzer alone cannot silently diverge, and corrects the docstring
claim that a run can analyze to several tokens - the raw tokenizer emits one
token whatever the input, so only the remove_long zero-token case can fire.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 08:01:02 -07:00
stumpylogandClaude Opus 5 e311c84139 fix(search): stem wildcard patterns so prefix searches match again
Index terms are stemmed but query patterns were not, so invoice* matched nothing
while invoic* worked. v2's index was unstemmed (whoosh TEXT() defaults to
StandardAnalyzer), so this regressed against both baselines, not just dev. Uses
the typed run's stem unless the stem is longer than the run, since a stem can be
longer than a partial prefix and a shorter prefix only widens recall. Patterns
spanning the stem boundary (produ*name) still cannot match a stemmed index, so
usage.md loses that example rather than advertising a broken one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 07:37:26 -07:00
stumpylogandClaude Opus 5 3f6af15f7d revert: log every search misconfiguration, not one per field
This reverts ea883f416, which suppressed repeat MISCONFIGURED logs to
once per field per process.

A misconfigured field is a static condition an operator can fix in one
change, so the repetition is the prompt to fix it rather than noise to
suppress, and it stops on its own once the schema is corrected. Keeping
the suppression meant carrying machinery whose key boundedness and
check-then-add race both had to be reasoned about, to solve a problem
that ends when someone fixes the config.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 07:23:15 -07:00
stumpylogandClaude Opus 5 4de3711940 fix(search): let library-internal search errors surface as 500s, not 400s
DocumentViewSet.list's trailing except Exception clause was catching the
re-raised QueryParserError/INTERNAL-cause QueryError that _map_emit_error
and whoosh-compat's own parse() deliberately let escape, and turning them
into a generic 400 -- exactly the outcome that routing exists to prevent.
Remove the catch-all (and the now-redundant QueryParserError re-raise it
made pointless) so a whoosh-compat library defect surfaces as a
monitorable 500 instead of blaming the user for it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 07:20:55 -07:00
stumpylogandClaude Opus 5 12667d9754 fix(search): log a search misconfiguration once per field, not per request
EXISTS_REQUIRES_FAST is MISCONFIGURED and reachable from ordinary query text
(notes.user:*), so the operator alert added with the Cause routing fired on
every such request. An alert that repeats on every user query is one operators
learn to filter out, which defeats routing MISCONFIGURED to an operator at all.

The condition is a static configuration fact: it stays true until an operator
changes the schema and reindexes, so the first log carries the same information
as the ten-thousandth. Deduped on (kind, field) in a per-process set; a restart
re-logs, re-surfacing the condition after a config change. The 400 is not
deduped: every request still gets its response and its message.

The key is bounded by the registry, not by query text. emit() only reports
MISCONFIGURED for a field it resolved, and FieldRegistry.resolve returns None
for any name or JSON subpath the registry does not declare.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 06:56:54 -07:00
stumpylogandClaude Opus 5 63de40c54e fix(search): route emit diagnostics by Cause, own the user-facing wording
The except QueryError arm converted every kind to a 400 on the strength of a
comment asserting the INTERNAL kinds could not occur. SCHEMA_FIELD_MISSING
fires on registry/schema drift, which deriving both from PUBLIC_FIELDS newly
makes possible, so a defect in our own wiring was reported to the user as a bad
query and never reached monitoring.

Diagnostics now route on Cause: INVALID_INPUT/UNSUPPORTED are a 400,
MISCONFIGURED is logged at error level naming the field and then a 400 (the
registry and the schema disagree, which only an operator can fix, but a request
is still waiting and the query cannot run either way), and INTERNAL is
re-raised rather than converted.

Messages, parse-time as well as emit-time, are built from the Diagnostic's
structured fields; d.message is documented as unstable developer output and
PATTERN_TOO_COMPLEX embedded raw backend error text in the 400 body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 06:52:21 -07:00
stumpylog 48b36d4f16 fixup! refactor(search): derive build_schema() from shared PUBLIC_FIELDS table 2026-08-20 06:23:59 -07:00
stumpylogandClaude Opus 5 414b26a374 chore: gitignore agent workflow scratch
.superpowers/ holds per-plan ledgers, task briefs and review diffs. Untracked
scratch in the tree is how unrelated files get swept into commits, and in a
sibling repo the same directory was silently pulled into an sdist by the build
backend's default include.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 06:16:22 -07:00
Trenton HolmesandClaude Sonnet 5 bb0d0fc6c0 fix(search): migrate off whoosh-compat's removed QueryEmitError/UnsupportedQueryError
whoosh-compat replaced both exception classes with a single QueryError
carrying a structured Diagnostic (kind/cause/field_kind); the old
message-text regex stripping is now dead weight since the library no
longer embeds host-facing wording (DIVERGENCES refs, fast=True advice)
in Diagnostic.message. Branch on diagnostic.kind instead.

Also trims a test that was re-asserting whoosh-compat's own message
contract (now covered by its own test_kind_matrix.py) down to the one
rewrite paperless still owns: EXISTS_REQUIRES_FAST's user-facing message.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:36:53 -07:00
Trenton Holmes 5abd568209 docs(search): document the quote-blindness trade-off in the pre-parse rewrites
_quote_date_keyword_phrases and _rewrite_bare_json_field_prefixes both
regex-match anywhere in raw_query, with no awareness of whether the match
falls inside an already-quoted phrase on an unrelated field. Unlikely in
practice and not fixed (quote-aware scanning is real work for an edge
case), but now called out explicitly like this file's other accepted
trade-offs, instead of being the one undocumented one.
2026-08-19 13:36:53 -07:00
Trenton Holmes 1fc5a8a8a3 refactor(search): log the CJK clause's skip path like the fuzzy clause's
_build_cjk_query silently swallowed a parse failure with no log line,
while _try_parse_fuzzy_query logs at debug for the same "skip this
optional clause" situation. Add the matching debug log.

Deliberately NOT narrowing except Exception to except ValueError here to
match the fuzzy path: the fuzzy blend's word string is pre-filtered to
\\w+-only tokens before it ever reaches index.parse_query, so ValueError
is the only realistic failure mode there. cjk_text has no equivalent
filter, so narrowing this catch without verifying tantivy's actual
exception behavior for CJK input would risk letting something other than
ValueError propagate uncaught - the same class of mistake as the fuzzy
blend regression this migration already fixed once, in the other
direction.
2026-08-19 13:36:53 -07:00
Trenton Holmes ddf8287072 refactor(search): split error classes and build_permission_filter out of _query.py
_query.py mixed three unrelated responsibilities: the SearchQueryError
family (paperless's public error-surface API, re-exported by __init__.py),
the actual query rewrite/parse/emit/blend pipeline, and
build_permission_filter, which has nothing to do with query parsing and
is consumed only by _backend.py.

- New _errors.py: SearchQueryError, InvalidDateQuery, InvalidNumberQuery,
  MultipleSearchQueryErrors, search_query_error_messages. _query.py now
  imports these instead of defining them.
- build_permission_filter moves to _backend.py, next to its one caller
  (TantivyBackend._build_permission_filter).
- __init__.py re-exports the error classes from _errors.py instead of
  _query.py; the package's public API (documents.search import ...) is
  unchanged for every caller going through it (views.py etc.).

_query.py now reads top-to-bottom as rewrite -> parse -> emit -> blend,
matching what parse_user_query's own docstring already claimed the file
was.
2026-08-19 13:36:53 -07:00
Trenton Holmes bf7eec7168 refactor(search): underscore-prefix and Final-type the module-private field lists
DEFAULT_SEARCH_FIELDS/SIMPLE_SEARCH_FIELDS/TITLE_SEARCH_FIELDS looked
public but are only ever used inside _query.py itself, sitting next to
underscore-prefixed constants at the same scope (_CJK_ALL_FIELDS etc.).
Rename to match, and add Final like their neighbors already have.
2026-08-19 13:36:53 -07:00
Trenton Holmes 6721651e82 test(search): dedupe next(f for f in PUBLIC_FIELDS...) lookups, collapse table tests
Both test_fields.py and test_registry.py repeated the same generator-next
lookup by field name. Add a module-level {name: field} dict in each and
use it instead.

Also collapse test_fields.py's five single-attribute tests
(document_type/storage_path aliases, tag's comma_values, notes/
custom_fields subpaths) into one parametrized test_field_attributes -
they were really one table-consistency check split into five copies of
the same three-line shape.
2026-08-19 13:36:53 -07:00
Trenton Holmes 9da2bb35cd test(api): dedupe the four archive-metadata search tests via a helper
test_search_by_asn/page_count/original_filename/checksum were all
create-doc -> index -> GET -> assert 200 and doc.id in results, repeated
verbatim four times. Extract _assert_query_finds() so each test states
only its distinguishing field and query.
2026-08-19 13:36:53 -07:00
Trenton Holmes 31b0ef1f25 test(search): remove redundant local imports in TestSearchQueryErrors
InvalidDateQuery/InvalidNumberQuery/MultipleSearchQueryErrors/
SearchQueryError are all already imported at module top; three test
bodies re-imported them locally for no reason.
2026-08-19 13:36:53 -07:00
Trenton Holmes f2f4d04ab6 test(search): hoist deferred imports, add _index() helper in test_acceptance.py
User/DocumentType/StoragePath were imported inside individual test bodies
despite the module already importing documents.models at top level -
nothing here needed deferred import. Also add an _index() helper
(Document.objects.create + backend.add_or_update in one call) for the many
sites where nothing needs to happen between creating a document and
indexing it; the two-step ceremony was outweighing the fixture data at
every call site. Left as two explicit steps wherever a Note or
CustomFieldInstance genuinely has to be attached before indexing.
2026-08-19 13:36:53 -07:00
Trenton Holmes e535e3b859 refactor(views): merge split local-import block in _get_search_document_ids
get_backend was imported alone, three statements ran, then
SearchQueryError/search_query_error_messages were imported separately -
one function's imports split across two blocks with code between them.
Merge into the single existing local-import block.
2026-08-19 13:36:53 -07:00
Trenton Holmes 8cf4b0c997 refactor(search): make PUBLIC_FIELDS a tuple[FieldSpec, ...], drop PublicField
PublicField duplicated seven fields whoosh-compat's own FieldSpec already
has (name/kind/aliases/comma_values/date_only/fast/subpaths), and
_registry.py hand-copied all of them across on every registry build.
FieldSpec is a frozen dataclass with analyzer/pattern_normalizer already
optional (default None), so PUBLIC_FIELDS can just BE the FieldSpec tuple -
_schema.py only ever read name/kind/fast off it and needs no changes.
_registry.py now attaches the per-language analyzer/pattern_normalizer via
dataclasses.replace() instead of reconstructing every field from scratch.

FieldSpec.__post_init__ normalizes subpaths into a MappingProxyType, so
test_fields.py's exact-tuple-equality subpath assertions become set
comparisons; a genuinely empty subpaths is now `not field.subpaths` rather
than `== ()`.

Verified test_api_trash.py::test_api_trash's "Schema error: An index exists
but the schema does not match" failure is a pre-existing, unrelated local
environment issue (a stale, untracked data/index/ directory in this
checkout) - reproduces identically with this commit's changes stashed out.
2026-08-19 13:36:53 -07:00
Trenton Holmes 45e6b1dcc6 refactor(search): delete the _simple_query_tokens pass-through wrapper
It called simple_search_tokens() and nothing else, with a comment
duplicating that function's own docstring. Call sites now call
simple_search_tokens() directly.
2026-08-19 13:36:53 -07:00
Trenton Holmes 170cf476d1 refactor(search): extract _any_of to collapse the single-clause boolean idiom
Four call sites in _query.py each hand-rolled "no clauses -> empty, one
clause -> return it bare, many -> wrap in boolean_query" - one of them also
handling the empty case, one written as a ternary, one returning a captured
variable instead of clauses[0][1] (same value, different spelling). Extract
_any_of() so the collapsing logic and its rationale (skip a wasted
single-clause boolean_query wrap) live in one place.
2026-08-19 13:36:53 -07:00
Trenton Holmes 622317345a test(search): unify the two Schema.__reduce__() introspection sites
_schema_field_names and TestFastFlagAgreement each independently reached
into Schema.__reduce__()[1][0] - the one fragile, version-coupled
expression this test suite depends on. Rename to _schema_fields, return
{name: field-state} instead of just names, and have both call sites use
it, so a tantivy-py upgrade that changes this shape breaks in one place.
2026-08-19 13:36:53 -07:00
Trenton Holmes 248e82cc61 test(search): promote query_index to a single module-scoped fixture
Three classes in test_query.py each defined an identical query_index
fixture. None of these tests write documents to the index, so consolidate
into one module-level, module-scoped fixture (mirroring conftest.py's
index fixture rationale) instead of three copies to keep in sync.

Deliberately NOT merged with conftest.py's own index fixture: that one
registers tokenizers with "english" (stemming on), while these tests rely
on "" (stemming off) - a real behavioral difference, not incidental.
2026-08-19 13:36:53 -07:00
Trenton Holmes e7a10fc153 test(search): dedupe the resolve-and-assert boilerplate in test_registry.py
Every test repeated the same 5-line make_ref/resolve/assert-not-None
sequence before its one real assertion. Extract a registry fixture and a
typed _resolve() helper so each test states one fact in one line.
2026-08-19 13:36:53 -07:00
Trenton Holmes 3e1e0aefe1 refactor(search): reuse extract_cjk_text in _build_cjk_query
_build_cjk_query re-derived the same CJK-run extraction extract_cjk_text
already implements, despite a docstring claiming they mirror each other.
Call it directly so the mirroring is structural, not a copy to keep in sync.
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Sonnet 5 90c8494c08 test(search): trim acceptance tests that pin whoosh-compat behavior, not ours
Deletes TestCommaValueLists, TestMultitokenInNestedOr, TestRfc3339TZDateRange,
TestCreatedTimezoneInvariance, and TestReversedDateRange: none of them
exercise any paperless-specific pre/post-processing code. Comma-list AND
semantics, multitoken resolution, RFC3339 T/Z UTC math, date-only timezone
invariance, and reversed-range disambiguation are all entirely
whoosh-compat's own grammar/semantics, already covered by its own test
suite. The comma_values flag paperless does own is still covered cheaply in
test_fields.py; the date_only flag is still covered in test_registry.py.

Also trims verbose docstrings/comments across _query.py and the surviving
acceptance tests: cuts references to whoosh-compat's internal
DIVERGENCES.md entry numbers and paperless v2/Whoosh-era implementation
history down to the user-facing behavior that actually matters, without
losing the substance.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RVj8NFy821G3YhNf68PF6X
2026-08-19 13:36:53 -07:00
Trenton Holmes 5e68318309 docs: drop dangling reference to a branch-only deleted test file
test_date_grammar_parity.py was added and deleted entirely within this
feature branch's own history; it never existed in dev. Referencing its
deletion in a docstring only makes sense while reading this branch's
intermediate commits, not once this merges - unlike PR #13010 or
whoosh-compat's DIVERGENCES.md, which are permanent, externally
verifiable references.
2026-08-19 13:36:53 -07:00
Trenton Holmes e50421542e chore: remove whoosh-compat transition planning artifacts
The design spec, implementation plan, and dev-skill for this migration
are no longer needed now that the migration is complete and merged into
this branch.
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Fable 5 ff1d3163fd test(search): harden coverage for aliases, fast flags, and date edges
Four targeted additions, no production code:

The type-alias test asserted only that a query object was built, and a
naive result-level replacement turned out equally vacuous for a subtle
reason: document_type is itself a default search field, so a broken
alias resolution demoting "type:invoice" to unfielded text STILL
matches the typed document through the field value under test. Both
alias tests (type/document_type, path/storage_path) now use
discriminating decoys carrying the query word in content, so demotion
matches the decoy and fails the exact-set assertion; the old
parse-shape test is deleted.

A new schema test pins that every PublicField.fast flag equals the
built tantivy schema's per-field fast option, in both drift directions:
whoosh-compat trusts the declared flag when resolving field:* existence
checks, and build_schema() only honors it for U64 and DATE kinds, so a
future fast=True TEXT/KEYWORD/JSON entry would otherwise make those
searches silently match nothing at query time.

Two result-level date pins restore behaviors whose assertions were lost
in the test migration: a created date matches regardless of the active
timezone (the America/New_York leg is the discriminating one: a
tz-applying implementation shifts the window past the naive-midnight
indexed value), and a reversed created:[2025 TO 2020] range still
matches its span through the joint-disambiguation swap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WMsn6DgzbvSqh1pwy66VVF
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Fable 5 efb6e4b0f1 fix(search): complete the query error surface across every endpoint
Four pieces of the same surface:

whoosh-compat's emit() documents a two-part host contract: both a parse
diagnostic and the QueryEmitError/UnsupportedQueryError pair are
user-input errors. Only the latter half was caught; QueryEmitError now
maps to SearchQueryError too. Messages pass through a cleanup that
strips the library's DIVERGENCES.md references and replaces the
fast=True host-configuration advice with user language, so no
library-internal vocabulary reaches a searching user.

The bulk selection paths (bulk edit, the legacy bulk endpoint, bulk
download) reached the backend with no SearchQueryError handler, so a
bad date or number in a selection filter raised straight to a DRF 500.
They now share the search list endpoint's exact mapping (a new
search_query_error_messages helper flattens MultipleSearchQueryErrors
in one place), returning the same 400 body for the same bad query.

QueryParserError means a whoosh-compat parser bug, not user-fixable
input, per its own contract; the list endpoint's blanket handler was
converting it to a generic 400. It now re-raises and surfaces as a 500
that monitoring can see.

All behavior is pinned test-first: bulk edit and bulk download API
tests assert 400s naming the bad value (previously unhandled
exceptions), a unit test pins the QueryEmitError mapping, three
parametrized checks assert no internal vocabulary leaks for the
unsupported query shapes, and a mocked parser-bug test asserts the 500.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WMsn6DgzbvSqh1pwy66VVF
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Fable 5 98edd75220 test(search): replace stale docs reference in id-field fold docstring
The TestUnregisteredIdFieldFoldsToLiteralText docstring pointed readers
at docs/usage.md's advanced-search note about the dropped *_id
prefixes, which a later commit removed. State the rationale directly
instead: the *_id fields were always internal index columns (v2
consumed them for permission filtering and its own criteria), and their
queryability as search syntax was an accident of whoosh resolving any
schema field name.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WMsn6DgzbvSqh1pwy66VVF
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Fable 5 bd87624b29 fix(search): rewrite bare notes:/custom_fields: prefixes to their subpaths
The v2 whoosh schema had plural notes/custom_fields TEXT fields (notes
indexed the joined note texts, custom_fields indexed joined
"name : value" strings), so "notes:foo" and "custom_fields:foo" were
valid fielded searches in released paperless and through the deleted
translation layer. On the whoosh-compat registry those names are JSON
fields addressable only via subpaths, and the bare spelling silently
demoted to an unfielded text search of the words themselves, matching
unrelated documents that merely contain "notes" or "custom".

parse_user_query now rewrites the bare prefixes live to the same
targets migration 0017 chose for the singular whoosh-era spellings:
notes: becomes notes.note: and custom_fields: becomes
custom_fields.value:, with 0017's lookbehind guard so subpath spellings
and words merely ending in the prefix are untouched. Prefix
substitution only; values ride through unchanged, and every value shape
lands in a documented outcome downstream (ranges, wildcards and exists
on JSON subpaths are typed errors, not crashes). The inherited
trade-off stands: custom_fields.value: drops the name-matching half of
v2's combined indexing, with custom_fields.name: available for it.

Acceptance tests pin the rewrite with decoy documents whose content
contains the literal prefix words, which the old demotion matched and
the fielded search must not, plus untouched-subpath controls.
docs/usage.md documents the bare prefixes as subpath shorthand.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WMsn6DgzbvSqh1pwy66VVF
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Fable 5 ff540041b5 fix(search): restore unquoted multi-word date keywords via pre-parse quoting
"added:previous month" returned HTTP 400 after the whoosh-compat
migration. The unquoted spelling was never parser-native anywhere: v2
rewrote it to explicit bracket ranges app-side before whoosh saw the
string, and the deleted translation layer consumed it itself, so users
and saved views have relied on it continuously while whoosh-compat
deliberately scopes it out of its parser (its DIVERGENCES.md entry 19)
and understands the phrases natively only as quoted values.

parse_user_query now quotes the closed six-phrase vocabulary (previous
week/month/quarter/year, this month/year) when it directly follows a
date field's colon, before parsing. Only quoting happens app-side; every
date computation stays in whoosh-compat's grammar, unlike v2's rewrite,
which computed the ranges itself. Date field names derive from
PUBLIC_FIELDS, the field name matches case-sensitively (the parser's own
field tagging is case-sensitive), the phrase case-insensitively (the
grammar accepts any case in the quoted form), and already-quoted
spellings, TEXT fields, unfielded words and bracketed ranges are
untouched.

The previously xfailed end-to-end regression test now passes as a plain
test, and a new acceptance class pins unquoted == quoted == mixed-case
result sets on a boundary fixture, no-error parsing for the whole
vocabulary across all three date fields, and that "title:previous month"
stays an ordinary text search. docs/usage.md now states the two
spellings are equivalent after a date field.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WMsn6DgzbvSqh1pwy66VVF
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Fable 5 8789ff55f5 fix(search): build the fuzzy blend from parsed free-text tokens, not the raw query
The fuzzy blend clause handed the raw query string to tantivy's own
parser, which rejects whoosh-only grammar (date keywords, whoosh ranges,
aliases needing resolution), so any mixed query silently lost its fuzzy
clause: a typo'd word beside "added:today" stopped matching the moment
the date keyword appeared, while the same typo without it still matched.
Before the whoosh-compat migration the parser received the translated
string, so fuzzy survived mixed queries.

The clause is now built from whoosh_compat.free_text_tokens over the
already-parsed AST: the query's free-text words, analyzed, deduplicated,
with negated terms excluded so a NOT'd word cannot resurface through the
fuzzy clause. The joined word string is always plain tokens, so tantivy
always parses it; a defensive word-character filter guards any future
field whose analyzer passes punctuation through, and the ValueError skip
remains as insurance. One chosen trade-off is documented in the
docstring: a term fielded on a default search field contributes its text
unfielded, widening fuzzy recall on the 0.1-boosted secondary clause.

Two result-level acceptance tests pin the behavior: the mixed
typo-plus-date-keyword query matches its document again, and a NOT'd
word does not fuzzy-resurface (shaped so the assertion genuinely fails
under a naive all-words implementation: the excluded word's document is
the only candidate hit, so score normalization cannot mask it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WMsn6DgzbvSqh1pwy66VVF
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Sonnet 5 2b90a577c7 docs: drop *_id field-removal note from usage.md
These prefixes were never documented public API (undocumented internal
fields the old KNOWN_FIELDS happened to accept), so their removal isn't a
user-facing regression worth calling out in usage.md. The behavior is still
covered by test_acceptance.py's TestUnregisteredIdFieldFoldsToLiteralText.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RVj8NFy821G3YhNf68PF6X
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Sonnet 5 c9adf8dd45 test(search): add result-level coverage for RFC3339 T/Z date-range queries
A prior commit deleted test_query.py's parametrized "doesn't raise" coverage
for this shape (created:[...T...Z TO ...] and comma-combined ranges), which
was also the only place PR #13010's T/Z backward-compat guarantee was
exercised. Nothing in paperless's suite proved the full parse_user_query() ->
tantivy Query -> matched-document pipeline still honors it after the
whoosh-compat grammar fix (commit f936143 in the whoosh-compat repo). Add
result-level acceptance cases: an in/out-of-range T/Z bracket range, PR
#13010's original comma-combined two-field shape, and an exact-boundary case
proving a Z-suffixed bound is absolute UTC, not shifted by the local search
timezone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RVj8NFy821G3YhNf68PF6X
2026-08-19 13:36:53 -07:00
Trenton Holmes 8d86db35ca refactor: minor cleanup from final whoosh-compat migration review
- Update stale test comments in test_query.py that described string
  rewriting / raw-query fallback behavior that no longer exists post
  whoosh-compat migration; rename
  test_date_rewriting_applied_before_tantivy_parse to
  test_date_keyword_resolves_without_raising to match.
- views.py: move the local MultipleSearchQueryErrors import up into the
  existing local-import block near the top of list(), consistent with
  the other documents.search imports there, instead of importing it
  again inside the except SearchQueryError clause.
- test_api_search.py: assert response.status_code explicitly before
  indexing into response.data["results"] in
  test_search_added_previous_month_excludes_next_period_start, and tie
  the xfail marker to AssertionError instead of the incidental KeyError
  that indexing a 400 response's missing "results" key produced.
2026-08-19 13:36:53 -07:00
Trenton Holmes 943e60fecc docs: clarify quoted date-keyword phrases and dropped *_id field aliases
Add a sentence to the "Supported date keywords" advanced-search section
noting that multi-word date keywords must be quoted (e.g.
added:"previous month") -- whoosh-compat requires quoting where the
unquoted form used to work. Also document that the old undocumented
*_id field aliases (tag_id, owner_id, viewer_id, correspondent_id,
document_type_id, storage_path_id, type_id, path_id) are no longer
recognized: a query using one now silently folds to a literal-text
search instead of matching the intended structured field.
2026-08-19 13:36:53 -07:00
Trenton Holmes 35917be2d4 test: assert unregistered id-field queries actually match nothing
test_unregistered_id_field_folds_to_literal_text_not_error only checked
that parse_user_query() didn't raise for a query like tag_id:5. Add a
result-level acceptance test (matching test_acceptance.py's
_matched_ids pattern, indexed against real documents) that asserts the
matched-document-ID set is genuinely empty, not just that the parse
step succeeds.
2026-08-19 13:36:53 -07:00
Trenton Holmes f5e7d309af fix: skip fuzzy search blend when raw query isn't tantivy-parseable
The fuzzy blend clause in parse_user_query() fed the raw, whoosh-syntax
query string directly to tantivy's own query parser. Since the
whoosh-compat migration, raw_query still contains whoosh grammar (date
keywords, whoosh-style ranges, bracket-class wildcards) that tantivy's
parser rejects with ValueError, which escaped parse_user_query and
turned into a generic HTTP 400 for the entire query whenever
ADVANCED_FUZZY_SEARCH_THRESHOLD was configured.

Deriving a clean plain-text-only extraction for the fuzzy clause was
ruled out: wc.parse() already expands unfielded terms into per-default-
field copies in the AST, so there's no "still unfielded" marker left to
walk without duplicating whoosh-compat's own expansion logic. Instead,
scope a narrow try/except ValueError around exactly the
index.parse_query() call and skip the fuzzy clause (logged at debug)
when it can't parse, leaving the exact/CJK clauses unaffected.
2026-08-19 13:36:53 -07:00
Trenton Holmes 5941c19fb4 docs: document asn/page_count/checksum/original_filename advanced search fields 2026-08-19 13:36:53 -07:00
Trenton Holmes f0af5efe00 refactor(search): delete _translate.py/_dates.py, superseded by whoosh-compat 2026-08-19 13:36:53 -07:00
Trenton Holmes 46e89928af test(api): add end-to-end search coverage for asn/page_count/original_filename/checksum 2026-08-19 13:36:53 -07:00
Trenton Holmes 1e7b9a966d test(search): add result-level acceptance corpus, trim internals-only test_query.py classes
Replaces test_query.py's intermediate-AST/query-string checks with a
result-level acceptance corpus that indexes real documents and asserts
matched-ID sets through parse_user_query(), covering the #13568
bracket-wildcard regression, comma value lists, field boosts, JSON subpaths,
and Multitoken-in-OR nesting. Removes TestCreatedDateField, TestDateTimeFields,
TestWhooshQueryRewriting, TestYearRangeRewriting, TestNonDateFieldsNotRewritten,
TestPassthrough, TestNormalizeQuery, and TestParseUserQuery's
test_advanced_search_queries_do_not_raise from test_query.py, since they test
translate_query/_dates.py internals or a diagnostics-free-parse guarantee
whoosh-compat's own suite already covers.
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Sonnet 5 eb2c770e39 feat(api): surface every search query error, not just the first
When parse_user_query() raises MultipleSearchQueryErrors due to multiple
field parsing failures (e.g. both an invalid date and an invalid number
in a single query), the exception handler now surfaces all error messages
in the 400 response, allowing users to fix them all in one round-trip
instead of discovering them one at a time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:36:53 -07:00
Trenton Holmes 083e72e6da Chore: remove obsolete xfail for RFC3339 T/Z date-range queries
whoosh-compat's date grammar now accepts "T" as a date/time separator
and a trailing "Z" UTC designator (paperless-ngx PR #13010
back-compat), so these advanced-search queries no longer raise.
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Sonnet 5 af0f010b79 feat(search): route parse_user_query through whoosh-compat
Rewires parse_user_query() to parse via wc.parse()/tantivy_emit() against
the shared FieldRegistry instead of the string-based translate_query()
pipeline, so diagnostics map to typed SearchQueryError subclasses
(InvalidDateQuery/InvalidNumberQuery/MultipleSearchQueryErrors) and every
bad field is reported, not just the first.

Marks three pre-existing tests xfail (2 in test_query.py, 1 in
test_api_search.py) for confirmed whoosh-compat grammar gaps found while
verifying this rewrite: unquoted multi-word date keywords (e.g.
`added:previous month`) and RFC3339 T/Z datetime range bounds no longer
parse.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Sonnet 5 70e2422735 refactor(search): move SearchQueryError family to _query.py, add InvalidNumberQuery/MultipleSearchQueryErrors
Move SearchQueryError and InvalidDateQuery from _translate.py to _query.py and
add two new exception classes: InvalidNumberQuery and MultipleSearchQueryErrors.
Update _translate.py to re-export the exceptions for backward compatibility
until the translation module is removed. Update __init__.py to export all
four exception classes from _query.py.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:36:53 -07:00
Trenton Holmes 0d2c4eb214 test(search): add transitional date-grammar parity audit against whoosh-compat 2026-08-19 13:36:53 -07:00
Trenton Holmes ab427a2c37 test(search): guard JSON subpath/dict-key coupling between _fields.py and _backend.py 2026-08-19 13:36:53 -07:00
Trenton Holmes 2c6f89bab1 feat(search): add whoosh-compat FieldRegistry construction 2026-08-19 13:36:53 -07:00
Trenton HolmesandClaude Sonnet 5 b30c7f42cb build: add whoosh-compat as a local-path dependency
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:36:53 -07:00
Trenton Holmes 87d786f3b9 refactor(search): derive build_schema() from shared PUBLIC_FIELDS table 2026-08-19 13:33:23 -07:00
Trenton HolmesandClaude Sonnet 5 38c82a674f feat(search): add shared PUBLIC_FIELDS table
Create the shared field-definition table consumed by the schema builder
(_schema.py) and the whoosh-compat field registry (_registry.py). This
eliminates drift between what the index exposes and what queries can address.

- Create PublicField frozen dataclass with field metadata
- Define PUBLIC_FIELDS tuple with 16 searchable fields
- Add comprehensive test suite covering field properties

The whoosh-compat pyproject.toml dependency addition is added in a
follow-up commit, with the correct [tantivy] extra, source comment, and a
matching uv.lock update.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:33:23 -07:00
stumpylog a1705ff185 Updates after reviewing and updating compat lirary 2026-08-19 13:33:23 -07:00
stumpylog d92286c76c docs: flag where the transition spec and plan describe a moved API
The library changed after these were written and more changes are already
decided upstream. Records what is wrong today, what to write toward, the
one question still open, and the fast-JSON-field trap, rather than
silently leaving code that would fail on contact.
2026-08-19 13:33:23 -07:00
stumpylog 7df20c86c8 chore: update transition guidance for the current whoosh-compat API
Field references became a typed value rather than a dotted string, so
diagnostics carry one too and the registry exposes a single resolver.
Also records that the JSON fields must stay non-fast while existence
checks against a fast JSON field return inverted results.
2026-08-19 13:33:23 -07:00
stumpylogandClaude Fable 5 2328fefc29 chore: build typed search errors from structured diagnostic data
Diagnostics now carry field and raw_value, so the transition guidance
points at those instead of parsing human-readable message text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 13:33:23 -07:00
stumpylogandClaude Sonnet 5 074d905238 docs: add whoosh-compat transition implementation plan
16 bite-sized, TDD tasks across the design spec's 4-PR stack, each with
a suggested subagent type/model for delegated execution. Test/fixture
code in the acceptance-corpus and API-expansion tasks was verified
against the real codebase (documents/tests/search/conftest.py's
existing backend/index fixtures, test_backend.py's pytestmark
convention, CustomFieldInstance's typed value_text field) rather than
guessed, and the date-grammar parity test's AST-shape assumption was
confirmed by actually running whoosh_compat.parse() against a real
DATE FieldRegistry.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:33:23 -07:00
stumpylogandClaude Sonnet 5 9a4b1b00d5 docs: fold agent-review findings into whoosh-compat transition spec
Agent review (source-verified against both repos) confirmed the spec's
claims accurate throughout, with one real gap: the JSON-subpath
tantivy-py carve-out (index.parse_query fallback for notes.*/
custom_fields.* until tantivy-py#716 ships) interacts with paperless's
pinned tantivy~=0.26.0 and wasn't mentioned. Also added two footnotes:
FieldRegistry forces date_only=True on any DATE spec regardless of the
PublicField default, and the date-grammar parity audit implicitly
grants new keyword vocabulary as a side effect.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:33:23 -07:00
stumpylogandClaude Sonnet 5 273154ed15 docs: add whoosh-compat transition design spec
Design for replacing _translate.py/_dates.py with whoosh-compat: shared
field-definition table driving both the Tantivy schema and the query
FieldRegistry, diagnostics->exception mapping (aggregating all errors,
not just the first), a 4-PR stack with no rollout flag, and a
result-level acceptance corpus + date-grammar parity audit as the
safety net instead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 13:33:23 -07:00
stumpylogandClaude Fable 5 4beea6464e chore: add whoosh-compat transition skill
Encodes the settled integration decisions for replacing the
hand-maintained search translation layer with whoosh-compat:
user-typed query surface policy, analyzer seam, diagnostics-before-emit
contract, mandatory date parity audit, test churn, and rollout plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 13:33:23 -07:00
48 changed files with 5592 additions and 2568 deletions
+3
View File
@@ -115,3 +115,6 @@ celerybeat-schedule*
# Git worktree local folder
.worktrees
# Agent workflow scratch (ledgers, briefs, review packages)
.superpowers/
+2 -1
View File
@@ -521,7 +521,8 @@ Pass `--recreate` to wipe the existing index before rebuilding. Use this when th
index is corrupted or you want a fully clean rebuild.
Pass `--if-needed` to skip the rebuild if the index is already up to date (schema
version and search language match). Safe to run on every startup or upgrade.
version, schema fingerprint and search language all match). Safe to run on every
startup or upgrade.
Specify `optimize` to optimize the index. This command is regularly invoked by the
task scheduler.
+90 -4
View File
@@ -886,6 +886,19 @@ Matching documents with logical expressions:
```
shopname AND (product1 OR product2)
invoice NOT draft
```
`AND`, `OR` and `NOT` must be written in capitals, and parentheses group sub-expressions. Terms written next to each other with no operator between them are combined with `AND`.
!!! warning
A leading `-` does **not** exclude a term. Separators are stripped during indexing, so `invoice -secret` searches for `invoice` and `secret`, which is the opposite of what you probably intended. Use `NOT` to exclude a term: `invoice NOT secret`.
Matching an exact phrase, in order, by quoting it:
```
"quick brown fox"
```
Matching specific tags, correspondents or types:
@@ -893,8 +906,12 @@ Matching specific tags, correspondents or types:
```
type:invoice tag:unpaid
correspondent:university certificate
tag:bills,unpaid
```
- `document_type` may be abbreviated to `type`, and `storage_path` to `path`.
- A comma-separated list after `tag:` requires **all** of the listed tags, so `tag:bills,unpaid` matches only documents tagged both `bills` and `unpaid`.
Matching dates:
```
@@ -903,14 +920,58 @@ added:yesterday
modified:today
```
Matching by archive metadata:
```
asn:100
page_count:12
num_notes:0
checksum:9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08
original_filename:invoice.pdf
```
- `asn` matches a document's Archive Serial Number.
- `page_count` matches a document's page count.
- `num_notes` matches how many notes a document has.
- `checksum` matches the checksum of the original document file (not the archived/processed version). Unlike the text fields, this one is stored verbatim rather than tokenized, so only a complete, lowercase checksum matches. To search by the first few characters instead, use a wildcard: `checksum:9f86d081*`. Wildcard patterns on the text fields are also tried stemmed, to line up with the stemmed index, but `checksum` is indexed without stemming, so its patterns are not stemmed either: a wildcard prefix is matched literally, apart from being lowercased first. `checksum:9F86D081*` therefore does find the document, even though the plain uppercase term does not.
- `original_filename` matches the filename of the document as originally consumed.
`asn`, `page_count` and `num_notes` are numeric and also accept ranges, for example `asn:[50 to 150]`.
Matching inexact words:
```
produ*name
invoice*
title:Invoice*
```
Wildcards are matched against the _stemmed_ terms stored in the index, not
against the words as they appear in the document. Each literal part of a
pattern is tried both as you typed it and in its stemmed form, so a trailing
`*` matches a word and its inflections (`invoice*` finds "invoice", "invoices"
and "invoiced") as well as longer words whose stored term still begins with
what you typed (`copy*` finds "copyright" alongside "copy" and "copies").
It is still not a plain prefix search over the original text. A trailing `*`
matches a stored term when either the run you typed or its stemmed form is a
prefix of that term, so a fragment that stops part-way between the two matches
neither: `universities*` finds "university" and "universities", which are both
stored as `univers`, while the shorter `universit*` finds nothing at all. For
the same reason `happine*` does not find "happiness", which is stored as
`happi`. And a pattern that requires letters after the wildcard which stemming
has removed cannot match either: `productname` is stored as `productnam`, so
`produ*name` finds nothing.
Matching natural date keywords:
The multi-word date keywords listed below work quoted or unquoted after a
date field (`added:"previous month"` and `added:previous month` are
equivalent); elsewhere in a query the same words are treated as ordinary
search text. Other date expressions the parser accepts (relative offsets
like `-1 week`, or specific dates like `12 december 2019`) must be quoted when
they stand alone as a value; inside a range's brackets they work unquoted, as
in `added:[-1 week to now]`.
```
added:today
modified:yesterday
@@ -923,6 +984,30 @@ Supported date keywords: `today`, `yesterday`, `previous week`,
`this month`, `previous month`, `this year`, `previous year`,
`previous quarter`.
These other date forms also work after a date field:
```
added:tomorrow
created:2005-03-04
added:january
modified:"next monday"
added:"last monday"
added:"2005-01-01T00:00:00Z"
created:[2005-01-01 to 2005-01-31]
added:[2005-06-15T09:00:00Z to 2005-06-15T17:00:00Z]
```
- `tomorrow`, like `today` and `yesterday`, covers that whole day.
- An ISO date such as `2005-03-04` covers that whole day, and `2005-01` covers that whole month.
- A month name such as `january` covers that whole month in the current year.
- `next <weekday>` and `last <weekday>` each cover that whole day and must be quoted. A bare weekday name such as `monday` is not accepted.
- A full timestamp such as `2005-01-01T00:00:00Z` matches that exact instant. Like the other expressions above, it has to be quoted when it stands on its own: `added:"2005-01-01T00:00:00Z"`. The unquoted spelling is rejected with an error rather than searched, because only part of it can be read as a date.
- A range takes two of the above as its bounds, for example `created:[2005 to 2009]` or `added:[2005-01-01 to 2005-01-31]`. Bounds may carry a time of day. A bound is normally written without quotes; if you do quote one, use single quotes (`added:['-1 week' to now]`), because a double-quoted bound is rejected with an error.
!!! warning
As a value on its own, `now`, `noon`, `midnight` and relative offsets such as `"-3 days"` or `"-1 week"` are accepted by the parser but resolve to a single instant rather than to a span of time, so they match only a document whose timestamp is exactly that instant, which in practice means no documents at all. Quoting does not change this. As a *range bound* they are the opposite of a trap and are what you want: `added:['-1 week' to now]` covers the whole of the last seven days. Spellings like `now-3days` and `"3 days ago"` are rejected outright wherever they appear.
#### Searching custom fields
Custom field names and values are included in the full-text index, but they
@@ -938,6 +1023,7 @@ custom_fields.name:Insurance custom_fields.value:policy
- `custom_fields.value` matches against the value of any custom field.
- `custom_fields.name` matches the name of the field (use quotes for multi-word names).
- Combine both to find documents where a specific named field contains a specific value.
- The bare `custom_fields:` prefix is shorthand for `custom_fields.value:`.
Because separators are stripped during indexing, individual parts of formatted
codes are searchable on their own. A value stored as `A-1312/99.50` produces the
@@ -965,9 +1051,9 @@ notes.note:reminder
notes.user:alice notes.note:insurance
```
All of these constructs can be combined as you see fit. If you want to
learn more about the query language used by paperless, see the
[Tantivy query language documentation](https://docs.rs/tantivy/latest/tantivy/query/struct.QueryParser.html).
The bare `notes:` prefix is shorthand for `notes.note:`.
All of these constructs can be combined as you see fit. What is described above is the whole of the query language paperless supports. It resembles other search query languages without being identical to any of them, so a construct that is not documented here is most likely treated as ordinary search text rather than as syntax, and an unrecognized field name is searched as text too.
!!! note
+4
View File
@@ -77,6 +77,7 @@ dependencies = [
"torch~=2.13.0",
"watchfiles>=1.2",
"whitenoise~=6.11",
"whoosh-compat[tantivy]",
"zxing-cpp~=3.1.0",
]
[project.optional-dependencies]
@@ -168,6 +169,9 @@ psycopg-c = [
torch = [
{ index = "pytorch-cpu" },
]
# TODO: switch to a pinned PyPI version once whoosh-compat releases; fall back
# to a pinned git commit SHA if that release slips.
whoosh-compat = { path = "../whoosh-compat" }
[tool.ruff]
target-version = "py311"
+10 -2
View File
@@ -6,13 +6,20 @@ from documents.search._backend import TantivyRelevanceList
from documents.search._backend import WriteBatch
from documents.search._backend import get_backend
from documents.search._backend import reset_backend
from documents.search._errors import InvalidDateQuery
from documents.search._errors import InvalidNumberQuery
from documents.search._errors import MultipleSearchQueryErrors
from documents.search._errors import QueryTooLongError
from documents.search._errors import SearchQueryError
from documents.search._errors import search_query_error_messages
from documents.search._schema import needs_rebuild
from documents.search._schema import wipe_index
from documents.search._translate import InvalidDateQuery
from documents.search._translate import SearchQueryError
__all__ = [
"InvalidDateQuery",
"InvalidNumberQuery",
"MultipleSearchQueryErrors",
"QueryTooLongError",
"SearchHit",
"SearchIndexLockError",
"SearchMode",
@@ -23,5 +30,6 @@ __all__ = [
"get_backend",
"needs_rebuild",
"reset_backend",
"search_query_error_messages",
"wipe_index",
]
+59 -9
View File
@@ -25,7 +25,6 @@ from django.utils.timezone import get_current_timezone
from guardian.shortcuts import get_groups_with_perms
from guardian.shortcuts import get_users_with_perms
from documents.search._query import build_permission_filter
from documents.search._query import extract_cjk_text
from documents.search._query import parse_simple_text_highlight_query
from documents.search._query import parse_simple_text_query
@@ -43,6 +42,7 @@ from documents.utils import QuerySetStream
from documents.utils import identity
if TYPE_CHECKING:
from collections.abc import Iterable
from collections.abc import Iterator
from collections.abc import Sequence
from pathlib import Path
@@ -294,6 +294,47 @@ class WriteBatch:
)
def build_permission_filter(
schema: tantivy.Schema,
user: AbstractUser,
viewer_group_ids: Iterable[int] = (),
) -> tantivy.Query:
"""
Build a query filter for user document permissions.
Creates a query that matches only documents visible to the specified user
according to paperless-ngx permission rules:
- Public documents (no owner) are visible to all users
- Private documents are visible to their owner
- Documents explicitly shared with the user are visible
- Documents shared with one of the user's current groups are visible
Args:
schema: Tantivy schema for field validation
user: User to check permissions for
viewer_group_ids: Current group memberships for the user
Returns:
Tantivy query that filters results to visible documents
"""
owner_any = tantivy.Query.exists_query("owner_id")
no_owner = tantivy.Query.boolean_query(
[
(tantivy.Occur.Must, tantivy.Query.all_query()),
(tantivy.Occur.MustNot, owner_any),
],
)
owned = tantivy.Query.term_query(schema, "owner_id", user.pk)
shared = tantivy.Query.term_query(schema, "viewer_id", user.pk)
group_shared = [
tantivy.Query.term_query(schema, "viewer_group_id", group_id)
for group_id in viewer_group_ids
]
return tantivy.Query.disjunction_max_query(
[no_owner, owned, shared, *group_shared],
)
class TantivyBackend:
"""
Tantivy search backend with explicit lifecycle management.
@@ -465,7 +506,6 @@ class TantivyBackend:
doc.add_text("correspondent_sort", document.correspondent.name)
if cjk_corr := extract_cjk_text(document.correspondent.name):
doc.add_text("bigram_correspondent", cjk_corr)
doc.add_unsigned("correspondent_id", document.correspondent_id)
# Document type
if document.document_type:
@@ -473,12 +513,10 @@ class TantivyBackend:
doc.add_text("type_sort", document.document_type.name)
if cjk_type := extract_cjk_text(document.document_type.name):
doc.add_text("bigram_document_type", cjk_type)
doc.add_unsigned("document_type_id", document.document_type_id)
# Storage path
if document.storage_path:
doc.add_text("storage_path", document.storage_path.name)
doc.add_unsigned("storage_path_id", document.storage_path_id)
# Tags — collect names for autocomplete in the same pass
tag_names: list[str] = []
@@ -486,12 +524,13 @@ class TantivyBackend:
doc.add_text("tag", tag.name)
if cjk_tag := extract_cjk_text(tag.name):
doc.add_text("bigram_tag", cjk_tag)
doc.add_unsigned("tag_id", tag.pk)
tag_names.append(tag.name)
# Notes — JSON for structured queries (notes.user:alice, notes.note:text).
# notes_text is a plain-text companion for snippet/highlight generation;
# tantivy's SnippetGenerator does not support JSON fields.
# tantivy's SnippetGenerator does not support JSON fields. It is not in
# _DEFAULT_SEARCH_FIELDS, so an unqualified query never searches it: a
# note matches through the JSON field or not at all.
num_notes = 0
note_texts: list[str] = []
for note in document.notes.all():
@@ -507,8 +546,9 @@ class TantivyBackend:
if note_texts:
doc.add_text("notes_text", " ".join(note_texts))
# Custom fields — JSON for structured queries (custom_fields.name:x, custom_fields.value:y),
# companion text field for default full-text search.
# Custom fields — JSON for structured queries (custom_fields.name:x,
# custom_fields.value:y). There is no companion text field here, unlike
# notes: custom field values are reachable only through the JSON field.
for cfi in document.custom_fields.all():
search_value = cfi.value_for_search
# Skip fields where there is no value yet
@@ -680,7 +720,17 @@ class TantivyBackend:
user_query = self._parse_query(query, search_mode)
highlight_query = user_query
if search_mode is SearchMode.TEXT:
highlight_query = parse_simple_text_highlight_query(self._index, query)
try:
highlight_query = parse_simple_text_highlight_query(
self._index,
query,
)
except ValueError:
logger.debug(
"Skipping simple text highlight query: token string is not "
"valid tantivy query syntax: %r",
query,
)
# For notes_text snippet generation, we need a query that targets the
# notes_text field directly. user_query may contain JSON-field terms
-171
View File
@@ -1,171 +0,0 @@
from __future__ import annotations
from datetime import UTC
from datetime import date
from datetime import datetime
from datetime import timedelta
from typing import TYPE_CHECKING
from typing import Final
from dateutil.relativedelta import relativedelta
if TYPE_CHECKING:
from datetime import tzinfo
_DATE_ONLY_FIELDS = frozenset({"created"})
_TODAY: Final[str] = "today"
_YESTERDAY: Final[str] = "yesterday"
_PREVIOUS_WEEK: Final[str] = "previous week"
_THIS_MONTH: Final[str] = "this month"
_PREVIOUS_MONTH: Final[str] = "previous month"
_THIS_YEAR: Final[str] = "this year"
_PREVIOUS_YEAR: Final[str] = "previous year"
_PREVIOUS_QUARTER: Final[str] = "previous quarter"
_DATE_KEYWORDS = frozenset(
{
_TODAY,
_YESTERDAY,
_PREVIOUS_WEEK,
_THIS_MONTH,
_PREVIOUS_MONTH,
_THIS_YEAR,
_PREVIOUS_YEAR,
_PREVIOUS_QUARTER,
},
)
def _fmt(dt: datetime) -> str:
"""Format a datetime as an ISO 8601 UTC string for use in Tantivy range queries."""
return dt.astimezone(UTC).strftime("%Y-%m-%dT%H:%M:%SZ")
def _iso_range(lo: datetime, hi: datetime) -> str:
"""
Format a half-open ``[lo TO hi)`` range in ISO 8601 for Tantivy query syntax.
``hi`` is always the exclusive ceiling of a computed period (the start of
the *next* day/week/month/quarter/year), so the closing bracket must be
the Tantivy exclusive-range brace ``}`` rather than ``]`` — otherwise the
first instant of the following period (e.g. the 1st of next month) is
incorrectly included in the match.
"""
return f"[{_fmt(lo)} TO {_fmt(hi)}}}"
def _quarter_start(d: date) -> date:
"""Return the first day of the calendar quarter containing ``d``."""
return date(d.year, ((d.month - 1) // 3) * 3 + 1, 1)
def _midnight(d: date, tz: tzinfo) -> datetime:
"""Convert a calendar date at local-timezone midnight to a UTC datetime."""
return datetime(d.year, d.month, d.day, tzinfo=tz).astimezone(UTC)
def _keyword_bounds(keyword: str, tz: tzinfo) -> tuple[date, date]:
"""
Map a relative date keyword to ``(start, exclusive_end)`` calendar dates.
``tz`` only determines what "today" is; the caller decides how the returned
dates become UTC datetime boundaries (date-only vs. local-midnight offset).
"""
today = datetime.now(tz).date()
if keyword == _TODAY:
return today, today + timedelta(days=1)
if keyword == _YESTERDAY:
return today - timedelta(days=1), today
if keyword == _PREVIOUS_WEEK:
this_monday = today - timedelta(days=today.weekday())
return this_monday - timedelta(weeks=1), this_monday
if keyword == _THIS_MONTH:
first = today.replace(day=1)
return first, first + relativedelta(months=1)
if keyword == _PREVIOUS_MONTH:
this_first = today.replace(day=1)
return this_first - relativedelta(months=1), this_first
if keyword == _THIS_YEAR:
return date(today.year, 1, 1), date(today.year + 1, 1, 1)
if keyword == _PREVIOUS_YEAR:
return date(today.year - 1, 1, 1), date(today.year, 1, 1)
if keyword == _PREVIOUS_QUARTER:
this_quarter = _quarter_start(today)
return this_quarter - relativedelta(months=3), this_quarter
raise ValueError(f"Unknown keyword: {keyword}")
def _date_only_range(keyword: str, tz: tzinfo) -> str:
"""
For `created` (DateField): use the local calendar date, converted to
midnight UTC boundaries. No offset arithmetic — date only.
"""
start, end = _keyword_bounds(keyword, tz)
lo = datetime(start.year, start.month, start.day, tzinfo=UTC)
hi = datetime(end.year, end.month, end.day, tzinfo=UTC)
return _iso_range(lo, hi)
def _datetime_range(keyword: str, tz: tzinfo) -> str:
"""
For `added` / `modified` (DateTimeField, stored as UTC): convert local day
boundaries to UTC — full offset arithmetic required.
"""
start, end = _keyword_bounds(keyword, tz)
return _iso_range(_midnight(start, tz), _midnight(end, tz))
def _precision_bounds(digits: str) -> tuple[date, date] | None:
"""
Map a 4/6/8-digit date token to (start, exclusive_end) calendar dates.
YYYY -> whole year, YYYYMM -> whole month, YYYYMMDD -> single day.
Returns None for any unparsable or out-of-range value (e.g. month 23),
so callers can emit a no-match clause instead of erroring (Whoosh parity).
"""
try:
if len(digits) == 4:
year = int(digits)
return date(year, 1, 1), date(year + 1, 1, 1)
if len(digits) == 6:
year, month = int(digits[:4]), int(digits[4:6])
start = date(year, month, 1)
end = date(year + 1, 1, 1) if month == 12 else date(year, month + 1, 1)
return start, end
if len(digits) == 8:
start = date(int(digits[:4]), int(digits[4:6]), int(digits[6:8]))
return start, start + timedelta(days=1)
except ValueError:
return None
return None
def _utc_bounds_for_field(
field: str,
start: date,
end: date,
tz: tzinfo,
) -> tuple[datetime, datetime]:
"""
Convert calendar-date bounds to UTC datetimes per the field's storage type.
For DateField (``created``) the bounds are UTC midnight (no offset). For
DateTimeField (``added``/``modified``) the bounds are local-tz midnight
converted to UTC, matching how each field is indexed.
"""
if field in _DATE_ONLY_FIELDS:
return (
datetime(start.year, start.month, start.day, tzinfo=UTC),
datetime(end.year, end.month, end.day, tzinfo=UTC),
)
return (
datetime(start.year, start.month, start.day, tzinfo=tz).astimezone(UTC),
datetime(end.year, end.month, end.day, tzinfo=tz).astimezone(UTC),
)
def _field_range_from_dates(field: str, start: date, end: date, tz: tzinfo) -> str:
"""Build a Tantivy ``field:[lo TO hi]`` ISO range from calendar-date bounds."""
lo, hi = _utc_bounds_for_field(field, start, end, tz)
return f"{field}:{_iso_range(lo, hi)}"
+71
View File
@@ -0,0 +1,71 @@
from __future__ import annotations
from typing import TYPE_CHECKING
if TYPE_CHECKING:
from collections.abc import Sequence
class SearchQueryError(ValueError):
"""
Base for user-fixable search query errors.
Carries a message safe to surface to the user (no internal details). The
view layer catches this and returns an HTTP 400, so any future subclass
gets the same treatment.
"""
class InvalidDateQuery(SearchQueryError):
"""Raised when a date field value or range bound cannot be parsed."""
def __init__(self, field: str | None, value: str | None) -> None:
self.field = field
self.value = value
super().__init__(f"Invalid date value {value!r} for field {field!r}.")
class InvalidNumberQuery(SearchQueryError):
"""Raised when a numeric field value or range bound cannot be parsed."""
def __init__(self, field: str | None, value: str | None) -> None:
self.field = field
self.value = value
super().__init__(f"Invalid numeric value {value!r} for field {field!r}.")
class QueryTooLongError(SearchQueryError):
"""Raised when a query string exceeds the maximum allowed length.
whoosh-compat's fieldname tagger is O(n^2) in plain word characters, so an
unbounded query is a CPU-exhaustion vector against a single request
handler. This is a hard boundary, not a validation nicety.
"""
def __init__(self, length: int, limit: int) -> None:
self.length = length
self.limit = limit
super().__init__(
f"The search query is too long ({length} characters). "
f"The maximum allowed length is {limit} characters.",
)
class MultipleSearchQueryErrors(SearchQueryError):
"""Aggregates every user-fixable error from one parse, not just the first."""
def __init__(self, errors: Sequence[SearchQueryError]) -> None:
self.errors = tuple(errors)
super().__init__("; ".join(str(e) for e in self.errors))
def search_query_error_messages(e: SearchQueryError) -> list[str]:
"""The user-facing message list for a SearchQueryError.
Every offending value's message, not just the first, so the user can
fix them all in one round-trip. Shared by every view that maps
SearchQueryError to an HTTP 400.
"""
if isinstance(e, MultipleSearchQueryErrors):
return [str(sub) for sub in e.errors]
return [str(e)]
+42
View File
@@ -0,0 +1,42 @@
from __future__ import annotations
from whoosh_compat import FieldKind
from whoosh_compat import FieldSpec
from whoosh_compat import SubpathSpec
# Internal-only schema fields with no query-syntax meaning of their own
# (sort shadow fields, bigram CJK fields, simple_title/simple_content,
# autocomplete_word, notes_text) are NOT represented here — they are
# declared in _schema.py's field_descriptors().
#
# analyzer/pattern_normalizer are deliberately left at FieldSpec's default
# (None): they're language-specific and only meaningful to whoosh-compat's
# parser, so _registry.py attaches them per-language via dataclasses.replace()
# rather than PUBLIC_FIELDS declaring them itself. _schema.py only reads
# name/kind/fast and never sees the analyzer at all.
PUBLIC_FIELDS: tuple[FieldSpec, ...] = (
FieldSpec("title", FieldKind.TEXT),
FieldSpec("content", FieldKind.TEXT),
FieldSpec("correspondent", FieldKind.TEXT),
FieldSpec("document_type", FieldKind.TEXT, aliases=("type",)),
FieldSpec("storage_path", FieldKind.TEXT, aliases=("path",)),
FieldSpec("original_filename", FieldKind.TEXT),
FieldSpec("tag", FieldKind.TEXT, comma_values=True),
FieldSpec("checksum", FieldKind.KEYWORD),
FieldSpec("asn", FieldKind.U64, fast=True),
FieldSpec("page_count", FieldKind.U64, fast=True),
FieldSpec("num_notes", FieldKind.U64, fast=True),
FieldSpec("created", FieldKind.DATE, date_only=True, fast=True),
FieldSpec("modified", FieldKind.DATETIME, fast=True),
FieldSpec("added", FieldKind.DATETIME, fast=True),
FieldSpec(
"notes",
FieldKind.JSON,
subpaths={"user": SubpathSpec(), "note": SubpathSpec(default=True)},
),
FieldSpec(
"custom_fields",
FieldKind.JSON,
subpaths={"name": SubpathSpec(), "value": SubpathSpec(default=True)},
),
)
+442 -145
View File
@@ -6,22 +6,30 @@ from typing import Final
import regex
import tantivy
import whoosh_compat as wc
from django.conf import settings
from whoosh_compat.emitters.tantivy_ import emit as tantivy_emit
from whoosh_compat.errors import Cause
from whoosh_compat.errors import Diagnostic
from whoosh_compat.errors import DiagnosticKind
from whoosh_compat.errors import QueryError
from documents.search._errors import InvalidDateQuery
from documents.search._errors import InvalidNumberQuery
from documents.search._errors import MultipleSearchQueryErrors
from documents.search._errors import SearchQueryError
from documents.search._registry import get_field_registry
from documents.search._tokenizer import simple_search_tokens
from documents.search._translate import SearchQueryError
from documents.search._translate import translate_query
if TYPE_CHECKING:
from collections.abc import Iterable
from datetime import tzinfo
from django.contrib.auth.base_user import AbstractBaseUser
logger = logging.getLogger("paperless.search")
# Maximum seconds any single regex substitution may run.
# Prevents ReDoS on adversarial user-supplied query strings.
# Maximum seconds any single regex substitution over user-supplied query text
# may run. The one remaining use is a character class, which cannot backtrack,
# so the bound is an upper limit on that substitution's cost, not the ReDoS
# guard it was originally written as.
_REGEX_TIMEOUT: Final[float] = 1.0
# Matches CJK/Hangul characters so queries can be routed to bigram fields.
@@ -29,6 +37,64 @@ _REGEX_TIMEOUT: Final[float] = 1.0
_CJK_RE: Final = regex.compile(r"[\p{Han}\p{Hiragana}\p{Katakana}\p{Hangul}]+")
def _user_facing_emit_message(d: Diagnostic) -> str:
"""A user-safe message for an emit-time QueryError's Diagnostic.
Built from the Diagnostic's structured fields (kind, field), never from
d.message: whoosh-compat documents that as developer/log output with no
stability guarantee, and PATTERN_TOO_COMPLEX embeds the raw backend
error text in it.
"""
field = str(d.field) if d.field is not None else None
if d.kind is DiagnosticKind.EXISTS_REQUIRES_FAST:
return f"Existence searches (field:*) are not supported for field {field!r}."
if d.kind is DiagnosticKind.TEXT_RANGE:
return f"Range searches are not supported for field {field!r}."
if d.kind is DiagnosticKind.PATTERN_TOO_COMPLEX:
return f"The wildcard pattern for field {field!r} is too complex."
if d.kind is DiagnosticKind.SCHEMA_FIELD_MISSING:
return f"Field {field!r} is not available in the search index."
logger.warning("Unmapped emit diagnostic %s: %s", d.kind, d.message)
return "The search query could not be executed."
def _map_emit_error(e: QueryError) -> SearchQueryError:
"""Route an emit-time QueryError by its Diagnostic's Cause.
INVALID_INPUT/UNSUPPORTED are user-input errors, exactly like a parse
diagnostic, and map to a 400. INTERNAL means a defect in whoosh-compat
or in our own AST handling, never the user's query, so the QueryError is
re-raised rather than converted, reaching the generic 500 handler instead
of blaming the query. MISCONFIGURED is deliberately both: the registry and
the index schema disagree, which only an operator can fix, so it is logged
as an error, but a request is still waiting and the query cannot run
either way, so it also returns a 400.
EXISTS_REQUIRES_FAST is the one MISCONFIGURED kind that is not a
disagreement. whoosh-compat derives it from the registry's own FieldSpec
(kind plus fast) without ever consulting the index schema, so it fires
whenever a non-fast field of a kind that cannot answer "exists" is asked
to: for us that is only the JSON fields, which field_descriptors() builds
non-fast on purpose. "notes:*" and the five other spellings of it are
ordinary user error that no operator action can clear, so they get the
400 without the alert.
"""
d = e.diagnostic
if d.cause is Cause.INTERNAL:
raise e
if (
d.cause is Cause.MISCONFIGURED
and d.kind is not DiagnosticKind.EXISTS_REQUIRES_FAST
):
logger.error(
"Search index misconfiguration for field %s (%s): %s",
d.field,
d.kind.name,
d.message,
)
return SearchQueryError(_user_facing_emit_message(d))
def _has_cjk(text: str) -> bool:
"""Return True if text contains any CJK characters."""
return bool(_CJK_RE.search(text))
@@ -37,14 +103,36 @@ def _has_cjk(text: str) -> bool:
def extract_cjk_text(text: str) -> str:
"""Join the CJK runs in ``text`` for indexing into bigram (char-ngram) fields.
Mirrors the query side (``_build_cjk_query``): only CJK runs are ever searched
against the bigram fields, so only CJK runs are worth indexing there. Latin
text fed to a character-bigram field is never matched and only bloats the
Mirrors the query side, which extracts the CJK runs of whatever it is
about to search for (the raw string in simple modes, the parsed query's
free-text tokens in query mode): only CJK runs are ever searched against
the bigram fields, so only CJK runs are worth indexing there. Latin text
fed to a character-bigram field is never matched and only bloats the
index and slows indexing/merge. Returns "" when there is no CJK text.
"""
return " ".join(_CJK_RE.findall(text))
def _parse_cjk_text(
index: tantivy.Index,
cjk_text: str,
fields: list[str],
) -> tantivy.Query | None:
"""Parse a plain CJK run string against ``fields``, or None if it won't parse."""
try:
return index.parse_query(cjk_text, fields)
except Exception:
# Broad on purpose, unlike _try_parse_fuzzy_query's narrower
# ValueError: cjk_text isn't filtered to a guaranteed-safe token
# set the way the fuzzy blend's word string is, so the exact
# failure mode tantivy could raise here isn't pinned down.
logger.debug(
"Skipping CJK search clause: could not parse CJK text: %r",
cjk_text,
)
return None
def _build_cjk_query(
index: tantivy.Index,
raw_query: str,
@@ -52,91 +140,259 @@ def _build_cjk_query(
) -> tantivy.Query | None:
"""Build a bigram-field query from the CJK runs in ``raw_query``.
Only the CJK character runs are extracted and parsed; ASCII field prefixes,
boolean operators and date keywords are discarded. This keeps the CJK clause
plain-text and consistent across query/simple modes (no leaked ``field:``
semantics, no parse failures from spaced ``-``/``+``), and avoids feeding
Latin tokens into the character-bigram matcher (which would produce spurious
matches against unrelated Latin text). Returns None when there is no CJK
text or the parse fails.
For the simple (TEXT/TITLE) modes, whose input is plain text and carries
no query grammar to respect. Only the CJK character runs are extracted, so
a stray ``field:`` prefix or ``-``/``+`` in the input can neither leak
field semantics nor fail the parse, and no Latin token reaches the
character-bigram matcher (where it would produce spurious matches against
unrelated Latin text). Returns None when there is no CJK text or the parse
fails.
"""
cjk_text = " ".join(_CJK_RE.findall(raw_query))
cjk_text = extract_cjk_text(raw_query)
if not cjk_text:
return None
return _parse_cjk_text(index, cjk_text, fields)
def _build_ast_cjk_query(
index: tantivy.Index,
ast: wc.ast.Node,
registry: wc.FieldRegistry,
) -> tantivy.Query | None:
"""Build the bigram clause of a QUERY-mode search from the parsed AST.
Same discipline as the fuzzy clause (see _try_parse_fuzzy_query): the CJK
runs come from whoosh_compat's ``free_text_tokens`` over the parsed tree,
never from the raw query string, so a term the user negated or restricted
to a field outside the default search fields contributes nothing, instead
of resurfacing as a top-level clause matching every bigram field.
``free_text_tokens`` reports no field of its own, so the tokens are
collected one default field at a time: a bare term, which the parser has
already copied onto every default field, is therefore searched across
every bigram field, while ``title:東京`` reaches ``bigram_title`` alone.
Fields whose CJK text is identical (the bare-term case) share a single
parse over all of their bigram fields at once.
Raw (``analyzed=False``) tokens are used because the bigram fields have
their own character-ngram analyzer: the default fields' word analyzers
have no useful say over a CJK run, and running them first would only
risk dropping it (remove_long) before the run is ever extracted.
Returns None when the query has no CJK free text.
"""
fields_by_text: dict[str, list[str]] = {}
for field, bigram_field in _CJK_BIGRAM_FIELDS.items():
tokens = wc.free_text_tokens(
ast,
registry=registry,
fields=[field],
analyzed=False,
)
cjk_text = extract_cjk_text(" ".join(tokens))
if cjk_text:
fields_by_text.setdefault(cjk_text, []).append(bigram_field)
clauses: list[tuple[tantivy.Occur, tantivy.Query]] = [
(tantivy.Occur.Should, query)
for cjk_text, bigram_fields in fields_by_text.items()
if (query := _parse_cjk_text(index, cjk_text, bigram_fields)) is not None
]
return _any_of(clauses) if clauses else None
# A joined fuzzy word string must stay plain words: it goes back through
# tantivy's own query parser, and the raw query text the clause collects
# routinely carries characters that parser reads as grammar (a colon, a
# bracket, a quote, a leading -). Each token is cut into its word runs and
# only those are kept, so no field syntax, pattern, range or grouping can
# reach the parser. Cutting rather than dropping the whole token is what
# keeps ordinary hyphenated, dotted and quoted input ("COVID-19",
# "hello@example.com", "tax reports") contributing to the clause at all.
_WORD_RUN_RE = regex.compile(r"\w+")
# The one piece of tantivy grammar that survives the cut: its boolean
# keywords are themselves word runs. Only these exact spellings are
# grammar there ("And"/"and" are ordinary terms), so lowercasing exactly
# these turns them back into the ordinary terms the field analyzer used to
# make of them, before the clause switched to raw text. Left alone, a
# quoted phrase would silently restructure the clause ("tax AND reports"
# becoming a conjunction) or fail to parse and drop it entirely
# ("tax AND", or "IN" anywhere).
#
# Only these words are touched: tantivy lowercases query terms with the
# field's own analyzer, and doing it ourselves first is not always the
# same operation (Python folds a final sigma to a different letter than
# tantivy does, and turns Turkish 'İ' into a sequence tantivy then splits
# in two), which would search for terms the index does not contain.
_TANTIVY_KEYWORDS: Final[frozenset[str]] = frozenset({"AND", "OR", "NOT", "IN"})
def _try_parse_fuzzy_query(
index: tantivy.Index,
ast: wc.ast.Node,
registry: wc.FieldRegistry,
) -> tantivy.Query | None:
"""Build the fuzzy blend clause from the parsed query's free-text
words, or None if it has none.
The clause is built by handing tantivy's own query parser a plain
word string (there's no clean AST-level fuzzy equivalent to
whoosh-compat's parse tree, and fuzzy matching was always an
approximate, secondary, 0.1-boosted clause). The words come from
whoosh_compat's ``free_text_tokens`` over the already-parsed AST,
never from the raw query string: raw whoosh grammar (date keywords,
``[2005 to 2009]`` ranges, bracket-class wildcards) is not tantivy
syntax, and feeding it here used to knock the fuzzy clause out for
the whole query the moment any such construct appeared alongside a
typo'd word. The helper also keeps excluded terms out: a ``NOT``'d
word must not resurface through the fuzzy clause.
Chosen trade-off: a term explicitly fielded on one of the default
search fields (``correspondent:acme``) contributes its text to the
word string UNFIELDED, so the fuzzy clause searches it across all
default fields rather than just the one the user named. That is
recall-only widening on a secondary 0.1-boosted clause the score
threshold already disciplines, accepted in exchange for never feeding
field syntax to tantivy's parser. What the word string guarantees is
exactly that: no field prefix, pattern, range, grouping or quoting
survives, and the boolean keywords that do survive (they are word
runs) are lowercased into ordinary terms; see _TANTIVY_KEYWORDS.
The words are the query's RAW text, not the analyzer's output
(``analyzed=False``), because ``index.parse_query`` analyzes whatever
it is given and analysis is not idempotent: ``universities`` stems to
``univers``, and handing that back stems it again to ``univ``, a term
the index does not contain. ``prefix=True`` hid this as over-broad
matching (``univ`` also prefixes ``unicycle``) rather than as no
matches at all. Raw text is untokenized, which is why it is cut into
word runs above rather than taken whole.
The ValueError guard stays as insurance (the word string is plain
tokens, so tantivy accepting it is expected, not assumed): on a parse
failure the fuzzy clause is skipped and the exact/CJK clauses stand,
rather than the whole query failing.
"""
tokens = wc.free_text_tokens(
ast,
registry=registry,
fields=_DEFAULT_SEARCH_FIELDS,
analyzed=False,
)
words = list(
dict.fromkeys(
word.lower() if word in _TANTIVY_KEYWORDS else word
for token in tokens
for word in _WORD_RUN_RE.findall(token)
),
)
if not words:
return None
fuzzy_text = " ".join(words)
try:
return index.parse_query(cjk_text, fields)
except Exception:
return index.parse_query(
fuzzy_text,
_DEFAULT_SEARCH_FIELDS,
field_boosts=_FIELD_BOOSTS,
fuzzy_fields={f: (True, 1, True) for f in _DEFAULT_SEARCH_FIELDS},
)
except ValueError:
logger.debug(
"Skipping fuzzy search clause: token string is not valid "
"tantivy query syntax: %r",
fuzzy_text,
)
return None
def build_permission_filter(
schema: tantivy.Schema,
user: AbstractBaseUser,
viewer_group_ids: Iterable[int] = (),
) -> tantivy.Query:
"""
Build a query filter for user document permissions.
Creates a query that matches only documents visible to the specified user
according to paperless-ngx permission rules:
- Public documents (no owner) are visible to all users
- Private documents are visible to their owner
- Documents explicitly shared with the user are visible
- Documents shared with one of the user's current groups are visible
Args:
schema: Tantivy schema for field validation
user: User to check permissions for
viewer_group_ids: Current group memberships for the user
Returns:
Tantivy query that filters results to visible documents
"""
owner_any = tantivy.Query.exists_query("owner_id")
no_owner = tantivy.Query.boolean_query(
[
(tantivy.Occur.Must, tantivy.Query.all_query()),
(tantivy.Occur.MustNot, owner_any),
],
)
owned = tantivy.Query.term_query(schema, "owner_id", user.pk)
shared = tantivy.Query.term_query(schema, "viewer_id", user.pk)
group_shared = [
tantivy.Query.term_query(schema, "viewer_group_id", group_id)
for group_id in viewer_group_ids
]
return tantivy.Query.disjunction_max_query(
[no_owner, owned, shared, *group_shared],
)
DEFAULT_SEARCH_FIELDS = [
_DEFAULT_SEARCH_FIELDS: Final[list[str]] = [
"title",
"content",
"correspondent",
"document_type",
"tag",
]
SIMPLE_SEARCH_FIELDS = ["simple_title", "simple_content"]
TITLE_SEARCH_FIELDS = ["simple_title"]
_CJK_ALL_FIELDS: Final[list[str]] = [
"bigram_content",
"bigram_title",
"bigram_correspondent",
"bigram_document_type",
"bigram_tag",
]
_SIMPLE_SEARCH_FIELDS: Final[list[str]] = ["simple_title", "simple_content"]
_TITLE_SEARCH_FIELDS: Final[list[str]] = ["simple_title"]
# The bigram (character-ngram) companion of each default search field.
_CJK_BIGRAM_FIELDS: Final[dict[str, str]] = {
field: f"bigram_{field}" for field in _DEFAULT_SEARCH_FIELDS
}
_CJK_CONTENT_FIELDS: Final[list[str]] = ["bigram_content"]
_CJK_TITLE_FIELDS: Final[list[str]] = ["bigram_title"]
_FIELD_BOOSTS = {"title": 2.0}
_SIMPLE_FIELD_BOOSTS = {"simple_title": 2.0}
def _simple_query_tokens(raw_query: str) -> list[str]:
# Tokenize and fold via the same analyzer used to index simple_title /
# simple_content, so query terms fold identically to the indexed terms
# (single source of truth for ASCII folding).
return simple_search_tokens(raw_query)
class _ConjunctiveNegations(wc.ast.Visitor[tuple["wc.ast.Node", ...]]):
"""Collect the subtrees an AST excludes from every document it matches.
A negation reached through ``And``/``AndNot``/``Require`` (and through
the required half of an ``AndMaybe``) constrains the whole query, so it
can be re-stated above the blend. ``Or`` is deliberately not descended
into: in ``invoice OR NOT secret`` the negation is one branch's own
condition, and hoisting it would throw away documents the other branch
matches. Nor is a collected subtree descended into, since a negation
inside a negation is not an exclusion.
Node types with no negation to contribute (every leaf, ``Or``) fall
through to ``generic_visit``.
"""
def generic_visit(self, node: wc.ast.Node) -> tuple[wc.ast.Node, ...]:
return ()
def visit_not(self, node: wc.ast.Not) -> tuple[wc.ast.Node, ...]:
return (node.child,)
def visit_andnot(self, node: wc.ast.AndNot) -> tuple[wc.ast.Node, ...]:
return (*self.visit(node.positive), node.negative)
def visit_and(self, node: wc.ast.And) -> tuple[wc.ast.Node, ...]:
return tuple(
negation for child in node.children for negation in self.visit(child)
)
def visit_boosted(self, node: wc.ast.Boosted) -> tuple[wc.ast.Node, ...]:
return self.visit(node.child)
def visit_andmaybe(self, node: wc.ast.AndMaybe) -> tuple[wc.ast.Node, ...]:
return self.visit(node.required)
def visit_require(self, node: wc.ast.Require) -> tuple[wc.ast.Node, ...]:
return (*self.visit(node.scored), *self.visit(node.filter_only))
def _negation_clauses(
index: tantivy.Index,
ast: wc.ast.Node,
registry: wc.FieldRegistry,
) -> list[tuple[tantivy.Occur, tantivy.Query]]:
"""MustNot clauses for everything ``ast`` excludes conjunctively.
Each excluded subtree is emitted as its own positive query and attached
with ``MustNot``, rather than emitting a negative query and hoping
tantivy accepts a bare one.
"""
try:
return [
(
tantivy.Occur.MustNot,
tantivy_emit(negation, index=index, registry=registry),
)
for negation in _ConjunctiveNegations().visit(ast)
]
except QueryError as e:
raise _map_emit_error(e) from e
def _any_of(clauses: list[tuple[tantivy.Occur, tantivy.Query]]) -> tantivy.Query:
"""Collapse a clause list: none -> empty, one -> itself (no wasted
single-clause boolean_query wrapping), many -> boolean_query(clauses)."""
if not clauses:
return tantivy.Query.empty_query()
if len(clauses) == 1:
return clauses[0][1]
return tantivy.Query.boolean_query(clauses)
def _build_simple_token_query(
@@ -168,9 +424,7 @@ def _build_simple_token_query(
query = tantivy.Query.boost_query(query, boost)
field_queries.append((tantivy.Occur.Should, query))
if len(field_queries) == 1:
return field_queries[0][1]
return tantivy.Query.boolean_query(field_queries)
return _any_of(field_queries)
def parse_user_query(
@@ -179,52 +433,53 @@ def parse_user_query(
tz: tzinfo,
) -> tantivy.Query:
"""
Parse user query through the complete preprocessing pipeline.
Parse user query through whoosh-compat, then blend in fuzzy/CJK clauses.
Transforms the raw user query through multiple stages:
1. Date keyword rewriting (today ISO 8601 ranges)
2. Query normalization (comma expansion, whitespace cleanup)
3. Tantivy parsing with field boosts
4. Optional fuzzy query blending (if ADVANCED_FUZZY_SEARCH_THRESHOLD set)
Args:
index: Tantivy index with registered tokenizers
raw_query: Original user query string
tz: Timezone for date boundary calculations
Returns:
Parsed Tantivy query ready for execution
Note:
When ADVANCED_FUZZY_SEARCH_THRESHOLD is configured, adds a low-priority
fuzzy query as a Should clause (0.1 boost) to catch approximate matches
while keeping exact matches ranked higher. The threshold value is applied
as a post-search score filter, not during query construction.
1. wc.parse() against the shared FieldRegistry (whoosh grammar -> AST).
Bare notes:/custom_fields: prefixes resolve to their default subpath
(notes.note:/custom_fields.value:) directly in the registry, via
each JSON field's SubpathSpec(default=True).
2. Any diagnostics (bad dates/numbers) map to SearchQueryError subclasses
and raise the view returns HTTP 400 with every offending field
listed, not just the first.
3. emit() turns the AST into a tantivy.Query directly (no string
round-trip). A QueryError is routed by its Diagnostic's Cause
(_map_emit_error): a construct that parses but can't execute against
tantivy (e.g. a text-field range) is a 400, a registry/schema
mismatch is logged and a 400, and an INTERNAL defect is re-raised.
4. Optional fuzzy blend (ADVANCED_FUZZY_SEARCH_THRESHOLD) builds a
plain word string from the parsed AST's free-text tokens
(whoosh_compat.free_text_tokens) and feeds THAT to
index.parse_query never raw_query, whose whoosh grammar (date
keywords, bracket-class wildcards, etc.) tantivy's parser rejects,
which used to silently knock the fuzzy clause out of any mixed
query (see _try_parse_fuzzy_query).
5. Optional CJK bigram clause, built from the same parsed AST for the
same reason (see _build_ast_cjk_query): a CJK term the query negated
or fielded must not resurface through it.
6. When any optional clause was added, the query's conjunctive
exclusions are restated as MustNot above the blend
(_negation_clauses): a clause built from positive terms cannot
express them, and as a bare Should it would undo them.
"""
registry = get_field_registry(settings.SEARCH_LANGUAGE)
result = wc.parse(
raw_query,
registry=registry,
default_fields=_DEFAULT_SEARCH_FIELDS,
field_boosts=_FIELD_BOOSTS,
tz=tz,
)
if result.diagnostics:
raise _diagnostics_to_error(result.diagnostics)
try:
query_str = translate_query(raw_query, tz)
except SearchQueryError:
# Intentional, user-fixable error (e.g. an unparsable date). Propagate so
# the view can return a 400 with a helpful message rather than falling
# back to the raw (still-invalid) query.
raise
except Exception: # pragma: no cover - defensive
logger.warning("Query translation failed; using raw query", exc_info=True)
query_str = raw_query
exact = tantivy_emit(result.ast, index=index, registry=registry)
except QueryError as e:
raise _map_emit_error(e) from e
exact = index.parse_query(
query_str,
DEFAULT_SEARCH_FIELDS,
field_boosts=_FIELD_BOOSTS,
)
# The standard analyzer keeps a whitespace-free CJK run as a single token,
# so substring queries can't match content/title (and long runs are dropped
# by remove_long). Route CJK queries to the bigram fields, whose ngram
# tokenizer indexes overlapping 2-grams for substring matching.
cjk_query = (
_build_cjk_query(index, raw_query, _CJK_ALL_FIELDS)
_build_ast_cjk_query(index, result.ast, registry)
if _has_cjk(raw_query)
else None
)
@@ -235,22 +490,65 @@ def parse_user_query(
threshold = settings.ADVANCED_FUZZY_SEARCH_THRESHOLD
if threshold is not None:
fuzzy = index.parse_query(
query_str,
DEFAULT_SEARCH_FIELDS,
field_boosts=_FIELD_BOOSTS,
# (prefix=True, distance=1, transposition_cost_one=True) — edit-distance fuzziness
fuzzy_fields={f: (True, 1, True) for f in DEFAULT_SEARCH_FIELDS},
)
# 0.1 boost keeps fuzzy hits ranked below exact matches (intentional)
clauses.append((tantivy.Occur.Should, tantivy.Query.boost_query(fuzzy, 0.1)))
fuzzy = _try_parse_fuzzy_query(index, result.ast, registry)
if fuzzy is not None:
clauses.append(
(tantivy.Occur.Should, tantivy.Query.boost_query(fuzzy, 0.1)),
)
if cjk_query is not None:
clauses.append((tantivy.Occur.Should, cjk_query))
if len(clauses) == 1:
return exact
return tantivy.Query.boolean_query(clauses)
# The fuzzy and CJK clauses are built from positive terms only, so as
# plain Shoulds beside the exact clause they re-admit exactly the
# documents the query excluded. Restate the exclusions once, above the
# whole blend. Redundant against the exact clause, which already
# carries them, but idempotently so.
negations = _negation_clauses(index, result.ast, registry)
if not negations:
return _any_of(clauses)
return tantivy.Query.boolean_query(
[(tantivy.Occur.Must, _any_of(clauses)), *negations],
)
# The three whoosh-compat kinds for a wildcard on a field that cannot
# carry one. d.field_kind supplies the discriminator, so naming the field's
# type needs no second trip through the registry.
_PATTERN_ON_KINDS: Final = frozenset(
{
DiagnosticKind.PATTERN_ON_NUMERIC,
DiagnosticKind.PATTERN_ON_BOOLEAN_EXISTS,
DiagnosticKind.PATTERN_ON_SUBPATH,
},
)
def _diagnostics_to_error(diagnostics: tuple[Diagnostic, ...]) -> SearchQueryError:
errors = [_single_diagnostic_to_error(d) for d in diagnostics]
return errors[0] if len(errors) == 1 else MultipleSearchQueryErrors(errors)
def _single_diagnostic_to_error(d: Diagnostic) -> SearchQueryError:
# d.field is a FieldRef, not a str: str(d.field) gives the canonical
# dotted name (an aliased query, e.g. type:, reports document_type).
field_name = str(d.field) if d.field is not None else None
if d.kind is DiagnosticKind.BAD_DATE:
return InvalidDateQuery(field_name, d.raw_value)
if d.kind is DiagnosticKind.BAD_NUMBER:
return InvalidNumberQuery(field_name, d.raw_value)
if d.kind is DiagnosticKind.TOO_DEEP:
return SearchQueryError("The search query is nested too deeply.")
if d.kind in _PATTERN_ON_KINDS:
kind_label = f" ({d.field_kind.name.lower()})" if d.field_kind else ""
return SearchQueryError(
f"Wildcard patterns are not supported for field "
f"{field_name!r}{kind_label}.",
)
logger.warning("Unmapped parse diagnostic %s: %s", d.kind, d.message)
return SearchQueryError("The search query could not be executed.")
def parse_simple_query(
@@ -268,7 +566,7 @@ def parse_simple_query(
CJK substrings the simple analyzer can't (long whitespace-free runs are
dropped by remove_long).
"""
tokens = _simple_query_tokens(raw_query)
tokens = simple_search_tokens(raw_query)
clauses: list[tuple[tantivy.Occur, tantivy.Query]] = []
if tokens:
@@ -291,23 +589,14 @@ def parse_simple_query(
)
for token in tokens
]
simple_query = (
token_queries[0][1]
if len(token_queries) == 1
else tantivy.Query.boolean_query(token_queries)
)
clauses.append((tantivy.Occur.Should, simple_query))
clauses.append((tantivy.Occur.Should, _any_of(token_queries)))
if cjk_fields and _has_cjk(raw_query):
cjk_q = _build_cjk_query(index, raw_query, cjk_fields)
if cjk_q is not None:
clauses.append((tantivy.Occur.Should, cjk_q))
if not clauses:
return tantivy.Query.empty_query()
if len(clauses) == 1:
return clauses[0][1]
return tantivy.Query.boolean_query(clauses)
return _any_of(clauses)
def parse_simple_text_highlight_query(
@@ -322,13 +611,21 @@ def parse_simple_text_highlight_query(
# Strip Tantivy operator chars before tokenizing: this is a plain-text
# highlight query, not a structured boolean query, so +/- are separators.
tokens = _simple_query_tokens(
tokens = simple_search_tokens(
regex.sub(r"[-+]", " ", raw_query, timeout=_REGEX_TIMEOUT),
)
if not tokens:
return tantivy.Query.empty_query()
return index.parse_query(" ".join(tokens), ["content"])
# Quote each token as its own phrase, escaping backslashes and embedded
# quotes. simple search tokens can carry arbitrary Tantivy syntax
# characters (`"`, `:`, `(`, `[`, `/`, ...) that the query-string parser
# would otherwise interpret as query grammar rather than literal text.
quoted_tokens = [
'"' + token.replace("\\", "\\\\").replace('"', '\\"') + '"' for token in tokens
]
return index.parse_query(" ".join(quoted_tokens), ["content"])
def parse_simple_text_query(
@@ -342,7 +639,7 @@ def parse_simple_text_query(
return parse_simple_query(
index,
raw_query,
SIMPLE_SEARCH_FIELDS,
_SIMPLE_SEARCH_FIELDS,
cjk_fields=_CJK_CONTENT_FIELDS,
)
@@ -358,6 +655,6 @@ def parse_simple_title_query(
return parse_simple_query(
index,
raw_query,
TITLE_SEARCH_FIELDS,
_TITLE_SEARCH_FIELDS,
cjk_fields=_CJK_TITLE_FIELDS,
)
+91
View File
@@ -0,0 +1,91 @@
from __future__ import annotations
import dataclasses
from typing import TYPE_CHECKING
from whoosh_compat import FieldKind
from whoosh_compat import FieldRegistry
from documents.search._fields import PUBLIC_FIELDS
from documents.search._tokenizer import ascii_fold
from documents.search._tokenizer import paperless_text_analyzer
from documents.search._tokenizer import stem_pattern_text
if TYPE_CHECKING:
from whoosh_compat import PatternNormalizer
_registry_cache: dict[str | None, FieldRegistry] = {}
def _identity_analyzer(text: str) -> list[str]:
"""Analyzer for KEYWORD fields indexed with the raw tokenizer (no splitting)."""
return [text]
def _fold_normalizer(text: str) -> str:
"""Wildcard/regex literal-run normalizer for fields indexed without stemming."""
return ascii_fold(text.lower())
def _make_pattern_normalizer(language: str | None) -> PatternNormalizer:
"""Build the wildcard/regex literal-run normalizer for a search language."""
def _pattern_normalizer(text: str) -> tuple[str, ...]:
"""Normalize a literal run into the forms a term may match.
TEXT index terms go through lowercase -> ascii_fold -> stem, so a
pattern that skips stemming can never match one: "invoice*" would look
for a term starting with "invoice" while the index holds "invoic". The
run is therefore offered stemmed as well. KEYWORD fields are indexed
raw and get _fold_normalizer instead, so their patterns stay literal.
Both forms are returned, as alternatives, because neither is a prefix
of the other in general: English stemming substitutes as well as
truncates ("copy" -> "copi"), so the stem alone loses the compounds
the typed run reaches ("copyright") while the typed run alone loses
the inflections the stem reaches ("copies"). whoosh-compat ORs the
alternatives per literal run and deduplicates them, so a run the
stemmer leaves alone costs exactly the one branch it did before.
Inside a bracket class the emitter calls this once per character and
uses the answer only if it is a single one-character form; two forms
there leave the character as typed. A stemmer does not change a lone
character, so the two forms deduplicate to one and the class body is
folded as before.
"""
folded = ascii_fold(text.lower())
stemmed = stem_pattern_text(folded, language)
return (folded, stemmed)
return _pattern_normalizer
def get_field_registry(language: str | None) -> FieldRegistry:
"""Build (or return the cached) FieldRegistry for the given search language.
Cached keyed by language, rebuilt on the same trigger register_tokenizers()
uses (settings.SEARCH_LANGUAGE change) a fresh call with a new language
builds and caches a new registry rather than mutating the old one.
"""
if language in _registry_cache:
return _registry_cache[language]
text_analyzer = paperless_text_analyzer(language).analyze
pattern_normalizer = _make_pattern_normalizer(language)
specs = [
dataclasses.replace(
field,
analyzer=_identity_analyzer
if field.kind is FieldKind.KEYWORD
else text_analyzer,
pattern_normalizer=_fold_normalizer
if field.kind is FieldKind.KEYWORD
else pattern_normalizer,
)
for field in PUBLIC_FIELDS
]
registry = FieldRegistry(specs)
_registry_cache[language] = registry
return registry
+238 -83
View File
@@ -1,14 +1,19 @@
from __future__ import annotations
import hashlib
import json
import logging
import shutil
from typing import TYPE_CHECKING
from typing import Final
from typing import NamedTuple
from typing import cast
import tantivy
from django.conf import settings
from whoosh_compat import FieldKind
from documents.search._fields import PUBLIC_FIELDS
if TYPE_CHECKING:
from pathlib import Path
@@ -16,7 +21,201 @@ if TYPE_CHECKING:
logger = logging.getLogger("paperless.search")
# v1 - Initial tantivy schema format
SCHEMA_VERSION: Final[int] = 1
# v2 - build_schema() derived from PUBLIC_FIELDS, changing the field declaration
# order, and the write-only correspondent/document_type/storage_path/tag id
# columns dropped. tantivy compares schemas by ordered field list, so an
# index built by v1 rejects every write against the v2 schema.
SCHEMA_VERSION: Final[int] = 2
class FieldDescriptor(NamedTuple):
"""One tantivy field, in declaration order.
The descriptor vocabulary is paperless', not tantivy-py's: it is both the
input to the SchemaBuilder and the input to schema_fingerprint(), so the
persisted fingerprint cannot move under a tantivy-py upgrade.
"""
name: str
kind: str
stored: bool
indexed: bool
fast: bool
tokenizer: str | None
def _public_field_descriptors() -> list[FieldDescriptor]:
"""Descriptors for the query-visible fields declared in PUBLIC_FIELDS."""
descriptors: list[FieldDescriptor] = []
for field in PUBLIC_FIELDS:
if field.kind is FieldKind.TEXT:
descriptors.append(
FieldDescriptor(
field.name,
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
)
elif field.kind is FieldKind.KEYWORD:
descriptors.append(
FieldDescriptor(
field.name,
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="raw",
),
)
elif field.kind is FieldKind.U64:
descriptors.append(
FieldDescriptor(
field.name,
"u64",
stored=True,
indexed=True,
fast=field.fast,
tokenizer=None,
),
)
elif field.kind in (FieldKind.DATE, FieldKind.DATETIME):
descriptors.append(
FieldDescriptor(
field.name,
"date",
stored=True,
indexed=True,
fast=field.fast,
tokenizer=None,
),
)
elif field.kind is FieldKind.JSON:
descriptors.append(
FieldDescriptor(
field.name,
"json",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
)
if field.name == "notes":
# Plain-text companion for snippet generation — tantivy's
# SnippetGenerator does not support JSON fields. Schema-only,
# no query-syntax meaning, not in PUBLIC_FIELDS.
descriptors.append(
FieldDescriptor(
"notes_text",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
)
return descriptors
def field_descriptors() -> list[FieldDescriptor]:
"""Every field of the document index, in the order tantivy declares them.
tantivy compares schemas by *ordered* field list, so the order here is
part of the on-disk contract: schema_fingerprint() hashes it and
needs_rebuild() acts on the result.
"""
return [
FieldDescriptor(
"id",
"u64",
stored=True,
indexed=True,
fast=True,
tokenizer=None,
),
*_public_field_descriptors(),
# Shadow sort fields - fast, not stored
*(
FieldDescriptor(
name,
"text",
stored=False,
indexed=True,
fast=True,
tokenizer="simple_analyzer",
)
for name in ("title_sort", "correspondent_sort", "type_sort")
),
# CJK support - not stored, indexed only
*(
FieldDescriptor(
name,
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="bigram_analyzer",
)
for name in (
"bigram_content",
"bigram_title",
"bigram_correspondent",
"bigram_document_type",
"bigram_tag",
)
),
# Simple substring search support for title/content - not stored,
# indexed only
*(
FieldDescriptor(
name,
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="simple_search_analyzer",
)
for name in ("simple_title", "simple_content")
),
# Autocomplete prefix scan via terms_with_prefix, which walks the
# field's term dictionary - so the field must be indexed (term dict),
# not stored. The stored value is never read back, so storing it only
# wastes space.
FieldDescriptor(
"autocomplete_word",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="raw",
),
# Permission filter columns, read by build_permission_filter.
*(
FieldDescriptor(
name,
"u64",
stored=False,
indexed=True,
fast=True,
tokenizer=None,
)
for name in ("owner_id", "viewer_id", "viewer_group_id")
),
]
def schema_fingerprint() -> str:
"""Hash of the field descriptors, stamped into .index_settings.json.
Changes whenever a field is added, removed, retyped, re-optioned or
reordered, so an index built from a different schema shape is detected
even when SCHEMA_VERSION was not bumped.
"""
payload = json.dumps([list(descriptor) for descriptor in field_descriptors()])
return hashlib.blake2b(payload.encode()).hexdigest()
def build_schema() -> tantivy.Schema:
@@ -32,85 +231,37 @@ def build_schema() -> tantivy.Schema:
"""
sb = tantivy.SchemaBuilder()
sb.add_unsigned_field("id", stored=True, indexed=True, fast=True)
sb.add_text_field("checksum", stored=True, tokenizer_name="raw")
for field in (
"title",
"correspondent",
"document_type",
"storage_path",
"original_filename",
"content",
):
sb.add_text_field(field, stored=True, tokenizer_name="paperless_text")
# Shadow sort fields - fast, not stored/indexed
for field in ("title_sort", "correspondent_sort", "type_sort"):
sb.add_text_field(
field,
stored=False,
tokenizer_name="simple_analyzer",
fast=True,
)
# CJK support - not stored, indexed only
sb.add_text_field("bigram_content", stored=False, tokenizer_name="bigram_analyzer")
sb.add_text_field("bigram_title", stored=False, tokenizer_name="bigram_analyzer")
sb.add_text_field(
"bigram_correspondent",
stored=False,
tokenizer_name="bigram_analyzer",
)
sb.add_text_field(
"bigram_document_type",
stored=False,
tokenizer_name="bigram_analyzer",
)
sb.add_text_field("bigram_tag", stored=False, tokenizer_name="bigram_analyzer")
# Simple substring search support for title/content - not stored, indexed only
sb.add_text_field(
"simple_title",
stored=False,
tokenizer_name="simple_search_analyzer",
)
sb.add_text_field(
"simple_content",
stored=False,
tokenizer_name="simple_search_analyzer",
)
# Autocomplete prefix scan via terms_with_prefix, which walks the field's
# term dictionary - so the field must be indexed (term dict), not stored.
# The stored value is never read back, so storing it only wastes space.
sb.add_text_field("autocomplete_word", stored=False, tokenizer_name="raw")
sb.add_text_field("tag", stored=True, tokenizer_name="paperless_text")
# JSON fields — structured queries: notes.user:alice, custom_fields.name:invoice
sb.add_json_field("notes", stored=True, tokenizer_name="paperless_text")
# Plain-text companion for notes — tantivy's SnippetGenerator does not support
# JSON fields, so highlights require a text field with the same content.
sb.add_text_field("notes_text", stored=True, tokenizer_name="paperless_text")
sb.add_json_field("custom_fields", stored=True, tokenizer_name="paperless_text")
for field in (
"correspondent_id",
"document_type_id",
"storage_path_id",
"tag_id",
"owner_id",
"viewer_id",
"viewer_group_id",
):
sb.add_unsigned_field(field, stored=False, indexed=True, fast=True)
for field in ("created", "modified", "added"):
sb.add_date_field(field, stored=True, indexed=True, fast=True)
for field in ("asn", "page_count", "num_notes"):
sb.add_unsigned_field(field, stored=True, indexed=True, fast=True)
for descriptor in field_descriptors():
if descriptor.kind == "text":
sb.add_text_field(
descriptor.name,
stored=descriptor.stored,
fast=descriptor.fast,
tokenizer_name=cast("str", descriptor.tokenizer),
)
elif descriptor.kind == "json":
sb.add_json_field(
descriptor.name,
stored=descriptor.stored,
fast=descriptor.fast,
tokenizer_name=cast("str", descriptor.tokenizer),
)
elif descriptor.kind == "u64":
sb.add_unsigned_field(
descriptor.name,
stored=descriptor.stored,
indexed=descriptor.indexed,
fast=descriptor.fast,
)
elif descriptor.kind == "date":
sb.add_date_field(
descriptor.name,
stored=descriptor.stored,
indexed=descriptor.indexed,
fast=descriptor.fast,
)
else:
raise ValueError(f"Unknown schema field kind: {descriptor.kind}")
return sb.build()
@@ -119,9 +270,9 @@ def needs_rebuild(index_dir: Path) -> bool:
"""
Check if the search index needs rebuilding.
Reads .index_settings.json to compare the stored schema version and
search language against the current configuration. Returns True if the
file is missing, unparsable, or either value mismatches.
Reads .index_settings.json to compare the stored schema version, search
language and schema fingerprint against the current configuration. Returns
True if the file is missing, unparsable, or any value mismatches.
Args:
index_dir: Path to the search index directory
@@ -140,6 +291,9 @@ def needs_rebuild(index_dir: Path) -> bool:
if "language" not in data or data["language"] != settings.SEARCH_LANGUAGE:
logger.info("Search index language changed - rebuilding.")
return True
if data.get("schema_fingerprint") != schema_fingerprint():
logger.info("Search index schema fingerprint mismatch - rebuilding.")
return True
except ValueError:
return True
return False
@@ -170,6 +324,7 @@ def _write_sentinels(index_dir: Path) -> None:
{
"schema_version": SCHEMA_VERSION,
"language": settings.SEARCH_LANGUAGE,
"schema_fingerprint": schema_fingerprint(),
},
),
)
+51 -2
View File
@@ -1,6 +1,7 @@
from __future__ import annotations
import logging
from functools import cache
from typing import Final
import tantivy
@@ -71,7 +72,7 @@ def register_tokenizers(index: tantivy.Index, language: str | None) -> None:
use fast=True and Tantivy requires fast-field tokenizers to exist
even for documents that omit those fields.
"""
index.register_tokenizer("paperless_text", _paperless_text(language))
index.register_tokenizer("paperless_text", paperless_text_analyzer(language))
index.register_tokenizer("simple_analyzer", _simple_analyzer())
index.register_tokenizer("bigram_analyzer", _bigram_analyzer())
index.register_tokenizer("simple_search_analyzer", _simple_search_analyzer())
@@ -79,7 +80,7 @@ def register_tokenizers(index: tantivy.Index, language: str | None) -> None:
index.register_fast_field_tokenizer("simple_analyzer", _simple_analyzer())
def _paperless_text(language: str | None) -> tantivy.TextAnalyzer:
def paperless_text_analyzer(language: str | None) -> tantivy.TextAnalyzer:
"""Main full-text tokenizer for content, title, etc: simple -> remove_long(129) -> lowercase -> ascii_fold [-> stemmer]"""
builder = (
tantivy.TextAnalyzerBuilder(tantivy.Tokenizer.simple())
@@ -100,6 +101,54 @@ def _paperless_text(language: str | None) -> tantivy.TextAnalyzer:
return builder.build()
@cache
def _pattern_stemmer(language: str | None) -> tantivy.TextAnalyzer | None:
"""The stemming tail of paperless_text_analyzer, over a whole literal run.
Same language gate and same Snowball stemmer paperless_text_analyzer
applies at index time, so query patterns follow SEARCH_LANGUAGE. Returns
None when that gate disables stemming; paperless_text_analyzer already
warns about an unsupported language, so this stays quiet.
The raw tokenizer keeps the run whole (a wildcard literal is a fragment,
not necessarily a word), and remove_long is kept so an over-long run is
treated the same way the index treats it.
"""
if not language:
return None
tantivy_lang = _LANGUAGE_MAP.get(language.lower())
if tantivy_lang is None:
return None
return (
tantivy.TextAnalyzerBuilder(tantivy.Tokenizer.raw())
.filter(tantivy.Filter.remove_long(_TOKEN_REMOVE_LONG_LIMIT))
.filter(tantivy.Filter.stemmer(tantivy_lang))
.build()
)
def stem_pattern_text(text: str, language: str | None) -> str:
"""Stem an already lowercased/ascii-folded run the way index terms are.
Returns text unchanged when stemming is disabled for language, and also
when the stem step does not yield exactly one token: remove_long drops a run
past the length limit, leaving no stem to substitute. Falling back to the
text as typed is the safe direction for a pattern prefix, since it can only
be as narrow as it was before stemming was considered.
The raw tokenizer emits one token whatever the input and the stemmer is
1-to-1, so only the zero-token case can fire today; the guard covers both
counts so a tokenizer change cannot turn this into an IndexError.
"""
analyzer = _pattern_stemmer(language)
if analyzer is None:
return text
tokens = analyzer.analyze(text)
if len(tokens) != 1:
return text
return tokens[0]
def _simple_analyzer() -> tantivy.TextAnalyzer:
"""Tokenizer for shadow sort fields (title_sort, correspondent_sort, type_sort): simple -> lowercase -> ascii_fold."""
return (
-610
View File
@@ -1,610 +0,0 @@
from __future__ import annotations
from dataclasses import dataclass
from datetime import UTC
from datetime import datetime
from datetime import timedelta
from typing import TYPE_CHECKING
from typing import TypeAlias
import regex
from dateutil.relativedelta import relativedelta
from documents.search._dates import _DATE_KEYWORDS
from documents.search._dates import _DATE_ONLY_FIELDS
from documents.search._dates import _date_only_range
from documents.search._dates import _datetime_range
from documents.search._dates import _field_range_from_dates
from documents.search._dates import _fmt
from documents.search._dates import _precision_bounds
from documents.search._dates import _utc_bounds_for_field
# Compiled regex that matches any known multi-word (or single-word) date keyword
# at the start of a match position, longest alternatives first so "previous week"
# wins over a hypothetical shorter "previous".
_KEYWORD_VALUE_RE = regex.compile(
"|".join(sorted((regex.escape(k) for k in _DATE_KEYWORDS), key=len, reverse=True)),
regex.IGNORECASE,
)
if TYPE_CHECKING:
from datetime import tzinfo
# TODO: this module translates date queries into Tantivy *string* syntax, which
# forces a workaround for something Tantivy's string parser cannot express on
# date fields: open-ended ranges use far-past/far-future string sentinels
# (OPEN_LO/OPEN_HI). These can be replaced with a real tantivy.Query object
# (Query.range_query(..., None) for open bounds) once tantivy-py accepts Python
# datetimes in range_query/term_query on Date fields. That support exists on
# tantivy-py master (PRs #655 + #666) but postdates the pinned 0.26.0 wheel, so
# it is blocked only on a published release > 0.26.0 and a dependency bump.
# (Unparsable dates now raise InvalidDateQuery -> HTTP 400 rather than using a
# no-match string sentinel.)
# Fields that store exact, non-analyzed comma-joined tokens in the index and so
# need explicit comma->AND expansion (Whoosh KEYWORD(commas=True) set).
MULTI_VALUE_FIELDS = frozenset({"tag", "tag_id", "viewer_id"})
# Date fields whose values/ranges get rewritten to RFC3339 Tantivy ranges.
DATE_FIELDS = frozenset({"created", "modified", "added"})
# Field aliases: Whoosh (v2) field names that were renamed in the Tantivy schema.
# Preserved here so v2 queries using the old names continue to work without 400
# errors instead of silently failing. Applied by _render to non-date field tokens.
FIELD_ALIASES: dict[str, str] = {
"type": "document_type",
"type_id": "document_type_id",
"path": "storage_path",
"path_id": "storage_path_id",
}
# Known schema fields: a comma immediately followed by ``<known>:`` is a clause
# separator. Restricting to known fields prevents URL-like ``http:`` misfires.
KNOWN_FIELDS = frozenset(
{
"title",
"content",
"correspondent",
"document_type",
"type", # v2 alias -> document_type
"storage_path",
"path", # v2 alias -> storage_path
"tag",
"tag_id",
"correspondent_id",
"document_type_id",
"type_id", # v2 alias -> document_type_id
"storage_path_id",
"path_id", # v2 alias -> storage_path_id
"owner_id",
"viewer_id",
"asn",
"page_count",
"num_notes",
"created",
"modified",
"added",
"original_filename",
"checksum",
"notes",
"custom_fields",
},
)
_FIELD_RE = regex.compile(r"(?P<field>\w+):")
# Matches the TO separator inside a range bracket. Handles three forms:
# middle: "lo TO hi" (either lo or hi may be empty)
# trailing: "lo TO" (open upper bound)
# leading: "TO hi" (open lower bound)
# Bounds MAY contain internal spaces (e.g. "-7 days"), so we use .*? / .+?
# and split on the whitespace-delimited " TO " / " to " separator.
_RANGE_RE = regex.compile(
r"^\s*(?P<lo>.*?)\s+[Tt][Oo]\s+(?P<hi>.+?)\s*$"
r"|"
r"^\s*(?P<lo2>.+?)\s+[Tt][Oo]\s*$"
r"|"
r"^\s*[Tt][Oo]\s+(?P<hi2>.+?)\s*$",
)
@dataclass(frozen=True, slots=True)
class FieldValue:
field: str
value: str
# Produced by the comma-resolution pass (not by scan()).
@dataclass(frozen=True, slots=True)
class FieldValueList:
field: str
values: tuple[str, ...]
@dataclass(frozen=True, slots=True)
class FieldRange:
field: str
open: str
lo: str
hi: str
close: str
# Produced by the comma-resolution pass (not by scan()).
@dataclass(frozen=True, slots=True)
class Comma:
pass
@dataclass(frozen=True, slots=True)
class Passthrough:
raw: str
Token: TypeAlias = FieldValue | FieldValueList | FieldRange | Comma | Passthrough
_CLOSE: dict[str, str] = {"[": "]", "{": "}"}
def scan(query: str) -> list[Token]:
"""
Tokenize a raw query into date/comma-aware tokens, leaving everything else
as verbatim ``Passthrough`` runs. Non-recursive: finds the first matching
close bracket/quote. Nested brackets are not valid Tantivy range syntax and
pass through verbatim on mismatch.
"""
tokens: list[Token] = []
buf: list[str] = [] # accumulates passthrough chars
i, n = 0, len(query)
while i < n:
matched = _match_field_token(query, i)
if matched is None:
buf.append(query[i])
i += 1
continue
token, i = matched
if buf and buf[-1] == ",":
buf.pop()
_flush(buf, tokens)
tokens.append(Comma())
else:
_flush(buf, tokens)
tokens.append(token)
i = _maybe_comma(query, i, tokens)
_flush(buf, tokens)
return tokens
def _flush(buf: list[str], tokens: list[Token]) -> None:
"""Emit any accumulated passthrough characters as a single token."""
if buf:
tokens.append(Passthrough("".join(buf)))
buf.clear()
def _at_word_boundary(query: str, i: int) -> bool:
"""A field token may begin only at the start or after a non-word character."""
return i == 0 or not (query[i - 1].isalnum() or query[i - 1] == "_")
def _match_field_token(query: str, i: int) -> tuple[Token, int] | None:
"""
If a known ``field:`` token starts at ``i``, consume it and return
``(token, end_index)``; otherwise return None so the caller treats the
character as passthrough. Handles both ``field:[range]`` and ``field:value``,
and returns None when the range/value cannot be consumed.
"""
m = _FIELD_RE.match(query, i)
if m is None or m.group("field") not in KNOWN_FIELDS:
return None
if not _at_word_boundary(query, i):
return None
field = m.group("field")
j = m.end()
if j < len(query) and query[j] in "[{":
return _consume_range(query, j, field)
consumed = _consume_field_value(query, field, j)
if consumed is None:
return None
value, end = consumed
return FieldValue(field, value), end
def _consume_field_value(query: str, field: str, start: int) -> tuple[str, int] | None:
"""
Consume a field value starting at ``start``: a multi-word date keyword phrase
(date fields only), or a bare/quoted value, then absorb any comma-joined
continuation that is not a clause separator. ``resolve_commas`` later splits a
multi-value field's joined value into a ``FieldValueList``; for other fields
the comma stays literal.
"""
n = len(query)
consumed = None
if field in DATE_FIELDS:
km = _KEYWORD_VALUE_RE.match(query, start)
if km is not None and (km.end() >= n or query[km.end()] in " \t),"):
consumed = (km.group(0), km.end())
if consumed is None:
consumed = _consume_value(query, start)
if consumed is None:
return None
value, k = consumed
while k < n and query[k] == ",":
if _looks_like_known_field(query, k + 1):
break # clause separator: left for _maybe_comma to emit a Comma()
more = _consume_value(query, k + 1)
if more is None:
break
value = f"{value},{more[0]}"
k = more[1]
return value, k
def _consume_range(
query: str,
start: int,
field: str,
) -> tuple[FieldRange, int] | None:
"""Consume ``[lo TO hi]`` / ``{lo TO hi}`` from ``start`` (the bracket)."""
open_br = query[start]
close_br = _CLOSE[open_br]
end = query.find(close_br, start + 1)
if end == -1:
return None
inner = query[start + 1 : end]
m = _RANGE_RE.match(inner)
if m is not None:
if m.group("lo") is not None or m.group("hi") is not None:
# Middle form: "lo TO hi" (either may be empty string)
lo = (m.group("lo") or "").strip()
hi = (m.group("hi") or "").strip()
elif m.group("lo2") is not None:
# Trailing form: "lo TO"
lo = m.group("lo2").strip()
hi = ""
else:
# Leading form: "TO hi"
lo = ""
hi = (m.group("hi2") or "").strip()
else:
lo, hi = inner.strip(), ""
return FieldRange(field, open_br, lo, hi, close_br), end + 1
def _consume_value(query: str, start: int) -> tuple[str, int] | None:
"""Consume a bare or quoted field value from ``start``, stopping at comma."""
n = len(query)
if start >= n or query[start] in " \t":
return None
if query[start] in "\"'":
quote = query[start]
end = query.find(quote, start + 1)
if end == -1:
return None
return query[start : end + 1], end + 1
j = start
while j < n and query[j] not in " \t),":
j += 1
return query[start:j], j
def _looks_like_known_field(query: str, pos: int) -> bool:
"""True if a known ``field:`` token starts at ``pos``."""
m = _FIELD_RE.match(query, pos)
return bool(m and m.group("field") in KNOWN_FIELDS)
def _maybe_comma(query: str, i: int, tokens: list) -> int:
"""If a clause-separator comma follows at ``i``, emit ``Comma()`` and advance."""
if i < len(query) and query[i] == "," and _looks_like_known_field(query, i + 1):
tokens.append(Comma())
return i + 1
return i
def resolve_commas(tokens: list) -> list:
"""
Collapse value-list commas into ``FieldValueList`` and keep clause-separator
commas as ``Comma``. (Clause-sep commas are already emitted by ``scan`` via
the value-stop logic; this pass folds value-lists.)
"""
out: list = []
for tok in tokens:
if (
isinstance(tok, FieldValue)
and tok.field in MULTI_VALUE_FIELDS
and "," in tok.value
):
values = tuple(v for v in tok.value.split(",") if v)
out.append(FieldValueList(tok.field, values))
else:
out.append(tok)
return out
class SearchQueryError(ValueError):
"""
Base for user-fixable search query errors.
Carries a message safe to surface to the user (no internal details). The view
layer catches this and returns an HTTP 400, so any future subclass (unknown
field, malformed range, wrapped parser errors) gets the same treatment.
"""
class InvalidDateQuery(SearchQueryError):
"""Raised when a date field value or range bound cannot be parsed."""
def __init__(self, field: str, value: str) -> None:
self.field = field
self.value = value
super().__init__(f"Invalid date value {value!r} for field {field!r}.")
_DIGITS_RE = regex.compile(r"^\d{4}(?:\d{2}){0,2}$")
_ISO_RE = regex.compile(r"^\d{4}(?:-\d{2}(?:-\d{2})?)?$")
def translate_scalar(field: str, value: str, tz: tzinfo) -> str:
"""Translate a bare date-field value to a Tantivy range string."""
bare = value.strip("\"'").lower()
if bare in _DATE_KEYWORDS:
if field in _DATE_ONLY_FIELDS:
return f"{field}:{_date_only_range(bare, tz)}"
return f"{field}:{_datetime_range(bare, tz)}"
digits = value.replace("-", "")
if _DIGITS_RE.match(value) or _ISO_RE.match(value):
bounds = _precision_bounds(digits)
if bounds is None:
raise InvalidDateQuery(field, value)
return _field_range_from_dates(field, bounds[0], bounds[1], tz)
if regex.fullmatch(r"\d{14}", value):
try:
dt = datetime(
int(value[0:4]),
int(value[4:6]),
int(value[6:8]),
int(value[8:10]),
int(value[10:12]),
int(value[12:14]),
tzinfo=UTC,
)
except ValueError:
raise InvalidDateQuery(field, value) from None
iso = _fmt(dt)
return f"{field}:[{iso} TO {iso}]"
# Unrecognized shape -> tell the user their date is malformed rather than
# silently matching nothing or emitting invalid Tantivy syntax.
raise InvalidDateQuery(field, value)
# Open-bound sentinels for date ranges. These far-past/far-future strings allow
# open-ended ranges to be expressed as Tantivy string queries until tantivy-py
# exposes Query.range_query(..., None) on Date fields (see module TODO).
OPEN_LO = "0001-01-01T00:00:00Z"
OPEN_HI = "9999-12-31T23:59:59Z"
# Matches compact now-offset tokens like now-7d, now+1h, now-30m.
_NOW_COMPACT_RE = regex.compile(
r"^now(?P<sign>[+-])(?P<n>\d+)(?P<unit>[dhm])$",
regex.IGNORECASE,
)
# Matches "±N <unit>" Whoosh-style offsets (e.g. -7 days, -1 week, +3 hours).
# Whoosh's own date parser (qparser.dateparse.PlusMinus) additionally accepted
# abbreviated unit spellings (e.g. "yrs", "yr", "y", "mos", "wks", "hrs", "mins",
# "secs"); saved views/searches created under the old Whoosh backend can still
# contain those tokens (e.g. "-999yrs"), so they are accepted here too and
# normalized to a canonical unit via _UNIT_ALIASES below.
_NOW_SPACED_RE = regex.compile(
r"^(?P<sign>[+-])(?P<n>\d+)\s*"
r"(?P<unit>years|year|yrs|yr|ys|y"
r"|months|month|mons|mon|mos|mo"
r"|weeks|week|wks|wk|ws|w"
r"|days|day|dys|dy|ds|d"
r"|hours|hour|hrs|hr|hs|h"
r"|minutes|minute|mins|min|ms|m"
r"|seconds|second|secs|sec|s)$",
regex.IGNORECASE,
)
# Maps every accepted unit spelling (including Whoosh-era abbreviations) to the
# canonical unit name used as a key into the delta map in _resolve_relative_bound.
_UNIT_ALIASES: dict[str, str] = {
alias: canonical
for canonical, aliases in {
"year": ("years", "year", "yrs", "yr", "ys", "y"),
"month": ("months", "month", "mons", "mon", "mos", "mo"),
"week": ("weeks", "week", "wks", "wk", "ws", "w"),
"day": ("days", "day", "dys", "dy", "ds", "d"),
"hour": ("hours", "hour", "hrs", "hr", "hs", "h"),
"minute": ("minutes", "minute", "mins", "min", "ms", "m"),
"second": ("seconds", "second", "secs", "sec", "s"),
}.items()
for alias in aliases
}
def _resolve_relative_bound(token: str) -> datetime | None:
"""
Resolve a relative bound token to an exact UTC instant, or return None.
Supported forms:
- ``now`` -> current UTC instant
- ``now+/-<n>d/h/m`` -> now +/- timedelta (d=days, h=hours, m=minutes)
- ``±N <unit>`` -> now +/- delta; month/year use relativedelta;
unit also accepts Whoosh-era abbreviations
(e.g. "yrs", "mos", "wks", "hrs", "mins", "secs")
"""
stripped = token.strip()
low = stripped.lower()
now = datetime.now(UTC)
if low == "now":
return now
m = _NOW_COMPACT_RE.match(stripped)
if m:
sign = 1 if m.group("sign") == "+" else -1
n = int(m.group("n"))
unit = m.group("unit").lower()
delta = (
sign
* {
"d": timedelta(days=n),
"h": timedelta(hours=n),
"m": timedelta(minutes=n),
}[unit]
)
return now + delta
m = _NOW_SPACED_RE.match(stripped)
if m:
sign = 1 if m.group("sign") == "+" else -1
n = int(m.group("n"))
unit = _UNIT_ALIASES[m.group("unit").lower()]
delta_map: dict[str, timedelta | relativedelta] = {
"second": timedelta(seconds=n),
"minute": timedelta(minutes=n),
"hour": timedelta(hours=n),
"day": timedelta(days=n),
"week": timedelta(weeks=n),
"month": relativedelta(months=n),
"year": relativedelta(years=n),
}
return now - delta_map[unit] if sign == -1 else now + delta_map[unit]
return None
def _bound_datetimes(
field: str,
token: str,
tz: tzinfo,
) -> tuple[datetime, datetime] | None:
"""
Return (floor_dt, ceil_dt) UTC datetimes for a single range bound token, or
None if the token is unparsable. ``now`` and relative offsets resolve to the
current instant (floor == ceil == that instant; no day-flooring).
"""
token = token.strip()
# Try relative/now forms first (before stripping hyphens which would mangle them).
rel = _resolve_relative_bound(token)
if rel is not None:
return rel, rel
# Full ISO datetime token (contains "T"): parse directly and return an exact
# instant (floor == ceil). Python 3.11+ datetime.fromisoformat accepts trailing Z.
if "T" in token:
try:
dt = datetime.fromisoformat(token)
# Ensure timezone-aware UTC result.
dt = dt.replace(tzinfo=UTC) if dt.tzinfo is None else dt.astimezone(UTC)
return dt, dt
except ValueError:
return None
digits = token.replace("-", "")
bounds = _precision_bounds(digits)
if bounds is None:
return None
start, end = bounds
return _utc_bounds_for_field(field, start, end, tz)
def _render(tok: Token, tz: tzinfo) -> str:
"""Render a single token back to a Tantivy query string fragment."""
if isinstance(tok, Passthrough):
return tok.raw
if isinstance(tok, Comma):
return " AND "
if isinstance(tok, FieldValueList):
field = FIELD_ALIASES.get(tok.field, tok.field)
return " AND ".join(f"{field}:{v}" for v in tok.values)
if isinstance(tok, FieldValue):
field = FIELD_ALIASES.get(tok.field, tok.field)
if field in DATE_FIELDS:
return translate_scalar(field, tok.value, tz)
return f"{field}:{tok.value}"
if isinstance(tok, FieldRange):
field = FIELD_ALIASES.get(tok.field, tok.field)
if field in DATE_FIELDS:
return translate_range(field, tok.lo, tok.hi, tz)
return f"{field}:{tok.open}{tok.lo} TO {tok.hi}{tok.close}"
return "" # pragma: no cover
# Post-render operator normalization patterns: collapse repeated whitespace and
# strip spaced/trailing Tantivy boolean operators that would otherwise be invalid.
_MULTI_SPACE_RE = regex.compile(r" {2,}")
_TRAILING_OP_RE = regex.compile(r"\s+[-+]+\s*$")
_SPACED_OP_RE = regex.compile(r"\s+[-+]\s+")
def _normalize_operators(text: str) -> str:
"""
Collapse multiple spaces, strip trailing dangling operators, and replace
spaced operators (`` - `` / `` + ``) with a single space.
Applied only to Passthrough fragments (the rendered output is scanned for
operator artifacts outside bracketed ranges) via a post-render pass on the
full rendered string. This preserves date ranges (``[... TO ...]``) verbatim
while cleaning natural-language separators in the surrounding text.
"""
text = _MULTI_SPACE_RE.sub(" ", text)
text = _TRAILING_OP_RE.sub("", text).strip()
text = _SPACED_OP_RE.sub(" ", text).strip()
return text
def translate_query(raw: str, tz: tzinfo) -> str:
"""Translate a raw Whoosh-style query into Tantivy-compatible syntax."""
tokens = resolve_commas(scan(raw))
rendered = "".join(_render(t, tz) for t in tokens)
return _normalize_operators(rendered)
def translate_range(field: str, lo: str, hi: str, tz: tzinfo) -> str:
"""Translate a date-field ``[lo TO hi]`` range to a Tantivy ISO range string.
Handles partial-date bounds (YYYY, YYYYMM, YYYYMMDD, ISO dash variants),
open bounds (empty string -> OPEN_LO/OPEN_HI), ``now``, and reversed ranges
(swaps tokens before computing floor/ceil so the span is always correct).
"""
lo_s = lo.strip()
hi_s = hi.strip()
# Parse both bounds to (floor, ceil) pairs when present.
lo_pair: tuple[datetime, datetime] | None = None
hi_pair: tuple[datetime, datetime] | None = None
if lo_s:
lo_pair = _bound_datetimes(field, lo_s, tz)
if lo_pair is None:
raise InvalidDateQuery(field, lo_s)
if hi_s:
hi_pair = _bound_datetimes(field, hi_s, tz)
if hi_pair is None:
raise InvalidDateQuery(field, hi_s)
# Detect a reversed range: only swap when BOTH bounds are present.
if lo_pair is not None and hi_pair is not None and lo_pair[0] > hi_pair[0]:
lo_pair, hi_pair = hi_pair, lo_pair
lo_iso = _fmt(lo_pair[0]) if lo_pair is not None else OPEN_LO
# A bound resolves to (floor, ceil) where floor == ceil for an exact instant
# (a full ISO datetime, "now", or a "+/-N unit" offset) and floor != ceil for
# a coarser period token (year/month/day precision). Only the latter needs a
# half-open close: its ceil is the start of the *next* period and must be
# excluded, or that instant (e.g. the 1st of next month) wrongly matches.
if hi_pair is not None:
hi_iso = _fmt(hi_pair[1])
hi_close = "]" if hi_pair[0] == hi_pair[1] else "}"
else:
hi_iso = OPEN_HI
hi_close = "]"
return f"{field}:[{lo_iso} TO {hi_iso}{hi_close}"
-12
View File
@@ -1,15 +1,11 @@
from __future__ import annotations
import tempfile
from typing import TYPE_CHECKING
import pytest
import tantivy
from documents.search._backend import TantivyBackend
from documents.search._backend import reset_backend
from documents.search._schema import build_schema
from documents.search._tokenizer import register_tokenizers
if TYPE_CHECKING:
from collections.abc import Generator
@@ -35,11 +31,3 @@ def backend() -> Generator[TantivyBackend, None, None]:
finally:
b.close()
reset_backend()
@pytest.fixture(scope="module")
def index() -> tantivy.Index:
"""A real Tantivy index for parse-acceptance tests (module scope for speed)."""
idx = tantivy.Index(build_schema(), path=tempfile.mkdtemp())
register_tokenizers(idx, "english")
return idx
@@ -0,0 +1,411 @@
"""Result-level acceptance corpus: real documents indexed via build_schema(),
real queries run through parse_user_query(), matched-document-ID sets
asserted not intermediate ASTs or query strings. This is paperless-ngx's
analogue of whoosh-compat's own tests/emitter/test_acceptance_e2e.py.
Supersedes test_query.py's TestParseUserQuery result-level cases.
"""
from __future__ import annotations
from datetime import UTC
from datetime import datetime
from typing import TYPE_CHECKING
import pytest
import time_machine
from django.contrib.auth.models import User
from documents.models import CustomField
from documents.models import CustomFieldInstance
from documents.models import Document
from documents.models import DocumentType
from documents.models import Note
from documents.models import StoragePath
from documents.search._query import parse_user_query
if TYPE_CHECKING:
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
FROZEN_NOW = datetime(2026, 6, 15, 12, 0, tzinfo=UTC)
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
"""Create a Document and index it in one step, for the common case
where nothing needs to happen between the two (no related Note/
CustomFieldInstance to attach first)."""
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
@pytest.fixture
def indexed_documents(backend: TantivyBackend) -> dict[str, int]:
"""Index a small fixture set, return {label: doc_id} for corpus queries."""
docs = {
"invoice_2020": _index(
backend,
title="Invoice 2020",
content="invoice total due",
checksum="acc-invoice-2020",
archive_serial_number=100,
),
"invoice_2021": _index(
backend,
title="Invoice 2021",
content="invoice total due",
checksum="acc-invoice-2021",
archive_serial_number=101,
),
"invoice_2023": _index(
backend,
title="Invoice 2023",
content="invoice total due",
checksum="acc-invoice-2023",
archive_serial_number=102,
),
"receipt_2022": _index(
backend,
title="Receipt 2022",
content="receipt total due",
checksum="acc-receipt-2022",
archive_serial_number=103,
),
}
return {label: doc.pk for label, doc in docs.items()}
class TestIssue13568BracketWildcard:
"""paperless-ngx#13568: title:202[0-3]* must keep its character class,
not fold to a prefix query that silently drops it."""
def test_bracket_class_wildcard_matches_only_in_range_years(
self,
backend: TantivyBackend,
indexed_documents: dict[str, int],
) -> None:
# [0-1] (not [0-3]) is deliberate: the fixture's four years are
# 2020/2021/2022/2023, i.e. their trailing digit is 0/1/2/3
# respectively - a [0-3] class would match all four and the test
# would pass even if the character class were silently dropped and
# folded to an unconstrained "202*" prefix. [0-1] partitions the
# fixture into a genuine in-range/out-of-range split.
matched = _matched_ids(backend, "title:202[0-1]*")
expected = {
indexed_documents["invoice_2020"],
indexed_documents["invoice_2021"],
}
assert matched == expected, (
"title:202[0-1]* must match 2020/2021 titles and exclude 2022/2023 "
"- if this matches everything, the wildcard's character class was "
"silently dropped (issue #13568's original bug)"
)
class TestFieldBoosts:
def test_title_boost_ranks_title_match_above_content_only_match(
self,
backend: TantivyBackend,
) -> None:
title_match = _index(
backend,
title="urgent",
content="nothing else relevant",
checksum="acc-boost-title",
)
_index(
backend,
title="nothing",
content="urgent matter here",
checksum="acc-boost-content",
)
query = parse_user_query(backend._index, "urgent", UTC)
searcher = backend._index.searcher()
results = searcher.search(query, limit=10)
ranked_ids = [
searcher.doc(addr).to_dict()["id"][0] for _score, addr in results.hits
]
assert ranked_ids[0] == title_match.pk
class TestJsonSubpaths:
def test_notes_user_matches_document_with_that_note_author(
self,
backend: TantivyBackend,
) -> None:
alice = User.objects.create_user(username="alice")
doc_with_note = Document.objects.create(
title="Has note",
content="x",
checksum="acc-note-with",
)
Note.objects.create(document=doc_with_note, user=alice, note="reminder")
backend.add_or_update(doc_with_note)
_index(backend, title="No note", content="x", checksum="acc-note-without")
matched = _matched_ids(backend, "notes.user:alice")
assert matched == {doc_with_note.pk}
def test_custom_fields_name_and_value_combine(
self,
backend: TantivyBackend,
) -> None:
field = CustomField.objects.create(
name="Contract Number",
data_type=CustomField.FieldDataType.STRING,
)
other_field = CustomField.objects.create(
name="Other Field",
data_type=CustomField.FieldDataType.STRING,
)
matching = Document.objects.create(
title="Matching",
content="x",
checksum="acc-cf-matching",
)
CustomFieldInstance.objects.create(
document=matching,
field=field,
value_text="policy",
)
backend.add_or_update(matching)
non_matching = Document.objects.create(
title="Non-matching",
content="x",
checksum="acc-cf-nonmatching",
)
CustomFieldInstance.objects.create(
document=non_matching,
field=other_field,
value_text="policy",
)
backend.add_or_update(non_matching)
matched = _matched_ids(
backend,
'custom_fields.name:"Contract Number" custom_fields.value:policy',
)
assert matched == {matching.pk}
class TestUnregisteredIdFieldFoldsToLiteralText:
"""tag_id, owner_id, etc. are intentionally excluded from the
FieldRegistry - always internal index columns, never meant to be
query-addressable. Prove an unregistered field folds to a literal
text search that matches nothing, rather than erroring."""
def test_tag_id_query_matches_nothing(
self,
backend: TantivyBackend,
indexed_documents: dict[str, int],
) -> None:
matched = _matched_ids(backend, "tag_id:5")
assert matched == set()
class TestFuzzyBlendSurvivesWhooshGrammar:
"""A query mixing whoosh-only grammar (a date keyword) with a typo'd
free-text word must still fuzzy-match the intended document when
ADVANCED_FUZZY_SEARCH_THRESHOLD is enabled. The fuzzy clause is built
from the parsed query's free-text tokens (whoosh_compat's
free_text_tokens), never from the raw query string, so whoosh grammar
that tantivy's own parser rejects cannot knock the fuzzy clause out."""
def test_typo_fuzzy_matches_alongside_date_keyword(
self,
backend: TantivyBackend,
settings,
) -> None:
settings.ADVANCED_FUZZY_SEARCH_THRESHOLD = 0.5
with time_machine.travel(FROZEN_NOW, tick=False):
doc = _index(
backend,
title="Receipt March",
content="receipt total due",
checksum="fuzzy-blend-1",
archive_serial_number=900,
)
# Sanity: the exact spelling matches through the exact clause.
assert doc.pk in _matched_ids(backend, "added:today receipt")
# The regression: the misspelling (one transposition) only
# matches via the fuzzy clause, and "added:today" is
# whoosh-only grammar tantivy's parser rejects, so raw-string
# fuzzy parsing skips the clause entirely and this returns
# nothing. The typo is deliberate; keep codespell away from it.
typo_query = "added:today reciept" # codespell:ignore reciept
assert doc.pk in _matched_ids(backend, typo_query)
def test_negated_words_do_not_fuzzy_match(
self,
backend: TantivyBackend,
settings,
) -> None:
# A term the user excluded must not resurface through the fuzzy
# clause. The shape is chosen so this genuinely discriminates: the
# indexed document contains the NOT'd word but NOT the positive
# word, so nothing matches the exact clause, and a fuzzy string
# naively built from ALL words (including the NOT'd one) would
# make this document the sole hit, normalize its score to 1.0,
# and survive any threshold. (A shape with an exact-matching
# sibling document does NOT discriminate: normalization ranks the
# resurfaced doc far below the exact match and the threshold cuts
# it even for a naive implementation.)
settings.ADVANCED_FUZZY_SEARCH_THRESHOLD = 0.5
with time_machine.travel(FROZEN_NOW, tick=False):
_index(
backend,
title="Receipt Archive",
content="receipt archived stack",
checksum="fuzzy-blend-2",
archive_serial_number=901,
)
assert _matched_ids(backend, "added:today total NOT receipt") == set()
class TestUnquotedDateKeywordPhrases:
"""The unquoted spelling (added:previous month) is honored natively by
whoosh-compat's own grammar for this closed phrase vocabulary — no
app-level rewrite is involved. Pins that the historically supported
spelling keeps working now that paperless no longer pre-quotes it."""
@pytest.fixture
def period_documents(self, backend: TantivyBackend) -> dict[str, int]:
with time_machine.travel(FROZEN_NOW, tick=False):
in_may = _index(
backend,
title="May Doc",
content="statement",
checksum="kw-may",
archive_serial_number=910,
added=datetime(2026, 5, 20, 12, 0, tzinfo=UTC),
)
in_june = _index(
backend,
title="June Doc",
content="statement",
checksum="kw-june",
archive_serial_number=911,
added=datetime(2026, 6, 10, 12, 0, tzinfo=UTC),
)
return {"in_may": in_may.pk, "in_june": in_june.pk}
@pytest.mark.parametrize(
"query",
[
pytest.param("added:previous month", id="unquoted"),
pytest.param('added:"previous month"', id="quoted"),
pytest.param("added:Previous Month", id="unquoted-mixed-case"),
],
)
def test_unquoted_matches_the_same_documents_as_quoted(
self,
backend: TantivyBackend,
period_documents: dict[str, int],
query: str,
) -> None:
with time_machine.travel(FROZEN_NOW, tick=False):
assert _matched_ids(backend, query) == {period_documents["in_may"]}
@pytest.mark.parametrize(
"query",
[
pytest.param("added:this month", id="this-month"),
pytest.param("added:this year", id="this-year"),
pytest.param("added:previous week", id="previous-week"),
pytest.param("added:previous quarter", id="previous-quarter"),
pytest.param("added:previous year", id="previous-year"),
pytest.param("created:previous month", id="created-field"),
pytest.param("modified:previous month", id="modified-field"),
],
)
def test_every_phrase_and_date_field_parses_without_error(
self,
backend: TantivyBackend,
period_documents: dict[str, int],
query: str,
) -> None:
# The whole vocabulary times every date field must at least parse
# and search cleanly (no SearchQueryError -> no HTTP 400); exact
# window semantics are whoosh-compat's, pinned in its own suite.
with time_machine.travel(FROZEN_NOW, tick=False):
_matched_ids(backend, query)
def test_text_field_keyword_words_are_ordinary_text(
self,
backend: TantivyBackend,
period_documents: dict[str, int],
) -> None:
# "previous month" after a TEXT field (or unfielded) is ordinary
# text, not a date phrase: a title actually containing the words
# matches, and the date-window documents do not.
with time_machine.travel(FROZEN_NOW, tick=False):
wordy = _index(
backend,
title="Notes from the previous month",
content="meeting notes",
checksum="kw-text",
archive_serial_number=912,
)
assert _matched_ids(backend, "title:previous month") == {wordy.pk}
class TestFieldAliases:
"""type:/path: are registry aliases for document_type:/storage_path:.
The only other alias coverage is parse-shape; these prove resolution
end-to-end against a real index."""
def test_type_alias_and_canonical_name_match_the_same_document(
self,
backend: TantivyBackend,
) -> None:
invoice_type = DocumentType.objects.create(name="invoice")
# Discriminating shape: document_type is itself a default search
# field, so if alias resolution ever broke and "type:invoice"
# demoted to unfielded text, the token would STILL match the typed
# document through the field value. The decoy carries the query
# word in content, so a demoted search matches BOTH documents and
# the exact-set assertions fail. (The title avoids stemming to
# "type": english stems Typed -> type.)
typed = _index(
backend,
title="First",
content="quarterly statement",
checksum="alias-type-1",
document_type=invoice_type,
)
_index(
backend,
title="Second",
content="invoice mentioned in body",
checksum="alias-type-2",
)
assert _matched_ids(backend, "type:invoice") == {typed.pk}
assert _matched_ids(backend, "document_type:invoice") == {typed.pk}
def test_path_alias_and_canonical_name_match_the_same_document(
self,
backend: TantivyBackend,
) -> None:
archive = StoragePath.objects.create(name="archive", path="archive/{title}")
stored = _index(
backend,
title="Stored",
content="quarterly statement",
checksum="alias-path-1",
storage_path=archive,
)
# storage_path is NOT a default search field today, so a demoted
# "path:archive" already matches nothing; the content decoy keeps
# this test discriminating even if it ever joins the defaults.
_index(
backend,
title="Loose",
content="archive mentioned in body",
checksum="alias-path-2",
)
assert _matched_ids(backend, "path:archive") == {stored.pk}
assert _matched_ids(backend, "storage_path:archive") == {stored.pk}
@@ -0,0 +1,133 @@
"""The CJK bigram clause blended into QUERY-mode searches.
The clause exists so CJK runs are matchable at all (the default analyzers
keep a whitespace-free CJK run as one indivisible token), but it must not
widen the query beyond what the user asked for: a CJK term the query
excludes, or restricts to one field, must not come back through it.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from documents.models import Document
if TYPE_CHECKING:
from pytest_django.fixtures import SettingsWrapper
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
class TestCjkClauseFollowsTheParsedQuery:
def test_negated_cjk_term_is_excluded(self, backend: TantivyBackend) -> None:
"""'invoice NOT 漢字' must not return the document containing 漢字."""
with_cjk = _index(
backend,
title="Invoice A",
content="invoice total 漢字",
checksum="cjk-neg-1",
)
without_cjk = _index(
backend,
title="Invoice B",
content="invoice total only",
checksum="cjk-neg-2",
)
assert _matched_ids(backend, "invoice") == {with_cjk.pk, without_cjk.pk}
assert _matched_ids(backend, "invoice NOT 漢字") == {without_cjk.pk}
@pytest.mark.parametrize(
("threshold", "expected"),
[
pytest.param(None, {"titled"}, id="fuzzy_off"),
pytest.param(0.0, {"titled", "content_only"}, id="fuzzy_on"),
],
)
def test_fielded_cjk_term_searches_only_that_field(
self,
backend: TantivyBackend,
settings: SettingsWrapper,
threshold: float | None,
expected: set[str],
) -> None:
"""'title:東京' must not match a document whose 東京 is in the content.
The CJK clause honours the field. The fuzzy clause, when enabled,
does not: it contributes every free-text term UNFIELDED by design
(see _try_parse_fuzzy_query), so it brings the content-only
document back on its own 0.1-boosted terms. That is the documented
trade-off, pinned here so it stays deliberate.
"""
settings.ADVANCED_FUZZY_SEARCH_THRESHOLD = threshold
content_only = _index(
backend,
title="Tokyo report",
content="東京都の人口は約1400万人です",
checksum="cjk-field-1",
)
titled = _index(
backend,
title="東京都の報告書",
content="an english summary",
checksum="cjk-field-2",
)
pks = {"titled": titled.pk, "content_only": content_only.pk}
assert _matched_ids(backend, "東京") == set(pks.values())
assert _matched_ids(backend, "title:東京") == {pks[label] for label in expected}
def test_cjk_on_a_non_default_field_builds_no_clause(
self,
backend: TantivyBackend,
) -> None:
"""A CJK term restricted to a field outside the default search fields
has nothing to contribute to the bigram clause: 'notes:東京' must not
fall back to matching 東京 in the content."""
_index(
backend,
title="Tokyo report",
content="東京都の人口は約1400万人です",
checksum="cjk-notes-1",
)
assert _matched_ids(backend, "notes:東京") == set()
def test_bare_cjk_term_still_matches_every_default_field(
self,
backend: TantivyBackend,
) -> None:
"""The clause's reason for existing: an unfielded CJK run matches
wherever it is indexed, and does so alongside a latin term."""
in_content = _index(
backend,
title="report",
content="本文に重要な情報",
checksum="cjk-bare-1",
)
in_title = _index(
backend,
title="重要な報告書",
content="english only",
checksum="cjk-bare-2",
)
assert _matched_ids(backend, "重要") == {in_content.pk, in_title.pk}
assert _matched_ids(backend, "重要 OR report") == {
in_content.pk,
in_title.pk,
}
@@ -0,0 +1,75 @@
"""Whoosh's compact, separator-free date spelling, resolved end to end.
whoosh-compat owns both widths of this spelling and asserts both of each
form's bounds directly: ``test_compact_numeric_datetime`` pins the 8-digit
form as a whole calendar day (lower bound, upper bound and exclusivity), and
``test_compact_numeric_datetime_full_width_is_a_single_second_instant`` pins
the 14-digit form as one instant. The 14-digit form is kept here as the single
representative because it is the one that exercises paperless's ``added``
DATETIME fast field at full precision: the corpus separates a document at
the named instant from one on the same calendar day at another hour and one
on the next day at the same hour, so a query that degrades into a whole-day
window, or drops the time of day, matches the wrong set rather than passing
on a corpus that could not tell the difference.
"""
from __future__ import annotations
from datetime import UTC
from datetime import datetime
from typing import TYPE_CHECKING
import pytest
from documents.models import Document
if TYPE_CHECKING:
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
@pytest.fixture
def docs(backend: TantivyBackend) -> dict[str, int]:
return {
"instant": _index(
backend,
title="On the instant",
content="x",
checksum="compact-date-instant",
added=datetime(2005, 3, 4, 15, 30, tzinfo=UTC),
).pk,
"same_day": _index(
backend,
title="Same day, other hour",
content="x",
checksum="compact-date-same-day",
added=datetime(2005, 3, 4, 9, 0, tzinfo=UTC),
).pk,
"next_day": _index(
backend,
title="Next day, same hour",
content="x",
checksum="compact-date-next-day",
added=datetime(2005, 3, 5, 15, 30, tzinfo=UTC),
).pk,
}
def test_fourteen_digits_is_a_single_instant(
backend: TantivyBackend,
docs: dict[str, int],
) -> None:
# same_day is what tells this apart from the 8-digit day-window form,
# next_day from a form that ignored the time altogether.
assert _matched_ids(backend, "added:20050304153000") == {docs["instant"]}
@@ -0,0 +1,72 @@
"""Pins the correctness gained by deleting the pre-parse
_quote_date_keyword_phrases rewrite.
That rewrite matched date-keyword phrases (e.g. "previous month" after a
date field) anywhere in the raw query string, including inside an
unrelated quoted string, and inserted quotes mid-phrase there too its
own docstring gave ``title:"see added:previous month notes"`` as the
example of what it corrupted. whoosh-compat's grammar accepts the same
phrase vocabulary unquoted natively (see TestUnquotedDateKeywordPhrases
in test_acceptance.py), so the rewrite was redundant everywhere it was
safe and actively wrong everywhere it was not. This is the one case that
tells the two apart: a literal title phrase that happens to contain
"added:previous month" as running text.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from documents.models import Document
if TYPE_CHECKING:
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
class TestQuotedStringContainingDateKeywordText:
"""A quoted title phrase containing the literal text
"added:previous month" as running words must match on that literal
text alone, never spill into an unfielded search for "previous" and
"month" across the default search fields the way the deleted rewrite
would have decomposed it into."""
def test_matches_only_the_literal_phrase(
self,
backend: TantivyBackend,
) -> None:
literal = _index(
backend,
title="see added:previous month notes",
content="quarterly filing",
checksum="dkp-literal",
archive_serial_number=920,
)
# Under the deleted rewrite, this decoy would incorrectly match:
# its title contains the "see added:" and " notes" fragments the
# corrupted parse required as title phrases, and its content
# supplies "previous" and "month" as the decomposed word-match
# clauses the rewrite turned the middle of the phrase into.
decoy = _index(
backend,
title="see added: quarterly report notes",
content="we reviewed the previous statement about month end",
checksum="dkp-decoy",
archive_serial_number=921,
)
query = 'title:"see added:previous month notes"'
assert _matched_ids(backend, query) == {literal.pk}
assert decoy.pk not in _matched_ids(backend, query)
@@ -0,0 +1,83 @@
"""Date keyword phrases (``today``, etc.) resolved in a non-UTC timezone,
end to end.
paperless's own ``tz=get_current_timezone()`` plumbing
(``TantivyBackend._parse_query``) is exercised elsewhere only for
relative *ranges* (``added:[-1 week to now]``, in
documents/tests/test_api_search.py). This covers a date *keyword*
(``today``), whose day boundary depends on the active timezone the same
way but goes through whoosh-compat's DateParserPlugin resolution instead
of an explicit range.
Discriminating shape: frozen at 2026-06-15T02:00 UTC, which is
2026-06-14T22:00 in America/New_York -- still "today" (06-14) there, but
already "today" (06-15) in UTC. Two documents pin both directions of the
mistake a hardcoded-UTC bug would make:
- ``in_ny_today`` (added 2026-06-14T20:00 UTC = 2026-06-14T16:00 NY) is
inside New York's "today" window and outside a naive UTC-calendar-day
window. A ``tz``-ignoring bug would miss it.
- ``in_utc_calendar_day_only`` (added 2026-06-15T10:00 UTC =
2026-06-15T06:00 NY) is inside a naive UTC-calendar-day window but
outside New York's actual "today" window. A ``tz``-ignoring bug would
wrongly match it.
"""
from __future__ import annotations
from datetime import UTC
from datetime import datetime
from typing import TYPE_CHECKING
import pytest
import time_machine
from documents.models import Document
if TYPE_CHECKING:
from pytest_django.fixtures import SettingsWrapper
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
FROZEN_NOW = datetime(2026, 6, 15, 2, 0, tzinfo=UTC)
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
class TestDateKeywordUsesTheActiveTimezone:
def test_today_matches_the_new_york_calendar_day_not_the_utc_one(
self,
backend: TantivyBackend,
settings: SettingsWrapper,
) -> None:
settings.TIME_ZONE = "America/New_York"
with time_machine.travel(FROZEN_NOW, tick=False):
in_ny_today = _index(
backend,
title="NY today",
content="x",
checksum="tz-keyword-ny-today",
added=datetime(2026, 6, 14, 20, 0, tzinfo=UTC),
)
# Not captured: the exact-set assertion below already proves
# this document (inside a naive UTC-calendar-day window, but
# outside New York's actual "today") does not match.
_index(
backend,
title="UTC calendar day only",
content="x",
checksum="tz-keyword-utc-calendar-day-only",
added=datetime(2026, 6, 15, 10, 0, tzinfo=UTC),
)
assert _matched_ids(backend, "added:today") == {in_ny_today.pk}
@@ -0,0 +1,20 @@
"""``_DEFAULT_SEARCH_FIELDS`` must stay a subset of the registered public
field names.
Nothing enforced this before: a rename in PUBLIC_FIELDS not mirrored in
``_DEFAULT_SEARCH_FIELDS`` (documents/search/_query.py) would 400 every
unfielded search at request time, since ``index.parse_query`` and the
fuzzy/CJK clause builders are handed a field name the schema no longer
has.
"""
from __future__ import annotations
from documents.search._fields import PUBLIC_FIELDS
from documents.search._query import _DEFAULT_SEARCH_FIELDS
class TestDefaultSearchFieldsAreRegistered:
def test_every_default_search_field_is_a_public_field(self) -> None:
public_field_names = {f.name for f in PUBLIC_FIELDS}
assert set(_DEFAULT_SEARCH_FIELDS) <= public_field_names
@@ -0,0 +1,356 @@
"""Pins the search syntax that ``docs/usage.md`` promises users.
Every query here appears verbatim, or as a direct paraphrase, in the
"Document searches" section of ``docs/usage.md``. Each case indexes real
documents and asserts on matched document IDs rather than on the parsed
query, because a query that parses cleanly is not necessarily a query that
means what the documentation says it means: ``added:now`` parses without a
single diagnostic and then matches nothing, because it resolves to an
instant rather than to a span.
The negative cases matter as much as the positive ones. They pin the
behaviours the docs explicitly warn about, so that if any of them ever
starts working the warning can be removed deliberately rather than being
left standing as a lie.
"""
from __future__ import annotations
from datetime import UTC
from datetime import datetime
from typing import TYPE_CHECKING
import pytest
import time_machine
from documents.models import Document
from documents.models import Note
from documents.models import Tag
from documents.search._errors import InvalidDateQuery
if TYPE_CHECKING:
from collections.abc import Generator
from django.contrib.auth.models import User
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
# A Monday, so that "next monday"/"last monday" land a clean week either side.
FROZEN_NOW = datetime(2026, 6, 15, 12, 0, tzinfo=UTC)
# The checksum used in the docs' `checksum:` example.
DOC_CHECKSUM = "9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08"
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
class TestLogicalExpressions:
@pytest.fixture
def docs(self, backend: TantivyBackend) -> dict[str, int]:
return {
"secret": _index(
backend,
title="Invoice one",
content="invoice secret contents",
checksum="doc-syntax-secret",
).pk,
"plain": _index(
backend,
title="Invoice two",
content="invoice ordinary contents",
checksum="doc-syntax-plain",
).pk,
}
def test_not_excludes_a_term(
self,
backend: TantivyBackend,
docs: dict[str, int],
) -> None:
assert _matched_ids(backend, "invoice NOT secret") == {docs["plain"]}
def test_leading_hyphen_requires_the_term_instead_of_excluding_it(
self,
backend: TantivyBackend,
docs: dict[str, int],
) -> None:
# The docs warn about exactly this: separators are stripped at index
# time, so "-secret" is the term "secret" and the query is an AND.
assert _matched_ids(backend, "invoice -secret") == {docs["secret"]}
def test_or_inside_parentheses_matches_either_branch(
self,
backend: TantivyBackend,
docs: dict[str, int],
) -> None:
matched = _matched_ids(backend, "invoice AND (secret OR ordinary)")
assert matched == {docs["secret"], docs["plain"]}
class TestPhraseSearch:
def test_quoted_phrase_requires_the_words_in_order(
self,
backend: TantivyBackend,
) -> None:
doc = _index(
backend,
title="Phrase",
content="the quick brown fox jumps",
checksum="doc-syntax-phrase",
)
assert _matched_ids(backend, '"quick brown fox"') == {doc.pk}
assert _matched_ids(backend, '"brown quick fox"') == set()
class TestTagCommaList:
"""``tag:bills,unpaid`` is published syntax (docs/usage.md), so this checks
that the documented spelling still returns what the docs promise: only the
document carrying every listed tag.
It is deliberately not proof of paperless's field configuration, and must
not be read as such. Removing ``comma_values`` from the ``tag`` FieldSpec
leaves this test passing, because paperless's analyzer splits the literal
value "bills,unpaid" into the same two tokens the value-list reading
produces, so the two readings select the same documents. The registry fact
-- that ``tag`` opts in and no other field does -- is observable only at
the registry, and is owned by test_registry.py's
``test_tag_is_comma_values``/``test_correspondent_is_not_comma_values``.
"""
def test_comma_list_requires_every_listed_tag(
self,
backend: TantivyBackend,
) -> None:
bills = Tag.objects.create(name="bills")
unpaid = Tag.objects.create(name="unpaid")
archived = Tag.objects.create(name="archived")
both = Document.objects.create(
title="Both tags",
content="body",
checksum="doc-syntax-tag-both",
)
both.tags.add(bills, unpaid)
backend.add_or_update(both)
one = Document.objects.create(
title="One tag",
content="body",
checksum="doc-syntax-tag-one",
)
one.tags.add(bills, archived)
backend.add_or_update(one)
assert _matched_ids(backend, "tag:bills,unpaid") == {both.pk}
assert _matched_ids(backend, "tag:bills") == {both.pk, one.pk}
class TestArchiveMetadataFields:
@pytest.fixture
def doc(self, backend: TantivyBackend, admin_user: User) -> Document:
doc = Document.objects.create(
title="Metadata",
content="body",
checksum=DOC_CHECKSUM,
archive_serial_number=100,
page_count=12,
original_filename="invoice.pdf",
)
Note.objects.create(document=doc, user=admin_user, note="a note")
backend.add_or_update(doc)
return doc
@pytest.mark.parametrize(
"query",
[
"asn:100",
"asn:[50 to 150]",
"page_count:12",
"page_count:[10 to 20]",
"num_notes:1",
"num_notes:[1 to 5]",
"original_filename:invoice.pdf",
f"checksum:{DOC_CHECKSUM}",
"checksum:9f86d081*",
# A checksum term is stored verbatim, but a checksum *pattern* is
# lowercased before it is matched, which the docs now say outright
# next to the "only a complete, lowercase checksum matches" rule
# that the uppercase term in the negative list below pins.
"checksum:9F86D081*",
],
)
def test_documented_metadata_query_matches(
self,
backend: TantivyBackend,
doc: Document,
query: str,
) -> None:
assert _matched_ids(backend, query) == {doc.pk}
@pytest.mark.parametrize(
"query",
[
# The docs say only a complete, lowercase checksum matches.
"checksum:9f86d081",
f"checksum:{DOC_CHECKSUM.upper()}",
],
)
def test_partial_or_uppercase_checksum_matches_nothing(
self,
backend: TantivyBackend,
doc: Document,
query: str,
) -> None:
assert _matched_ids(backend, query) == set()
class TestDocumentedDateForms:
@pytest.fixture(autouse=True)
def frozen_now(self) -> Generator[None, None, None]:
with time_machine.travel(FROZEN_NOW, tick=False):
yield
@pytest.fixture
def dated(self, backend: TantivyBackend) -> dict[str, int]:
stamps = {
"today": datetime(2026, 6, 15, 9, 0, tzinfo=UTC),
"yesterday": datetime(2026, 6, 14, 9, 0, tzinfo=UTC),
"tomorrow": datetime(2026, 6, 16, 9, 0, tzinfo=UTC),
"next_monday": datetime(2026, 6, 22, 10, 0, tzinfo=UTC),
"last_monday": datetime(2026, 6, 8, 10, 0, tzinfo=UTC),
"january": datetime(2026, 1, 10, 10, 0, tzinfo=UTC),
"old": datetime(2005, 3, 4, 15, 30, tzinfo=UTC),
}
return {
label: _index(
backend,
title=label,
content="dated body",
checksum=f"doc-syntax-date-{label}",
added=stamp,
).pk
for label, stamp in stamps.items()
}
@pytest.mark.parametrize(
("query", "label"),
[
("added:today", "today"),
("added:yesterday", "yesterday"),
("added:tomorrow", "tomorrow"),
('added:"next monday"', "next_monday"),
('added:"last monday"', "last_monday"),
("added:january", "january"),
("added:2005-03-04", "old"),
("added:2005-03", "old"),
("added:[2005-01-01 to 2005-12-31]", "old"),
("added:[2005 to 2009]", "old"),
# A full timestamp works, but only quoted when it stands alone,
# and only unquoted when it is a range bound. The bare standalone
# spelling is pinned as a non-match below.
('added:"2005-03-04T15:30:00Z"', "old"),
("added:[2005-03-04T09:00:00Z to 2005-03-04T17:00:00Z]", "old"),
# A quoted range bound works when the quotes are single ones; the
# double-quoted spelling is pinned as an error below.
("added:['2005-03-04' to 2005-03-05]", "old"),
],
)
def test_documented_date_form_matches_its_day_or_month(
self,
backend: TantivyBackend,
dated: dict[str, int],
query: str,
label: str,
) -> None:
assert _matched_ids(backend, query) == {dated[label]}
@pytest.mark.parametrize(
"query",
[
# Zero-width: these resolve to a single instant, not a span, so
# nothing in a realistic corpus lands on them. The docs warn
# about them rather than presenting them as usable.
"added:now",
"added:noon",
"added:midnight",
# Quoting is what rescues the other multi-word date expressions,
# so pin that it does not rescue these: the problem is the width
# of the resulting range, not the way the value is delimited.
# One quoted spelling is enough for that; which keyword sits
# inside the quotes is grammar whoosh-compat owns.
'added:"now"',
# A relative offset, which the warning in the docs names by this
# exact spelling. Standing alone it is an instant like the rest of
# this list; the same offset used as a range bound is a real
# window, pinned by the test below.
'added:"-1 week"',
],
)
def test_forms_the_docs_warn_about_match_nothing(
self,
backend: TantivyBackend,
dated: dict[str, int],
query: str,
) -> None:
assert _matched_ids(backend, query) == set()
def test_bare_timestamp_is_rejected_rather_than_matching_nothing(
self,
backend: TantivyBackend,
dated: dict[str, int],
) -> None:
"""The bare, unquoted spelling of a full timestamp. The quoted and
range-bound spellings pinned above do work and match this fixture's
document; this one is a user-fixable error rather than an empty
result set, so the docs tell the user to quote it.
The reported value is the prefix the date grammar could consume, not
the whole of what the user typed: paperless's own message, not this
value, is what has to carry the "quote it" guidance.
"""
with pytest.raises(InvalidDateQuery) as exc_info:
_matched_ids(backend, "added:2005-03-04T15:30:00Z")
assert exc_info.value.field == "added"
assert exc_info.value.value == "2005-03-"
def test_relative_offset_as_a_range_bound_is_a_real_window(
self,
backend: TantivyBackend,
dated: dict[str, int],
) -> None:
"""The same offset that matches nothing on its own spans the last
seven days as a lower bound. The docs say so, next to the warning
about the standalone form, so both readings are pinned together.
"last_monday" is indexed at 2026-06-08T10:00, two hours before the
window opens, so its exclusion is what shows the bound is the offset
and not a whole-day rounding of it.
"""
assert _matched_ids(backend, "added:['-1 week' to now]") == {
dated["today"],
dated["yesterday"],
}
def test_double_quoted_range_bound_is_rejected(
self,
backend: TantivyBackend,
dated: dict[str, int],
) -> None:
"""Quoting a range bound is allowed, but only with single quotes: the
double-quoted spelling reaches the date grammar with its quotes still
attached and is not a recognizable date. The docs say so, so pin which
of the two quote characters is the one that fails.
"""
with pytest.raises(InvalidDateQuery) as exc_info:
_matched_ids(backend, 'added:["2005-03-04" to 2005-03-05]')
assert exc_info.value.value == '"2005-03-04"'
@@ -0,0 +1,221 @@
"""Diagnostics route by Cause, and user-facing messages are host-owned.
whoosh-compat documents ``Diagnostic.message`` as developer output with no
stability guarantee, so it must never reach an HTTP response body.
"""
from __future__ import annotations
import logging
from datetime import UTC
import pytest
import tantivy
from whoosh_compat.errors import Diagnostic
from whoosh_compat.errors import DiagnosticKind
from whoosh_compat.errors import QueryError
from whoosh_compat.errors import cause_for
from whoosh_compat.fields import FieldKind
from whoosh_compat.fields import FieldRef
from documents.search._errors import SearchQueryError
from documents.search._query import _map_emit_error
from documents.search._query import _single_diagnostic_to_error
from documents.search._query import parse_user_query
from documents.search._schema import build_schema
from documents.search._tokenizer import register_tokenizers
pytestmark = pytest.mark.search
_LIBRARY_PROSE = "INTERNAL LIBRARY WORDING WITH raw tantivy detail"
@pytest.fixture(scope="module")
def query_index() -> tantivy.Index:
"""An in-memory, unstemmed index; these tests only parse, never index."""
idx = tantivy.Index(build_schema(), path=None)
register_tokenizers(idx, "")
return idx
def _diagnostic(
kind: DiagnosticKind,
*,
field: FieldRef | None = FieldRef("title"),
field_kind: FieldKind | None = FieldKind.TEXT,
) -> Diagnostic:
"""A Diagnostic shaped like the emitter's, with the library's own
kind -> cause mapping rather than a hand-picked cause."""
return Diagnostic(
kind=kind,
cause=cause_for(kind),
message=_LIBRARY_PROSE,
field=field,
field_kind=field_kind,
)
class TestEmitErrorRouting:
"""Every Cause gets a distinguishable treatment, not just "a 400"."""
@pytest.mark.parametrize(
"kind",
[
DiagnosticKind.BACKEND_REJECTED,
DiagnosticKind.AST_INVALID_SHAPE,
DiagnosticKind.AST_UNKNOWN_FIELD,
],
)
def test_internal_cause_is_not_converted(self, kind: DiagnosticKind) -> None:
"""A library defect must surface as a 500 monitoring can see, not a
400 blaming the user."""
error = QueryError(_diagnostic(kind))
with pytest.raises(QueryError) as excinfo:
_map_emit_error(error)
assert excinfo.value is error
def test_misconfigured_cause_is_logged_and_becomes_a_400(
self,
caplog: pytest.LogCaptureFixture,
) -> None:
kind = DiagnosticKind.SCHEMA_FIELD_MISSING
with caplog.at_level(logging.ERROR, logger="paperless.search"):
error = _map_emit_error(
QueryError(_diagnostic(kind, field=FieldRef("asn"))),
)
assert isinstance(error, SearchQueryError)
errors = [r for r in caplog.records if r.levelno == logging.ERROR]
assert len(errors) == 1
assert "asn" in errors[0].getMessage()
assert kind.name in errors[0].getMessage()
@pytest.mark.parametrize(
"kind",
[
DiagnosticKind.TEXT_RANGE,
DiagnosticKind.PATTERN_TOO_COMPLEX,
DiagnosticKind.EXISTS_REQUIRES_FAST,
],
)
def test_unsupported_cause_is_a_400_with_no_operator_log(
self,
kind: DiagnosticKind,
caplog: pytest.LogCaptureFixture,
) -> None:
"""A query tantivy cannot run is the user's to fix; it must not page
an operator the way a registry/schema mismatch does.
EXISTS_REQUIRES_FAST is nominally MISCONFIGURED but belongs here: it
is decided from the registry's own FieldSpec, so it never reports a
disagreement anyone could resolve."""
with caplog.at_level(logging.WARNING, logger="paperless.search"):
error = _map_emit_error(QueryError(_diagnostic(kind)))
assert isinstance(error, SearchQueryError)
assert caplog.records == []
@pytest.mark.parametrize(
"kind",
[
DiagnosticKind.TEXT_RANGE,
DiagnosticKind.PATTERN_TOO_COMPLEX,
DiagnosticKind.EXISTS_REQUIRES_FAST,
DiagnosticKind.SCHEMA_FIELD_MISSING,
],
)
def test_user_facing_message_never_echoes_library_prose(
self,
kind: DiagnosticKind,
) -> None:
error = _map_emit_error(QueryError(_diagnostic(kind)))
assert _LIBRARY_PROSE not in str(error)
@pytest.mark.parametrize(
"kind",
[
DiagnosticKind.TEXT_RANGE,
DiagnosticKind.PATTERN_TOO_COMPLEX,
DiagnosticKind.EXISTS_REQUIRES_FAST,
DiagnosticKind.SCHEMA_FIELD_MISSING,
],
)
def test_user_facing_message_names_the_field(
self,
kind: DiagnosticKind,
) -> None:
"""FieldRef.__str__ yields the canonical dotted name, including a
JSON subpath, so every user-reachable emit kind can name it."""
diagnostic = _diagnostic(
kind,
field=FieldRef("custom_fields", "value"),
field_kind=FieldKind.JSON,
)
error = _map_emit_error(QueryError(diagnostic))
assert "custom_fields.value" in str(error)
class TestParseDiagnosticMessages:
"""Parse-time diagnostics are host-worded too, off field_kind."""
def test_too_deep_is_a_400_without_library_prose(self) -> None:
error = _single_diagnostic_to_error(
_diagnostic(DiagnosticKind.TOO_DEEP, field=None, field_kind=None),
)
assert isinstance(error, SearchQueryError)
assert _LIBRARY_PROSE not in str(error)
@pytest.mark.parametrize(
("kind", "field_kind"),
[
(DiagnosticKind.PATTERN_ON_NUMERIC, FieldKind.U64),
(DiagnosticKind.PATTERN_ON_BOOLEAN_EXISTS, FieldKind.BOOLEAN_EXISTS),
(DiagnosticKind.PATTERN_ON_SUBPATH, FieldKind.JSON),
],
)
def test_pattern_on_kinds_name_the_field_and_its_kind(
self,
kind: DiagnosticKind,
field_kind: FieldKind,
) -> None:
error = _single_diagnostic_to_error(
_diagnostic(kind, field=FieldRef("asn"), field_kind=field_kind),
)
message = str(error)
assert _LIBRARY_PROSE not in message
assert "asn" in message
assert field_kind.name.lower() in message
class TestRealQueriesRouteCorrectly:
"""The routing table against diagnostics emit() really produces."""
def test_text_range_is_a_400_naming_the_field(
self,
query_index: tantivy.Index,
) -> None:
with pytest.raises(SearchQueryError) as excinfo:
parse_user_query(query_index, "title:[a to b]", UTC)
assert "title" in str(excinfo.value)
def test_wildcard_on_a_numeric_field_is_a_400_naming_the_field(
self,
query_index: tantivy.Index,
) -> None:
with pytest.raises(SearchQueryError) as excinfo:
parse_user_query(query_index, "asn:12*", UTC)
assert "asn" in str(excinfo.value)
def test_internal_diagnostic_escapes_as_a_query_error(
self,
query_index: tantivy.Index,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The one case with no query text that reaches it: emit() reporting
a defect in itself must not be converted to a user-facing 400."""
import documents.search._query as query_mod
def raise_internal(*args: object, **kwargs: object) -> None:
raise QueryError(_diagnostic(DiagnosticKind.BACKEND_REJECTED))
monkeypatch.setattr(query_mod, "tantivy_emit", raise_internal)
with pytest.raises(QueryError):
parse_user_query(query_index, "invoice", UTC)
@@ -0,0 +1,92 @@
"""``field:*`` on a JSON field is user error, not an operator alert.
whoosh-compat classifies EXISTS_REQUIRES_FAST as MISCONFIGURED, and
_map_emit_error used to route every MISCONFIGURED diagnostic to an ERROR log.
But the kind is decided from the registry's own FieldSpec (kind plus fast)
without consulting the index schema, and field_descriptors() builds the JSON
fields non-fast deliberately, so nothing is misconfigured and no operator
action can clear the condition. Any authenticated user could otherwise emit
ERROR lines in a loop by repeating ``notes:*``.
SCHEMA_FIELD_MISSING, the other MISCONFIGURED kind, does compare the registry
against the live schema, so it stays an ERROR.
"""
from __future__ import annotations
import logging
from datetime import UTC
import pytest
import tantivy
from whoosh_compat.errors import Diagnostic
from whoosh_compat.errors import DiagnosticKind
from whoosh_compat.errors import QueryError
from whoosh_compat.errors import cause_for
from whoosh_compat.fields import FieldKind
from whoosh_compat.fields import FieldRef
from documents.search._errors import SearchQueryError
from documents.search._query import _map_emit_error
from documents.search._query import parse_user_query
from documents.search._schema import build_schema
from documents.search._tokenizer import register_tokenizers
pytestmark = pytest.mark.search
# Every spelling of "does this JSON field have a value" a user can type.
EXISTS_QUERIES = [
"notes:*",
"notes.note:*",
"notes.user:*",
"custom_fields:*",
"custom_fields.name:*",
"custom_fields.value:*",
]
@pytest.fixture(scope="module")
def query_index() -> tantivy.Index:
idx = tantivy.Index(build_schema(), path=None)
register_tokenizers(idx, "")
return idx
class TestJsonExistsIsUserError:
@pytest.mark.parametrize("query", EXISTS_QUERIES)
def test_query_is_a_400_that_emits_no_error_log(
self,
query_index: tantivy.Index,
caplog: pytest.LogCaptureFixture,
query: str,
) -> None:
with caplog.at_level(logging.WARNING, logger="paperless.search"):
with pytest.raises(SearchQueryError) as excinfo:
parse_user_query(query_index, query, UTC)
assert query.split(":", maxsplit=1)[0] in str(excinfo.value)
assert [r for r in caplog.records if r.levelno >= logging.ERROR] == []
class TestGenuineMisconfigurationStillLogs:
def test_schema_field_missing_is_an_error_log(
self,
caplog: pytest.LogCaptureFixture,
) -> None:
"""The registry naming a field the index schema does not have is a
real mismatch an operator can fix, so it keeps the alert."""
kind = DiagnosticKind.SCHEMA_FIELD_MISSING
error = QueryError(
Diagnostic(
kind=kind,
cause=cause_for(kind),
message="field 'asn' is not defined in the index schema",
field=FieldRef("asn"),
field_kind=FieldKind.U64,
),
)
with caplog.at_level(logging.ERROR, logger="paperless.search"):
mapped = _map_emit_error(error)
assert isinstance(mapped, SearchQueryError)
records = [r for r in caplog.records if r.levelno == logging.ERROR]
assert len(records) == 1
assert kind.name in records[0].getMessage()
+10
View File
@@ -0,0 +1,10 @@
from whoosh_compat import FieldKind
from documents.search._fields import PUBLIC_FIELDS
class TestPublicFields:
def test_json_fields_have_subpaths(self) -> None:
for field in PUBLIC_FIELDS:
if field.kind is FieldKind.JSON:
assert field.subpaths, f"{field.name} is JSON but has no subpaths"
@@ -0,0 +1,174 @@
"""The words the fuzzy blend clause hands back to tantivy's parser.
The clause re-parses a word string through tantivy, which analyzes it
again, so the words must be the query's raw text rather than the analyzed
text (analysis is not idempotent), and must still be split into plain
words so that hyphenated, dotted and quoted terms keep contributing.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from documents.models import Document
if TYPE_CHECKING:
from pytest_django.fixtures import SettingsWrapper
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
@pytest.fixture(autouse=True)
def fuzzy_enabled(settings: SettingsWrapper) -> None:
"""Enable the fuzzy blend clause. The threshold doubles as a minimum
score filter, so it is set to 0.0: every hit passes and the test sees
the clause's matching behaviour, not the filter's."""
settings.ADVANCED_FUZZY_SEARCH_THRESHOLD = 0.0
class TestFuzzyClauseWords:
def test_a_stemmed_word_is_not_stemmed_a_second_time(
self,
backend: TantivyBackend,
) -> None:
"""'universities' stems to 'univers'; feeding that back to tantivy
stems it again to 'univ', whose fuzzy prefix reaches unrelated
words. The clause must stay wide enough for a typo and no wider."""
wanted = _index(
backend,
title="A",
content="universities of europe",
checksum="fuzz-stem-1",
)
typo = _index(
backend,
title="B",
content="universties of europe",
checksum="fuzz-stem-2",
)
_index(
backend,
title="C",
content="univalent chemical bonds",
checksum="fuzz-stem-3",
)
_index(
backend,
title="D",
content="unicycle repair manual",
checksum="fuzz-stem-4",
)
assert _matched_ids(backend, "universities") == {wanted.pk, typo.pk}
def test_a_hyphenated_term_still_reaches_the_clause(
self,
backend: TantivyBackend,
) -> None:
"""'COVID-19' is one raw token: unless it is split into words, it
carries characters the re-parse would read as grammar, is dropped,
and the whole query loses its fuzzy clause."""
misspelled = _index(
backend,
title="A",
content="covidx testing results",
checksum="fuzz-hyphen-1",
)
assert _matched_ids(backend, "COVID-19") == {misspelled.pk}
def test_a_phrase_still_reaches_the_clause(
self,
backend: TantivyBackend,
) -> None:
"""A phrase is one raw token carrying a space, and is the whole
query's only free text here."""
near_miss = _index(
backend,
title="A",
content="taxation reportage weekly",
checksum="fuzz-phrase-1",
)
assert _matched_ids(backend, '"tax reports"') == {near_miss.pk}
class TestBooleanKeywordsInRawText:
"""Tantivy's boolean keywords are word runs, so they survive the cut
into words and its own parser reads them as grammar. Raw query text
reaches that parser with its case intact, so a quoted phrase can carry
them in."""
@pytest.fixture
def corpus(self, backend: TantivyBackend) -> dict[str, int]:
both = _index(
backend,
title="A",
content="taxation reportage weekly",
checksum="fuzz-kw-1",
)
tax_only = _index(
backend,
title="B",
content="taxation only here",
checksum="fuzz-kw-2",
)
report_only = _index(
backend,
title="C",
content="reportage only here",
checksum="fuzz-kw-3",
)
return {
"both": both.pk,
"tax_only": tax_only.pk,
"report_only": report_only.pk,
}
@pytest.mark.parametrize(
"query",
[
pytest.param('"tax AND reports"', id="and"),
pytest.param('"tax OR reports"', id="or"),
pytest.param('"tax NOT reports"', id="not"),
pytest.param('"tax IN reports"', id="in"),
],
)
def test_a_keyword_inside_a_phrase_stays_an_ordinary_word(
self,
backend: TantivyBackend,
corpus: dict[str, int],
query: str,
) -> None:
"""The phrase asks for three words, so the clause must stay the
disjunction it is for '"tax reports"': AND must not turn it into a
conjunction, NOT must not give it its own exclusion, IN must not
fail the parse."""
assert _matched_ids(backend, '"tax reports"') == set(corpus.values())
assert _matched_ids(backend, query) == set(corpus.values())
def test_a_trailing_keyword_does_not_drop_the_clause(
self,
backend: TantivyBackend,
corpus: dict[str, int],
) -> None:
"""'tax AND' is a syntax error to tantivy's parser, which would
cost the whole query its fuzzy clause."""
assert _matched_ids(backend, '"tax AND"') == {
corpus["both"],
corpus["tax_only"],
}
@@ -0,0 +1,192 @@
"""Regression coverage for the unguarded TEXT-mode highlight query.
parse_simple_text_highlight_query re-parses simple-search tokens through
Tantivy's query-string parser to build a SnippetGenerator-compatible query.
Simple-search tokens keep arbitrary punctuation (quotes, colons, brackets,
slashes), so any token carrying Tantivy query grammar raised an unguarded
ValueError. The search itself had already succeeded by the time this ran:
only the highlight step failed, and with the DocumentViewSet.list
exception handler narrowed elsewhere on this branch, that ValueError now
reaches the client as a bare 500 rather than a 400.
Covers three angles:
- the query builder itself: quoting each token as its own escaped phrase
should let it parse instead of raising, for every failure mode a plain-
text query can trigger (syntax error, unknown field, unsupported regex).
- highlight_hits: even when a token still can't be expressed as a
highlight query, the guard must fall back to a query that still
produces usable highlight HTML, not silently empty ones.
- the real API endpoint: pinning the previously-500 status to 200.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
import tantivy
from rest_framework import status
from documents.search._backend import SearchMode
from documents.search._query import parse_simple_text_highlight_query
from documents.search._schema import build_schema
from documents.search._tokenizer import register_tokenizers
from documents.tests.factories import DocumentFactory
if TYPE_CHECKING:
from rest_framework.test import APIClient
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
# Each spelling below trips a different Tantivy parser failure mode:
# 'a"b' -> Syntax Error (unterminated quote)
# foo:bar -> unknown field
# (a -> Syntax Error (unbalanced group)
# [a -> Syntax Error (unbalanced range)
# /a/ -> Unsupported query (regex queries disallowed)
_MALFORMED_QUERIES = [
pytest.param('a"b', id="unterminated_quote"),
pytest.param("foo:bar", id="unknown_field"),
pytest.param("(a", id="unbalanced_group"),
pytest.param("[a", id="unbalanced_range"),
pytest.param("/a/", id="unsupported_regex"),
]
@pytest.fixture(scope="module")
def query_index() -> tantivy.Index:
"""An in-memory, unstemmed index for parse-only tests."""
schema = build_schema()
idx = tantivy.Index(schema, path=None)
register_tokenizers(idx, "")
return idx
class TestParseSimpleTextHighlightQueryDoesNotRaise:
"""The query builder itself must tolerate Tantivy syntax in its tokens."""
@pytest.mark.parametrize("raw_query", _MALFORMED_QUERIES)
def test_malformed_token_does_not_raise(
self,
query_index: tantivy.Index,
raw_query: str,
) -> None:
assert isinstance(
parse_simple_text_highlight_query(query_index, raw_query),
tantivy.Query,
)
class TestHighlightHitsProducesUsableHighlights:
"""highlight_hits must keep producing real <b>-wrapped snippet HTML for
these queries, not merely avoid raising."""
@pytest.mark.parametrize(
"raw_query",
[*_MALFORMED_QUERIES, pytest.param("plain text", id="plain_text_sanity")],
)
def test_highlight_still_contains_matched_text(
self,
backend: TantivyBackend,
raw_query: str,
) -> None:
doc = DocumentFactory.create(
title="probe",
content=f"needle content containing {raw_query} literally here",
)
backend.add_or_update(doc)
hits = backend.highlight_hits(
raw_query,
[doc.pk],
search_mode=SearchMode.TEXT,
)
assert len(hits) == 1
highlights = hits[0]["highlights"]
assert "content" in highlights, (
f"Expected a content highlight for {raw_query!r}, got: {highlights!r}"
)
assert "<b>" in highlights["content"], (
f"Highlight for {raw_query!r} carries no matched-term markup: "
f"{highlights['content']!r}"
)
class TestHighlightGuardDiscriminatesOnValueError:
"""The guard added to highlight_hits must catch exactly ValueError, the
same shape as the sibling notes_text guard, and let anything else
through -- so a real library defect is never mistaken for a harmless
syntax error."""
def test_non_value_error_is_not_swallowed(
self,
backend: TantivyBackend,
monkeypatch: pytest.MonkeyPatch,
) -> None:
import documents.search._backend as backend_mod
def raise_runtime_error(*args: object, **kwargs: object) -> object:
raise RuntimeError("synthetic bug, unrelated to query syntax")
monkeypatch.setattr(
backend_mod,
"parse_simple_text_highlight_query",
raise_runtime_error,
)
doc = DocumentFactory.create(title="probe", content="anything here")
backend.add_or_update(doc)
with pytest.raises(RuntimeError):
backend.highlight_hits(
"anything",
[doc.pk],
search_mode=SearchMode.TEXT,
)
@pytest.mark.usefixtures("_search_index")
class TestApiNoLongerReturns500:
"""Pins the actual regression: a matching TEXT-mode search whose query
string carries Tantivy syntax must return results, not a server error."""
@pytest.mark.parametrize("raw_query", _MALFORMED_QUERIES)
def test_malformed_text_query_returns_200(
self,
admin_client: APIClient,
raw_query: str,
) -> None:
from documents.search import get_backend
doc = DocumentFactory.create(
title="probe",
content=f"needle content containing {raw_query} literally here",
)
get_backend().add_or_update(doc)
response = admin_client.get(f"/api/documents/?text={raw_query}")
assert response.status_code == status.HTTP_200_OK
assert response.data["count"] == 1
def test_plain_text_query_still_returns_200(
self,
admin_client: APIClient,
) -> None:
"""Sanity check: the guard must not mask a total failure of the
ordinary highlight path."""
from documents.search import get_backend
doc = DocumentFactory.create(
title="probe",
content="needle content containing plain text literally here",
)
get_backend().add_or_update(doc)
response = admin_client.get("/api/documents/?text=plain text")
assert response.status_code == status.HTTP_200_OK
assert response.data["count"] == 1
@@ -0,0 +1,148 @@
"""Bare notes:/custom_fields: prefix resolution.
"notes:foo"/"custom_fields:foo" were valid fielded searches before the
whoosh-compat migration. The registry only exposes them as JSON subpaths, so
each JSON FieldSpec declares a default subpath (SubpathSpec(default=True)):
notes: resolves to notes.note:, custom_fields: resolves to
custom_fields.value:. This replaced an earlier regex-based rewrite
(_rewrite_bare_json_field_prefixes) that ran on the raw query string before
parsing and was blind to quoting, so a phrase like
content:"payment notes: none" was silently corrupted into a notes-field
search and matched nothing. Resolving the default subpath inside the parser
instead means quoting is already understood by the time it happens.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from django.contrib.auth.models import User
from documents.models import CustomField
from documents.models import CustomFieldInstance
from documents.models import Document
from documents.models import Note
if TYPE_CHECKING:
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
class TestBareJsonFieldPrefixes:
def test_bare_notes_prefix_searches_note_text(
self,
backend: TantivyBackend,
) -> None:
alice = User.objects.create_user(username="alice")
with_note = Document.objects.create(
title="Has note",
content="x",
checksum="bare-notes-with",
)
Note.objects.create(document=with_note, user=alice, note="crocodile")
backend.add_or_update(with_note)
# This document's CONTENT contains the words a demoted text search
# would match; it must NOT match once the prefix addresses notes.
_index(
backend,
title="Notes about things",
content="notes crocodile mention",
checksum="bare-notes-decoy",
)
assert _matched_ids(backend, "notes:crocodile") == {with_note.pk}
def test_bare_custom_fields_prefix_searches_values(
self,
backend: TantivyBackend,
) -> None:
field = CustomField.objects.create(
name="Policy Number",
data_type=CustomField.FieldDataType.STRING,
)
with_value = Document.objects.create(
title="Has field",
content="x",
checksum="bare-cf-with",
)
CustomFieldInstance.objects.create(
document=with_value,
field=field,
value_text="crocodile",
)
backend.add_or_update(with_value)
_index(
backend,
title="Custom things",
content="custom fields crocodile",
checksum="bare-cf-decoy",
)
assert _matched_ids(backend, "custom_fields:crocodile") == {with_value.pk}
def test_subpath_spellings_are_untouched(
self,
backend: TantivyBackend,
) -> None:
bob = User.objects.create_user(username="bob")
doc = Document.objects.create(
title="Bob note",
content="x",
checksum="bare-subpath",
)
Note.objects.create(document=doc, user=bob, note="remark")
backend.add_or_update(doc)
assert _matched_ids(backend, "notes.user:bob") == {doc.pk}
assert _matched_ids(backend, "notes.note:remark") == {doc.pk}
class TestQuotedPhraseContainingNotesColonIsNotCorrupted:
"""The regex rewrite this migration removes was blind to quoting: it
matched "notes:" anywhere in the raw query string, including inside an
already-quoted phrase on an unrelated field, silently turning
content:"payment notes: none" into a notes-field search that matched
nothing. Resolving the default subpath during parsing (which is
quote-aware) fixes this."""
def test_quoted_phrase_with_notes_colon_matches_by_content(
self,
backend: TantivyBackend,
) -> None:
target = _index(
backend,
title="Statement",
content="payment notes: none",
checksum="quoted-phrase-notes-colon",
)
assert _matched_ids(
backend,
'content:"payment notes: none"',
) == {target.pk}
def test_quoted_phrase_matches_the_same_document_unquoted(
self,
backend: TantivyBackend,
) -> None:
# Same document, phrasing without the colon: this proves the fix is
# about quote-awareness, not about the words themselves being
# unsearchable.
target = _index(
backend,
title="Statement",
content="payment notes none",
checksum="quoted-phrase-no-colon",
)
assert _matched_ids(
backend,
'content:"payment notes none"',
) == {target.pk}
@@ -0,0 +1,83 @@
"""Every declared JSON subpath must actually be written to the index.
PUBLIC_FIELDS declares each JSON field's subpaths (e.g. ``notes`` ->
{"user", "note"}), but nothing coupled that declaration to what
``_backend.py``'s document builder actually writes into the JSON blob at
index time. A subpath declared but never written would be
queryable-but-always-empty -- syntactically valid, silently matching
nothing -- with no test failure anywhere.
This indexes one real document carrying values for every JSON field
(a Note, a CustomFieldInstance) and inspects the document's own stored
JSON payload, rather than running field-specific queries: that way a
future JSON field's subpaths are covered automatically, without a new
per-subpath query having to be added by hand each time.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
import tantivy
from django.contrib.auth.models import User
from whoosh_compat import FieldKind
from documents.models import CustomField
from documents.models import CustomFieldInstance
from documents.models import Document
from documents.models import Note
from documents.search._fields import PUBLIC_FIELDS
if TYPE_CHECKING:
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
class TestJsonSubpathsAreWrittenAtIndexTime:
def test_every_declared_json_subpath_appears_in_the_stored_document(
self,
backend: TantivyBackend,
) -> None:
user = User.objects.create_user(username="completeness-user")
field = CustomField.objects.create(
name="Completeness Field",
data_type=CustomField.FieldDataType.STRING,
)
doc = Document.objects.create(
title="Completeness doc",
content="x",
checksum="json-subpath-completeness",
)
Note.objects.create(document=doc, user=user, note="a note")
CustomFieldInstance.objects.create(
document=doc,
field=field,
value_text="a value",
)
backend.add_or_update(doc)
index = backend._index
searcher = index.searcher()
hits = searcher.search(
tantivy.Query.term_query(index.schema, "id", doc.pk),
limit=1,
).hits
assert hits, "the document was not indexed"
stored = searcher.doc(hits[0][1]).to_dict()
json_fields = [f for f in PUBLIC_FIELDS if f.kind is FieldKind.JSON]
assert json_fields, "no JSON fields declared - fixture is stale"
for field_spec in json_fields:
stored_values = stored.get(field_spec.name)
assert stored_values, (
f"{field_spec.name} was not written to the index at all"
)
written_keys = stored_values[0].keys()
for subpath in field_spec.subpaths:
assert subpath in written_keys, (
f"{field_spec.name}.{subpath} is declared in PUBLIC_FIELDS "
"but _backend.py's document builder never writes it - it "
"would be queryable but always empty"
)
@@ -0,0 +1,92 @@
"""Wildcard patterns on KEYWORD fields must stay literal.
``checksum`` is the only KEYWORD field: it is indexed with the raw tokenizer,
so its terms are never lowercased, folded or stemmed. Running its wildcard
patterns through the stemming normalizer rewrote hex prefixes ("ceded" ->
"cede") and returned documents whose checksum did not start with what the user
typed, which for an identity field is a wrong answer.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from documents.models import Document
from documents.search._registry import get_field_registry
if TYPE_CHECKING:
from whoosh_compat import FieldRegistry
from whoosh_compat import PatternNormalizer
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
CEDEF00D = "cedef00ddeadbeef0123456789abcdef01234567"
CEDEDEAD = "cededeadbeef567801234567" + "89abcdef01234567"
def _normalizer(registry: FieldRegistry, name: str) -> PatternNormalizer:
ref = registry.make_ref(name)
assert ref is not None
resolved = registry.resolve(ref)
assert resolved is not None
assert resolved.spec.pattern_normalizer is not None
return resolved.spec.pattern_normalizer
class TestKeywordPatternNormalizer:
@pytest.mark.parametrize(
"run",
[
pytest.param("ceded", id="stems_to_cede"),
pytest.param("added", id="stems_to_ad"),
pytest.param("cafed", id="stems_to_cafe"),
],
)
def test_keyword_runs_are_folded_not_stemmed(self, run: str) -> None:
"""One form, the run as typed: a KEYWORD pattern must never be widened
to a stem, which would return checksums that do not start with what
the user typed."""
normalize = _normalizer(get_field_registry("en"), "checksum")
assert normalize(run) == run
def test_text_runs_still_offer_their_stem(self) -> None:
"""A TEXT field offers the stem alongside the typed run, so a term
matching either one is reachable."""
normalize = _normalizer(get_field_registry("en"), "title")
assert tuple(normalize("Running")) == ("running", "run")
class TestChecksumPrefixQueries:
@pytest.fixture
def indexed(self, backend: TantivyBackend) -> None:
for i, checksum in enumerate((CEDEF00D, CEDEDEAD)):
doc = Document.objects.create(
title=f"Checksum doc {i}",
content="invoices for the quarter",
checksum=checksum,
archive_serial_number=940 + i,
)
backend.add_or_update(doc)
def _ids(self, backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def test_prefix_matches_only_the_document_that_starts_with_it(
self,
backend: TantivyBackend,
indexed: None,
) -> None:
matched = self._ids(backend, "checksum:ceded*")
expected = Document.objects.get(checksum=CEDEDEAD).pk
assert matched == {expected}
def test_text_prefix_still_reaches_the_stemmed_index(
self,
backend: TantivyBackend,
indexed: None,
) -> None:
assert len(self._ids(backend, "invoice*")) == 2
@@ -0,0 +1,233 @@
"""Wildcard patterns must match a stemmed index.
Query patterns are normalized but were not stemmed, while index terms are
stemmed, so the natural spelling of a prefix search matched nothing:
``invoice*`` found no document although ``invoic*`` did. v2's index was
UNSTEMMED (whoosh ``TEXT()`` defaults to ``StandardAnalyzer``), so this
regressed against both baselines.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from documents.models import Document
from documents.search._registry import _make_pattern_normalizer
from documents.search._tokenizer import ascii_fold
from documents.search._tokenizer import paperless_text_analyzer
from documents.search._tokenizer import stem_pattern_text
if TYPE_CHECKING:
from whoosh_compat import PatternNormalizer
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
CONTENT = (
"invoice total due for electricity from both companies, "
"payments made to the university library, copies attached"
)
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
@pytest.fixture
def indexed_doc(backend: TantivyBackend) -> Document:
doc = Document.objects.create(
title="Invoice 2020 productname",
content=CONTENT,
checksum="pattern-stemming-1",
archive_serial_number=900,
)
backend.add_or_update(doc)
return doc
class TestPrefixStemming:
@pytest.mark.parametrize(
"query",
[
"invoice*",
"electricity*",
"companies*",
"payments*",
"library*",
"title:Invoice*",
],
)
def test_full_word_prefix_matches_its_stem(
self,
backend: TantivyBackend,
indexed_doc: Document,
query: str,
) -> None:
assert _matched_ids(backend, query) == {indexed_doc.id}
@pytest.mark.parametrize("query", ["invoic*", "electr*", "payment*"])
def test_already_stemmed_prefix_still_matches(
self,
backend: TantivyBackend,
indexed_doc: Document,
query: str,
) -> None:
assert _matched_ids(backend, query) == {indexed_doc.id}
@pytest.mark.parametrize("query", ["univers*", "librar*"])
def test_partial_prefix_reaches_the_stemmed_term(
self,
backend: TantivyBackend,
indexed_doc: Document,
query: str,
) -> None:
"""A prefix shorter than a whole word still matches, and neither of
these needs the two-alternative path to do it.
Measured under "en": the stemmer leaves "librar" alone, so it has one
form, and that form is a prefix of the "librari" the index holds for
"library". "univers" stems to the *shorter* "univ", and the run as
typed and its stem are both prefixes of the "univers" the index holds
for "university". The case where the two forms genuinely diverge, and
only one of them matches, is
test_stem_substitution_reaches_both_the_inflection_and_the_compound.
"""
assert _matched_ids(backend, query) == {indexed_doc.id}
def test_full_word_reaches_the_stem_but_a_fragment_of_it_does_not(
self,
backend: TantivyBackend,
indexed_doc: Document,
) -> None:
"""The alternatives widen recall without turning a wildcard into a
prefix search over the original text.
"university" is stored as "univers". The stem of "universities" is
that same "univers", so the longer word matches; "universit" is a
prefix of neither its own stem nor the stored term, so the *shorter*
fragment matches nothing. usage.md names this pair, so a reader told
that `universit*` fails is also told which spelling works.
"""
assert _matched_ids(backend, "universities*") == {indexed_doc.id}
assert _matched_ids(backend, "universit*") == set()
def test_pattern_past_the_stem_boundary_is_documented_not_fixed(
self,
backend: TantivyBackend,
indexed_doc: Document,
) -> None:
"""produ*name cannot match a stemmed index ("productname" is indexed as
"productnam"); usage.md must not advertise it. Pinned so the limitation
is deliberate, not accidental."""
assert _matched_ids(backend, "produ*name") == set()
def test_stem_substitution_reaches_both_the_inflection_and_the_compound(
self,
backend: TantivyBackend,
indexed_doc: Document,
) -> None:
"""English stemming substitutes as well as truncates: "copy" and
"copies" both index as "copi", while "copyright" keeps its literal "y".
Neither form is a prefix of the other, so no single normalized string
reaches both. The run is therefore emitted as a disjunction of the
folded and stemmed forms, and "copy*" reaches the base word, its
inflections and the compound alike.
"""
compound = Document.objects.create(
title="Copyright notice",
content="copyright notice for the work",
checksum="pattern-stemming-2",
archive_serial_number=901,
)
backend.add_or_update(compound)
assert _matched_ids(backend, "copy*") == {indexed_doc.id, compound.id}
assert _matched_ids(backend, "copyright*") == {compound.id}
class TestStemsMatchTheIndexAnalyzer:
"""stem_pattern_text rebuilds paperless_text_analyzer's stemming tail rather
than sharing it, so a filter added to the index analyzer alone would silently
stop patterns from reaching the terms it produces.
"""
@pytest.mark.parametrize(
"language",
["en", "de", "fr", "es", "sv", None, "klingon"],
)
@pytest.mark.parametrize(
"word",
["Copies", "copyright", "Companies", "Invoices", "laufen", "casas", "Straße"],
)
def test_stem_equals_the_index_term(self, word: str, language: str | None) -> None:
indexed = paperless_text_analyzer(language).analyze(word)[0]
assert stem_pattern_text(ascii_fold(word.lower()), language) == indexed
def _forms(normalize: PatternNormalizer, text: str) -> tuple[str, ...]:
"""The distinct forms a term may match, in order, the way the emitter reads
the normalizer's answer (see whoosh_compat.PatternNormalizer)."""
result = normalize(text)
if isinstance(result, str):
return (result,)
return tuple(dict.fromkeys(result))
class TestPatternNormalizer:
@pytest.mark.parametrize(
("text", "expected"),
[
("Invoice", ("invoice", "invoic")),
("companies", ("companies", "compani")),
# y -> i is a substitution, so both forms are needed: the index
# holds "librari" for "library" and "library" for "librarian".
("library", ("library", "librari")),
# A run the stemmer leaves alone collapses back to one form, so it
# costs exactly the one regex branch it did before.
("invoic", ("invoic",)),
("Universit", ("universit",)),
("Café", ("cafe",)),
],
)
def test_offers_the_typed_run_and_its_stem(
self,
text: str,
expected: tuple[str, ...],
) -> None:
assert _forms(_make_pattern_normalizer("en"), text) == expected
def test_run_that_yields_no_token_falls_back_to_the_typed_run(self) -> None:
"""A run past the remove_long limit analyzes to zero tokens, so there is
no stem to offer and only the folded run remains."""
over_long = "invoices" * 20
assert _forms(_make_pattern_normalizer("en"), over_long) == (over_long,)
@pytest.mark.parametrize("language", [None, "klingon"])
def test_unstemmed_language_folds_only(self, language: str | None) -> None:
"""With no stemmer configured, or one this build has no stemmer for, the
index holds surface forms and the pattern must keep them too."""
assert _forms(_make_pattern_normalizer(language), "Invoices") == ("invoices",)
@pytest.mark.parametrize("char", ["a", "Z", "é"])
def test_a_single_character_collapses_to_one_folded_form(self, char: str) -> None:
"""A bracket class body is normalized one character at a time and the
answer is used only when it is a single one-character form, so a
stemmer that changed a lone character would silently disable folding
inside classes."""
forms = _forms(_make_pattern_normalizer("en"), char)
assert len(forms) == 1
assert len(forms[0]) == 1
class TestBracketClassStillFolds:
def test_class_body_matches_case_insensitively(
self,
backend: TantivyBackend,
indexed_doc: Document,
) -> None:
"""The class body is folded per character, which the alternatives
contract preserves only because a lone character stems to itself."""
assert _matched_ids(backend, "title:[IP]nvoice*") == {indexed_doc.id}
@@ -0,0 +1,148 @@
"""Permission filtering must hold against the real indexed document shape.
Only three of the index's unsigned ``*_id`` columns are load-bearing:
``owner_id``, ``viewer_id`` and ``viewer_group_id``, all read by
build_permission_filter. The rest (correspondent/document_type/storage_path/tag
ids) were written on every document and read by nothing, and were dropped.
These tests index real Documents through the backend's own document builder and
assert result-level visibility per user, so a mistake about which columns are
load-bearing shows up as documents leaking across users rather than as a passing
unit test over a hand-built index.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from django.contrib.auth.models import Group
from django.contrib.auth.models import User
from guardian.shortcuts import assign_perm
from documents.models import Correspondent
from documents.models import Document
from documents.models import DocumentType
from documents.models import StoragePath
from documents.models import Tag
if TYPE_CHECKING:
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
@pytest.fixture
def owner() -> User:
return User.objects.create_user(username="owner")
@pytest.fixture
def stranger() -> User:
return User.objects.create_user(username="stranger")
@pytest.fixture
def viewer() -> User:
return User.objects.create_user(username="viewer")
@pytest.fixture
def group_member() -> User:
user = User.objects.create_user(username="group_member")
user.groups.add(Group.objects.create(name="accounting"))
return user
class TestPermissionFilteringOnIndexedDocuments:
def test_unowned_document_is_visible_to_everyone(
self,
backend: TantivyBackend,
stranger: User,
) -> None:
doc = Document.objects.create(
title="Public Invoice",
content="invoice total due",
checksum="perm-unowned",
)
backend.add_or_update(doc)
assert backend.search_ids("invoice", user=stranger) == [doc.pk]
def test_owned_document_is_visible_only_to_its_owner(
self,
backend: TantivyBackend,
owner: User,
stranger: User,
) -> None:
doc = Document.objects.create(
title="Private Invoice",
content="invoice total due",
checksum="perm-owned",
owner=owner,
)
backend.add_or_update(doc)
assert backend.search_ids("invoice", user=owner) == [doc.pk]
assert backend.search_ids("invoice", user=stranger) == []
def test_explicitly_shared_document_is_visible_to_the_viewer(
self,
backend: TantivyBackend,
owner: User,
viewer: User,
stranger: User,
) -> None:
doc = Document.objects.create(
title="Shared Invoice",
content="invoice total due",
checksum="perm-shared-user",
owner=owner,
)
assign_perm("view_document", viewer, doc)
backend.add_or_update(doc)
assert backend.search_ids("invoice", user=viewer) == [doc.pk]
assert backend.search_ids("invoice", user=stranger) == []
def test_group_shared_document_is_visible_to_group_members(
self,
backend: TantivyBackend,
owner: User,
group_member: User,
stranger: User,
) -> None:
doc = Document.objects.create(
title="Group Invoice",
content="invoice total due",
checksum="perm-shared-group",
owner=owner,
)
assign_perm("view_document", group_member.groups.first(), doc)
backend.add_or_update(doc)
assert backend.search_ids("invoice", user=group_member) == [doc.pk]
assert backend.search_ids("invoice", user=stranger) == []
def test_metadata_does_not_widen_visibility(
self,
backend: TantivyBackend,
owner: User,
stranger: User,
) -> None:
"""A document carrying correspondent/type/storage-path/tag metadata is
still filtered by owner alone."""
doc = Document.objects.create(
title="Tagged Invoice",
content="invoice total due",
checksum="perm-metadata",
owner=owner,
correspondent=Correspondent.objects.create(name="ACME"),
document_type=DocumentType.objects.create(name="Bill"),
storage_path=StoragePath.objects.create(name="Archive", path="archive/"),
)
doc.tags.add(Tag.objects.create(name="paid"))
backend.add_or_update(doc)
assert backend.search_ids("invoice", user=owner) == [doc.pk]
assert backend.search_ids("invoice", user=stranger) == []
+149 -687
View File
@@ -1,448 +1,96 @@
from __future__ import annotations
import re
from datetime import UTC
from datetime import datetime
from datetime import tzinfo
from typing import TYPE_CHECKING
from zoneinfo import ZoneInfo
import pytest
import tantivy
import time_machine
from documents.search._dates import _date_only_range
from documents.search._dates import _datetime_range
from documents.search._query import build_permission_filter
from documents.search._backend import build_permission_filter
from documents.search._errors import InvalidDateQuery
from documents.search._errors import InvalidNumberQuery
from documents.search._errors import MultipleSearchQueryErrors
from documents.search._errors import SearchQueryError
from documents.search._query import parse_simple_text_highlight_query
from documents.search._query import parse_user_query
from documents.search._schema import build_schema
from documents.search._tokenizer import register_tokenizers
from documents.search._translate import InvalidDateQuery
from documents.search._translate import translate_query
if TYPE_CHECKING:
from django.contrib.auth.base_user import AbstractBaseUser
pytestmark = pytest.mark.search
EASTERN = ZoneInfo("America/New_York") # UTC-5 / UTC-4 (DST)
AUCKLAND = ZoneInfo("Pacific/Auckland") # UTC+13 in southern-hemisphere summer
@pytest.fixture(scope="module")
def query_index() -> tantivy.Index:
"""An in-memory, unstemmed index shared read-only across this module's
parse-only tests (none of them index documents)."""
schema = build_schema()
idx = tantivy.Index(schema, path=None)
register_tokenizers(idx, "")
return idx
def _range(result: str, field: str) -> tuple[str, str]:
# Half-open period ranges close with "}" (exclusive); exact-instant ranges
# (full ISO datetimes, "now", relative offsets) close with "]" (inclusive).
m = re.search(rf"{field}:\[(.+?) TO (.+?)[\]}}]", result)
assert m, f"No range for {field!r} in: {result!r}"
return m.group(1), m.group(2)
@pytest.fixture(scope="module")
def populated_index() -> tantivy.Index:
"""An index holding one document, so a query matching nothing is
distinguishable from one matching everything."""
idx = tantivy.Index(build_schema(), path=None)
register_tokenizers(idx, "")
writer = idx.writer()
doc = tantivy.Document()
doc.add_unsigned("id", 1)
doc.add_text("content", "needle in indexed content")
writer.add_document(doc)
writer.commit()
idx.reload()
return idx
class TestCreatedDateField:
"""
created is a Django DateField: indexed as midnight UTC of the local calendar
date. No offset arithmetic needed - the local calendar date is what matters.
"""
@pytest.mark.parametrize(
("tz", "expected_lo", "expected_hi"),
[
pytest.param(UTC, "2026-03-28T00:00:00Z", "2026-03-29T00:00:00Z", id="utc"),
pytest.param(
EASTERN,
"2026-03-28T00:00:00Z",
"2026-03-29T00:00:00Z",
id="eastern_same_calendar_date",
),
],
)
@time_machine.travel(datetime(2026, 3, 28, 15, 30, tzinfo=UTC), tick=False)
def test_today(self, tz: tzinfo, expected_lo: str, expected_hi: str) -> None:
lo, hi = _range(translate_query("created:today", tz), "created")
assert lo == expected_lo
assert hi == expected_hi
@time_machine.travel(datetime(2026, 3, 28, 3, 0, tzinfo=UTC), tick=False)
def test_today_auckland_ahead_of_utc(self) -> None:
# UTC 03:00 -> Auckland (UTC+13) = 16:00 same date; local date = 2026-03-28
lo, _ = _range(
translate_query("created:today", AUCKLAND),
"created",
)
assert lo == "2026-03-28T00:00:00Z"
@pytest.mark.parametrize(
("field", "keyword", "expected_lo", "expected_hi"),
[
pytest.param(
"created",
"yesterday",
"2026-03-27T00:00:00Z",
"2026-03-28T00:00:00Z",
id="yesterday",
),
pytest.param(
"created",
"previous week",
"2026-03-16T00:00:00Z",
"2026-03-23T00:00:00Z",
id="previous_week",
),
pytest.param(
"created",
"this month",
"2026-03-01T00:00:00Z",
"2026-04-01T00:00:00Z",
id="this_month",
),
pytest.param(
"created",
"previous month",
"2026-02-01T00:00:00Z",
"2026-03-01T00:00:00Z",
id="previous_month",
),
pytest.param(
"created",
"this year",
"2026-01-01T00:00:00Z",
"2027-01-01T00:00:00Z",
id="this_year",
),
pytest.param(
"created",
"previous year",
"2025-01-01T00:00:00Z",
"2026-01-01T00:00:00Z",
id="previous_year",
),
],
)
@time_machine.travel(datetime(2026, 3, 28, 15, 0, tzinfo=UTC), tick=False)
def test_date_keywords(
self,
field: str,
keyword: str,
expected_lo: str,
expected_hi: str,
) -> None:
# 2026-03-28 is Saturday; Mon-Sun week calculation built into expectations
query = f"{field}:{keyword}"
lo, hi = _range(translate_query(query, UTC), field)
assert lo == expected_lo
assert hi == expected_hi
@time_machine.travel(datetime(2026, 12, 15, 12, 0, tzinfo=UTC), tick=False)
def test_this_month_december_wraps_to_next_year(self) -> None:
# December: next month must roll over to January 1 of next year
lo, hi = _range(
translate_query("created:this month", UTC),
"created",
)
assert lo == "2026-12-01T00:00:00Z"
assert hi == "2027-01-01T00:00:00Z"
@time_machine.travel(datetime(2026, 1, 15, 12, 0, tzinfo=UTC), tick=False)
def test_last_month_january_wraps_to_previous_year(self) -> None:
# January: last month must roll back to December 1 of previous year
lo, hi = _range(
translate_query("created:previous month", UTC),
"created",
)
assert lo == "2025-12-01T00:00:00Z"
assert hi == "2026-01-01T00:00:00Z"
@time_machine.travel(datetime(2026, 7, 15, 12, 0, tzinfo=UTC), tick=False)
def test_previous_quarter(self) -> None:
lo, hi = _range(
translate_query('created:"previous quarter"', UTC),
"created",
)
assert lo == "2026-04-01T00:00:00Z"
assert hi == "2026-07-01T00:00:00Z"
def test_unknown_keyword_raises(self) -> None:
with pytest.raises(ValueError, match="Unknown keyword"):
_date_only_range("bogus_keyword", UTC)
class TestDateTimeFields:
"""
added/modified store full UTC datetimes. Natural keywords must convert
the local day boundaries to UTC - timezone offset arithmetic IS required.
"""
@time_machine.travel(datetime(2026, 3, 28, 15, 30, tzinfo=UTC), tick=False)
def test_added_today_eastern(self) -> None:
# EDT = UTC-4; local midnight 2026-03-28 00:00 EDT = 2026-03-28 04:00 UTC
lo, hi = _range(translate_query("added:today", EASTERN), "added")
assert lo == "2026-03-28T04:00:00Z"
assert hi == "2026-03-29T04:00:00Z"
@time_machine.travel(datetime(2026, 3, 29, 2, 0, tzinfo=UTC), tick=False)
def test_added_today_auckland_midnight_crossing(self) -> None:
# UTC 02:00 on 2026-03-29 -> Auckland (UTC+13) = 2026-03-29 15:00 local
# Auckland midnight = UTC 2026-03-28 11:00
lo, hi = _range(translate_query("added:today", AUCKLAND), "added")
assert lo == "2026-03-28T11:00:00Z"
assert hi == "2026-03-29T11:00:00Z"
@time_machine.travel(datetime(2026, 3, 28, 15, 0, tzinfo=UTC), tick=False)
def test_modified_today_utc(self) -> None:
lo, hi = _range(
translate_query("modified:today", UTC),
"modified",
)
assert lo == "2026-03-28T00:00:00Z"
assert hi == "2026-03-29T00:00:00Z"
@pytest.mark.parametrize(
("keyword", "expected_lo", "expected_hi"),
[
pytest.param(
"yesterday",
"2026-03-27T00:00:00Z",
"2026-03-28T00:00:00Z",
id="yesterday",
),
pytest.param(
"previous week",
"2026-03-16T00:00:00Z",
"2026-03-23T00:00:00Z",
id="previous_week",
),
pytest.param(
"this month",
"2026-03-01T00:00:00Z",
"2026-04-01T00:00:00Z",
id="this_month",
),
pytest.param(
"previous month",
"2026-02-01T00:00:00Z",
"2026-03-01T00:00:00Z",
id="previous_month",
),
pytest.param(
"this year",
"2026-01-01T00:00:00Z",
"2027-01-01T00:00:00Z",
id="this_year",
),
pytest.param(
"previous year",
"2025-01-01T00:00:00Z",
"2026-01-01T00:00:00Z",
id="previous_year",
),
],
)
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_datetime_keywords_utc(
self,
keyword: str,
expected_lo: str,
expected_hi: str,
) -> None:
# 2026-03-28 is Saturday; weekday()==5 so Monday=2026-03-23
lo, hi = _range(translate_query(f"added:{keyword}", UTC), "added")
assert lo == expected_lo
assert hi == expected_hi
@time_machine.travel(datetime(2026, 12, 15, 12, 0, tzinfo=UTC), tick=False)
def test_this_month_december_wraps_to_next_year(self) -> None:
# December: next month wraps to January of next year
lo, hi = _range(translate_query("added:this month", UTC), "added")
assert lo == "2026-12-01T00:00:00Z"
assert hi == "2027-01-01T00:00:00Z"
@time_machine.travel(datetime(2026, 1, 15, 12, 0, tzinfo=UTC), tick=False)
def test_last_month_january_wraps_to_previous_year(self) -> None:
# January: last month wraps back to December of previous year
lo, hi = _range(
translate_query("added:previous month", UTC),
"added",
)
assert lo == "2025-12-01T00:00:00Z"
assert hi == "2026-01-01T00:00:00Z"
@pytest.mark.parametrize(
("query", "expected_lo", "expected_hi"),
[
pytest.param(
'added:"previous quarter"',
"2026-04-01T00:00:00Z",
"2026-07-01T00:00:00Z",
id="quoted_previous_quarter",
),
pytest.param(
"added:previous month",
"2026-06-01T00:00:00Z",
"2026-07-01T00:00:00Z",
id="bare_previous_month",
),
pytest.param(
"added:this month",
"2026-07-01T00:00:00Z",
"2026-08-01T00:00:00Z",
id="bare_this_month",
),
],
)
@time_machine.travel(datetime(2026, 7, 15, 12, 0, tzinfo=UTC), tick=False)
def test_legacy_natural_language_aliases(
self,
query: str,
expected_lo: str,
expected_hi: str,
) -> None:
lo, hi = _range(translate_query(query, UTC), "added")
assert lo == expected_lo
assert hi == expected_hi
def test_unknown_keyword_raises(self) -> None:
with pytest.raises(ValueError, match="Unknown keyword"):
_datetime_range("bogus_keyword", UTC)
class TestWhooshQueryRewriting:
"""All Whoosh query syntax variants must be rewritten to ISO 8601 before Tantivy parses them."""
@time_machine.travel(datetime(2026, 3, 28, 15, 0, tzinfo=UTC), tick=False)
def test_compact_date_shim_rewrites_to_iso(self) -> None:
result = translate_query("created:20240115120000", UTC)
assert "2024-01-15" in result
assert "20240115120000" not in result
@time_machine.travel(datetime(2026, 3, 28, 15, 0, tzinfo=UTC), tick=False)
def test_relative_range_shim_removes_now(self) -> None:
result = translate_query("added:[now-7d TO now]", UTC)
assert "now" not in result
assert "2026-03-" in result
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_bracket_minus_7_days(self) -> None:
lo, hi = _range(
translate_query("added:[-7 days to now]", UTC),
"added",
)
assert lo == "2026-03-21T12:00:00Z"
assert hi == "2026-03-28T12:00:00Z"
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_bracket_minus_1_week(self) -> None:
lo, hi = _range(
translate_query("added:[-1 week to now]", UTC),
"added",
)
assert lo == "2026-03-21T12:00:00Z"
assert hi == "2026-03-28T12:00:00Z"
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_bracket_minus_1_month_uses_relativedelta(self) -> None:
# relativedelta(months=1) from 2026-03-28 = 2026-02-28 (not 29)
lo, hi = _range(
translate_query("created:[-1 month to now]", UTC),
"created",
)
assert lo == "2026-02-28T12:00:00Z"
assert hi == "2026-03-28T12:00:00Z"
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_bracket_minus_1_year(self) -> None:
lo, hi = _range(
translate_query("modified:[-1 year to now]", UTC),
"modified",
)
assert lo == "2025-03-28T12:00:00Z"
assert hi == "2026-03-28T12:00:00Z"
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_bracket_plural_unit_hours(self) -> None:
lo, hi = _range(
translate_query("added:[-3 hours to now]", UTC),
"added",
)
assert lo == "2026-03-28T09:00:00Z"
assert hi == "2026-03-28T12:00:00Z"
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_bracket_case_insensitive(self) -> None:
result = translate_query("added:[-1 WEEK TO NOW]", UTC)
assert "now" not in result.lower()
lo, hi = _range(result, "added")
assert lo == "2026-03-21T12:00:00Z"
assert hi == "2026-03-28T12:00:00Z"
@time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False)
def test_relative_range_swaps_bounds_when_lo_exceeds_hi(self) -> None:
# [now+1h TO now-1h] has lo > hi before substitution; they must be swapped
lo, hi = _range(
translate_query("added:[now+1h TO now-1h]", UTC),
"added",
)
assert lo == "2026-03-28T11:00:00Z"
assert hi == "2026-03-28T13:00:00Z"
def test_8digit_created_date_field_always_uses_utc_midnight(self) -> None:
# created is a DateField: boundaries are always UTC midnight, no TZ offset
result = translate_query("created:20231201", EASTERN)
lo, hi = _range(result, "created")
assert lo == "2023-12-01T00:00:00Z"
assert hi == "2023-12-02T00:00:00Z"
def test_8digit_added_datetime_field_converts_local_midnight_to_utc(self) -> None:
# added is DateTimeField: midnight Dec 1 Eastern (EST = UTC-5) = 05:00 UTC
result = translate_query("added:20231201", EASTERN)
lo, hi = _range(result, "added")
assert lo == "2023-12-01T05:00:00Z"
assert hi == "2023-12-02T05:00:00Z"
def test_8digit_modified_datetime_field_converts_local_midnight_to_utc(
self,
) -> None:
result = translate_query("modified:20231201", EASTERN)
lo, hi = _range(result, "modified")
assert lo == "2023-12-01T05:00:00Z"
assert hi == "2023-12-02T05:00:00Z"
def test_8digit_invalid_date_raises(self) -> None:
# The translation pipeline raises InvalidDateQuery for unparsable dates
# (e.g. month=13) so the API can surface a 400 telling the user the date
# is malformed instead of silently returning zero results.
with pytest.raises(InvalidDateQuery) as exc_info:
translate_query("added:20231340", UTC)
assert exc_info.value.field == "added"
assert exc_info.value.value == "20231340"
def _highlight_hit_count(index: tantivy.Index, raw_query: str) -> int:
query = parse_simple_text_highlight_query(index, raw_query)
return index.searcher().search(query, limit=1).count
class TestParseUserQuery:
"""parse_user_query runs the full preprocessing pipeline."""
@pytest.fixture
def query_index(self) -> tantivy.Index:
schema = build_schema()
idx = tantivy.Index(schema, path=None)
register_tokenizers(idx, "")
return idx
def test_returns_tantivy_query(self, query_index: tantivy.Index) -> None:
assert isinstance(parse_user_query(query_index, "invoice", UTC), tantivy.Query)
@pytest.mark.parametrize(
"raw_query",
[
pytest.param("invoice", id="plain_text"),
pytest.param("created:today", id="date_keyword"),
pytest.param("created:[2005 to 2009]", id="whoosh_date_range"),
pytest.param('added:"previous month"', id="quoted_date_phrase"),
pytest.param("title:202[0-1]*", id="bracket_class_wildcard"),
],
)
def test_fuzzy_mode_does_not_raise(
self,
query_index: tantivy.Index,
settings,
raw_query: str,
) -> None:
# These are all valid whoosh grammar that tantivy's own query parser
# (used only by the fuzzy blend clause) cannot parse; the fuzzy
# clause must degrade gracefully instead of raising and failing the
# whole query. See _try_parse_fuzzy_query.
settings.ADVANCED_FUZZY_SEARCH_THRESHOLD = 0.5
assert isinstance(parse_user_query(query_index, "invoice", UTC), tantivy.Query)
assert isinstance(parse_user_query(query_index, raw_query, UTC), tantivy.Query)
def test_date_rewriting_applied_before_tantivy_parse(
def test_date_keyword_resolves_without_raising(
self,
query_index: tantivy.Index,
) -> None:
# created:today must be rewritten to an ISO range before Tantivy parses it;
# if passed raw, Tantivy would reject "today" as an invalid date value
# whoosh-compat's DateParserPlugin resolves "today" against the AST
# directly (no string rewrite to an ISO range happens anywhere in
# this pipeline); the emitted tantivy query must still build cleanly.
with time_machine.travel(datetime(2026, 3, 28, 12, 0, tzinfo=UTC), tick=False):
q = parse_user_query(query_index, "created:today", UTC)
assert isinstance(q, tantivy.Query)
@@ -466,302 +114,58 @@ class TestParseUserQuery:
) -> None:
assert isinstance(parse_user_query(query_index, raw_query, UTC), tantivy.Query)
@pytest.mark.parametrize(
"raw_query",
[
# Partial date scalar (year only)
pytest.param("created:2020", id="created_year_scalar"),
# 8-digit compact date range in brackets
pytest.param(
"created:[20200101 TO 20201231]",
id="created_8digit_bracket_range",
),
# Comma-separated field + date range (Whoosh v2 multi-clause syntax)
pytest.param(
"title:x,created:[2020 TO 2021]",
id="title_comma_created_range",
),
# Field alias: type -> document_type
pytest.param("type:invoice", id="type_alias"),
# Multi-word date keyword
pytest.param("created:previous week", id="created_previous_week"),
# Full ISO datetime range
pytest.param(
"created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z]",
id="created_iso_range",
),
# Comma-separated ISO ranges (Whoosh v2 syntax)
pytest.param(
"created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z],"
"added:[2026-05-01T00:00:00Z TO 2026-06-01T00:00:00Z]",
id="comma_iso_ranges",
),
],
)
def test_advanced_search_queries_do_not_raise(
self,
query_index: tantivy.Index,
raw_query: str,
) -> None:
"""
End-to-end: queries that the frontend sends must parse without raising.
This tests the full pipeline: translate_query -> tantivy parse_query.
Equivalent to asserting HTTP 200 (not 400) for each query form.
"""
with time_machine.travel(datetime(2026, 6, 15, 12, 0, tzinfo=UTC), tick=False):
assert isinstance(
parse_user_query(query_index, raw_query, UTC),
tantivy.Query,
)
def test_invalid_date_propagates_not_swallowed(
self,
query_index: tantivy.Index,
) -> None:
# parse_user_query falls back to the raw query on unexpected translation
# errors, but an InvalidDateQuery is intentional and must propagate so the
# view can return a 400 instead of silently parsing the raw (invalid) date.
# parse_user_query never falls back to the raw query string on a parse
# error — a bad date diagnostic from whoosh-compat always maps to an
# InvalidDateQuery and must propagate, so the view can return a 400
# instead of silently parsing the raw (invalid) date.
with pytest.raises(InvalidDateQuery) as exc_info:
parse_user_query(query_index, "created:202023", UTC)
assert exc_info.value.field == "created"
assert exc_info.value.value == "202023"
class TestYearRangeRewriting:
"""Whoosh-style year-only date ranges must be rewritten to ISO 8601."""
@pytest.mark.parametrize(
("query", "field", "expected_lo", "expected_hi"),
[
pytest.param(
"created:[2020 TO 2020]",
"created",
"2020-01-01T00:00:00Z",
"2021-01-01T00:00:00Z",
id="single_year_created",
),
pytest.param(
"created:[2018 TO 2021]",
"created",
"2018-01-01T00:00:00Z",
"2022-01-01T00:00:00Z",
id="multi_year_range_created",
),
pytest.param(
"added:[2022 TO 2023]",
"added",
"2022-01-01T00:00:00Z",
"2024-01-01T00:00:00Z",
id="added_field",
),
pytest.param(
"modified:[2021 TO 2021]",
"modified",
"2021-01-01T00:00:00Z",
"2022-01-01T00:00:00Z",
id="modified_field",
),
pytest.param(
"created:[2020 to 2020]",
"created",
"2020-01-01T00:00:00Z",
"2021-01-01T00:00:00Z",
id="lowercase_to_keyword",
),
],
)
def test_year_range_rewritten(
def test_invalid_number_raises_invalid_number_query(
self,
query: str,
field: str,
expected_lo: str,
expected_hi: str,
query_index: tantivy.Index,
) -> None:
result = translate_query(query, UTC)
lo, hi = _range(result, field)
assert lo == expected_lo
assert hi == expected_hi
with pytest.raises(InvalidNumberQuery) as exc_info:
parse_user_query(query_index, "asn:notanumber", UTC)
assert exc_info.value.field == "asn"
assert exc_info.value.value == "notanumber"
def test_reversed_year_range_is_swapped(self) -> None:
# A reversed range must not yield lo > hi, which Tantivy treats as an
# empty range (silently zero results). The bounds are swapped instead.
result = translate_query("created:[2025 TO 2020]", UTC)
lo, hi = _range(result, "created")
assert lo == "2020-01-01T00:00:00Z"
assert hi == "2026-01-01T00:00:00Z"
def test_year_range_in_complex_boolean_query(self) -> None:
query = "tag:steuer AND (title:2020 OR (NOT title:2019 AND NOT title:2018 AND created:[2020 TO 2020]))"
result = translate_query(query, UTC)
lo, hi = _range(result, "created")
assert lo == "2020-01-01T00:00:00Z"
assert hi == "2021-01-01T00:00:00Z"
assert "title:2020" in result
assert "title:2019" in result
assert "title:2018" in result
def test_already_iso_date_range_passes_through_unchanged(self) -> None:
original = "created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z]"
assert translate_query(original, UTC) == original
def test_8digit_in_brackets_not_matched_as_year_range(self) -> None:
# [YYYYMMDD TO YYYYMMDD]: the translation layer converts 8-digit bounds to
# ISO day ranges. 20200101 -> 2020-01-01T00:00:00Z (lo of that day);
# 20201231 -> the ceil of Dec 31 = 2021-01-01T00:00:00Z (exclusive end).
# This is the correct and accepted behavior: old compact form becomes a
# proper Tantivy-parseable ISO range.
original = "created:[20200101 TO 20201231]"
result = translate_query(original, UTC)
lo, hi = _range(result, "created")
assert lo == "2020-01-01T00:00:00Z"
assert hi == "2021-01-01T00:00:00Z"
class TestNonDateFieldsNotRewritten:
"""Date rewriters must only fire on the date fields (created/modified/added).
Integer fields like asn/id/page_count and unknown fields would otherwise be
rewritten into date ranges and rejected by Tantivy as type mismatches.
"""
@pytest.mark.parametrize(
"query",
[
pytest.param("asn:20240101", id="asn_8digit"),
pytest.param("id:20240101", id="id_8digit"),
pytest.param("page_count:12345678", id="page_count_8digit"),
pytest.param("num_notes:20231201", id="num_notes_8digit"),
],
)
def test_8digit_on_integer_field_passes_through_unchanged(self, query: str) -> None:
assert translate_query(query, EASTERN) == query
@pytest.mark.parametrize(
"query",
[
pytest.param("asn:[2000 TO 2024]", id="asn_year_range"),
pytest.param("id:[2000 TO 2024]", id="id_year_range"),
pytest.param("page_count:[2000 TO 2024]", id="page_count_year_range"),
],
)
def test_year_range_on_integer_field_passes_through_unchanged(
def test_multiple_bad_fields_raise_multiple_search_query_errors(
self,
query: str,
query_index: tantivy.Index,
) -> None:
assert translate_query(query, UTC) == query
with pytest.raises(MultipleSearchQueryErrors) as exc_info:
parse_user_query(
query_index,
"created:notadate AND asn:notanumber",
UTC,
)
assert len(exc_info.value.errors) == 2
kinds = {type(e) for e in exc_info.value.errors}
assert kinds == {InvalidDateQuery, InvalidNumberQuery}
def test_unknown_field_keyword_passes_through_unchanged(self) -> None:
# foobar is not a date field: 'foobar:today' must not become a date range,
# which Tantivy would otherwise reject as an unknown/typed field.
assert translate_query("foobar:today", UTC) == "foobar:today"
class TestPassthrough:
"""Queries without field prefixes or unrelated content pass through unchanged."""
def test_bare_keyword_no_field_prefix_unchanged(self) -> None:
# Bare 'today' with no field: prefix passes through unchanged
result = translate_query("bank statement today", UTC)
assert "today" in result
def test_unrelated_query_unchanged(self) -> None:
assert translate_query("title:invoice", UTC) == "title:invoice"
class TestNormalizeQuery:
"""translate_query expands comma-separated values and collapses whitespace."""
def test_normalize_expands_comma_separated_tags(self) -> None:
assert translate_query("tag:foo,bar", UTC) == "tag:foo AND tag:bar"
def test_normalize_comma_between_range_expressions(self) -> None:
# Comma-separated field range expressions (Whoosh v2 syntax) must be
# converted to AND so Tantivy does not receive an invalid comma.
q = "created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z],added:[2026-05-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
assert translate_query(q, UTC) == (
"created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
" AND "
"added:[2026-05-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
)
def test_normalize_expands_three_values(self) -> None:
assert (
translate_query("tag:foo,bar,baz", UTC) == "tag:foo AND tag:bar AND tag:baz"
)
def test_normalize_collapses_whitespace(self) -> None:
assert translate_query("bank statement", UTC) == "bank statement"
def test_normalize_no_commas_unchanged(self) -> None:
assert translate_query("bank statement", UTC) == "bank statement"
@pytest.mark.parametrize(
("raw", "expected"),
[
pytest.param(
"h52.1 - kurzsichtigkeit",
"h52.1 kurzsichtigkeit",
id="icd_code_dash_description",
),
pytest.param(
"H52.1 - asd",
"H52.1 asd",
id="icd_code_uppercase_dash",
),
pytest.param(
"h52.1 -",
"h52.1",
id="trailing_minus",
),
pytest.param(
". -",
".",
id="dot_trailing_minus",
),
pytest.param(
"h52. -",
"h52.",
id="partial_code_trailing_minus",
),
pytest.param(
"foo - bar - baz",
"foo bar baz",
id="multiple_dashes",
),
pytest.param(
"foo + bar",
"foo bar",
id="spaced_plus_operator",
),
],
)
def test_normalize_strips_dangling_operators(self, raw: str, expected: str) -> None:
assert translate_query(raw, UTC) == expected
@pytest.mark.parametrize(
"query",
[
pytest.param("term -other", id="adjacent_not_operator"),
pytest.param("-term", id="leading_not_operator"),
pytest.param("+term", id="leading_must_operator"),
pytest.param("foo -bar +baz", id="mixed_adjacent_operators"),
],
)
def test_normalize_preserves_valid_operators(self, query: str) -> None:
assert translate_query(query, UTC) == query
def test_unregistered_id_field_folds_to_literal_text_not_error(
self,
query_index: tantivy.Index,
) -> None:
# tag_id is intentionally excluded from the FieldRegistry — whoosh-compat
# parity leniency folds it into literal text, not a diagnostic/400.
# A result-level assertion that this fold actually matches nothing
# against real documents lives in
# test_acceptance.py::TestUnregisteredIdFieldFoldsToLiteralText.
q = parse_user_query(query_index, "tag_id:5", UTC)
assert isinstance(q, tantivy.Query)
class TestParseSimpleTextHighlightQuery:
"""parse_simple_text_highlight_query must not raise on natural-language queries."""
@pytest.fixture
def query_index(self) -> tantivy.Index:
schema = build_schema()
idx = tantivy.Index(schema, path=None)
register_tokenizers(idx, "")
return idx
@pytest.mark.parametrize(
"raw_query",
[
@@ -783,16 +187,25 @@ class TestParseSimpleTextHighlightQuery:
tantivy.Query,
)
def test_empty_query_returns_empty_query(self, query_index: tantivy.Index) -> None:
result = parse_simple_text_highlight_query(query_index, "")
assert isinstance(result, tantivy.Query)
def test_all_operators_returns_empty_query(
def test_a_real_token_matches_the_corpus(
self,
query_index: tantivy.Index,
populated_index: tantivy.Index,
) -> None:
result = parse_simple_text_highlight_query(query_index, "- +")
assert isinstance(result, tantivy.Query)
"""Without this, an empty corpus would make the two assertions below
pass for a query that matches every document."""
assert _highlight_hit_count(populated_index, "needle") == 1
def test_empty_query_matches_no_document(
self,
populated_index: tantivy.Index,
) -> None:
assert _highlight_hit_count(populated_index, "") == 0
def test_all_operators_query_matches_no_document(
self,
populated_index: tantivy.Index,
) -> None:
assert _highlight_hit_count(populated_index, "- +") == 0
class TestPermissionFilter:
@@ -884,3 +297,52 @@ class TestPermissionFilter:
user = django_user_model(pk=20)
perm = build_permission_filter(perm_index.schema, user)
assert perm_index.searcher().search(perm, limit=10).count == 1 # only unowned
class TestSearchQueryErrors:
def test_invalid_date_query_is_a_search_query_error(self) -> None:
err = InvalidDateQuery("created", "notadate")
assert isinstance(err, SearchQueryError)
assert err.field == "created"
assert err.value == "notadate"
assert "created" in str(err)
assert "notadate" in str(err)
def test_invalid_number_query_is_a_search_query_error(self) -> None:
err = InvalidNumberQuery("asn", "notanumber")
assert isinstance(err, SearchQueryError)
assert err.field == "asn"
assert err.value == "notanumber"
assert "asn" in str(err)
assert "notanumber" in str(err)
def test_multiple_search_query_errors_aggregates(self) -> None:
sub_errors = [
InvalidDateQuery("created", "notadate"),
InvalidNumberQuery("asn", "notanumber"),
]
err = MultipleSearchQueryErrors(sub_errors)
assert isinstance(err, SearchQueryError)
assert err.errors == tuple(sub_errors)
assert "created" in str(err)
assert "asn" in str(err)
class TestEmitErrorContract:
"""A QueryError from emit() surfaces as a SearchQueryError (HTTP 400).
The Cause-based routing table itself is covered in test_error_routing.py.
"""
def test_exists_requires_fast_gets_the_user_facing_rewrite(
self,
query_index: tantivy.Index,
) -> None:
# whoosh-compat's own message advises a host-side fast=True config
# change the user can't act on, so this checks OUR wording, not
# whoosh-compat's (that's its own test suite's job now).
with pytest.raises(SearchQueryError) as exc_info:
parse_user_query(query_index, "notes.user:*", UTC)
assert str(exc_info.value) == (
"Existence searches (field:*) are not supported for field 'notes.user'."
)
@@ -0,0 +1,153 @@
"""Negation must survive the blended query.
parse_user_query ORs an exact clause with optional fuzzy and CJK clauses.
Each of those is built from positive terms only, so unless the query's
exclusions are applied to the blend as a whole, a document the exact
clause excluded is re-admitted by whichever other clause is enabled.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from documents.models import Document
if TYPE_CHECKING:
from pytest_django.fixtures import SettingsWrapper
from documents.search._backend import TantivyBackend
pytestmark = [pytest.mark.search, pytest.mark.django_db]
def _matched_ids(backend: TantivyBackend, query: str) -> set[int]:
return set(backend.search_ids(query, user=None))
def _index(backend: TantivyBackend, **kwargs: object) -> Document:
doc = Document.objects.create(**kwargs)
backend.add_or_update(doc)
return doc
@pytest.fixture
def fuzzy_enabled(settings: SettingsWrapper) -> None:
"""Enable the fuzzy blend clause. The threshold doubles as a minimum
score filter, so it is set to 0.0: every hit passes and the test sees
the clause's matching behaviour, not the filter's."""
settings.ADVANCED_FUZZY_SEARCH_THRESHOLD = 0.0
class TestNegationConstrainsEveryClause:
@pytest.mark.usefixtures("fuzzy_enabled")
def test_fuzzy_clause_does_not_readmit_an_excluded_document(
self,
backend: TantivyBackend,
) -> None:
secret = _index(
backend,
title="Invoice A",
content="invoice total secret",
checksum="neg-fuzzy-1",
)
public = _index(
backend,
title="Invoice B",
content="invoice total public",
checksum="neg-fuzzy-2",
)
assert _matched_ids(backend, "invoice") == {secret.pk, public.pk}
assert _matched_ids(backend, "invoice NOT secret") == {public.pk}
def test_cjk_clause_does_not_readmit_an_excluded_document(
self,
backend: TantivyBackend,
) -> None:
"""The CJK clause legitimately carries 東京 here, so rebuilding it
from the AST cannot help: only applying the exclusion above the
blend keeps the secret document out."""
secret = _index(
backend,
title="Tokyo A",
content="東京都の秘密です secret",
checksum="neg-cjk-1",
)
public = _index(
backend,
title="Tokyo B",
content="東京都の報告書です public",
checksum="neg-cjk-2",
)
assert _matched_ids(backend, "東京") == {secret.pk, public.pk}
assert _matched_ids(backend, "東京 NOT secret") == {public.pk}
@pytest.mark.usefixtures("fuzzy_enabled")
def test_disjunctive_negation_still_admits_the_other_branch(
self,
backend: TantivyBackend,
) -> None:
"""'invoice OR NOT secret' excludes nothing on its own: a document
matching the left branch stays in even though it contains secret."""
secret_invoice = _index(
backend,
title="Invoice A",
content="invoice total secret",
checksum="neg-or-1",
)
unrelated = _index(
backend,
title="Recipe",
content="flour and water",
checksum="neg-or-2",
)
assert _matched_ids(backend, "invoice OR NOT secret") == {
secret_invoice.pk,
unrelated.pk,
}
def test_a_negation_under_or_does_not_constrain_the_cjk_clause(
self,
backend: TantivyBackend,
) -> None:
"""The limit of the hoist, pinned deliberately.
An exclusion that is one branch's own condition cannot be restated
above the blend without dropping documents the other branch
matches, so it is left where it is and the CJK clause stays
unconstrained by it. That shows through here in a way it does not
for latin text: the exact clause cannot match a CJK run at all, so
the CJK clause is the only thing matching the tokyo documents, and
the secret one comes with it.
"""
secret = _index(
backend,
title="Tokyo A",
content="東京都の秘密です secret",
checksum="neg-or-cjk-1",
)
public = _index(
backend,
title="Tokyo B",
content="東京都の報告書です public",
checksum="neg-or-cjk-2",
)
bill = _index(
backend,
title="Bill",
content="bill payment received",
checksum="neg-or-cjk-3",
)
assert _matched_ids(backend, "(東京 AND NOT secret) OR bill") == {
bill.pk,
public.pk,
secret.pk,
}
# The same exclusion in conjunctive position is hoisted, and does
# constrain the CJK clause.
assert _matched_ids(backend, "東京 AND NOT secret") == {public.pk}
+149
View File
@@ -0,0 +1,149 @@
from collections.abc import Sequence
import pytest
from whoosh_compat import FieldKind
from whoosh_compat import FieldRegistry
from whoosh_compat.fields import ResolvedField
from documents.search._fields import PUBLIC_FIELDS
from documents.search._registry import get_field_registry
@pytest.fixture
def registry() -> FieldRegistry:
return get_field_registry(None)
def _resolve(registry: FieldRegistry, name: str) -> ResolvedField:
ref = registry.make_ref(name)
assert ref is not None, f"{name} is not a valid field ref"
resolved = registry.resolve(ref)
assert resolved is not None, f"{name} did not resolve"
return resolved
def _distinct_forms(result: str | Sequence[str]) -> tuple[str, ...]:
"""The forms a term may match, in order, the way whoosh-compat's emitter
reads a pattern_normalizer's answer: a bare str is one form, a sequence is
several, deduplicated."""
if isinstance(result, str):
return (result,)
return tuple(dict.fromkeys(result))
class TestFieldRegistry:
def test_internal_id_fields_are_not_registered(
self,
registry: FieldRegistry,
) -> None:
for name in (
"tag_id",
"owner_id",
"viewer_id",
"correspondent_id",
"document_type_id",
"storage_path_id",
"viewer_group_id",
):
assert name not in registry
def test_no_queryable_field_name_ends_in_id(self) -> None:
# The list above names the seven that were dropped; this catches the
# eighth. Internal *_id columns are written for permission filtering
# and joins, and whoosh only exposed them as query fields by accident,
# so a new one reaching the query surface is a leak rather than a
# feature. Checked against PUBLIC_FIELDS rather than the registry so
# an internal field is caught where it is declared.
leaked = [f.name for f in PUBLIC_FIELDS if f.name.endswith("_id")]
assert not leaked, f"internal id fields reached the query surface: {leaked}"
def test_type_alias_resolves_to_document_type(
self,
registry: FieldRegistry,
) -> None:
assert _resolve(registry, "type").spec.name == "document_type"
def test_path_alias_resolves_to_storage_path(self, registry: FieldRegistry) -> None:
assert _resolve(registry, "path").spec.name == "storage_path"
def test_notes_json_subpaths_resolve(self, registry: FieldRegistry) -> None:
resolved = _resolve(registry, "notes.user")
assert resolved.spec.name == "notes"
assert resolved.json_path == "user"
assert resolved.is_subpath is True
def test_custom_fields_json_subpaths_resolve(self, registry: FieldRegistry) -> None:
for raw in ("custom_fields.name", "custom_fields.value"):
_resolve(registry, raw)
def test_unregistered_json_subpath_does_not_resolve(
self,
registry: FieldRegistry,
) -> None:
# An unregistered subpath is not even a valid FieldRef: make_ref
# returns None for a dotted name whose subpath isn't registered
# (it doesn't produce a ref for resolve() to then reject).
assert registry.make_ref("notes.bogus") is None
def test_tag_is_comma_values(self, registry: FieldRegistry) -> None:
assert _resolve(registry, "tag").spec.comma_values is True
def test_correspondent_is_not_comma_values(self, registry: FieldRegistry) -> None:
# "tag" is the only field that opts in. This is only observable here:
# end to end the two readings of "correspondent:foo,bar" agree,
# because the analyzer splits the literal value on the comma anyway,
# so a result-level test cannot tell a value list from literal text.
assert _resolve(registry, "correspondent").spec.comma_values is False
def test_created_is_date_kind(self, registry: FieldRegistry) -> None:
resolved = _resolve(registry, "created")
assert resolved.spec.kind is FieldKind.DATE
assert resolved.spec.date_only is True
def test_analyzer_lowercases_and_ascii_folds(self, registry: FieldRegistry) -> None:
# title uses the paperless_text analyzer: simple -> remove_long ->
# lowercase -> ascii_fold [-> stemmer]. With no language configured
# (None), no stemmer runs, so "Café" folds to the single token "cafe".
resolved = _resolve(registry, "title")
assert resolved.spec.analyzer is not None
assert resolved.spec.analyzer("Café") == ["cafe"]
def test_checksum_analyzer_is_identity_single_token(
self,
registry: FieldRegistry,
) -> None:
# checksum uses the raw tokenizer at index time (no splitting).
resolved = _resolve(registry, "checksum")
assert resolved.spec.analyzer is not None
assert resolved.spec.analyzer("ABC-123") == ["ABC-123"]
def test_pattern_normalizer_follows_the_registry_language(
self,
registry: FieldRegistry,
) -> None:
# Index terms are stemmed, so patterns offer their stem too, using the
# registry's own language: "Running" has to reach the indexed "run".
# Without a language the index holds surface forms, so there is no
# second form and the run is only case/accent-folded.
resolved = _resolve(registry, "title")
assert resolved.spec.pattern_normalizer is not None
assert _distinct_forms(resolved.spec.pattern_normalizer("Running")) == (
"running",
)
resolved_en = _resolve(get_field_registry("en"), "title")
assert resolved_en.spec.pattern_normalizer is not None
assert _distinct_forms(resolved_en.spec.pattern_normalizer("Running")) == (
"running",
"run",
)
def test_registry_is_cached_per_language(self) -> None:
a = get_field_registry("en")
b = get_field_registry("en")
assert a is b
def test_registry_rebuilds_on_language_change(self) -> None:
a = get_field_registry("en")
b = get_field_registry("de")
assert a is not b
+84 -1
View File
@@ -1,12 +1,20 @@
from __future__ import annotations
import json
from datetime import UTC
from datetime import datetime
from typing import TYPE_CHECKING
import pytest
import tantivy
from documents.search._fields import PUBLIC_FIELDS
from documents.search._schema import SCHEMA_VERSION
from documents.search._schema import build_schema
from documents.search._schema import field_descriptors
from documents.search._schema import needs_rebuild
from documents.search._schema import schema_fingerprint
from documents.search._tokenizer import register_tokenizers
if TYPE_CHECKING:
from pathlib import Path
@@ -29,7 +37,13 @@ class TestNeedsRebuild:
) -> None:
settings.SEARCH_LANGUAGE = "en"
(index_dir / ".index_settings.json").write_text(
json.dumps({"schema_version": SCHEMA_VERSION, "language": "en"}),
json.dumps(
{
"schema_version": SCHEMA_VERSION,
"language": "en",
"schema_fingerprint": schema_fingerprint(),
},
),
)
assert needs_rebuild(index_dir) is False
@@ -76,3 +90,72 @@ class TestNeedsRebuild:
json.dumps({"schema_version": SCHEMA_VERSION, "language": "en"}),
)
assert needs_rebuild(index_dir) is True
def _schema_fields(schema: tantivy.Schema) -> dict[str, dict]:
"""{name: field-state} for every field declared on a tantivy Schema.
tantivy-py 0.26 exposes no public introspection API on Schema (no
__iter__, get_field, to_json, etc.) -- __reduce__() (used internally for
pickling) is the only way to recover the field list, so we lean on it
here for test assertions only.
"""
state = schema.__reduce__()[1][0]
return {field["name"]: field for field in state["inner"]}
class TestSchemaMatchesPublicFields:
def test_every_public_field_is_in_the_schema(self) -> None:
schema = build_schema()
schema_field_names = set(_schema_fields(schema))
for field in PUBLIC_FIELDS:
assert field.name in schema_field_names, (
f"{field.name} is in PUBLIC_FIELDS but missing from build_schema()"
)
def test_asn_page_count_num_notes_are_fast_unsigned_fields(self) -> None:
# Spot-check kind-derived construction for the U64 fields.
schema = build_schema()
doc = tantivy.Document()
doc.add_unsigned("id", 1)
doc.add_text("checksum", "x")
doc.add_unsigned("asn", 42)
doc.add_unsigned("page_count", 3)
doc.add_unsigned("num_notes", 0)
doc.add_date("created", datetime(2020, 1, 1, tzinfo=UTC))
doc.add_date("modified", datetime(2020, 1, 1, tzinfo=UTC))
doc.add_date("added", datetime(2020, 1, 1, tzinfo=UTC))
index = tantivy.Index(schema)
register_tokenizers(index, None)
writer = index.writer()
writer.add_document(doc)
writer.commit()
index.reload()
searcher = index.searcher()
results = searcher.search(tantivy.Query.term_query(schema, "asn", 42), limit=1)
assert len(results.hits) == 1
class TestFastFlagAgreement:
def test_every_public_field_fast_flag_matches_the_built_schema(self) -> None:
# whoosh-compat's registry trusts PUBLIC_FIELDS' fast flag when resolving
# field:* existence checks (its FAST_FIELD strategy); a fast=True
# entry whose actual tantivy column is not fast would make those
# searches silently match nothing at search time. Only the U64 and
# DATE descriptors can carry the flag today, so this
# pins the agreement for EVERY kind: a future fast=True
# TEXT/KEYWORD/JSON entry the builder silently ignores fails here
# instead of at a user's query.
#
# field_descriptors() (not tantivy-py's __reduce__() pickling
# internals) is used as the probe here: it is exactly the input
# build_schema()'s SchemaBuilder consumes for the `fast` kwarg on
# every field kind, so it pins the same agreement without depending
# on a private pickled representation surviving a tantivy-py
# upgrade.
descriptor_fast = {d.name: d.fast for d in field_descriptors()}
for public_field in PUBLIC_FIELDS:
assert descriptor_fast[public_field.name] == public_field.fast, (
f"{public_field.name}: PUBLIC_FIELDS says fast={public_field.fast} but"
f" field_descriptors() says fast={descriptor_fast[public_field.name]}"
)
@@ -0,0 +1,492 @@
"""The schema fingerprint stamped into .index_settings.json.
tantivy compares schemas by *ordered* field list, and `tantivy.Index(schema,
path=...)` (what every write path does) raises on any difference. SCHEMA_VERSION
is the manual guard against that, but build_schema() is edited for *parser*
reasons - adding an alias, flipping fast=True, adding a subpath - by people not
thinking about the on-disk index, and forgetting the bump is exactly how this
branch's bug happened.
The fingerprint is the automatic guard: it hashes the field descriptor list that
build_schema() itself iterates, so any change to a field's name, kind, options
or *position* forces a rebuild on its own.
"""
from __future__ import annotations
import hashlib
import json
from typing import TYPE_CHECKING
import pytest
import tantivy
from documents.search import _schema
from documents.search._schema import SCHEMA_VERSION
from documents.search._schema import FieldDescriptor
from documents.search._schema import _write_sentinels
from documents.search._schema import build_schema
from documents.search._schema import field_descriptors
from documents.search._schema import needs_rebuild
from documents.search._schema import schema_fingerprint
if TYPE_CHECKING:
from pathlib import Path
from pytest_django.fixtures import SettingsWrapper
pytestmark = pytest.mark.search
# The on-disk field layout of a v2 index, pinned as data. Any edit here is an
# index-format change: it must come with a rebuild, which the fingerprint now
# forces automatically. Reproduced from build_schema()'s output as it stood
# before the descriptor refactor, so it also pins that the refactor changed
# nothing.
PINNED_DESCRIPTORS: tuple[FieldDescriptor, ...] = (
FieldDescriptor("id", "u64", stored=True, indexed=True, fast=True, tokenizer=None),
FieldDescriptor(
"title",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"content",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"correspondent",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"document_type",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"storage_path",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"original_filename",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"tag",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"checksum",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="raw",
),
FieldDescriptor("asn", "u64", stored=True, indexed=True, fast=True, tokenizer=None),
FieldDescriptor(
"page_count",
"u64",
stored=True,
indexed=True,
fast=True,
tokenizer=None,
),
FieldDescriptor(
"num_notes",
"u64",
stored=True,
indexed=True,
fast=True,
tokenizer=None,
),
FieldDescriptor(
"created",
"date",
stored=True,
indexed=True,
fast=True,
tokenizer=None,
),
FieldDescriptor(
"modified",
"date",
stored=True,
indexed=True,
fast=True,
tokenizer=None,
),
FieldDescriptor(
"added",
"date",
stored=True,
indexed=True,
fast=True,
tokenizer=None,
),
FieldDescriptor(
"notes",
"json",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"notes_text",
"text",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"custom_fields",
"json",
stored=True,
indexed=True,
fast=False,
tokenizer="paperless_text",
),
FieldDescriptor(
"title_sort",
"text",
stored=False,
indexed=True,
fast=True,
tokenizer="simple_analyzer",
),
FieldDescriptor(
"correspondent_sort",
"text",
stored=False,
indexed=True,
fast=True,
tokenizer="simple_analyzer",
),
FieldDescriptor(
"type_sort",
"text",
stored=False,
indexed=True,
fast=True,
tokenizer="simple_analyzer",
),
FieldDescriptor(
"bigram_content",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="bigram_analyzer",
),
FieldDescriptor(
"bigram_title",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="bigram_analyzer",
),
FieldDescriptor(
"bigram_correspondent",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="bigram_analyzer",
),
FieldDescriptor(
"bigram_document_type",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="bigram_analyzer",
),
FieldDescriptor(
"bigram_tag",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="bigram_analyzer",
),
FieldDescriptor(
"simple_title",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="simple_search_analyzer",
),
FieldDescriptor(
"simple_content",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="simple_search_analyzer",
),
FieldDescriptor(
"autocomplete_word",
"text",
stored=False,
indexed=True,
fast=False,
tokenizer="raw",
),
FieldDescriptor(
"owner_id",
"u64",
stored=False,
indexed=True,
fast=True,
tokenizer=None,
),
FieldDescriptor(
"viewer_id",
"u64",
stored=False,
indexed=True,
fast=True,
tokenizer=None,
),
FieldDescriptor(
"viewer_group_id",
"u64",
stored=False,
indexed=True,
fast=True,
tokenizer=None,
),
)
def _schema_fields(schema: tantivy.Schema) -> list[dict]:
"""The tantivy-level field list, in declaration order.
tantivy-py 0.26 exposes no public introspection API on Schema, so
__reduce__() (its pickling hook) is the only way to recover the field list.
It is used here, in a test, precisely because it is the representation the
persisted fingerprint must NOT depend on.
"""
return schema.__reduce__()[1][0]["inner"]
def _sentinels(index_dir: Path, **overrides: object) -> None:
data = {
"schema_version": SCHEMA_VERSION,
"language": None,
"schema_fingerprint": schema_fingerprint(),
}
data.update(overrides)
(index_dir / ".index_settings.json").write_text(json.dumps(data))
class TestDescriptorsDescribeTheBuiltSchema:
def test_descriptors_match_the_pinned_field_layout(self) -> None:
assert tuple(field_descriptors()) == PINNED_DESCRIPTORS
def test_built_schema_matches_the_descriptors(self) -> None:
"""The descriptors are not a parallel description - they are the input.
Reading the built schema back proves the loop honours every option, so
a descriptor edit cannot claim a shape the SchemaBuilder did not build.
"""
kinds = {"text": "text", "json": "json_object", "u64": "u64", "date": "date"}
built = [
(
field["name"],
field["type"],
field["options"]["stored"],
bool(field["options"].get("fast")),
(field["options"].get("indexing") or {}).get("tokenizer"),
)
for field in _schema_fields(build_schema())
]
expected = [
(
descriptor.name,
kinds[descriptor.kind],
descriptor.stored,
descriptor.fast,
descriptor.tokenizer,
)
for descriptor in field_descriptors()
]
assert built == expected
class TestFingerprintSensitivity:
def test_a_field_option_change_moves_the_fingerprint(
self,
monkeypatch: pytest.MonkeyPatch,
) -> None:
before = schema_fingerprint()
changed = field_descriptors()
changed[1] = changed[1]._replace(fast=True)
monkeypatch.setattr(_schema, "field_descriptors", lambda: changed)
assert schema_fingerprint() != before
def test_reordering_alone_moves_the_fingerprint(
self,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The original bug: same fields, different declaration order.
A set- or dict-based fingerprint would be blind to this, and tantivy
would reject every write against the existing index.
"""
before = schema_fingerprint()
swapped = field_descriptors()
swapped[1], swapped[2] = swapped[2], swapped[1]
monkeypatch.setattr(_schema, "field_descriptors", lambda: swapped)
assert schema_fingerprint() != before
def test_repeated_calls_agree(self) -> None:
assert schema_fingerprint() == schema_fingerprint()
class TestFingerprintIsIndependentOfTantivy:
def test_a_tantivy_option_key_addition_would_not_move_it(self) -> None:
"""A tantivy-py upgrade must not force a global reindex.
Hashing schema.__reduce__() would do exactly that: the simulated new
option key below changes that payload for every user with no schema
change at all.
"""
fields = _schema_fields(build_schema())
upgraded = [
{**field, "options": {**field["options"], "coerce": True}}
for field in fields
]
assert _hash(upgraded) != _hash(fields)
assert schema_fingerprint() == _fingerprint_of(field_descriptors())
def test_fingerprint_never_touches_the_schema_builder(
self,
monkeypatch: pytest.MonkeyPatch,
) -> None:
before = schema_fingerprint()
class _RemovedSchemaBuilder:
def __init__(self) -> None:
raise AssertionError("tantivy.SchemaBuilder was consulted")
monkeypatch.setattr(tantivy, "SchemaBuilder", _RemovedSchemaBuilder)
with pytest.raises(AssertionError):
build_schema()
assert schema_fingerprint() == before
def _hash(payload: object) -> str:
return hashlib.blake2b(json.dumps(payload).encode()).hexdigest()
def _fingerprint_of(descriptors: list[FieldDescriptor]) -> str:
return _hash([list(descriptor) for descriptor in descriptors])
class TestNeedsRebuildOnFingerprint:
def test_matching_fingerprint_does_not_rebuild(
self,
index_dir: Path,
settings: SettingsWrapper,
) -> None:
settings.SEARCH_LANGUAGE = None
_sentinels(index_dir)
assert needs_rebuild(index_dir) is False
def test_stale_fingerprint_rebuilds_despite_a_matching_version(
self,
index_dir: Path,
settings: SettingsWrapper,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The failure this task exists to prevent: schema edited, version not
bumped. Without the fingerprint check, `reindex --if-needed` reports the
index up to date and every write then raises."""
settings.SEARCH_LANGUAGE = None
_sentinels(index_dir)
extended = [
*field_descriptors(),
FieldDescriptor(
"new_field",
"u64",
stored=False,
indexed=True,
fast=True,
tokenizer=None,
),
]
monkeypatch.setattr(_schema, "field_descriptors", lambda: extended)
assert needs_rebuild(index_dir) is True
def test_reordered_schema_rebuilds(
self,
index_dir: Path,
settings: SettingsWrapper,
monkeypatch: pytest.MonkeyPatch,
) -> None:
settings.SEARCH_LANGUAGE = None
_sentinels(index_dir)
reordered = field_descriptors()
reordered[1], reordered[2] = reordered[2], reordered[1]
monkeypatch.setattr(_schema, "field_descriptors", lambda: reordered)
assert needs_rebuild(index_dir) is True
def test_missing_fingerprint_rebuilds(
self,
index_dir: Path,
settings: SettingsWrapper,
) -> None:
"""No seeding: an index whose schema shape nobody recorded is rebuilt
rather than trusted."""
settings.SEARCH_LANGUAGE = None
(index_dir / ".index_settings.json").write_text(
json.dumps({"schema_version": SCHEMA_VERSION, "language": None}),
)
assert needs_rebuild(index_dir) is True
def test_written_sentinels_satisfy_the_check(
self,
index_dir: Path,
settings: SettingsWrapper,
) -> None:
settings.SEARCH_LANGUAGE = "en"
_write_sentinels(index_dir)
assert needs_rebuild(index_dir) is False
@@ -0,0 +1,164 @@
"""SCHEMA_VERSION must change whenever build_schema()'s field list or order does.
tantivy compares schemas by *ordered* field list. ``Index.open()`` loads the
schema from the index's own ``meta.json``, so reads against an index built by an
older release keep working after a field reorder. Writes do not:
``WriteBatch.__enter__`` calls ``tantivy.Index(build_schema(), path=...)``, an
open-or-create that raises ``ValueError`` on any schema difference. Nothing
catches that ValueError, so consumption, index_document and bulk edit all
hard-fail while ``/api/status/`` still reports the index healthy.
The only thing that saves such an install is ``needs_rebuild()`` noticing the
version stamped in ``.index_settings.json`` is stale.
"""
from __future__ import annotations
import json
from typing import TYPE_CHECKING
import pytest
import tantivy
from django.conf import settings as django_settings
from documents.search._schema import build_schema
from documents.search._schema import needs_rebuild
from documents.search._schema import open_or_rebuild_index
if TYPE_CHECKING:
from pathlib import Path
pytestmark = [pytest.mark.search]
RELEASED_V1_SCHEMA_VERSION = 1
def _build_released_v1_schema() -> tantivy.Schema:
"""Frozen copy of build_schema() as shipped in v3.0.x (schema version 1).
Deliberately duplicated rather than imported: it must keep describing the
on-disk layout of already-deployed indexes even as build_schema() evolves.
"""
sb = tantivy.SchemaBuilder()
sb.add_unsigned_field("id", stored=True, indexed=True, fast=True)
sb.add_text_field("checksum", stored=True, tokenizer_name="raw")
for field in (
"title",
"correspondent",
"document_type",
"storage_path",
"original_filename",
"content",
):
sb.add_text_field(field, stored=True, tokenizer_name="paperless_text")
for field in ("title_sort", "correspondent_sort", "type_sort"):
sb.add_text_field(
field,
stored=False,
tokenizer_name="simple_analyzer",
fast=True,
)
for field in (
"bigram_content",
"bigram_title",
"bigram_correspondent",
"bigram_document_type",
"bigram_tag",
):
sb.add_text_field(field, stored=False, tokenizer_name="bigram_analyzer")
for field in ("simple_title", "simple_content"):
sb.add_text_field(field, stored=False, tokenizer_name="simple_search_analyzer")
sb.add_text_field("autocomplete_word", stored=False, tokenizer_name="raw")
sb.add_text_field("tag", stored=True, tokenizer_name="paperless_text")
sb.add_json_field("notes", stored=True, tokenizer_name="paperless_text")
sb.add_text_field("notes_text", stored=True, tokenizer_name="paperless_text")
sb.add_json_field("custom_fields", stored=True, tokenizer_name="paperless_text")
for field in (
"correspondent_id",
"document_type_id",
"storage_path_id",
"tag_id",
"owner_id",
"viewer_id",
"viewer_group_id",
):
sb.add_unsigned_field(field, stored=False, indexed=True, fast=True)
for field in ("created", "modified", "added"):
sb.add_date_field(field, stored=True, indexed=True, fast=True)
for field in ("asn", "page_count", "num_notes"):
sb.add_unsigned_field(field, stored=True, indexed=True, fast=True)
return sb.build()
@pytest.fixture
def released_v1_index(tmp_path: Path) -> Path:
"""An index directory as a v3.0.x install would leave it on disk."""
index_dir = tmp_path / "index"
index_dir.mkdir()
tantivy.Index(_build_released_v1_schema(), path=str(index_dir))
(index_dir / ".index_settings.json").write_text(
json.dumps(
{
"schema_version": RELEASED_V1_SCHEMA_VERSION,
"language": django_settings.SEARCH_LANGUAGE,
},
),
)
return index_dir
class TestUpgradeFromReleasedV1Index:
def test_released_v1_index_is_flagged_for_rebuild(
self,
released_v1_index: Path,
) -> None:
"""The current schema differs from v1's, so the sentinel must be stale.
If this fails, `document_index reindex --if-needed` prints "Search index
is up to date" and skips, leaving the mismatched index in place.
"""
assert needs_rebuild(released_v1_index) is True
def test_v1_index_rejects_writes_against_the_current_schema(
self,
released_v1_index: Path,
) -> None:
"""The failure mode the version bump exists to prevent.
This is exactly what WriteBatch.__enter__ does on every index write.
"""
with pytest.raises(ValueError, match="schema does not match"):
tantivy.Index(build_schema(), path=str(released_v1_index))
def test_opening_a_v1_index_leaves_it_writable(
self,
released_v1_index: Path,
) -> None:
"""End to end: open_or_rebuild_index must hand back an index that the
write path can reopen. Before the version bump, needs_rebuild() returned
False here, the stale directory survived untouched, and every subsequent
write raised the ValueError above."""
open_or_rebuild_index(released_v1_index)
tantivy.Index(build_schema(), path=str(released_v1_index))
def test_rebuilt_index_is_not_rebuilt_again(
self,
released_v1_index: Path,
) -> None:
"""The rebuild must stamp the version it actually wrote, otherwise every
startup wipes and reindexes the whole corpus."""
open_or_rebuild_index(released_v1_index)
assert needs_rebuild(released_v1_index) is False
+2 -2
View File
@@ -7,8 +7,8 @@ import pytest
import tantivy
from documents.search._tokenizer import _bigram_analyzer
from documents.search._tokenizer import _paperless_text
from documents.search._tokenizer import _simple_search_analyzer
from documents.search._tokenizer import paperless_text_analyzer
from documents.search._tokenizer import register_tokenizers
if TYPE_CHECKING:
@@ -25,7 +25,7 @@ class TestTokenizers:
sb.add_text_field("content", stored=True, tokenizer_name="paperless_text")
schema = sb.build()
idx = tantivy.Index(schema, path=None)
idx.register_tokenizer("paperless_text", _paperless_text(""))
idx.register_tokenizer("paperless_text", paperless_text_analyzer(""))
return idx
@pytest.fixture
@@ -1,810 +0,0 @@
from __future__ import annotations
from datetime import UTC
from datetime import datetime
from typing import TYPE_CHECKING
from zoneinfo import ZoneInfo
import pytest
import time_machine
from documents.search._dates import _precision_bounds
if TYPE_CHECKING:
import tantivy
from documents.search._query import _FIELD_BOOSTS
from documents.search._query import DEFAULT_SEARCH_FIELDS
from documents.search._translate import OPEN_HI
from documents.search._translate import OPEN_LO
from documents.search._translate import Comma
from documents.search._translate import FieldRange
from documents.search._translate import FieldValue
from documents.search._translate import FieldValueList
from documents.search._translate import InvalidDateQuery
from documents.search._translate import Passthrough
from documents.search._translate import resolve_commas
from documents.search._translate import scan
from documents.search._translate import translate_query
from documents.search._translate import translate_range
from documents.search._translate import translate_scalar
@pytest.mark.search
class TestPrecisionBounds:
@pytest.mark.parametrize(
("digits", "expected"),
[
("2020", ((2020, 1, 1), (2021, 1, 1))),
("202003", ((2020, 3, 1), (2020, 4, 1))),
("202012", ((2020, 12, 1), (2021, 1, 1))),
("20200115", ((2020, 1, 15), (2020, 1, 16))),
("20201231", ((2020, 12, 31), (2021, 1, 1))),
],
)
def test_valid(self, digits, expected):
lo, hi = _precision_bounds(digits)
assert (lo.year, lo.month, lo.day) == expected[0]
assert (hi.year, hi.month, hi.day) == expected[1]
@pytest.mark.parametrize("digits", ["202023", "20200230", "20201301", "20", "abcd"])
def test_invalid_returns_none(self, digits):
assert _precision_bounds(digits) is None
@pytest.mark.search
class TestScan:
def test_plain_words_are_passthrough(self):
assert scan("bank statement") == [Passthrough("bank statement")]
def test_field_value(self):
assert scan("created:2020") == [FieldValue("created", "2020")]
def test_field_value_in_boolean(self):
toks = scan("created:2020 OR foo")
assert toks == [
FieldValue("created", "2020"),
Passthrough(" OR foo"),
]
def test_field_value_in_parens(self):
toks = scan("(created:2020 OR foo)")
assert toks == [
Passthrough("("),
FieldValue("created", "2020"),
Passthrough(" OR foo)"),
]
def test_quoted_value(self):
assert scan('correspondent:"A B"') == [FieldValue("correspondent", '"A B"')]
def test_field_range(self):
assert scan("created:[2020 TO 2021]") == [
FieldRange("created", "[", "2020", "2021", "]"),
]
@pytest.mark.parametrize(
("query", "expected"),
[
pytest.param(
"created:[2020 to]",
FieldRange("created", "[", "2020", "", "]"),
id="open_upper",
),
pytest.param(
"created:[to 2020]",
FieldRange("created", "[", "", "2020", "]"),
id="open_lower",
),
],
)
def test_open_range(self, query, expected):
assert scan(query) == [expected]
def test_comma_inside_range_not_split(self):
# No depth-0 comma here; the whole thing is one range token.
toks = scan("created:[2020 TO 2021]")
assert len(toks) == 1
# --- Edge-case / regression tests (scan must never raise) ---
def test_url_is_passthrough(self):
# "http" is not a known field; the whole URL must pass through verbatim.
assert scan("http://example.com") == [Passthrough("http://example.com")]
def test_unterminated_quote_is_passthrough(self):
# title is a known field but the quoted value has no closing quote;
# _consume_value returns None so the whole string falls into passthrough.
assert scan('title:"abc') == [Passthrough('title:"abc')]
def test_unterminated_bracket_is_passthrough(self):
# created is a known field but the range bracket is never closed;
# _consume_range returns None so the whole string falls into passthrough.
assert scan("created:[2020") == [Passthrough("created:[2020")]
def test_empty_value_at_end_is_passthrough(self):
# created is a known field but there is no value after the colon
# (_consume_value returns None for start >= n), so passthrough.
assert scan("created:") == [Passthrough("created:")]
def test_value_containing_colon(self):
# The bare-word value reader stops at whitespace/paren, not at colon,
# so "2020:30" is consumed as a single value token.
assert scan("created:2020:30") == [FieldValue("created", "2020:30")]
def test_comma_followed_by_unconsumable_value_stops(self):
# A comma followed by whitespace is neither a value-list continuation nor a
# clause separator: the value stops and the comma stays as passthrough.
assert scan("tag:foo, bar") == [
FieldValue("tag", "foo"),
Passthrough(", bar"),
]
def test_bracket_without_to_is_open_upper_bound(self):
# A bracketed value with no TO falls back to (value, "") -> open upper bound.
assert scan("created:[2020]") == [
FieldRange("created", "[", "2020", "", "]"),
]
def test_known_field_name_midword_is_passthrough(self):
# A known field name embedded mid-word is not a field token (the
# word-boundary guard); the whole run stays passthrough.
assert scan("xtag:foo") == [Passthrough("xtag:foo")]
@pytest.mark.search
class TestCommaResolution:
def test_value_list_multi_value_field(self):
toks = resolve_commas(scan("tag:foo,bar"))
assert toks == [FieldValueList("tag", ("foo", "bar"))]
def test_value_list_three(self):
toks = resolve_commas(scan("tag_id:1,2,3"))
assert toks == [FieldValueList("tag_id", ("1", "2", "3"))]
def test_text_field_comma_is_literal(self):
# correspondent is not multi-value: comma stays inside the value.
toks = resolve_commas(scan("correspondent:foo,bar"))
assert toks == [FieldValue("correspondent", "foo,bar")]
def test_clause_separator_before_known_field(self):
toks = resolve_commas(scan("tag:foo,type:bar"))
assert toks == [FieldValue("tag", "foo"), Comma(), FieldValue("type", "bar")]
def test_clause_separator_after_range(self):
toks = resolve_commas(scan("created:[2020 TO 2021],added:[2022 TO 2023]"))
assert toks == [
FieldRange("created", "[", "2020", "2021", "]"),
Comma(),
FieldRange("added", "[", "2022", "2023", "]"),
]
def test_clause_separator_after_quote(self):
toks = resolve_commas(scan('correspondent:"A B",created:[2020 TO 2021]'))
assert toks == [
FieldValue("correspondent", '"A B"'),
Comma(),
FieldRange("created", "[", "2020", "2021", "]"),
]
def test_url_comma_is_literal_passthrough(self):
toks = resolve_commas(scan("http://example.com/a,b"))
assert toks == [Passthrough("http://example.com/a,b")]
def test_non_multi_value_comma_is_literal(self):
# title is not in MULTI_VALUE_FIELDS: comma stays inside the value.
toks = resolve_commas(scan("title:10,20"))
assert toks == [FieldValue("title", "10,20")]
def test_clause_separator_before_known_date_field(self):
# The comma between a bare value and a known date field acts as a
# clause separator; both sides survive as distinct tokens.
toks = resolve_commas(scan("correspondent:foo,created:[2020 TO 2021]"))
assert toks == [
FieldValue("correspondent", "foo"),
Comma(),
FieldRange("created", "[", "2020", "2021", "]"),
]
@pytest.mark.search
class TestTranslateScalar:
@pytest.mark.parametrize(
("field", "value", "expected"),
[
(
"created",
"2020",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}",
),
(
"created",
"202003",
"created:[2020-03-01T00:00:00Z TO 2020-04-01T00:00:00Z}",
),
(
"created",
"20200115",
"created:[2020-01-15T00:00:00Z TO 2020-01-16T00:00:00Z}",
),
(
"created",
"2020-01-15",
"created:[2020-01-15T00:00:00Z TO 2020-01-16T00:00:00Z}",
),
(
"created",
"2020-03",
"created:[2020-03-01T00:00:00Z TO 2020-04-01T00:00:00Z}",
),
],
)
def test_partial_and_iso_dates(self, field: str, value: str, expected: str) -> None:
assert translate_scalar(field, value, UTC) == expected
def test_invalid_date_raises(self) -> None:
with pytest.raises(InvalidDateQuery) as exc_info:
translate_scalar("created", "202023", UTC)
assert exc_info.value.field == "created"
assert exc_info.value.value == "202023"
def test_keyword_delegates(self) -> None:
# keyword path produces a half-open range; just assert it is a created range
out = translate_scalar("created", "today", UTC)
assert out.startswith("created:[") and out.endswith("}")
def test_14digit_compact_datetime(self) -> None:
out = translate_scalar("created", "20240115120000", UTC)
assert "20240115120000" not in out
assert out.startswith("created:")
assert out == "created:[2024-01-15T12:00:00Z TO 2024-01-15T12:00:00Z]"
def test_14digit_invalid_month_raises(self) -> None:
with pytest.raises(InvalidDateQuery) as exc_info:
translate_scalar("created", "20231300120000", UTC)
assert exc_info.value.field == "created"
assert exc_info.value.value == "20231300120000"
def test_unrecognized_value_raises(self) -> None:
# A value that is not a keyword, digits, ISO date, or compact timestamp
# raises rather than producing invalid Tantivy syntax or silently matching
# nothing.
with pytest.raises(InvalidDateQuery) as exc_info:
translate_scalar("created", "garbage", UTC)
assert exc_info.value.field == "created"
assert exc_info.value.value == "garbage"
@pytest.mark.search
class TestTranslateRange:
@pytest.mark.parametrize(
("lo", "hi", "expected"),
[
("2005", "2009", "created:[2005-01-01T00:00:00Z TO 2010-01-01T00:00:00Z}"),
(
"202001",
"202006",
"created:[2020-01-01T00:00:00Z TO 2020-07-01T00:00:00Z}",
),
(
"20200101",
"20201231",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}",
),
(
"2020-01-01",
"2020-12-31",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}",
),
],
)
def test_absolute_ranges(self, lo, hi, expected):
assert translate_range("created", lo, hi, UTC) == expected
def test_reversed_swaps(self):
assert translate_range("created", "2009", "2005", UTC) == (
"created:[2005-01-01T00:00:00Z TO 2010-01-01T00:00:00Z}"
)
def test_open_upper(self):
out = translate_range("created", "2020", "", UTC)
assert out == f"created:[2020-01-01T00:00:00Z TO {OPEN_HI}]"
def test_open_lower(self):
out = translate_range("created", "", "2020", UTC)
assert out == f"created:[{OPEN_LO} TO 2021-01-01T00:00:00Z}}"
def test_invalid_bound_raises(self):
with pytest.raises(InvalidDateQuery) as exc_info:
translate_range("created", "202023", "2025", UTC)
assert exc_info.value.field == "created"
assert exc_info.value.value == "202023"
def test_invalid_high_bound_raises(self):
# Low bound parses, high bound does not -> raise on the high bound.
with pytest.raises(InvalidDateQuery) as exc_info:
translate_range("created", "2020", "garbage", UTC)
assert exc_info.value.field == "created"
assert exc_info.value.value == "garbage"
@pytest.mark.search
class TestTranslateQuery:
@pytest.mark.parametrize(
("raw", "expected"),
[
(
"created:2020",
"created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z}",
),
("tag:foo,bar", "tag:foo AND tag:bar"),
# 'type' is a user-facing alias rewritten to 'document_type' (the real schema field)
("tag:foo,type:bar", "tag:foo AND document_type:bar"),
(
"created:[2020 TO 2021],added:[2022 TO 2023]",
(
"created:[2020-01-01T00:00:00Z TO 2022-01-01T00:00:00Z}"
" AND "
"added:[2022-01-01T00:00:00Z TO 2024-01-01T00:00:00Z}"
),
),
# correspondent is not multi-value: comma stays literal inside the value
("correspondent:foo,bar", "correspondent:foo,bar"),
],
)
def test_golden(self, raw: str, expected: str) -> None:
assert translate_query(raw, UTC) == expected
@pytest.mark.parametrize(
"raw",
[
"created:2020",
"created:202003",
"created:[20200101 TO 20201231]",
"created:[2020-01-01 TO 2020-12-31]",
"created:[2020 to]",
"created:[to 2020]",
"title:x,created:[2020 TO 2021]",
"created:2020 OR foo",
"(created:2020 OR invoice)",
"tag:foo,type:bar",
"bank statement",
],
)
def test_parse_acceptance(self, index: tantivy.Index, raw: str) -> None:
translated = translate_query(raw, UTC)
# Must not raise:
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
@pytest.mark.search
class TestFieldAliasing:
"""Whoosh->Tantivy field-name aliasing (type/path -> document_type/storage_path)."""
def test_type_alias(self) -> None:
assert translate_query("type:invoice", UTC) == "document_type:invoice"
def test_path_alias(self) -> None:
assert translate_query("path:/foo/bar", UTC) == "storage_path:/foo/bar"
def test_type_id_alias(self) -> None:
assert translate_query("type_id:5", UTC) == "document_type_id:5"
def test_path_id_alias(self) -> None:
assert translate_query("path_id:7", UTC) == "storage_path_id:7"
def test_clause_separator_plus_alias(self) -> None:
# Comma between known fields acts as AND separator; alias still applied.
assert (
translate_query("tag:foo,type:bar", UTC) == "tag:foo AND document_type:bar"
)
def test_type_range_alias(self) -> None:
# type is not a date field; range passes through verbatim with alias applied.
assert (
translate_query("type:[2020 TO 2021]", UTC)
== "document_type:[2020 TO 2021]"
)
def test_parse_acceptance_type(self, index: tantivy.Index) -> None:
# Translated output must be accepted by the real Tantivy parser.
translated = translate_query("type:invoice", UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
def test_parse_acceptance_path(self, index: tantivy.Index) -> None:
translated = translate_query("path:foo", UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
# Freeze time so relative-date tests are deterministic.
_FROZEN_NOW = datetime(2026, 3, 28, 12, 0, 0, tzinfo=UTC)
@pytest.mark.search
class TestRelativeRanges:
"""Relative date-range tokens resolved against a frozen clock."""
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_minus_7_days_to_now(self) -> None:
assert translate_query("added:[-7 days to now]", UTC) == (
"added:[2026-03-21T12:00:00Z TO 2026-03-28T12:00:00Z]"
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_minus_1_week_to_now(self) -> None:
assert translate_query("added:[-1 week to now]", UTC) == (
"added:[2026-03-21T12:00:00Z TO 2026-03-28T12:00:00Z]"
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_minus_1_month_to_now(self) -> None:
assert translate_query("created:[-1 month to now]", UTC) == (
"created:[2026-02-28T12:00:00Z TO 2026-03-28T12:00:00Z]"
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_minus_1_year_to_now(self) -> None:
assert translate_query("modified:[-1 year to now]", UTC) == (
"modified:[2025-03-28T12:00:00Z TO 2026-03-28T12:00:00Z]"
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_minus_3_hours_to_now(self) -> None:
assert translate_query("added:[-3 hours to now]", UTC) == (
"added:[2026-03-28T09:00:00Z TO 2026-03-28T12:00:00Z]"
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_uppercase_units(self) -> None:
assert translate_query("added:[-1 WEEK TO NOW]", UTC) == (
"added:[2026-03-21T12:00:00Z TO 2026-03-28T12:00:00Z]"
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_now_minus_7d_compact(self) -> None:
assert translate_query("added:[now-7d TO now]", UTC) == (
"added:[2026-03-21T12:00:00Z TO 2026-03-28T12:00:00Z]"
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_reversed_range_swapped(self) -> None:
# now+1h TO now-1h is reversed; translate_range swaps -> lo=now-1h, hi=now+1h
assert translate_query("added:[now+1h TO now-1h]", UTC) == (
"added:[2026-03-28T11:00:00Z TO 2026-03-28T13:00:00Z]"
)
@pytest.mark.parametrize(
"raw",
[
"added:[-7 days to now]",
"added:[-1 week to now]",
"created:[-1 month to now]",
"modified:[-1 year to now]",
"added:[-3 hours to now]",
"added:[now-7d TO now]",
"added:[now+1h TO now-1h]",
],
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_parse_acceptance(self, index: tantivy.Index, raw: str) -> None:
translated = translate_query(raw, UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
@pytest.mark.search
class TestWhooshUnitAbbreviations:
"""
Whoosh's PlusMinus date grammar accepted abbreviated unit spellings
(e.g. "yrs", "mos", "wks", "hrs", "mins", "secs"); saved views/searches
created under the old Whoosh backend can contain those tokens (see
https://github.com/paperless-ngx/paperless-ngx/issues/13482), so the
Tantivy translator must still accept them.
"""
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_minus_999_yrs(self) -> None:
assert translate_query("created:[-999yrs to now]", UTC) == (
"created:[1027-03-28T12:00:00Z TO 2026-03-28T12:00:00Z]"
)
@pytest.mark.parametrize(
("token", "expected_lo"),
[
("-1y", "2025-03-28T12:00:00Z"),
("-1yr", "2025-03-28T12:00:00Z"),
("-3mos", "2025-12-28T12:00:00Z"),
("-3mo", "2025-12-28T12:00:00Z"),
("-2wks", "2026-03-14T12:00:00Z"),
("-2wk", "2026-03-14T12:00:00Z"),
("-5dys", "2026-03-23T12:00:00Z"),
("-5dy", "2026-03-23T12:00:00Z"),
("-1hrs", "2026-03-28T11:00:00Z"),
("-1hr", "2026-03-28T11:00:00Z"),
("-10mins", "2026-03-28T11:50:00Z"),
("-10min", "2026-03-28T11:50:00Z"),
("-30secs", "2026-03-28T11:59:30Z"),
("-30sec", "2026-03-28T11:59:30Z"),
],
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_abbreviated_units(self, token: str, expected_lo: str) -> None:
assert translate_query(f"added:[{token} to now]", UTC) == (
f"added:[{expected_lo} TO 2026-03-28T12:00:00Z]"
)
@pytest.mark.parametrize(
"raw",
[
"created:[-999yrs to now]",
"added:[-1y to now]",
"created:[-3mos to now]",
"added:[-2wks to now]",
"added:[-5dys to now]",
"added:[-1hrs to now]",
"added:[-10mins to now]",
"added:[-30secs to now]",
],
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_parse_acceptance(self, index: tantivy.Index, raw: str) -> None:
translated = translate_query(raw, UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
@pytest.mark.search
class TestOperatorNormalization:
"""Post-render operator normalization in translate_query."""
def test_spaced_dash_removed(self) -> None:
assert (
translate_query("H52.1 - Kurzsichtigkeit", UTC) == "H52.1 Kurzsichtigkeit"
)
def test_spaced_dash_simple(self) -> None:
assert translate_query("bar - baz", UTC) == "bar baz"
def test_trailing_operator_stripped(self) -> None:
assert translate_query("foo -", UTC) == "foo"
def test_date_range_preserved(self) -> None:
out = translate_query("created:[2020 TO 2021]", UTC)
# Must not corrupt the ISO range
assert out == "created:[2020-01-01T00:00:00Z TO 2022-01-01T00:00:00Z}"
def test_date_scalar_with_or(self) -> None:
out = translate_query("created:2020 OR foo", UTC)
# The created scalar becomes a range; " OR foo" passes through verbatim.
assert out.startswith("created:[")
assert "OR foo" in out
def test_parse_acceptance_spaced_dash(self, index: tantivy.Index) -> None:
translated = translate_query("H52.1 - Kurzsichtigkeit", UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
def test_parse_acceptance_trailing_op(self, index: tantivy.Index) -> None:
translated = translate_query("foo -", UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
@pytest.mark.search
class TestMultiWordDateKeywords:
"""scan() must consume multi-word date keywords as a single value."""
def test_scan_previous_week_as_single_token(self) -> None:
# "created:previous week" must produce one FieldValue with value "previous week",
# not FieldValue("created","previous") + Passthrough(" week").
toks = scan("created:previous week")
assert toks == [FieldValue("created", "previous week")]
def test_scan_this_month_as_single_token(self) -> None:
toks = scan("added:this month")
assert toks == [FieldValue("added", "this month")]
def test_scan_previous_month_as_single_token(self) -> None:
toks = scan("created:previous month")
assert toks == [FieldValue("created", "previous month")]
def test_scan_this_year_as_single_token(self) -> None:
toks = scan("added:this year")
assert toks == [FieldValue("added", "this year")]
def test_scan_previous_year_as_single_token(self) -> None:
toks = scan("created:previous year")
assert toks == [FieldValue("created", "previous year")]
def test_scan_previous_quarter_as_single_token(self) -> None:
toks = scan("created:previous quarter")
assert toks == [FieldValue("created", "previous quarter")]
def test_quoted_multi_word_keyword_still_works(self) -> None:
# The quoted form must continue to work as before.
toks = scan('created:"previous week"')
assert toks == [FieldValue("created", '"previous week"')]
def test_non_date_field_not_affected(self) -> None:
# "previous" stops at the space for non-date fields; " week" passes through.
toks = scan("correspondent:previous week")
assert toks == [
FieldValue("correspondent", "previous"),
Passthrough(" week"),
]
@pytest.mark.search
class TestKeywordDateResolution:
"""Relative date keywords resolve to exact ISO ranges against a frozen clock.
Frozen at 2026-03-28 12:00 UTC (a Saturday in Q1) so the week, month,
quarter and year rollovers are all exercised by a single anchor.
"""
# created is a DateField: bounds are UTC midnight, no timezone offset.
@pytest.mark.parametrize(
("keyword", "expected"),
[
pytest.param(
"today",
"created:[2026-03-28T00:00:00Z TO 2026-03-29T00:00:00Z}",
id="today",
),
pytest.param(
"yesterday",
"created:[2026-03-27T00:00:00Z TO 2026-03-28T00:00:00Z}",
id="yesterday",
),
pytest.param(
"previous week",
"created:[2026-03-16T00:00:00Z TO 2026-03-23T00:00:00Z}",
id="previous-week",
),
pytest.param(
"this month",
"created:[2026-03-01T00:00:00Z TO 2026-04-01T00:00:00Z}",
id="this-month",
),
pytest.param(
"previous month",
"created:[2026-02-01T00:00:00Z TO 2026-03-01T00:00:00Z}",
id="previous-month",
),
pytest.param(
"this year",
"created:[2026-01-01T00:00:00Z TO 2027-01-01T00:00:00Z}",
id="this-year",
),
pytest.param(
"previous year",
"created:[2025-01-01T00:00:00Z TO 2026-01-01T00:00:00Z}",
id="previous-year",
),
pytest.param(
"previous quarter",
"created:[2025-10-01T00:00:00Z TO 2026-01-01T00:00:00Z}",
id="previous-quarter",
),
],
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_date_only_field_keyword_ranges(
self,
keyword: str,
expected: str,
) -> None:
assert translate_query(f"created:{keyword}", UTC) == expected
# added is a DateTimeField: local-tz midnight converted to UTC. Tokyo
# (+09:00, no DST) shifts each midnight boundary back to 15:00Z the day
# before, so this also exercises the local-midnight offset path.
@pytest.mark.parametrize(
("keyword", "expected"),
[
pytest.param(
"today",
"added:[2026-03-27T15:00:00Z TO 2026-03-28T15:00:00Z}",
id="today",
),
pytest.param(
"yesterday",
"added:[2026-03-26T15:00:00Z TO 2026-03-27T15:00:00Z}",
id="yesterday",
),
pytest.param(
"previous week",
"added:[2026-03-15T15:00:00Z TO 2026-03-22T15:00:00Z}",
id="previous-week",
),
pytest.param(
"this month",
"added:[2026-02-28T15:00:00Z TO 2026-03-31T15:00:00Z}",
id="this-month",
),
pytest.param(
"previous month",
"added:[2026-01-31T15:00:00Z TO 2026-02-28T15:00:00Z}",
id="previous-month",
),
pytest.param(
"this year",
"added:[2025-12-31T15:00:00Z TO 2026-12-31T15:00:00Z}",
id="this-year",
),
pytest.param(
"previous year",
"added:[2024-12-31T15:00:00Z TO 2025-12-31T15:00:00Z}",
id="previous-year",
),
pytest.param(
"previous quarter",
"added:[2025-09-30T15:00:00Z TO 2025-12-31T15:00:00Z}",
id="previous-quarter",
),
],
)
@time_machine.travel(_FROZEN_NOW, tick=False)
def test_datetime_field_keyword_ranges_local_tz(
self,
keyword: str,
expected: str,
) -> None:
assert translate_query(f"added:{keyword}", ZoneInfo("Asia/Tokyo")) == expected
@pytest.mark.search
class TestISODatetimeBounds:
"""Full ISO datetime tokens in range bounds must be parsed directly."""
def test_translate_range_iso_bounds_passthrough(self) -> None:
# Already-ISO datetime bounds must pass through as-is (exact instant).
result = translate_range(
"created",
"2020-01-01T00:00:00Z",
"2021-01-01T00:00:00Z",
UTC,
)
assert result == "created:[2020-01-01T00:00:00Z TO 2021-01-01T00:00:00Z]"
def test_translate_query_iso_range_preserved(self) -> None:
q = "created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
assert translate_query(q, UTC) == q
def test_translate_query_comma_separated_iso_ranges(self) -> None:
q = (
"created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z],"
"added:[2026-05-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
)
result = translate_query(q, UTC)
assert result == (
"created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
" AND "
"added:[2026-05-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
)
def test_translate_query_text_before_comma_separated_date_clause(self) -> None:
result = translate_query("schäfersee,created:previous year", UTC)
assert result == (
"schäfersee AND created:[2025-01-01T00:00:00Z TO 2026-01-01T00:00:00Z}"
)
def test_invalid_iso_datetime_raises(self) -> None:
# A token with "T" that is not valid ISO datetime -> raise.
with pytest.raises(InvalidDateQuery) as exc_info:
translate_range(
"created",
"2020-01-01T99:00:00Z",
"2021-01-01T00:00:00Z",
UTC,
)
assert exc_info.value.field == "created"
assert exc_info.value.value == "2020-01-01T99:00:00Z"
def test_parse_acceptance_iso_bounds(self, index: tantivy.Index) -> None:
q = "created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
translated = translate_query(q, UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
def test_parse_acceptance_comma_iso_ranges(self, index: tantivy.Index) -> None:
q = (
"created:[2026-01-01T00:00:00Z TO 2026-06-01T00:00:00Z],"
"added:[2026-05-01T00:00:00Z TO 2026-06-01T00:00:00Z]"
)
translated = translate_query(q, UTC)
index.parse_query(translated, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS)
@@ -339,3 +339,21 @@ class TestBulkDownload(DirectoriesMixin, SampleDirMixin, APITestCase):
self.assertEqual(response.status_code, status.HTTP_403_FORBIDDEN)
self.assertEqual(response.content, b"Insufficient permissions")
def test_bad_search_query_returns_400(self) -> None:
response = self.client.post(
self.ENDPOINT,
json.dumps(
{
"all": True,
"filters": {"query": "added:notadate"},
"content": "originals",
},
),
content_type="application/json",
)
# A user-fixable query error must surface as a 400 naming the bad
# value, exactly like the search list endpoint, never a 500.
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
self.assertIn(b"notadate", response.content)
+19
View File
@@ -1976,3 +1976,22 @@ class TestBulkEditAPI(DirectoriesMixin, APITestCase):
self.assertEqual(response.status_code, status.HTTP_200_OK)
self.assertEqual(LogEntry.objects.filter(object_pk=self.doc1.id).count(), 2)
def test_api_bulk_edit_with_bad_search_query_returns_400(self) -> None:
response = self.client.post(
"/api/documents/bulk_edit/",
json.dumps(
{
"all": True,
"filters": {"query": "added:notadate"},
"method": "set_storage_path",
"parameters": {"storage_path": self.sp1.id},
},
),
content_type="application/json",
)
# A user-fixable query error must surface as a 400 naming the bad
# value, exactly like the search list endpoint, never a 500.
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
self.assertIn(b"notadate", response.content)
+89
View File
@@ -756,6 +756,10 @@ class TestDocumentSearchApi(DirectoriesMixin, APITestCase):
tick=False,
):
response = self.client.get("/api/documents/?query=added:previous month")
assert response.status_code == 200, (
f"expected a successful search response, got {response.status_code}: "
f"{response.data!r}"
)
results = response.data["results"]
self.assertEqual(len(results), 1)
@@ -788,6 +792,26 @@ class TestDocumentSearchApi(DirectoriesMixin, APITestCase):
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
self.assertIn("invalid-date", str(response.data["query"]))
def test_search_multiple_bad_fields_returns_all_messages(self) -> None:
"""
GIVEN:
- One document added
WHEN:
- Query with multiple bad fields (e.g. invalid date and invalid number)
THEN:
- 400 Bad Request with error messages for every bad field,
so the user can fix them all in one round-trip
"""
response = self.client.get(
"/api/documents/",
{"query": "created:notadate AND asn:notanumber"},
)
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
messages = response.data["query"]
self.assertEqual(len(messages), 2)
self.assertTrue(any("created" in m for m in messages))
self.assertTrue(any("asn" in m for m in messages))
@override_settings(
TIME_ZONE="UTC",
)
@@ -831,6 +855,29 @@ class TestDocumentSearchApi(DirectoriesMixin, APITestCase):
results = response.data["results"]
self.assertEqual({r["id"] for r in results}, {1, 2})
@mock.patch("documents.search._backend.parse_user_query")
def test_search_parser_bug_surfaces_as_500_not_400(self, m) -> None:
"""
GIVEN:
- The query parser itself fails (a whoosh-compat bug, per
QueryParserError's own contract: not user-fixable input)
WHEN:
- Any search request runs
THEN:
- The error surfaces as a 500 monitoring can see, never a 400
blaming the user for a library defect
"""
from whoosh_compat.errors import QueryParserError
m.side_effect = QueryParserError("synthetic parser bug")
self.client.raise_request_exception = False
response = self.client.get("/api/documents/?query=anything")
self.assertEqual(
response.status_code,
status.HTTP_500_INTERNAL_SERVER_ERROR,
)
@mock.patch("documents.search._backend.TantivyBackend.autocomplete")
def test_search_autocomplete_limits(self, m) -> None:
"""
@@ -2005,3 +2052,45 @@ class TestDocumentSearchApi(DirectoriesMixin, APITestCase):
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
response = self.client.get("/api/search/?query=no")
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
def _assert_query_finds(self, doc: Document, query: str) -> None:
get_backend().add_or_update(doc)
response = self.client.get("/api/documents/", {"query": query})
self.assertEqual(response.status_code, status.HTTP_200_OK)
ids = [r["id"] for r in response.data["results"]]
self.assertIn(doc.id, ids)
def test_search_by_asn(self) -> None:
doc = Document.objects.create(
title="Has ASN",
content="content",
checksum="asn-checksum",
archive_serial_number=555,
)
self._assert_query_finds(doc, "asn:555")
def test_search_by_page_count(self) -> None:
doc = Document.objects.create(
title="Multi-page",
content="content",
checksum="page-count-checksum",
page_count=42,
)
self._assert_query_finds(doc, "page_count:42")
def test_search_by_original_filename(self) -> None:
doc = Document.objects.create(
title="Named file",
content="content",
checksum="filename-checksum",
original_filename="quarterly-report.pdf",
)
self._assert_query_finds(doc, "original_filename:quarterly-report.pdf")
def test_search_by_checksum(self) -> None:
doc = Document.objects.create(
title="Checksum doc",
content="content",
checksum="deadbeef1234",
)
self._assert_query_finds(doc, "checksum:deadbeef1234")
@@ -0,0 +1,275 @@
"""The search list endpoint's exception handling: what becomes a 400 and
what a library defect surfaces as instead.
Companion to documents/tests/search/test_error_routing.py, which pins the
Cause -> SearchQueryError/QueryError routing inside documents/search/_query.py.
These tests pin the layer above it: DocumentViewSet.list's own except clauses,
which decide what an already-routed error becomes on the wire.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from rest_framework import status
from whoosh_compat.errors import Cause
from whoosh_compat.errors import Diagnostic
from whoosh_compat.errors import DiagnosticKind
from whoosh_compat.errors import QueryError
from documents.search import SearchQueryError
from documents.tests.factories import DocumentFactory
if TYPE_CHECKING:
from rest_framework.test import APIClient
from documents.models import Document
pytestmark = [pytest.mark.django_db, pytest.mark.usefixtures("_search_index")]
@pytest.fixture
def indexed_document() -> Document:
from documents.search import get_backend
doc = DocumentFactory.create(title="quarterly invoice", content="acme corp")
get_backend().add_or_update(doc)
return doc
class TestSearchQueryErrorStillBecomesA400:
def test_search_query_error_becomes_a_400_naming_the_field(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
import documents.search._backend as backend_mod
def raise_search_query_error(*args: object, **kwargs: object) -> object:
raise SearchQueryError("bad value for field 'added'")
monkeypatch.setattr(
backend_mod,
"parse_user_query",
raise_search_query_error,
)
response = admin_client.get("/api/documents/?query=anything")
assert response.status_code == status.HTTP_400_BAD_REQUEST
assert "added" in str(response.data["query"])
class TestLibraryDefectsPropagate:
"""The exact regression this task exists to fix: an unexpected or
INTERNAL-cause library error must not be relabeled a 400."""
def test_unexpected_exception_is_not_converted_to_a_400(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
import documents.search._backend as backend_mod
def raise_zero_division(*args: object, **kwargs: object) -> object:
raise ZeroDivisionError("synthetic bug, unrelated to search grammar")
monkeypatch.setattr(
backend_mod,
"parse_user_query",
raise_zero_division,
)
with pytest.raises(ZeroDivisionError):
admin_client.get("/api/documents/?query=anything")
def test_internal_cause_query_error_is_not_converted_to_a_400(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
"""Forces the one library-internal failure mode reachable from a real
query: emit() reporting a defect in itself (Cause.INTERNAL) after a
real query string went through the real parse and routing pipeline.
``tantivy_emit`` (the whoosh-compat emitter) is monkeypatched rather
than ``parse_user_query`` itself, so everything upstream of it --
the pre-parse rewrites, ``wc.parse()``, and ``_map_emit_error``'s own
Cause routing in documents/search/_query.py -- runs for real; only
the final emit call is forced to report the defect.
"""
import documents.search._query as query_mod
def raise_internal(*args: object, **kwargs: object) -> object:
raise QueryError(
Diagnostic(
kind=DiagnosticKind.BACKEND_REJECTED,
cause=Cause.INTERNAL,
message="synthetic whoosh-compat emitter defect",
),
)
monkeypatch.setattr(query_mod, "tantivy_emit", raise_internal)
with pytest.raises(QueryError):
admin_client.get("/api/documents/?query=invoice")
class TestSelectionPathsAgreeWithSearch:
"""DocumentSelectionMixin backs bulk edit, bulk download, and a
more_like_id selection filter. It catches only SearchQueryError -- the
same contract the search list endpoint enforces above -- so all three
must map SearchQueryError to a 400 and let anything else surface."""
def test_bulk_edit_maps_search_query_error_to_a_400(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
import documents.search._backend as backend_mod
def raise_search_query_error(*args: object, **kwargs: object) -> object:
raise SearchQueryError("bad value for field 'added'")
monkeypatch.setattr(
backend_mod,
"parse_user_query",
raise_search_query_error,
)
response = admin_client.post(
"/api/documents/bulk_edit/",
{
"documents": [],
"all": True,
"filters": {"query": "anything"},
"method": "set_document_type",
"parameters": {"document_type": None},
},
format="json",
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
assert "added" in str(response.data["query"])
def test_bulk_edit_lets_an_unexpected_exception_surface(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
import documents.search._backend as backend_mod
def raise_zero_division(*args: object, **kwargs: object) -> object:
raise ZeroDivisionError("synthetic bug, unrelated to search grammar")
monkeypatch.setattr(
backend_mod,
"parse_user_query",
raise_zero_division,
)
with pytest.raises(ZeroDivisionError):
admin_client.post(
"/api/documents/bulk_edit/",
{
"documents": [],
"all": True,
"filters": {"query": "anything"},
"method": "set_document_type",
"parameters": {"document_type": None},
},
format="json",
)
def test_bulk_download_maps_search_query_error_to_a_400(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
import documents.search._backend as backend_mod
def raise_search_query_error(*args: object, **kwargs: object) -> object:
raise SearchQueryError("bad value for field 'added'")
monkeypatch.setattr(
backend_mod,
"parse_user_query",
raise_search_query_error,
)
response = admin_client.post(
"/api/documents/bulk_download/",
{
"documents": [],
"all": True,
"filters": {"query": "anything"},
},
format="json",
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
assert "added" in str(response.data["query"])
def test_more_like_id_selection_filter_maps_search_query_error_to_a_400(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
import documents.search._backend as backend_mod
def raise_search_query_error(*args: object, **kwargs: object) -> object:
raise SearchQueryError("similar-document lookup is unavailable")
monkeypatch.setattr(
backend_mod.TantivyBackend,
"more_like_this_ids",
raise_search_query_error,
)
response = admin_client.post(
"/api/documents/bulk_download/",
{
"documents": [],
"all": True,
"filters": {"more_like_id": indexed_document.pk},
},
format="json",
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
def test_more_like_id_selection_filter_lets_an_unexpected_exception_surface(
self,
admin_client: APIClient,
monkeypatch: pytest.MonkeyPatch,
indexed_document: Document,
) -> None:
import documents.search._backend as backend_mod
def raise_zero_division(*args: object, **kwargs: object) -> object:
raise ZeroDivisionError("synthetic bug, unrelated to similarity lookup")
monkeypatch.setattr(
backend_mod.TantivyBackend,
"more_like_this_ids",
raise_zero_division,
)
with pytest.raises(ZeroDivisionError):
admin_client.post(
"/api/documents/bulk_download/",
{
"documents": [],
"all": True,
"filters": {"more_like_id": indexed_document.pk},
},
format="json",
)
@@ -0,0 +1,185 @@
"""The query-length cap in ``_get_tantivy_query_and_mode`` (F3).
whoosh-compat's fieldname tagger is O(n^2) in plain word characters, so an
unbounded ``query`` (SearchMode.QUERY) string is a CPU-exhaustion vector
against a single request handler. The GET search endpoint is incidentally
bounded by the web server's header limit, but the POST selection-filter
path (bulk edit, bulk download) is not -- that is the real vector, so it
must be pinned here too, not just the GET path.
The cap is enforced once, in the shared helper both entry points call, so
these tests exercise the real endpoints rather than the helper directly:
a construct that looks right in isolation has repeatedly behaved
differently end to end on this branch.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from rest_framework import status
from documents.tests.factories import DocumentFactory
from documents.views import _MAX_QUERY_LENGTH
if TYPE_CHECKING:
from rest_framework.test import APIClient
from documents.models import Document
pytestmark = [pytest.mark.django_db, pytest.mark.usefixtures("_search_index")]
@pytest.fixture
def indexed_document() -> Document:
from documents.search import get_backend
doc = DocumentFactory.create(title="quarterly invoice", content="acme corp")
get_backend().add_or_update(doc)
return doc
class TestGetSearchEndpointEnforcesTheCap:
def test_query_one_over_the_cap_is_a_400(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
query = "a" * (_MAX_QUERY_LENGTH + 1)
response = admin_client.get("/api/documents/", {"query": query})
assert response.status_code == status.HTTP_400_BAD_REQUEST
message = str(response.data["query"])
assert str(_MAX_QUERY_LENGTH) in message
assert str(_MAX_QUERY_LENGTH + 1) in message
def test_query_at_exactly_the_cap_is_accepted(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
query = "a" * _MAX_QUERY_LENGTH
response = admin_client.get("/api/documents/", {"query": query})
assert response.status_code == status.HTTP_200_OK
def test_an_ordinary_query_is_unaffected(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
response = admin_client.get("/api/documents/", {"query": "invoice"})
assert response.status_code == status.HTTP_200_OK
assert response.data["count"] == 1
class TestPostSelectionPathsEnforceTheCap:
"""The bulk-edit and bulk-download selection filters share the same
helper the GET search path uses. This is the path that actually
matters: it is not bounded by a web server's header-length limit the
way the GET path incidentally is."""
def test_bulk_edit_query_one_over_the_cap_is_a_400(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
query = "a" * (_MAX_QUERY_LENGTH + 1)
response = admin_client.post(
"/api/documents/bulk_edit/",
{
"documents": [],
"all": True,
"filters": {"query": query},
"method": "set_document_type",
"parameters": {"document_type": None},
},
format="json",
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
message = str(response.data["query"])
assert str(_MAX_QUERY_LENGTH) in message
assert str(_MAX_QUERY_LENGTH + 1) in message
def test_bulk_edit_query_at_exactly_the_cap_is_accepted(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
query = "a" * _MAX_QUERY_LENGTH
response = admin_client.post(
"/api/documents/bulk_edit/",
{
"documents": [],
"all": True,
"filters": {"query": query},
"method": "set_document_type",
"parameters": {"document_type": None},
},
format="json",
)
assert response.status_code == status.HTTP_200_OK
def test_bulk_download_query_one_over_the_cap_is_a_400(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
query = "a" * (_MAX_QUERY_LENGTH + 1)
response = admin_client.post(
"/api/documents/bulk_download/",
{
"documents": [],
"all": True,
"filters": {"query": query},
},
format="json",
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
message = str(response.data["query"])
assert str(_MAX_QUERY_LENGTH) in message
assert str(_MAX_QUERY_LENGTH + 1) in message
class TestGlobalSearchEnforcesTheCapToo:
"""GlobalSearchView calls the backend directly, not through the shared helper.
It hardcodes SearchMode.TEXT, which is linear rather than quadratic, so it
was never the CPU-exhaustion vector. It is capped anyway so that "every
user query string reaching the backend passes a length check" is an
invariant rather than a claim with an exception: the view already bounds
the query from below, and a later change letting it select a mode would
otherwise reopen the hole silently.
"""
def test_query_one_over_the_cap_is_a_400(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
response = admin_client.get(
"/api/search/",
{"query": "a" * (_MAX_QUERY_LENGTH + 1)},
)
assert response.status_code == status.HTTP_400_BAD_REQUEST
def test_query_at_exactly_the_cap_is_accepted(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
response = admin_client.get(
"/api/search/",
{"query": "a" * _MAX_QUERY_LENGTH},
)
assert response.status_code == status.HTTP_200_OK
@@ -0,0 +1,67 @@
"""An unterminated ``[`` date range bracket at the API level.
``created:[2020`` (with or without a dangling ``to <value>``) now raises
BAD_DATE and the search endpoint returns HTTP 400, where it used to parse
past the missing ``]`` and silently pass the malformed range through.
A 400 is correct: malformed input should fail loudly rather than silently
matching an unintended query. Pinned at the API level -- the layer a user
or client actually sees -- rather than only against the parser directly.
The properly closed decoy proves the bracket is what matters, not
whoosh-compat's date grammar generally: ``created:[2020 to 2021]`` parses
and searches cleanly.
"""
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from rest_framework import status
from documents.tests.factories import DocumentFactory
if TYPE_CHECKING:
from rest_framework.test import APIClient
from documents.models import Document
pytestmark = [pytest.mark.django_db, pytest.mark.usefixtures("_search_index")]
@pytest.fixture
def indexed_document() -> Document:
from documents.search import get_backend
doc = DocumentFactory.create(title="quarterly invoice", content="acme corp")
get_backend().add_or_update(doc)
return doc
class TestUnterminatedBracketReturnsA400:
@pytest.mark.parametrize(
"query",
[
pytest.param("created:[2020", id="missing_upper_bound_and_bracket"),
pytest.param("created:[2020 to 2021", id="missing_closing_bracket"),
],
)
def test_unterminated_bracket_is_a_400(
self,
admin_client: APIClient,
indexed_document: Document,
query: str,
) -> None:
response = admin_client.get(f"/api/documents/?query={query}")
assert response.status_code == status.HTTP_400_BAD_REQUEST
assert "created" in str(response.data["query"])
def test_properly_closed_bracket_still_searches_cleanly(
self,
admin_client: APIClient,
indexed_document: Document,
) -> None:
response = admin_client.get(
"/api/documents/?query=created:[2020 to 2021]",
)
assert response.status_code == status.HTTP_200_OK
+60 -25
View File
@@ -16,6 +16,7 @@ from time import mktime
from time import sleep
from typing import TYPE_CHECKING
from typing import Any
from typing import Final
from typing import Literal
from typing import NamedTuple
from unicodedata import normalize
@@ -279,17 +280,40 @@ logger = logging.getLogger("paperless.api")
_TANTIVY_INTERSECT_THRESHOLD = 5_000
_TANTIVY_SEARCH_PARAM_NAMES = ("text", "title_search", "query", "more_like_id")
# whoosh-compat's fieldname tagger (used only for SearchMode.QUERY, via the
# whoosh grammar in parse_user_query) is O(n^2) in plain word characters:
# measured at ~0.96s/10k chars, ~3.67s/20k, ~14.4s/40k against the real field
# registry. Django's DATA_UPLOAD_MAX_MEMORY_SIZE default (2.5 MB) does not
# bound this on the POST-body selection-filter path, so an unbounded query
# is a single-request CPU exhaustion vector. 4096 chars caps the worst case
# at roughly 0.16s (quadratic extrapolation from the measurements above),
# far beyond any plausible hand-typed advanced query, while still being fast
# enough to absorb inside a request handler. Applied to all three modes at
# this shared choke point: TEXT and TITLE route through simple_search_tokens
# instead and measure linear even at 20k chars, so the cap is hygiene for
# them, not a fix, but a single limit here is simpler than one exemption.
# Not exposed as a PAPERLESS_* setting: this is a hard security boundary,
# not a tunable, and a raisable ceiling would let a misconfiguration
# reintroduce the exact hazard this exists to close.
_MAX_QUERY_LENGTH: Final[int] = 4096
def _get_tantivy_query_and_mode(params):
from documents.search import QueryTooLongError
from documents.search import SearchMode
if "text" in params:
return str(params["text"]), SearchMode.TEXT
if "title_search" in params:
return str(params["title_search"]), SearchMode.TITLE
if "query" in params:
return str(params["query"]), SearchMode.QUERY
return None # pragma: no cover
raw, mode = str(params["text"]), SearchMode.TEXT
elif "title_search" in params:
raw, mode = str(params["title_search"]), SearchMode.TITLE
elif "query" in params:
raw, mode = str(params["query"]), SearchMode.QUERY
else:
return None # pragma: no cover
if len(raw) > _MAX_QUERY_LENGTH:
raise QueryTooLongError(len(raw), _MAX_QUERY_LENGTH)
return raw, mode
def _get_more_like_id(query_params: dict[str, Any], user: User | None) -> int:
@@ -2421,6 +2445,7 @@ class UnifiedSearchViewSet(DocumentViewSet):
from documents.search import TantivyBackend
from documents.search import TantivyRelevanceList
from documents.search import get_backend
from documents.search import search_query_error_messages
def parse_search_params() -> SearchParams:
"""Extract query string, search mode, and ordering from request."""
@@ -2611,15 +2636,10 @@ class UnifiedSearchViewSet(DocumentViewSet):
except ValidationError:
raise
except SearchQueryError as e:
# User-fixable query error (e.g. an unparsable date): surface the
# specific message so the user can correct it, rather than a generic
# 400 or silently empty results.
raise ValidationError({"query": [str(e)]}) from e
except Exception as e:
logger.warning(f"An error occurred listing search results: {e!s}")
return HttpResponseBadRequest(
"Error listing search results, check logs for more detail.",
)
# User-fixable query error(s) (e.g. unparsable dates/numbers):
# surface every offending field's message, not just the first,
# so the user can fix them all in one round-trip.
raise ValidationError({"query": search_query_error_messages(e)}) from e
@action(detail=False, methods=["GET"], name="Get Next ASN")
def next_asn(self, request, *args, **kwargs):
@@ -2755,23 +2775,34 @@ class DocumentSelectionMixin:
},
)
from documents.search import SearchQueryError
from documents.search import get_backend
from documents.search import search_query_error_messages
filter_name = search_filters[0]
backend = get_backend()
search_user = None if user.is_superuser else user
if filter_name == "more_like_id":
more_like_doc_id = _get_more_like_id(filters, user)
try:
if filter_name == "more_like_id":
more_like_doc_id = _get_more_like_id(filters, user)
search_ids = backend.more_like_this_ids(more_like_doc_id, user=search_user)
else:
query_str, search_mode = _get_tantivy_query_and_mode(filters)
search_ids = backend.search_ids(
query_str,
user=search_user,
search_mode=search_mode,
)
search_ids = backend.more_like_this_ids(
more_like_doc_id,
user=search_user,
)
else:
query_str, search_mode = _get_tantivy_query_and_mode(filters)
search_ids = backend.search_ids(
query_str,
user=search_user,
search_mode=search_mode,
)
except SearchQueryError as e:
# Same user-fixable-query mapping as the search list endpoint:
# a bad date/number in a bulk selection filter is a 400 naming
# the value, never a 500.
raise ValidationError({"query": search_query_error_messages(e)}) from e
return search_ids
@@ -3608,6 +3639,10 @@ class GlobalSearchView(PassUserMixin):
return HttpResponseBadRequest("Query required")
if len(query) < 3:
return HttpResponseBadRequest("Query must be at least 3 characters")
if len(query) > _MAX_QUERY_LENGTH:
return HttpResponseBadRequest(
f"Query must be at most {_MAX_QUERY_LENGTH} characters",
)
db_only = request.query_params.get("db_only", False)
Generated
+40 -4
View File
@@ -4,11 +4,11 @@ requires-python = ">=3.11"
resolution-markers = [
"python_full_version >= '3.15' and sys_platform == 'darwin'",
"python_full_version >= '3.15' and sys_platform == 'linux'",
"python_full_version == '3.14.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
"python_full_version == '3.14.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
"python_full_version >= '3.12' and python_full_version < '3.15' and sys_platform == 'darwin'",
"python_full_version == '3.12.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
"python_full_version == '3.12.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
"python_full_version == '3.14.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
"python_full_version == '3.14.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
"(python_full_version >= '3.12' and python_full_version < '3.15' and platform_machine != 'aarch64' and platform_machine != 'x86_64' and sys_platform == 'linux') or (python_full_version == '3.13.*' and platform_machine == 'aarch64' and sys_platform == 'linux') or (python_full_version == '3.13.*' and platform_machine == 'x86_64' and sys_platform == 'linux')",
"python_full_version < '3.12' and sys_platform == 'darwin'",
"python_full_version < '3.12' and sys_platform == 'linux'",
@@ -2933,6 +2933,7 @@ dependencies = [
{ name = "torch", version = "2.13.0+cpu", source = { registry = "https://download.pytorch.org/whl/cpu" }, marker = "sys_platform == 'linux'" },
{ name = "watchfiles" },
{ name = "whitenoise" },
{ name = "whoosh-compat", extra = ["tantivy"] },
{ name = "zxing-cpp" },
]
@@ -3091,6 +3092,7 @@ requires-dist = [
{ name = "torch", specifier = "~=2.13.0", index = "https://download.pytorch.org/whl/cpu" },
{ name = "watchfiles", specifier = ">=1.2" },
{ name = "whitenoise", specifier = "~=6.11" },
{ name = "whoosh-compat", extras = ["tantivy"], directory = "../whoosh-compat" },
{ name = "zxing-cpp", specifier = "~=3.1.0" },
]
provides-extras = ["mariadb", "postgres", "webserver"]
@@ -5014,10 +5016,10 @@ version = "2.13.0+cpu"
source = { registry = "https://download.pytorch.org/whl/cpu" }
resolution-markers = [
"python_full_version >= '3.15' and sys_platform == 'linux'",
"python_full_version == '3.12.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
"python_full_version == '3.12.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
"python_full_version == '3.14.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
"python_full_version == '3.14.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
"python_full_version == '3.12.*' and platform_machine == 'x86_64' and sys_platform == 'linux'",
"python_full_version == '3.12.*' and platform_machine == 'aarch64' and sys_platform == 'linux'",
"(python_full_version >= '3.12' and python_full_version < '3.15' and platform_machine != 'aarch64' and platform_machine != 'x86_64' and sys_platform == 'linux') or (python_full_version == '3.13.*' and platform_machine == 'aarch64' and sys_platform == 'linux') or (python_full_version == '3.13.*' and platform_machine == 'x86_64' and sys_platform == 'linux')",
"python_full_version < '3.12' and sys_platform == 'linux'",
]
@@ -5652,6 +5654,40 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/db/eb/d5583a11486211f3ebd4b385545ae787f32363d453c19fffd81106c9c138/whitenoise-6.12.0-py3-none-any.whl", hash = "sha256:fc5e8c572e33ebf24795b47b6a7da8da3c00cff2349f5b04c02f28d0cc5a3cc2", size = 20302, upload-time = "2026-02-27T00:05:40.086Z" },
]
[[package]]
name = "whoosh-compat"
version = "0.1.0"
source = { directory = "../whoosh-compat" }
dependencies = [
{ name = "python-dateutil" },
]
[package.optional-dependencies]
tantivy = [
{ name = "tantivy" },
]
[package.metadata]
requires-dist = [
{ name = "python-dateutil", specifier = ">=2.8" },
{ name = "tantivy", marker = "extra == 'tantivy'", specifier = ">=0.26.0" },
]
provides-extras = ["tantivy"]
[package.metadata.requires-dev]
dev = [
{ name = "hypothesis", specifier = ">=6" },
{ name = "mypy" },
{ name = "pyrefly", specifier = ">=1.2.0" },
{ name = "pytest", specifier = ">=8" },
{ name = "pytest-cov" },
{ name = "ruff" },
{ name = "types-python-dateutil" },
{ name = "tzdata", specifier = ">=2026.3" },
{ name = "whoosh", git = "https://github.com/whoosh-community/whoosh?rev=baa4d577fdb34bfcf30547c8c6bf853fffeb7fe0" },
{ name = "whoosh-compat", extras = ["tantivy"] },
]
[[package]]
name = "wrapt"
version = "2.0.1"