* Feature: match fuzzy terms in place inside the parsed query
Fuzzy matching was a separate clause OR'd in above the query: a flat bag
of the query's words, re-parsed through tantivy's own parser, blended
beside the exact clause. Nothing around a term reached it, so a fielded
term fuzzed across every default field, a filter did not constrain it, and
an exclusion had to be hoisted back over the whole blend to stop the
clause re-admitting what the query had just excluded.
Widen each leaf where it sits instead, through emit()'s rewrite_leaf hook,
so fielding, negation, AND, REQUIRE and positive filters constrain the
fuzzy match exactly as they constrain the exact one. Each of a leaf's
words becomes a Fuzzy leaf on the leaf's own field, boosted to 0.1, beside
the leaf and any CJK alternative it already had.
* Hello?
QUERY-mode searches blended a separate bigram clause in at the top of the
query, built from the parsed AST's free-text tokens. Because it sat beside
the exact clause rather than inside the query, nothing around a CJK term
constrained its bigram match: an exclusion that was one OR branch's own
condition could never reach it, so "(東京 AND NOT secret) OR bill" still
returned the secret document.
Widen each CJK leaf where it sits instead, through emit()'s rewrite_leaf
hook, so every AND, NOT, REQUIRE, boost and field restriction around the
leaf applies to its bigram match too. Negated leaves are widened on
purpose, so "NOT X" excludes exactly what "X" matches.
* Feature: parse advanced search with whoosh-compat and delete the hand-written translator
* Don't cover these, they exist for defensive, but there's no other current diagnostic
* A mix of more no cover and tests
* Route SCHEMA_FIELD_MISSING into an error, not an HTTP 400
* feat(search): add whoosh-compat, the shared field table and the field registry
* Sonar being useful actually
* Adds the given/when/then commenting
* Trims tests I don't think cover our logic or code or are redundant
* Coverage
* test(search): add pattern normalizer stem-alternates unit tests
* build: bump whoosh-compat to 0.2.0
* fix(search): resolve index-write permissions and effective content in bulk
Add WriteBatch.add_or_update_ids() and use it in bulk_update_documents
and trash restore, cutting index writes from ~8 queries per document
to a constant handful per batch
* Always these new ones with xdist, try a better condition
* Fix: exclude next-period start from relative date-range filters
Tantivy's [lo TO hi] range is inclusive on both ends, but computed upper
bounds (keyword ranges, YYYY/YYYYMM/YYYYMMDD tokens) represent the start of
the next period. Use half-open [lo TO hi} for those so e.g. "previous month"
no longer matches the 1st of the current month.
* Adds a regression test down to the second check for the hi range
* Tantivy: get permissions by chunks
-40% indexing time compared to previous commit
* Make progress bar process one by one with chunk
-15% indexing time compared to previous commit
* Prefetch FK + iterate over chunk from SQL
Prefetch additional needed data (note user, custom field content)
-20% indexing time compared to previous commit
* Reindex: increase Tantivy heap size from 128 to 512MB
Gains probably vary depending on the machine,
but it seems a sweet spot compatible with low-end hardware.
* Reindex: optimization on permission fetching and autocomplete word set
-10% indexing time compared to previous commit
* Autocomplete analyzer python->rust
Splits words with underscore compared to the python analyzer.
E.g.: "blue_print" -> ["blue", "print"]
It can still be found with the "blue_print" keyword,
as the search string is also split in two words.
-50% indexing time compared to previous commit (indexing is twice faster!)
* Index bigram for CJK content only
Inedxing time slightly longer (~3%),
but since the non-CJK content is not indexed,
bigram searchs will be slightly optimized.
* Fix group-based view_document permissions missing from bulk rebuild
_bulk_get_viewer_ids only queried UserObjectPermission, dropping the
group-permission expansion that get_users_with_perms(with_group_users=True)
performs for the non-batched per-document indexing path. A user who could
only see a document via group membership would lose search access to it
after any full reindex.
Also query GroupObjectPermission and expand group membership to user ids,
matching the existing single-document behavior.
* Yield (document, viewer_ids) pairs from _DocumentViewerStream
Previously _DocumentViewerStream.__iter__ yielded plain Document objects
while the matching viewer ids were exposed through a separate mutable
attribute (viewer_ids_by_pk), overwritten each time the generator crossed
a chunk boundary. rebuild() read that attribute out-of-band per document.
This only worked because the current iter_wrapper (a plain progress-bar
passthrough) happens to consume the stream in strict lock-step with no
lookahead. Any wrapper that buffers, batches, or reorders would silently
pair a document with the wrong chunk's viewer ids. Yield the pair directly
so the association travels with the document regardless of how iter_wrapper
consumes the stream, and drop the now-unneeded viewer_ids_by_pk attribute.
* Add --heap-size-mb CLI arg to document_index reindex
writer_heap_bytes was hardcoded at 512MB with no way to tune it. Expose it
as a manual-rebuild-only CLI arg rather than a settings/env var, per review
feedback, so lower-memory hosts can reduce it without a wider config
surface. Defaults to unset so TantivyBackend.rebuild's own default stays
the single source of truth.
---------
Co-authored-by: stumpylog <797416+stumpylog@users.noreply.github.com>
Implements and tests a retry with backoff + jitter for aquring the index update lock. If we still can't get it, dispatch a celery task to handle it later instead (also with retry)
Signed-off-by: stumpylog <797416+stumpylog@users.noreply.github.com>