mirror of
https://github.com/paperless-ngx/paperless-ngx.git
synced 2026-08-14 14:53:18 +00:00
Compare commits
26
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6fda0460c0 | ||
|
|
b3101d68b1 | ||
|
|
568fc3c9a9 | ||
|
|
898ed02637 | ||
|
|
96f3741475 | ||
|
|
3fab83dff8 | ||
|
|
13376d0c12 | ||
|
|
d801dc7dc4 | ||
|
|
1a76c59238 | ||
|
|
330491eeed | ||
|
|
0a082b677b | ||
|
|
27f9443c93 | ||
|
|
aad5f7fac6 | ||
|
|
3b0a619d38 | ||
|
|
985c92914e | ||
|
|
778d08def0 | ||
|
|
18675da008 | ||
|
|
c63bb46ac4 | ||
|
|
fa6f4a0533 | ||
|
|
154e3f99a9 | ||
|
|
2d6370e734 | ||
|
|
b91c183cae | ||
|
|
0f4f875289 | ||
|
|
99a055cf09 | ||
|
|
cba7729cd5 | ||
|
|
4d9ec9d912 |
@@ -1,66 +0,0 @@
|
|||||||
---
|
|
||||||
name: whoosh-compat-transition
|
|
||||||
description: Use when integrating the whoosh-compat library into paperless-ngx search, replacing src/documents/search/_translate.py or _dates.py, building the search FieldRegistry, or changing user query parsing during the whoosh-to-tantivy transition
|
|
||||||
---
|
|
||||||
|
|
||||||
# whoosh-compat transition
|
|
||||||
|
|
||||||
## Overview
|
|
||||||
|
|
||||||
whoosh-compat (github.com/stumpylog/whoosh-compat; local checkout usually at `../whoosh-compat`) replaces the hand-maintained translation layer (`src/documents/search/_translate.py`, `_dates.py`): it parses user queries with a faithful fork of whoosh's real grammar into a typed AST and emits programmatic tantivy queries. Read its README and ARCHITECTURE.md before wiring anything; its DIVERGENCES.md lists intended behavior differences and is the authority on "is this difference a bug".
|
|
||||||
|
|
||||||
## Decisions already made (do not re-derive)
|
|
||||||
|
|
||||||
- **Queries are user-typed free text.** The advanced search box passes whatever the user types straight to the parser (that is how the issue #13568 queries exist). Do NOT try to infer the supported field surface from frontend code; the frontend only generates a few date filter strings, everything else is typed by users.
|
|
||||||
- **The field surface is a policy decision, not `KNOWN_FIELDS`.** Today's `KNOWN_FIELDS` accepts internal ID fields (`tag_id`, `owner_id`, `viewer_id`, other `*_id`) that are undocumented in `docs/usage.md` and were ruled not user-searchable by the maintainer: exclude them from the `FieldRegistry` (they stay as programmatic permission/filter fields in `build_permission_filter`, which never touches user query text). The registry is built from documented syntax in `docs/usage.md` plus the v2-compat aliases (`type`, `path`, `type_id`-style aliases follow their canonical field's fate). Undocumented-but-working fields (`asn`, `page_count`, `num_notes`, `original_filename`, `checksum`) need an explicit maintainer yes/no; since users type freely, silently dropping one breaks any saved view using it, so a drop must be a visible, documented decision.
|
|
||||||
- **Analyzer seam:** `FieldSpec.analyzer` binds the live registered tantivy analyzer's `.analyze` (the same Rust analyzer used at index time; language-keyed, so rebuild the registry when `SEARCH_LANGUAGE` changes, on the same trigger as `register_tokenizers`). `pattern_normalizer` is `_tokenizer.ascii_fold`: character-level lowercase+fold only, NEVER stemming.
|
|
||||||
- **Diagnostics before emit:** `whoosh_compat.parse()` never raises on bad input. Check `ParseResult.diagnostics` and map to `SearchQueryError`/`InvalidDateQuery` (HTTP 400) BEFORE calling `emit()`; also catch the emitter's `UnsupportedQueryError` into a 400. Never carry forward the legacy raw-string fallback (`except Exception: query_str = raw_query`) into the new path; it masks integration bugs.
|
|
||||||
- **Build typed errors from structured diagnostic data, never by parsing `message`.** Each `Diagnostic` carries `kind`, `startchar`/`endchar`, and `field`/`raw_value`. `message` is human-readable text whose wording can change. For a range that fails on one bound, `raw_value` is the bound that actually failed.
|
|
||||||
- `diagnostic.field` is a **`FieldRef`**, not a string: use `str(diagnostic.field)` for the canonical dotted name (`created`, `notes.user`) or `diagnostic.field.name` for the field alone. Note the name is canonical, so an aliased query (`type:`) reports the field it resolves to (`document_type`), and the diagnostic span covers the offending value rather than the field name, so the text the user typed for the field is not recoverable.
|
|
||||||
- **The registry has one resolver.** `registry.make_ref(raw)` turns a raw field string into a `FieldRef` or `None` for an unknown field, and `registry.resolve(ref)` returns a `ResolvedField | None`, not a bare spec: read `.spec` for the `FieldSpec`, `.json_path` for the subpath (or `None`), `.is_subpath` and `.dotted_name` are convenience properties. There is no `resolve_json()`; a dotted name is interpreted only inside `make_ref`. Write `resolved = registry.resolve(ref)` then `resolved.spec.kind`, not `spec = registry.resolve(ref)` then `spec.kind`.
|
|
||||||
- **`notes` and `custom_fields` are JSON fields** with fixed subpaths (`notes.user`/`notes.note`, `custom_fields.name`/`custom_fields.value`); the registry stays a static, language-keyed singleton, never per-request.
|
|
||||||
- **`emit()`'s signature is `emit(node, *, index, registry)`, with no `schema` parameter.** Do not write a call site passing `schema=`. `emit()` calls the library's own `analyze()` pipeline stage internally (token analysis, multitoken resolution, zero-token drop), so paperless-ngx never needs to call `analyze()` itself.
|
|
||||||
- **A wildcard/prefix pattern on a JSON subpath reports a parse-time diagnostic**, not a silent whole-field query: `custom_fields.value:abc*` reports `DiagnosticKind.UNSUPPORTED_PATTERN` (the same kind used for a wildcard on a numeric or BOOLEAN_EXISTS field) instead of matching against the wrong encoded bytes. Relevant here because `custom_fields.value` is exactly the kind of JSON subpath a user might expect to pattern-match; the error-mapping code needs a case for `UNSUPPORTED_PATTERN`, not just `BAD_DATE`/`BAD_NUMBER`.
|
|
||||||
|
|
||||||
## Mandatory before deleting old code
|
|
||||||
|
|
||||||
- Date-grammar parity audit, line by line: every keyword, relative unit, and abbreviation `_dates.py` and `_translate.py` accept today (including the whoosh-era abbreviations kept for old saved views) must have an accepted form in whoosh-compat's dateparse grammar. Silent keyword loss is the saved-view breakage class behind issue #13568.
|
|
||||||
- Acceptance corpus compared by matched-document-ID sets, not query strings: the #13568 queries verbatim, real saved-view strings, every date keyword, field aliases, comma lists, date and numeric ranges, wildcards with bracket classes, boosts, JSON subpaths.
|
|
||||||
|
|
||||||
## Tests: what goes, what comes
|
|
||||||
|
|
||||||
Removed with their modules (do not port their string-level assertions):
|
|
||||||
|
|
||||||
- `src/documents/tests/search/test_translate.py`: its subject is deleted; string-translation unit cases are whoosh-compat's own responsibility now. Cases that encode real user-visible behavior get reincarnated as result-level acceptance cases, not string assertions.
|
|
||||||
- Date-keyword unit tests tied to `_dates.py` internals: same treatment.
|
|
||||||
- `test_query.py` cases asserting `parse_user_query` internals or intermediate query strings: rewritten against the new pipeline, asserting on matched results.
|
|
||||||
|
|
||||||
Kept: `test_migration_fulltext_query_field_prefixes.py` (data migration, orthogonal), `test_schema.py`, `test_tokenizer.py`, permission-filter and simple-search tests.
|
|
||||||
|
|
||||||
Added:
|
|
||||||
|
|
||||||
- A result-level acceptance module (paperless's analogue of whoosh-compat's `test_acceptance_e2e.py`): the corpus above against a real index built from `build_schema()`, asserting document-ID sets. Use `pytest.param(..., id="...")` for every case.
|
|
||||||
- Registry unit tests: internal `*_id` names rejected, aliases resolve to canonical fields, JSON subpaths match `docs/usage.md`, construction deterministic per language.
|
|
||||||
- One `Multitoken` case nested inside a top-level `OR` (whoosh-compat DIVERGENCES entry on Multitoken.DEFAULT) to prove it does not matter for paperless's data.
|
|
||||||
- If acceptance work surfaces a new whoosh-compat divergence, that is a whoosh-compat-repo change (its `differential-triage` skill applies), not a silent paperless workaround.
|
|
||||||
|
|
||||||
## Fast JSON field existence checks
|
|
||||||
|
|
||||||
Existence checks against a fast JSON field work correctly, both whole-field (`notes:*`, which internally requires `json_subpaths=True`) and subpath-scoped (`custom_fields.value:*`, which checks only that subpath's own fast column). Both are covered by whoosh-compat's own test suite; see `DIVERGENCES.md` entry 20 for the exists-strategy design and its subpath-scoping note.
|
|
||||||
|
|
||||||
Whether to mark `notes`/`custom_fields` fast is a paperless-ngx-side tradeoff (fast fields cost index size/build time for cheaper existence/range queries) independent of whoosh-compat's correctness — worth a maintainer decision, not assumed by this document.
|
|
||||||
|
|
||||||
## Coordination
|
|
||||||
|
|
||||||
- whoosh-compat is pre-1.0: pin an exact version or git SHA; upgrades are deliberate, reviewed changes.
|
|
||||||
- JSON subpath emission depends on the installed tantivy-py version (fallback until quickwit-oss/tantivy-py#716 ships). The whoosh-compat repo has a `carve-out-retirement` skill; coordinate tantivy pin bumps with it, in a separate PR from the parser migration.
|
|
||||||
- Rollout: no feature flag, no shadow-compare period. Safety comes from the date-grammar parity audit and the result-level acceptance corpus instead; `_translate.py`/`_dates.py` are deleted once those are green.
|
|
||||||
|
|
||||||
## Common mistakes
|
|
||||||
|
|
||||||
- Inferring the field surface from frontend code (users type queries directly).
|
|
||||||
- Copying `KNOWN_FIELDS` into the registry wholesale (resurfaces internal fields).
|
|
||||||
- Wiring stemming into `pattern_normalizer`.
|
|
||||||
- Calling `emit()` unconditionally, or porting the legacy raw-string fallback.
|
|
||||||
- Deleting `_dates.py` without the parity audit.
|
|
||||||
- Porting `test_translate.py`'s string assertions instead of writing result-level tests.
|
|
||||||
+2
-1
@@ -301,7 +301,8 @@ The following methods are supported:
|
|||||||
- `delete`
|
- `delete`
|
||||||
- No `parameters` required
|
- No `parameters` required
|
||||||
- `reprocess`
|
- `reprocess`
|
||||||
- No `parameters` required
|
- Optional `parameters`: `{ "remote_ocr": true }` to send the documents to the
|
||||||
|
remote OCR engine, see [Remote OCR](usage.md#remote-ocr). Defaults to false.
|
||||||
- `set_permissions`
|
- `set_permissions`
|
||||||
- Requires `parameters`:
|
- Requires `parameters`:
|
||||||
- `"set_permissions": PERMISSIONS_OBJ` (see format [above](#permissions)) and / or
|
- `"set_permissions": PERMISSIONS_OBJ` (see format [above](#permissions)) and / or
|
||||||
|
|||||||
@@ -2048,6 +2048,18 @@ password. All of these options come from their similarly-named [Django settings]
|
|||||||
|
|
||||||
Defaults to None.
|
Defaults to None.
|
||||||
|
|
||||||
|
#### [`PAPERLESS_REMOTE_OCR_MODE=<str>`](#PAPERLESS_REMOTE_OCR_MODE) {#PAPERLESS_REMOTE_OCR_MODE}
|
||||||
|
|
||||||
|
: Which documents are sent to the remote OCR engine.
|
||||||
|
|
||||||
|
- `always`: every document of a supported file type is sent to the remote
|
||||||
|
engine, bypassing the local OCR engine.
|
||||||
|
- `workflow_only`: documents are processed locally unless a workflow
|
||||||
|
explicitly enables remote OCR for them, letting you use the remote engine
|
||||||
|
selectively.
|
||||||
|
|
||||||
|
Defaults to "always".
|
||||||
|
|
||||||
## AI {#ai}
|
## AI {#ai}
|
||||||
|
|
||||||
#### [`PAPERLESS_AI_ENABLED=<bool>`](#PAPERLESS_AI_ENABLED) {#PAPERLESS_AI_ENABLED}
|
#### [`PAPERLESS_AI_ENABLED=<bool>`](#PAPERLESS_AI_ENABLED) {#PAPERLESS_AI_ENABLED}
|
||||||
|
|||||||
@@ -456,6 +456,20 @@ def score(
|
|||||||
return 10
|
return 10
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**Remote services**
|
||||||
|
|
||||||
|
If your parser sends document content to a remote service, declare it:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class MyCustomParser:
|
||||||
|
uses_remote_service = True
|
||||||
|
```
|
||||||
|
|
||||||
|
Paperless-ngx excludes such parsers when the document being consumed has not
|
||||||
|
been marked for remote processing, so users can keep remote OCR off by default
|
||||||
|
and enable it selectively with a workflow. Parsers that do not declare the
|
||||||
|
attribute are treated as fully local and are always considered.
|
||||||
|
|
||||||
**Archive and rendition flags**
|
**Archive and rendition flags**
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -1,502 +0,0 @@
|
|||||||
# whoosh-compat transition design
|
|
||||||
|
|
||||||
Date: 2026-08-07
|
|
||||||
Status: approved
|
|
||||||
Related skill: `whoosh-compat-transition`
|
|
||||||
|
|
||||||
> **API reference used by this design.** `FieldRegistry` exposes one
|
|
||||||
> resolution path: `registry.make_ref(raw) -> FieldRef | None` interprets a
|
|
||||||
> raw, possibly dotted field string (an unknown field or an unknown subpath
|
|
||||||
> both return `None`), and `registry.resolve(ref) -> ResolvedField | None`
|
|
||||||
> looks up the resolved ref. `ResolvedField` carries `.spec` (the
|
|
||||||
> `FieldSpec`), `.json_path` (the subpath, or `None`), `.is_subpath`, and
|
|
||||||
> `.dotted_name` — read `resolved.spec.kind`, not `spec.kind` off a bare
|
|
||||||
> `FieldSpec`. `Diagnostic.field` is a `FieldRef`, not a string: use
|
|
||||||
> `str(d.field)` for the canonical dotted name, or `d.field.name` for the
|
|
||||||
> field alone (the name is canonical, so an aliased query like `type:`
|
|
||||||
> reports `document_type`). Every field-carrying AST leaf holds a `FieldRef`.
|
|
||||||
> `emit()`'s signature is `emit(node, *, index, registry) -> tantivy.Query`,
|
|
||||||
> with no `schema` parameter; it calls the library's own `analyze()` pipeline
|
|
||||||
> stage internally (token analysis, multitoken resolution, zero-token drop)
|
|
||||||
> before visiting the tree, so this design's call sites never invoke
|
|
||||||
> `analyze()` themselves. `FieldSpec.subpaths` is stored internally as
|
|
||||||
> `Mapping[str, SubpathSpec]`, though construction still accepts a plain
|
|
||||||
> `tuple[str, ...]` as sugar and normalizes it automatically — this design's
|
|
||||||
> own `PublicField.subpaths: tuple[str, ...]` (below) passes a tuple into
|
|
||||||
> `FieldSpec(..., subpaths=...)` and needs nothing further.
|
|
||||||
> `DiagnosticKind` has four members: `BAD_DATE`, `BAD_NUMBER`, `TOO_DEEP`, and
|
|
||||||
> `UNSUPPORTED_PATTERN`; the error-mapping code below needs cases for all
|
|
||||||
> four.
|
|
||||||
>
|
|
||||||
> A few library behaviors worth knowing before writing code against it:
|
|
||||||
> `parse()` validates its own configuration eagerly — an empty or unknown
|
|
||||||
> `default_fields`, or a `field_boosts` key that resolves to neither a known
|
|
||||||
> field nor an alias, raises `ValueError` at the `parse()` call itself, and an
|
|
||||||
> alias in either argument resolves normally. A naive `basedate` is rejected
|
|
||||||
> (`ValueError`) rather than silently read in the host machine's local
|
|
||||||
> timezone; pass an aware datetime. A wildcard/prefix pattern on a numeric
|
|
||||||
> (`U64`) field, a `BOOLEAN_EXISTS` field, or a JSON subpath produces a
|
|
||||||
> parse-time `Diagnostic(kind=UNSUPPORTED_PATTERN)` instead of silently
|
|
||||||
> mangling to an exact-match term or matching the wrong encoded bytes — this
|
|
||||||
> is directly relevant to `custom_fields.value`: a user typing
|
|
||||||
> `custom_fields.value:abc*` gets a diagnostic, not a query that silently
|
|
||||||
> matches the wrong documents. A bare JSON field name with no subpath
|
|
||||||
> (`notes:foo`) demotes to an ordinary text search for the literal string,
|
|
||||||
> the same treatment an unknown field or unknown subpath gets. Registry
|
|
||||||
> construction validates its input eagerly: exists-target cycles, empty
|
|
||||||
> field/alias names, duplicate aliases, dotted canonical names,
|
|
||||||
> invalid-character or empty JSON subpath strings, and a subpath that would
|
|
||||||
> shadow a registered plain field are all rejected at `FieldRegistry.__init__`
|
|
||||||
> with an actionable message, not deferred to query time.
|
|
||||||
>
|
|
||||||
> Fast-field existence checks against a JSON field are correct, both for
|
|
||||||
> whole-field existence (`notes:*`) and the per-subpath case
|
|
||||||
> (`custom_fields.value:*`, which checks only that subpath's own fast
|
|
||||||
> column). Marking `notes`/`custom_fields` fast is therefore a plain
|
|
||||||
> paperless-ngx-side index-size/query-cost tradeoff, independent of
|
|
||||||
> whoosh-compat correctness — worth a maintainer decision, not something this
|
|
||||||
> document settles.
|
|
||||||
|
|
||||||
## Summary
|
|
||||||
|
|
||||||
Replace paperless-ngx's hand-maintained query-translation layer
|
|
||||||
(`src/documents/search/_translate.py`, `src/documents/search/_dates.py`)
|
|
||||||
with [whoosh-compat](https://github.com/stumpylog/whoosh-compat): a typed
|
|
||||||
Whoosh-grammar parser that emits programmatically constructed
|
|
||||||
`tantivy.Query` objects instead of building an intermediate Tantivy query
|
|
||||||
_string_. The integration point is narrow: `parse_user_query()` in
|
|
||||||
`src/documents/search/_query.py` is the only function whose implementation
|
|
||||||
changes; `_backend.py`, `_tokenizer.py`, simple/title search, CJK handling,
|
|
||||||
and permission filtering are all unaffected.
|
|
||||||
|
|
||||||
Delivered as a stack of four paperless-ngx PRs plus one prerequisite change
|
|
||||||
in whoosh-compat itself (same maintainer, no cross-repo coordination
|
|
||||||
overhead), landed with no feature flag and no shadow-compare rollout period
|
|
||||||
— safety comes from a result-level acceptance test corpus and a
|
|
||||||
date-grammar parity audit instead.
|
|
||||||
|
|
||||||
## Architecture
|
|
||||||
|
|
||||||
```
|
|
||||||
raw_query (user-typed)
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
wc.parse(raw_query, registry=FIELD_REGISTRY, default_fields=DEFAULT_SEARCH_FIELDS,
|
|
||||||
field_boosts=_FIELD_BOOSTS, tz=tz)
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
ParseResult(ast, diagnostics)
|
|
||||||
│
|
|
||||||
├─ diagnostics non-empty? → map ALL diagnostics to SearchQueryError
|
|
||||||
│ subclass(es) → HTTP 400 (never just the first diagnostic)
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
emit(ast, index=index, registry=FIELD_REGISTRY)
|
|
||||||
│ (calls whoosh-compat's own analyze() pipeline stage internally, then
|
|
||||||
│ raises UnsupportedQueryError → mapped to SearchQueryError → 400,
|
|
||||||
│ for constructs that parse but can't execute against tantivy)
|
|
||||||
▼
|
|
||||||
tantivy.Query
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
existing clause assembly in parse_user_query(): Should(exact) + optional
|
|
||||||
fuzzy re-parse of raw_query + optional CJK bigram query, unchanged from today
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
_apply_permission_filter() in _backend.py wraps the result with
|
|
||||||
build_permission_filter() — entirely independent of whoosh-compat, unchanged
|
|
||||||
```
|
|
||||||
|
|
||||||
Permission filtering is explicitly out of scope for this migration:
|
|
||||||
`build_permission_filter()` builds its `tantivy.Query` directly against
|
|
||||||
`owner_id`/`viewer_id`/`viewer_group_id`, never through the parser or
|
|
||||||
registry, and those fields are exactly the internal `*_id` fields excluded
|
|
||||||
from the `FieldRegistry` (see "Field surface" below). Nothing in this
|
|
||||||
migration's diff touches it.
|
|
||||||
|
|
||||||
## PR stack
|
|
||||||
|
|
||||||
Each PR is independently buildable, reviewable, and CI-able; later PRs
|
|
||||||
rebase on earlier ones. No PR depends on whoosh-compat behavior it hasn't
|
|
||||||
already proven correct in isolation.
|
|
||||||
|
|
||||||
1. **Refactor `_schema.py` to a shared field-definition table.** Pure
|
|
||||||
refactor — `build_schema()`'s output is byte-identical before and after.
|
|
||||||
`test_schema.py` (existing) proves it.
|
|
||||||
2. **Pin whoosh-compat as a real dependency; build `FieldRegistry`.** New
|
|
||||||
`_registry.py` built from the same table PR 1 introduced. Registry unit
|
|
||||||
tests only — no wiring into search yet.
|
|
||||||
3. **Date-grammar parity audit.** A transitional, executable differential
|
|
||||||
test using the still-present `_dates.py`/`_translate.py` as the oracle.
|
|
||||||
Any gap found is fixed in whoosh-compat directly before this PR closes.
|
|
||||||
A whoosh-compat PyPI release is expected around this point (see
|
|
||||||
"Dependency pinning").
|
|
||||||
4. **Wire it in; delete the old path.** Rewrite `parse_user_query()`,
|
|
||||||
diagnostics→exception mapping, add the result-level acceptance corpus,
|
|
||||||
expand `test_api_search.py`, delete `_translate.py`/`_dates.py`/
|
|
||||||
`test_translate.py` and the internals-testing classes in `test_query.py`,
|
|
||||||
update `docs/usage.md` and changelog.
|
|
||||||
|
|
||||||
`Diagnostic` carries `field: FieldRef | None` and `raw_value: str | None`,
|
|
||||||
populated at its construction sites (`dateparse.py`'s `_error()`,
|
|
||||||
`default.py`'s `BAD_NUMBER` sites), so paperless can build typed exceptions
|
|
||||||
without parsing whoosh-compat's human-readable `message` text. `field` is a
|
|
||||||
`FieldRef`, not a plain string; see the API reference at the top of this
|
|
||||||
document.
|
|
||||||
|
|
||||||
## Field surface
|
|
||||||
|
|
||||||
The `FieldRegistry` covers only query-syntax-addressable fields — a subset
|
|
||||||
of the full Tantivy schema. Internal-only schema fields with no query-syntax
|
|
||||||
meaning of their own (`title_sort`/`correspondent_sort`/`type_sort` shadow
|
|
||||||
sort fields, `bigram_*` CJK fields, `simple_title`/`simple_content`,
|
|
||||||
`autocomplete_word`, `notes_text`) stay hardcoded `sb.add_*` calls in
|
|
||||||
`_schema.py`, untouched by the shared table.
|
|
||||||
|
|
||||||
**Decision: keep and document all five currently-undocumented-but-working
|
|
||||||
fields** (`asn`, `page_count`, `num_notes`, `original_filename`,
|
|
||||||
`checksum`) rather than dropping them — least risk of silently breaking an
|
|
||||||
existing saved view. `docs/usage.md`'s advanced-search section gets these
|
|
||||||
added with examples, as part of PR 4.
|
|
||||||
|
|
||||||
**Decision: `archive_checksum` stays out of scope.** Unlike `checksum`, it
|
|
||||||
isn't indexed in the Tantivy schema at all today (confirmed: `_schema.py`
|
|
||||||
only adds `checksum`; `_build_tantivy_doc` only calls
|
|
||||||
`doc.add_text("checksum", document.checksum)`). Making it searchable is a
|
|
||||||
schema-level change (new indexed field, new document population code), not
|
|
||||||
a parser-migration concern — left as a separate follow-up.
|
|
||||||
|
|
||||||
**Decision: internal `*_id` fields (`tag_id`, `correspondent_id`,
|
|
||||||
`document_type_id`, `storage_path_id`, `owner_id`, `viewer_id`,
|
|
||||||
`viewer_group_id`) are excluded from the `FieldRegistry` entirely.** They
|
|
||||||
remain Tantivy-schema-only, used exclusively by `build_permission_filter()`.
|
|
||||||
Because whoosh-compat folds any unrecognized `field:` prefix into literal
|
|
||||||
text (Whoosh-parity leniency, confirmed in `FieldsPlugin.do_fieldnames` —
|
|
||||||
not an error), a saved view typed as `tag_id:5` won't 400: it silently
|
|
||||||
becomes a text search for the literal string `tag_id:5`, most likely
|
|
||||||
returning zero results. This is a real behavior change and gets a
|
|
||||||
**changelog callout**, not just a docs update, since a docs addition alone
|
|
||||||
wouldn't surface it to someone skimming release notes.
|
|
||||||
|
|
||||||
## Shared field-definition table (`_fields.py`)
|
|
||||||
|
|
||||||
```python
|
|
||||||
from whoosh_compat import FieldKind # reused directly — no parallel enum
|
|
||||||
|
|
||||||
@dataclass(frozen=True, slots=True)
|
|
||||||
class PublicField:
|
|
||||||
name: str
|
|
||||||
kind: FieldKind
|
|
||||||
aliases: tuple[str, ...] = ()
|
|
||||||
comma_values: bool = False
|
|
||||||
date_only: bool = False
|
|
||||||
fast: bool = False
|
|
||||||
subpaths: tuple[str, ...] = () # JSON kind only
|
|
||||||
|
|
||||||
PUBLIC_FIELDS = (
|
|
||||||
PublicField("title", FieldKind.TEXT),
|
|
||||||
PublicField("content", FieldKind.TEXT),
|
|
||||||
PublicField("correspondent", FieldKind.TEXT),
|
|
||||||
PublicField("document_type", FieldKind.TEXT, aliases=("type",)),
|
|
||||||
PublicField("storage_path", FieldKind.TEXT, aliases=("path",)),
|
|
||||||
PublicField("original_filename", FieldKind.TEXT),
|
|
||||||
PublicField("tag", FieldKind.TEXT, comma_values=True),
|
|
||||||
PublicField("checksum", FieldKind.KEYWORD),
|
|
||||||
PublicField("asn", FieldKind.U64, fast=True),
|
|
||||||
PublicField("page_count", FieldKind.U64, fast=True),
|
|
||||||
PublicField("num_notes", FieldKind.U64, fast=True),
|
|
||||||
PublicField("created", FieldKind.DATE, date_only=True, fast=True),
|
|
||||||
PublicField("modified", FieldKind.DATETIME, fast=True),
|
|
||||||
PublicField("added", FieldKind.DATETIME, fast=True),
|
|
||||||
PublicField("notes", FieldKind.JSON, subpaths=("user", "note")),
|
|
||||||
PublicField("custom_fields", FieldKind.JSON, subpaths=("name", "value")),
|
|
||||||
)
|
|
||||||
```
|
|
||||||
|
|
||||||
`build_schema()` derives its `sb.add_*` call and tokenizer from `kind`
|
|
||||||
(TEXT/KEYWORD → `add_text_field` with `paperless_text`/`raw` tokenizer
|
|
||||||
respectively; U64 → `add_unsigned_field`; DATE/DATETIME → `add_date_field`;
|
|
||||||
JSON → `add_json_field`). The `notes_text` snippet-companion field stays a
|
|
||||||
separate hardcoded line right after the `notes` entry — schema-only
|
|
||||||
plumbing with no query-syntax meaning.
|
|
||||||
|
|
||||||
`_registry.py` maps each `PublicField` to a `whoosh_compat.FieldSpec`,
|
|
||||||
kept as one flat dataclass (no kind-specific subclassing) to mirror
|
|
||||||
whoosh-compat's own `FieldSpec` design, which validates kind-conditional
|
|
||||||
attributes (e.g. JSON requires non-empty `subpaths`) at
|
|
||||||
`FieldRegistry.__init__` rather than in the type system.
|
|
||||||
|
|
||||||
Footnote for whoever writes `_registry.py`: `FieldRegistry.__init__` forces
|
|
||||||
`date_only=True` on _any_ `FieldKind.DATE` spec regardless of what's
|
|
||||||
passed, unconditionally — `PublicField.date_only` isn't an independent
|
|
||||||
knob for DATE fields the way it might look; it only matters in the sense
|
|
||||||
that `created` sets it explicitly for clarity, while `modified`/`added`
|
|
||||||
use `FieldKind.DATETIME` instead of relying on that override.
|
|
||||||
|
|
||||||
**`PublicField.subpaths` stays `tuple[str, ...]`, not a nested structure.**
|
|
||||||
Confirmed against whoosh-compat's own `FieldRegistry.make_ref()`: it splits a
|
|
||||||
dotted query term on the _first_ dot only and matches the remainder as an
|
|
||||||
exact string against `spec.subpaths` — even the docstring's own
|
|
||||||
`"metadata.author.name"` example is a single opaque string in the tuple,
|
|
||||||
not a recursive tree. (`FieldSpec.subpaths` itself now stores a `Mapping[str,
|
|
||||||
SubpathSpec]` internally, normalized from whatever tuple is passed at
|
|
||||||
construction; that's an implementation detail of `FieldSpec.__post_init__`,
|
|
||||||
not something `PublicField`'s own table needs to mirror — passing a plain
|
|
||||||
tuple into `FieldSpec(..., subpaths=...)` still works exactly as written
|
|
||||||
here.) A tuple of strings is exactly as expressive as the library it feeds;
|
|
||||||
inventing richer structure in `PublicField` now would just get flattened
|
|
||||||
back to strings at the registry-construction boundary. Real recursive
|
|
||||||
nesting, if ever needed, is new whoosh-compat capability first (the
|
|
||||||
per-subpath `SubpathSpec` container exists specifically to make that a
|
|
||||||
later, additive change).
|
|
||||||
|
|
||||||
**JSON document population stays separate from `subpaths`.** `subpaths` is
|
|
||||||
query-side only — it declares which dotted names are legal to type and
|
|
||||||
which JSON keys the emitter should address. It says nothing about how
|
|
||||||
`_backend.py::_build_tantivy_doc` builds the JSON documents at index-write
|
|
||||||
time, and that logic isn't uniform attribute access (`note.user.username`
|
|
||||||
needs a null guard and isn't `note.user`; `cfi.value_for_search` is a
|
|
||||||
property, not a literal `value` attribute), so a generic
|
|
||||||
`getattr(obj, subpath_name)` scheme would silently do the wrong thing for
|
|
||||||
both. That code stays hand-written, unchanged by this migration. Mitigation
|
|
||||||
instead: a coupling test (PR 2, alongside the registry unit tests) asserting
|
|
||||||
the literal JSON keys used in `_build_tantivy_doc`'s `doc.add_json(...)`
|
|
||||||
calls match `PUBLIC_FIELDS`' `notes`/`custom_fields` `subpaths` exactly, so
|
|
||||||
drift between the two is caught rather than silently becoming an
|
|
||||||
unqueryable (or silently unindexed) field.
|
|
||||||
|
|
||||||
**JSON subpath queries (`notes.*`, `custom_fields.*`) route through
|
|
||||||
`index.parse_query()`, not programmatic construction, given paperless's
|
|
||||||
pinned tantivy version.** Installed `tantivy-py`'s `Query.term_query`
|
|
||||||
cannot resolve a JSON subpath by exact field name — it raises as if the
|
|
||||||
field didn't exist. Until
|
|
||||||
[tantivy-py#716](https://github.com/quickwit-oss/tantivy-py/pull/716) lands
|
|
||||||
and ships, whoosh-compat's `TantivyEmitter._json_paths_supported()` feature-
|
|
||||||
detects this per process and falls back to a strictly escaped, single-leaf
|
|
||||||
`index.parse_query()` call for just that one leaf (whoosh-compat's README/
|
|
||||||
ARCHITECTURE.md call this out as "the JSON subpath carve-out"). Paperless
|
|
||||||
pins `tantivy~=0.26.0`, squarely inside the affected range (whoosh-compat's
|
|
||||||
`tantivy` extra only requires `tantivy>=0.24`, so nothing prevents this
|
|
||||||
combination). Nothing needs to change in this design because of it — the
|
|
||||||
carve-out is self-retiring on whoosh-compat's side once tantivy-py catches
|
|
||||||
up — but the acceptance corpus's `notes.user:`/`custom_fields.name:` cases
|
|
||||||
(PR 4) are exercising that fallback escaping path specifically, not the
|
|
||||||
programmatic path every other field goes through, and that's worth knowing
|
|
||||||
if one of those cases ever behaves oddly around quoting/escaping. A
|
|
||||||
multi-token JSON subpath value with `Multitoken.AND`/`OR` now gets correct
|
|
||||||
combinator semantics through this fallback (each token becomes its own
|
|
||||||
`index.parse_query()`-backed leaf, `Must`/`Should`-combined normally,
|
|
||||||
instead of collapsing into one space-joined phrase-shaped query); a genuine
|
|
||||||
quoted phrase on a JSON subpath still cannot carry an explicit slop through
|
|
||||||
this fallback (silently ignored, `~N` has no effect) until the carve-out
|
|
||||||
retires. Also worth knowing given `custom_fields.value` is JSON: this
|
|
||||||
fallback's `index.parse_query()` call gives a JSON subpath term free
|
|
||||||
numeric/boolean type inference tantivy's own query grammar provides (a
|
|
||||||
query like `custom_fields.value:100` matches both a stored JSON number `100`
|
|
||||||
and a stored JSON string `"100"`); the future programmatic path (once
|
|
||||||
tantivy-py#716 ships) has no equivalent union and would need this
|
|
||||||
re-evaluated for numeric/boolean custom field values specifically
|
|
||||||
(whoosh-compat's `DIVERGENCES.md` entry 22 tracks this open question).
|
|
||||||
|
|
||||||
**Analyzer wiring**: `FieldSpec.analyzer` reuses the same `tantivy
|
|
||||||
.TextAnalyzer` objects `_tokenizer.py` already builds (`_paperless_text
|
|
||||||
(language)`, etc.) — standalone objects not dependent on index
|
|
||||||
registration, so `_registry.py` calls the same builder functions and binds
|
|
||||||
`.analyze` directly; `checksum` (KEYWORD, `raw` tokenizer) gets an identity
|
|
||||||
analyzer (`lambda t: [t]`). `pattern_normalizer` for every field is
|
|
||||||
`_tokenizer.ascii_fold` (character-fold only, never stemming) per the
|
|
||||||
skill's explicit instruction. The whole `FieldRegistry` is built once,
|
|
||||||
cached keyed by `settings.SEARCH_LANGUAGE`, rebuilt on the same trigger
|
|
||||||
`register_tokenizers()` already uses.
|
|
||||||
|
|
||||||
## Error handling
|
|
||||||
|
|
||||||
```python
|
|
||||||
class SearchQueryError(ValueError): ... # unchanged, base
|
|
||||||
|
|
||||||
class InvalidDateQuery(SearchQueryError): # unchanged
|
|
||||||
def __init__(self, field, value): ...
|
|
||||||
|
|
||||||
class InvalidNumberQuery(SearchQueryError): # new
|
|
||||||
def __init__(self, field: str | None, value: str | None) -> None:
|
|
||||||
self.field = field
|
|
||||||
self.value = value
|
|
||||||
super().__init__(f"Invalid numeric value {value!r} for field {field!r}.")
|
|
||||||
|
|
||||||
class MultipleSearchQueryErrors(SearchQueryError): # new
|
|
||||||
"""Aggregates every user-fixable error from one parse, not just the first."""
|
|
||||||
def __init__(self, errors: Sequence[SearchQueryError]) -> None:
|
|
||||||
self.errors = tuple(errors)
|
|
||||||
super().__init__("; ".join(str(e) for e in self.errors))
|
|
||||||
```
|
|
||||||
|
|
||||||
```python
|
|
||||||
def parse_user_query(index, raw_query, tz):
|
|
||||||
registry = get_field_registry(settings.SEARCH_LANGUAGE)
|
|
||||||
result = wc.parse(
|
|
||||||
raw_query, registry=registry, default_fields=DEFAULT_SEARCH_FIELDS,
|
|
||||||
field_boosts=_FIELD_BOOSTS, tz=tz,
|
|
||||||
)
|
|
||||||
if result.diagnostics:
|
|
||||||
raise _diagnostics_to_error(result.diagnostics) # ALL diagnostics, not [0]
|
|
||||||
|
|
||||||
try:
|
|
||||||
exact = tantivy_emit(result.ast, index=index, registry=registry)
|
|
||||||
except UnsupportedQueryError as e:
|
|
||||||
raise SearchQueryError(str(e)) from e
|
|
||||||
|
|
||||||
# CJK: unchanged — already re-parses raw_query directly via index.parse_query,
|
|
||||||
# never went through translate_query, so nothing here changes.
|
|
||||||
cjk_query = _build_cjk_query(index, raw_query, _CJK_ALL_FIELDS) if _has_cjk(raw_query) else None
|
|
||||||
|
|
||||||
clauses = [(tantivy.Occur.Should, exact)]
|
|
||||||
threshold = settings.ADVANCED_FUZZY_SEARCH_THRESHOLD
|
|
||||||
if threshold is not None:
|
|
||||||
# Fuzzy re-parses raw_query (not the AST) — no clean AST-level fuzzy
|
|
||||||
# equivalent exists; fuzzy matching was always an approximate,
|
|
||||||
# secondary clause, so this divergence from the exact-match path is
|
|
||||||
# acceptable.
|
|
||||||
fuzzy = index.parse_query(raw_query, DEFAULT_SEARCH_FIELDS, field_boosts=_FIELD_BOOSTS,
|
|
||||||
fuzzy_fields={f: (True, 1, True) for f in DEFAULT_SEARCH_FIELDS})
|
|
||||||
clauses.append((tantivy.Occur.Should, tantivy.Query.boost_query(fuzzy, 0.1)))
|
|
||||||
if cjk_query is not None:
|
|
||||||
clauses.append((tantivy.Occur.Should, cjk_query))
|
|
||||||
|
|
||||||
return exact if len(clauses) == 1 else tantivy.Query.boolean_query(clauses)
|
|
||||||
|
|
||||||
|
|
||||||
def _diagnostics_to_error(diagnostics: tuple[Diagnostic, ...]) -> SearchQueryError:
|
|
||||||
errors = [_single_diagnostic_to_error(d) for d in diagnostics]
|
|
||||||
return errors[0] if len(errors) == 1 else MultipleSearchQueryErrors(errors)
|
|
||||||
|
|
||||||
|
|
||||||
def _single_diagnostic_to_error(d: Diagnostic) -> SearchQueryError:
|
|
||||||
# d.field is a FieldRef, not a string: str(d.field) gives the canonical
|
|
||||||
# dotted name (e.g. "created", "custom_fields.value"); an aliased query
|
|
||||||
# (type:) reports the field it resolves to (document_type).
|
|
||||||
field_name = str(d.field) if d.field is not None else None
|
|
||||||
if d.kind is DiagnosticKind.BAD_DATE:
|
|
||||||
return InvalidDateQuery(field_name, d.raw_value)
|
|
||||||
if d.kind is DiagnosticKind.BAD_NUMBER:
|
|
||||||
return InvalidNumberQuery(field_name, d.raw_value)
|
|
||||||
# TOO_DEEP (pathological paren nesting) and UNSUPPORTED_PATTERN (a
|
|
||||||
# wildcard/prefix pattern on a numeric, BOOLEAN_EXISTS, or JSON-subpath
|
|
||||||
# field) both fall through to the generic message; a typed subclass for
|
|
||||||
# either isn't warranted unless a caller needs to branch on it.
|
|
||||||
return SearchQueryError(d.message)
|
|
||||||
```
|
|
||||||
|
|
||||||
No `except Exception: query_str = raw_query` fallback — per the skill, that
|
|
||||||
legacy defensive branch is explicitly not carried forward. A bug in the new
|
|
||||||
path must surface as a real error, not silently degrade to stale behavior.
|
|
||||||
|
|
||||||
`views.py`'s existing `except SearchQueryError as e: raise
|
|
||||||
ValidationError({"query": [str(e)]}) from e` handler gets one added branch
|
|
||||||
to surface every aggregated message instead of just one:
|
|
||||||
|
|
||||||
```python
|
|
||||||
except SearchQueryError as e:
|
|
||||||
messages = [str(sub) for sub in e.errors] if isinstance(e, MultipleSearchQueryErrors) else [str(e)]
|
|
||||||
raise ValidationError({"query": messages}) from e
|
|
||||||
```
|
|
||||||
|
|
||||||
`d.field`/`d.raw_value` are populated for `BAD_DATE` and `BAD_NUMBER`
|
|
||||||
diagnostics; if either is `None` for a diagnostic kind that doesn't populate
|
|
||||||
them, `_single_diagnostic_to_error`'s fallthrough to
|
|
||||||
`SearchQueryError(d.message)` still applies.
|
|
||||||
|
|
||||||
Deferred, explicitly out of scope for this PR stack: any frontend use of
|
|
||||||
`startchar`/`endchar` (already present on `Diagnostic` today) to highlight
|
|
||||||
the offending span in the search box. Backend-only for now, per explicit
|
|
||||||
decision.
|
|
||||||
|
|
||||||
## Testing
|
|
||||||
|
|
||||||
Existing test inventory (`src/documents/tests/search/` and
|
|
||||||
`test_api_search.py`):
|
|
||||||
|
|
||||||
| File | Fate |
|
|
||||||
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
||||||
| `test_translate.py` | Deleted (PR 4) — subject deleted |
|
|
||||||
| `test_query.py`: `TestCreatedDateField`, `TestDateTimeFields`, `TestWhooshQueryRewriting`, `TestYearRangeRewriting`, `TestNonDateFieldsNotRewritten`, `TestPassthrough`, `TestNormalizeQuery` | Deleted (PR 4) — test `_translate.py`/`_dates.py` internals or intermediate query strings |
|
|
||||||
| `test_query.py`: `TestParseUserQuery` | Reviewed at plan time; result-level assertions folded into the new acceptance module, internals-only assertions dropped |
|
|
||||||
| `test_query.py`: `TestParseSimpleTextHighlightQuery`, `TestPermissionFilter` | Unchanged — never touched `translate_query` |
|
|
||||||
| `test_schema.py`, `test_tokenizer.py`, `test_backend.py`, `test_lock_backoff.py`, `test_migration_fulltext_query_field_prefixes.py` | Unchanged |
|
|
||||||
| `test_api_search.py` (`TestDocumentSearchApi`, 43 tests) | **Stays green across every PR in the stack** (hard gate, not just PR 4) — full HTTP+DB+index integration coverage catches wiring mistakes none of the narrower tests would |
|
|
||||||
|
|
||||||
New tests per PR:
|
|
||||||
|
|
||||||
- **PR 2**: `test_registry.py` — internal `*_id` names rejected; `type`/
|
|
||||||
`path` aliases resolve to canonical fields; JSON subpaths match
|
|
||||||
`docs/usage.md`; registry construction deterministic per language; the
|
|
||||||
`notes`/`custom_fields` dict-key coupling test described above.
|
|
||||||
- **PR 3**: `test_date_grammar_parity.py` — transitional, parametrized over
|
|
||||||
every keyword/unit `_dates.py`/`_translate.py` accept today
|
|
||||||
(`_DATE_KEYWORDS`, all of `_UNIT_ALIASES`'s Whoosh-era abbreviations —
|
|
||||||
`yrs`/`mos`/`wks`/`hrs`/`mins`/`secs` etc. — digit-precision forms, ISO
|
|
||||||
dash forms, `now-7d`/`now+1h`/`now-30m` compact offsets, open/reversed
|
|
||||||
ranges). Each case parses through `wc.parse()` against a DATE-kind
|
|
||||||
`FieldRegistry` and asserts no diagnostics come back — a coverage check
|
|
||||||
only (does whoosh-compat accept this input at all), not a check on the
|
|
||||||
bounds or AST shape it parses to, which is whoosh-compat's own
|
|
||||||
differential-testing responsibility against a real whoosh oracle, not
|
|
||||||
something to re-verify here against `_translate.py` as a second, weaker
|
|
||||||
oracle. If the team wants confidence that actual search _behavior_ at a
|
|
||||||
given keyword didn't change, that belongs in the PR 4 result-level
|
|
||||||
acceptance corpus (real indexed documents at date boundaries, matched-ID
|
|
||||||
assertions), not an AST/bounds comparison.
|
|
||||||
Deleted again in PR 4 along with the legacy code it audits, superseded by
|
|
||||||
the permanent acceptance corpus. This audit is scoped to _parity_ only —
|
|
||||||
whoosh-compat's date grammar is a strict superset of what `_dates.py`
|
|
||||||
accepts today (e.g. `tomorrow`, `now`, `midnight`, `noon`, weekday names
|
|
||||||
like `next monday`), so the migration also grants new date vocabulary for
|
|
||||||
free. That's a nice side effect, not something this PR needs to test or
|
|
||||||
document beyond noting it in the changelog alongside the other behavior
|
|
||||||
changes.
|
|
||||||
- **PR 4**:
|
|
||||||
- Result-level acceptance module (paperless's analogue of whoosh-compat's
|
|
||||||
`test_acceptance_e2e.py`): a real index built via `build_schema()`, the
|
|
||||||
issue #13568 queries verbatim, real saved-view strings, every date
|
|
||||||
keyword/unit, field aliases, comma lists, numeric/date ranges,
|
|
||||||
bracket-class wildcards, boosts, JSON subpaths — asserted by matched
|
|
||||||
document-ID set, `pytest.param(..., id=...)` per case. Plus a
|
|
||||||
multi-diagnostic case (two bad fields → `MultipleSearchQueryErrors`
|
|
||||||
with both messages present) and one `Multitoken` case nested inside a
|
|
||||||
top-level `OR` (proves DIVERGENCES entry 15 doesn't matter for
|
|
||||||
paperless's data, per the skill).
|
|
||||||
- `test_api_search.py` expanded: a multi-bad-field query (e.g.
|
|
||||||
`created:notadate AND asn:notanumber`) asserting the 400 response's
|
|
||||||
`query` list contains both messages; end-to-end searches on the five
|
|
||||||
newly-documented fields (`asn:`, `page_count:`, `num_notes:`,
|
|
||||||
`original_filename:`, `checksum:`) returning the right documents
|
|
||||||
through the real index.
|
|
||||||
|
|
||||||
## Dependency pinning
|
|
||||||
|
|
||||||
Stays `path = "../whoosh-compat"` in `[tool.uv.sources]` through the whole
|
|
||||||
PR stack — both repos are being actively co-developed. The final swap
|
|
||||||
happens at PR 4:
|
|
||||||
|
|
||||||
- **Primary plan**: whoosh-compat is released to PyPI around PR 3 (per
|
|
||||||
your stated intent), assuming the parity audit doesn't turn up anything
|
|
||||||
needing a second round. PR 4 switches to a pinned PyPI version
|
|
||||||
(`whoosh-compat[tantivy]==X.Y.Z` in `dependencies`, the
|
|
||||||
`[tool.uv.sources]` override removed entirely).
|
|
||||||
- **Fallback**: if the PyPI release slips past PR 4's start, pin an exact
|
|
||||||
git commit SHA instead (`whoosh-compat[tantivy] @ git+https://github.com/
|
|
||||||
stumpylog/whoosh-compat@<sha>`), per the skill's "pre-1.0: pin an exact
|
|
||||||
version or git SHA, upgrades are deliberate" guidance.
|
|
||||||
|
|
||||||
The `TODO` comment already sitting in `pyproject.toml` (from the earlier
|
|
||||||
smoke-test setup) gets updated to reflect this — "release, else pinned SHA"
|
|
||||||
— rather than committing hard to one path before it's known which applies.
|
|
||||||
|
|
||||||
## Explicitly out of scope
|
|
||||||
|
|
||||||
- `archive_checksum` indexing/search (separate schema-level follow-up).
|
|
||||||
- Frontend consumption of `Diagnostic.startchar`/`endchar` for in-box error
|
|
||||||
highlighting (backend-only for this PR stack).
|
|
||||||
- A feature flag or shadow-compare rollout period — explicitly decided
|
|
||||||
against; safety comes from the acceptance corpus and parity audit instead.
|
|
||||||
- Any change to `build_permission_filter()`/`_apply_permission_filter()` —
|
|
||||||
confirmed untouched by this migration.
|
|
||||||
+22
-1
@@ -650,6 +650,19 @@ happened while it was still encrypted, that original version will likewise be mi
|
|||||||
**Current limitation**: Passwords are stored as a simple list without descriptions. To handle
|
**Current limitation**: Passwords are stored as a simple list without descriptions. To handle
|
||||||
multiple PDF types with different passwords, create separate workflows for each use case.
|
multiple PDF types with different passwords, create separate workflows for each use case.
|
||||||
|
|
||||||
|
##### Remote OCR {#workflow-action-remote-ocr}
|
||||||
|
|
||||||
|
"Remote OCR" actions send the document to the configured remote OCR engine instead of processing it
|
||||||
|
locally. To use remote OCR selectively, set the [remote OCR mode](configuration.md#PAPERLESS_REMOTE_OCR_MODE)
|
||||||
|
to `workflow_only` then add this action to a workflow that matches only the documents you
|
||||||
|
want sent to the remote engine. See [Remote OCR](#remote-ocr) for the engine setup. The action only works with
|
||||||
|
a **Consumption Started** trigger.
|
||||||
|
|
||||||
|
The action takes no options, its presence is what enables remote OCR for a matching document.
|
||||||
|
|
||||||
|
If the remote engine is not configured, or does not support the document's file type, the document is
|
||||||
|
processed locally instead and a warning is written to the log.
|
||||||
|
|
||||||
#### Workflow placeholders
|
#### Workflow placeholders
|
||||||
|
|
||||||
Titles and webhook payloads can be generated by workflows using [Jinja templates](https://jinja.palletsprojects.com/en/3.1.x/templates/).
|
Titles and webhook payloads can be generated by workflows using [Jinja templates](https://jinja.palletsprojects.com/en/3.1.x/templates/).
|
||||||
@@ -1086,11 +1099,19 @@ Paperless-ngx supports performing OCR on documents using remote services. At the
|
|||||||
[Microsoft's Azure "Document Intelligence" service](https://azure.microsoft.com/en-us/products/ai-services/ai-document-intelligence).
|
[Microsoft's Azure "Document Intelligence" service](https://azure.microsoft.com/en-us/products/ai-services/ai-document-intelligence).
|
||||||
This is of course a paid service (with a free tier) which requires an Azure account and subscription. Azure AI is not affiliated with
|
This is of course a paid service (with a free tier) which requires an Azure account and subscription. Azure AI is not affiliated with
|
||||||
Paperless-ngx in any way. When enabled, Paperless-ngx will automatically send appropriate documents to Azure for OCR processing, bypassing
|
Paperless-ngx in any way. When enabled, Paperless-ngx will automatically send appropriate documents to Azure for OCR processing, bypassing
|
||||||
the local OCR engine. See the [configuration](configuration.md#PAPERLESS_REMOTE_OCR_ENGINE) options for more details.
|
the local OCR engine. See the [configuration](configuration.md#PAPERLESS_REMOTE_OCR_ENGINE) options for more details. These
|
||||||
|
settings can be supplied as environment variables or via **Application Configuration**.
|
||||||
|
|
||||||
Additionally, when using a commercial service with this feature, consider both potential costs as well as any associated file size
|
Additionally, when using a commercial service with this feature, consider both potential costs as well as any associated file size
|
||||||
or page limitations (e.g. with a free tier).
|
or page limitations (e.g. with a free tier).
|
||||||
|
|
||||||
|
By default, every document of a supported file type is sent to the remote engine. To use it more selectively, set the
|
||||||
|
[remote OCR mode](configuration.md#PAPERLESS_REMOTE_OCR_MODE) to `workflow_only`. Documents are then processed locally
|
||||||
|
unless a [remote OCR workflow action](#workflow-action-remote-ocr) enables it for them, so you can limit the remote
|
||||||
|
engine to particular documents.
|
||||||
|
|
||||||
|
Setting the mode to `workflow_only` also allows the **Reprocess** actions to selectively use remote OCR for individual documents.
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
Paperless-ngx consists of the following components:
|
Paperless-ngx consists of the following components:
|
||||||
|
|||||||
@@ -14,43 +14,48 @@
|
|||||||
<a ngbNavLink>{{category}}</a>
|
<a ngbNavLink>{{category}}</a>
|
||||||
<ng-template ngbNavContent>
|
<ng-template ngbNavContent>
|
||||||
<div class="p-3">
|
<div class="p-3">
|
||||||
<div class="row row-cols-1 row-cols-md-2 row-cols-lg-3 g-2">
|
@for (section of getCategorySections(category); track section) {
|
||||||
@for (option of getCategoryOptions(category); track option.key) {
|
@if (section) {
|
||||||
<div class="col">
|
<h5 class="mt-4 mb-3">{{section}}</h5>
|
||||||
<div class="card bg-light">
|
}
|
||||||
<div class="card-body">
|
<div class="row row-cols-1 row-cols-md-2 row-cols-lg-3 g-2">
|
||||||
<div class="card-title d-flex align-items-center">
|
@for (option of getCategoryOptions(category, section); track option.key) {
|
||||||
<h6 class="mb-0">
|
<div class="col">
|
||||||
{{option.title}}
|
<div class="card bg-light">
|
||||||
</h6>
|
<div class="card-body">
|
||||||
<a class="btn btn-sm btn-link" title="Read the documentation about this setting" i18n-title [href]="getDocsUrl(option.config_key)" target="_blank" referrerpolicy="no-referrer">
|
<div class="card-title d-flex align-items-center">
|
||||||
<i-bs name="info-circle"></i-bs>
|
<h6 class="mb-0">
|
||||||
</a>
|
{{option.title}}
|
||||||
@if (isSet(option.key)) {
|
</h6>
|
||||||
<button type="button" class="btn btn-sm btn-link text-danger ms-auto pe-0" title="Reset" i18n-title (click)="resetOption(option.key)">
|
<a class="btn btn-sm btn-link" title="Read the documentation about this setting" i18n-title [href]="getDocsUrl(option.config_key)" target="_blank" referrerpolicy="no-referrer">
|
||||||
<i-bs class="me-1" name="x"></i-bs><ng-container i18n>Reset</ng-container>
|
<i-bs name="info-circle"></i-bs>
|
||||||
</button>
|
</a>
|
||||||
|
@if (isSet(option.key)) {
|
||||||
|
<button type="button" class="btn btn-sm btn-link text-danger ms-auto pe-0" title="Reset" i18n-title (click)="resetOption(option.key)">
|
||||||
|
<i-bs class="me-1" name="x"></i-bs><ng-container i18n>Reset</ng-container>
|
||||||
|
</button>
|
||||||
|
}
|
||||||
|
</div>
|
||||||
|
<div class="mb-n3">
|
||||||
|
@switch (option.type) {
|
||||||
|
@case (ConfigOptionType.Select) { <pngx-input-select [formControlName]="option.key" [error]="errors[option.key]" [items]="option.choices" [allowNull]="true"></pngx-input-select> }
|
||||||
|
@case (ConfigOptionType.Number) { <pngx-input-number [formControlName]="option.key" [error]="errors[option.key]" [showAdd]="false"></pngx-input-number> }
|
||||||
|
@case (ConfigOptionType.Boolean) { <pngx-input-switch [formControlName]="option.key" [error]="errors[option.key]" [showUnsetNote]="true" [horizontal]="true" title="Enable" i18n-title></pngx-input-switch> }
|
||||||
|
@case (ConfigOptionType.String) { <pngx-input-text [formControlName]="option.key" [error]="errors[option.key]"></pngx-input-text> }
|
||||||
|
@case (ConfigOptionType.JSON) { <pngx-input-text [formControlName]="option.key" [error]="errors[option.key]"></pngx-input-text> }
|
||||||
|
@case (ConfigOptionType.File) { <pngx-input-file [formControlName]="option.key" (upload)="uploadFile($event, option.key)" [error]="errors[option.key]"></pngx-input-file> }
|
||||||
|
@case (ConfigOptionType.Password) { <pngx-input-password [formControlName]="option.key" [error]="errors[option.key]"></pngx-input-password> }
|
||||||
|
}
|
||||||
|
</div>
|
||||||
|
@if (option.note) {
|
||||||
|
<div class="form-text fst-italic">{{option.note}}</div>
|
||||||
}
|
}
|
||||||
</div>
|
</div>
|
||||||
<div class="mb-n3">
|
|
||||||
@switch (option.type) {
|
|
||||||
@case (ConfigOptionType.Select) { <pngx-input-select [formControlName]="option.key" [error]="errors[option.key]" [items]="option.choices" [allowNull]="true"></pngx-input-select> }
|
|
||||||
@case (ConfigOptionType.Number) { <pngx-input-number [formControlName]="option.key" [error]="errors[option.key]" [showAdd]="false"></pngx-input-number> }
|
|
||||||
@case (ConfigOptionType.Boolean) { <pngx-input-switch [formControlName]="option.key" [error]="errors[option.key]" [showUnsetNote]="true" [horizontal]="true" title="Enable" i18n-title></pngx-input-switch> }
|
|
||||||
@case (ConfigOptionType.String) { <pngx-input-text [formControlName]="option.key" [error]="errors[option.key]"></pngx-input-text> }
|
|
||||||
@case (ConfigOptionType.JSON) { <pngx-input-text [formControlName]="option.key" [error]="errors[option.key]"></pngx-input-text> }
|
|
||||||
@case (ConfigOptionType.File) { <pngx-input-file [formControlName]="option.key" (upload)="uploadFile($event, option.key)" [error]="errors[option.key]"></pngx-input-file> }
|
|
||||||
@case (ConfigOptionType.Password) { <pngx-input-password [formControlName]="option.key" [error]="errors[option.key]"></pngx-input-password> }
|
|
||||||
}
|
|
||||||
</div>
|
|
||||||
@if (option.note) {
|
|
||||||
<div class="form-text fst-italic">{{option.note}}</div>
|
|
||||||
}
|
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
}
|
||||||
}
|
</div>
|
||||||
</div>
|
}
|
||||||
</div>
|
</div>
|
||||||
</ng-template>
|
</ng-template>
|
||||||
</li>
|
</li>
|
||||||
|
|||||||
@@ -8,7 +8,11 @@ import { NgbModule } from '@ng-bootstrap/ng-bootstrap'
|
|||||||
import { NgSelectModule } from '@ng-select/ng-select'
|
import { NgSelectModule } from '@ng-select/ng-select'
|
||||||
import { NgxBootstrapIconsModule, allIcons } from 'ngx-bootstrap-icons'
|
import { NgxBootstrapIconsModule, allIcons } from 'ngx-bootstrap-icons'
|
||||||
import { of, throwError } from 'rxjs'
|
import { of, throwError } from 'rxjs'
|
||||||
import { OutputTypeConfig } from 'src/app/data/paperless-config'
|
import {
|
||||||
|
ConfigCategory,
|
||||||
|
ConfigSection,
|
||||||
|
OutputTypeConfig,
|
||||||
|
} from 'src/app/data/paperless-config'
|
||||||
import { ConfigService } from 'src/app/services/config.service'
|
import { ConfigService } from 'src/app/services/config.service'
|
||||||
import { SettingsService } from 'src/app/services/settings.service'
|
import { SettingsService } from 'src/app/services/settings.service'
|
||||||
import { ToastService } from 'src/app/services/toast.service'
|
import { ToastService } from 'src/app/services/toast.service'
|
||||||
@@ -158,4 +162,24 @@ describe('ConfigComponent', () => {
|
|||||||
component.resetOption('barcodes_enabled')
|
component.resetOption('barcodes_enabled')
|
||||||
expect(component.configForm.get('barcodes_enabled').value).toBeNull()
|
expect(component.configForm.get('barcodes_enabled').value).toBeNull()
|
||||||
})
|
})
|
||||||
|
|
||||||
|
it('should group options into sections within a category, or not', () => {
|
||||||
|
const sections = component.getCategorySections(ConfigCategory.OCR)
|
||||||
|
expect(sections).toEqual([null, ConfigSection.RemoteOCR])
|
||||||
|
expect(
|
||||||
|
component
|
||||||
|
.getCategoryOptions(ConfigCategory.OCR)
|
||||||
|
.map((option) => option.key)
|
||||||
|
).toContain('output_type')
|
||||||
|
expect(
|
||||||
|
component
|
||||||
|
.getCategoryOptions(ConfigCategory.OCR, ConfigSection.RemoteOCR)
|
||||||
|
.map((option) => option.key)
|
||||||
|
).toEqual([
|
||||||
|
'remote_ocr_engine',
|
||||||
|
'remote_ocr_api_key',
|
||||||
|
'remote_ocr_endpoint',
|
||||||
|
'remote_ocr_mode',
|
||||||
|
])
|
||||||
|
})
|
||||||
})
|
})
|
||||||
|
|||||||
@@ -74,8 +74,20 @@ export class ConfigComponent
|
|||||||
return Object.values(ConfigCategory)
|
return Object.values(ConfigCategory)
|
||||||
}
|
}
|
||||||
|
|
||||||
getCategoryOptions(category: string): ConfigOption[] {
|
getCategorySections(category: string): string[] {
|
||||||
return PaperlessConfigOptions.filter((o) => o.category === category)
|
return [
|
||||||
|
...new Set(
|
||||||
|
PaperlessConfigOptions.filter((o) => o.category === category).map(
|
||||||
|
(o) => o.section ?? null // null means no section
|
||||||
|
)
|
||||||
|
),
|
||||||
|
]
|
||||||
|
}
|
||||||
|
|
||||||
|
getCategoryOptions(category: string, section: string = null): ConfigOption[] {
|
||||||
|
return PaperlessConfigOptions.filter(
|
||||||
|
(o) => o.category === category && (o.section ?? null) === section
|
||||||
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
initialConfig: PaperlessConfig
|
initialConfig: PaperlessConfig
|
||||||
|
|||||||
+28
@@ -0,0 +1,28 @@
|
|||||||
|
<div class="modal-header">
|
||||||
|
<h4 class="modal-title" id="modal-basic-title">{{title}}</h4>
|
||||||
|
<button type="button" class="btn-close" aria-label="Close" (click)="cancel()">
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
|
<div class="modal-body">
|
||||||
|
@if (messageBold) {
|
||||||
|
<p class="text-break"><b>{{messageBold}}</b></p>
|
||||||
|
}
|
||||||
|
@if (message) {
|
||||||
|
<p class="mb-0 text-break" [innerHTML]="message"></p>
|
||||||
|
}
|
||||||
|
@if (showRemoteOcr) {
|
||||||
|
<div class="form-check mt-3">
|
||||||
|
<input class="form-check-input" type="checkbox" id="reprocessRemoteOcr" [(ngModel)]="remoteOcr" />
|
||||||
|
<label class="form-check-label" for="reprocessRemoteOcr" i18n>Use remote OCR</label>
|
||||||
|
<div class="form-text" i18n>Sends the document to the configured remote OCR service, which may incur costs.</div>
|
||||||
|
</div>
|
||||||
|
}
|
||||||
|
</div>
|
||||||
|
<div class="modal-footer">
|
||||||
|
<button type="button" class="btn" [class]="cancelBtnClass" (click)="cancel()" [disabled]="!buttonsEnabled">
|
||||||
|
<span class="d-inline-block" style="padding-bottom: 1px;">{{cancelBtnCaption}}</span>
|
||||||
|
</button>
|
||||||
|
<button type="button" class="btn" [class]="btnClass" (click)="confirm()" [disabled]="!confirmButtonEnabled || !buttonsEnabled">
|
||||||
|
{{btnCaption}}
|
||||||
|
</button>
|
||||||
|
</div>
|
||||||
+72
@@ -0,0 +1,72 @@
|
|||||||
|
import { provideHttpClient, withInterceptorsFromDi } from '@angular/common/http'
|
||||||
|
import { provideHttpClientTesting } from '@angular/common/http/testing'
|
||||||
|
import { ComponentFixture, TestBed } from '@angular/core/testing'
|
||||||
|
import { NgbActiveModal } from '@ng-bootstrap/ng-bootstrap'
|
||||||
|
import { RemoteOCRModeConfig } from 'src/app/data/paperless-config'
|
||||||
|
import { SETTINGS_KEYS } from 'src/app/data/ui-settings'
|
||||||
|
import { SettingsService } from 'src/app/services/settings.service'
|
||||||
|
import { ReprocessConfirmDialogComponent } from './reprocess-confirm-dialog.component'
|
||||||
|
|
||||||
|
describe('ReprocessConfirmDialogComponent', () => {
|
||||||
|
let component: ReprocessConfirmDialogComponent
|
||||||
|
let fixture: ComponentFixture<ReprocessConfirmDialogComponent>
|
||||||
|
let settingsService: SettingsService
|
||||||
|
|
||||||
|
const createComponent = (configured: boolean, mode: string) => {
|
||||||
|
settingsService.set(SETTINGS_KEYS.REMOTE_OCR_CONFIGURED, configured)
|
||||||
|
settingsService.set(SETTINGS_KEYS.REMOTE_OCR_MODE, mode)
|
||||||
|
|
||||||
|
fixture = TestBed.createComponent(ReprocessConfirmDialogComponent)
|
||||||
|
component = fixture.componentInstance
|
||||||
|
fixture.detectChanges()
|
||||||
|
}
|
||||||
|
|
||||||
|
beforeEach(async () => {
|
||||||
|
TestBed.configureTestingModule({
|
||||||
|
providers: [
|
||||||
|
NgbActiveModal,
|
||||||
|
provideHttpClient(withInterceptorsFromDi()),
|
||||||
|
provideHttpClientTesting(),
|
||||||
|
],
|
||||||
|
imports: [ReprocessConfirmDialogComponent],
|
||||||
|
}).compileComponents()
|
||||||
|
|
||||||
|
settingsService = TestBed.inject(SettingsService)
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should not request remote OCR by default', () => {
|
||||||
|
createComponent(true, RemoteOCRModeConfig.WORKFLOW_ONLY)
|
||||||
|
|
||||||
|
expect(component.remoteOcr).toBeFalsy()
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should not offer remote OCR when no engine is configured', () => {
|
||||||
|
createComponent(false, RemoteOCRModeConfig.WORKFLOW_ONLY)
|
||||||
|
|
||||||
|
expect(component.showRemoteOcr).toBeFalsy()
|
||||||
|
expect(
|
||||||
|
fixture.nativeElement.querySelector('#reprocessRemoteOcr')
|
||||||
|
).toBeNull()
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should not offer remote OCR when it already handles every document', () => {
|
||||||
|
createComponent(true, RemoteOCRModeConfig.ALWAYS)
|
||||||
|
|
||||||
|
expect(component.showRemoteOcr).toBeFalsy()
|
||||||
|
expect(
|
||||||
|
fixture.nativeElement.querySelector('#reprocessRemoteOcr')
|
||||||
|
).toBeNull()
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should offer remote OCR when configured and selective', () => {
|
||||||
|
createComponent(true, RemoteOCRModeConfig.WORKFLOW_ONLY)
|
||||||
|
|
||||||
|
expect(component.showRemoteOcr).toBeTruthy()
|
||||||
|
const checkbox = fixture.nativeElement.querySelector('#reprocessRemoteOcr')
|
||||||
|
expect(checkbox).not.toBeNull()
|
||||||
|
|
||||||
|
checkbox.click()
|
||||||
|
fixture.detectChanges()
|
||||||
|
expect(component.remoteOcr).toBeTruthy()
|
||||||
|
})
|
||||||
|
})
|
||||||
+20
@@ -0,0 +1,20 @@
|
|||||||
|
import { Component, inject } from '@angular/core'
|
||||||
|
import { FormsModule } from '@angular/forms'
|
||||||
|
import { SettingsService } from 'src/app/services/settings.service'
|
||||||
|
import { ConfirmDialogComponent } from '../confirm-dialog.component'
|
||||||
|
|
||||||
|
@Component({
|
||||||
|
selector: 'pngx-reprocess-confirm-dialog',
|
||||||
|
templateUrl: './reprocess-confirm-dialog.component.html',
|
||||||
|
imports: [FormsModule],
|
||||||
|
})
|
||||||
|
export class ReprocessConfirmDialogComponent extends ConfirmDialogComponent {
|
||||||
|
private settings = inject(SettingsService)
|
||||||
|
|
||||||
|
remoteOcr: boolean = false
|
||||||
|
|
||||||
|
public get showRemoteOcr(): boolean {
|
||||||
|
// Hidden when it is not configured, or when it already handles every document anyway.
|
||||||
|
return this.settings.remoteOCRIsSelectable
|
||||||
|
}
|
||||||
|
}
|
||||||
+7
@@ -455,6 +455,13 @@
|
|||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
}
|
}
|
||||||
|
@case (WorkflowActionType.RemoteOcr) {
|
||||||
|
<div class="row">
|
||||||
|
<div class="col">
|
||||||
|
<p class="text-muted small" i18n>The document will be sent to the configured remote OCR service. May incur costs.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
}
|
||||||
}
|
}
|
||||||
</div>
|
</div>
|
||||||
</ng-template>
|
</ng-template>
|
||||||
|
|||||||
+105
-2
@@ -29,6 +29,7 @@ import {
|
|||||||
DocumentSource,
|
DocumentSource,
|
||||||
WorkflowTriggerType,
|
WorkflowTriggerType,
|
||||||
} from 'src/app/data/workflow-trigger'
|
} from 'src/app/data/workflow-trigger'
|
||||||
|
import { SETTINGS_KEYS } from 'src/app/data/ui-settings'
|
||||||
import { IfOwnerDirective } from 'src/app/directives/if-owner.directive'
|
import { IfOwnerDirective } from 'src/app/directives/if-owner.directive'
|
||||||
import { IfPermissionsDirective } from 'src/app/directives/if-permissions.directive'
|
import { IfPermissionsDirective } from 'src/app/directives/if-permissions.directive'
|
||||||
import { CorrespondentService } from 'src/app/services/rest/correspondent.service'
|
import { CorrespondentService } from 'src/app/services/rest/correspondent.service'
|
||||||
@@ -224,7 +225,12 @@ describe('WorkflowEditDialogComponent', () => {
|
|||||||
).toEqual('Document Added')
|
).toEqual('Document Added')
|
||||||
expect(component.getTriggerTypeOptionName(null)).toEqual('')
|
expect(component.getTriggerTypeOptionName(null)).toEqual('')
|
||||||
expect(component.sourceOptions).toEqual(DOCUMENT_SOURCE_OPTIONS)
|
expect(component.sourceOptions).toEqual(DOCUMENT_SOURCE_OPTIONS)
|
||||||
expect(component.actionTypeOptions).toEqual(WORKFLOW_ACTION_OPTIONS)
|
// Remote OCR is absent until the workflow has a consumption trigger
|
||||||
|
expect(component.actionTypeOptions).toEqual(
|
||||||
|
WORKFLOW_ACTION_OPTIONS.filter(
|
||||||
|
(a) => a.id !== WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
)
|
||||||
expect(
|
expect(
|
||||||
component.getActionTypeOptionName(WorkflowActionType.Assignment)
|
component.getActionTypeOptionName(WorkflowActionType.Assignment)
|
||||||
).toEqual('Assignment')
|
).toEqual('Assignment')
|
||||||
@@ -237,7 +243,104 @@ describe('WorkflowEditDialogComponent', () => {
|
|||||||
jest.spyOn(settingsService, 'get').mockReturnValue(false)
|
jest.spyOn(settingsService, 'get').mockReturnValue(false)
|
||||||
component.ngOnInit()
|
component.ngOnInit()
|
||||||
expect(component.actionTypeOptions).toEqual(
|
expect(component.actionTypeOptions).toEqual(
|
||||||
WORKFLOW_ACTION_OPTIONS.filter((a) => a.id !== WorkflowActionType.Email)
|
WORKFLOW_ACTION_OPTIONS.filter(
|
||||||
|
(a) =>
|
||||||
|
a.id !== WorkflowActionType.Email &&
|
||||||
|
a.id !== WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
)
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should offer remote OCR only for consumption workflows', () => {
|
||||||
|
jest.spyOn(settingsService, 'get').mockReturnValue(true)
|
||||||
|
|
||||||
|
// A consumption trigger makes the action reachable
|
||||||
|
component.object = {
|
||||||
|
name: 'Workflow 1',
|
||||||
|
order: 0,
|
||||||
|
enabled: true,
|
||||||
|
triggers: [{ type: WorkflowTriggerType.Consumption }],
|
||||||
|
actions: [],
|
||||||
|
} as Workflow
|
||||||
|
component.ngOnInit()
|
||||||
|
expect(component.actionTypeOptions.map((a) => a.id)).toContain(
|
||||||
|
WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
|
||||||
|
// Any other trigger type runs after the document has been parsed
|
||||||
|
component.object = {
|
||||||
|
name: 'Workflow 2',
|
||||||
|
order: 0,
|
||||||
|
enabled: true,
|
||||||
|
triggers: [{ type: WorkflowTriggerType.DocumentAdded }],
|
||||||
|
actions: [],
|
||||||
|
} as Workflow
|
||||||
|
component.ngOnInit()
|
||||||
|
expect(component.actionTypeOptions.map((a) => a.id)).not.toContain(
|
||||||
|
WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should offer remote OCR on a trigger added to a new workflow', () => {
|
||||||
|
jest.spyOn(settingsService, 'get').mockReturnValue(true)
|
||||||
|
component.ngOnInit()
|
||||||
|
|
||||||
|
// Nothing for the action to apply to yet
|
||||||
|
expect(component.actionTypeOptions.map((a) => a.id)).not.toContain(
|
||||||
|
WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
|
||||||
|
// addTrigger creates the form field with emitEvent false, so the options
|
||||||
|
// have to be computed on read rather than cached from valueChanges
|
||||||
|
component.addTrigger()
|
||||||
|
expect(component.actionTypeOptions.map((a) => a.id)).toContain(
|
||||||
|
WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
|
||||||
|
// Switching that trigger to a type that runs after parsing removes it
|
||||||
|
component.triggerFields
|
||||||
|
.at(0)
|
||||||
|
.get('type')
|
||||||
|
.setValue(WorkflowTriggerType.DocumentAdded)
|
||||||
|
expect(component.actionTypeOptions.map((a) => a.id)).not.toContain(
|
||||||
|
WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should keep remote OCR listed when an action already uses it', () => {
|
||||||
|
jest.spyOn(settingsService, 'get').mockReturnValue(true)
|
||||||
|
|
||||||
|
// Otherwise changing the trigger would silently blank the selection
|
||||||
|
component.object = {
|
||||||
|
name: 'Workflow 1',
|
||||||
|
order: 0,
|
||||||
|
enabled: true,
|
||||||
|
triggers: [{ type: WorkflowTriggerType.DocumentAdded }],
|
||||||
|
actions: [{ type: WorkflowActionType.RemoteOcr }],
|
||||||
|
} as Workflow
|
||||||
|
component.ngOnInit()
|
||||||
|
|
||||||
|
expect(component.actionTypeOptions.map((a) => a.id)).toContain(
|
||||||
|
WorkflowActionType.RemoteOcr
|
||||||
|
)
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should not offer remote OCR when no engine is configured', () => {
|
||||||
|
jest
|
||||||
|
.spyOn(settingsService, 'get')
|
||||||
|
.mockImplementation((key) => key !== SETTINGS_KEYS.REMOTE_OCR_CONFIGURED)
|
||||||
|
|
||||||
|
component.object = {
|
||||||
|
name: 'Workflow 1',
|
||||||
|
order: 0,
|
||||||
|
enabled: true,
|
||||||
|
triggers: [{ type: WorkflowTriggerType.Consumption }],
|
||||||
|
actions: [],
|
||||||
|
} as Workflow
|
||||||
|
component.ngOnInit()
|
||||||
|
|
||||||
|
expect(component.actionTypeOptions.map((a) => a.id)).not.toContain(
|
||||||
|
WorkflowActionType.RemoteOcr
|
||||||
)
|
)
|
||||||
})
|
})
|
||||||
|
|
||||||
|
|||||||
+40
-10
@@ -148,6 +148,10 @@ export const WORKFLOW_ACTION_OPTIONS = [
|
|||||||
id: WorkflowActionType.MoveToTrash,
|
id: WorkflowActionType.MoveToTrash,
|
||||||
name: $localize`Move to trash`,
|
name: $localize`Move to trash`,
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
id: WorkflowActionType.RemoteOcr,
|
||||||
|
name: $localize`Remote OCR`,
|
||||||
|
},
|
||||||
]
|
]
|
||||||
|
|
||||||
export enum TriggerFilterType {
|
export enum TriggerFilterType {
|
||||||
@@ -504,8 +508,6 @@ export class WorkflowEditDialogComponent
|
|||||||
|
|
||||||
expandedItem: number = null
|
expandedItem: number = null
|
||||||
|
|
||||||
readonly allowedActionTypes = signal([])
|
|
||||||
|
|
||||||
private readonly triggerFilterOptionsMap = new WeakMap<
|
private readonly triggerFilterOptionsMap = new WeakMap<
|
||||||
FormArray,
|
FormArray,
|
||||||
TriggerFilterOption[]
|
TriggerFilterOption[]
|
||||||
@@ -548,13 +550,40 @@ export class WorkflowEditDialogComponent
|
|||||||
this.checkRemovalActionFields.bind(this)
|
this.checkRemovalActionFields.bind(this)
|
||||||
)
|
)
|
||||||
this.checkRemovalActionFields(this.objectForm.value)
|
this.checkRemovalActionFields(this.objectForm.value)
|
||||||
this.allowedActionTypes.set(
|
}
|
||||||
this.settingsService.get(SETTINGS_KEYS.EMAIL_ENABLED)
|
|
||||||
? WORKFLOW_ACTION_OPTIONS
|
private allowedActionTypes: typeof WORKFLOW_ACTION_OPTIONS = null
|
||||||
: WORKFLOW_ACTION_OPTIONS.filter(
|
|
||||||
(a) => a.id !== WorkflowActionType.Email
|
private getAllowedActionTypes() {
|
||||||
)
|
let allowed = WORKFLOW_ACTION_OPTIONS
|
||||||
)
|
|
||||||
|
if (!this.settingsService.get(SETTINGS_KEYS.EMAIL_ENABLED)) {
|
||||||
|
allowed = allowed.filter((a) => a.id !== WorkflowActionType.Email)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Remote OCR is decided before the document is parsed, so it is only
|
||||||
|
// offered for workflows that run at consumption.
|
||||||
|
const formWorkflow: Workflow = this.objectForm?.value
|
||||||
|
const remoteOcrUsable =
|
||||||
|
this.settingsService.get(SETTINGS_KEYS.REMOTE_OCR_CONFIGURED) &&
|
||||||
|
(formWorkflow?.triggers?.some(
|
||||||
|
(trigger) => trigger.type === WorkflowTriggerType.Consumption
|
||||||
|
) ||
|
||||||
|
formWorkflow?.actions?.some(
|
||||||
|
(action) => action.type === WorkflowActionType.RemoteOcr
|
||||||
|
))
|
||||||
|
if (!remoteOcrUsable) {
|
||||||
|
allowed = allowed.filter((a) => a.id !== WorkflowActionType.RemoteOcr)
|
||||||
|
}
|
||||||
|
|
||||||
|
if (
|
||||||
|
this.allowedActionTypes?.length === allowed.length &&
|
||||||
|
this.allowedActionTypes.every((a, i) => a.id === allowed[i].id)
|
||||||
|
) {
|
||||||
|
return this.allowedActionTypes
|
||||||
|
}
|
||||||
|
this.allowedActionTypes = allowed
|
||||||
|
return allowed
|
||||||
}
|
}
|
||||||
|
|
||||||
private checkRemovalActionFields(formWorkflow: Workflow) {
|
private checkRemovalActionFields(formWorkflow: Workflow) {
|
||||||
@@ -1279,7 +1308,8 @@ export class WorkflowEditDialogComponent
|
|||||||
|
|
||||||
get actionTypeOptions() {
|
get actionTypeOptions() {
|
||||||
this.settingsService.trackChanges()
|
this.settingsService.trackChanges()
|
||||||
return this.allowedActionTypes()
|
// Computed on read rather than cached
|
||||||
|
return this.getAllowedActionTypes()
|
||||||
}
|
}
|
||||||
|
|
||||||
getActionTypeOptionName(type: WorkflowActionType): string {
|
getActionTypeOptionName(type: WorkflowActionType): string {
|
||||||
|
|||||||
@@ -963,12 +963,24 @@ describe('DocumentDetailComponent', () => {
|
|||||||
component.reprocess()
|
component.reprocess()
|
||||||
const modalCloseSpy = jest.spyOn(openModal, 'close')
|
const modalCloseSpy = jest.spyOn(openModal, 'close')
|
||||||
openModal.componentInstance.confirmClicked.next()
|
openModal.componentInstance.confirmClicked.next()
|
||||||
expect(reprocessSpy).toHaveBeenCalledWith({ documents: [doc.id] })
|
expect(reprocessSpy).toHaveBeenCalledWith({ documents: [doc.id] }, false)
|
||||||
expect(modalSpy).toHaveBeenCalled()
|
expect(modalSpy).toHaveBeenCalled()
|
||||||
expect(toastSpy).toHaveBeenCalled()
|
expect(toastSpy).toHaveBeenCalled()
|
||||||
expect(modalCloseSpy).toHaveBeenCalled()
|
expect(modalCloseSpy).toHaveBeenCalled()
|
||||||
})
|
})
|
||||||
|
|
||||||
|
it('should pass remote OCR choice when reprocessing', () => {
|
||||||
|
initNormally()
|
||||||
|
const reprocessSpy = jest.spyOn(documentService, 'reprocessDocuments')
|
||||||
|
reprocessSpy.mockReturnValue(of(true))
|
||||||
|
let openModal: NgbModalRef
|
||||||
|
modalService.activeInstances.subscribe((modal) => (openModal = modal[0]))
|
||||||
|
component.reprocess()
|
||||||
|
openModal.componentInstance.remoteOcr = true
|
||||||
|
openModal.componentInstance.confirmClicked.next()
|
||||||
|
expect(reprocessSpy).toHaveBeenCalledWith({ documents: [doc.id] }, true)
|
||||||
|
})
|
||||||
|
|
||||||
it('should show error if redo ocr call fails', () => {
|
it('should show error if redo ocr call fails', () => {
|
||||||
initNormally()
|
initNormally()
|
||||||
const reprocessSpy = jest.spyOn(documentService, 'reprocessDocuments')
|
const reprocessSpy = jest.spyOn(documentService, 'reprocessDocuments')
|
||||||
|
|||||||
@@ -97,6 +97,7 @@ import { ISODateAdapter } from 'src/app/utils/ngb-iso-date-adapter'
|
|||||||
import * as UTIF from 'utif'
|
import * as UTIF from 'utif'
|
||||||
import { DocumentDetailFieldID } from '../admin/settings/settings.component'
|
import { DocumentDetailFieldID } from '../admin/settings/settings.component'
|
||||||
import { ConfirmDialogComponent } from '../common/confirm-dialog/confirm-dialog.component'
|
import { ConfirmDialogComponent } from '../common/confirm-dialog/confirm-dialog.component'
|
||||||
|
import { ReprocessConfirmDialogComponent } from '../common/confirm-dialog/reprocess-confirm-dialog/reprocess-confirm-dialog.component'
|
||||||
import { PasswordRemovalConfirmDialogComponent } from '../common/confirm-dialog/password-removal-confirm-dialog/password-removal-confirm-dialog.component'
|
import { PasswordRemovalConfirmDialogComponent } from '../common/confirm-dialog/password-removal-confirm-dialog/password-removal-confirm-dialog.component'
|
||||||
import { CustomFieldsDropdownComponent } from '../common/custom-fields-dropdown/custom-fields-dropdown.component'
|
import { CustomFieldsDropdownComponent } from '../common/custom-fields-dropdown/custom-fields-dropdown.component'
|
||||||
import { CorrespondentEditDialogComponent } from '../common/edit-dialog/correspondent-edit-dialog/correspondent-edit-dialog.component'
|
import { CorrespondentEditDialogComponent } from '../common/edit-dialog/correspondent-edit-dialog/correspondent-edit-dialog.component'
|
||||||
@@ -1402,7 +1403,7 @@ export class DocumentDetailComponent
|
|||||||
}
|
}
|
||||||
|
|
||||||
reprocess() {
|
reprocess() {
|
||||||
let modal = this.modalService.open(ConfirmDialogComponent, {
|
let modal = this.modalService.open(ReprocessConfirmDialogComponent, {
|
||||||
backdrop: 'static',
|
backdrop: 'static',
|
||||||
})
|
})
|
||||||
modal.componentInstance.title = $localize`Reprocess confirm`
|
modal.componentInstance.title = $localize`Reprocess confirm`
|
||||||
@@ -1413,7 +1414,10 @@ export class DocumentDetailComponent
|
|||||||
modal.componentInstance.confirmClicked.subscribe(() => {
|
modal.componentInstance.confirmClicked.subscribe(() => {
|
||||||
modal.componentInstance.buttonsEnabled = false
|
modal.componentInstance.buttonsEnabled = false
|
||||||
this.documentsService
|
this.documentsService
|
||||||
.reprocessDocuments({ documents: [this.document().id] })
|
.reprocessDocuments(
|
||||||
|
{ documents: [this.document().id] },
|
||||||
|
modal.componentInstance.remoteOcr
|
||||||
|
)
|
||||||
.subscribe({
|
.subscribe({
|
||||||
next: () => {
|
next: () => {
|
||||||
this.toastService.showInfo(
|
this.toastService.showInfo(
|
||||||
|
|||||||
@@ -1122,6 +1122,7 @@ describe('BulkEditorComponent', () => {
|
|||||||
req.flush(true)
|
req.flush(true)
|
||||||
expect(req.request.body).toEqual({
|
expect(req.request.body).toEqual({
|
||||||
documents: [3, 4],
|
documents: [3, 4],
|
||||||
|
remote_ocr: false,
|
||||||
})
|
})
|
||||||
httpTestingController.match(
|
httpTestingController.match(
|
||||||
`${environment.apiBaseUrl}documents/?page=1&page_size=50&ordering=-created&truncate_content=true&include_selection_data=true`
|
`${environment.apiBaseUrl}documents/?page=1&page_size=50&ordering=-created&truncate_content=true&include_selection_data=true`
|
||||||
|
|||||||
@@ -51,6 +51,7 @@ import { ToastService } from 'src/app/services/toast.service'
|
|||||||
import { flattenTags } from 'src/app/utils/flatten-tags'
|
import { flattenTags } from 'src/app/utils/flatten-tags'
|
||||||
import { queryParamsFromFilterRules } from 'src/app/utils/query-params'
|
import { queryParamsFromFilterRules } from 'src/app/utils/query-params'
|
||||||
import { MergeConfirmDialogComponent } from '../../common/confirm-dialog/merge-confirm-dialog/merge-confirm-dialog.component'
|
import { MergeConfirmDialogComponent } from '../../common/confirm-dialog/merge-confirm-dialog/merge-confirm-dialog.component'
|
||||||
|
import { ReprocessConfirmDialogComponent } from '../../common/confirm-dialog/reprocess-confirm-dialog/reprocess-confirm-dialog.component'
|
||||||
import { RotateConfirmDialogComponent } from '../../common/confirm-dialog/rotate-confirm-dialog/rotate-confirm-dialog.component'
|
import { RotateConfirmDialogComponent } from '../../common/confirm-dialog/rotate-confirm-dialog/rotate-confirm-dialog.component'
|
||||||
import { CorrespondentEditDialogComponent } from '../../common/edit-dialog/correspondent-edit-dialog/correspondent-edit-dialog.component'
|
import { CorrespondentEditDialogComponent } from '../../common/edit-dialog/correspondent-edit-dialog/correspondent-edit-dialog.component'
|
||||||
import { CustomFieldEditDialogComponent } from '../../common/edit-dialog/custom-field-edit-dialog/custom-field-edit-dialog.component'
|
import { CustomFieldEditDialogComponent } from '../../common/edit-dialog/custom-field-edit-dialog/custom-field-edit-dialog.component'
|
||||||
@@ -909,7 +910,7 @@ export class BulkEditorComponent
|
|||||||
}
|
}
|
||||||
|
|
||||||
reprocessSelected() {
|
reprocessSelected() {
|
||||||
let modal = this.modalService.open(ConfirmDialogComponent, {
|
let modal = this.modalService.open(ReprocessConfirmDialogComponent, {
|
||||||
backdrop: 'static',
|
backdrop: 'static',
|
||||||
})
|
})
|
||||||
modal.componentInstance.title = $localize`Reprocess confirm`
|
modal.componentInstance.title = $localize`Reprocess confirm`
|
||||||
@@ -923,7 +924,10 @@ export class BulkEditorComponent
|
|||||||
modal.componentInstance.buttonsEnabled = false
|
modal.componentInstance.buttonsEnabled = false
|
||||||
this.executeDocumentAction(
|
this.executeDocumentAction(
|
||||||
modal,
|
modal,
|
||||||
this.documentService.reprocessDocuments(this.getSelectionQuery())
|
this.documentService.reprocessDocuments(
|
||||||
|
this.getSelectionQuery(),
|
||||||
|
modal.componentInstance.remoteOcr
|
||||||
|
)
|
||||||
)
|
)
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -54,6 +54,10 @@ export const ConfigCategory = {
|
|||||||
AI: $localize`AI Settings`,
|
AI: $localize`AI Settings`,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export const ConfigSection = {
|
||||||
|
RemoteOCR: $localize`Remote OCR`,
|
||||||
|
}
|
||||||
|
|
||||||
export const LLMEmbeddingBackendConfig = {
|
export const LLMEmbeddingBackendConfig = {
|
||||||
OPENAI_LIKE: 'openai-like',
|
OPENAI_LIKE: 'openai-like',
|
||||||
HUGGINGFACE: 'huggingface',
|
HUGGINGFACE: 'huggingface',
|
||||||
@@ -65,6 +69,15 @@ export const LLMBackendConfig = {
|
|||||||
OLLAMA: 'ollama',
|
OLLAMA: 'ollama',
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export const RemoteOCREngineConfig = {
|
||||||
|
AZURE_AI: 'azureai',
|
||||||
|
}
|
||||||
|
|
||||||
|
export const RemoteOCRModeConfig = {
|
||||||
|
ALWAYS: 'always',
|
||||||
|
WORKFLOW_ONLY: 'workflow_only',
|
||||||
|
}
|
||||||
|
|
||||||
export interface ConfigOption {
|
export interface ConfigOption {
|
||||||
key: string
|
key: string
|
||||||
title: string
|
title: string
|
||||||
@@ -72,6 +85,7 @@ export interface ConfigOption {
|
|||||||
choices?: Array<{ id: string; name: string }>
|
choices?: Array<{ id: string; name: string }>
|
||||||
config_key?: string
|
config_key?: string
|
||||||
category: string
|
category: string
|
||||||
|
section?: string
|
||||||
note?: string
|
note?: string
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -181,6 +195,43 @@ export const PaperlessConfigOptions: ConfigOption[] = [
|
|||||||
config_key: 'PAPERLESS_OCR_USER_ARGS',
|
config_key: 'PAPERLESS_OCR_USER_ARGS',
|
||||||
category: ConfigCategory.OCR,
|
category: ConfigCategory.OCR,
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
key: 'remote_ocr_engine',
|
||||||
|
title: $localize`Remote OCR Engine`,
|
||||||
|
type: ConfigOptionType.Select,
|
||||||
|
choices: mapToItems(RemoteOCREngineConfig),
|
||||||
|
config_key: 'PAPERLESS_REMOTE_OCR_ENGINE',
|
||||||
|
category: ConfigCategory.OCR,
|
||||||
|
section: ConfigSection.RemoteOCR,
|
||||||
|
note: $localize`Enabling remote OCR sends documents to a third-party service for processing. Consider the privacy implications as well as potential costs before enabling.`,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
key: 'remote_ocr_api_key',
|
||||||
|
title: $localize`Remote OCR API Key`,
|
||||||
|
type: ConfigOptionType.Password,
|
||||||
|
config_key: 'PAPERLESS_REMOTE_OCR_API_KEY',
|
||||||
|
category: ConfigCategory.OCR,
|
||||||
|
section: ConfigSection.RemoteOCR,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
key: 'remote_ocr_endpoint',
|
||||||
|
title: $localize`Remote OCR Endpoint`,
|
||||||
|
type: ConfigOptionType.String,
|
||||||
|
config_key: 'PAPERLESS_REMOTE_OCR_ENDPOINT',
|
||||||
|
category: ConfigCategory.OCR,
|
||||||
|
section: ConfigSection.RemoteOCR,
|
||||||
|
note: $localize`Required when using the Azure AI engine.`,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
key: 'remote_ocr_mode',
|
||||||
|
title: $localize`Remote OCR Mode`,
|
||||||
|
type: ConfigOptionType.Select,
|
||||||
|
choices: mapToItems(RemoteOCRModeConfig),
|
||||||
|
config_key: 'PAPERLESS_REMOTE_OCR_MODE',
|
||||||
|
category: ConfigCategory.OCR,
|
||||||
|
section: ConfigSection.RemoteOCR,
|
||||||
|
note: $localize`Which documents are sent to the remote engine. Use 'workflow_only' to keep remote OCR off unless a workflow enables it for a document.`,
|
||||||
|
},
|
||||||
{
|
{
|
||||||
key: 'app_logo',
|
key: 'app_logo',
|
||||||
title: $localize`Application Logo`,
|
title: $localize`Application Logo`,
|
||||||
@@ -398,6 +449,10 @@ export interface PaperlessConfig extends ObjectWithId {
|
|||||||
barcode_enable_tag: boolean
|
barcode_enable_tag: boolean
|
||||||
barcode_tag_mapping: object
|
barcode_tag_mapping: object
|
||||||
barcode_tag_split: boolean
|
barcode_tag_split: boolean
|
||||||
|
remote_ocr_engine: string
|
||||||
|
remote_ocr_api_key: string
|
||||||
|
remote_ocr_endpoint: string
|
||||||
|
remote_ocr_mode: string
|
||||||
ai_enabled: boolean
|
ai_enabled: boolean
|
||||||
llm_embedding_backend: string
|
llm_embedding_backend: string
|
||||||
llm_embedding_model: string
|
llm_embedding_model: string
|
||||||
|
|||||||
@@ -1,5 +1,6 @@
|
|||||||
import { PdfEditorEditMode } from '../components/common/pdf-editor/pdf-editor-edit-mode'
|
import { PdfEditorEditMode } from '../components/common/pdf-editor/pdf-editor-edit-mode'
|
||||||
import { PdfZoomScale } from '../components/common/pdf-viewer/pdf-viewer.types'
|
import { PdfZoomScale } from '../components/common/pdf-viewer/pdf-viewer.types'
|
||||||
|
import { RemoteOCRModeConfig } from './paperless-config'
|
||||||
import { User } from './user'
|
import { User } from './user'
|
||||||
|
|
||||||
export interface UiSettings {
|
export interface UiSettings {
|
||||||
@@ -94,6 +95,8 @@ export const SETTINGS_KEYS = {
|
|||||||
OUTLOOK_OAUTH_URL: 'outlook_oauth_url',
|
OUTLOOK_OAUTH_URL: 'outlook_oauth_url',
|
||||||
EMAIL_ENABLED: 'email_enabled',
|
EMAIL_ENABLED: 'email_enabled',
|
||||||
AI_ENABLED: 'ai_enabled',
|
AI_ENABLED: 'ai_enabled',
|
||||||
|
REMOTE_OCR_CONFIGURED: 'remote_ocr:configured',
|
||||||
|
REMOTE_OCR_MODE: 'remote_ocr:mode',
|
||||||
}
|
}
|
||||||
|
|
||||||
export const SETTINGS: UiSetting[] = [
|
export const SETTINGS: UiSetting[] = [
|
||||||
@@ -347,4 +350,14 @@ export const SETTINGS: UiSetting[] = [
|
|||||||
type: 'string',
|
type: 'string',
|
||||||
default: PdfEditorEditMode.Create,
|
default: PdfEditorEditMode.Create,
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
key: SETTINGS_KEYS.REMOTE_OCR_CONFIGURED,
|
||||||
|
type: 'boolean',
|
||||||
|
default: false,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
key: SETTINGS_KEYS.REMOTE_OCR_MODE,
|
||||||
|
type: 'string',
|
||||||
|
default: RemoteOCRModeConfig.ALWAYS,
|
||||||
|
},
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ export enum WorkflowActionType {
|
|||||||
Webhook = 4,
|
Webhook = 4,
|
||||||
PasswordRemoval = 5,
|
PasswordRemoval = 5,
|
||||||
MoveToTrash = 6,
|
MoveToTrash = 6,
|
||||||
|
RemoteOcr = 7,
|
||||||
}
|
}
|
||||||
|
|
||||||
export interface WorkflowActionEmail extends ObjectWithId {
|
export interface WorkflowActionEmail extends ObjectWithId {
|
||||||
|
|||||||
@@ -284,6 +284,21 @@ describe(`DocumentService`, () => {
|
|||||||
expect(req.request.method).toEqual('POST')
|
expect(req.request.method).toEqual('POST')
|
||||||
expect(req.request.body).toEqual({
|
expect(req.request.body).toEqual({
|
||||||
documents: ids,
|
documents: ids,
|
||||||
|
remote_ocr: false,
|
||||||
|
})
|
||||||
|
})
|
||||||
|
|
||||||
|
it('should request remote OCR when reprocessing with it enabled', () => {
|
||||||
|
const ids = [1, 2, 3]
|
||||||
|
subscription = service
|
||||||
|
.reprocessDocuments({ documents: ids }, true)
|
||||||
|
.subscribe()
|
||||||
|
const req = httpTestingController.expectOne(
|
||||||
|
`${environment.apiBaseUrl}${endpoint}/reprocess/`
|
||||||
|
)
|
||||||
|
expect(req.request.body).toEqual({
|
||||||
|
documents: ids,
|
||||||
|
remote_ocr: true,
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|
||||||
|
|||||||
@@ -349,9 +349,13 @@ export class DocumentService extends AbstractPaperlessService<Document> {
|
|||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
reprocessDocuments(selection: DocumentSelectionQuery) {
|
reprocessDocuments(
|
||||||
|
selection: DocumentSelectionQuery,
|
||||||
|
remoteOcr: boolean = false
|
||||||
|
) {
|
||||||
return this.http.post(this.getResourceUrl(null, 'reprocess'), {
|
return this.http.post(this.getResourceUrl(null, 'reprocess'), {
|
||||||
...selection,
|
...selection,
|
||||||
|
remote_ocr: remoteOcr,
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -13,6 +13,7 @@ import { environment } from 'src/environments/environment'
|
|||||||
import { CustomFieldDataType } from '../data/custom-field'
|
import { CustomFieldDataType } from '../data/custom-field'
|
||||||
import { DEFAULT_DISPLAY_FIELDS, DisplayField } from '../data/document'
|
import { DEFAULT_DISPLAY_FIELDS, DisplayField } from '../data/document'
|
||||||
import { SavedView } from '../data/saved-view'
|
import { SavedView } from '../data/saved-view'
|
||||||
|
import { RemoteOCRModeConfig } from '../data/paperless-config'
|
||||||
import { SETTINGS_KEYS, UiSettings } from '../data/ui-settings'
|
import { SETTINGS_KEYS, UiSettings } from '../data/ui-settings'
|
||||||
import { PermissionsService } from './permissions.service'
|
import { PermissionsService } from './permissions.service'
|
||||||
import { CustomFieldsService } from './rest/custom-fields.service'
|
import { CustomFieldsService } from './rest/custom-fields.service'
|
||||||
@@ -434,4 +435,26 @@ describe('SettingsService', () => {
|
|||||||
).name
|
).name
|
||||||
).toEqual(customFields[0].name)
|
).toEqual(customFields[0].name)
|
||||||
})
|
})
|
||||||
|
it('should offer remote OCR only when configured and selective', () => {
|
||||||
|
settingsService.set(SETTINGS_KEYS.REMOTE_OCR_CONFIGURED, false)
|
||||||
|
settingsService.set(
|
||||||
|
SETTINGS_KEYS.REMOTE_OCR_MODE,
|
||||||
|
RemoteOCRModeConfig.WORKFLOW_ONLY
|
||||||
|
)
|
||||||
|
expect(settingsService.remoteOCRIsSelectable).toBeFalsy()
|
||||||
|
|
||||||
|
// configured, but already handling every document
|
||||||
|
settingsService.set(SETTINGS_KEYS.REMOTE_OCR_CONFIGURED, true)
|
||||||
|
settingsService.set(
|
||||||
|
SETTINGS_KEYS.REMOTE_OCR_MODE,
|
||||||
|
RemoteOCRModeConfig.ALWAYS
|
||||||
|
)
|
||||||
|
expect(settingsService.remoteOCRIsSelectable).toBeFalsy()
|
||||||
|
|
||||||
|
settingsService.set(
|
||||||
|
SETTINGS_KEYS.REMOTE_OCR_MODE,
|
||||||
|
RemoteOCRModeConfig.WORKFLOW_ONLY
|
||||||
|
)
|
||||||
|
expect(settingsService.remoteOCRIsSelectable).toBeTruthy()
|
||||||
|
})
|
||||||
})
|
})
|
||||||
|
|||||||
@@ -19,6 +19,7 @@ import {
|
|||||||
} from 'src/app/utils/color'
|
} from 'src/app/utils/color'
|
||||||
import { DEFAULT_APP_TITLE, environment } from 'src/environments/environment'
|
import { DEFAULT_APP_TITLE, environment } from 'src/environments/environment'
|
||||||
import { DEFAULT_DISPLAY_FIELDS, DisplayField } from '../data/document'
|
import { DEFAULT_DISPLAY_FIELDS, DisplayField } from '../data/document'
|
||||||
|
import { RemoteOCRModeConfig } from '../data/paperless-config'
|
||||||
import { SavedView } from '../data/saved-view'
|
import { SavedView } from '../data/saved-view'
|
||||||
import {
|
import {
|
||||||
PAPERLESS_GREEN_HEX,
|
PAPERLESS_GREEN_HEX,
|
||||||
@@ -687,6 +688,17 @@ export class SettingsService {
|
|||||||
return this.settingIsSet(SETTINGS_KEYS.UPDATE_CHECKING_ENABLED)
|
return this.settingIsSet(SETTINGS_KEYS.UPDATE_CHECKING_ENABLED)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* Offering remote OCR as a choice only makes sense when an engine
|
||||||
|
* is configured but is not already handling every document.
|
||||||
|
*/
|
||||||
|
get remoteOCRIsSelectable(): boolean {
|
||||||
|
return (
|
||||||
|
this.get(SETTINGS_KEYS.REMOTE_OCR_CONFIGURED) &&
|
||||||
|
this.get(SETTINGS_KEYS.REMOTE_OCR_MODE) !== RemoteOCRModeConfig.ALWAYS
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
offerTour(): boolean {
|
offerTour(): boolean {
|
||||||
return this.dashboardIsEmpty() && !this.get(SETTINGS_KEYS.TOUR_COMPLETE)
|
return this.dashboardIsEmpty() && !this.get(SETTINGS_KEYS.TOUR_COMPLETE)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -394,10 +394,16 @@ def delete(doc_ids: list[int]) -> Literal["OK"]:
|
|||||||
return "OK"
|
return "OK"
|
||||||
|
|
||||||
|
|
||||||
def reprocess(doc_ids: list[int]) -> Literal["OK"]:
|
def reprocess(doc_ids: list[int], *, remote_ocr: bool = False) -> Literal["OK"]:
|
||||||
|
"""
|
||||||
|
Re-run parsing for the given documents.
|
||||||
|
|
||||||
|
Consumption workflows do not run here, so ``remote_ocr`` is how the user
|
||||||
|
asks for the remote engine when it is not configured to handle everything.
|
||||||
|
"""
|
||||||
for document_id in doc_ids:
|
for document_id in doc_ids:
|
||||||
update_document_content_maybe_archive_file.apply_async(
|
update_document_content_maybe_archive_file.apply_async(
|
||||||
kwargs={"document_id": document_id},
|
kwargs={"document_id": document_id, "remote_ocr": remote_ocr},
|
||||||
headers={"trigger_source": PaperlessTask.TriggerSource.MANUAL},
|
headers={"trigger_source": PaperlessTask.TriggerSource.MANUAL},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|||||||
@@ -53,6 +53,7 @@ from documents.utils import copy_basic_file_stats
|
|||||||
from documents.utils import copy_file_with_basic_stats
|
from documents.utils import copy_file_with_basic_stats
|
||||||
from documents.utils import run_subprocess
|
from documents.utils import run_subprocess
|
||||||
from paperless.config import OcrConfig
|
from paperless.config import OcrConfig
|
||||||
|
from paperless.config import RemoteOCRConfig
|
||||||
from paperless.models import ArchiveFileGenerationChoices
|
from paperless.models import ArchiveFileGenerationChoices
|
||||||
from paperless.parsers import ParserContext
|
from paperless.parsers import ParserContext
|
||||||
from paperless.parsers import ParserProtocol
|
from paperless.parsers import ParserProtocol
|
||||||
@@ -451,12 +452,19 @@ class ConsumerPlugin(
|
|||||||
except Exception as e:
|
except Exception as e:
|
||||||
self.log.error(f"Error attempting to clean PDF: {e}")
|
self.log.error(f"Error attempting to clean PDF: {e}")
|
||||||
|
|
||||||
|
# Workflows have already run at this point, so the metadata knows
|
||||||
|
# whether this document was singled out for remote OCR
|
||||||
|
allow_remote = (
|
||||||
|
self.metadata.remote_ocr or RemoteOCRConfig().remote_ocr_by_default
|
||||||
|
)
|
||||||
|
|
||||||
# Based on the mime type, get the parser for that type
|
# Based on the mime type, get the parser for that type
|
||||||
parser_class: type[ParserProtocol] | None = (
|
parser_class: type[ParserProtocol] | None = (
|
||||||
get_parser_registry().get_parser_for_file(
|
get_parser_registry().get_parser_for_file(
|
||||||
mime_type,
|
mime_type,
|
||||||
self.filename,
|
self.filename,
|
||||||
self.working_copy,
|
self.working_copy,
|
||||||
|
allow_remote=allow_remote,
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
if not parser_class:
|
if not parser_class:
|
||||||
@@ -465,6 +473,16 @@ class ConsumerPlugin(
|
|||||||
f"Unsupported mime type {mime_type}",
|
f"Unsupported mime type {mime_type}",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
if self.metadata.remote_ocr and not getattr(
|
||||||
|
parser_class,
|
||||||
|
"uses_remote_service",
|
||||||
|
False,
|
||||||
|
):
|
||||||
|
self.log.warning(
|
||||||
|
"Remote OCR was requested for this document but no remote "
|
||||||
|
"parser is available for it, processing locally instead.",
|
||||||
|
)
|
||||||
|
|
||||||
# Notify all listeners that we're going to do some work.
|
# Notify all listeners that we're going to do some work.
|
||||||
|
|
||||||
document_consumption_started.send(
|
document_consumption_started.send(
|
||||||
|
|||||||
@@ -34,6 +34,7 @@ class DocumentMetadataOverrides:
|
|||||||
skip_asn_if_exists: bool = False
|
skip_asn_if_exists: bool = False
|
||||||
version_label: str | None = None
|
version_label: str | None = None
|
||||||
actor_id: int | None = None
|
actor_id: int | None = None
|
||||||
|
remote_ocr: bool = False
|
||||||
|
|
||||||
def update(self, other: "DocumentMetadataOverrides") -> "DocumentMetadataOverrides":
|
def update(self, other: "DocumentMetadataOverrides") -> "DocumentMetadataOverrides":
|
||||||
"""
|
"""
|
||||||
@@ -57,6 +58,8 @@ class DocumentMetadataOverrides:
|
|||||||
self.actor_id = other.actor_id
|
self.actor_id = other.actor_id
|
||||||
if other.skip_asn_if_exists:
|
if other.skip_asn_if_exists:
|
||||||
self.skip_asn_if_exists = True
|
self.skip_asn_if_exists = True
|
||||||
|
if other.remote_ocr:
|
||||||
|
self.remote_ocr = True
|
||||||
if other.version_label is not None:
|
if other.version_label is not None:
|
||||||
self.version_label = other.version_label
|
self.version_label = other.version_label
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,30 @@
|
|||||||
|
# Generated by Django 5.2.16 on 2026-08-10 17:27
|
||||||
|
|
||||||
|
from django.db import migrations
|
||||||
|
from django.db import models
|
||||||
|
|
||||||
|
|
||||||
|
class Migration(migrations.Migration):
|
||||||
|
dependencies = [
|
||||||
|
("documents", "0023_savedview_icon"),
|
||||||
|
]
|
||||||
|
|
||||||
|
operations = [
|
||||||
|
migrations.AlterField(
|
||||||
|
model_name="workflowaction",
|
||||||
|
name="type",
|
||||||
|
field=models.PositiveSmallIntegerField(
|
||||||
|
choices=[
|
||||||
|
(1, "Assignment"),
|
||||||
|
(2, "Removal"),
|
||||||
|
(3, "Email"),
|
||||||
|
(4, "Webhook"),
|
||||||
|
(5, "Password removal"),
|
||||||
|
(6, "Move to trash"),
|
||||||
|
(7, "Remote OCR"),
|
||||||
|
],
|
||||||
|
default=1,
|
||||||
|
verbose_name="Workflow Action Type",
|
||||||
|
),
|
||||||
|
),
|
||||||
|
]
|
||||||
@@ -1668,6 +1668,10 @@ class WorkflowAction(models.Model):
|
|||||||
6,
|
6,
|
||||||
_("Move to trash"),
|
_("Move to trash"),
|
||||||
)
|
)
|
||||||
|
REMOTE_OCR = (
|
||||||
|
7,
|
||||||
|
_("Remote OCR"),
|
||||||
|
)
|
||||||
|
|
||||||
type = models.PositiveSmallIntegerField(
|
type = models.PositiveSmallIntegerField(
|
||||||
_("Workflow Action Type"),
|
_("Workflow Action Type"),
|
||||||
|
|||||||
@@ -1745,7 +1745,7 @@ class DeleteDocumentsSerializer(DocumentSelectionSerializer):
|
|||||||
|
|
||||||
|
|
||||||
class ReprocessDocumentsSerializer(DocumentSelectionSerializer):
|
class ReprocessDocumentsSerializer(DocumentSelectionSerializer):
|
||||||
pass
|
remote_ocr = serializers.BooleanField(required=False, default=False)
|
||||||
|
|
||||||
|
|
||||||
class BulkEditSerializer(
|
class BulkEditSerializer(
|
||||||
@@ -2087,6 +2087,13 @@ class BulkEditSerializer(
|
|||||||
f"Page {op['page']} is out of bounds for document with {doc.page_count} pages.",
|
f"Page {op['page']} is out of bounds for document with {doc.page_count} pages.",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
def _validate_parameters_reprocess(self, parameters) -> None:
|
||||||
|
if "remote_ocr" in parameters:
|
||||||
|
if not isinstance(parameters["remote_ocr"], bool):
|
||||||
|
raise serializers.ValidationError("remote_ocr must be a boolean")
|
||||||
|
else:
|
||||||
|
parameters["remote_ocr"] = False
|
||||||
|
|
||||||
def validate_parameters_remove_password(self, parameters):
|
def validate_parameters_remove_password(self, parameters):
|
||||||
if "password" not in parameters:
|
if "password" not in parameters:
|
||||||
raise serializers.ValidationError("password not specified")
|
raise serializers.ValidationError("password not specified")
|
||||||
@@ -2151,6 +2158,8 @@ class BulkEditSerializer(
|
|||||||
self._validate_parameters_edit_pdf(parameters, attrs["documents"][0])
|
self._validate_parameters_edit_pdf(parameters, attrs["documents"][0])
|
||||||
elif method == bulk_edit.remove_password:
|
elif method == bulk_edit.remove_password:
|
||||||
self.validate_parameters_remove_password(parameters)
|
self.validate_parameters_remove_password(parameters)
|
||||||
|
elif method == bulk_edit.reprocess:
|
||||||
|
self._validate_parameters_reprocess(parameters)
|
||||||
|
|
||||||
return attrs
|
return attrs
|
||||||
|
|
||||||
@@ -3276,6 +3285,41 @@ class WorkflowSerializer(serializers.ModelSerializer[Workflow]):
|
|||||||
"actions",
|
"actions",
|
||||||
]
|
]
|
||||||
|
|
||||||
|
def validate(self, attrs):
|
||||||
|
attrs = super().validate(attrs)
|
||||||
|
|
||||||
|
if "actions" in attrs:
|
||||||
|
has_remote_ocr_action = any(
|
||||||
|
action.get("type") == WorkflowAction.WorkflowActionType.REMOTE_OCR
|
||||||
|
for action in attrs["actions"]
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
has_remote_ocr_action = self.instance is not None and (
|
||||||
|
self.instance.actions.filter(
|
||||||
|
type=WorkflowAction.WorkflowActionType.REMOTE_OCR,
|
||||||
|
).exists()
|
||||||
|
)
|
||||||
|
|
||||||
|
if "triggers" in attrs:
|
||||||
|
has_consumption_trigger = any(
|
||||||
|
trigger.get("type") == WorkflowTrigger.WorkflowTriggerType.CONSUMPTION
|
||||||
|
for trigger in attrs["triggers"]
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
has_consumption_trigger = self.instance is not None and (
|
||||||
|
self.instance.triggers.filter(
|
||||||
|
type=WorkflowTrigger.WorkflowTriggerType.CONSUMPTION,
|
||||||
|
).exists()
|
||||||
|
)
|
||||||
|
|
||||||
|
# Remote OCR can only work with consumption triggers
|
||||||
|
if has_remote_ocr_action and not has_consumption_trigger:
|
||||||
|
raise serializers.ValidationError(
|
||||||
|
"Remote OCR actions require a consumption started trigger",
|
||||||
|
)
|
||||||
|
|
||||||
|
return attrs
|
||||||
|
|
||||||
def update_triggers_and_actions(
|
def update_triggers_and_actions(
|
||||||
self,
|
self,
|
||||||
instance: Workflow,
|
instance: Workflow,
|
||||||
|
|||||||
@@ -971,6 +971,17 @@ def run_workflows(
|
|||||||
)
|
)
|
||||||
elif action.type == WorkflowAction.WorkflowActionType.MOVE_TO_TRASH:
|
elif action.type == WorkflowAction.WorkflowActionType.MOVE_TO_TRASH:
|
||||||
has_move_to_trash_action = True
|
has_move_to_trash_action = True
|
||||||
|
elif action.type == WorkflowAction.WorkflowActionType.REMOTE_OCR:
|
||||||
|
if use_overrides and overrides:
|
||||||
|
overrides.remote_ocr = True
|
||||||
|
else:
|
||||||
|
# If a workflow has a consumption trigger *and* another type,
|
||||||
|
# the document has already been parsed by the time the other one fires
|
||||||
|
logger.debug(
|
||||||
|
"Remote OCR action only applies to consumption "
|
||||||
|
"triggers, ignoring",
|
||||||
|
extra={"group": logging_group},
|
||||||
|
)
|
||||||
|
|
||||||
if not use_overrides:
|
if not use_overrides:
|
||||||
# limit title to 128 characters
|
# limit title to 128 characters
|
||||||
|
|||||||
+10
-1
@@ -66,6 +66,7 @@ from documents.utils import compute_checksum
|
|||||||
from documents.utils import identity
|
from documents.utils import identity
|
||||||
from documents.workflows.utils import get_workflows_for_trigger
|
from documents.workflows.utils import get_workflows_for_trigger
|
||||||
from paperless.config import AIConfig
|
from paperless.config import AIConfig
|
||||||
|
from paperless.config import RemoteOCRConfig
|
||||||
from paperless.logging import consume_task_id
|
from paperless.logging import consume_task_id
|
||||||
from paperless.parsers import ParserContext
|
from paperless.parsers import ParserContext
|
||||||
from paperless.parsers.registry import get_parser_registry
|
from paperless.parsers.registry import get_parser_registry
|
||||||
@@ -337,10 +338,17 @@ def bulk_update_documents(document_ids) -> None:
|
|||||||
|
|
||||||
|
|
||||||
@shared_task
|
@shared_task
|
||||||
def update_document_content_maybe_archive_file(document_id) -> None:
|
def update_document_content_maybe_archive_file(
|
||||||
|
document_id,
|
||||||
|
*,
|
||||||
|
remote_ocr: bool = False,
|
||||||
|
) -> None:
|
||||||
"""
|
"""
|
||||||
Re-creates OCR content and thumbnail for a document, and archive file if
|
Re-creates OCR content and thumbnail for a document, and archive file if
|
||||||
it exists.
|
it exists.
|
||||||
|
|
||||||
|
Remote OCR is used only when the engine is configured to handle everything
|
||||||
|
or if explicitly asked for via ``remote_ocr``.
|
||||||
"""
|
"""
|
||||||
document = Document.objects.get(id=document_id)
|
document = Document.objects.get(id=document_id)
|
||||||
|
|
||||||
@@ -350,6 +358,7 @@ def update_document_content_maybe_archive_file(document_id) -> None:
|
|||||||
mime_type,
|
mime_type,
|
||||||
document.original_filename or "",
|
document.original_filename or "",
|
||||||
document.source_path,
|
document.source_path,
|
||||||
|
allow_remote=remote_ocr or RemoteOCRConfig().remote_ocr_by_default,
|
||||||
)
|
)
|
||||||
|
|
||||||
if not parser_class:
|
if not parser_class:
|
||||||
|
|||||||
@@ -72,6 +72,10 @@ class TestApiAppConfig(DirectoriesMixin, APITestCase):
|
|||||||
"barcode_enable_tag": None,
|
"barcode_enable_tag": None,
|
||||||
"barcode_tag_mapping": None,
|
"barcode_tag_mapping": None,
|
||||||
"barcode_tag_split": None,
|
"barcode_tag_split": None,
|
||||||
|
"remote_ocr_engine": None,
|
||||||
|
"remote_ocr_api_key": None,
|
||||||
|
"remote_ocr_endpoint": None,
|
||||||
|
"remote_ocr_mode": None,
|
||||||
"ai_enabled": False,
|
"ai_enabled": False,
|
||||||
"llm_embedding_backend": None,
|
"llm_embedding_backend": None,
|
||||||
"llm_embedding_model": None,
|
"llm_embedding_model": None,
|
||||||
@@ -870,6 +874,49 @@ class TestApiAppConfig(DirectoriesMixin, APITestCase):
|
|||||||
config.refresh_from_db()
|
config.refresh_from_db()
|
||||||
self.assertEqual(config.llm_api_key, None)
|
self.assertEqual(config.llm_api_key, None)
|
||||||
|
|
||||||
|
def test_update_remote_ocr_api_key(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- Existing config with remote_ocr_api_key specified
|
||||||
|
WHEN:
|
||||||
|
- API to update remote_ocr_api_key is called with all *s
|
||||||
|
- API to update remote_ocr_api_key is called with empty string
|
||||||
|
THEN:
|
||||||
|
- remote_ocr_api_key is unchanged
|
||||||
|
- remote_ocr_api_key is set to None
|
||||||
|
"""
|
||||||
|
config = ApplicationConfiguration.objects.first()
|
||||||
|
assert config is not None
|
||||||
|
config.remote_ocr_api_key = "1234567890"
|
||||||
|
config.save()
|
||||||
|
|
||||||
|
# Test with all *
|
||||||
|
response = self.client.patch(
|
||||||
|
f"{self.ENDPOINT}1/",
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"remote_ocr_api_key": "*" * 32,
|
||||||
|
},
|
||||||
|
),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_200_OK)
|
||||||
|
config.refresh_from_db()
|
||||||
|
self.assertEqual(config.remote_ocr_api_key, "1234567890")
|
||||||
|
# Test with empty string
|
||||||
|
response = self.client.patch(
|
||||||
|
f"{self.ENDPOINT}1/",
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"remote_ocr_api_key": "",
|
||||||
|
},
|
||||||
|
),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_200_OK)
|
||||||
|
config.refresh_from_db()
|
||||||
|
self.assertEqual(config.remote_ocr_api_key, None)
|
||||||
|
|
||||||
def test_enable_ai_index_triggers_update(self) -> None:
|
def test_enable_ai_index_triggers_update(self) -> None:
|
||||||
"""
|
"""
|
||||||
GIVEN:
|
GIVEN:
|
||||||
|
|||||||
@@ -532,7 +532,29 @@ class TestBulkEditAPI(DirectoriesMixin, APITestCase):
|
|||||||
m.assert_called_once()
|
m.assert_called_once()
|
||||||
args, kwargs = m.call_args
|
args, kwargs = m.call_args
|
||||||
self.assertEqual(args[0], [self.doc1.id])
|
self.assertEqual(args[0], [self.doc1.id])
|
||||||
self.assertEqual(len(kwargs), 0)
|
self.assertEqual(kwargs, {"remote_ocr": False})
|
||||||
|
|
||||||
|
@mock.patch("documents.views.bulk_edit.reprocess")
|
||||||
|
def test_reprocess_documents_endpoint_remote_ocr(self, m) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- API data to reprocess a document with remote OCR requested
|
||||||
|
WHEN:
|
||||||
|
- API is called
|
||||||
|
THEN:
|
||||||
|
- reprocess is called with remote_ocr=True
|
||||||
|
"""
|
||||||
|
self.setup_mock(m, "reprocess")
|
||||||
|
response = self.client.post(
|
||||||
|
"/api/documents/reprocess/",
|
||||||
|
json.dumps({"documents": [self.doc1.id], "remote_ocr": True}),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_200_OK)
|
||||||
|
m.assert_called_once()
|
||||||
|
args, kwargs = m.call_args
|
||||||
|
self.assertEqual(args[0], [self.doc1.id])
|
||||||
|
self.assertEqual(kwargs, {"remote_ocr": True})
|
||||||
|
|
||||||
@mock.patch("documents.serialisers.bulk_edit.set_storage_path")
|
@mock.patch("documents.serialisers.bulk_edit.set_storage_path")
|
||||||
def test_api_set_storage_path(self, m) -> None:
|
def test_api_set_storage_path(self, m) -> None:
|
||||||
@@ -1553,6 +1575,29 @@ class TestBulkEditAPI(DirectoriesMixin, APITestCase):
|
|||||||
),
|
),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
def test_legacy_bulk_edit_reprocess_invalid_remote_ocr(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- The deprecated bulk_edit endpoint with a non-boolean remote_ocr
|
||||||
|
WHEN:
|
||||||
|
- API is called
|
||||||
|
THEN:
|
||||||
|
- The request is rejected rather than passed through to the task
|
||||||
|
"""
|
||||||
|
response = self.client.post(
|
||||||
|
"/api/documents/bulk_edit/",
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"documents": [self.doc1.id],
|
||||||
|
"method": "reprocess",
|
||||||
|
"parameters": {"remote_ocr": "yes please"},
|
||||||
|
},
|
||||||
|
),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
|
||||||
|
|
||||||
@mock.patch("documents.views.bulk_edit.edit_pdf")
|
@mock.patch("documents.views.bulk_edit.edit_pdf")
|
||||||
def test_edit_pdf(self, m) -> None:
|
def test_edit_pdf(self, m) -> None:
|
||||||
self.setup_mock(m, "edit_pdf")
|
self.setup_mock(m, "edit_pdf")
|
||||||
|
|||||||
@@ -60,6 +60,10 @@ class TestApiUiSettings(DirectoriesMixin, APITestCase):
|
|||||||
},
|
},
|
||||||
"email_enabled": False,
|
"email_enabled": False,
|
||||||
"ai_enabled": False,
|
"ai_enabled": False,
|
||||||
|
"remote_ocr": {
|
||||||
|
"configured": False,
|
||||||
|
"mode": "always",
|
||||||
|
},
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -154,6 +158,50 @@ class TestApiUiSettings(DirectoriesMixin, APITestCase):
|
|||||||
str(response.data["settings"]),
|
str(response.data["settings"]),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
@override_settings(
|
||||||
|
REMOTE_OCR_ENGINE="azureai",
|
||||||
|
REMOTE_OCR_API_KEY="somekey",
|
||||||
|
REMOTE_OCR_ENDPOINT="https://example.cognitiveservices.azure.com",
|
||||||
|
REMOTE_OCR_MODE="workflow_only",
|
||||||
|
)
|
||||||
|
def test_settings_reports_remote_ocr_when_configured(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- A fully configured remote OCR engine in workflow_only mode
|
||||||
|
WHEN:
|
||||||
|
- The ui_settings endpoint is called
|
||||||
|
THEN:
|
||||||
|
- The UI is told remote OCR is available and selective, so it can
|
||||||
|
offer it where it would actually change something
|
||||||
|
"""
|
||||||
|
response = self.client.get(self.ENDPOINT, format="json")
|
||||||
|
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_200_OK)
|
||||||
|
self.assertEqual(
|
||||||
|
response.data["settings"]["remote_ocr"],
|
||||||
|
{"configured": True, "mode": "workflow_only"},
|
||||||
|
)
|
||||||
|
|
||||||
|
@override_settings(
|
||||||
|
REMOTE_OCR_ENGINE="azureai",
|
||||||
|
REMOTE_OCR_API_KEY=None,
|
||||||
|
REMOTE_OCR_ENDPOINT=None,
|
||||||
|
)
|
||||||
|
def test_settings_reports_remote_ocr_incompletely_configured(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- An engine named but missing its endpoint and API key
|
||||||
|
WHEN:
|
||||||
|
- The ui_settings endpoint is called
|
||||||
|
THEN:
|
||||||
|
- It is reported as not configured, matching what the parser
|
||||||
|
registry will actually do
|
||||||
|
"""
|
||||||
|
response = self.client.get(self.ENDPOINT, format="json")
|
||||||
|
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_200_OK)
|
||||||
|
self.assertFalse(response.data["settings"]["remote_ocr"]["configured"])
|
||||||
|
|
||||||
@override_settings(
|
@override_settings(
|
||||||
OAUTH_CALLBACK_BASE_URL="http://localhost:8000",
|
OAUTH_CALLBACK_BASE_URL="http://localhost:8000",
|
||||||
GMAIL_OAUTH_CLIENT_ID="abc123",
|
GMAIL_OAUTH_CLIENT_ID="abc123",
|
||||||
|
|||||||
@@ -390,6 +390,141 @@ class TestApiWorkflows(DirectoriesMixin, APITestCase):
|
|||||||
|
|
||||||
self.assertEqual(Workflow.objects.count(), 1)
|
self.assertEqual(Workflow.objects.count(), 1)
|
||||||
|
|
||||||
|
def test_api_create_remote_ocr_action_requires_consumption_trigger(
|
||||||
|
self,
|
||||||
|
) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- API request to create a workflow with a remote OCR action
|
||||||
|
- No consumption started trigger, so the action could never run
|
||||||
|
WHEN:
|
||||||
|
- API is called
|
||||||
|
THEN:
|
||||||
|
- Correct HTTP 400 response
|
||||||
|
- No objects are created
|
||||||
|
"""
|
||||||
|
existing_count = Workflow.objects.count()
|
||||||
|
|
||||||
|
response = self.client.post(
|
||||||
|
self.ENDPOINT,
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"name": "Remote OCR too late",
|
||||||
|
"order": 1,
|
||||||
|
"triggers": [
|
||||||
|
{
|
||||||
|
"type": WorkflowTrigger.WorkflowTriggerType.DOCUMENT_ADDED,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
"actions": [
|
||||||
|
{
|
||||||
|
"type": WorkflowAction.WorkflowActionType.REMOTE_OCR,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
},
|
||||||
|
),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
|
||||||
|
self.assertEqual(Workflow.objects.count(), existing_count)
|
||||||
|
|
||||||
|
def test_api_create_remote_ocr_action_with_consumption_trigger(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- API request to create a workflow with a remote OCR action
|
||||||
|
- A consumption started trigger alongside another trigger type
|
||||||
|
WHEN:
|
||||||
|
- API is called
|
||||||
|
THEN:
|
||||||
|
- The workflow is created, the action applies to consumption only
|
||||||
|
"""
|
||||||
|
response = self.client.post(
|
||||||
|
self.ENDPOINT,
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"name": "Remote OCR on consume",
|
||||||
|
"order": 1,
|
||||||
|
"triggers": [
|
||||||
|
{
|
||||||
|
"type": WorkflowTrigger.WorkflowTriggerType.CONSUMPTION,
|
||||||
|
"filter_filename": "*.pdf",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": WorkflowTrigger.WorkflowTriggerType.DOCUMENT_ADDED,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
"actions": [
|
||||||
|
{
|
||||||
|
"type": WorkflowAction.WorkflowActionType.REMOTE_OCR,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
},
|
||||||
|
),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_201_CREATED)
|
||||||
|
|
||||||
|
def test_api_partial_update_adds_remote_ocr_action(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- An existing workflow with a consumption started trigger
|
||||||
|
WHEN:
|
||||||
|
- A partial update adds a remote OCR action without resubmitting triggers
|
||||||
|
THEN:
|
||||||
|
- The existing trigger is considered and the update succeeds
|
||||||
|
"""
|
||||||
|
response = self.client.patch(
|
||||||
|
f"{self.ENDPOINT}{self.workflow.id}/",
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"actions": [
|
||||||
|
{
|
||||||
|
"type": WorkflowAction.WorkflowActionType.REMOTE_OCR,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
},
|
||||||
|
),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_200_OK)
|
||||||
|
self.assertEqual(
|
||||||
|
self.workflow.actions.get().type,
|
||||||
|
WorkflowAction.WorkflowActionType.REMOTE_OCR,
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_api_partial_update_cannot_remove_remote_ocr_trigger(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- An existing workflow with a remote OCR action
|
||||||
|
- An existing consumption started trigger
|
||||||
|
WHEN:
|
||||||
|
- A partial update replaces the trigger without resubmitting actions
|
||||||
|
THEN:
|
||||||
|
- The existing action is considered and the update is rejected
|
||||||
|
"""
|
||||||
|
self.action.type = WorkflowAction.WorkflowActionType.REMOTE_OCR
|
||||||
|
self.action.save()
|
||||||
|
|
||||||
|
response = self.client.patch(
|
||||||
|
f"{self.ENDPOINT}{self.workflow.id}/",
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"triggers": [
|
||||||
|
{
|
||||||
|
"type": WorkflowTrigger.WorkflowTriggerType.DOCUMENT_UPDATED,
|
||||||
|
},
|
||||||
|
],
|
||||||
|
},
|
||||||
|
),
|
||||||
|
content_type="application/json",
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(response.status_code, status.HTTP_400_BAD_REQUEST)
|
||||||
|
self.assertEqual(self.workflow.triggers.get(), self.trigger)
|
||||||
|
|
||||||
def test_api_create_workflow_trigger_action_empty_fields(self) -> None:
|
def test_api_create_workflow_trigger_action_empty_fields(self) -> None:
|
||||||
"""
|
"""
|
||||||
GIVEN:
|
GIVEN:
|
||||||
|
|||||||
@@ -1782,3 +1782,56 @@ class TestPDFActions(DirectoriesMixin, TestCase):
|
|||||||
|
|
||||||
self.assertIn("wrong password", str(exc.exception))
|
self.assertIn("wrong password", str(exc.exception))
|
||||||
self.assertIn("Error removing password from document", cm.output[0])
|
self.assertIn("Error removing password from document", cm.output[0])
|
||||||
|
|
||||||
|
|
||||||
|
class TestBulkEditReprocess(DirectoriesMixin, TestCase):
|
||||||
|
def setUp(self) -> None:
|
||||||
|
super().setUp()
|
||||||
|
|
||||||
|
self.doc = Document.objects.create(
|
||||||
|
title="test",
|
||||||
|
checksum="A",
|
||||||
|
mime_type="application/pdf",
|
||||||
|
)
|
||||||
|
|
||||||
|
@mock.patch("documents.bulk_edit.update_document_content_maybe_archive_file")
|
||||||
|
def test_reprocess_defaults_to_local(self, mock_task: mock.Mock) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- A reprocess request that says nothing about remote OCR
|
||||||
|
WHEN:
|
||||||
|
- reprocess is called
|
||||||
|
THEN:
|
||||||
|
- The task is queued without asking for the remote engine
|
||||||
|
"""
|
||||||
|
result = bulk_edit.reprocess([self.doc.id])
|
||||||
|
|
||||||
|
self.assertEqual(result, "OK")
|
||||||
|
mock_task.apply_async.assert_called_once()
|
||||||
|
_, kwargs = mock_task.apply_async.call_args
|
||||||
|
self.assertEqual(
|
||||||
|
kwargs["kwargs"],
|
||||||
|
{"document_id": self.doc.id, "remote_ocr": False},
|
||||||
|
)
|
||||||
|
|
||||||
|
@mock.patch("documents.bulk_edit.update_document_content_maybe_archive_file")
|
||||||
|
def test_reprocess_passes_remote_ocr(self, mock_task: mock.Mock) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- A reprocess request that explicitly asks for remote OCR
|
||||||
|
WHEN:
|
||||||
|
- reprocess is called
|
||||||
|
THEN:
|
||||||
|
- The request is forwarded to the task for every document
|
||||||
|
"""
|
||||||
|
other = Document.objects.create(
|
||||||
|
title="test2",
|
||||||
|
checksum="B",
|
||||||
|
mime_type="application/pdf",
|
||||||
|
)
|
||||||
|
|
||||||
|
bulk_edit.reprocess([self.doc.id, other.id], remote_ocr=True)
|
||||||
|
|
||||||
|
self.assertEqual(mock_task.apply_async.call_count, 2)
|
||||||
|
for call in mock_task.apply_async.call_args_list:
|
||||||
|
self.assertTrue(call.kwargs["kwargs"]["remote_ocr"])
|
||||||
|
|||||||
@@ -1559,6 +1559,72 @@ class PostConsumeTestCase(DirectoriesMixin, GetConsumerMixin, TestCase):
|
|||||||
consumer.run_post_consume_script(doc)
|
consumer.run_post_consume_script(doc)
|
||||||
|
|
||||||
|
|
||||||
|
class TestConsumerRemoteOCR(
|
||||||
|
DirectoriesMixin,
|
||||||
|
FileSystemAssertsMixin,
|
||||||
|
GetConsumerMixin,
|
||||||
|
TestCase,
|
||||||
|
):
|
||||||
|
"""
|
||||||
|
The consumer resolves the remote OCR mode and the per-document request from
|
||||||
|
workflows into the allow_remote flag it hands to the parser registry.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def setUp(self) -> None:
|
||||||
|
super().setUp()
|
||||||
|
|
||||||
|
patcher = mock.patch("documents.consumer.get_parser_registry")
|
||||||
|
self.mock_registry = patcher.start()
|
||||||
|
self.mock_registry.return_value.get_parser_for_file.return_value = DummyParser
|
||||||
|
self.addCleanup(patcher.stop)
|
||||||
|
|
||||||
|
def _consume(self, *, overrides: DocumentMetadataOverrides | None = None) -> bool:
|
||||||
|
src = (
|
||||||
|
Path(__file__).parent
|
||||||
|
/ "samples"
|
||||||
|
/ "documents"
|
||||||
|
/ "originals"
|
||||||
|
/ "0000001.pdf"
|
||||||
|
)
|
||||||
|
dst = self.dirs.scratch_dir / "sample.pdf"
|
||||||
|
shutil.copy(src, dst)
|
||||||
|
|
||||||
|
with self.get_consumer(dst, overrides=overrides) as consumer:
|
||||||
|
consumer.run()
|
||||||
|
|
||||||
|
_, kwargs = self.mock_registry.return_value.get_parser_for_file.call_args
|
||||||
|
return kwargs["allow_remote"]
|
||||||
|
|
||||||
|
@override_settings(REMOTE_OCR_MODE="always")
|
||||||
|
def test_always_mode_allows_remote(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: Remote OCR mode is 'always'.
|
||||||
|
WHEN: A document is consumed without any workflow asking for it.
|
||||||
|
THEN: The registry is allowed to pick the remote parser.
|
||||||
|
"""
|
||||||
|
self.assertTrue(self._consume())
|
||||||
|
|
||||||
|
@override_settings(REMOTE_OCR_MODE="workflow_only")
|
||||||
|
def test_workflow_only_mode_denies_remote_by_default(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: Remote OCR mode is 'workflow_only'.
|
||||||
|
WHEN: A document is consumed and nothing asked for remote OCR.
|
||||||
|
THEN: The remote parser is excluded.
|
||||||
|
"""
|
||||||
|
self.assertFalse(self._consume())
|
||||||
|
|
||||||
|
@override_settings(REMOTE_OCR_MODE="workflow_only")
|
||||||
|
def test_workflow_only_mode_allows_remote_when_requested(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: Remote OCR mode is 'workflow_only'.
|
||||||
|
WHEN: A workflow set remote_ocr on the metadata overrides.
|
||||||
|
THEN: The registry is allowed to pick the remote parser.
|
||||||
|
"""
|
||||||
|
self.assertTrue(
|
||||||
|
self._consume(overrides=DocumentMetadataOverrides(remote_ocr=True)),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
class TestMetadataOverrides(TestCase):
|
class TestMetadataOverrides(TestCase):
|
||||||
def test_update_skip_asn_if_exists(self) -> None:
|
def test_update_skip_asn_if_exists(self) -> None:
|
||||||
base = DocumentMetadataOverrides()
|
base = DocumentMetadataOverrides()
|
||||||
@@ -1566,6 +1632,20 @@ class TestMetadataOverrides(TestCase):
|
|||||||
base.update(incoming)
|
base.update(incoming)
|
||||||
self.assertTrue(base.skip_asn_if_exists)
|
self.assertTrue(base.skip_asn_if_exists)
|
||||||
|
|
||||||
|
def test_update_remote_ocr(self) -> None:
|
||||||
|
base = DocumentMetadataOverrides()
|
||||||
|
base.update(DocumentMetadataOverrides(remote_ocr=True))
|
||||||
|
self.assertTrue(base.remote_ocr)
|
||||||
|
|
||||||
|
def test_update_remote_ocr_is_not_unset(self) -> None:
|
||||||
|
"""
|
||||||
|
A later workflow that says nothing must not undo an earlier one that
|
||||||
|
asked for remote OCR.
|
||||||
|
"""
|
||||||
|
base = DocumentMetadataOverrides(remote_ocr=True)
|
||||||
|
base.update(DocumentMetadataOverrides())
|
||||||
|
self.assertTrue(base.remote_ocr)
|
||||||
|
|
||||||
def test_update_actor_and_version_label(self) -> None:
|
def test_update_actor_and_version_label(self) -> None:
|
||||||
base = DocumentMetadataOverrides(
|
base = DocumentMetadataOverrides(
|
||||||
actor_id=1,
|
actor_id=1,
|
||||||
|
|||||||
@@ -287,6 +287,45 @@ class TestUpdateContent(DirectoriesMixin, TestCase):
|
|||||||
self.assertNotEqual(Document.objects.get(pk=doc.pk).content, "test")
|
self.assertNotEqual(Document.objects.get(pk=doc.pk).content, "test")
|
||||||
|
|
||||||
|
|
||||||
|
class TestUpdateContentRemoteOCR(DirectoriesMixin, TestCase):
|
||||||
|
"""
|
||||||
|
Consumption workflows do not run on reprocess, so the remote parser is
|
||||||
|
used only in 'always' mode or when the caller explicitly asks for it.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def setUp(self) -> None:
|
||||||
|
super().setUp()
|
||||||
|
|
||||||
|
patcher = mock.patch("documents.tasks.get_parser_registry")
|
||||||
|
self.mock_registry = patcher.start()
|
||||||
|
self.mock_registry.return_value.get_parser_for_file.return_value = None
|
||||||
|
self.addCleanup(patcher.stop)
|
||||||
|
|
||||||
|
self.doc = Document.objects.create(
|
||||||
|
title="test",
|
||||||
|
content="my document",
|
||||||
|
checksum="wow",
|
||||||
|
mime_type="application/pdf",
|
||||||
|
)
|
||||||
|
|
||||||
|
def _allow_remote(self, **kwargs) -> bool:
|
||||||
|
tasks.update_document_content_maybe_archive_file(self.doc.pk, **kwargs)
|
||||||
|
_, call_kwargs = self.mock_registry.return_value.get_parser_for_file.call_args
|
||||||
|
return call_kwargs["allow_remote"]
|
||||||
|
|
||||||
|
@override_settings(REMOTE_OCR_MODE="always")
|
||||||
|
def test_always_mode_allows_remote(self) -> None:
|
||||||
|
self.assertTrue(self._allow_remote())
|
||||||
|
|
||||||
|
@override_settings(REMOTE_OCR_MODE="workflow_only")
|
||||||
|
def test_workflow_only_mode_denies_remote_by_default(self) -> None:
|
||||||
|
self.assertFalse(self._allow_remote())
|
||||||
|
|
||||||
|
@override_settings(REMOTE_OCR_MODE="workflow_only")
|
||||||
|
def test_workflow_only_mode_allows_remote_when_requested(self) -> None:
|
||||||
|
self.assertTrue(self._allow_remote(remote_ocr=True))
|
||||||
|
|
||||||
|
|
||||||
class TestAIIndex(DirectoriesMixin, TestCase):
|
class TestAIIndex(DirectoriesMixin, TestCase):
|
||||||
@override_settings(
|
@override_settings(
|
||||||
AI_ENABLED=True,
|
AI_ENABLED=True,
|
||||||
|
|||||||
@@ -5409,3 +5409,82 @@ class TestDateWorkflowLocalization(
|
|||||||
document = Document.objects.first()
|
document = Document.objects.first()
|
||||||
assert document is not None
|
assert document is not None
|
||||||
assert document.title == expected_title
|
assert document.title == expected_title
|
||||||
|
|
||||||
|
|
||||||
|
class TestRemoteOCRWorkflowAction(DirectoriesMixin, SampleDirMixin, APITestCase):
|
||||||
|
def _make_workflow(self, trigger_type) -> None:
|
||||||
|
trigger = WorkflowTrigger.objects.create(type=trigger_type)
|
||||||
|
action = WorkflowAction.objects.create(
|
||||||
|
type=WorkflowAction.WorkflowActionType.REMOTE_OCR,
|
||||||
|
)
|
||||||
|
w = Workflow.objects.create(name="Remote OCR", order=0)
|
||||||
|
w.triggers.add(trigger)
|
||||||
|
w.actions.add(action)
|
||||||
|
w.save()
|
||||||
|
|
||||||
|
def test_consumption_trigger_requests_remote_ocr(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- A consumption workflow with a remote OCR action
|
||||||
|
WHEN:
|
||||||
|
- A matching document is consumed
|
||||||
|
THEN:
|
||||||
|
- The overrides ask for remote OCR, which is what the consumer
|
||||||
|
reads when choosing a parser
|
||||||
|
"""
|
||||||
|
self._make_workflow(WorkflowTrigger.WorkflowTriggerType.CONSUMPTION)
|
||||||
|
|
||||||
|
test_file = shutil.copy(
|
||||||
|
self.SAMPLE_DIR / "simple.pdf",
|
||||||
|
self.dirs.scratch_dir / "simple.pdf",
|
||||||
|
)
|
||||||
|
overrides = DocumentMetadataOverrides()
|
||||||
|
|
||||||
|
run_workflows(
|
||||||
|
WorkflowTrigger.WorkflowTriggerType.CONSUMPTION,
|
||||||
|
ConsumableDocument(
|
||||||
|
source=DocumentSource.ConsumeFolder,
|
||||||
|
original_file=test_file,
|
||||||
|
),
|
||||||
|
overrides=overrides,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertTrue(overrides.remote_ocr)
|
||||||
|
|
||||||
|
def test_other_trigger_types_are_ignored(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN:
|
||||||
|
- A workflow with a remote OCR action that also has a
|
||||||
|
non-consumption trigger, which is a valid combination
|
||||||
|
WHEN:
|
||||||
|
- The non-consumption trigger fires
|
||||||
|
THEN:
|
||||||
|
- The action is skipped, since the document has already been
|
||||||
|
parsed by this point
|
||||||
|
"""
|
||||||
|
trigger = WorkflowTrigger.objects.create(
|
||||||
|
type=WorkflowTrigger.WorkflowTriggerType.CONSUMPTION,
|
||||||
|
)
|
||||||
|
updated_trigger = WorkflowTrigger.objects.create(
|
||||||
|
type=WorkflowTrigger.WorkflowTriggerType.DOCUMENT_UPDATED,
|
||||||
|
)
|
||||||
|
action = WorkflowAction.objects.create(
|
||||||
|
type=WorkflowAction.WorkflowActionType.REMOTE_OCR,
|
||||||
|
)
|
||||||
|
w = Workflow.objects.create(name="Remote OCR", order=0)
|
||||||
|
w.triggers.add(trigger, updated_trigger)
|
||||||
|
w.actions.add(action)
|
||||||
|
w.save()
|
||||||
|
|
||||||
|
doc = Document.objects.create(
|
||||||
|
title="sample test",
|
||||||
|
original_filename="sample.pdf",
|
||||||
|
)
|
||||||
|
|
||||||
|
with self.assertLogs("paperless.handlers", level="DEBUG") as cm:
|
||||||
|
run_workflows(
|
||||||
|
WorkflowTrigger.WorkflowTriggerType.DOCUMENT_UPDATED,
|
||||||
|
doc,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertIn("only applies to consumption triggers", "".join(cm.output))
|
||||||
|
|||||||
@@ -236,8 +236,10 @@ from paperless import version
|
|||||||
from paperless.celery import app as celery_app
|
from paperless.celery import app as celery_app
|
||||||
from paperless.config import AIConfig
|
from paperless.config import AIConfig
|
||||||
from paperless.config import GeneralConfig
|
from paperless.config import GeneralConfig
|
||||||
|
from paperless.config import RemoteOCRConfig
|
||||||
from paperless.models import ApplicationConfiguration
|
from paperless.models import ApplicationConfiguration
|
||||||
from paperless.parsers.registry import get_parser_registry
|
from paperless.parsers.registry import get_parser_registry
|
||||||
|
from paperless.parsers.remote import RemoteEngineConfig
|
||||||
from paperless.serialisers import GroupSerializer
|
from paperless.serialisers import GroupSerializer
|
||||||
from paperless.serialisers import UserSerializer
|
from paperless.serialisers import UserSerializer
|
||||||
from paperless.views import StandardPagination
|
from paperless.views import StandardPagination
|
||||||
@@ -4010,6 +4012,11 @@ class UiSettingsView(GenericAPIView[Any]):
|
|||||||
|
|
||||||
ui_settings["auditlog_enabled"] = settings.AUDIT_LOG_ENABLED
|
ui_settings["auditlog_enabled"] = settings.AUDIT_LOG_ENABLED
|
||||||
|
|
||||||
|
ui_settings["remote_ocr"] = {
|
||||||
|
"configured": RemoteEngineConfig.from_app_config().engine_is_valid(),
|
||||||
|
"mode": RemoteOCRConfig().remote_ocr_mode,
|
||||||
|
}
|
||||||
|
|
||||||
if settings.GMAIL_OAUTH_ENABLED or settings.OUTLOOK_OAUTH_ENABLED:
|
if settings.GMAIL_OAUTH_ENABLED or settings.OUTLOOK_OAUTH_ENABLED:
|
||||||
manager = PaperlessMailOAuth2Manager()
|
manager = PaperlessMailOAuth2Manager()
|
||||||
if settings.GMAIL_OAUTH_ENABLED:
|
if settings.GMAIL_OAUTH_ENABLED:
|
||||||
|
|||||||
@@ -338,13 +338,16 @@ def check_deprecated_v2_ocr_env_vars(
|
|||||||
|
|
||||||
|
|
||||||
@register()
|
@register()
|
||||||
def check_remote_parser_configured(app_configs: Any, **kwargs: Any) -> list[Error]:
|
def check_remote_ocr_mode(app_configs: Any, **kwargs: Any) -> list[Error]:
|
||||||
if settings.REMOTE_OCR_ENGINE == "azureai" and not (
|
# Import here because checks.py runs before the app registry is ready
|
||||||
settings.REMOTE_OCR_ENDPOINT and settings.REMOTE_OCR_API_KEY
|
from paperless.models import RemoteOCRMode
|
||||||
):
|
|
||||||
|
valid_modes = {mode.value for mode in RemoteOCRMode}
|
||||||
|
if settings.REMOTE_OCR_MODE not in valid_modes:
|
||||||
return [
|
return [
|
||||||
Error(
|
Error(
|
||||||
"Azure AI remote parser requires endpoint and API key to be configured.",
|
f"PAPERLESS_REMOTE_OCR_MODE is set to {settings.REMOTE_OCR_MODE!r}, "
|
||||||
|
f"expected one of {sorted(valid_modes)}.",
|
||||||
),
|
),
|
||||||
]
|
]
|
||||||
|
|
||||||
|
|||||||
@@ -9,6 +9,7 @@ from paperless.models import CleanChoices
|
|||||||
from paperless.models import ColorConvertChoices
|
from paperless.models import ColorConvertChoices
|
||||||
from paperless.models import ModeChoices
|
from paperless.models import ModeChoices
|
||||||
from paperless.models import OutputTypeChoices
|
from paperless.models import OutputTypeChoices
|
||||||
|
from paperless.models import RemoteOCRMode
|
||||||
|
|
||||||
|
|
||||||
@dataclasses.dataclass
|
@dataclasses.dataclass
|
||||||
@@ -185,6 +186,45 @@ class GeneralConfig(BaseConfig):
|
|||||||
self.app_logo = app_config.app_logo.url if app_config.app_logo else None
|
self.app_logo = app_config.app_logo.url if app_config.app_logo else None
|
||||||
|
|
||||||
|
|
||||||
|
@dataclasses.dataclass
|
||||||
|
class RemoteOCRConfig(BaseConfig):
|
||||||
|
"""
|
||||||
|
Settings for the remote (cloud) OCR parser
|
||||||
|
"""
|
||||||
|
|
||||||
|
remote_ocr_engine: str | None = dataclasses.field(init=False)
|
||||||
|
remote_ocr_api_key: str | None = dataclasses.field(init=False)
|
||||||
|
remote_ocr_endpoint: str | None = dataclasses.field(init=False)
|
||||||
|
remote_ocr_mode: RemoteOCRMode = dataclasses.field(init=False)
|
||||||
|
|
||||||
|
def __post_init__(self) -> None:
|
||||||
|
app_config = self._get_config_instance()
|
||||||
|
|
||||||
|
self.remote_ocr_engine = (
|
||||||
|
app_config.remote_ocr_engine or settings.REMOTE_OCR_ENGINE
|
||||||
|
)
|
||||||
|
self.remote_ocr_api_key = (
|
||||||
|
app_config.remote_ocr_api_key or settings.REMOTE_OCR_API_KEY
|
||||||
|
)
|
||||||
|
self.remote_ocr_endpoint = (
|
||||||
|
app_config.remote_ocr_endpoint or settings.REMOTE_OCR_ENDPOINT
|
||||||
|
)
|
||||||
|
self.remote_ocr_mode = app_config.remote_ocr_mode or RemoteOCRMode(
|
||||||
|
settings.REMOTE_OCR_MODE,
|
||||||
|
)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def remote_ocr_by_default(self) -> bool:
|
||||||
|
"""
|
||||||
|
Whether every supported document goes to the remote engine.
|
||||||
|
|
||||||
|
When False the remote engine is used only for documents that
|
||||||
|
explicitly asked for it, i.e. a workflow matched during consumption or
|
||||||
|
the user ticked the box when reprocessing.
|
||||||
|
"""
|
||||||
|
return self.remote_ocr_mode == RemoteOCRMode.ALWAYS
|
||||||
|
|
||||||
|
|
||||||
@dataclasses.dataclass
|
@dataclasses.dataclass
|
||||||
class AIConfig(BaseConfig):
|
class AIConfig(BaseConfig):
|
||||||
"""
|
"""
|
||||||
|
|||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# Generated by Django 5.2.16 on 2026-08-10 14:37
|
||||||
|
|
||||||
|
from django.db import migrations
|
||||||
|
from django.db import models
|
||||||
|
|
||||||
|
|
||||||
|
class Migration(migrations.Migration):
|
||||||
|
dependencies = [
|
||||||
|
("paperless", "0013_applicationconfiguration_llm_request_timeout"),
|
||||||
|
]
|
||||||
|
|
||||||
|
operations = [
|
||||||
|
migrations.AddField(
|
||||||
|
model_name="applicationconfiguration",
|
||||||
|
name="remote_ocr_api_key",
|
||||||
|
field=models.CharField(
|
||||||
|
blank=True,
|
||||||
|
max_length=1024,
|
||||||
|
null=True,
|
||||||
|
verbose_name="Sets the remote OCR API key",
|
||||||
|
),
|
||||||
|
),
|
||||||
|
migrations.AddField(
|
||||||
|
model_name="applicationconfiguration",
|
||||||
|
name="remote_ocr_endpoint",
|
||||||
|
field=models.CharField(
|
||||||
|
blank=True,
|
||||||
|
max_length=256,
|
||||||
|
null=True,
|
||||||
|
verbose_name="Sets the remote OCR endpoint",
|
||||||
|
),
|
||||||
|
),
|
||||||
|
migrations.AddField(
|
||||||
|
model_name="applicationconfiguration",
|
||||||
|
name="remote_ocr_engine",
|
||||||
|
field=models.CharField(
|
||||||
|
blank=True,
|
||||||
|
choices=[("azureai", "Azure AI Document Intelligence")],
|
||||||
|
max_length=32,
|
||||||
|
null=True,
|
||||||
|
verbose_name="Sets the remote OCR engine",
|
||||||
|
),
|
||||||
|
),
|
||||||
|
]
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
# Generated by Django 5.2.16 on 2026-08-10 15:43
|
||||||
|
|
||||||
|
from django.db import migrations
|
||||||
|
from django.db import models
|
||||||
|
|
||||||
|
|
||||||
|
class Migration(migrations.Migration):
|
||||||
|
dependencies = [
|
||||||
|
("paperless", "0014_applicationconfiguration_remote_ocr_api_key_and_more"),
|
||||||
|
]
|
||||||
|
|
||||||
|
operations = [
|
||||||
|
migrations.AddField(
|
||||||
|
model_name="applicationconfiguration",
|
||||||
|
name="remote_ocr_mode",
|
||||||
|
field=models.CharField(
|
||||||
|
blank=True,
|
||||||
|
choices=[
|
||||||
|
("always", "All supported documents"),
|
||||||
|
("workflow_only", "Only when a workflow enables it"),
|
||||||
|
],
|
||||||
|
max_length=32,
|
||||||
|
null=True,
|
||||||
|
verbose_name="Sets which documents are sent to the remote OCR engine",
|
||||||
|
),
|
||||||
|
),
|
||||||
|
]
|
||||||
@@ -74,6 +74,23 @@ class ColorConvertChoices(models.TextChoices):
|
|||||||
CMYK = ("CMYK", _("CMYK"))
|
CMYK = ("CMYK", _("CMYK"))
|
||||||
|
|
||||||
|
|
||||||
|
class RemoteOCREngine(models.TextChoices):
|
||||||
|
"""
|
||||||
|
Matches to PAPERLESS_REMOTE_OCR_ENGINE
|
||||||
|
"""
|
||||||
|
|
||||||
|
AZURE_AI = ("azureai", _("Azure AI Document Intelligence"))
|
||||||
|
|
||||||
|
|
||||||
|
class RemoteOCRMode(models.TextChoices):
|
||||||
|
"""
|
||||||
|
Matches to PAPERLESS_REMOTE_OCR_MODE
|
||||||
|
"""
|
||||||
|
|
||||||
|
ALWAYS = ("always", _("All supported documents"))
|
||||||
|
WORKFLOW_ONLY = ("workflow_only", _("Only when a workflow enables it"))
|
||||||
|
|
||||||
|
|
||||||
class LLMEmbeddingBackend(models.TextChoices):
|
class LLMEmbeddingBackend(models.TextChoices):
|
||||||
OPENAI_LIKE = ("openai-like", _("OpenAI-compatible"))
|
OPENAI_LIKE = ("openai-like", _("OpenAI-compatible"))
|
||||||
HUGGINGFACE = ("huggingface", _("Huggingface"))
|
HUGGINGFACE = ("huggingface", _("Huggingface"))
|
||||||
@@ -286,6 +303,44 @@ class ApplicationConfiguration(AbstractSingletonModel):
|
|||||||
null=True,
|
null=True,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
"""
|
||||||
|
Settings for the remote OCR parser
|
||||||
|
"""
|
||||||
|
|
||||||
|
# PAPERLESS_REMOTE_OCR_ENGINE
|
||||||
|
remote_ocr_engine = models.CharField(
|
||||||
|
verbose_name=_("Sets the remote OCR engine"),
|
||||||
|
blank=True,
|
||||||
|
null=True,
|
||||||
|
max_length=32,
|
||||||
|
choices=RemoteOCREngine.choices,
|
||||||
|
)
|
||||||
|
|
||||||
|
# PAPERLESS_REMOTE_OCR_API_KEY
|
||||||
|
remote_ocr_api_key = models.CharField(
|
||||||
|
verbose_name=_("Sets the remote OCR API key"),
|
||||||
|
blank=True,
|
||||||
|
null=True,
|
||||||
|
max_length=1024,
|
||||||
|
)
|
||||||
|
|
||||||
|
# PAPERLESS_REMOTE_OCR_ENDPOINT
|
||||||
|
remote_ocr_endpoint = models.CharField(
|
||||||
|
verbose_name=_("Sets the remote OCR endpoint"),
|
||||||
|
blank=True,
|
||||||
|
null=True,
|
||||||
|
max_length=256,
|
||||||
|
)
|
||||||
|
|
||||||
|
# PAPERLESS_REMOTE_OCR_MODE
|
||||||
|
remote_ocr_mode = models.CharField(
|
||||||
|
verbose_name=_("Sets which documents are sent to the remote OCR engine"),
|
||||||
|
blank=True,
|
||||||
|
null=True,
|
||||||
|
max_length=32,
|
||||||
|
choices=RemoteOCRMode.choices,
|
||||||
|
)
|
||||||
|
|
||||||
"""
|
"""
|
||||||
AI related settings
|
AI related settings
|
||||||
"""
|
"""
|
||||||
|
|||||||
@@ -134,6 +134,11 @@ class ParserProtocol(Protocol):
|
|||||||
Author or organisation name.
|
Author or organisation name.
|
||||||
url : str
|
url : str
|
||||||
URL for documentation, source code, or issue tracker.
|
URL for documentation, source code, or issue tracker.
|
||||||
|
|
||||||
|
Parsers that send document content to a remote service should additionally
|
||||||
|
set ``uses_remote_service = True`` so the registry can exclude them when
|
||||||
|
remote processing has not been requested for a document. The attribute is
|
||||||
|
optional so a parser that omits it is treated as fully local.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
@@ -145,6 +150,10 @@ class ParserProtocol(Protocol):
|
|||||||
author: str
|
author: str
|
||||||
url: str
|
url: str
|
||||||
|
|
||||||
|
# NOTE: uses_remote_service is not declared here, the registry reads it
|
||||||
|
# with getattr(cls, ..., False) for backwards-compatibility with existing
|
||||||
|
# parsers
|
||||||
|
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
# Class methods
|
# Class methods
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
|
|||||||
@@ -334,6 +334,8 @@ class ParserRegistry:
|
|||||||
mime_type: str,
|
mime_type: str,
|
||||||
filename: str,
|
filename: str,
|
||||||
path: Path | None = None,
|
path: Path | None = None,
|
||||||
|
*,
|
||||||
|
allow_remote: bool = True,
|
||||||
) -> type[ParserProtocol] | None:
|
) -> type[ParserProtocol] | None:
|
||||||
"""Return the best parser class for the given file, or None.
|
"""Return the best parser class for the given file, or None.
|
||||||
|
|
||||||
@@ -359,6 +361,11 @@ class ParserRegistry:
|
|||||||
path:
|
path:
|
||||||
Optional filesystem path to the file. Forwarded to each
|
Optional filesystem path to the file. Forwarded to each
|
||||||
parser's score method.
|
parser's score method.
|
||||||
|
allow_remote:
|
||||||
|
When False, parsers that declare ``uses_remote_service = True``
|
||||||
|
are excluded from consideration, so a document is never sent to
|
||||||
|
a remote service. Parsers that do not declare the attribute
|
||||||
|
are treated as local and are always considered.
|
||||||
|
|
||||||
Returns
|
Returns
|
||||||
-------
|
-------
|
||||||
@@ -374,6 +381,13 @@ class ParserRegistry:
|
|||||||
if mime_type not in parser_class.supported_mime_types():
|
if mime_type not in parser_class.supported_mime_types():
|
||||||
continue
|
continue
|
||||||
|
|
||||||
|
if not allow_remote and getattr(
|
||||||
|
parser_class,
|
||||||
|
"uses_remote_service",
|
||||||
|
False,
|
||||||
|
):
|
||||||
|
continue
|
||||||
|
|
||||||
score = parser_class.score(mime_type, filename, path)
|
score = parser_class.score(mime_type, filename, path)
|
||||||
if score is None:
|
if score is None:
|
||||||
continue
|
continue
|
||||||
|
|||||||
@@ -61,6 +61,18 @@ class RemoteEngineConfig:
|
|||||||
self.api_key = api_key
|
self.api_key = api_key
|
||||||
self.endpoint = endpoint
|
self.endpoint = endpoint
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def from_app_config(cls) -> Self:
|
||||||
|
"""Build the config from the app config, falling back to the env."""
|
||||||
|
from paperless.config import RemoteOCRConfig
|
||||||
|
|
||||||
|
app_config = RemoteOCRConfig()
|
||||||
|
return cls(
|
||||||
|
engine=app_config.remote_ocr_engine,
|
||||||
|
api_key=app_config.remote_ocr_api_key,
|
||||||
|
endpoint=app_config.remote_ocr_endpoint,
|
||||||
|
)
|
||||||
|
|
||||||
def engine_is_valid(self) -> bool:
|
def engine_is_valid(self) -> bool:
|
||||||
"""Return True when the engine is known and fully configured."""
|
"""Return True when the engine is known and fully configured."""
|
||||||
return (
|
return (
|
||||||
@@ -90,6 +102,9 @@ class RemoteDocumentParser:
|
|||||||
Maintainer name.
|
Maintainer name.
|
||||||
url : str
|
url : str
|
||||||
Issue tracker / source URL.
|
Issue tracker / source URL.
|
||||||
|
uses_remote_service : bool
|
||||||
|
Content is sent to a remote service, True so that the registry
|
||||||
|
can skip this parser if remote processing was not requested.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
name: str = "Paperless-ngx Remote OCR Parser"
|
name: str = "Paperless-ngx Remote OCR Parser"
|
||||||
@@ -97,6 +112,8 @@ class RemoteDocumentParser:
|
|||||||
author: str = "Paperless-ngx Contributors"
|
author: str = "Paperless-ngx Contributors"
|
||||||
url: str = "https://github.com/paperless-ngx/paperless-ngx"
|
url: str = "https://github.com/paperless-ngx/paperless-ngx"
|
||||||
|
|
||||||
|
uses_remote_service: bool = True
|
||||||
|
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
# Class methods
|
# Class methods
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
@@ -145,11 +162,7 @@ class RemoteDocumentParser:
|
|||||||
20 when the remote engine is configured and the MIME type is
|
20 when the remote engine is configured and the MIME type is
|
||||||
supported, otherwise None.
|
supported, otherwise None.
|
||||||
"""
|
"""
|
||||||
config = RemoteEngineConfig(
|
config = RemoteEngineConfig.from_app_config()
|
||||||
engine=settings.REMOTE_OCR_ENGINE,
|
|
||||||
api_key=settings.REMOTE_OCR_API_KEY,
|
|
||||||
endpoint=settings.REMOTE_OCR_ENDPOINT,
|
|
||||||
)
|
|
||||||
if not config.engine_is_valid():
|
if not config.engine_is_valid():
|
||||||
return None
|
return None
|
||||||
if mime_type not in _SUPPORTED_MIME_TYPES:
|
if mime_type not in _SUPPORTED_MIME_TYPES:
|
||||||
@@ -244,11 +257,7 @@ class RemoteDocumentParser:
|
|||||||
Whether an archive copy is wanted. For PDFs, False skips the
|
Whether an archive copy is wanted. For PDFs, False skips the
|
||||||
remote engine and uses locally-extracted text instead.
|
remote engine and uses locally-extracted text instead.
|
||||||
"""
|
"""
|
||||||
config = RemoteEngineConfig(
|
config = RemoteEngineConfig.from_app_config()
|
||||||
engine=settings.REMOTE_OCR_ENGINE,
|
|
||||||
api_key=settings.REMOTE_OCR_API_KEY,
|
|
||||||
endpoint=settings.REMOTE_OCR_ENDPOINT,
|
|
||||||
)
|
|
||||||
|
|
||||||
if not config.engine_is_valid():
|
if not config.engine_is_valid():
|
||||||
logger.warning(
|
logger.warning(
|
||||||
|
|||||||
@@ -219,6 +219,13 @@ class ApplicationConfigurationSerializer(
|
|||||||
allow_null=True,
|
allow_null=True,
|
||||||
max_length=1024,
|
max_length=1024,
|
||||||
)
|
)
|
||||||
|
remote_ocr_api_key = ObfuscatedPasswordField(
|
||||||
|
required=False,
|
||||||
|
allow_null=True,
|
||||||
|
max_length=1024,
|
||||||
|
)
|
||||||
|
|
||||||
|
OBFUSCATED_FIELDS = ("llm_api_key", "remote_ocr_api_key")
|
||||||
|
|
||||||
def run_validation(self, data):
|
def run_validation(self, data):
|
||||||
# Empty strings treated as None to avoid unexpected behavior
|
# Empty strings treated as None to avoid unexpected behavior
|
||||||
@@ -230,11 +237,13 @@ class ApplicationConfigurationSerializer(
|
|||||||
data["language"] = None
|
data["language"] = None
|
||||||
if "llm_output_language" in data and data["llm_output_language"] == "":
|
if "llm_output_language" in data and data["llm_output_language"] == "":
|
||||||
data["llm_output_language"] = None
|
data["llm_output_language"] = None
|
||||||
if "llm_api_key" in data and data["llm_api_key"] is not None:
|
for field in self.OBFUSCATED_FIELDS:
|
||||||
if data["llm_api_key"] == "":
|
if field in data and data[field] is not None:
|
||||||
data["llm_api_key"] = None
|
if data[field] == "":
|
||||||
elif len(data["llm_api_key"].replace("*", "")) == 0:
|
data[field] = None
|
||||||
del data["llm_api_key"]
|
# Not a real value, don't overwrite the stored one
|
||||||
|
elif len(data[field].replace("*", "")) == 0:
|
||||||
|
del data[field]
|
||||||
return super().run_validation(data)
|
return super().run_validation(data)
|
||||||
|
|
||||||
def update(self, instance, validated_data):
|
def update(self, instance, validated_data):
|
||||||
|
|||||||
@@ -1197,6 +1197,7 @@ WEBHOOKS_ALLOW_INTERNAL_REQUESTS = get_bool_from_env(
|
|||||||
REMOTE_OCR_ENGINE = os.getenv("PAPERLESS_REMOTE_OCR_ENGINE")
|
REMOTE_OCR_ENGINE = os.getenv("PAPERLESS_REMOTE_OCR_ENGINE")
|
||||||
REMOTE_OCR_API_KEY = os.getenv("PAPERLESS_REMOTE_OCR_API_KEY")
|
REMOTE_OCR_API_KEY = os.getenv("PAPERLESS_REMOTE_OCR_API_KEY")
|
||||||
REMOTE_OCR_ENDPOINT = os.getenv("PAPERLESS_REMOTE_OCR_ENDPOINT")
|
REMOTE_OCR_ENDPOINT = os.getenv("PAPERLESS_REMOTE_OCR_ENDPOINT")
|
||||||
|
REMOTE_OCR_MODE = os.getenv("PAPERLESS_REMOTE_OCR_MODE", "always")
|
||||||
|
|
||||||
################################################################################
|
################################################################################
|
||||||
# AI Settings #
|
# AI Settings #
|
||||||
|
|||||||
@@ -21,6 +21,7 @@ from unittest.mock import Mock
|
|||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
from documents.parsers import ParseError
|
from documents.parsers import ParseError
|
||||||
|
from paperless.models import ApplicationConfiguration
|
||||||
from paperless.parsers import ParserContext
|
from paperless.parsers import ParserContext
|
||||||
from paperless.parsers import ParserProtocol
|
from paperless.parsers import ParserProtocol
|
||||||
from paperless.parsers.remote import RemoteDocumentParser
|
from paperless.parsers.remote import RemoteDocumentParser
|
||||||
@@ -33,6 +34,10 @@ if TYPE_CHECKING:
|
|||||||
from pytest_mock import MockerFixture
|
from pytest_mock import MockerFixture
|
||||||
|
|
||||||
|
|
||||||
|
# Remote ocr config from ApplicationConfiguration needs DB access
|
||||||
|
pytestmark = pytest.mark.django_db
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Module-local fixtures
|
# Module-local fixtures
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
@@ -227,6 +232,18 @@ class TestRemoteParserScore:
|
|||||||
score = RemoteDocumentParser.score("application/pdf", "doc.pdf")
|
score = RemoteDocumentParser.score("application/pdf", "doc.pdf")
|
||||||
assert score is not None and score > 10
|
assert score is not None and score > 10
|
||||||
|
|
||||||
|
@pytest.mark.usefixtures("no_engine_settings")
|
||||||
|
def test_score_uses_app_config_when_env_unset(self) -> None:
|
||||||
|
"""The app config alone is enough to activate the parser."""
|
||||||
|
config = ApplicationConfiguration.objects.first()
|
||||||
|
assert config is not None
|
||||||
|
config.remote_ocr_engine = "azureai"
|
||||||
|
config.remote_ocr_api_key = "app-config-key"
|
||||||
|
config.remote_ocr_endpoint = "https://config.cognitiveservices.azure.com"
|
||||||
|
config.save()
|
||||||
|
|
||||||
|
assert RemoteDocumentParser.score("application/pdf", "doc.pdf") == 20
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Properties
|
# Properties
|
||||||
|
|||||||
@@ -1277,6 +1277,8 @@ class TestParserFileTypes:
|
|||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
|
||||||
|
# Remote ocr config from ApplicationConfiguration needs DB access
|
||||||
|
@pytest.mark.django_db
|
||||||
class TestRasterisedDocumentParserRegistry:
|
class TestRasterisedDocumentParserRegistry:
|
||||||
def test_registered_in_defaults(self) -> None:
|
def test_registered_in_defaults(self) -> None:
|
||||||
from paperless.parsers.registry import ParserRegistry
|
from paperless.parsers.registry import ParserRegistry
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ from paperless.checks import audit_log_check
|
|||||||
from paperless.checks import binaries_check
|
from paperless.checks import binaries_check
|
||||||
from paperless.checks import check_default_language_available
|
from paperless.checks import check_default_language_available
|
||||||
from paperless.checks import check_deprecated_db_settings
|
from paperless.checks import check_deprecated_db_settings
|
||||||
from paperless.checks import check_remote_parser_configured
|
from paperless.checks import check_remote_ocr_mode
|
||||||
from paperless.checks import check_v3_minimum_upgrade_version
|
from paperless.checks import check_v3_minimum_upgrade_version
|
||||||
from paperless.checks import debug_mode_check
|
from paperless.checks import debug_mode_check
|
||||||
from paperless.checks import paths_check
|
from paperless.checks import paths_check
|
||||||
@@ -631,29 +631,21 @@ class TestV3MinimumUpgradeVersionCheck:
|
|||||||
assert check_v3_minimum_upgrade_version(None) == []
|
assert check_v3_minimum_upgrade_version(None) == []
|
||||||
|
|
||||||
|
|
||||||
class TestRemoteParserChecks:
|
class TestRemoteOCRModeCheck:
|
||||||
def test_no_engine(self, settings: SettingsWrapper) -> None:
|
def test_valid_mode(self, settings: SettingsWrapper) -> None:
|
||||||
settings.REMOTE_OCR_ENGINE = None
|
settings.REMOTE_OCR_MODE = "workflow_only"
|
||||||
msgs = check_remote_parser_configured(None)
|
|
||||||
|
msgs = check_remote_ocr_mode(None)
|
||||||
|
|
||||||
assert len(msgs) == 0
|
assert len(msgs) == 0
|
||||||
|
|
||||||
def test_azure_no_endpoint(self, settings: SettingsWrapper) -> None:
|
def test_invalid_mode(self, settings: SettingsWrapper) -> None:
|
||||||
|
settings.REMOTE_OCR_MODE = "sometimes"
|
||||||
|
|
||||||
settings.REMOTE_OCR_ENGINE = "azureai"
|
msgs = check_remote_ocr_mode(None)
|
||||||
settings.REMOTE_OCR_API_KEY = "somekey"
|
|
||||||
settings.REMOTE_OCR_ENDPOINT = None
|
|
||||||
|
|
||||||
msgs = check_remote_parser_configured(None)
|
|
||||||
|
|
||||||
assert len(msgs) == 1
|
assert len(msgs) == 1
|
||||||
|
assert "PAPERLESS_REMOTE_OCR_MODE is set to 'sometimes'" in msgs[0].msg
|
||||||
msg = msgs[0]
|
|
||||||
|
|
||||||
assert (
|
|
||||||
"Azure AI remote parser requires endpoint and API key to be configured."
|
|
||||||
in msg.msg
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
class TestTesseractChecks:
|
class TestTesseractChecks:
|
||||||
|
|||||||
@@ -468,6 +468,124 @@ class TestParserRegistryGetParserForFile:
|
|||||||
assert result is AcceptingBuiltin
|
assert result is AcceptingBuiltin
|
||||||
|
|
||||||
|
|
||||||
|
class TestParserRegistryRemoteParsers:
|
||||||
|
"""Verify the allow_remote filter in ParserRegistry.get_parser_for_file()."""
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _remote_parser_cls() -> type:
|
||||||
|
class RemoteParser:
|
||||||
|
name = "remote"
|
||||||
|
version = "1.0"
|
||||||
|
author = "A"
|
||||||
|
url = "https://example.com/remote"
|
||||||
|
uses_remote_service = True
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def supported_mime_types(cls):
|
||||||
|
return {"text/plain": ".txt"}
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def score(cls, mime_type, filename, path=None):
|
||||||
|
return 20
|
||||||
|
|
||||||
|
return RemoteParser
|
||||||
|
|
||||||
|
def test_remote_parser_wins_when_remote_allowed(
|
||||||
|
self,
|
||||||
|
dummy_parser_cls: type,
|
||||||
|
) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: A remote parser scoring 20 and a local parser scoring 10.
|
||||||
|
WHEN: get_parser_for_file() is called with allow_remote=True.
|
||||||
|
THEN: The remote parser is returned.
|
||||||
|
"""
|
||||||
|
remote_parser_cls = self._remote_parser_cls()
|
||||||
|
registry = ParserRegistry()
|
||||||
|
registry.register_builtin(dummy_parser_cls)
|
||||||
|
registry.register_builtin(remote_parser_cls)
|
||||||
|
|
||||||
|
result = registry.get_parser_for_file(
|
||||||
|
"text/plain",
|
||||||
|
"readme.txt",
|
||||||
|
allow_remote=True,
|
||||||
|
)
|
||||||
|
assert result is remote_parser_cls
|
||||||
|
|
||||||
|
def test_remote_parser_skipped_when_remote_not_allowed(
|
||||||
|
self,
|
||||||
|
dummy_parser_cls: type,
|
||||||
|
) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: A remote parser scoring 20 and a local parser scoring 10.
|
||||||
|
WHEN: get_parser_for_file() is called with allow_remote=False.
|
||||||
|
THEN: The local parser is returned despite its lower score.
|
||||||
|
"""
|
||||||
|
registry = ParserRegistry()
|
||||||
|
registry.register_builtin(dummy_parser_cls)
|
||||||
|
registry.register_builtin(self._remote_parser_cls())
|
||||||
|
|
||||||
|
result = registry.get_parser_for_file(
|
||||||
|
"text/plain",
|
||||||
|
"readme.txt",
|
||||||
|
allow_remote=False,
|
||||||
|
)
|
||||||
|
assert result is dummy_parser_cls
|
||||||
|
|
||||||
|
def test_no_parser_when_only_remote_available_and_not_allowed(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: A registry whose only candidate declares uses_remote_service.
|
||||||
|
WHEN: get_parser_for_file() is called with allow_remote=False.
|
||||||
|
THEN: None is returned — the remote parser is never used as a
|
||||||
|
fallback when remote processing was not requested.
|
||||||
|
"""
|
||||||
|
registry = ParserRegistry()
|
||||||
|
registry.register_builtin(self._remote_parser_cls())
|
||||||
|
|
||||||
|
result = registry.get_parser_for_file(
|
||||||
|
"text/plain",
|
||||||
|
"readme.txt",
|
||||||
|
allow_remote=False,
|
||||||
|
)
|
||||||
|
assert result is None
|
||||||
|
|
||||||
|
def test_parser_without_attribute_treated_as_local(
|
||||||
|
self,
|
||||||
|
dummy_parser_cls: type,
|
||||||
|
) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: A third-party parser predating uses_remote_service, so it does
|
||||||
|
not declare the attribute at all.
|
||||||
|
WHEN: get_parser_for_file() is called with allow_remote=False.
|
||||||
|
THEN: It is still considered, i.e. treated as fully local, rather
|
||||||
|
than raising AttributeError.
|
||||||
|
"""
|
||||||
|
assert not hasattr(dummy_parser_cls, "uses_remote_service")
|
||||||
|
|
||||||
|
registry = ParserRegistry()
|
||||||
|
registry.register_builtin(dummy_parser_cls)
|
||||||
|
|
||||||
|
result = registry.get_parser_for_file(
|
||||||
|
"text/plain",
|
||||||
|
"readme.txt",
|
||||||
|
allow_remote=False,
|
||||||
|
)
|
||||||
|
assert result is dummy_parser_cls
|
||||||
|
|
||||||
|
def test_remote_allowed_by_default(self) -> None:
|
||||||
|
"""
|
||||||
|
GIVEN: A registry containing only a remote parser.
|
||||||
|
WHEN: get_parser_for_file() is called without allow_remote.
|
||||||
|
THEN: The remote parser is returned — callers that do not opt in to
|
||||||
|
the filter keep the previous behaviour.
|
||||||
|
"""
|
||||||
|
remote_parser_cls = self._remote_parser_cls()
|
||||||
|
registry = ParserRegistry()
|
||||||
|
registry.register_builtin(remote_parser_cls)
|
||||||
|
|
||||||
|
result = registry.get_parser_for_file("text/plain", "readme.txt")
|
||||||
|
assert result is remote_parser_cls
|
||||||
|
|
||||||
|
|
||||||
class TestDiscover:
|
class TestDiscover:
|
||||||
"""Verify entrypoint discovery in ParserRegistry.discover()."""
|
"""Verify entrypoint discovery in ParserRegistry.discover()."""
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,113 @@
|
|||||||
|
"""Tests for RemoteOCRConfig precedence between app config and Django settings."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import TYPE_CHECKING
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
from django.test import override_settings
|
||||||
|
|
||||||
|
from paperless.config import RemoteOCRConfig
|
||||||
|
from paperless.models import RemoteOCRMode
|
||||||
|
|
||||||
|
if TYPE_CHECKING:
|
||||||
|
from unittest.mock import MagicMock
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def null_app_config(mocker) -> MagicMock:
|
||||||
|
"""Mock ApplicationConfiguration with all fields None → falls back to Django settings."""
|
||||||
|
return mocker.MagicMock(
|
||||||
|
remote_ocr_engine=None,
|
||||||
|
remote_ocr_api_key=None,
|
||||||
|
remote_ocr_endpoint=None,
|
||||||
|
remote_ocr_mode=None,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture()
|
||||||
|
def make_remote_ocr_config(mocker):
|
||||||
|
def _make(app_config, **django_settings_overrides):
|
||||||
|
mocker.patch(
|
||||||
|
"paperless.config.BaseConfig._get_config_instance",
|
||||||
|
return_value=app_config,
|
||||||
|
)
|
||||||
|
with override_settings(**django_settings_overrides):
|
||||||
|
return RemoteOCRConfig()
|
||||||
|
|
||||||
|
return _make
|
||||||
|
|
||||||
|
|
||||||
|
class TestRemoteOCRConfig:
|
||||||
|
def test_falls_back_to_settings(
|
||||||
|
self,
|
||||||
|
make_remote_ocr_config,
|
||||||
|
null_app_config,
|
||||||
|
) -> None:
|
||||||
|
cfg = make_remote_ocr_config(
|
||||||
|
null_app_config,
|
||||||
|
REMOTE_OCR_ENGINE="azureai",
|
||||||
|
REMOTE_OCR_API_KEY="env-key",
|
||||||
|
REMOTE_OCR_ENDPOINT="https://env.cognitiveservices.azure.com",
|
||||||
|
REMOTE_OCR_MODE=RemoteOCRMode.WORKFLOW_ONLY,
|
||||||
|
)
|
||||||
|
assert cfg.remote_ocr_engine == "azureai"
|
||||||
|
assert cfg.remote_ocr_api_key == "env-key"
|
||||||
|
assert cfg.remote_ocr_endpoint == "https://env.cognitiveservices.azure.com"
|
||||||
|
assert cfg.remote_ocr_mode == RemoteOCRMode.WORKFLOW_ONLY
|
||||||
|
|
||||||
|
def test_app_config_takes_precedence(
|
||||||
|
self,
|
||||||
|
make_remote_ocr_config,
|
||||||
|
mocker,
|
||||||
|
) -> None:
|
||||||
|
app_config = mocker.MagicMock(
|
||||||
|
remote_ocr_engine="azureai",
|
||||||
|
remote_ocr_api_key="db-key",
|
||||||
|
remote_ocr_endpoint="https://db.cognitiveservices.azure.com",
|
||||||
|
remote_ocr_mode=RemoteOCRMode.WORKFLOW_ONLY,
|
||||||
|
)
|
||||||
|
cfg = make_remote_ocr_config(
|
||||||
|
app_config,
|
||||||
|
REMOTE_OCR_ENGINE=None,
|
||||||
|
REMOTE_OCR_API_KEY="env-key",
|
||||||
|
REMOTE_OCR_ENDPOINT="https://env.cognitiveservices.azure.com",
|
||||||
|
REMOTE_OCR_MODE=RemoteOCRMode.ALWAYS,
|
||||||
|
)
|
||||||
|
assert cfg.remote_ocr_engine == "azureai"
|
||||||
|
assert cfg.remote_ocr_api_key == "db-key"
|
||||||
|
assert cfg.remote_ocr_endpoint == "https://db.cognitiveservices.azure.com"
|
||||||
|
assert cfg.remote_ocr_mode == RemoteOCRMode.WORKFLOW_ONLY
|
||||||
|
|
||||||
|
def test_unset_everywhere(
|
||||||
|
self,
|
||||||
|
make_remote_ocr_config,
|
||||||
|
null_app_config,
|
||||||
|
) -> None:
|
||||||
|
cfg = make_remote_ocr_config(
|
||||||
|
null_app_config,
|
||||||
|
REMOTE_OCR_ENGINE=None,
|
||||||
|
REMOTE_OCR_API_KEY=None,
|
||||||
|
REMOTE_OCR_ENDPOINT=None,
|
||||||
|
)
|
||||||
|
assert cfg.remote_ocr_engine is None
|
||||||
|
assert cfg.remote_ocr_api_key is None
|
||||||
|
assert cfg.remote_ocr_endpoint is None
|
||||||
|
|
||||||
|
|
||||||
|
class TestRemoteOCRByDefault:
|
||||||
|
def test_always_mode(self, make_remote_ocr_config, null_app_config) -> None:
|
||||||
|
cfg = make_remote_ocr_config(
|
||||||
|
null_app_config,
|
||||||
|
REMOTE_OCR_MODE=RemoteOCRMode.ALWAYS,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert cfg.remote_ocr_by_default is True
|
||||||
|
|
||||||
|
def test_workflow_only_mode(self, make_remote_ocr_config, null_app_config) -> None:
|
||||||
|
cfg = make_remote_ocr_config(
|
||||||
|
null_app_config,
|
||||||
|
REMOTE_OCR_MODE=RemoteOCRMode.WORKFLOW_ONLY,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert cfg.remote_ocr_by_default is False
|
||||||
Reference in New Issue
Block a user