Merge branch 'dev'

This commit is contained in:
shamoon
2026-10-05 20:33:36 -07:00
294 changed files with 54074 additions and 52966 deletions
+27 -9
View File
@@ -136,13 +136,15 @@ for suggested generation and embedding models.
### AI-assisted suggestions
With AI enabled, Paperless-ngx can suggest a title, tags, correspondent, document type,
storage path and dates by sending the document to the LLM. This is **opt-in per request**
and surfaces through the "Suggest" control on the document detail page, alongside the
classic classifier-based suggestions — it does not disable them. Suggestions are requested
automatically when you open a document that carries an inbox tag unless "Automatically request
suggestions for inbox documents" under Settings > Documents is disabled. Suggestion output
language can be steered with
[`PAPERLESS_AI_LLM_OUTPUT_LANGUAGE`](configuration.md#PAPERLESS_AI_LLM_OUTPUT_LANGUAGE)
storage path and dates by sending the document to the LLM using "Suggest" button on the document
detail page. You can choose which type of suggestions are requested by default under Settings >
Documents, either ML (classifier-based) suggestions, AI suggestions, or both. When both are requested
the results are combined.
Suggestions are requested automatically when you open a document that carries an inbox tag
unless "Automatically request suggestions for inbox documents" under Settings > Documents is disabled.
Suggestion output language can be steered with [`PAPERLESS_AI_LLM_OUTPUT_LANGUAGE`](configuration.md#PAPERLESS_AI_LLM_OUTPUT_LANGUAGE)
(otherwise it follows the user's UI language).
### The LLM index (RAG) and similar documents
@@ -153,8 +155,11 @@ in similar existing documents, and the document chat can retrieve relevant conte
Enable it by setting
[`PAPERLESS_AI_LLM_EMBEDDING_BACKEND`](configuration.md#PAPERLESS_AI_LLM_EMBEDDING_BACKEND)
(`huggingface` for fully-local embeddings, or `ollama` / `openai-like`). The index is only
built when AI is enabled **and** an embedding backend is set.
(`huggingface` for fully-local embeddings, or `ollama` / `openai-like`). By default, the main
LLM API key and endpoint are used, but an optional embedding-specific[API key](configuration.md#PAPERLESS_AI_LLM_EMBEDDING_API_KEY)
and [endpoint](configuration.md#PAPERLESS_AI_LLM_EMBEDDING_ENDPOINT) can be configured.
The index is only built when AI is enabled **and** an embedding backend is set.
The index is updated automatically on a schedule controlled by
[`PAPERLESS_LLM_INDEX_TASK_CRON`](configuration.md#PAPERLESS_LLM_INDEX_TASK_CRON) (daily by
@@ -1005,6 +1010,19 @@ documents to both separate and categorize them in a single operation.
**Example:** A 6-page scan with TAG:invoice on page 3 and TAG:receipt on page 5 will create
three documents: pages 1-2 (no tags), pages 3-4 (tagged "invoice"), and pages 5-6 (tagged "receipt").
### Barcode Contents {#barcode-contents}
By default, Paperless only uses barcodes for splitting, ASNs and tags. With
[`PAPERLESS_CONSUMER_STORE_BARCODE_VALUES`](configuration.md#PAPERLESS_CONSUMER_STORE_BARCODE_VALUES)
enabled, it stores the content of every barcode with the document, e.g. payment codes or QR codes.
- Barcodes are listed on the **Metadata** tab with page, type and content, and can be copied.
- The API returns them in the `barcodes` field of `/api/documents/{id}/metadata/`.
- They can be [searched](usage.md#searching-barcodes), e.g. `barcodes:DE89370400440532013000`.
- Only the first [`PAPERLESS_CONSUMER_BARCODE_MAX_PAGES`](configuration.md#PAPERLESS_CONSUMER_BARCODE_MAX_PAGES)
pages are scanned. Reprocessing reads the barcodes of existing documents.
- Each version keeps its own barcodes, and the newest version's are shown and searched.
## Automatic collation of double-sided documents {#collate}
!!! note
+27
View File
@@ -1796,6 +1796,13 @@ assigns or creates tags if a properly formatted barcode is detected.
Defaults to false.
#### [`PAPERLESS_CONSUMER_STORE_BARCODE_VALUES=<bool>`](#PAPERLESS_CONSUMER_STORE_BARCODE_VALUES) {#PAPERLESS_CONSUMER_STORE_BARCODE_VALUES}
: Stores the content of every barcode found during consumption, see
[Barcode Contents](advanced_usage.md#barcode-contents).
Defaults to false.
## Audit Trail
#### [`PAPERLESS_AUDIT_LOG_ENABLED=<bool>`](#PAPERLESS_AUDIT_LOG_ENABLED) {#PAPERLESS_AUDIT_LOG_ENABLED}
@@ -2133,6 +2140,13 @@ for language and resource considerations.
Defaults to None.
#### [`PAPERLESS_AI_LLM_EMBEDDING_API_KEY=<str>`](#PAPERLESS_AI_LLM_EMBEDDING_API_KEY) {#PAPERLESS_AI_LLM_EMBEDDING_API_KEY}
: The API key to use for the embedding backend. If not supplied, embeddings use
`PAPERLESS_AI_LLM_API_KEY`.
Defaults to None.
#### [`PAPERLESS_AI_LLM_EMBEDDING_ENDPOINT=<str>`](#PAPERLESS_AI_LLM_EMBEDDING_ENDPOINT) {#PAPERLESS_AI_LLM_EMBEDDING_ENDPOINT}
: The endpoint / url to use for the embedding backend. If not supplied, embeddings use
@@ -2217,6 +2231,19 @@ used with the OpenAI-compatible backend to target a custom provider or local gat
Defaults to true, which allows internal endpoints.
#### [`PAPERLESS_AI_LLM_EXTRA_PARAMS=<json>`](#PAPERLESS_AI_LLM_EXTRA_PARAMS) {#PAPERLESS_AI_LLM_EXTRA_PARAMS}
: A JSON object of extra parameters sent with every LLM request, for providers that require a parameter Paperless does not
set itself. Values here override Paperless' own, and no validation is performed. Whatever you put here is passed to the
backend as-is, so an invalid parameter will simply be rejected by your provider. For example, current OpenAI reasoning
models refuse tool calls on the chat completions API unless reasoning is off:
```
PAPERLESS_AI_LLM_EXTRA_PARAMS={"reasoning_effort": "none"}
```
Defaults to empty, which adds nothing to requests.
#### [`PAPERLESS_LLM_INDEX_TASK_CRON=<cron expression>`](#PAPERLESS_LLM_INDEX_TASK_CRON) {#PAPERLESS_LLM_INDEX_TASK_CRON}
: Configures the schedule to update the AI embeddings of text content and metadata for all documents. Only performed if
+1
View File
@@ -150,6 +150,7 @@ pnpm ng build --configuration production
is loaded as well. However, the tests rely on the default
configuration. This is not ideal. But for now, make sure no settings
except for DEBUG are overridden when testing.
- Tests run in a random order each session, so that one test cannot quietly depend on another having run first. The seed is printed at the top of the run; pass `--randomly-seed=<seed>` to replay that exact order, or `--randomly-seed=last` to repeat the previous run.
!!! note
+14
View File
@@ -1060,6 +1060,20 @@ notes.user:alice notes.note:insurance
The bare `notes:` prefix is shorthand for `notes.note:`.
#### Searching barcodes
If [barcode contents are stored](advanced_usage.md#barcode-contents), they can be searched by
content or type, but only with a field name:
```
barcodes.value:DE89370400440532013000
barcodes.format:qrcode
barcodes:wifi barcodes:guest
```
`barcodes:` is shorthand for `barcodes.value:`. Separators are stripped, so each part of e.g.
`WIFI:S:Guest;P:secret;;` can be searched on its own.
All of these can be combined. Syntax not described here may not work as expected, and an unknown field name is searched as ordinary text.
!!! note