Feature: store barcode contents, list and search them (#14276)

* Feature: store barcode contents, list and search them

New setting PAPERLESS_CONSUMER_STORE_BARCODE_VALUES (off by default)
stores all barcodes found during consumption with the document: page,
type and content. They are listed on the metadata tab with a copy
button, returned by the documents API and searchable with barcodes:
in the advanced search. Versions keep their own barcodes, reprocessing
reads them again. Refs #9898

* Tests: cover the remaining barcode branches

Covers unchanged and failing barcode reads on reprocessing, unsupported files and DocumentBarcode.__str__, and uses toHaveLength in the barcode list spec as suggested by SonarCloud.

* Address review: keep barcodes when they can't be read, format choices

- Reprocessing and new versions keep the stored barcodes when the file
  can't be scanned or the scan fails, and replace them atomically.
- Format is a TextChoices of the zxing-cpp formats, with a test.
- Shared scan code, latest_version helper, TypedDict, serializer reuse.
- Barcodes in the split manifest, export/import tests, pytest-style tests.

* Barcode tests: TIFF reprocess, format check both ways, fixtures

- Reprocessing a TIFF with TIFF support off keeps the stored barcodes.
- The format test also fails when zxing-cpp drops a format.
- Fixtures in place of the sample dir mixin, plugin-level disabled test.
- Format labels aren't translated, OpenAPI enum named BarcodeFormatEnum.

* Review: module-level zxing reader, shorter barcode docs

- read_barcodes_zxing is a module-level function used by scan_pdf.
- Drop the trivial __str__ test, mark it no cover.
- Shorten the barcode docs and remove the duplicate in configuration.md.

* Fix header

* Use utility class

* Return barcodes from the metadata endpoint only

Drop the barcodes field and its prefetches from the document serializer,
as agreed in the review. Also remove the now empty component stylesheet.

---------

Co-authored-by: shamoon <4887959+shamoon@users.noreply.github.com>
This commit is contained in:
jurassicparkicecreamandshamoon authored and GitHub committed 2026-10-05 16:25:19 +00:00
1 parent 312684aef3
commit 62cc31fbdb
37 files changed
+1196 -80

No files matched your search

@@ -8,7 +8,7 @@ queryable-but-always-empty -- syntactically valid, silently matching
nothing -- with no test failure anywhere.
This indexes one real document carrying values for every JSON field
(a Note, a CustomFieldInstance) and inspects the document's own stored
(a Note, a CustomFieldInstance, a DocumentBarcode) and inspects the document's own stored
JSON payload, rather than running field-specific queries: that way a
future JSON field's subpaths are covered automatically, without a new
per-subpath query having to be added by hand each time.
@@ -27,6 +27,7 @@ from documents.models import CustomFieldInstance
from documents.models import Document
from documents.models import Note
from documents.search._fields import PUBLIC_FIELDS
from paperless_testing.factories import DocumentBarcodeFactory
from paperless_testing.factories import UserFactory
if TYPE_CHECKING:
@@ -42,11 +43,12 @@ class TestJsonSubpathsAreWrittenAtIndexTime:
) -> None:
"""
GIVEN:
- A document with a Note and a CustomFieldInstance attached
- A document with a Note, a CustomFieldInstance and a
DocumentBarcode attached
WHEN:
- The document is indexed via TantivyBackend.add_or_update
THEN:
- Every subpath PUBLIC_FIELDS declares for notes/custom_fields
- Every subpath PUBLIC_FIELDS declares for notes/custom_fields/barcodes
is present as a key in the document's stored JSON payload
"""
user = UserFactory(username="completeness-user")
@@ -65,6 +67,7 @@ class TestJsonSubpathsAreWrittenAtIndexTime:
field=field,
value_text="a value",
)
DocumentBarcodeFactory(document=doc, value="a barcode")
backend.add_or_update(doc)
index = backend._index