Feature: store barcode contents, list and search them (#14276)

* Feature: store barcode contents, list and search them

New setting PAPERLESS_CONSUMER_STORE_BARCODE_VALUES (off by default)
stores all barcodes found during consumption with the document: page,
type and content. They are listed on the metadata tab with a copy
button, returned by the documents API and searchable with barcodes:
in the advanced search. Versions keep their own barcodes, reprocessing
reads them again. Refs #9898

* Tests: cover the remaining barcode branches

Covers unchanged and failing barcode reads on reprocessing, unsupported files and DocumentBarcode.__str__, and uses toHaveLength in the barcode list spec as suggested by SonarCloud.

* Address review: keep barcodes when they can't be read, format choices

- Reprocessing and new versions keep the stored barcodes when the file
  can't be scanned or the scan fails, and replace them atomically.
- Format is a TextChoices of the zxing-cpp formats, with a test.
- Shared scan code, latest_version helper, TypedDict, serializer reuse.
- Barcodes in the split manifest, export/import tests, pytest-style tests.

* Barcode tests: TIFF reprocess, format check both ways, fixtures

- Reprocessing a TIFF with TIFF support off keeps the stored barcodes.
- The format test also fails when zxing-cpp drops a format.
- Fixtures in place of the sample dir mixin, plugin-level disabled test.
- Format labels aren't translated, OpenAPI enum named BarcodeFormatEnum.

* Review: module-level zxing reader, shorter barcode docs

- read_barcodes_zxing is a module-level function used by scan_pdf.
- Drop the trivial __str__ test, mark it no cover.
- Shorten the barcode docs and remove the duplicate in configuration.md.

* Fix header

* Use utility class

* Return barcodes from the metadata endpoint only

Drop the barcodes field and its prefetches from the document serializer,
as agreed in the review. Also remove the now empty component stylesheet.

---------

Co-authored-by: shamoon <4887959+shamoon@users.noreply.github.com>
This commit is contained in:
jurassicparkicecreamandshamoon authored and GitHub committed 2026-10-05 16:25:19 +00:00
1 parent 312684aef3
commit 62cc31fbdb
37 files changed
+1196 -80

No files matched your search

+17 -1
View File
@@ -311,7 +311,13 @@ class WriteBatch:
queryset = annotate_effective_content(
Document.objects.filter(pk__in=ids)
.select_related("correspondent", "document_type", "storage_path", "owner")
.prefetch_related("tags", "notes__user", "custom_fields__field"),
.prefetch_related(
"tags",
"notes__user",
"custom_fields__field",
"barcodes",
"versions__barcodes",
),
)
for document, grant in _DocumentViewerStream(queryset, chunk_size=1000):
self.remove(document.pk)
@@ -604,6 +610,16 @@ class TantivyBackend:
},
)
# Barcodes: JSON field like custom_fields, only filled when stored
for barcode in document.get_effective_barcodes():
doc.add_json(
"barcodes",
{
"value": normalize_search_text(barcode.value),
"format": normalize_search_text(barcode.format),
},
)
# Dates
created_date = datetime(
document.created.year,
+5
View File
@@ -39,4 +39,9 @@ PUBLIC_FIELDS: tuple[FieldSpec, ...] = (
FieldKind.JSON,
subpaths={"name": SubpathSpec(), "value": SubpathSpec(default=True)},
),
FieldSpec(
"barcodes",
FieldKind.JSON,
subpaths={"value": SubpathSpec(default=True), "format": SubpathSpec()},
),
)
+2 -1
View File
@@ -25,7 +25,8 @@ logger = logging.getLogger("paperless.search")
# order, and the write-only correspondent/document_type/storage_path/tag id
# columns dropped. tantivy compares schemas by ordered field list, so an
# index built by v1 rejects every write against the v2 schema.
SCHEMA_VERSION: Final[int] = 2
# v3 - barcodes JSON field for stored barcode contents
SCHEMA_VERSION: Final[int] = 3
class FieldDescriptor(NamedTuple):