Compare commits

..
Author SHA1 Message Date
stumpylog 2a8df9dca0 Chore: Trim thumbnail comments and docstrings
The comments and docstrings around PDF thumbnail generation were wordier than the code they describe. This cuts them back to the essential facts.
2026-10-07 14:56:47 -07:00
stumpylog f1e85fff3e Chore: Simplify the PDF thumbnail helpers and their tests
The qpdf repair step in the thumbnail fallback is now its own helper, so the fallback has a single ParseError handler instead of a nested try. The fallback also stops re-wrapping temp_dir in Path and annotating obvious locals.

Parameters no caller used are removed: the cropbox toggle on the rasterizer, which now always passes -cropbox, and the width and height limits on the WebP encoder, which uses the module constants directly. The comments around the 300 DPI cap are consolidated so the purpose is stated once.

In the tests, the clamp and supersample encoder cases are merged into one parametrized test, the PDF thumbnail tests share a work directory fixture and a blank PDF helper, redundant assertions are dropped, and the tesseract fallback test uses a plain import of the rasterizer. A garbled docstring is rewritten.
2026-10-07 14:51:09 -07:00
stumpylog f549201b7c Chore: Tighten thumbnail supersampling edge cases
Pages so large that the computed DPI hit the floor of 1 were still supersampled to 2 DPI, quadrupling the pixel count pdftoppm had to render for an already oversized page. Supersampling is now skipped at that floor.

The downsample test also did not prove the factor was applied, because the final 500x5000 clamp hid it in most cases. A 900x1200 render now has to come out at 450x600, which only happens when the downsample runs. The comments on the DPI cap and the clamp are reworded to describe the supersampled pipeline accurately.
2026-10-07 14:32:54 -07:00
stumpylog 26efad77ac Chore: Supersample the PDF thumbnail render 2x
Rendering page 1 straight at the thumbnail density left text slightly soft, since pdftoppm antialiases at the final size.

The page is now rendered at twice the computed DPI and downsampled by that factor with Lanczos before the existing 500x5000 clamp, which gives crisper text at about the same file size for a negligible memory cost. When the page geometry cannot be read, the fixed 150 DPI fallback is kept as a single unsupersampled render, because doubling a render that is not bounded by the thumbnail size would quadruple its pixel count. The qpdf repair path shares the same render step and gets the same behavior.
2026-10-07 14:30:50 -07:00
stumpylog 901dd4be93 Chore: Drop system check for deprecated convert variables
A dedicated system check for the two deprecated ImageMagick convert variables is more machinery than a simple deprecation needs.

Remove the check and its tests. The settings removal, the deprecation notes in the configuration docs, and the example configuration cleanup stay as they were.
2026-10-07 13:47:50 -07:00
stumpylog 8f73514fbb Chore: Deprecate PAPERLESS_CONVERT_MEMORY_LIMIT and PAPERLESS_CONVERT_TMPDIR
These two options set ImageMagick's memory limit and scratch directory, but their only reader was the PDF thumbnail conversion helper. PDF thumbnails no longer use ImageMagick, so the settings did nothing while still being documented and advertised in the example configuration.

Remove the unused settings, add system warnings for anyone who still has either variable set, mark both options as deprecated and without effect in the configuration docs, and drop them from paperless.conf.example. The manual setup guide is also corrected so it no longer claims ImageMagick is needed for PDF conversion, and the ImageMagick policy step now only covers hardening.
2026-10-07 13:46:40 -07:00
stumpylog f30a65f440 Fix: Correct ImageMagick PDF policy note in setup docs
The note claimed that steps still relying on ImageMagick, such as TIFF to
PDF conversion, could fail without the PDF policy change. The remaining
convert calls only strip alpha from raster images, and the TIFF to PDF
step uses img2pdf, so the PDF coder policy no longer matters.

The passage now says Paperless-ngx no longer passes PDF documents to
ImageMagick, so enabling PDF processing is not required, while keeping the
policy hardening guidance in place.
2026-10-07 12:59:05 -07:00
stumpylog 2136659e2b Fix: Harden PDF thumbnail encoding and align CI and docs
Pillow raises DecompressionBombError, which is not an OSError, so an
enormous render from the unknown-geometry fallback DPI skipped the
default thumbnail fallback. It is now converted to a ParseError like other
encoding failures, with a test covering both error types.

CI did not install qpdf, which the new thumbnail repair tests need, so it
is added to the backend package list. The setup docs now list poppler-utils
as used for thumbnail generation and no longer claim PDF thumbnails fall
back to Ghostscript when the ImageMagick PDF policy is not enabled.
2026-10-07 12:58:40 -07:00
stumpylog 26baf52a78 Chore: Remove the unused run_convert helper
run_convert wrapped ImageMagick for the PDF thumbnail path, but thumbnails are now produced with pdftoppm and Pillow, and the remaining convert callers invoke the binary directly through run_subprocess. The helper had no callers left, so it and its now unused os import are removed. CONVERT_BINARY is still used by those callers and stays.
2026-10-07 12:54:40 -07:00
stumpylog 3143d1936a Chore: Cover the qpdf repair and double failure thumbnail paths
The existing fallback test fails the first rasterization artificially and only proves the retry plumbing. Nothing showed that a PDF which pdftoppm genuinely cannot read is repaired by qpdf into a real rendered thumbnail, and nothing covered the case where both the render and the repair fail.

Add a test which builds a PDF with its cross reference table and trailer cut off, confirms pdftoppm rejects it, and checks that the thumbnail produced via the qpdf copy is a real 500px page while the original file is untouched. Add a parametrized test for the double failure, with qpdf itself failing and with the repaired copy still unrenderable, asserting the result is a copy of the default thumbnail.
2026-10-07 12:51:21 -07:00
stumpylog 1a9b08394c Chore: Keep PDF thumbnails at 500px wide for small and odd-sized pages
The new pdftoppm thumbnail path capped its computed DPI at 72, on the
assumption that the old "-scale 500x5000>" never enlarged a page past its
natural size. The old pipeline actually rendered at 300 DPI before
shrinking, so any page wider than about 120pt used to produce a 500px wide
thumbnail, while the 72 DPI cap made A5 pages, receipts and other small
documents come out noticeably narrower. Rounding the DPI to the nearest
integer could also land a few pixels short of 500px (a landscape Letter page
rendered 495px wide at 45 DPI), and the Pillow clamp only ever shrinks.

The DPI is now capped at 300 to match the old render density, and rounded
up so the render always lands at or just above the target, letting the
shrink-only clamp trim it to exactly 500px. The page size helper also no
longer logs a full traceback for every encrypted or corrupt PDF, matching
the page count helper.
2026-10-07 12:46:49 -07:00
stumpylog ae3d8680fa Chore: Generate PDF thumbnails with pdftoppm and Pillow instead of ImageMagick
PDF thumbnails were produced by handing the original PDF to ImageMagick's
convert, which delegates PDF handling to Ghostscript. When that failed, a
fallback invoked gs directly and then ran convert a second time just to
encode the WebP, so a single thumbnail could take three subprocess calls
through two general purpose tools for what is only "render page one small".

The first page is now rasterized with Poppler's pdftoppm, which is already
installed for pdftotext, at a DPI computed from the page's own CropBox and
effective rotation as read by pikepdf. The DPI is chosen so the page fits
500x5000 pixels in a single render and is capped at 72 so small pages are
never enlarged, matching the previous shrink-only scale. Pillow then
flattens any alpha onto white, applies a no-enlarge safety clamp and saves
the WebP in process. If the geometry cannot be read, a fixed 150 DPI is
used and the clamp keeps the output in bounds.

The Ghostscript fallback is replaced with a qpdf repair and retry: the PDF
is copied, repaired in place with qpdf (treating its "repaired with
warnings" exit status as success), and rasterized again. If that also
fails, the default thumbnail is used as before.
2026-10-07 12:43:18 -07:00
17 changed files with 872 additions and 545 deletions

No files matched your search

+1 -1
View File
@@ -111,7 +111,7 @@ jobs:
timeout-minutes: 12
uses: $/.github/actions/apt-install
with:
packages: unpaper tesseract-ocr imagemagick ghostscript poppler-utils
packages: unpaper tesseract-ocr imagemagick ghostscript poppler-utils qpdf
- name: Configure ImageMagick
run: |
sudo cp docker/rootfs/etc/ImageMagick-6/paperless-policy.xml /etc/ImageMagick-7/policy.xml
+6 -18
View File
@@ -1315,29 +1315,17 @@ valid crontab(5) expression describing when to run.
#### [`PAPERLESS_CONVERT_MEMORY_LIMIT=<num>`](#PAPERLESS_CONVERT_MEMORY_LIMIT) {#PAPERLESS_CONVERT_MEMORY_LIMIT}
: On smaller systems, or even in the case of Very Large Documents, the
consumer may explode, complaining about how it's "unable to extend
pixel cache". In such cases, try setting this to a reasonably low
value, like 32. The default is to use whatever is necessary to do
everything without writing to disk, and units are in megabytes.
!!! warning
For more information on how to use this value, you should search the
web for "MAGICK_MEMORY_LIMIT".
Defaults to 0, which disables the limit.
Deprecated and has no effect, since PDF thumbnails no longer use
ImageMagick. It will be removed in a future release.
#### [`PAPERLESS_CONVERT_TMPDIR=<path>`](#PAPERLESS_CONVERT_TMPDIR) {#PAPERLESS_CONVERT_TMPDIR}
: Similar to the memory limit, if you've got a small system and your
OS mounts /tmp as tmpfs, you should set this to a path that's on a
physical disk, like /home/your_user/tmp or something. ImageMagick
will use this as scratch space when crunching through very large
documents.
!!! warning
For more information on how to use this value, you should search the
web for "MAGICK_TMPDIR".
Default is none, which disables the temporary directory.
Deprecated and has no effect, since PDF thumbnails no longer use
ImageMagick. It will be removed in a future release.
#### [`PAPERLESS_APPS=<string>`](#PAPERLESS_APPS) {#PAPERLESS_APPS}
+4 -7
View File
@@ -177,12 +177,12 @@ to a positive number to enable polling and disable native filesystem notificatio
- `pkg-config` for mysqlclient (python dependency)
- `fonts-liberation` for generating thumbnails for plain text
files
- `imagemagick` >= 6 for PDF conversion
- `imagemagick` >= 6 for image alpha handling
- `gnupg` for decrypting GPG-encrypted email
- `libpq-dev` for PostgreSQL
- `libmagic-dev` for mime type detection
- `mariadb-client` for MariaDB compile time
- `poppler-utils` for barcode detection
- `poppler-utils` for thumbnail generation and barcode detection
Use this list for your preferred package management:
@@ -416,11 +416,8 @@ to a positive number to enable polling and disable native filesystem notificatio
You may need to change the path in the files. Example:
`ExecStart=/opt/paperless/.local/bin/celery --app paperless worker --loglevel INFO`
12. Configure ImageMagick to allow processing of PDF documents and disable
formats that Paperless-ngx does not use. Most distributions disable PDF
processing by default, since PDF documents can contain malware. If you
don't enable it, Paperless-ngx will fall back to Ghostscript for certain
steps such as thumbnail generation.
12. Harden ImageMagick by disabling formats that Paperless-ngx does not use.
PDF processing is not needed and should stay disabled.
Configure the active ImageMagick policy file (commonly
`/etc/ImageMagick-6/policy.xml` or `/etc/ImageMagick-7/policy.xml`) and
-2
View File
@@ -50,8 +50,6 @@ PAPERLESS_SECRET_KEY=change-me
#PAPERLESS_OCR_ROTATE_PAGES=true
#PAPERLESS_OCR_ROTATE_PAGES_THRESHOLD=12.0
#PAPERLESS_OCR_USER_ARGS={}
#PAPERLESS_CONVERT_MEMORY_LIMIT=0
#PAPERLESS_CONVERT_TMPDIR=/var/tmp/paperless
# Software tweaks
+161 -98
View File
@@ -1,8 +1,8 @@
from __future__ import annotations
import logging
import math
import mimetypes
import os
import shutil
import subprocess
import tempfile
@@ -68,58 +68,6 @@ def get_supported_file_extensions() -> set[str]:
return extensions
def run_convert(
input_file,
output_file,
*,
density=None,
scale=None,
alpha=None,
strip=False,
trim=False,
type=None,
depth=None,
auto_orient=False,
use_cropbox=False,
extra=None,
logging_group=None,
) -> None:
environment = os.environ.copy()
if settings.CONVERT_MEMORY_LIMIT:
# MAGICK_MEMORY_LIMIT sets the maximum amount of RAM the pixel cache can use.
# MAGICK_MAP_LIMIT sets the maximum amount of memory-mapped I/O allowed.
#
# For large-format documents ImageMagick will hit the RAM limit and
# immediately try to "map" the remaining data. If MAGICK_MAP_LIMIT isn't
# also set, the process may trigger an OOM kill because the default
# system/policy map limit is often too restrictive for these massive bitmaps.
environment["MAGICK_MEMORY_LIMIT"] = settings.CONVERT_MEMORY_LIMIT
environment["MAGICK_MAP_LIMIT"] = settings.CONVERT_MEMORY_LIMIT
if settings.CONVERT_TMPDIR:
environment["MAGICK_TMPDIR"] = settings.CONVERT_TMPDIR
args = [settings.CONVERT_BINARY]
args += ["-density", str(density)] if density else []
args += ["-scale", str(scale)] if scale else []
args += ["-alpha", str(alpha)] if alpha else []
args += ["-strip"] if strip else []
args += ["-trim"] if trim else []
args += ["-type", str(type)] if type else []
args += ["-depth", str(depth)] if depth else []
args += ["-auto-orient"] if auto_orient else []
args += ["-define", "pdf:use-cropbox=true"] if use_cropbox else []
args += [str(input_file), str(output_file)]
logger.debug("Execute: " + " ".join(args), extra={"group": logging_group})
try:
run_subprocess(args, environment, logger)
except subprocess.CalledProcessError as e:
raise ParseError(f"Convert failed at {args}") from e
except Exception as e: # pragma: no cover
raise ParseError("Unknown error running convert") from e
def get_default_thumbnail() -> Path:
"""
Returns the path to a generic thumbnail
@@ -127,46 +75,168 @@ def get_default_thumbnail() -> Path:
return (Path(__file__).parent / "resources" / "document.webp").resolve()
def make_thumbnail_from_pdf_gs_fallback(in_path, temp_dir, logging_group=None) -> Path:
out_path: Path = Path(temp_dir) / "convert_gs.webp"
_THUMBNAIL_MAX_WIDTH = 500
_THUMBNAIL_MAX_HEIGHT = 5000
# Used only when the page geometry cannot be read
_THUMBNAIL_FALLBACK_DPI = 150
# Applied before supersampling, so tiny pages are not enlarged
_THUMBNAIL_MAX_DPI = 300
# Rendering at a multiple and downsampling keeps text crisper
_THUMBNAIL_SUPERSAMPLE = 2
# if convert fails, fall back to extracting
# the first PDF page as a PNG using Ghostscript
logger.warning(
"Thumbnail generation with ImageMagick failed, falling back "
"to ghostscript. Check your /etc/ImageMagick-x/policy.xml!",
extra={"group": logging_group},
)
# Ghostscript doesn't handle WebP outputs
gs_out_path: Path = Path(temp_dir) / "gs_out.png"
cmd = [settings.GS_BINARY, "-q", "-sDEVICE=pngalpha", "-o", gs_out_path, in_path]
def rasterize_pdf_page_to_png(
in_path: Path,
out_path: Path,
*,
dpi: int,
logging_group=None,
) -> None:
"""
Rasterizes the first page of a PDF to a PNG with pdftoppm.
"""
# -singlefile drops the page number and -png appends ".png", so pass the
# path without its suffix
args = [
"pdftoppm",
"-f",
"1",
"-l",
"1",
"-r",
str(dpi),
"-png",
"-singlefile",
"-cropbox",
str(in_path),
str(out_path.with_suffix("")),
]
logger.debug("Execute: " + " ".join(args), extra={"group": logging_group})
try:
try:
run_subprocess(cmd, logger=logger)
except subprocess.CalledProcessError as e:
raise ParseError(f"Thumbnail (gs) failed at {cmd}") from e
# then run convert on the output from gs to make WebP
run_convert(
density=300,
scale="500x5000>",
alpha="remove",
strip=True,
trim=False,
auto_orient=True,
input_file=gs_out_path,
output_file=out_path,
logging_group=logging_group,
)
run_subprocess(args, logger=logger)
except subprocess.CalledProcessError as e:
raise ParseError(f"pdftoppm failed at {args}") from e
except Exception as e: # pragma: no cover
raise ParseError("Unknown error running pdftoppm") from e
def encode_thumbnail_webp(
png_path: Path,
out_path: Path,
*,
supersample: int = 1,
) -> None:
"""
Flattens alpha onto white, undoes supersampling, shrinks to fit and saves as WebP.
"""
from PIL import Image
try:
with Image.open(png_path) as im:
if im.mode in ("RGBA", "LA"):
flattened = Image.new("RGB", im.size, (255, 255, 255))
flattened.paste(im, mask=im.split()[-1])
else:
flattened = im.convert("RGB")
if supersample > 1:
flattened = flattened.resize(
(
max(1, round(flattened.width / supersample)),
max(1, round(flattened.height / supersample)),
),
Image.Resampling.LANCZOS,
)
flattened.thumbnail((_THUMBNAIL_MAX_WIDTH, _THUMBNAIL_MAX_HEIGHT))
flattened.save(out_path, format="WEBP")
except (OSError, Image.DecompressionBombError) as e:
raise ParseError(f"Unable to encode thumbnail from {png_path}") from e
def _compute_thumbnail_dpi(in_path: Path, logging_group=None) -> tuple[int, int]:
"""
Returns (dpi, supersample). Unknown geometry is not supersampled, since
the render size cannot be bounded.
"""
from paperless.parsers.utils import get_pdf_first_page_size_points
size = get_pdf_first_page_size_points(in_path)
if size is None:
logger.debug(
"Could not read PDF page size, using fallback DPI",
extra={"group": logging_group},
)
return _THUMBNAIL_FALLBACK_DPI, 1
width_pts, height_pts = size
dpi_for_width = _THUMBNAIL_MAX_WIDTH * 72 / width_pts
dpi_for_height = _THUMBNAIL_MAX_HEIGHT * 72 / height_pts
# Round up so the downsampled render is never a few pixels short of the
# target; the shrink-only clamp in encode_thumbnail_webp trims the excess.
dpi = max(
1,
math.ceil(min(_THUMBNAIL_MAX_DPI, dpi_for_width, dpi_for_height)),
)
# At the 1 DPI floor the page is already oversized, so do not supersample
return dpi, 1 if dpi == 1 else _THUMBNAIL_SUPERSAMPLE
def _render_pdf_thumbnail(
in_path: Path,
png_path: Path,
out_path: Path,
logging_group=None,
) -> None:
dpi, supersample = _compute_thumbnail_dpi(in_path, logging_group=logging_group)
rasterize_pdf_page_to_png(
in_path,
png_path,
dpi=dpi * supersample,
logging_group=logging_group,
)
encode_thumbnail_webp(png_path, out_path, supersample=supersample)
def _repair_pdf_with_qpdf(in_path: Path, out_path: Path) -> None:
# qpdf exits 3 after a repair; --warning-exit-0 keeps that from failing
try:
shutil.copy(in_path, out_path)
run_subprocess(
["qpdf", "--warning-exit-0", "--replace-input", str(out_path)],
logger=logger,
)
except (subprocess.CalledProcessError, OSError) as e:
raise ParseError(f"qpdf repair failed for {in_path}") from e
def make_thumbnail_from_pdf_qpdf_fallback(
in_path: Path,
temp_dir: Path,
logging_group=None,
) -> Path:
png_path = temp_dir / "page1_repaired.png"
out_path = temp_dir / "convert_qpdf.webp"
repaired_path = temp_dir / "repaired.pdf"
logger.warning(
"Thumbnail generation with pdftoppm failed, attempting qpdf repair and retry.",
extra={"group": logging_group},
)
try:
_repair_pdf_with_qpdf(in_path, repaired_path)
_render_pdf_thumbnail(repaired_path, png_path, out_path, logging_group)
return out_path
except ParseError as e:
logger.error(f"Unable to make thumbnail with Ghostscript: {e}")
logger.error(f"Unable to make thumbnail after qpdf repair: {e}")
# The caller might expect a generated thumbnail that can be moved,
# so we need to copy it before it gets moved.
# https://github.com/paperless-ngx/paperless-ngx/issues/3631
default_thumbnail_path: Path = Path(temp_dir) / "document.webp"
default_thumbnail_path = temp_dir / "document.webp"
copy_file_with_basic_stats(get_default_thumbnail(), default_thumbnail_path)
return default_thumbnail_path
@@ -175,25 +245,18 @@ def make_thumbnail_from_pdf(in_path: Path, temp_dir: Path, logging_group=None) -
"""
The thumbnail of a PDF is just a 500px wide image of the first page.
"""
png_path: Path = temp_dir / "page1.png"
out_path: Path = temp_dir / "convert.webp"
# Run convert to get a decent thumbnail
try:
run_convert(
density=300,
scale="500x5000>",
alpha="remove",
strip=True,
trim=False,
auto_orient=True,
use_cropbox=True,
input_file=f"{in_path}[0]",
output_file=str(out_path),
logging_group=logging_group,
)
_render_pdf_thumbnail(in_path, png_path, out_path, logging_group)
except ParseError as e:
logger.error(f"Unable to make thumbnail with convert: {e}")
out_path = make_thumbnail_from_pdf_gs_fallback(in_path, temp_dir, logging_group)
logger.error(f"Unable to make thumbnail with pdftoppm: {e}")
out_path = make_thumbnail_from_pdf_qpdf_fallback(
in_path,
temp_dir,
logging_group,
)
return out_path
+32 -38
View File
@@ -31,7 +31,6 @@ from documents.search._query import parse_user_query
from documents.search._schema import _write_sentinels
from documents.search._schema import build_schema
from documents.search._schema import open_or_rebuild_index
from documents.search._schema import rebuild_in_progress
from documents.search._schema import wipe_index
from documents.search._tokenizer import ascii_fold
from documents.search._tokenizer import autocomplete_tokens
@@ -1111,44 +1110,39 @@ class TantivyBackend:
flushing a segment, deferring merge work; they do not avoid it.
"""
wipe_index(self._path)
# The marker covers the window where the empty index is already stamped
# as current but not yet populated, so an interrupted rebuild is retried.
with rebuild_in_progress(self._path):
new_index = tantivy.Index(build_schema(), path=str(self._path))
_write_sentinels(self._path)
register_tokenizers(new_index, settings.SEARCH_LANGUAGE)
new_index = tantivy.Index(build_schema(), path=str(self._path))
_write_sentinels(self._path)
register_tokenizers(new_index, settings.SEARCH_LANGUAGE)
# Point instance at the new index so _build_tantivy_doc uses it
old_index, old_schema = self._raw_index, self._raw_schema
self._raw_index = new_index
self._raw_schema = new_index.schema
# Stream documents one-by-one (so the progress bar advances per
# document) while fetching viewer permissions one SQL query per
# chunk. The stream is Sized, so iter_wrapper can still discover
# the total.
documents_stream = _DocumentViewerStream(documents, chunk_size=1000)
try:
writer = new_index.writer(heap_size=writer_heap_bytes)
for document, (viewer_ids, viewer_group_ids) in iter_wrapper(
documents_stream,
):
doc = self._build_tantivy_doc(
document,
viewer_ids=viewer_ids,
viewer_group_ids=viewer_group_ids,
)
writer.add_document(doc)
writer.commit()
# Wait for background merge threads to finish so all segments
# are fully merged and persisted before the index is considered
# rebuilt.
writer.wait_merging_threads()
new_index.reload()
except BaseException: # pragma: no cover
# Restore old index on failure so the backend remains usable
self._raw_index = old_index
self._raw_schema = old_schema
raise
# Point instance at the new index so _build_tantivy_doc uses it
old_index, old_schema = self._raw_index, self._raw_schema
self._raw_index = new_index
self._raw_schema = new_index.schema
# Stream documents one-by-one (so the progress bar advances per
# document) while fetching viewer permissions one SQL query per chunk.
# The stream is Sized, so iter_wrapper can still discover the total.
documents_stream = _DocumentViewerStream(documents, chunk_size=1000)
try:
writer = new_index.writer(heap_size=writer_heap_bytes)
for document, (viewer_ids, viewer_group_ids) in iter_wrapper(
documents_stream,
):
doc = self._build_tantivy_doc(
document,
viewer_ids=viewer_ids,
viewer_group_ids=viewer_group_ids,
)
writer.add_document(doc)
writer.commit()
# Wait for background merge threads to finish so all segments are
# fully merged and persisted before the index is considered rebuilt.
writer.wait_merging_threads()
new_index.reload()
except BaseException: # pragma: no cover
# Restore old index on failure so the backend remains usable
self._raw_index = old_index
self._raw_schema = old_schema
raise
def chunked(iterable, size):
+4 -45
View File
@@ -4,7 +4,6 @@ import hashlib
import json
import logging
import shutil
from contextlib import contextmanager
from typing import TYPE_CHECKING
from typing import Final
from typing import NamedTuple
@@ -17,7 +16,6 @@ from whoosh_compat import FieldKind
from documents.search._fields import PUBLIC_FIELDS
if TYPE_CHECKING:
from collections.abc import Iterator
from pathlib import Path
logger = logging.getLogger("paperless.search")
@@ -30,11 +28,6 @@ logger = logging.getLogger("paperless.search")
# v3 - barcodes JSON field for stored barcode contents
SCHEMA_VERSION: Final[int] = 3
# Present in the index directory from the moment a full rebuild starts until it
# finishes. If a rebuild is interrupted it is left behind, so the half-built
# index is not mistaken for a complete one.
REBUILD_MARKER: Final[str] = ".rebuilding"
class FieldDescriptor(NamedTuple):
"""One tantivy field, in declaration order.
@@ -262,9 +255,9 @@ def needs_rebuild(index_dir: Path) -> bool:
"""
Check if the search index needs rebuilding.
True if a previous full rebuild never finished (the rebuild marker is still
present), or if the index's stamped settings no longer match the current
configuration. See _settings_mismatch().
Reads .index_settings.json to compare the stored schema version, search
language and schema fingerprint against the current configuration. Returns
True if the file is missing, unparsable, or any value mismatches.
Args:
index_dir: Path to the search index directory
@@ -272,40 +265,6 @@ def needs_rebuild(index_dir: Path) -> bool:
Returns:
True if the index needs rebuilding, False if it's up to date
"""
if (index_dir / REBUILD_MARKER).exists():
logger.warning("Previous search index rebuild did not finish - rebuilding.")
return True
return _settings_mismatch(index_dir)
@contextmanager
def rebuild_in_progress(index_dir: Path) -> Iterator[None]:
"""
Flag the index as incomplete for the duration of a full rebuild.
The marker is cleared only if the block exits cleanly. There is deliberately
no try/finally: an exception must leave the marker behind so the next
needs_rebuild() check retries the rebuild.
"""
marker = index_dir / REBUILD_MARKER
marker.touch()
yield
marker.unlink(missing_ok=True)
def _settings_mismatch(index_dir: Path) -> bool:
"""
Check the stamped settings against the current configuration.
Reads .index_settings.json to compare the stored schema version, search
language and schema fingerprint. Returns True if the file is missing,
unparsable, or any value mismatches.
This deliberately ignores the rebuild marker: open_or_rebuild_index() uses it
so that a process opening the index while another process is mid-rebuild
(or after one died) does not wipe the partial index out from under it.
Repopulating is the job of ``document_index reindex``.
"""
settings_file = index_dir / ".index_settings.json"
if not settings_file.exists():
return True
@@ -374,7 +333,7 @@ def open_or_rebuild_index(index_dir: Path | None = None) -> tantivy.Index:
index_dir = cast("Path", settings.INDEX_DIR)
if not index_dir.exists():
return tantivy.Index(build_schema())
if _settings_mismatch(index_dir):
if needs_rebuild(index_dir):
wipe_index(index_dir)
idx = tantivy.Index(build_schema(), path=str(index_dir))
_write_sentinels(index_dir)
+1 -2
View File
@@ -90,7 +90,6 @@ from documents.templating.utils import convert_format_str_to_template_format
from documents.templating.workflows import validate_workflow_template
from documents.validators import uri_validator
from documents.validators import url_validator
from documents.versioning import get_root_document
from documents.versioning import has_prefetched_effective_content
from documents.versioning import sort_versions_newest_first
@@ -2895,7 +2894,7 @@ class ShareLinkSerializer(OwnedObjectSerializer):
and has_perms_owner_aware(
self.user,
"view_document",
get_root_document(document),
document,
)
):
return document
@@ -16,10 +16,7 @@ from documents.search._backend import TantivyBackend
from documents.search._backend import WriteBatch
from documents.search._backend import get_backend
from documents.search._backend import reset_backend
from documents.search._schema import REBUILD_MARKER
from documents.search._schema import needs_rebuild
from documents.signals.handlers import add_to_index
from paperless_testing.dirs import PaperlessDirs
from paperless_testing.factories import CorrespondentFactory
from paperless_testing.factories import DocumentFactory
from paperless_testing.factories import DocumentTypeFactory
@@ -826,53 +823,6 @@ class TestRebuild:
backend.rebuild(Document.objects.all(), iter_wrapper=wrapper)
assert 30 in seen
def test_successful_rebuild_leaves_index_up_to_date(
self,
backend: TantivyBackend,
paperless_dirs: PaperlessDirs,
) -> None:
"""
GIVEN:
- A backend and one document
WHEN:
- rebuild() completes
THEN:
- needs_rebuild() is False and no rebuild marker remains
"""
DocumentFactory.create()
backend.rebuild(Document.objects.all())
assert needs_rebuild(paperless_dirs.index_dir) is False
assert not (paperless_dirs.index_dir / REBUILD_MARKER).exists()
def test_interrupted_rebuild_is_retried(
self,
backend: TantivyBackend,
paperless_dirs: PaperlessDirs,
) -> None:
"""
GIVEN:
- A rebuild that dies while indexing documents (e.g. the database
connection is lost)
WHEN:
- needs_rebuild() is checked afterwards
THEN:
- It is True, even though the empty index was already stamped with
current settings, so the next start rebuilds instead of reporting
the index as up to date
"""
DocumentFactory.create()
def die(pairs):
raise RuntimeError("terminating connection due to administrator command")
yield # pragma: no cover
with pytest.raises(RuntimeError):
backend.rebuild(Document.objects.all(), iter_wrapper=die)
assert needs_rebuild(paperless_dirs.index_dir) is True
def test_includes_group_granted_viewers(self, backend: TantivyBackend) -> None:
"""Rebuild must index viewer ids for group-only grants, not just direct ones.
@@ -8,7 +8,6 @@ from auditlog.models import LogEntry # type: ignore[import-untyped]
from django.contrib.contenttypes.models import ContentType
from django.core.files.uploadedfile import SimpleUploadedFile
from django.test import TestCase as DjangoTestCase
from django.test import override_settings
from django.utils import timezone
from rest_framework import status
from rest_framework.test import APITestCase
@@ -17,17 +16,13 @@ from documents.data_models import DocumentSource
from documents.filters import EffectiveContentFilter
from documents.filters import TitleContentFilter
from documents.models import Document
from documents.models import Note
from documents.models import ShareLink
from documents.versioning import annotate_effective_content
from documents.views import DocumentSelectionMixin
from paperless_testing.dirs import DirectoriesMixin
from paperless_testing.factories import DocumentFactory
from paperless_testing.factories import UserFactory
from paperless_testing.http import read_streaming_response
from paperless_testing.permissions import grant_all_global
from paperless_testing.permissions import grant_global
from paperless_testing.permissions import grant_object
if TYPE_CHECKING:
from pathlib import Path
@@ -1048,128 +1043,3 @@ class TestBulkSelectionExcludesVersions(DjangoTestCase):
)
self.assertEqual(selected, [root.id])
class TestVersionActionPermissions(DirectoriesMixin, APITestCase):
def setUp(self):
super().setUp()
self.user = UserFactory()
grant_all_global(self.user)
self.client.force_authenticate(self.user)
self.root = DocumentFactory(owner=UserFactory())
self.version = DocumentFactory(root_document=self.root, owner=None)
@override_settings(AUDIT_LOG_ENABLED=True)
def test_actions_reject_stale_version_ownership(self):
note = Note.objects.create(document=self.version, note="Version note")
for owner in (None, self.user):
self.version.owner = owner
self.version.save(update_fields=["owner"])
for action in (
"notes",
"suggestions",
"ai_suggestions",
"history",
"share_links",
):
with self.subTest(owner=owner, action=action):
response = self.client.get(
f"/api/documents/{self.version.pk}/{action}/",
)
self.assertEqual(response.status_code, 403)
response = self.client.post(
f"/api/documents/{self.version.pk}/notes/",
{"note": "New note"},
)
self.assertEqual(response.status_code, 403)
response = self.client.delete(
f"/api/documents/{self.version.pk}/notes/?id={note.pk}",
)
self.assertEqual(response.status_code, 403)
response = self.client.post(
"/api/share_links/",
{"document": self.version.pk, "file_version": "original"},
)
self.assertEqual(response.status_code, 403)
response = self.client.post(
"/api/share_link_bundles/",
{"document_ids": [self.version.pk], "file_version": "original"},
format="json",
)
self.assertEqual(response.status_code, 400)
response = self.client.post(
"/api/documents/email/",
{
"documents": [self.version.pk],
"addresses": "recipient@example.com",
"subject": "Version",
"message": "Version",
},
format="json",
)
self.assertEqual(response.status_code, 403)
with (
mock.patch("documents.views.AIConfig") as ai_config,
mock.patch("documents.views.stream_chat_with_documents") as chat,
):
ai_config.return_value.ai_enabled = True
response = self.client.post(
"/api/documents/chat/",
{"q": "Version?", "document_id": self.version.pk},
format="json",
)
self.assertEqual(response.status_code, 403)
chat.assert_not_called()
self.assertTrue(Note.objects.filter(pk=note.pk).exists())
self.assertFalse(ShareLink.objects.exists())
@mock.patch("documents.views.build_share_link_bundle.apply_async")
def test_root_permissions_allow_sharing_a_private_version(self, build_mock):
self.version.owner = UserFactory()
self.version.save(update_fields=["owner"])
grant_object(self.user, self.root, "view_document", "change_document")
note = Note.objects.create(document=self.version, note="Version note")
response = self.client.get(f"/api/documents/{self.version.pk}/notes/")
self.assertEqual(response.status_code, 200)
self.assertEqual(response.data[0]["id"], note.pk)
response = self.client.post(
"/api/share_links/",
{"document": self.version.pk, "file_version": "original"},
)
self.assertEqual(response.status_code, 201)
self.assertEqual(ShareLink.objects.get().document_id, self.version.pk)
response = self.client.get(f"/api/documents/{self.version.pk}/share_links/")
self.assertEqual(response.status_code, 200)
self.assertEqual(len(response.data), 1)
response = self.client.post(
"/api/share_link_bundles/",
{"document_ids": [self.version.pk], "file_version": "original"},
format="json",
)
self.assertEqual(response.status_code, 201)
build_mock.assert_called_once()
def test_root_view_permission_does_not_allow_note_changes(self):
grant_object(self.user, self.root, "view_document")
note = Note.objects.create(document=self.version, note="Version note")
response = self.client.get(f"/api/documents/{self.version.pk}/notes/")
self.assertEqual(response.status_code, 200)
response = self.client.post(
f"/api/documents/{self.version.pk}/notes/",
{"note": "New note"},
)
self.assertEqual(response.status_code, 403)
response = self.client.delete(
f"/api/documents/{self.version.pk}/notes/?id={note.pk}",
)
self.assertEqual(response.status_code, 403)
self.assertTrue(Note.objects.filter(pk=note.pk).exists())
@override_settings(AUDIT_LOG_ENABLED=True)
def test_history_uses_root_ownership(self):
self.root.owner = self.user
self.root.save(update_fields=["owner"])
self.version.owner = UserFactory()
self.version.save(update_fields=["owner"])
response = self.client.get(f"/api/documents/{self.version.pk}/history/")
self.assertEqual(response.status_code, 200)
-68
View File
@@ -279,71 +279,3 @@ class TestTrashAPI(DirectoriesMixin, APITestCase):
Document.objects.filter(root_document=root).values_list("id", flat=True),
[version.pk for version in versions],
)
def test_api_trash_version_follows_root_owner(self) -> None:
"""
GIVEN:
- A deleted version of user2's document, owned by nobody
- A deleted version of the user's document, owned by user2
WHEN:
- The user lists the trash and tries to restore or empty the versions
THEN:
- Only the version of the user's own document is listed
- The other version can't be restored or emptied
- The version of the user's own document can be restored
"""
user2 = UserFactory(username="user2")
other_root = Document.objects.create(
title="other root",
checksum="other-root",
mime_type="application/pdf",
owner=user2,
)
other_version = Document.objects.create(
title="other version",
checksum="other-version",
mime_type="application/pdf",
root_document=other_root,
version_index=1,
)
other_version.delete()
own_root = Document.objects.create(
title="own root",
checksum="own-root",
mime_type="application/pdf",
owner=self.user,
)
own_version = Document.objects.create(
title="own version",
checksum="own-version",
mime_type="application/pdf",
owner=user2,
root_document=own_root,
version_index=1,
)
own_version.delete()
resp = self.client.get("/api/trash/")
self.assertEqual(resp.status_code, status.HTTP_200_OK)
self.assertEqual(
[doc["id"] for doc in resp.data["results"]],
[own_version.pk],
)
for action in ("restore", "empty"):
with self.subTest(action=action):
resp = self.client.post(
"/api/trash/",
{"action": action, "documents": [other_version.pk]},
)
self.assertEqual(resp.status_code, status.HTTP_403_FORBIDDEN)
self.assertTrue(
Document.deleted_objects.filter(pk=other_version.pk).exists(),
)
resp = self.client.post(
"/api/trash/",
{"action": "restore", "documents": [own_version.pk]},
)
self.assertEqual(resp.status_code, status.HTTP_200_OK)
self.assertTrue(Document.objects.filter(pk=own_version.pk).exists())
+382
View File
@@ -1,11 +1,22 @@
import subprocess
from collections.abc import Generator
from pathlib import Path
import pikepdf
import pytest
from PIL import Image
from pytest_django.fixtures import Settings
from pytest_mock import MockerFixture
from documents.parsers import ParseError
from documents.parsers import _compute_thumbnail_dpi
from documents.parsers import encode_thumbnail_webp
from documents.parsers import get_default_file_extension
from documents.parsers import get_default_thumbnail
from documents.parsers import get_supported_file_extensions
from documents.parsers import is_file_ext_supported
from documents.parsers import make_thumbnail_from_pdf
from documents.parsers import rasterize_pdf_page_to_png
from paperless.parsers.registry import get_parser_registry
from paperless.parsers.registry import reset_parser_registry
from paperless.parsers.tesseract import RasterisedDocumentParser
@@ -125,3 +136,374 @@ class TestParserAvailability:
assert is_file_ext_supported(".pdf")
assert not is_file_ext_supported(".hsdfh")
assert not is_file_ext_supported("")
class TestComputeThumbnailDpi:
@pytest.mark.parametrize(
("size", "expected"),
[
pytest.param((612.0, 792.0), (59, 2), id="letter-width-bound"),
pytest.param((792.0, 612.0), (46, 2), id="landscape-rounded-up"),
pytest.param((612.0, 100000.0), (4, 2), id="tall-strip-height-bound"),
pytest.param((200.0, 300.0), (180, 2), id="small-page-width-bound"),
pytest.param((72.0, 72.0), (300, 2), id="tiny-page-capped-at-300"),
pytest.param(
(1000000.0, 1000000.0),
(1, 1),
id="huge-page-minimum-one-unsupersampled",
),
pytest.param(None, (150, 1), id="unreadable-geometry-fallback"),
],
)
def test_dpi_from_page_size(
self,
mocker: MockerFixture,
tmp_path: Path,
size: tuple[float, float] | None,
expected: tuple[int, int],
) -> None:
"""
GIVEN:
- A first page of the given size, or unreadable geometry
WHEN:
- The thumbnail DPI is computed
THEN:
- The expected DPI and supersample factor are returned
"""
mocker.patch(
"paperless.parsers.utils.get_pdf_first_page_size_points",
return_value=size,
)
assert _compute_thumbnail_dpi(tmp_path / "doc.pdf") == expected
class TestMakeThumbnailFromPdf:
@pytest.fixture
def work_dir(self, tmp_path: Path) -> Path:
path = tmp_path / "work"
path.mkdir()
return path
@staticmethod
def _write_blank_pdf(path: Path, page_size: tuple[int, int] = (612, 792)) -> Path:
pdf = pikepdf.new()
pdf.add_blank_page(page_size=page_size)
pdf.save(path, object_stream_mode=pikepdf.ObjectStreamMode.disable)
return path
@pytest.mark.parametrize(
("size", "expected_dpi", "expected_supersample"),
[
pytest.param((612.0, 792.0), 118, 2, id="known-geometry-2x"),
pytest.param(None, 150, 1, id="unreadable-geometry-plain-fallback"),
],
)
def test_render_dpi_requested(
self,
mocker: MockerFixture,
tmp_path: Path,
work_dir: Path,
size: tuple[float, float] | None,
expected_dpi: int,
expected_supersample: int,
) -> None:
"""
GIVEN:
- Readable or unreadable page geometry
WHEN:
- A thumbnail is made
THEN:
- Rasterize and encode get the matching DPI and supersample factor
"""
mocker.patch(
"paperless.parsers.utils.get_pdf_first_page_size_points",
return_value=size,
)
rasterize = mocker.patch("documents.parsers.rasterize_pdf_page_to_png")
encode = mocker.patch("documents.parsers.encode_thumbnail_webp")
make_thumbnail_from_pdf(tmp_path / "in.pdf", work_dir)
assert rasterize.call_args.kwargs["dpi"] == expected_dpi
assert encode.call_args.kwargs["supersample"] == expected_supersample
@pytest.mark.parametrize(
("page_size", "expected_width"),
[
pytest.param((612, 792), 500, id="letter"),
pytest.param((792, 612), 500, id="landscape-letter"),
pytest.param((595, 842), 500, id="a4"),
pytest.param((200, 300), 500, id="small-page"),
pytest.param((72, 72), 300, id="tiny-page-capped"),
],
)
def test_thumbnail_width(
self,
tmp_path: Path,
work_dir: Path,
page_size: tuple[int, int],
expected_width: int,
) -> None:
"""
GIVEN:
- A PDF whose first page has the given size in points
WHEN:
- A thumbnail is made from it
THEN:
- The thumbnail has the expected width
"""
pdf_path = self._write_blank_pdf(tmp_path / "in.pdf", page_size)
thumb = make_thumbnail_from_pdf(pdf_path, work_dir)
assert thumb == work_dir / "convert.webp"
with Image.open(thumb) as im:
assert im.format == "WEBP"
assert im.width == expected_width
@classmethod
def _write_pdf_without_xref(cls, path: Path) -> Path:
"""
Cuts off the xref and trailer, which pdftoppm cannot recover from but qpdf can.
"""
cls._write_blank_pdf(path)
data = path.read_bytes()
path.write_bytes(data[: data.rindex(b"\nxref")])
return path
def test_qpdf_repair_produces_real_thumbnail(
self,
tmp_path: Path,
work_dir: Path,
) -> None:
"""
GIVEN:
- A PDF with its xref table and trailer cut off
WHEN:
- A thumbnail is made from it
THEN:
- The thumbnail is rendered from a qpdf repaired copy
- The original file is unchanged
"""
pdf_path = self._write_pdf_without_xref(tmp_path / "broken.pdf")
original_bytes = pdf_path.read_bytes()
with pytest.raises(ParseError):
rasterize_pdf_page_to_png(pdf_path, work_dir / "probe.png", dpi=50)
thumb = make_thumbnail_from_pdf(pdf_path, work_dir)
assert thumb == work_dir / "convert_qpdf.webp"
with Image.open(thumb) as im:
assert im.format == "WEBP"
assert im.width == 500
assert pdf_path.read_bytes() == original_bytes
@pytest.mark.parametrize(
"qpdf_error",
[
pytest.param(subprocess.CalledProcessError(2, "qpdf"), id="qpdf-fails"),
pytest.param(None, id="repaired-still-unrenderable"),
],
)
def test_double_failure_uses_default_thumbnail(
self,
mocker: MockerFixture,
tmp_path: Path,
work_dir: Path,
qpdf_error: subprocess.CalledProcessError | None,
) -> None:
"""
GIVEN:
- A PDF that cannot be rendered, even after qpdf repair
WHEN:
- A thumbnail is made from it
THEN:
- A copy of the default thumbnail is returned
"""
mocker.patch(
"documents.parsers.rasterize_pdf_page_to_png",
side_effect=ParseError("Does not compute."),
)
if qpdf_error is not None:
mocker.patch("documents.parsers.run_subprocess", side_effect=qpdf_error)
pdf_path = self._write_blank_pdf(tmp_path / "in.pdf")
thumb = make_thumbnail_from_pdf(pdf_path, work_dir)
assert thumb == work_dir / "document.webp"
assert thumb.read_bytes() == get_default_thumbnail().read_bytes()
class TestRasterizePdfPageToPng:
@staticmethod
def _write_pdf(
path: Path,
*,
crop_box: tuple[float, float, float, float] | None = None,
rotate: int | None = None,
) -> Path:
pdf = pikepdf.new()
pdf.add_blank_page(page_size=(144, 72))
pdf.add_blank_page(page_size=(300, 300))
page = pdf.pages[0]
if crop_box is not None:
page.obj.CropBox = pikepdf.Array(crop_box)
if rotate is not None:
page.obj.Rotate = rotate
pdf.save(path)
return path
@pytest.mark.parametrize(
("crop_box", "rotate", "expected_size"),
[
pytest.param(None, None, (144, 72), id="plain"),
pytest.param(None, 90, (72, 144), id="rotated-90"),
pytest.param((0, 0, 72, 36), None, (72, 36), id="crop-box"),
],
)
def test_renders_first_page_to_exact_path(
self,
tmp_path: Path,
crop_box: tuple[float, float, float, float] | None,
rotate: int | None,
expected_size: tuple[int, int],
) -> None:
"""
GIVEN:
- A two page PDF, first page optionally cropped or rotated
WHEN:
- The first page is rasterized at 72 DPI
THEN:
- Only out_path is written, sized to the first page's crop and rotation
"""
pdf_path = self._write_pdf(
tmp_path / "in.pdf",
crop_box=crop_box,
rotate=rotate,
)
out_dir = tmp_path / "out"
out_dir.mkdir()
out_path = out_dir / "page1.png"
rasterize_pdf_page_to_png(pdf_path, out_path, dpi=72)
assert list(out_dir.iterdir()) == [out_path]
with Image.open(out_path) as im:
assert im.format == "PNG"
assert im.size == expected_size
def test_failure_raises_parse_error(self, tmp_path: Path) -> None:
"""
GIVEN:
- A file that is not a PDF
WHEN:
- Rasterization is attempted
THEN:
- A ParseError is raised
"""
bad = tmp_path / "bad.pdf"
bad.write_bytes(b"not a pdf")
with pytest.raises(ParseError):
rasterize_pdf_page_to_png(bad, tmp_path / "page1.png", dpi=72)
class TestEncodeThumbnailWebp:
@pytest.mark.parametrize(
("mode", "color"),
[
pytest.param("RGBA", (0, 0, 0, 0), id="rgba"),
pytest.param("LA", (0, 0), id="la"),
],
)
def test_alpha_flattened_onto_white(
self,
tmp_path: Path,
mode: str,
color: tuple[int, ...],
) -> None:
"""
GIVEN:
- A fully transparent PNG with an alpha channel
WHEN:
- It is encoded as a thumbnail
THEN:
- The WebP output is RGB with the transparency flattened to white
"""
png_path = tmp_path / "in.png"
Image.new(mode, (20, 10), color).save(png_path)
out_path = tmp_path / "out.webp"
encode_thumbnail_webp(png_path, out_path)
with Image.open(out_path) as im:
assert im.format == "WEBP"
assert im.mode == "RGB"
assert im.size == (20, 10)
red, green, blue = im.getpixel((10, 5))
assert min(red, green, blue) >= 250
@pytest.mark.parametrize(
("in_size", "supersample", "expected_size"),
[
pytest.param((1000, 2000), 1, (500, 1000), id="too-wide-shrunk"),
pytest.param((100, 10000), 1, (50, 5000), id="too-tall-shrunk"),
pytest.param((100, 200), 1, (100, 200), id="small-not-enlarged"),
pytest.param((1000, 1400), 2, (500, 700), id="2x-halved"),
pytest.param((1001, 1401), 2, (500, 700), id="2x-odd-rounded"),
pytest.param((600, 800), 2, (300, 400), id="2x-small-not-enlarged"),
pytest.param((900, 1200), 2, (450, 600), id="2x-below-clamp"),
],
)
def test_size_clamped_and_downsampled(
self,
tmp_path: Path,
in_size: tuple[int, int],
supersample: int,
expected_size: tuple[int, int],
) -> None:
"""
GIVEN:
- A rendered image and its supersample factor
WHEN:
- It is encoded as a thumbnail with that factor
THEN:
- It is downsampled, fit within 500x5000 and never enlarged
"""
png_path = tmp_path / "in.png"
Image.new("RGB", in_size, (255, 255, 255)).save(png_path)
out_path = tmp_path / "out.webp"
encode_thumbnail_webp(png_path, out_path, supersample=supersample)
with Image.open(out_path) as im:
assert im.size == expected_size
@pytest.mark.parametrize(
"error",
[
pytest.param(OSError("broken image"), id="os-error"),
pytest.param(Image.DecompressionBombError("too large"), id="bomb"),
],
)
def test_decode_failure_raises_parse_error(
self,
tmp_path: Path,
mocker: MockerFixture,
error: Exception,
) -> None:
"""
GIVEN:
- Opening the rendered image fails
WHEN:
- It is encoded as a thumbnail
THEN:
- A ParseError is raised
"""
png_path = tmp_path / "in.png"
Image.new("RGB", (10, 10)).save(png_path)
mocker.patch("PIL.Image.open", side_effect=error)
with pytest.raises(ParseError):
encode_thumbnail_webp(png_path, tmp_path / "out.webp")
+32 -74
View File
@@ -1550,16 +1550,13 @@ class DocumentViewSet(
)
def suggestions(self, request, pk=None):
doc = get_object_or_404(
Document.objects.select_related(
"owner",
"root_document__owner",
).prefetch_related("versions"),
Document.objects.select_related("owner").prefetch_related("versions"),
pk=pk,
)
if request.user is not None and not has_perms_owner_aware(
request.user,
"change_document",
get_root_document(doc),
doc,
):
return HttpResponseForbidden("Insufficient permissions")
@@ -1613,16 +1610,13 @@ class DocumentViewSet(
@method_decorator(cache_control(no_cache=True))
def ai_suggestions(self, request, pk=None):
doc = get_object_or_404(
Document.objects.select_related(
"owner",
"root_document__owner",
).prefetch_related("versions"),
Document.objects.select_related("owner").prefetch_related("versions"),
pk=pk,
)
if request.user is not None and not has_perms_owner_aware(
request.user,
"change_document",
get_root_document(doc),
doc,
):
return HttpResponseForbidden("Insufficient permissions")
@@ -1862,20 +1856,15 @@ class DocumentViewSet(
currentUser = request.user
try:
doc = (
Document.objects.select_related("owner", "root_document__owner")
Document.objects.select_related("owner")
.prefetch_related("notes")
.only(
"pk",
"owner__id",
"root_document__id",
"root_document__owner__id",
)
.only("pk", "owner__id")
.get(pk=pk)
)
if currentUser is not None and not has_perms_owner_aware(
currentUser,
"view_document",
get_root_document(doc),
doc,
):
return HttpResponseForbidden("Insufficient permissions to view notes")
except Document.DoesNotExist:
@@ -1897,7 +1886,7 @@ class DocumentViewSet(
if currentUser is not None and not has_perms_owner_aware(
currentUser,
"change_document",
get_root_document(doc),
doc,
):
return HttpResponseForbidden(
"Insufficient permissions to create notes",
@@ -1940,7 +1929,7 @@ class DocumentViewSet(
if currentUser is not None and not has_perms_owner_aware(
currentUser,
"change_document",
get_root_document(doc),
doc,
):
return HttpResponseForbidden("Insufficient permissions to delete notes")
@@ -1984,13 +1973,11 @@ class DocumentViewSet(
def share_links(self, request, pk=None):
currentUser = request.user
try:
doc = Document.objects.select_related("owner", "root_document__owner").get(
pk=pk,
)
doc = Document.objects.select_related("owner").get(pk=pk)
if currentUser is not None and not has_perms_owner_aware(
currentUser,
"change_document",
get_root_document(doc),
doc,
):
return HttpResponseForbidden(
"Insufficient permissions to add share link",
@@ -2021,11 +2008,10 @@ class DocumentViewSet(
if not settings.AUDIT_LOG_ENABLED:
return HttpResponseBadRequest("Audit log is disabled")
try:
doc = Document.objects.select_related("root_document__owner").get(pk=pk)
root_doc = get_root_document(doc)
doc = Document.objects.get(pk=pk)
if not request.user.has_perm("auditlog.view_logentry") or (
root_doc.owner is not None
and root_doc.owner != request.user
doc.owner is not None
and doc.owner != request.user
and not request.user.is_superuser
):
return HttpResponseForbidden(
@@ -2113,13 +2099,14 @@ class DocumentViewSet(
message = validated_data.get("message")
use_archive_version = validated_data.get("use_archive_version", True)
documents = Document.objects.filter(pk__in=document_ids).select_related(
"root_document__owner",
)
if request.user is not None:
permitted_ids = set(permitted_document_ids(request.user))
if any(get_root_document(doc).pk not in permitted_ids for doc in documents):
return HttpResponseForbidden("Insufficient permissions")
documents = Document.objects.filter(pk__in=document_ids)
if (
request.user is not None
and documents.exclude(
pk__in=permitted_document_ids(request.user),
).exists()
):
return HttpResponseForbidden("Insufficient permissions")
try:
attachments: list[EmailAttachment] = []
@@ -2443,17 +2430,11 @@ class ChatStreamingView(GenericAPIView[Any]):
if doc_id:
try:
document = Document.objects.select_related(
"root_document__owner",
).get(id=doc_id)
document = Document.objects.get(id=doc_id)
except Document.DoesNotExist:
return HttpResponseBadRequest("Document not found")
if not has_perms_owner_aware(
request.user,
"view_document",
get_root_document(document),
):
if not has_perms_owner_aware(request.user, "view_document", document):
return HttpResponseForbidden("Insufficient permissions")
documents = Document.objects.filter(pk=document.pk)
@@ -4796,7 +4777,6 @@ class ShareLinkBundleViewSet(PassUserMixin, ModelViewSet[ShareLinkBundle]):
document_ids = serializer.validated_data["document_ids"]
documents_qs = Document.objects.filter(pk__in=document_ids).select_related(
"owner",
"root_document__owner",
)
found_ids = set(documents_qs.values_list("pk", flat=True))
missing = sorted(set(document_ids) - found_ids)
@@ -4813,7 +4793,7 @@ class ShareLinkBundleViewSet(PassUserMixin, ModelViewSet[ShareLinkBundle]):
documents = list(documents_qs)
permitted_ids = set(permitted_document_ids(request.user))
for document in documents:
if get_root_document(document).pk not in permitted_ids:
if document.pk not in permitted_ids:
raise ValidationError(
{
"document_ids": _(
@@ -5607,23 +5587,6 @@ class TrashView(ListModelMixin, PassUserMixin):
class _TrashPermittedObjectsFilter(PermittedObjectsFilter):
include_granted = False
def filter_queryset(self, request, queryset, view):
if request.user.is_superuser or not request.user.is_active:
return super().filter_queryset(request, queryset, view)
# A version belongs to whoever owns its root
def owned_or_unowned(prefix: str) -> Q:
return Q(**{f"{prefix}owner": request.user}) | Q(
**{f"{prefix}owner__isnull": True},
)
return queryset.filter(
(Q(root_document__isnull=True) & owned_or_unowned(""))
| (
Q(root_document__isnull=False) & owned_or_unowned("root_document__")
),
)
filter_backends = (_TrashPermittedObjectsFilter,)
pagination_class = StandardPagination
@@ -5653,18 +5616,13 @@ class TrashView(ListModelMixin, PassUserMixin):
if doc_ids is not None
else self.filter_queryset(self.get_queryset()).all()
)
# Versions are authorized by their root document
if (
docs.annotate(root_id=Coalesce("root_document_id", "id"))
.exclude(
root_id__in=permitted_document_ids(
request.user,
perm="delete_document",
include_deleted=True,
),
)
.exists()
):
if docs.exclude(
pk__in=permitted_document_ids(
request.user,
perm="delete_document",
include_deleted=True,
),
).exists():
return HttpResponseForbidden("Insufficient permissions")
action = serializer.validated_data.get("action")
if action == "restore":
+46
View File
@@ -265,6 +265,52 @@ def get_page_count_for_pdf(
return None
def get_pdf_first_page_size_points(
path: Path,
log: logging.Logger | None = None,
) -> tuple[float, float] | None:
"""Return the first page's (width, height) in PDF points, post-rotation.
Uses the CropBox (MediaBox if absent), which must match pdftoppm's
``-cropbox`` or the computed DPI targets the wrong box.
Swaps width and height for 90/270 rotation. ``page.rotation`` resolves
inherited and negative ``/Rotate`` values, a raw lookup does not.
Parameters
----------
path:
Absolute path to the PDF file.
log:
Logger for warnings. Falls back to the module-level logger when omitted.
Returns
-------
tuple[float, float] | None
``(width_points, height_points)``, or ``None`` if the file cannot be
opened, has no pages, or the page box is degenerate.
"""
import pikepdf
_log = log or logger
try:
with pikepdf.Pdf.open(path) as pdf:
if len(pdf.pages) == 0:
return None
page = pdf.pages[0]
llx, lly, urx, ury = (float(v) for v in page.cropbox)
width = abs(urx - llx)
height = abs(ury - lly)
if width <= 0 or height <= 0:
return None
if page.rotation in (90, 270):
width, height = height, width
return width, height
except Exception as e:
_log.warning("Could not determine PDF page size for %s: %s", path, e)
return None
def extract_pdf_metadata(
document_path: Path,
log: logging.Logger | None = None,
-2
View File
@@ -986,8 +986,6 @@ GNUPG_HOME = os.getenv("HOME", "/tmp")
# Convert is part of the ImageMagick package
CONVERT_BINARY = os.getenv("PAPERLESS_CONVERT_BINARY", "convert")
CONVERT_TMPDIR = os.getenv("PAPERLESS_CONVERT_TMPDIR")
CONVERT_MEMORY_LIMIT = os.getenv("PAPERLESS_CONVERT_MEMORY_LIMIT")
GS_BINARY = os.getenv("PAPERLESS_GS_BINARY", "gs")
@@ -15,9 +15,10 @@ from typing import TYPE_CHECKING
import pytest
from ocrmypdf import SubprocessOutputError
from PIL import Image
from documents.parsers import ParseError
from documents.parsers import run_convert
from documents.parsers import rasterize_pdf_page_to_png
from paperless.models import ModeChoices
from paperless.parsers import ParserProtocol
from paperless.parsers.tesseract import RasterisedDocumentParser
@@ -280,24 +281,71 @@ class TestGetThumbnail:
)
assert thumb.is_file()
def test_thumbnail_fallback_on_convert_error(
@pytest.mark.parametrize(
("filename", "expected_height"),
[
pytest.param("simple-digital.pdf", 647, id="portrait-letter"),
pytest.param("rotated.pdf", 386, id="landscape"),
],
)
def test_thumbnail_is_correct_format_and_size(
self,
tesseract_parser: RasterisedDocumentParser,
tesseract_samples_dir: Path,
filename: str,
expected_height: int,
) -> None:
"""
GIVEN:
- A portrait or landscape PDF
WHEN:
- A thumbnail is generated
THEN:
- A 500px wide WebP keeping the page's aspect ratio
"""
thumb = tesseract_parser.get_thumbnail(
tesseract_samples_dir / filename,
"application/pdf",
)
with Image.open(thumb) as im:
assert im.format == "WEBP"
assert im.width == 500
assert im.height == pytest.approx(expected_height, abs=2)
def test_thumbnail_fallback_on_pdftoppm_error(
self,
mocker: MockerFixture,
tesseract_parser: RasterisedDocumentParser,
tesseract_samples_dir: Path,
) -> None:
def _raise_on_pdf(input_file, output_file, **kwargs) -> None:
if ".pdf" in str(input_file):
"""
GIVEN:
- Rasterizing the original PDF fails
WHEN:
- A thumbnail is generated
THEN:
- The PDF is repaired with qpdf and a real thumbnail is rendered
"""
original = tesseract_samples_dir / "simple-digital.pdf"
def _fail_on_original(in_path: Path, out_path: Path, **kwargs) -> None:
if in_path == original:
raise ParseError("Does not compute.")
run_convert(input_file=input_file, output_file=output_file, **kwargs)
rasterize_pdf_page_to_png(in_path, out_path, **kwargs)
mocker.patch("documents.parsers.run_convert", side_effect=_raise_on_pdf)
thumb = tesseract_parser.get_thumbnail(
tesseract_samples_dir / "simple-digital.pdf",
"application/pdf",
rasterize = mocker.patch(
"documents.parsers.rasterize_pdf_page_to_png",
side_effect=_fail_on_original,
)
thumb = tesseract_parser.get_thumbnail(original, "application/pdf")
assert rasterize.call_count == 2
assert thumb.is_file()
assert thumb.name == "convert_qpdf.webp"
with Image.open(thumb) as im:
assert im.format == "WEBP"
assert im.width == 500
def test_thumbnail_encrypted_pdf(
self,
+145
View File
@@ -6,8 +6,10 @@ import codecs
from pathlib import Path
from typing import TYPE_CHECKING
import pikepdf
import pytest
from paperless.parsers.utils import get_pdf_first_page_size_points
from paperless.parsers.utils import is_tagged_pdf
from paperless.parsers.utils import pdf_born_digital_text
from paperless.parsers.utils import post_process_text
@@ -70,6 +72,149 @@ class TestIsTaggedPdf:
assert is_tagged_pdf(bad) is False
class TestGetPdfFirstPageSizePoints:
@staticmethod
def _write_pdf(
path: Path,
*,
media_box: tuple[float, float, float, float] = (0, 0, 600, 800),
crop_box: tuple[float, float, float, float] | None = None,
page_rotate: int | None = None,
inherited_rotate: int | None = None,
) -> Path:
pdf = pikepdf.new()
pdf.add_blank_page(page_size=(media_box[2], media_box[3]))
page = pdf.pages[0]
page.obj.MediaBox = pikepdf.Array(media_box)
if crop_box is not None:
page.obj.CropBox = pikepdf.Array(crop_box)
if page_rotate is not None:
page.obj.Rotate = page_rotate
if inherited_rotate is not None:
pdf.Root.Pages.Rotate = inherited_rotate
pdf.save(path)
return path
def test_letter_sample(self) -> None:
"""
GIVEN:
- A US Letter sample PDF with no CropBox and no rotation
WHEN:
- The first page size is requested
THEN:
- The MediaBox size in points is returned
"""
assert get_pdf_first_page_size_points(SAMPLES / "simple-digital.pdf") == (
612.0,
792.0,
)
@pytest.mark.parametrize(
("rotate", "expected"),
[
pytest.param(0, (600.0, 800.0), id="rotate-0"),
pytest.param(90, (800.0, 600.0), id="rotate-90"),
pytest.param(180, (600.0, 800.0), id="rotate-180"),
pytest.param(270, (800.0, 600.0), id="rotate-270"),
pytest.param(-90, (800.0, 600.0), id="rotate-negative-90"),
],
)
def test_page_rotation_swaps_dimensions(
self,
tmp_path: Path,
rotate: int,
expected: tuple[float, float],
) -> None:
"""
GIVEN:
- A portrait PDF page with /Rotate set directly on the page
WHEN:
- The first page size is requested
THEN:
- Width and height are swapped for quarter-turn rotations only
"""
pdf_path = self._write_pdf(tmp_path / "rotated.pdf", page_rotate=rotate)
assert get_pdf_first_page_size_points(pdf_path) == expected
def test_inherited_rotation_swaps_dimensions(self, tmp_path: Path) -> None:
"""
GIVEN:
- A page inheriting /Rotate 90 from the /Pages node
WHEN:
- The first page size is requested
THEN:
- Width and height are swapped
"""
pdf_path = self._write_pdf(tmp_path / "inherited.pdf", inherited_rotate=90)
assert get_pdf_first_page_size_points(pdf_path) == (800.0, 600.0)
def test_crop_box_preferred_over_media_box(self, tmp_path: Path) -> None:
"""
GIVEN:
- A PDF page with a CropBox smaller than its MediaBox
WHEN:
- The first page size is requested
THEN:
- The CropBox size is returned
"""
pdf_path = self._write_pdf(
tmp_path / "cropped.pdf",
crop_box=(50, 100, 350, 500),
)
assert get_pdf_first_page_size_points(pdf_path) == (300.0, 400.0)
def test_degenerate_box_returns_none(self, tmp_path: Path) -> None:
"""
GIVEN:
- A PDF page whose box has zero width
WHEN:
- The first page size is requested
THEN:
- None is returned
"""
pdf_path = self._write_pdf(
tmp_path / "degenerate.pdf",
media_box=(0, 0, 600, 800),
crop_box=(100, 0, 100, 800),
)
assert get_pdf_first_page_size_points(pdf_path) is None
def test_nonexistent_path_returns_none(self) -> None:
"""
GIVEN:
- A path that does not exist
WHEN:
- The first page size is requested
THEN:
- None is returned and nothing is raised
"""
assert get_pdf_first_page_size_points(Path("/nonexistent/file.pdf")) is None
def test_corrupt_pdf_returns_none(self, tmp_path: Path) -> None:
"""
GIVEN:
- A file that is not a PDF
WHEN:
- The first page size is requested
THEN:
- None is returned and nothing is raised
"""
bad = tmp_path / "bad.pdf"
bad.write_bytes(b"not a pdf")
assert get_pdf_first_page_size_points(bad) is None
def test_encrypted_pdf_returns_none(self) -> None:
"""
GIVEN:
- A password protected PDF
WHEN:
- The first page size is requested
THEN:
- None is returned and nothing is raised
"""
assert get_pdf_first_page_size_points(SAMPLES / "encrypted.pdf") is None
class TestPostProcessText:
@pytest.mark.parametrize(
("source", "expected"),