These two options set ImageMagick's memory limit and scratch directory, but their only reader was the PDF thumbnail conversion helper. PDF thumbnails no longer use ImageMagick, so the settings did nothing while still being documented and advertised in the example configuration.
Remove the unused settings, add system warnings for anyone who still has either variable set, mark both options as deprecated and without effect in the configuration docs, and drop them from paperless.conf.example. The manual setup guide is also corrected so it no longer claims ImageMagick is needed for PDF conversion, and the ImageMagick policy step now only covers hardening.
The note claimed that steps still relying on ImageMagick, such as TIFF to
PDF conversion, could fail without the PDF policy change. The remaining
convert calls only strip alpha from raster images, and the TIFF to PDF
step uses img2pdf, so the PDF coder policy no longer matters.
The passage now says Paperless-ngx no longer passes PDF documents to
ImageMagick, so enabling PDF processing is not required, while keeping the
policy hardening guidance in place.
Pillow raises DecompressionBombError, which is not an OSError, so an
enormous render from the unknown-geometry fallback DPI skipped the
default thumbnail fallback. It is now converted to a ParseError like other
encoding failures, with a test covering both error types.
CI did not install qpdf, which the new thumbnail repair tests need, so it
is added to the backend package list. The setup docs now list poppler-utils
as used for thumbnail generation and no longer claim PDF thumbnails fall
back to Ghostscript when the ImageMagick PDF policy is not enabled.
* Preprocesses classifier content with Tantivy instead of NLTK
Tokenizing and stemming now happen in one Rust call instead of NLTK's
Python tokenizer and per word stemming, which also removes the Redis
backed stem cache from every preprocessing call. The output matches the
NLTK pipeline closely; tokens containing digits are now stemmed, and the
English stop words follow Snowball's list.
Stemming and stop word removal apply whenever the OCR language is one of
the supported classifier languages, so PAPERLESS_ENABLE_NLTK and
PAPERLESS_NLTK_DIR are removed.
* Copies packages instead of hardlinking them in backend CI, some NLTK thing
* Adds a normalization to NFC to better fit what Tantivy expects
To be consistent, let's provide an easily copy-pastable list of packages
even for the final set of build dependencies.
Signed-off-by: martin f. krafft <madduck@madduck.net>