Add images of json_documents: save() and load()

An image is a document stored so that loading it needs no
parsing: save() writes the node index, the text and the decoded
strings; the static load() reads an image written by save().

load() takes a pointer and size, a borrowed vector, or an owned
rvalue vector; the nodes are copied so they are aligned and can
be edited, while the text and decoded strings stay in the image.

image_check controls how much load() trusts the input: full
checks structure, bounds, strings and numbers, the parser's own
guarantees; bounds checks structure and bounds only; none skips
all checks, for images from a trusted source.

Layout is little-endian only ("NJVI" header, nodes, text, decoded
strings), following the idea of zero-copy formats such as
FlatBuffers and YaFF; the check follows FlatBuffers' Verifier.

New errors: parse_error.116 for a malformed image or a failed
check, type_error.320 for a discarded document or a big-endian
target.

A dedicated fuzzer and 6,000 seeded corruptions, checked under
ASan/UBSan, found and fixed two gaps: unchecked reserved header
fields, and unbounded null/boolean offsets that could make
dump() throw std::length_error.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann committed 2026-10-07 16:42:51 +02:00
1 parent c6ee5a64be
commit 73ceb71c95
27 files changed
+3064 -244

No files matched your search

+6
View File
@@ -9,6 +9,12 @@ Additionally, `parse_json_view_fuzzer` (`tests/src/fuzzer-parse_json_view.cpp`)
produces, and that a rejected input makes both parsers throw with an identical `what()`. It takes plain JSON text, so it
reuses the `corpus_json` corpus rather than a format of its own.
`json_view_image_fuzzer` (`tests/src/fuzzer-json_view_image.cpp`) tests the images of `json_document` (`save()` and
`load()`). It uses each input twice: as an image, which `load()` must either reject with `parse_error.116` or read
safely (with `image_check::full`, the document must also serialize to the JSON it reads as), and as a JSON text, whose
image must load and serialize to the same text. A corpus of images can be made from JSON files with a small program
that calls `json_document::parse(text).save()`; plain JSON files work as well.
## What the fuzzers check
Each fuzzer driver (`tests/src/fuzzer-parse_*.cpp`) parses its input twice: once with `allow_exceptions = false` and