Commit Graph
4 Commits
Author SHA1 Message Date
Niels Lohmann 64c3dbd152 Write json_view's dump() without a call or conversion per token
The default dump() (no indentation, no ensure_ascii) gets its own
writer that makes the same walk and produces the same output:

- the write position stays in a local variable instead of a
  member, so the compiler keeps it in a register across stores
  through aliasing char pointers;
- strings and number tokens are copied with fixed-size 32-byte
  moves wherever enough source bytes remain, instead of one
  memcpy call per token;
- the innermost open container lives in local variables; a stack
  that starts as a local array of 32 entries holds the rest;
- unedited documents are walked through the node array in order,
  and integer tokens are read from the source directly.

On top of that, float tokens of at most 15 significant digits are
written straight from their digits via zmij::to_shortest() and
write_shortest(), without converting to a double and back: such
decimals are farther apart than a double's rounding interval, so
the token's digits are the double's shortest digits. Tokens of
16+ digits, or edited values, still go through decimal_to_float().
The view's own NEON write_decimal() is removed in favor of the
shared writer, and the dump output now grows in 64 KiB steps
instead of being resized to its estimate at once.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-06 10:48:07 +02:00
Niels Lohmann 9cb9f7fbad Add images of json_documents: save() and load()
An image is a document stored so that loading it needs no
parsing: save() writes the node index, the text and the decoded
strings; the static load() reads an image written by save().

load() takes a pointer and size, a borrowed vector, or an owned
rvalue vector; the nodes are copied so they are aligned and can
be edited, while the text and decoded strings stay in the image.

image_check controls how much load() trusts the input: full
checks structure, bounds, strings and numbers, the parser's own
guarantees; bounds checks structure and bounds only; none skips
all checks, for images from a trusted source.

Layout is little-endian only ("NJVI" header, nodes, text, decoded
strings), following the idea of zero-copy formats such as
FlatBuffers and YaFF; the check follows FlatBuffers' Verifier.

New errors: parse_error.116 for a malformed image or a failed
check, type_error.320 for a discarded document or a big-endian
target.

A dedicated fuzzer and 6,000 seeded corruptions, checked under
ASan/UBSan, found and fixed two gaps: unchecked reserved header
fields, and unbounded null/boolean offsets that could make
dump() throw std::length_error.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-06 10:48:07 +02:00
Niels Lohmann a0b98325cf Add editable json_documents: set, push_back, insert, and erase
basic_json_document gets a second template parameter, Editable
(false by default), plus the aliases json_editable_document,
json_editable_view, ordered_json_editable_document and
ordered_json_editable_view.

Editable documents can change values and structure without
rewriting the source text: set()/push_back() on values, keys,
array indices and JSON pointers; insert() before an array
element; erase() of an object key, array index or JSON pointer.

New values and element sequences go into edit storage that the
document owns and never moves, so views keep referring to their
value across edits and a parsed node never moves. Read-only
documents walk the plain node array and are unaffected.

Strings are checked for UTF-8 on entry, so dump() of an editable
document never throws type_error.316. Binary values cannot be
stored (type_error.319).

A seeded differential test applies random edits to an editable
document and to the equivalent ordered_json and compares both
after every step.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-06 10:48:07 +02:00
Niels Lohmann d26cbef976 Add dump() and comparisons to json_view
Add basic_json_view::dump() and the comparison operators, and read
floats from the parser's digit layout instead of rescanning the
token.

dump(indent, indent_char, ensure_ascii, number_format) writes a
value the way ordered_json::parse(text).dump() writes it for the
same arguments: members in document order, all of them should a
key occur more than once; strings escaped by the same rules, using
the library's scanning kernels; floats written with the library's
to_chars conversion, so the output equals basic_json's byte for
byte; integers copied from the source, where they are already
canonical, except -0, which parse() reads as 0. There is no
error_handler argument, because the view only holds valid UTF-8.
number_format::source copies numbers exactly as they appear in the
source (e.g. "1.50", "1E2", "-0"), which basic_json cannot provide.
operator<< takes the indentation from the stream width, as for
basic_json. The writer walks iteratively, so nesting depth is
limited by memory only.

operator== and operator!= compare two views, or a view and a
basic_json value in either order, by the rules basic_json's
operator== uses: numbers compare by value across their types,
objects compare by their members with duplicate keys resolved as
parse() resolves them, member order matters only where the object
type keeps one, and discarded views compare as discarded basic_json
values do, including under JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON.
Nothing is materialized except single scalars.

While parsing, the view now records where the integer digits, the
fraction digits, and the exponent of a float token are, so floats
and doubles with at most 19 digits are read from that layout with
the library's decimal_to_float() instead of rescanning the token.
Both round correctly, so the values are those of parse(). get<double>(),
materialize(), dump(), and the comparisons all use it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-06 10:48:06 +02:00