An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).
- save() writes an edited document in its current state, in document
order (floats that are not finite become null, as in dump()); the
same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
load(vector&&) keeps it without a copy. The nodes are copied (aligned,
and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
bounds, UTF-8, strings of the source, number tokens and their values);
bounds checks structure and bounds, so that reading and serializing
stay safe; none trusts the image.
A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.
As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.
The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.
Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Preparation for editable documents, without a change in behavior: views and
documents get a template parameter Editable (false by default), and every
walk over the index (iterators, lookups, dump(), materialize()) goes
through detail::view::navigation<Editable>. For read-only documents it is
the plain node array, as before, so they compile without any of the edit
handling. For editable documents it also follows the representation of
edits, which this commit defines:
- node flags `edited` (a string or number token in the edit arena),
`moved` (the elements of an array/object live in a separate sequence),
and `is_new` (no source position), and link nodes (kind_link) that
stand for a value stored elsewhere
- document_data::edit_state: the moved sequences, the storage of new
values, and the edit arena
materialize() now keeps a frame per open container instead of returning
to the end of a closed one, as the serializer does, so that it can
follow moved sequences. Floats whose token lives in the edit arena (also
"nan", "inf", "-inf") are converted out of line. dump() copies only
strings of the source without escaping, and shrink_to_fit() leaves the
node array in place once there are edits, as they link into it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For floats and doubles with at most 19
digits, the value is now read from that layout: the digits eight at a time,
without scanning the token, and rounded by the library's conversion core
(detail::decimal_to_float(): Clinger's fast path where both operands are
exact, else the Eisel-Lemire algorithm, which needs no fallback for up to 19
digits). It rounds correctly, so the values are those of parse(); other
tokens and types keep the library's conversion of the whole token.
get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.95 -> 1.29 GB/s.
Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22, and
those of float) to the bit-for-bit comparison with parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The output buffer initializes its members in the initializer list, and the
escaping has no nested conditional operators; the test marks a fixed seed.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.
The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.
Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>