dump() writes a float token of at most 15 significant digits from its
digits, without converting it to a double and back: two decimals of at
most 15 digits are farther apart than the rounding interval of a
normal double (the argument behind DBL_DIG), so the token's digits are
the shortest ones of its double, which the library's conversion (Zmij)
writes. The exponent must keep the value away from subnormals and
overflow. Longer tokens are converted from the digits already read.
Doubles are written into the output directly instead of through a
local buffer. With NEON, the fixed layouts ("12.5", "0.001", "100.0")
are put together in vector registers by a table lookup of the digit
bytes: the portable layout copies the digits through a buffer at
another offset, and a load that spans several recent stores waits
until they reach the cache.
dump() of float-heavy documents: numbers -69%, marine_ik -62%,
mesh.pretty -34%, canada (mostly 16 or 17 digits) -14%.
Tests: 20,000 float tokens of 1 to 17 significant digits in every
spelling (point, exponent, leading and trailing zeros, sign), from about
1e-320 to 1e300, written as json::dump() writes them. On AArch64 they
check the NEON layout; x86 and JSON_VIEW_NO_SIMD use the library's.
Other float types, now the only ones on the general path, are tested
with non-finite values set by edits (written as null).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).
- save() writes an edited document in its current state, in document
order (floats that are not finite become null, as in dump()); the
same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
load(vector&&) keeps it without a copy. The nodes are copied (aligned,
and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
bounds, UTF-8, strings of the source, number tokens and their values);
bounds checks structure and bounds, so that reading and serializing
stay safe; none trusts the image.
A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.
As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.
The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.
Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Preparation for editable documents, without a change in behavior: views and
documents get a template parameter Editable (false by default), and every
walk over the index (iterators, lookups, dump(), materialize()) goes
through detail::view::navigation<Editable>. For read-only documents it is
the plain node array, as before, so they compile without any of the edit
handling. For editable documents it also follows the representation of
edits, which this commit defines:
- node flags `edited` (a string or number token in the edit arena),
`moved` (the elements of an array/object live in a separate sequence),
and `is_new` (no source position), and link nodes (kind_link) that
stand for a value stored elsewhere
- document_data::edit_state: the moved sequences, the storage of new
values, and the edit arena
materialize() now keeps a frame per open container instead of returning
to the end of a closed one, as the serializer does, so that it can
follow moved sequences. Floats whose token lives in the edit arena (also
"nan", "inf", "-inf") are converted out of line. dump() copies only
strings of the source without escaping, and shrink_to_fit() leaves the
node array in place once there are edits, as they link into it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For doubles with at most 19 digits, the
value is now read from that layout: the digits eight at a time, without
scanning the token, and rounded with Clinger's fast path where both
operands are exact, else with the Eisel-Lemire algorithm (which needs no
fallback for up to 19 digits). Both round correctly, so the values are
those of parse(); other tokens and types keep the library's conversion.
get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.53 -> 0.86 GB/s.
Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22) to the
bit-for-bit comparison with parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The public classes of the zero-copy view (#5295), in the new header
<nlohmann/json_view.hpp>:
- basic_json_document<BasicJsonType>: parse (borrowing contiguous byte
inputs, owning rvalue strings, streams, and other inputs), parse_copy,
accept, read, root, is_discarded, source, owns_source, node_count,
memory_usage, shrink_to_fit
- basic_json_view<BasicJsonType>: type and the is_* queries, size, empty,
materialize (the value parse() would produce, built by the same SAX
handler), source_offset
- the aliases json_document, json_view, ordered_json_document, and
ordered_json_view
A parse error throws the exception basic_json::parse would throw for the
same input: the library parser is run on the failing input, so messages,
positions, and exception ids are the same. Inputs of 4 GiB or more are
rejected with out_of_range.416.
The single header single_include/nlohmann/json_view.hpp keeps including
json.hpp; make amalgamate, check-amalgamation, include.zip, and release
handle it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>