Compare commits

...
Author SHA1 Message Date
Niels Lohmann cf2af0396d Scan strings vector-first on x86-64 and avoid a stall when nesting
- With SSE2, string runs are checked 16 bytes at a time from their first
  byte: one compare finds the end of most keys and short values, faster on
  x86-64 than a branch per byte. AArch64 keeps the byte steps (there, a
  NEON mask costs more and the branches predict well; vector-first was 20%
  slower on Apple M1).
- open() stores the parent's frame field by field. Built on the stack and
  copied, it was read back by loads wider than its stores, which waited for
  them (store forwarding fails): 18% of the time on citm_catalog.json.

Parse on x86-64 (GCC 13 / Clang 18, us): twitter 352 -> 306 / 316 -> 288,
citm_catalog 1000 -> 770 / 843 -> 692, canada 1941 -> 1687 / 1978 -> 1722.
Unchanged on Apple M1.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-06 07:22:15 +02:00
Niels Lohmann 44fba02ec2 Format the digits of numbers in vector registers and without reloads
- zmij::to_shortest() keeps the last digit apart from the 15 or 16 digits
  before it, and dtoa_impl::write_shortest() converts those at once: with
  SSE2 on x86-64 and NEON on 64-bit Arm (no CPU check needed), else eight
  digits at a time. The decimal point is inserted in the register; reading
  the digits back right after storing them stalled store forwarding.
- json::dump() writes floats and integers straight into its write buffer
  instead of copying them from number_buffer, and integers below 10^16
  eight digits at a time.
- json_view's dump() writes doubles from their bits and tokens of up to 15
  digits through the same code; its own NEON writer is removed.

to_chars() on canada.json: 35 -> 29 ns per double (x86-64), 18.8 -> 14.2 ns
(Apple M1); json::dump() of canada.json 16-29% faster, of citm_catalog.json
14-19%.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-05 22:28:58 +02:00
Niels Lohmann 42f8043413 Make the json_view comparison fair to fresh documents and robust
- compare.py: make --data, --corpus, and --build-dir absolute, since the
  benchmarks run in the build directory; download into a .part file and
  remove an archive whose SHA-256 does not match, so that an interrupted
  download is not kept
- bench_view/bench_corpus/bench_edit: report files that cannot be opened instead of
  aborting; run each engine once untimed before its timed call, so that
  no engine pays for the allocator cleaning up after the previous one
  (with glibc, json_view after json::parse looked 1.7x slower on
  citm_catalog traverse); add "simdjson DOM (fresh)" and time
  "json_view (reused)" for traverse and select too
- README: explain fresh vs. reused documents and page faults on Linux

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-05 20:25:36 +02:00
Niels Lohmann 606c1e710b Validate non-ASCII strings with SSSE3 on every x86-64 CPU that has it
The vector UTF-8 check of json_view needed SSSE3 at compile time
(JSON_VIEW_USE_SSSE3 with -mssse3), so default x86-64 builds validated
non-ASCII text one sequence at a time. The check is now compiled for
SSSE3 with a function attribute (GCC 4.9 and later, Clang; MSVC compiles
the intrinsics anyway) and used where CPUID reports SSSE3. The answer is
kept in an atomic that is initialized at compile time, so neither a
guard nor a global constructor is needed. The definitions do not depend
on compiler flags, so there is no ODR issue. JSON_VIEW_USE_SSSE3 now only
skips the CPU check.

On x86-64 Linux (Haswell), twitter.json parses 23% faster with Clang 18
and 25-34% faster with GCC 13, now ahead of yyjson.

Also always inline read_eight_bytes() and parse_eight_digits(): GCC
called both in the number loops of the lexer and of json_view (52 call
sites), which cost about 10% on citm_catalog.json at -O2.

Document that reusing a document with read() avoids the page faults of
a fresh node index (about 40% of a 55 MB parse on x86-64 Linux).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-05 20:16:24 +02:00
Niels Lohmann fbacf1f1a8 Merge remote-tracking branch 'origin/json-view/23-zmij' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:23:52 +02:00
Niels Lohmann 173fbc0c90 Merge remote-tracking branch 'origin/json-view/22-view-dump-fast' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	docs/mkdocs/docs/api/basic_json/dump.md
2026-10-04 17:23:43 +02:00
Niels Lohmann 2db0e307ca Merge remote-tracking branch 'origin/json-view/21-images' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:21:05 +02:00
Niels Lohmann 923917bb7d Merge remote-tracking branch 'origin/json-view/20-edit-structure' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:21:00 +02:00
Niels Lohmann 77730529ab Merge remote-tracking branch 'origin/json-view/19-edit-set' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:20:55 +02:00
Niels Lohmann 05afd8e8e0 Merge branch 'json-view/23-zmij' into json-view/24-view-token-digits
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:45:09 +02:00
Niels Lohmann c908630082 Merge branch 'json-view/22-view-dump-fast' into json-view/23-zmij
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:45:04 +02:00
Niels Lohmann c99f99ccca Merge branch 'json-view/21-images' into json-view/22-view-dump-fast
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:45:00 +02:00
Niels Lohmann 268ef063d3 Mark Infer false positives in unit-json_view_image
Infer 1.3.0 reports NULLPTR_DEREFERENCE for passing a freshly parsed document to check_round_trip(). Suppress it on those two lines, as develop does for its own Infer false positives (#5750).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:44:55 +02:00
Niels Lohmann 4068a5a4d8 Move image_check's description into load()'s Notes
check_structure.py (added to develop in #5638) only allows the standard API page sections, so the separate "image_check" section failed it. Its content now opens the Notes section, which directly follows; also wrapped a line that exceeded 160 characters.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:44:33 +02:00
Niels Lohmann e0ae5040dc Merge branch 'json-view/20-edit-structure' into json-view/21-images
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:44:22 +02:00
Niels Lohmann f579e7e055 Merge branch 'json-view/19-edit-set' into json-view/20-edit-structure
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:44:17 +02:00
Niels Lohmann 4a93680068 Merge branch 'json-view/23-zmij' into json-view/24-view-token-digits
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:29:56 +02:00
Niels Lohmann 1fdef286dc Merge branch 'json-view/22-view-dump-fast' into json-view/23-zmij
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:29:33 +02:00
Niels Lohmann 650f3c091e Merge branch 'json-view/21-images' into json-view/22-view-dump-fast
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:29:25 +02:00
Niels Lohmann 4dd4505e1c Merge branch 'json-view/20-edit-structure' into json-view/21-images
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:29:17 +02:00
Niels Lohmann d3ba92d4fb Merge branch 'json-view/19-edit-set' into json-view/20-edit-structure
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:29:07 +02:00
Niels Lohmann 56cf446840 Merge branch 'json-view/23-zmij' into json-view/24-view-token-digits
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:18 +02:00
Niels Lohmann babae3f35d Merge branch 'json-view/22-view-dump-fast' into json-view/23-zmij
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:16 +02:00
Niels Lohmann b3ce6fcce3 Merge branch 'json-view/21-images' into json-view/22-view-dump-fast
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:14 +02:00
Niels Lohmann cff2af5966 Merge branch 'json-view/20-edit-structure' into json-view/21-images
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:12 +02:00
Niels Lohmann ad248a3290 Merge branch 'json-view/19-edit-set' into json-view/20-edit-structure
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:20:10 +02:00
Niels Lohmann 414d378bb4 Fix CI: useless casts in the float token test of json_view
GCC -Werror=useless-cast on Linux x86-64 rejected
static_cast<std::size_t>(tokens() % n): std::mt19937_64 yields
std::uint_fast64_t, which is std::size_t there. Draw the numbers through
a lambda that casts a named std::uint64_t, which also makes the
conversions for std::string's count explicit where std::size_t is
32 bits wide. The sequence of draws is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:21 +02:00
Niels Lohmann 7e459c5366 Write the floats of json_view from their digits
dump() writes a float token of at most 15 significant digits from its
digits, without converting it to a double and back: two decimals of at
most 15 digits are farther apart than the rounding interval of a
normal double (the argument behind DBL_DIG), so the token's digits are
the shortest ones of its double, which the library's conversion (Zmij)
writes. The exponent must keep the value away from subnormals and
overflow. Longer tokens are converted from the digits already read.

Doubles are written into the output directly instead of through a
local buffer. With NEON, the fixed layouts ("12.5", "0.001", "100.0")
are put together in vector registers by a table lookup of the digit
bytes: the portable layout copies the digits through a buffer at
another offset, and a load that spans several recent stores waits
until they reach the cache.

dump() of float-heavy documents: numbers -69%, marine_ik -62%,
mesh.pretty -34%, canada (mostly 16 or 17 digits) -14%.

Tests: 20,000 float tokens of 1 to 17 significant digits in every
spelling (point, exponent, leading and trailing zeros, sign), from about
1e-320 to 1e300, written as json::dump() writes them. On AArch64 they
check the NEON layout; x86 and JSON_VIEW_NO_SIMD use the library's.
Other float types, now the only ones on the general path, are tested
with non-finite values set by edits (written as null).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:21 +02:00
Niels Lohmann 0a97d94497 Fix CI: useless cast in the Zmij digit writer and snprintf truncation
- ci_test_gcc (Linux x86-64): static_cast<std::size_t>(d.significand % 100)
  was a useless cast (a std::uint64_t prvalue, the same type as
  std::size_t there); cast a named variable instead.
- ci_test_gcc: -Werror=format-truncation for snprintf("%.*e") in
  unit-to_chars.cpp, whose precision GCC cannot bound; write the
  neighboring decimal with a stream (classic locale, std::scientific),
  which gives the same text.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:19 +02:00
Niels Lohmann bee66ec810 Write doubles with the shortest digits (Zmij)
dump() writes doubles with the conversion of Zmij by Victor Zverovich
(https://github.com/vitaut/zmij, MIT), ported to C++11 in
detail/conversions/zmij.hpp: the shortest decimal in the rounding
interval, the closest one if there are several. Grisu2 does not always
find the shortest digits; about 0.14% of random doubles are now written
differently (0.08% with fewer digits, 0.06% with the closest last
digit); short decimals such as 0.1 or 2555.56 are not affected. float
keeps Grisu2.

The layout of doubles is unchanged, but written differently: the digits
are converted eight at a time (the BCD conversion of Xiang JunBo, as in
Zmij) and stored with one byte swap per eight digits; leading and
trailing zeros are counted from those bytes; and the layouts of
format_buffer() are written with fixed-size moves instead of per-digit
loops and moves of the buffer (to_chars() uses a local buffer if the
caller's is shorter than the 41 bytes this may write).

The powers of ten come from the table for number parsing, adjusted
where it holds them rounded up, and from the compressed tables of Zmij
beyond 10^308. json::dump() gets faster on floats: canada -53%,
numbers -46%, mesh -37%, marine_ik -30%.

Tests: the powers of ten recomputed with a small big-integer; for random
doubles, all powers of two and of ten and their neighbors, and boundary
values: the output reads back as the same value, no decimal with one
digit fewer does, the layout equals that of format_buffer() for the same
digits, and (C++17) the digits equal those of std::to_chars.
The size ratios of canada.json in unit-binary_formats.cpp and one
expectation in unit-to_chars.cpp change with the shorter output.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:18 +02:00
Niels Lohmann 1df24a0bd9 Write compact dumps of json_view without a library call per token
The default dump() (no indentation, no ensure_ascii) gets its own
writer: the same walk and output, with the write position in a local
variable (stores through char pointers would otherwise force a reload
of the buffer's members after each one), strings and number tokens of
the source copied by fixed-size moves of 32 bytes where the source has
that many bytes left (the buffer keeps 64 bytes of slack), and decoded
strings copied in runs up to the next quote, backslash, or control
character. Documents that are not edited are walked through the node
array in order, so that a frame only needs the end of its container,
and integer tokens are read from the source directly. The innermost
open container is kept in local variables, and the stack holds only
the ones around it; the stack starts in a local array of 32 and moves
to the heap only for deeper nesting (its address does not escape, so
its pointers stay in registers). Dumps of shallow documents thus
allocate only the output, whose first size includes the slack, so it
does not grow just before the end.

The long copies are out of line: otherwise, the compiler merges the
fixed-size moves into the same library call.

These techniques come from the prototype; the writer lost them when the
view was split into pull requests, which made dump() 2 to 3 times
slower.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:17 +02:00
Niels Lohmann a279e6f4b9 Fix CI: image tests with GCC on Linux, without exceptions, and on clang 3.6
- ci_test_gcc (Linux x86-64): static_cast<std::size_t>(header_field(...))
  was a useless cast (std::uint64_t is std::size_t there), and returning
  std::mt19937::result_type (std::uint_fast32_t, unsigned long there) as
  std::uint32_t failed -Werror=conversion. Cast named variables instead.
- ci_test_noexceptions: the error, check, and damaged-image tests test the
  exceptions of load() and save() and catch outside a CHECK_THROWS, which
  aborts with JSON_NOEXCEPTION; compile them and their helpers only with
  exceptions.
- clang 3.6: value-initialize a const json_document (no user-provided
  default constructor, CWG 253).
- Format the image fuzzer with the pinned astyle, which the "check" job
  runs over tests/.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:16 +02:00
Niels Lohmann b61c04bbe4 Document how images store the node index of json_view
The node index section on the architecture page says how save() writes
the nodes and that a change of their layout must raise the image
version, and save's format note links to it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:15 +02:00
Niels Lohmann 779ae7fffc Address the cpplint findings of json_document images
The exponent of the overflow check is an std::int64_t instead of a long
(runtime/int).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:14 +02:00
Niels Lohmann a59d1f64e9 Document the images of json_documents
Add API pages for basic_json_document::save() and load(), including the
image_check enumeration and its three levels. Add examples that show
caching a parsed document as an image, loading it without parsing, the
difference between a borrowed and an owned image, and a full check
rejecting a damaged image that a bounds check still reads safely.

Add an "Images" section to the json_view feature page, register the new
pages in mkdocs.yml and docSet.sql, group basic_json_document's member
list by parsing/access/images/edits, and document parse_error.116 and
type_error.320 on the exceptions page, extending out_of_range.416 for
images.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:14 +02:00
Niels Lohmann 8db804ca8f Add images of json_documents: save() and load()
An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).

- save() writes an edited document in its current state, in document
  order (floats that are not finite become null, as in dump()); the
  same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
  load(vector&&) keeps it without a copy. The nodes are copied (aligned,
  and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
  bounds, UTF-8, strings of the source, number tokens and their values);
  bounds checks structure and bounds, so that reading and serializing
  stay safe; none trusts the image.

A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.

As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.

The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.

Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:13 +02:00
Niels Lohmann a496c709f4 Document insert() and erase() of editable json_documents
Add API reference pages for basic_json_document::insert and
basic_json_document::erase, matching the style of set.md and
push_back.md: signatures, parameters, return values, exception safety,
exceptions with their exact ids and messages, complexity, notes on
duplicate keys and view/iterator validity, and an example.

Add example programs basic_json_document__insert.cpp and
basic_json_document__erase.cpp with their expected output, each
comparing an edit on an editable document with the same edit on a
plain json value to show what is preserved: member order, the
spelling of untouched numbers, and, for insert, that a view taken
before the insert keeps referring to the same element after its index
shifts.

Register both new pages in mkdocs.yml, docSet.sql, and the member list
of basic_json_document/index.md, add cross-references to them from
set.md and push_back.md, and mention insert/erase in the "Editing a
document" section of features/json_view.md.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:12 +02:00
Niels Lohmann 57bdb2f67f Address the clang-tidy findings of the structural edit tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:11 +02:00
Niels Lohmann b6d6d4996a Benchmark editing documents with json_view, yyjson, Boost.JSON, and json
bench_edit.cpp joins the comparison: parse, apply the same logical edits
with each library's own API (a handful at fixed places, or one in every
record), and serialize; all outputs must describe the same value.
compare.py builds and runs it with the other two programs.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:10 +02:00
Niels Lohmann dcc81f43f0 Add insert() and erase() to editable documents
- insert(array, index, value): insert before an element (index <= size)
- erase(object, key): remove all members with the key; returns their
  number
- erase(array, index): remove an element
- erase(json_pointer): remove the member or element a pointer names

The errors are those of basic_json (type_error.307/309,
out_of_range.401/403/405). A view of an erased value keeps its last value,
and views of other values keep referring to them when elements move.

Tests: the differential test now also inserts and erases members and
elements, directly and through JSON pointers; plus the errors, views
across inserts and erasures, duplicate keys, and large objects.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:02:09 +02:00
62 changed files with 7452 additions and 434 deletions
+1 -1
View File
@@ -51,7 +51,7 @@ labels:
- "include/nlohmann/detail/view/.*" - "include/nlohmann/detail/view/.*"
- "single_include/nlohmann/json_view\\.hpp" - "single_include/nlohmann/json_view\\.hpp"
- "tests/src/unit-json_view.*" - "tests/src/unit-json_view.*"
- "tests/src/fuzzer-parse_json_view\\.cpp" - "tests/src/fuzzer-(parse_json_view|json_view_image)\\.cpp"
- "tests/benchmarks/json_view/.*" - "tests/benchmarks/json_view/.*"
- "tools/amalgamate/config_json_view\\.json" - "tools/amalgamate/config_json_view\\.json"
- "docs/mkdocs/docs/features/json_view\\.md" - "docs/mkdocs/docs/features/json_view\\.md"
+2
View File
@@ -26,6 +26,7 @@ cc_library(
"include/nlohmann/detail/conversions/from_json.hpp", "include/nlohmann/detail/conversions/from_json.hpp",
"include/nlohmann/detail/conversions/to_chars.hpp", "include/nlohmann/detail/conversions/to_chars.hpp",
"include/nlohmann/detail/conversions/to_json.hpp", "include/nlohmann/detail/conversions/to_json.hpp",
"include/nlohmann/detail/conversions/zmij.hpp",
"include/nlohmann/detail/exceptions.hpp", "include/nlohmann/detail/exceptions.hpp",
"include/nlohmann/detail/hash.hpp", "include/nlohmann/detail/hash.hpp",
"include/nlohmann/detail/input/binary_reader.hpp", "include/nlohmann/detail/input/binary_reader.hpp",
@@ -72,6 +73,7 @@ cc_library(
"include/nlohmann/detail/view/edit.hpp", "include/nlohmann/detail/view/edit.hpp",
"include/nlohmann/detail/view/edit_storage.hpp", "include/nlohmann/detail/view/edit_storage.hpp",
"include/nlohmann/detail/view/errors.hpp", "include/nlohmann/detail/view/errors.hpp",
"include/nlohmann/detail/view/image.hpp",
"include/nlohmann/detail/view/input.hpp", "include/nlohmann/detail/view/input.hpp",
"include/nlohmann/detail/view/iterator.hpp", "include/nlohmann/detail/view/iterator.hpp",
"include/nlohmann/detail/view/lookup.hpp", "include/nlohmann/detail/view/lookup.hpp",
+1
View File
@@ -1399,6 +1399,7 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I
- The class contains the UTF-8 Decoder from Bjoern Hoehrmann which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2008-2009 [Björn Hoehrmann](https://bjoern.hoehrmann.de/) <bjoern@hoehrmann.de> - The class contains the UTF-8 Decoder from Bjoern Hoehrmann which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2008-2009 [Björn Hoehrmann](https://bjoern.hoehrmann.de/) <bjoern@hoehrmann.de>
- The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/) - The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/)
- The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2025 [Victor Zverovich](https://github.com/vitaut)
- The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). - The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
- The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0). - The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0).
- The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors - The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
+4
View File
@@ -135,7 +135,10 @@ INSERT INTO searchIndex(name, type, path) VALUES ('basic_json::~basic_json', 'Me
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document', 'Class', 'api/basic_json_document/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document', 'Class', 'api/basic_json_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::basic_json_document', 'Constructor', 'api/basic_json_document/basic_json_document/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::basic_json_document', 'Constructor', 'api/basic_json_document/basic_json_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::accept', 'Function', 'api/basic_json_document/accept/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::accept', 'Function', 'api/basic_json_document/accept/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::erase', 'Method', 'api/basic_json_document/erase/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::insert', 'Method', 'api/basic_json_document/insert/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::is_discarded', 'Method', 'api/basic_json_document/is_discarded/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::is_discarded', 'Method', 'api/basic_json_document/is_discarded/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::load', 'Function', 'api/basic_json_document/load/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::memory_usage', 'Method', 'api/basic_json_document/memory_usage/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::memory_usage', 'Method', 'api/basic_json_document/memory_usage/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::node_count', 'Method', 'api/basic_json_document/node_count/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::node_count', 'Method', 'api/basic_json_document/node_count/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::owns_source', 'Method', 'api/basic_json_document/owns_source/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::owns_source', 'Method', 'api/basic_json_document/owns_source/index.html');
@@ -144,6 +147,7 @@ INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::parse_co
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::push_back', 'Method', 'api/basic_json_document/push_back/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::push_back', 'Method', 'api/basic_json_document/push_back/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::read', 'Method', 'api/basic_json_document/read/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::read', 'Method', 'api/basic_json_document/read/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::root', 'Method', 'api/basic_json_document/root/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::root', 'Method', 'api/basic_json_document/root/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::save', 'Method', 'api/basic_json_document/save/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::set', 'Method', 'api/basic_json_document/set/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::set', 'Method', 'api/basic_json_document/set/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::shrink_to_fit', 'Method', 'api/basic_json_document/shrink_to_fit/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::shrink_to_fit', 'Method', 'api/basic_json_document/shrink_to_fit/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::source', 'Method', 'api/basic_json_document/source/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::source', 'Method', 'api/basic_json_document/source/index.html');
+5
View File
@@ -62,6 +62,9 @@ Linear.
## Notes ## Notes
Floating-point numbers are written with the fewest digits that read back as the same value (for `#!cpp double`; see
[number handling](../../features/types/number_handling.md#number-serialization)).
Binary values are serialized as an object containing two keys: Binary values are serialized as an object containing two keys:
- "bytes": an array of bytes as integers - "bytes": an array of bytes as integers
@@ -99,3 +102,5 @@ Binary values are serialized as an object containing two keys:
- Error handlers added in version 3.4.0. - Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0. - Serialization of binary values added in version 3.8.0.
- Error handler `keep` added in version 3.13.0. - Error handler `keep` added in version 3.13.0.
- Doubles are written with the shortest digits (Żmij instead of Grisu2) since version 3.13.0; about 0.1% of doubles are
written differently, most of them with fewer digits.
@@ -0,0 +1,128 @@
# <small>nlohmann::basic_json_document::</small>erase
```cpp
// (1)
std::size_t erase(view_type object, string_view_t key);
// (2)
template<typename I>
void erase(view_type array, I idx);
// (3)
std::size_t erase(const json_pointer& ptr);
```
Only an **editable** document (`#!cpp Editable == true`, e.g. [`json_editable_document`](../json_editable_document.md))
has `erase`; calling it on a read-only `basic_json_document` fails to compile (`#!cpp static_assert`).
1. Removes every member of `object` whose key is `key` (see [Notes](#notes) on duplicate keys) and returns how many
were removed; `#!cpp 0` if `object` has no member with this key.
2. Removes the element at index `idx` of `array`, which must already exist (`#!cpp idx < array.size()`).
3. Removes the value the JSON pointer `ptr` refers to, relative to [`root()`](root.md), and returns how many values
were removed: the *parent* of the target must already exist, and the target itself is removed as in 1. (an object
member; `#!cpp 0` or more) or 2. (an array element; always `#!cpp 1`). `ptr` must not be empty -- [`root()`](root.md)
itself cannot be erased.
## Template parameters
`I`
: an integral type other than `#!cpp bool`, deduced (overloads taking a `#!cpp bool` or a non-integral type for
`idx` do not participate in overload resolution).
## Parameters
`object` (in)
: the object to remove a member of
`array` (in)
: the array to remove an element of
`key` (in)
: the key of the member(s) to remove
`idx` (in)
: the index of the element to remove; a negative value throws (see [Exceptions](#exceptions))
`ptr` (in)
: a JSON pointer to the value to remove, relative to `root()`
## Return value
1. the number of removed members (`#!cpp 0` if `object` had none with this `key`)
2. (nothing)
3. the number of removed values (`#!cpp 0` or more for an object member, always `#!cpp 1` for an array element)
## Exceptions
1. Throws [`type_error.307`](../../home/exceptions.md#jsonexceptiontype_error307) if `object` is not an object -- the
same message [`BasicJsonType::erase`](../basic_json/erase.md) throws for the same type.
2. Throws `type_error.307` if `array` is not an array. Throws
[`out_of_range.401`](../../home/exceptions.md#jsonexceptionout_of_range401) if `idx` is negative, or if
`#!cpp idx >= array.size()`.
3. Throws [`out_of_range.405`](../../home/exceptions.md#jsonexceptionout_of_range405) ("JSON pointer has no parent")
if `ptr` is empty. Throws what [`at`](../basic_json_view/at.md) throws (overload 3) for resolving `ptr`'s parent.
For the last reference token itself: if the parent is an array, throws what 2. throws for an index that is out of
range, or, for a token that is not a valid array index,
[`parse_error.106`](../../home/exceptions.md#jsonexceptionparse_error106) (a leading `#!cpp '0'`),
[`parse_error.109`](../../home/exceptions.md#jsonexceptionparse_error109) (not a number),
[`out_of_range.410`](../../home/exceptions.md#jsonexceptionout_of_range410) (too large for `size_type`), or
[`out_of_range.404`](../../home/exceptions.md#jsonexceptionout_of_range404) (an empty token); otherwise (an
object, or a primitive value the pointer's parent resolves to) throws what 1. throws.
Every overload also throws [`invalid_iterator.202`](../../home/exceptions.md#jsonexceptioninvalid_iterator202) ("view
does not belong to this document") if `object`/`array` is a [discarded](../basic_json_view/is_discarded.md) view or a
view of a *different* document (overloads 1-2 only; overload 3 always starts from this document's own
[`root()`](root.md)).
## Complexity
1. Linear in the number of members of `object`.
2. Linear in the number of elements of `array` at or after `idx` (they move one slot over).
3. Linear in the number of reference tokens of `ptr` and, for each token, in the number of members of the object at
that level or the index into the array (as [`at`](../basic_json_view/at.md)), plus the complexity of 1. or 2. for
the last token.
## Notes
!!! info "Duplicate keys"
Overload 1. removes *every* member with `key`, not just the first -- unlike [`set`](set.md), which assigns the
first occurrence and drops the rest. This is why it returns a count rather than a single view: there may be
more than one member removed, or none.
Like [`set`](set.md) and [`push_back`](push_back.md), `erase` never moves an element's *value*: a view still
referring to a removed member or element keeps showing what it last held (see [Edits](index.md#edits)) -- it just no
longer appears when `array`/`object` is read, dumped, or iterated. Removing an element of `array` (2.) does shift the
*links* to the elements after it, the same way `insert`, `set`, or `push_back` on the same array would; any iterator
already taken over `array`/`object` is invalidated by an erase, since it was walking the old layout.
## Examples
??? example
The example below drops a deprecated field and a decommissioned entry from a configuration document -- using all
three overloads -- and shows what stays intact that would not with a plain `json`/`ordered_json` value: the order
of the fields around the ones removed, and the exact spelling of a number that was never touched.
```cpp
--8<-- "examples/basic_json_document__erase.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__erase.output"
```
## See also
- [insert](insert.md) - insert an element into an array
- [set](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers to
- [push_back](push_back.md) - append to an array
- [root](root.md) - the view of the root value, the starting point of overload 3
- [`BasicJsonType::erase`](../basic_json/erase.md) - the corresponding function of `basic_json`
- [Edits](index.md#edits) - what an edit guarantees, for every overload
## Version history
- Added in version 3.13.0.
@@ -19,10 +19,10 @@ it (a copy, or an rvalue `#!cpp std::string` that was moved in); see [`owns_sour
is move-only: copying a document would either duplicate a potentially large index and text, or leave two documents is move-only: copying a document would either duplicate a potentially large index and text, or leave two documents
claiming to borrow the same buffer, so it is disabled. claiming to borrow the same buffer, so it is disabled.
With `#!cpp Editable == true`, the document also offers [`set`](set.md) and [`push_back`](push_back.md) to change With `#!cpp Editable == true`, the document also offers [`set`](set.md), [`push_back`](push_back.md),
values in place, see [Edits](#edits) below. The source text itself is never written; a read-only document [`insert`](insert.md), and [`erase`](erase.md) to change values in place, see [Edits](#edits) below. The source text
(`#!cpp Editable == false`, the default) does not carry any of the bookkeeping edits need, and calling `set` or itself is never written; a read-only document (`#!cpp Editable == false`, the default) does not carry any of the
`push_back` on one fails to compile (`#!cpp static_assert`). bookkeeping edits need, and calling any of them on one fails to compile (`#!cpp static_assert`).
## Template parameters ## Template parameters
@@ -32,8 +32,8 @@ values in place, see [Edits](#edits) below. The source text itself is never writ
is checked with a `static_assert`. is checked with a `static_assert`.
`Editable` `Editable`
: whether the document supports [`set`](set.md) and [`push_back`](push_back.md) (optional, `#!cpp false` by : whether the document supports [`set`](set.md), [`push_back`](push_back.md), [`insert`](insert.md), and
default). See [Edits](#edits) below. [`erase`](erase.md) (optional, `#!cpp false` by default). See [Edits](#edits) below.
## Specializations ## Specializations
@@ -52,10 +52,16 @@ values in place, see [Edits](#edits) below. The source text itself is never writ
## Member functions ## Member functions
- [(constructor)](basic_json_document.md) - [(constructor)](basic_json_document.md)
### Parsing
- [**parse**](parse.md) (_static_) - deserialize from a compatible input, borrowing or owning it as appropriate - [**parse**](parse.md) (_static_) - deserialize from a compatible input, borrowing or owning it as appropriate
- [**parse_copy**](parse_copy.md) (_static_) - deserialize a copy of a compatible input - [**parse_copy**](parse_copy.md) (_static_) - deserialize a copy of a compatible input
- [**accept**](accept.md) (_static_) - check whether the input is valid JSON - [**accept**](accept.md) (_static_) - check whether the input is valid JSON
- [**read**](read.md) - (re-)parse into this document, reusing its memory - [**read**](read.md) - (re-)parse into this document, reusing its memory
### Access
- [**root**](root.md) - the view of the root value - [**root**](root.md) - the view of the root value
- [**is_discarded**](is_discarded.md) - return whether the last parse failed - [**is_discarded**](is_discarded.md) - return whether the last parse failed
- [**source**](source.md) - the parsed text - [**source**](source.md) - the parsed text
@@ -63,14 +69,26 @@ values in place, see [Edits](#edits) below. The source text itself is never writ
- [**node_count**](node_count.md) - the number of index entries (values plus object keys) - [**node_count**](node_count.md) - the number of index entries (values plus object keys)
- [**memory_usage**](memory_usage.md) - the number of bytes held by the document - [**memory_usage**](memory_usage.md) - the number of bytes held by the document
- [**shrink_to_fit**](shrink_to_fit.md) - release unused index capacity - [**shrink_to_fit**](shrink_to_fit.md) - release unused index capacity
### Images
- [**save**](save.md) - the document as an image that `load()` reads without parsing
- [**load**](load.md) (_static_) - read an image written by `save()`
### Edits
- [**set**](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers - [**set**](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers
to (`#!cpp Editable` documents only) to (`#!cpp Editable` documents only)
- [**push_back**](push_back.md) - append to an array (`#!cpp Editable` documents only) - [**push_back**](push_back.md) - append to an array (`#!cpp Editable` documents only)
- [**insert**](insert.md) - insert an element into an array before a given position (`#!cpp Editable` documents only)
- [**erase**](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
(`#!cpp Editable` documents only)
## Edits ## Edits
An editable document (`#!cpp Editable == true`) can be changed after parsing, with [`set`](set.md) and An editable document (`#!cpp Editable == true`) can be changed after parsing, with [`set`](set.md),
[`push_back`](push_back.md); [`json_editable_document`](../json_editable_document.md) and [`push_back`](push_back.md), [`insert`](insert.md), and [`erase`](erase.md);
[`json_editable_document`](../json_editable_document.md) and
[`ordered_json_editable_document`](../ordered_json_editable_document.md) are the corresponding specializations. A few [`ordered_json_editable_document`](../ordered_json_editable_document.md) are the corresponding specializations. A few
points apply to every edit: points apply to every edit:
@@ -0,0 +1,114 @@
# <small>nlohmann::basic_json_document::</small>insert
```cpp
template<typename I, typename V>
view_type insert(view_type array, I idx, V&& value);
```
Only an **editable** document (`#!cpp Editable == true`, e.g. [`json_editable_document`](../json_editable_document.md))
has `insert`; calling it on a read-only `basic_json_document` fails to compile (`#!cpp static_assert`).
Inserts `value` into `array` as a new element before position `idx`, which must not be past the end
(`#!cpp idx <= array.size()`; `#!cpp idx == array.size()` appends, like [`push_back`](push_back.md)). Unlike
[`push_back`](push_back.md), a [null](../basic_json_view/is_null.md) `array` does *not* first become an empty array:
`array` must already be an array.
`value` is accepted three ways: a [`basic_json_view`](../basic_json_view/index.md) of *any* document -- read-only or
editable, and it does not have to be `array`'s own document -- which is copied so that nothing is shared with the
source document afterward; a `BasicJsonType` value; or anything `BasicJsonType` can be constructed from (numbers,
strings, `#!cpp bool`, `#!cpp nullptr`, containers, ...).
## Template parameters
`I`
: an integral type other than `#!cpp bool`, deduced (overloads taking a `#!cpp bool` or a non-integral type for
`idx` do not participate in overload resolution).
`V`
: the type of `value`, deduced; see above for what is accepted.
## Parameters
`array` (in)
: the array to insert into
`idx` (in)
: the position to insert `value` before; a negative value throws (see [Exceptions](#exceptions))
`value` (in)
: the value to insert
## Return value
a view of the new element, now holding `value`
## Exception safety
Basic exception safety: `value` is fully encoded -- including the checks below -- into storage owned by the document
before anything already reachable from [`root()`](root.md) is touched, so a failure while encoding `value` (an
invalid argument, or `#!cpp std::bad_alloc`) leaves the document completely unchanged, other than memory allocated
for the encoding that is not reclaimed. A failure of a later allocation -- while `array` switches from its parsed
layout to a growable block, or while that block grows, see [Notes](#notes) -- can still leave `array` already
switched to that layout even though `value` itself was not inserted.
## Exceptions
Throws [`type_error.309`](../../home/exceptions.md#jsonexceptiontype_error309) if `array` is not an array -- the same
message [`BasicJsonType::insert`](../basic_json/insert.md) throws for the same type; a null `array` throws this too
(see above). Throws [`out_of_range.401`](../../home/exceptions.md#jsonexceptionout_of_range401) if `idx` is negative,
or if `#!cpp idx > array.size()`. Throws
[`invalid_iterator.202`](../../home/exceptions.md#jsonexceptioninvalid_iterator202) ("view does not belong to this
document") if `array` is a [discarded](../basic_json_view/is_discarded.md) view or a view of a *different* document.
Throws [`type_error.302`](../../home/exceptions.md#jsonexceptiontype_error302) if `value` is a
[discarded](../basic_json_view/is_discarded.md) view or a [discarded](../basic_json/is_discarded.md) `BasicJsonType`
value, and [`type_error.319`](../../home/exceptions.md#jsonexceptiontype_error319) if `value` is (or contains) a
binary value -- `BasicJsonType` can hold one, but a `json_document` cannot. Throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) if `value` is (or contains) a string that is
not valid UTF-8, with the same message [`BasicJsonType::dump()`](../basic_json/dump.md) gives for that string.
## Complexity
Linear in the number of elements of `array` at or after `idx` (they move one slot over), plus time linear in the
size of `value` to encode it into the document's storage (constant for a scalar, linear in the number of nested
values for an array or object): like [`push_back`](push_back.md), the elements of `array` move to a growable block
of links the first time it is inserted into (or [`set`](set.md)/[`push_back`](push_back.md) on), and that block
grows in amortized constant time; inserting before the end within that block still shifts every later element.
## Notes
Like [`set`](set.md) on a member or an element, `insert` never moves an existing *element's value* -- only where
`array`'s *links* to its elements live -- so a view of an existing element of `array` stays valid across an
`insert`, and keeps referring to the same element even though its index shifts. Any iterator already taken over
`array` is invalidated, since it was walking the old layout. See [Edits](index.md#edits) for what stays valid across
an edit in general.
## Examples
??? example
The example below inserts a step into the middle of a deployment plan, without touching the steps that come
after it, and shows that a view taken before the insert keeps referring to the same element even though its
index shifts -- something a plain `json`/`ordered_json` array, or its `std::vector`-based storage, has no
equivalent for.
```cpp
--8<-- "examples/basic_json_document__insert.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__insert.output"
```
## See also
- [push_back](push_back.md) - append to an array
- [erase](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
- [set](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers to
- [`BasicJsonType::insert`](../basic_json/insert.md) - the corresponding function of `basic_json`
- [Edits](index.md#edits) - what an edit guarantees, for every overload
## Version history
- Added in version 3.13.0.
@@ -0,0 +1,169 @@
# <small>nlohmann::basic_json_document::</small>load
```cpp
// (1)
static basic_json_document load(const std::uint8_t* image, std::size_t size,
const image_check check = image_check::full);
// (2)
static basic_json_document load(const std::vector<std::uint8_t>& image,
const image_check check = image_check::full);
// (3)
static basic_json_document load(std::vector<std::uint8_t>&& image,
const image_check check = image_check::full);
```
1. Reads an image [`save()`](save.md) wrote, from a pointer and a byte count. The image is **borrowed**: `image`
must stay alive and unchanged for as long as the returned document, and any view taken from it, is used.
2. Reads an image from a `#!cpp std::vector`. Also **borrowed** -- equivalent to overload 1 called with
`#!cpp image.data()` and `#!cpp image.size()`.
3. Reads an image, keeping the vector instead of copying it: `image` is moved into the document (no copy), which
then owns it for as long as it needs the text and the decoded strings. [`owns_source()`](owns_source.md) is
`#!cpp true` afterward.
In every overload, the node index is copied into storage the document itself owns -- so that it is properly aligned,
and, for an [editable](index.md#edits) document, can be edited -- while the text and the decoded strings stay in
`image`. The hash indexes [large objects](../../features/json_view.md) use for lookup are rebuilt, exactly as after
parsing.
## Parameters
`image` (in)
: the image [`save()`](save.md) wrote (overloads 1 and 2), or one to take ownership of (overload 3)
`size` (in)
: the number of bytes at `image` (overload 1)
`check` (in)
: how thoroughly to validate `image` before trusting it; see [`image_check`](#image_check) below (optional,
`#!cpp image_check::full` by default)
## Return value
The document read from the image.
## Exception safety
Overloads 1 and 2 give the strong guarantee: `image` is only read, never written, so a thrown exception leaves the
caller's buffer untouched.
Overload 3 moves `image` into the document *before* validating it, so that a good image is kept without a copy. If
loading then fails, the partially built document -- and the vector now inside it -- is discarded along with the
exception, and `image` itself is left **empty**, not restored to what was passed in. Move a copy in instead, or
validate with overload 2 first, if the original vector must survive a failed load.
## Exceptions
On a big-endian target, throws [`type_error.320`](../../home/exceptions.md#jsonexceptiontype_error320) -- the same
exception [`save()`](save.md#exceptions) throws there, since the image format is little-endian only.
Otherwise throws [`parse_error.116`](../../home/exceptions.md#jsonexceptionparse_error116) if `image` is not one
`save()` could have written, or fails the requested `check`:
| message | when |
|------------------------|------------------------------------------------------------------------------------------------------|
| `too short` | `image` is `#!cpp nullptr`, or `size` is smaller than the 64-byte header |
| `unknown format` | the header's magic bytes or version do not match, or a reserved header field is not zero |
| `sizes out of range` | the node count, text size, or decoded-string size the header describes does not fit `size`, or the `#!cpp '\0'` after the text or after the decoded strings is missing |
| `the check failed` | `check` is not `#!cpp image_check::none`, and the image fails it -- see [`image_check`](#image_check) |
!!! failure "Example messages"
```
[json.exception.parse_error.116] parse error: invalid json_document image: too short
```
```
[json.exception.parse_error.116] parse error: invalid json_document image: unknown format
```
```
[json.exception.parse_error.116] parse error: invalid json_document image: sizes out of range
```
```
[json.exception.parse_error.116] parse error: invalid json_document image: the check failed
```
## Complexity
Linear in the number of nodes, which are always copied into the document. With `#!cpp check == image_check::full`,
additionally linear in the combined length of the text and the decoded strings; `#!cpp image_check::bounds` and
`#!cpp image_check::none` do not read them.
## Notes
**The `image_check` modes.**
```cpp
using image_check = detail::view::image_check;
enum class image_check
{
full,
bounds,
none
};
```
How thoroughly `load()` validates `image` before trusting it.
| value | checks | guarantees |
|----------|--------------------------------------------------------------------------------------------------------------|------------|
| `full` | everything the parser itself guarantees: structure and bounds; that every string is valid UTF-8 (and, for a string still in the source text, that it contains no quote, backslash, or control character); and that every number token is well-formed and matches the value stored for it | reading and serializing a checked image is safe and always produces valid JSON, exactly as for a parsed document |
| `bounds` | structure and bounds only -- that every offset and count in the node index stays inside the image | reading and serializing stay memory-safe, but a crafted image can hold strings that are not valid UTF-8 or that serialize to invalid JSON ([`dump()`](../basic_json_view/dump.md) writes them unchanged or throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)), and numbers whose values differ from their text |
| `none` | nothing | images from a trusted source only -- reading a damaged image is undefined behavior |
`full` is the default and the right choice for an image from anything you do not fully control -- a file, a cache
shared with other processes, a peer on the network. `bounds` skips scanning the text and the decoded strings, so it
fits a cache your own process just wrote and reads straight back, where damage would mean a bug or a hardware fault
rather than adversarial input; it still cannot crash or read out of bounds. `none` skips validation entirely and
should only be used for an image you trust as much as your own memory.
**Lifetime.** Overloads 1 and 2 borrow `image`: it must stay alive and byte-for-byte unchanged for as long as the
returned document, and any [view](../basic_json_view/index.md) taken from it, is used -- exactly like a document
[`parse()`](parse.md) borrowed its input for. Overload 3 avoids this by keeping the vector itself; see
[`owns_source`](owns_source.md).
!!! warning "Experimental"
The image format is versioned but not yet stable, and may change in an incompatible way before it is declared
stable; `load()` already rejects an image written by a different format version with `parse_error.116`
("unknown format"). Use images to cache a document within one build of the library, or to hand one to another
process running the *same* build on the *same* (little-endian) machine -- not as a long-term storage format.
**What `image_check::bounds` does not guarantee.** A bounds-checked image can never make `load()`,
[`root()`](root.md), element access, or [`materialize()`](../basic_json_view/materialize.md) read outside the image,
so those stay safe on a damaged one. It does *not* guarantee that the image describes valid JSON: a string
that a `full` check would have rejected can make [`dump()`](../basic_json_view/dump.md) write invalid UTF-8 or invalid
JSON, or throw `type_error.316`, and a number can read back with a value that does not match how it is spelled.
Reserve `bounds` for images you already trust to be well-formed, and use it only to skip the extra scan.
## Examples
??? example "Caching a document, ownership, and a rejected image"
The example below saves a parsed document as an image, checks that `load()` reproduces the original
[`dump()`](../basic_json_view/dump.md) without parsing, and shows the difference between
`load(std::move(image))` (owned) and `load(image)` (borrowed). It then damages one byte of the image and shows
`image_check::full` rejecting it with `parse_error.116`, while `image_check::bounds` -- meant for a cache the
process already trusts -- still reads it without going out of bounds.
```cpp
--8<-- "examples/basic_json_document__load.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__load.output"
```
## See also
- [save](save.md) - write the document as an image
- [owns_source](owns_source.md) - return whether the document holds its own copy of the text
- [parse](parse.md) - deserialize from JSON text instead of an image
- [Images](../../features/json_view.md#images) - why and when to use images
## Version history
- Added in version 3.13.0.
@@ -132,6 +132,7 @@ integer type becomes a floating-point value.
- [accept](accept.md) - check whether the input is valid JSON - [accept](accept.md) - check whether the input is valid JSON
- [read](read.md) - (re-)parse into this document, reusing its memory - [read](read.md) - (re-)parse into this document, reusing its memory
- [owns_source](owns_source.md) - return whether the document holds its own copy of the text - [owns_source](owns_source.md) - return whether the document holds its own copy of the text
- [load](load.md) - read a document from an image instead of parsing JSON text
- [`BasicJsonType::parse`](../basic_json/parse.md) - the corresponding function of `basic_json` - [`BasicJsonType::parse`](../basic_json/parse.md) - the corresponding function of `basic_json`
## Version history ## Version history
@@ -91,6 +91,8 @@ Like [`set`](set.md) on a member or an element, `push_back` never moves an exist
## See also ## See also
- [set](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers to - [set](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers to
- [insert](insert.md) - insert an element into an array before a given position
- [erase](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
- [root](root.md) - the view of the root value - [root](root.md) - the view of the root value
- [`BasicJsonType::push_back`](../basic_json/push_back.md) - the corresponding function of `basic_json` - [`BasicJsonType::push_back`](../basic_json/push_back.md) - the corresponding function of `basic_json`
- [Edits](index.md#edits) - what an edit guarantees, for every overload - [Edits](index.md#edits) - what an edit guarantees, for every overload
@@ -49,6 +49,11 @@ whether or not the new parse succeeds; take fresh views from [`root()`](root.md)
`input` is borrowed or owned by the same rules as [`parse()`](parse.md#notes); a document can borrow on one call and `input` is borrowed or owned by the same rules as [`parse()`](parse.md#notes); a document can borrow on one call and
own on the next, since ownership is decided freshly each time. own on the next, since ownership is decided freshly each time.
Reusing a document matters most for large inputs: the operating system provides the memory of a fresh node index one
page at a time, and every page costs a page fault the first time it is written. On x86-64 Linux (4 KiB pages), parsing
a 55 MB document into a reused document took about 40 % less time than parsing it into a fresh one. Programs that parse
many documents of similar size should therefore keep one document and call `read()`.
## Examples ## Examples
??? example ??? example
@@ -70,6 +75,7 @@ own on the next, since ownership is decided freshly each time.
- [parse](parse.md) - deserialize from a compatible input - [parse](parse.md) - deserialize from a compatible input
- [root](root.md) - the view of the root value - [root](root.md) - the view of the root value
- [load](load.md) - read a document from an image instead of parsing JSON text
## Version history ## Version history
@@ -0,0 +1,104 @@
# <small>nlohmann::basic_json_document::</small>save
```cpp
std::vector<std::uint8_t> save() const;
```
Writes the document as an *image*: a byte buffer that [`load`](load.md) reads back without parsing. The image holds
the node index, the source text (plus, for an edited document, the number tokens edits wrote), and the decoded
strings (plus the strings edits wrote) -- everything [`root()`](root.md) needs, with nothing left to parse.
An edited document is written in its *current* state, with its values in document order, the way the library's own
parser would have produced them for that JSON text: a member [`set`](set.md) added goes at the end, an
[`erase`](erase.md)d member leaves no trace, and a float that is not finite (NaN or positive/negative infinity)
becomes null, the same substitution [`dump()`](../basic_json_view/dump.md) makes. The same document always saves to
the same bytes -- also across `BasicJsonType` and `#!cpp Editable`, since the image reflects document order and
values only, not which specialization produced them.
## Return value
The image, as a `#!cpp std::vector<std::uint8_t>`. Pass it, or a pointer to its data together with its size, to
[`load`](load.md) to read the document back.
## Exception safety
Strong guarantee: `save()` does not modify `#!cpp *this` (it is `#!cpp const`), so if it throws, the document is left
exactly as it was, and the partially built image is discarded with the exception.
## Exceptions
Throws [`type_error.320`](../../home/exceptions.md#jsonexceptiontype_error320) if the document is
[discarded](is_discarded.md) -- a default-constructed document, or one a failed [`parse()`](parse.md)/
[`read()`](read.md) with `allow_exceptions == false` left discarded.
On a big-endian target, throws `type_error.320` with a different message instead: the image format is little-endian
only (see [Notes](#notes)).
Throws [`out_of_range.416`](../../home/exceptions.md#jsonexceptionout_of_range416) if the node count, the text, or
the decoded strings of the image would individually reach 4 GiB -- the same 32-bit offsets
[`parse()`](parse.md#exceptions) and, for edits, [`set`](set.md)/[`push_back`](push_back.md) are already limited to.
!!! failure "Example messages"
```
[json.exception.type_error.320] cannot save a discarded json_document
```
```
[json.exception.type_error.320] json_document images need a little-endian target
```
```
[json.exception.out_of_range.416] images of 4 GiB or more are not supported by json_document
```
## Complexity
Linear in the size of the document: the number of nodes, plus the length of the text and the decoded strings that end
up in the image.
## Notes
**Format.** The image begins with a 64-byte header (the magic bytes `#!cpp "NJVI"`, a version number, the node count,
and the sizes of the text and the decoded strings, all little-endian), followed by the nodes
([16 bytes each](../../home/architecture.md#node-index-of-json-views)), the text and a `#!cpp '\0'`, and the decoded
strings and a `#!cpp '\0'`. [`load`](load.md) checks the header, and the sizes it describes, before reading anything
else -- see [`load`'s Exceptions](load.md#exceptions).
!!! warning "Experimental"
The image format is versioned but not yet stable: it may change in an incompatible way before it is declared
stable. Use images to cache a document within one build of the library, or to hand one to another process running
the *same* build on the *same* (little-endian) machine -- not as a long-term storage format. Keep the original
JSON text if you need to read a saved document back with a future library version.
**Little-endian only.** The image is written as raw little-endian bytes, with no byte-swapping. `save()` (and
[`load`](load.md)) throw `type_error.320` on a big-endian target rather than silently produce bytes a big-endian
reader could not interpret correctly.
## Examples
??? example "Caching a parsed document as an image"
The example below saves a parsed configuration as an image -- the way a service might cache one to answer later
requests without parsing the text again -- and confirms that loading it back gives exactly the same result as
parsing did, and that saving is deterministic.
```cpp
--8<-- "examples/basic_json_document__save.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__save.output"
```
## See also
- [load](load.md) - read an image written by `save()`
- [owns_source](owns_source.md) - return whether the document holds its own copy of the text
- [`basic_json_view::dump`](../basic_json_view/dump.md) - serialize the document to JSON text instead of an image
- [Images](../../features/json_view.md#images) - why and when to use images
## Version history
- Added in version 3.13.0.
@@ -180,6 +180,8 @@ one-node scalar.
## See also ## See also
- [push_back](push_back.md) - append to an array - [push_back](push_back.md) - append to an array
- [insert](insert.md) - insert an element into an array
- [erase](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
- [root](root.md) - the view of the root value, the starting point of overload 4 - [root](root.md) - the view of the root value, the starting point of overload 4
- [`basic_json_view::dump`](../basic_json_view/dump.md) - serialize the document, keeping an untouched number's - [`basic_json_view::dump`](../basic_json_view/dump.md) - serialize the document, keeping an untouched number's
spelling with `#!cpp number_format::source` spelling with `#!cpp number_format::source`
@@ -5,13 +5,17 @@
``` ```
When defined on x86-64, the parser of [`basic_json_document`](../basic_json_document/index.md) When defined on x86-64, the parser of [`basic_json_document`](../basic_json_document/index.md)
(`<nlohmann/json_view.hpp>`) validates non-ASCII text in strings with SSSE3, 16 bytes at a time, using the "lookup4" (`<nlohmann/json_view.hpp>`) validates non-ASCII text in strings with SSSE3 without asking the CPU first.
algorithm of [simdjson](https://github.com/simdjson/simdjson). Without it, non-ASCII text is validated one UTF-8
sequence at a time on x86-64; on AArch64, the vector check uses NEON and is always on.
SSSE3 is not part of the x86-64 baseline, so the code must be compiled for it: define the macro only together with a By default, the parser checks once at run time whether the CPU has SSSE3 (all x86-64 CPUs since about 2011 have it)
compiler option that enables SSSE3 (e.g. `-mssse3`, or `-march=` with a CPU that has it), and only for programs that and then validates non-ASCII text 16 bytes at a time, using the "lookup4" algorithm of
run on such CPUs. The same input is accepted or rejected either way; only the speed of non-ASCII text differs. [simdjson](https://github.com/simdjson/simdjson); on CPUs without SSSE3, it validates one UTF-8 sequence at a time.
The vector check is compiled for SSSE3 with a function attribute (GCC 4.9 and later, Clang), so this needs no compiler
option. With MSVC, the check uses `__cpuid`. On AArch64, the vector check uses NEON and is always on.
Define the macro only together with a compiler option that enables SSSE3 (e.g. `-mssse3`, or `-march=` with a CPU that
has it), and only for programs that run on such CPUs. It saves the check of the CPU, which costs little. The same
input is accepted or rejected either way; only the speed of non-ASCII text differs.
!!! warning "Define consistently" !!! warning "Define consistently"
@@ -0,0 +1,38 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_editable_document = nlohmann::json_editable_document;
using json_editable_view = nlohmann::json_editable_view;
int main()
{
// a deprecated field is dropped from a configuration file, and a
// decommissioned replica is removed from the list -- "price" keeps its
// trailing zero, and the fields around the removed ones keep their order
const std::string text = R"({
"name": "cache",
"legacy_host": "db0",
"host": "db1",
"price": 19.90,
"replicas": ["db2", "db3", "db4"]
})";
json_editable_document doc = json_editable_document::parse(text);
doc.erase(doc.root(), "legacy_host"); // (1) an object member
doc.erase(doc.root()["replicas"], 1); // (2) an array element ("db3")
const std::size_t removed = doc.erase(json::json_pointer("/replicas/0")); // (3) via a JSON pointer
std::cout << removed << '\n';
std::cout << doc.root().dump(2, ' ', false, json_editable_view::number_format::source) << "\n\n";
// the same edits on a plain json value: object_t is a std::map, so
// parsing already sorted the keys, and dump() rewrites every number to
// its shortest form, even "price", which was never touched
json plain = json::parse(text);
plain.erase("legacy_host");
plain["replicas"].erase(1);
plain["replicas"].erase(0);
std::cout << plain.dump(2) << '\n';
}
@@ -0,0 +1,18 @@
1
{
"name": "cache",
"host": "db1",
"price": 19.90,
"replicas": [
"db4"
]
}
{
"host": "db1",
"name": "cache",
"price": 19.9,
"replicas": [
"db4"
]
}
@@ -0,0 +1,38 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_editable_document = nlohmann::json_editable_document;
using json_editable_view = nlohmann::json_editable_view;
int main()
{
// a deployment plan -- "budget" is written with a trailing zero that has
// no effect on its value
const std::string text = R"({
"release": "2026.09",
"steps": ["build", "test", "deploy"],
"budget": 19.90
})";
json_editable_document doc = json_editable_document::parse(text);
const std::size_t deploy_index = 2;
const auto deploy = doc.root()["steps"][deploy_index]; // held across the insert
doc.insert(doc.root()["steps"], deploy_index, "smoke-test"); // insert before "deploy"
// the held view still refers to "deploy", even though its index moved
// from 2 to 3, and nothing else in the document was touched
std::cout << deploy.dump() << '\n';
std::cout << doc.root().dump(2, ' ', false, json_editable_view::number_format::source) << "\n\n";
// the same edit on a plain json value: an index held from before the
// insert now refers to whatever moved into that slot, and dump()
// rewrites "budget" to its shortest form even though it was never
// touched
json plain = json::parse(text);
plain["steps"].insert(plain["steps"].begin() + static_cast<std::ptrdiff_t>(deploy_index), "smoke-test");
std::cout << plain["steps"][deploy_index].dump() << '\n';
std::cout << plain.dump(2) << '\n';
}
@@ -0,0 +1,23 @@
"deploy"
{
"release": "2026.09",
"steps": [
"build",
"test",
"smoke-test",
"deploy"
],
"budget": 19.90
}
"smoke-test"
{
"budget": 19.9,
"release": "2026.09",
"steps": [
"build",
"test",
"smoke-test",
"deploy"
]
}
@@ -0,0 +1,52 @@
#include <cstdint>
#include <iostream>
#include <string>
#include <vector>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_document = nlohmann::json_document;
using image_check = json_document::image_check;
int main()
{
std::cout << std::boolalpha;
// the image of a parsed document -- as if read back from a cache file or
// received from another process running the same build of the library
const std::string text = R"({"name": "cache", "note": "caf\u00e9", "replicas": ["db2", "db3"]})";
const json_document parsed = json_document::parse(text);
const std::vector<std::uint8_t> image = parsed.save();
// (1)/(2) load() needs no parsing, yet dumps exactly what parsing did
const json_document borrowed = json_document::load(image);
std::cout << (borrowed.root().dump() == parsed.root().dump()) << '\n';
std::cout << borrowed.owns_source() << '\n'; // borrowed: still points into `image`
// (3) load(std::move(image)) keeps the vector instead of copying it
std::vector<std::uint8_t> to_move = image;
const json_document owned = json_document::load(std::move(to_move));
std::cout << owned.owns_source() << '\n';
// a damaged image -- the last byte of the decoded string "note" holds
// (an escape sequence, so it was unescaped into the document's own
// buffer), flipped, as storage or transport corruption might do
std::vector<std::uint8_t> damaged = image;
damaged[damaged.size() - 2] = 0xFF;
// image_check::full inspects strings and numbers, so it catches the damage
try
{
static_cast<void>(json_document::load(damaged, image_check::full));
}
catch (const json::parse_error& e)
{
std::cout << e.id << '\n';
}
// image_check::bounds only checks structure and bounds, so a cache the
// process already trusts loads without the extra scan -- reading a value
// the damage did not touch is still safe
const json_document trusted = json_document::load(damaged, image_check::bounds);
std::cout << trusted.root()["name"].get<std::string>() << '\n';
}
@@ -0,0 +1,5 @@
true
false
true
116
cache
@@ -0,0 +1,30 @@
#include <cstdint>
#include <iostream>
#include <string>
#include <vector>
#include <nlohmann/json_view.hpp>
using json_document = nlohmann::json_document;
int main()
{
std::cout << std::boolalpha;
// a configuration a service parses once and then caches as an image, so
// that later requests can load() it instead of parsing the text again
const std::string text = R"({"name": "cache", "host": "db1", "port": 6379, "replicas": ["db2", "db3"]})";
const json_document config = json_document::parse(text);
// save() turns the parsed document into a byte buffer: a 64-byte header,
// the node index, the source text, and the decoded strings
const std::vector<std::uint8_t> image = config.save();
std::cout << image.size() << '\n';
// the same document always saves to the same bytes
std::cout << (image == json_document::parse(text).save()) << '\n';
// loading the image back needs no parsing, yet dumps exactly what
// parsing the text produced
const json_document reloaded = json_document::load(image);
std::cout << (reloaded.root().dump() == config.root().dump()) << '\n';
}
@@ -0,0 +1,3 @@
316
true
true
+2 -1
View File
@@ -14,7 +14,8 @@ C++ types, and finally serialize it again.
[parsing untrusted input](parsing/untrusted_input.md). [parsing untrusted input](parsing/untrusted_input.md).
- [Zero-copy JSON views](json_view.md) — read a JSON text through a flat index instead of building a `json` tree; - [Zero-copy JSON views](json_view.md) — read a JSON text through a flat index instead of building a `json` tree;
strings and numbers stay in the input and are only decoded when needed. strings and numbers stay in the input and are only decoded when needed.
[Editable documents](json_view.md#editing-a-document) can also be modified. [Editable documents](json_view.md#editing-a-document) can also be modified, and [images](json_view.md#images) load a
parsed document again without parsing it.
- [Comments](comments.md) and [trailing commas](trailing_commas.md) — opt-in relaxations of the JSON grammar. - [Comments](comments.md) and [trailing commas](trailing_commas.md) — opt-in relaxations of the JSON grammar.
## Accessing and modifying values ## Accessing and modifying values
+66 -10
View File
@@ -193,11 +193,13 @@ Everything above is read-only: a `json_document`/`json_view` lets you look at a
not change it. [`basic_json_document<BasicJsonType, true>`](../api/basic_json_document/index.md) -- more conveniently not change it. [`basic_json_document<BasicJsonType, true>`](../api/basic_json_document/index.md) -- more conveniently
spelled [`json_editable_document`](../api/json_editable_document.md) or spelled [`json_editable_document`](../api/json_editable_document.md) or
[`ordered_json_editable_document`](../api/ordered_json_editable_document.md) -- also lets you [`ordered_json_editable_document`](../api/ordered_json_editable_document.md) -- also lets you
[`set`](../api/basic_json_document/set.md) a value and [`push_back`](../api/basic_json_document/push_back.md) onto [`set`](../api/basic_json_document/set.md) a value, [`push_back`](../api/basic_json_document/push_back.md) onto or
an array, still without ever building a `basic_json` tree for parts you do not touch. [`insert`](../api/basic_json_document/insert.md) into an array, and [`erase`](../api/basic_json_document/erase.md)
an object member or an array element, still without ever building a `basic_json` tree for parts you do not touch.
`#!cpp Editable` defaults to `#!cpp false`, so `json_document`/`ordered_json_document` are unaffected -- they carry `#!cpp Editable` defaults to `#!cpp false`, so `json_document`/`ordered_json_document` are unaffected -- they carry
none of the bookkeeping edits need, and calling `set`/`push_back` on one is a compile error, not a runtime one. none of the bookkeeping edits need, and calling `set`/`push_back`/`insert`/`erase` on one is a compile error, not a
runtime one.
### Why: editing without reformatting ### Why: editing without reformatting
@@ -212,8 +214,9 @@ or `ordered_json` value in place lossy. Say you parse a configuration file, patc
An editable document keeps both. [`dump()`](../api/basic_json_view/dump.md) of an edited document writes members in An editable document keeps both. [`dump()`](../api/basic_json_view/dump.md) of an edited document writes members in
document order -- a member [`set`](../api/basic_json_document/set.md) added goes at the end, exactly where it was document order -- a member [`set`](../api/basic_json_document/set.md) added goes at the end, exactly where it was
inserted -- and [`number_format::source`](../api/basic_json_view/number_format.md) keeps the exact spelling of inserted, and an [`erase`](../api/basic_json_document/erase.md)d member simply leaves a gap: everything around it
every number an edit did not itself touch; a number an edit *did* touch is written the way keeps its place -- and [`number_format::source`](../api/basic_json_view/number_format.md) keeps the exact spelling
of every number an edit did not itself touch; a number an edit *did* touch is written the way
[`BasicJsonType::dump()`](../api/basic_json/dump.md) would write it, since there is no source spelling for a brand [`BasicJsonType::dump()`](../api/basic_json/dump.md) would write it, since there is no source spelling for a brand
new value. new value.
@@ -236,24 +239,77 @@ parsed into for as long as it is not itself replaced. So every [view](../api/bas
an edit, including a previously obtained [`root()`](../api/basic_json_document/root.md), stays valid and, if it an edit, including a previously obtained [`root()`](../api/basic_json_document/root.md), stays valid and, if it
still refers to the edited value, sees the edit; a view of a value a later edit drops or replaces just keeps showing still refers to the edited value, sees the edit; a view of a value a later edit drops or replaces just keeps showing
what it last held. New values go to storage the document allocates and owns on demand. The one thing an edit does what it last held. New values go to storage the document allocates and owns on demand. The one thing an edit does
invalidate is the **iterators** taken over an edited array or object: the first time one of its elements is set or invalidate is the **iterators** taken over an edited array or object: the first time one of its elements is set,
appended to, its elements move from the parsed, fixed layout to a growable block of links so that appended to, inserted into, or erased, its elements move from the parsed, fixed layout to a growable block of links
[`push_back`](../api/basic_json_document/push_back.md) can later grow it in amortized constant time -- existing so that [`push_back`](../api/basic_json_document/push_back.md) can later grow it in amortized constant time --
elements are not touched, but an iterator that was walking the old layout no longer matches. A string obtained with existing elements are not touched, but an iterator that was walking the old layout no longer matches. A string
obtained with
[`get_string()`](../api/basic_json_view/get_string.md) is unaffected either way and stays valid across further [`get_string()`](../api/basic_json_view/get_string.md) is unaffected either way and stays valid across further
edits. See [`basic_json_document`'s Edits](../api/basic_json_document/index.md#edits) for the details, and edits. See [`basic_json_document`'s Edits](../api/basic_json_document/index.md#edits) for the details, and
[`set`'s Exception safety](../api/basic_json_document/set.md#exception-safety) for what an edit guarantees if it [`set`'s Exception safety](../api/basic_json_document/set.md#exception-safety) for what an edit guarantees if it
throws (the *basic* guarantee, not the strong one `dump()` and the read-only functions provide). How edits are kept in throws (the *basic* guarantee, not the strong one `dump()` and the read-only functions provide). How edits are kept in
the index is described in the [architecture overview](../home/architecture.md#node-index-of-json-views). the index is described in the [architecture overview](../home/architecture.md#node-index-of-json-views).
## Images
[`save()`](../api/basic_json_document/save.md) writes a document as an *image*: a byte buffer that
[`load()`](../api/basic_json_document/load.md) reads back into a document without parsing -- no lexing, no building
the node index, nothing but copying the nodes and pointing the text and the decoded strings at the image. Where
[`parse_copy()`](../api/basic_json_document/parse_copy.md) still has to scan the whole input,
[`load()`](../api/basic_json_document/load.md) turns that scan into a copy of the node index alone.
**Why.** A document that is parsed once and then read many times -- a configuration loaded at startup, a template
rendered on every request, a large reference dataset a worker process needs in memory -- pays for parsing once but
can amortize [`save()`](../api/basic_json_document/save.md)'s cost across every later load. That makes images useful
for a cache: save a document the first time it is parsed (to a file, a shared-memory segment, an in-process cache),
and [`load()`](../api/basic_json_document/load.md) it on every later use instead of parsing the source text again.
They are just as useful for handing a parsed document to another process (or a forked worker) running the same build
of the library, since [`load()`](../api/basic_json_document/load.md) turns the transfer into a copy of the node index
plus pointers into the received bytes, not a re-parse.
**Choosing a check.** [`load()`](../api/basic_json_document/load.md) takes an
[`image_check`](../api/basic_json_document/load.md#image_check) that trades validation against speed:
`image_check::full` (the default) checks everything the parser itself guarantees, so a checked image is exactly as
safe to read and serialize as a freshly parsed document -- the right choice whenever the image did not come straight
from this process's own [`save()`](../api/basic_json_document/save.md), such as a file or a network peer.
`image_check::bounds` only checks structure and bounds -- cheaper, since it skips scanning the text and the decoded
strings -- and fits a cache the process trusts, one it wrote and reads back itself. `image_check::none` skips
validation entirely, for an image trusted as much as the process's own memory. See
[`load()`'s Notes](../api/basic_json_document/load.md#notes) for exactly what each level does and does not guarantee.
??? example "Example: cache a parsed configuration as an image"
```cpp
--8<-- "examples/basic_json_document__save.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__save.output"
```
!!! warning "Experimental"
The image format is versioned but not yet stable, and may change in an incompatible way before it is declared
stable. It is little-endian only, and tied to the library build that wrote it -- use it to cache a document or to
hand one to another process running the *same* build, not as a long-term storage format; keep the original JSON
text if a saved document needs to be readable by a future library version.
The idea of a document you can read without parsing comes from zero-copy formats such as
[FlatBuffers](https://github.com/google/flatbuffers) and [YaFF](https://github.com/yandex/yaff); the check
[`load()`](../api/basic_json_document/load.md) runs follows the idea of FlatBuffers' Verifier. No code is taken from
either.
## Choosing between `json`, `ordered_json`, the SAX interface, and `json_view` ## Choosing between `json`, `ordered_json`, the SAX interface, and `json_view`
| | [`json`](../api/json.md) / [`ordered_json`](../api/ordered_json.md) | [SAX interface](parsing/sax_interface.md) | [`json_document`](../api/json_document.md) / [`json_view`](../api/json_view.md) | [`json_editable_document`](../api/json_editable_document.md) / [`json_editable_view`](../api/json_editable_view.md) | | | [`json`](../api/json.md) / [`ordered_json`](../api/ordered_json.md) | [SAX interface](parsing/sax_interface.md) | [`json_document`](../api/json_document.md) / [`json_view`](../api/json_view.md) | [`json_editable_document`](../api/json_editable_document.md) / [`json_editable_view`](../api/json_editable_view.md) |
|---|---|---|---|---| |---|---|---|---|---|
| **Ownership** | owns every value | owns nothing; you decide what to keep, in your handler | borrows or owns the *text*; the index is always owned by the document | same as `json_document`; edits go to storage the document owns | | **Ownership** | owns every value | owns nothing; you decide what to keep, in your handler | borrows or owns the *text*; the index is always owned by the document | same as `json_document`; edits go to storage the document owns |
| **Mutability** | freely mutable | not applicable (a one-shot event stream) | read-only | [`set`](../api/basic_json_document/set.md)/[`push_back`](../api/basic_json_document/push_back.md) edit in place; the source text is never rewritten | | **Mutability** | freely mutable | not applicable (a one-shot event stream) | read-only | [`set`](../api/basic_json_document/set.md)/[`push_back`](../api/basic_json_document/push_back.md)/[`insert`](../api/basic_json_document/insert.md)/[`erase`](../api/basic_json_document/erase.md) edit in place; the source text is never rewritten |
| **What you get** | a full tree you can read, write, and keep as long as you like | a sequence of callbacks; whatever your handler builds from them | a flat index plus, on demand, [`materialize()`](../api/basic_json_view/materialize.md)d `json`/`ordered_json` values for the parts you actually use | the same, plus [`dump()`](../api/basic_json_view/dump.md) of an edited document that keeps the member order and, with [`number_format::source`](../api/basic_json_view/number_format.md), the spelling of every untouched number | | **What you get** | a full tree you can read, write, and keep as long as you like | a sequence of callbacks; whatever your handler builds from them | a flat index plus, on demand, [`materialize()`](../api/basic_json_view/materialize.md)d `json`/`ordered_json` values for the parts you actually use | the same, plus [`dump()`](../api/basic_json_view/dump.md) of an edited document that keeps the member order and, with [`number_format::source`](../api/basic_json_view/number_format.md), the spelling of every untouched number |
| **Typical use** | general-purpose JSON handling: config, request/response bodies you build or modify, anything you hold onto | validating or projecting a text into your own data structure without ever holding the whole thing as JSON | large or high-volume input where you only need part of it, or need it repeatedly, and can keep the source text (or a copy) alive for as long as the document lives | a document you read, patch a few fields of, and write back -- a configuration file, for instance -- where the rest of it should come back exactly as it was | | **Typical use** | general-purpose JSON handling: config, request/response bodies you build or modify, anything you hold onto | validating or projecting a text into your own data structure without ever holding the whole thing as JSON | large or high-volume input where you only need part of it, or need it repeatedly, and can keep the source text (or a copy) alive for as long as the document lives | a document you read, patch a few fields of, and write back -- a configuration file, for instance -- where the rest of it should come back exactly as it was |
| **Caching/reload** | not applicable -- re-parse, or roll your own serialization | not applicable | [`save()`](../api/basic_json_document/save.md)/[`load()`](../api/basic_json_document/load.md): cache the parsed index as an image and reload it without parsing | same, saving the document's current -- possibly edited -- state |
## Version history ## Version history
@@ -134,9 +134,10 @@ That is, `-0` is stored as a signed integer, but the serialization does not repr
### Number serialization ### Number serialization
- Integer numbers are serialized as is; that is, no scientific notation is used. - Integer numbers are serialized as is; that is, no scientific notation is used.
- Floating-point numbers are serialized as specified by the `#!c %g` printf modifier with - Floating-point numbers are serialized with the fewest digits that read back as the same value (the closest such
[`std::numeric_limits<double>::max_digits10`](https://en.cppreference.com/w/cpp/types/numeric_limits/max_digits10) digits if there are several), in the layout of the `#!c %g` printf modifier: `#!c 1.5`, `#!c 100.0`, `#!c 1e+100`.
significant digits. The rationale is to use the shortest representation while still allowing round-tripping. Doubles are converted with the algorithm of [Żmij](https://github.com/vitaut/zmij), floats with Grisu2, which
can write more digits than necessary.
!!! hint "Notes regarding precision of floating-point numbers" !!! hint "Notes regarding precision of floating-point numbers"
@@ -545,9 +545,10 @@ therefore silently changes parse results rather than raising an error. See
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
(`#!cpp float` is promoted to `#!cpp double`). (`#!cpp float` is promoted to `#!cpp double`).
If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary32 or binary64 number, `dump` uses the If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary64 number, `dump` uses the algorithm of
Grisu2 algorithm, which produces the shortest representation that round-trips. Otherwise the `snprintf` fallback with Żmij, which produces the shortest representation that round-trips. For IEEE 754 binary32 numbers, it uses Grisu2,
`max_digits10` digits is used. which produces a short representation that round-trips. Otherwise the `snprintf` fallback with `max_digits10` digits is
used.
### Required for the binary formats ### Required for the binary formats
@@ -559,7 +560,7 @@ binary32 or binary64 field and have no encoding for `#!cpp long double`.
| Type | Support | | Type | Support |
|--------------------------|-----------------------------------------------------------------------------------------------------------------------| |--------------------------|-----------------------------------------------------------------------------------------------------------------------|
| `#!cpp double` (default) | full; short round-trip output through Grisu2 | | `#!cpp double` (default) | full; shortest round-trip output through Żmij |
| `#!cpp float` | full; short round-trip output through Grisu2 | | `#!cpp float` | full; short round-trip output through Grisu2 |
| `#!cpp long double` | `dump` and `parse` only; the binary format writers do not compile, as they only handle IEEE 754 binary32 and binary64 | | `#!cpp long double` | `dump` and `parse` only; the binary format writers do not compile, as they only handle IEEE 754 binary32 and binary64 |
| any other type | not usable | | any other type | not usable |
+8
View File
@@ -259,6 +259,14 @@ edited:
- Views of read-only documents compile without any of this: how views walk the index is a template parameter - Views of read-only documents compile without any of this: how views walk the index is a template parameter
(`navigation<Editable>`). (`navigation<Editable>`).
Images ([`save`](../api/basic_json_document/save.md) and [`load`](../api/basic_json_document/load.md),
[`detail/view/image.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/view/image.hpp)) store
the nodes as they are: a 64-byte header (the magic bytes `NJVI`, a format version, the sizes, and reserved bytes that
must be zero), the nodes, the text, and the decoded strings. An edited document is first written in document order, as
the parser would have written it (without links), and the numbers of the hash indexes are cleared, since `load`
rebuilds the indexes. So a change of the node layout is a change of the image format: it must raise `image_version`,
and `load` then rejects images of other versions (`parse_error.116`) instead of misreading them.
## Input adapters ## Input adapters
Input is read via **input adapters** that abstract a source. Every input adapter provides this interface: Input is read via **input adapters** that abstract a source. Every input adapter provides this interface:
+43 -1
View File
@@ -391,6 +391,23 @@ A UBJSON high-precision number could not be parsed.
[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing UBJSON high-precision number: invalid number text: 1A [json.exception.parse_error.115] parse error at byte 5: syntax error while parsing UBJSON high-precision number: invalid number text: 1A
``` ```
### json.exception.parse_error.116
[`basic_json_document::load()`](../api/basic_json_document/load.md) rejected an
[image](../features/json_view.md#images): either the bytes are not one [`save()`](../api/basic_json_document/save.md)
could have written (too short, an unknown magic number or format version, or sizes that do not fit the buffer), or
they are, but fail the requested [`image_check`](../api/basic_json_document/load.md#image_check).
!!! failure "Example message"
```
[json.exception.parse_error.116] parse error: invalid json_document image: the check failed
```
!!! note
This exception was added in version 3.13.0, together with [images](../features/json_view.md#images).
## Iterator errors ## Iterator errors
This exception is thrown if iterators passed to a library function do not match This exception is thrown if iterators passed to a library function do not match
@@ -809,6 +826,26 @@ from JSON text.
This exception was added in version 3.13.0, together with editable [`json_document`s](../features/json_view.md). This exception was added in version 3.13.0, together with editable [`json_document`s](../features/json_view.md).
### json.exception.type_error.320
[`basic_json_document::save()`](../api/basic_json_document/save.md) cannot write an
[image](../features/json_view.md#images) of a [discarded](../api/basic_json_document/is_discarded.md) document.
[`save()`](../api/basic_json_document/save.md) and [`load()`](../api/basic_json_document/load.md) also throw this
exception on a big-endian target, since the image format is little-endian only.
!!! failure "Example messages"
```
[json.exception.type_error.320] cannot save a discarded json_document
```
```
[json.exception.type_error.320] json_document images need a little-endian target
```
!!! note
This exception was added in version 3.13.0, together with [images](../features/json_view.md#images).
## Out of range ## Out of range
This exception is thrown in case a library function is called on an input parameter that exceeds the expected range, for instance, in the case of array indices or nonexisting object keys. This exception is thrown in case a library function is called on an input parameter that exceeds the expected range, for instance, in the case of array indices or nonexisting object keys.
@@ -1036,7 +1073,9 @@ MessagePack's ext type and BSON's binary subtype are each stored in a single byt
so they do not support an input of 4 GiB or more. The same 32-bit limit applies to an **editable** document's own so they do not support an input of 4 GiB or more. The same 32-bit limit applies to an **editable** document's own
storage: [`set`](../api/basic_json_document/set.md) and [`push_back`](../api/basic_json_document/push_back.md) throw storage: [`set`](../api/basic_json_document/set.md) and [`push_back`](../api/basic_json_document/push_back.md) throw
this exception once the strings and number tokens written by edits reach 4 GiB in total, or once more than this exception once the strings and number tokens written by edits reach 4 GiB in total, or once more than
4294967295 arrays/objects have had an element set or appended to them. 4294967295 arrays/objects have had an element set or appended to them. The same limit applies to an
[image](../features/json_view.md#images): [`save()`](../api/basic_json_document/save.md) throws it if the node
count, the text, or the decoded strings it would write would individually reach 4 GiB.
!!! failure "Example messages" !!! failure "Example messages"
@@ -1046,6 +1085,9 @@ this exception once the strings and number tokens written by edits reach 4 GiB i
``` ```
[json.exception.out_of_range.416] edits of 4 GiB or more are not supported by json_document [json.exception.out_of_range.416] edits of 4 GiB or more are not supported by json_document
``` ```
```
[json.exception.out_of_range.416] images of 4 GiB or more are not supported by json_document
```
!!! note !!! note
+2
View File
@@ -18,6 +18,8 @@ The class contains the UTF-8 Decoder from Bjoern Hoehrmann which is licensed und
The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/) The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/)
The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2025 [Victor Zverovich](https://github.com/vitaut)
The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
+4
View File
@@ -237,7 +237,10 @@ nav:
- 'Overview': api/basic_json_document/index.md - 'Overview': api/basic_json_document/index.md
- '(Constructor)': api/basic_json_document/basic_json_document.md - '(Constructor)': api/basic_json_document/basic_json_document.md
- 'accept': api/basic_json_document/accept.md - 'accept': api/basic_json_document/accept.md
- 'erase': api/basic_json_document/erase.md
- 'insert': api/basic_json_document/insert.md
- 'is_discarded': api/basic_json_document/is_discarded.md - 'is_discarded': api/basic_json_document/is_discarded.md
- 'load': api/basic_json_document/load.md
- 'memory_usage': api/basic_json_document/memory_usage.md - 'memory_usage': api/basic_json_document/memory_usage.md
- 'node_count': api/basic_json_document/node_count.md - 'node_count': api/basic_json_document/node_count.md
- 'owns_source': api/basic_json_document/owns_source.md - 'owns_source': api/basic_json_document/owns_source.md
@@ -246,6 +249,7 @@ nav:
- 'push_back': api/basic_json_document/push_back.md - 'push_back': api/basic_json_document/push_back.md
- 'read': api/basic_json_document/read.md - 'read': api/basic_json_document/read.md
- 'root': api/basic_json_document/root.md - 'root': api/basic_json_document/root.md
- 'save': api/basic_json_document/save.md
- 'set': api/basic_json_document/set.md - 'set': api/basic_json_document/set.md
- 'shrink_to_fit': api/basic_json_document/shrink_to_fit.md - 'shrink_to_fit': api/basic_json_document/shrink_to_fit.md
- 'source': api/basic_json_document/source.md - 'source': api/basic_json_document/source.md
+4 -3
View File
@@ -86,8 +86,9 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc
} }
/// eight bytes as a little-endian word (compilers fold this into one load on /// eight bytes as a little-endian word (compilers fold this into one load on
/// little-endian targets) /// little-endian targets; always inlined, as GCC otherwise calls it in the
inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept /// number loops)
JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
{ {
return static_cast<std::uint64_t>(b[0]) | (static_cast<std::uint64_t>(b[1]) << 8u) return static_cast<std::uint64_t>(b[0]) | (static_cast<std::uint64_t>(b[1]) << 8u)
| (static_cast<std::uint64_t>(b[2]) << 16u) | (static_cast<std::uint64_t>(b[3]) << 24u) | (static_cast<std::uint64_t>(b[2]) << 16u) | (static_cast<std::uint64_t>(b[3]) << 24u)
@@ -96,7 +97,7 @@ inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
} }
/// eight bytes as a little-endian word /// eight bytes as a little-endian word
inline std::uint64_t read_eight_bytes(const char* p) noexcept JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const char* p) noexcept
{ {
return read_eight_bytes(reinterpret_cast<const unsigned char*>(p)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast) return read_eight_bytes(reinterpret_cast<const unsigned char*>(p)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
} }
+472 -23
View File
@@ -11,11 +11,32 @@
#include <array> // array #include <array> // array
#include <cmath> // signbit, isfinite #include <cmath> // signbit, isfinite
#include <cstddef> // size_t
#include <cstdint> // intN_t, uintN_t #include <cstdint> // intN_t, uintN_t
#include <cstring> // memcpy, memmove #include <cstring> // memcpy, memmove
#include <limits> // numeric_limits #include <limits> // numeric_limits
#include <type_traits> // conditional #include <type_traits> // conditional
#ifdef _MSC_VER
#include <cstdlib> // _byteswap_uint64
#endif
// SSE2 (every x86-64 CPU) and NEON (every 64-bit Arm CPU) convert the 16
// digits of a double at once
#if defined(__x86_64__) || (defined(_M_X64) && !defined(_M_ARM64EC))
#include <emmintrin.h>
#define JSON_DTOA_SSE2 1
#define JSON_DTOA_NEON 0
#elif (defined(__aarch64__) || defined(_M_ARM64)) && !defined(_M_ARM64EC) && !defined(__ARM_BIG_ENDIAN)
#include <arm_neon.h>
#define JSON_DTOA_SSE2 0
#define JSON_DTOA_NEON 1
#else
#define JSON_DTOA_SSE2 0
#define JSON_DTOA_NEON 0
#endif
#include <nlohmann/detail/conversions/zmij.hpp>
#include <nlohmann/detail/macro_scope.hpp> #include <nlohmann/detail/macro_scope.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN NLOHMANN_JSON_NAMESPACE_BEGIN
@@ -918,6 +939,88 @@ void grisu2(char* buf, int& len, int& decimal_exponent, FloatType value)
grisu2(buf, len, decimal_exponent, w.minus, w.w, w.plus); grisu2(buf, len, decimal_exponent, w.minus, w.w, w.plus);
} }
/*!
@brief the shortest digits of a positive finite float (other than double): Grisu2
*/
template<typename FloatType>
JSON_HEDLEY_NON_NULL(1)
void shortest_digits(char* buf, int& len, int& decimal_exponent, FloatType value)
{
grisu2(buf, len, decimal_exponent, value);
}
/*!
@brief the shortest digits of a positive finite double: the conversion of
Zmij (see zmij.hpp), which always finds the shortest digits that read back as
the same value (Grisu2 does not for about one double in a thousand), and the
closest of them if there are several
v = buf * 10^decimal_exponent, as for grisu2()
*/
JSON_HEDLEY_NON_NULL(1)
inline void shortest_digits(char* buf, int& len, int& decimal_exponent, double value)
{
static_assert(std::numeric_limits<double>::is_iec559 && std::numeric_limits<double>::digits == 53,
"internal error: the conversion of Zmij needs IEEE 754 binary64 doubles");
JSON_ASSERT(std::isfinite(value));
JSON_ASSERT(value > 0);
std::uint64_t bits = 0;
std::memcpy(&bits, &value, sizeof(bits));
zmij::decimal d = zmij::to_decimal(bits);
// without trailing zeros (up to 16): 8, 4, 2, 1 at a time
while (d.significand % 100000000 == 0)
{
d.significand /= 100000000;
d.exponent += 8;
}
if (d.significand % 10000 == 0)
{
d.significand /= 10000;
d.exponent += 4;
}
if (d.significand % 100 == 0)
{
d.significand /= 100;
d.exponent += 2;
}
if (d.significand % 10 == 0)
{
d.significand /= 10;
d.exponent += 1;
}
// at most 17 digits, written from the back two at a time
static constexpr const char* pairs =
"00010203040506070809101112131415161718192021222324252627282930313233343536373839"
"40414243444546474849505152535455565758596061626364656667686970717273747576777879"
"8081828384858687888990919293949596979899";
std::array<char, 20> digits{};
std::size_t n = digits.size();
while (d.significand >= 100)
{
const std::uint64_t two_digits = d.significand % 100; // a variable: GCC calls a cast of the remainder useless where std::uint64_t is std::size_t
const auto i = static_cast<std::size_t>(two_digits) * 2;
d.significand /= 100;
n -= 2;
digits[n] = pairs[i];
digits[n + 1] = pairs[i + 1];
}
if (d.significand >= 10)
{
const auto i = static_cast<std::size_t>(d.significand) * 2;
n -= 2;
digits[n] = pairs[i];
digits[n + 1] = pairs[i + 1];
}
else
{
digits[--n] = static_cast<char>('0' + d.significand);
}
len = static_cast<int>(digits.size() - n);
std::memcpy(buf, digits.data() + n, static_cast<std::size_t>(len));
decimal_exponent = d.exponent;
}
/*! /*!
@brief appends a decimal representation of e to buf @brief appends a decimal representation of e to buf
@return a pointer to the element following the exponent. @return a pointer to the element following the exponent.
@@ -1047,6 +1150,374 @@ inline char* format_buffer(char* buf, int len, int decimal_exponent,
return append_exponent(buf, n - 1); return append_exponent(buf, n - 1);
} }
/// eight decimal digits (a value below 10^8) as bytes 0..9, the first digit
/// in the most significant byte: three steps that divide all lanes at once
/// by a multiplication (the conversion of Xiang JunBo, as in Zmij)
inline std::uint64_t eight_digit_bytes(std::uint64_t abcdefgh) noexcept
{
const std::uint64_t abcd_efgh = abcdefgh + (((std::uint64_t{1} << 32u) - 10000u) * ((abcdefgh * (((std::uint64_t{1} << 40u) / 10000u) + 1u)) >> 40u));
const std::uint64_t ab_cd_ef_gh = abcd_efgh + (((std::uint64_t{1} << 16u) - 100u) * (((abcd_efgh * (((std::uint64_t{1} << 19u) / 100u) + 1u)) >> 19u) & 0x7F0000007Fu));
return ab_cd_ef_gh + (((std::uint64_t{1} << 8u) - 10u) * (((ab_cd_ef_gh * (((std::uint64_t{1} << 10u) / 10u) + 1u)) >> 10u) & 0x000F000F000F000Fu));
}
/// store the bytes of v, the most significant one first (one byte swap and
/// one store where the byte order is known: compilers do not reliably merge
/// the byte stores once this is inlined)
inline void store_msb_first(char* p, std::uint64_t v) noexcept
{
#if defined(__BYTE_ORDER__) && defined(__ORDER_LITTLE_ENDIAN__) && __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__
v = __builtin_bswap64(v);
std::memcpy(p, &v, sizeof(v));
#elif defined(__BYTE_ORDER__) && defined(__ORDER_BIG_ENDIAN__) && __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
std::memcpy(p, &v, sizeof(v));
#elif defined(_MSC_VER) // (little-endian on all its targets)
v = _byteswap_uint64(v);
std::memcpy(p, &v, sizeof(v));
#else
for (unsigned i = 0; i < 8; ++i)
{
p[i] = static_cast<char>(v >> (56u - (8u * i)));
}
#endif
}
/*!
@brief digits * 10^exp for a double, in the layout of format_buffer()
The layout is that of format_buffer() with min_exp -4 and max_exp 15 (the
digits10 of double). The digits are converted eight at a time and placed
with fixed-size moves instead of per-digit loops and moves of the buffer.
@param[in] digits the digits (not 0, at most 17 digits; trailing zeros allowed)
@param[in] exp the decimal exponent of the last digit
@return a pointer past the text; up to 41 bytes at @a first are written
(some beyond the returned end)
*/
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_decimal(char* first, std::uint64_t digits, int exp) noexcept
{
JSON_ASSERT(digits != 0 && digits < 100000000000000000u);
const std::uint64_t upper = digits / 100000000u;
const std::uint64_t b0 = upper / 100000000u; // (one digit: it is its own byte)
const std::uint64_t b1 = eight_digit_bytes(upper % 100000000u);
const std::uint64_t b2 = eight_digit_bytes(digits % 100000000u);
// leading and trailing zero digits: zero bytes, counted without division
int leading = 16;
int zeros = 16;
if (b0 != 0)
{
leading = count_leading_zeros(b0) / 8;
}
else if (b1 != 0)
{
leading = 8 + (count_leading_zeros(b1) / 8);
}
else
{
leading += count_leading_zeros(b2) / 8;
}
if (b2 != 0)
{
zeros = count_trailing_zeros(b2) / 8;
}
else if (b1 != 0)
{
zeros = 8 + (count_trailing_zeros(b1) / 8);
}
// (else: 16, b0 is the one digit that is not 0)
// the digits as text at text + leading, then '0's, so that fixed-size
// moves need not check how many digits there are
std::array<char, 64> text; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
store_msb_first(text.data(), b0 + 0x3030303030303030u);
store_msb_first(text.data() + 8, b1 + 0x3030303030303030u);
store_msb_first(text.data() + 16, b2 + 0x3030303030303030u);
std::memset(text.data() + 24, '0', 40);
const int k = 24 - leading - zeros; // significant digits
const int n = k + exp + zeros; // position of the decimal point after the first digit
const char* const s0 = text.data() + leading;
if (-4 < n && n <= 15)
{
// "0.[000]digits" (n <= 0) is the digits after 1 - n leading '0's
// with the point after the first; "digits[000].0" (n >= k) and
// "dig.its" put the point after n characters
const int pad = n <= 0 ? 1 - n : 0;
const char* const s = s0 - pad;
const int len = k + pad;
const int point = n + pad;
std::memcpy(first, s, 16);
std::memcpy(first + point + 1, s + point, 24);
first[point] = '.';
return first + (point >= len ? point + 2 : len + 1);
}
// d.igitse+XX, with at least two exponent digits (as append_exponent())
std::memcpy(first, s0, 16);
std::memcpy(first + 2, s0 + 1, 16);
first[1] = '.';
char* const end = first + (k == 1 ? 1 : k + 1);
const int e = n - 1;
const auto ea = static_cast<unsigned>(e < 0 ? -e : e);
const bool three = ea >= 100;
end[0] = 'e';
end[1] = e < 0 ? '-' : '+';
end[2] = static_cast<char>('0' + (three ? ea / 100 : (ea / 10) % 10));
end[3] = static_cast<char>('0' + (three ? (ea / 10) % 10 : ea % 10));
end[4] = static_cast<char>('0' + (ea % 10));
return end + (three ? 5 : 4);
}
/*!
@brief the shortest decimal of a positive double (Zmij), as write_decimal()
writes it
For a normal double, the shorter candidate has 15 or 16 digits: they are
converted at once (two halves of eight digits) and followed by the digit
after them, if there is one, without the multiplication and division by 10
that counting the digits of one number would take. The fixed layouts move
the digits after the point by one byte.
@return a pointer past the text; up to 41 bytes at @a first are written
(some beyond the returned end)
*/
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_shortest(char* first, const zmij::shortest_decimal d) noexcept
{
const std::uint64_t sig = d.integral;
if (JSON_HEDLEY_UNLIKELY(sig < 100000000000000u || sig >= 10000000000000000u))
{
// (subnormals)
return d.has_digit ? write_decimal(first, (sig * 10) + d.digit, d.exponent) : write_decimal(first, sig, d.exponent + 1);
}
const bool sixteen = sig >= 1000000000000000u; // (else 15 digits)
const int last = d.has_digit ? d.digit : 0;
const std::uint64_t upper = sig / 100000000u;
#if JSON_DTOA_SSE2
// NOLINTBEGIN(portability-simd-intrinsics)
// the two halves in the 64-bit lanes, each as abcd * 2^32 + efgh, then as
// bytes (as eight_digit_bytes(), one lane each)
const __m128i x = _mm_set_epi64x(static_cast<long long>(sig - (upper * 100000000u)), static_cast<long long>(upper));
const __m128i abcd = _mm_srli_epi64(_mm_mul_epu32(x, _mm_set1_epi64x(109951163)), 40); // 2^40 / 10000 + 1
const __m128i abcd_efgh = _mm_add_epi64(x, _mm_mul_epu32(abcd, _mm_set1_epi64x(4294957296))); // 2^32 - 10000
// 32-bit lanes in the order of the text: abcd, efgh of both halves
const __m128i fours = _mm_shuffle_epi32(abcd_efgh, _MM_SHUFFLE(2, 3, 0, 1));
const __m128i ab = _mm_srli_epi16(_mm_mulhi_epu16(fours, _mm_set1_epi32(5243)), 3);
const __m128i ab_cd = _mm_or_si128(_mm_slli_epi32(_mm_sub_epi16(fours, _mm_mullo_epi16(ab, _mm_set1_epi32(100))), 16), ab);
// 16-bit lanes ab (< 100) -> bytes a, b: 256 * ab - 2559 * (ab / 10)
const __m128i bytes = _mm_sub_epi16(_mm_slli_epi16(ab_cd, 8), _mm_mullo_epi16(_mm_set1_epi16(2559), _mm_mulhi_epu16(ab_cd, _mm_set1_epi16(6554))));
// the last digit that is not 0 (sig is not 0)
const auto nonzero = static_cast<std::uint64_t>(_mm_movemask_epi8(_mm_cmpgt_epi8(bytes, _mm_setzero_si128())));
const int digits = 63 - count_leading_zeros(nonzero) + (sixteen ? 1 : 0); // without trailing zeros
const __m128i chars = _mm_add_epi8(bytes, _mm_set1_epi8('0'));
// the 16 characters from the first digit
const __m128i s = sixteen ? chars : _mm_or_si128(_mm_srli_si128(chars, 1), _mm_slli_si128(_mm_cvtsi32_si128('0' + last), 15));
const char s16 = static_cast<char>(sixteen ? '0' + last : '0'); // the 17th
const auto store_16 = [&s](char* p) noexcept
{
std::memcpy(p, &s, 16);
};
const char first_digit = static_cast<char>(_mm_cvtsi128_si32(s));
// NOLINTEND(portability-simd-intrinsics)
#elif JSON_DTOA_NEON
// as with SSE2: the halves in 32-bit lanes, then abcd, efgh of both
const uint32x2_t halves = vcreate_u32(upper | ((sig - (upper * 100000000u)) << 32u));
const uint32x2_t abcd = vmovn_u64(vshrq_n_u64(vmull_n_u32(halves, static_cast<std::uint32_t>(((std::uint64_t{1} << 40u) / 10000u) + 1u)), 40));
const uint32x2_t efgh = vmls_n_u32(halves, abcd, 10000u);
const uint32x4_t fours = vcombine_u32(vzip1_u32(abcd, efgh), vzip2_u32(abcd, efgh));
const uint32x4_t ab = vshrq_n_u32(vmulq_n_u32(fours, 5243u), 19);
const uint16x8_t ab_cd = vreinterpretq_u16_u32(vorrq_u32(ab, vshlq_n_u32(vmlsq_n_u32(fours, ab, 100u), 16)));
const uint16x8_t tens = vshrq_n_u16(vmulq_n_u16(ab_cd, 103u), 10);
const uint8x16_t bytes = vreinterpretq_u8_u16(vorrq_u16(tens, vshlq_n_u16(vmlsq_n_u16(ab_cd, tens, 10u), 8)));
// the last digit that is not 0 (sig is not 0): a nibble per byte
const std::uint64_t nonzero = vget_lane_u64(vreinterpret_u64_u8(vshrn_n_u16(vreinterpretq_u16_u8(vtstq_u8(bytes, bytes)), 4)), 0);
const int digits = ((63 - count_leading_zeros(nonzero)) / 4) + (sixteen ? 1 : 0); // without trailing zeros
const uint8x16_t chars = vaddq_u8(bytes, vdupq_n_u8('0'));
// the 16 characters from the first digit
const uint8x16_t s = sixteen ? chars : vextq_u8(chars, vdupq_n_u8(static_cast<std::uint8_t>('0' + last)), 1);
const char s16 = static_cast<char>(sixteen ? '0' + last : '0'); // the 17th
const auto store_16 = [&s](char* p) noexcept
{
vst1q_u8(reinterpret_cast<std::uint8_t*>(p), s); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
};
const auto first_digit = static_cast<char>(vgetq_lane_u8(s, 0));
#else
const std::uint64_t hi = eight_digit_bytes(upper);
const std::uint64_t lo = eight_digit_bytes(sig - (upper * 100000000u));
// trailing zero digits: zero bytes (sig is not 0)
const int zeros = lo != 0 ? count_trailing_zeros(lo) / 8 : 8 + (count_trailing_zeros(hi) / 8);
const int digits = 15 - zeros + (sixteen ? 1 : 0); // without trailing zeros
// the 16 characters from the first digit
const std::uint64_t s_hi = (sixteen ? hi : (hi << 8u) | (lo >> 56u)) + 0x3030303030303030u;
const std::uint64_t s_lo = (sixteen ? lo : (lo << 8u) | static_cast<std::uint64_t>(last)) + 0x3030303030303030u;
const char s16 = static_cast<char>(sixteen ? '0' + last : '0'); // the 17th
const auto store_16 = [s_hi, s_lo](char* p) noexcept
{
store_msb_first(p, s_hi);
store_msb_first(p + 8, s_lo);
};
const auto first_digit = static_cast<char>(s_hi >> 56u);
#endif
const int len = d.has_digit ? 16 + (sixteen ? 1 : 0) : digits; // significant digits
const int n = 16 + (sixteen ? 1 : 0) + d.exponent; // digits before the point
if (JSON_HEDLEY_LIKELY(n >= 1 && n <= 15))
{
// "dig.its" and "digits[000].0": the digits after the point move by
// one byte ('0's follow the digits)
#if JSON_DTOA_SSE2
// NOLINTBEGIN(portability-simd-intrinsics)
// (in the register: reading the digits back from memory right after
// storing them waits until the stores are done)
const __m128i at = _mm_set1_epi8(static_cast<char>(n));
const __m128i index = _mm_setr_epi8(0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15);
const __m128i before = _mm_cmpgt_epi8(at, index);
const __m128i after = _mm_cmpgt_epi8(index, at);
const __m128i text = _mm_or_si128(_mm_or_si128(_mm_and_si128(s, before), _mm_and_si128(_mm_slli_si128(s, 1), after)),
_mm_andnot_si128(_mm_or_si128(before, after), _mm_set1_epi8('.')));
std::memcpy(first, &text, 16);
first[16] = static_cast<char>(_mm_extract_epi16(s, 7) >> 8);
first[17] = s16;
// NOLINTEND(portability-simd-intrinsics)
#elif JSON_DTOA_NEON
const uint8x16_t index = vcombine_u8(vcreate_u8(0x0706050403020100u), vcreate_u8(0x0F0E0D0C0B0A0908u));
const uint8x16_t at = vdupq_n_u8(static_cast<std::uint8_t>(n));
const uint8x16_t after_point = vbslq_u8(vcgtq_u8(index, at), vextq_u8(vdupq_n_u8(0), s, 15), vdupq_n_u8('.'));
vst1q_u8(reinterpret_cast<std::uint8_t*>(first), vbslq_u8(vcltq_u8(index, at), s, after_point)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
first[16] = static_cast<char>(vgetq_lane_u8(s, 15));
first[17] = s16;
#else
store_16(first);
first[16] = s16;
std::uint64_t after_point[2]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays,cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
std::memcpy(after_point, first + n, 16);
std::memcpy(first + n + 1, after_point, 16);
first[n] = '.';
#endif
return first + (n >= len ? n + 2 : len + 1);
}
if (n <= 0 && n > -4)
{
// "0.[000]digits"
std::memset(first, '0', 8);
first[1] = '.';
store_16(first + 2 - n);
first[18 - n] = s16;
return first + 2 - n + len;
}
// d.igitse+XX, with at least two exponent digits (as append_exponent())
store_16(first + 1);
first[17] = s16;
first[0] = first_digit;
first[1] = '.';
char* const end = first + (len == 1 ? 1 : len + 1);
const int e = n - 1;
const auto ea = static_cast<unsigned>(e < 0 ? -e : e);
const bool three = ea >= 100;
end[0] = 'e';
end[1] = e < 0 ? '-' : '+';
end[2] = static_cast<char>('0' + (three ? ea / 100 : (ea / 10) % 10));
end[3] = static_cast<char>('0' + (three ? (ea / 10) % 10 : ea % 10));
end[4] = static_cast<char>('0' + (ea % 10));
return end + (three ? 5 : 4);
}
/// the powers of ten up to 10^16
inline const std::array<std::uint64_t, 17>& powers_of_ten_16() noexcept
{
static const std::array<std::uint64_t, 17> powers =
{
{
1u, 10u, 100u, 1000u, 10000u, 100000u, 1000000u, 10000000u, 100000000u, 1000000000u, 10000000000u,
100000000000u, 1000000000000u, 10000000000000u, 100000000000000u, 1000000000000000u, 10000000000000000u
}
};
return powers;
}
/*!
@brief digits * 10^exp, as write_decimal() writes it, for the digits of a
double that need no conversion (count digits, at most 15, the first not 0;
trailing zeros allowed): extended to 16 digits and written by write_shortest()
@return a pointer past the text; up to 41 bytes at @a first are written
(some beyond the returned end)
*/
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_short_decimal(char* first, std::uint64_t digits, int count, int exp) noexcept
{
JSON_ASSERT(digits >= powers_of_ten_16()[static_cast<std::size_t>(count - 1)] && count <= 15);
const int scale = 16 - count;
return write_shortest(first, zmij::shortest_decimal{digits * powers_of_ten_16()[static_cast<std::size_t>(scale)], exp - scale - 1, 0, false});
}
/// as write_short_decimal(), counting the digits (not 0, less than 10^15)
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_short_decimal(char* first, std::uint64_t digits, int exp) noexcept
{
JSON_ASSERT(digits != 0 && digits < 1000000000000000u);
// floor(log10(2^bits)) + 1 digits, or one less
const int log2_bound = ((64 - count_leading_zeros(digits)) * 1233) >> 12;
const int count = log2_bound + (digits >= powers_of_ten_16()[static_cast<std::size_t>(log2_bound)] ? 1 : 0);
return write_short_decimal(first, digits, count, exp);
}
/// a positive finite float (other than double): Grisu2 and format_buffer()
template<typename FloatType>
JSON_HEDLEY_NON_NULL(1, 2)
JSON_HEDLEY_RETURNS_NON_NULL
char* write_positive(char* first, const char* last, FloatType value)
{
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10);
// Compute v = buffer * 10^decimal_exponent.
// The decimal digits are stored in the buffer, which needs to be interpreted
// as an unsigned decimal integer.
// len is the length of the buffer, i.e., the number of decimal digits.
int len = 0;
int decimal_exponent = 0;
shortest_digits(first, len, decimal_exponent, value);
JSON_ASSERT(len <= std::numeric_limits<FloatType>::max_digits10);
// Format the buffer like printf("%.*g", prec, value)
constexpr int kMinExp = -4;
// Use digits10 here to increase compatibility with version 2.
constexpr int kMaxExp = std::numeric_limits<FloatType>::digits10;
JSON_ASSERT(last - first >= kMaxExp + 2);
JSON_ASSERT(last - first >= 2 + (-kMinExp - 1) + std::numeric_limits<FloatType>::max_digits10);
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10 + 6);
return format_buffer(first, len, decimal_exponent, kMinExp, kMaxExp);
}
/// a positive finite double: the shortest digits (Zmij), laid out by
/// write_shortest() (through a local buffer if [first, last) is shorter than
/// the 41 bytes it may write)
JSON_HEDLEY_NON_NULL(1, 2)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_positive(char* first, const char* last, double value)
{
static_assert(std::numeric_limits<double>::is_iec559 && std::numeric_limits<double>::digits == 53,
"internal error: the conversion of Zmij needs IEEE 754 binary64 doubles");
std::uint64_t bits = 0;
std::memcpy(&bits, &value, sizeof(bits));
const zmij::shortest_decimal d = zmij::to_shortest(bits);
if (JSON_HEDLEY_LIKELY(last - first >= 41))
{
return write_shortest(first, d);
}
std::array<char, 64> buf; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
const auto len = static_cast<std::size_t>(write_shortest(buf.data(), d) - buf.data());
JSON_ASSERT(static_cast<std::size_t>(last - first) >= len);
std::memcpy(first, buf.data(), len);
return first + len;
}
} // namespace dtoa_impl } // namespace dtoa_impl
/*! /*!
@@ -1064,7 +1535,6 @@ JSON_HEDLEY_NON_NULL(1, 2)
JSON_HEDLEY_RETURNS_NON_NULL JSON_HEDLEY_RETURNS_NON_NULL
char* to_chars(char* first, const char* last, FloatType value) char* to_chars(char* first, const char* last, FloatType value)
{ {
static_cast<void>(last); // maybe unused - fix warning
JSON_ASSERT(std::isfinite(value)); JSON_ASSERT(std::isfinite(value));
// Use signbit(value) instead of (value < 0) since signbit works for -0. // Use signbit(value) instead of (value < 0) since signbit works for -0.
@@ -1090,28 +1560,7 @@ char* to_chars(char* first, const char* last, FloatType value)
JSON_HEDLEY_DIAGNOSTIC_POP JSON_HEDLEY_DIAGNOSTIC_POP
#endif #endif
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10); return dtoa_impl::write_positive(first, last, value);
// Compute v = buffer * 10^decimal_exponent.
// The decimal digits are stored in the buffer, which needs to be interpreted
// as an unsigned decimal integer.
// len is the length of the buffer, i.e., the number of decimal digits.
int len = 0;
int decimal_exponent = 0;
dtoa_impl::grisu2(first, len, decimal_exponent, value);
JSON_ASSERT(len <= std::numeric_limits<FloatType>::max_digits10);
// Format the buffer like printf("%.*g", prec, value)
constexpr int kMinExp = -4;
// Use digits10 here to increase compatibility with version 2.
constexpr int kMaxExp = std::numeric_limits<FloatType>::digits10;
JSON_ASSERT(last - first >= kMaxExp + 2);
JSON_ASSERT(last - first >= 2 + (-kMinExp - 1) + std::numeric_limits<FloatType>::max_digits10);
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10 + 6);
return dtoa_impl::format_buffer(first, len, decimal_exponent, kMinExp, kMaxExp);
} }
} // namespace detail } // namespace detail
@@ -0,0 +1,238 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2025 Victor Zverovich <https://github.com/vitaut/zmij>
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint32_t, uint64_t
#include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/bit_ops.hpp>
#include <nlohmann/detail/input/pow5_table.hpp>
#include <nlohmann/detail/macro_scope.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief the shortest decimal representation of a double
A C++11 port of the conversion of Zmij by Victor Zverovich
(https://github.com/vitaut/zmij, MIT license): the shortest decimal in the
rounding interval of a double, the closest one if there are several. Zmij
credits Xiang JunBo (producing the shorter candidate without a division) and
Dougall Johnson (the compressed powers of ten). The powers of ten are taken
from the table for number parsing (pow5_table.hpp) where it holds them, and
computed from the compressed tables of Zmij beyond it.
*/
namespace zmij
{
/// significand * 10^exponent
struct decimal
{
std::uint64_t significand;
int exponent;
};
/// the compressed powers of ten of Zmij
inline const std::array<std::uint64_t, 28>& pow10_minor() noexcept
{
static const std::array<std::uint64_t, 28> table =
{
{
0x8000000000000000u, 0xa000000000000000u, 0xc800000000000000u, 0xfa00000000000000u, 0x9c40000000000000u,
0xc350000000000000u, 0xf424000000000000u, 0x9896800000000000u, 0xbebc200000000000u, 0xee6b280000000000u,
0x9502f90000000000u, 0xba43b74000000000u, 0xe8d4a51000000000u, 0x9184e72a00000000u, 0xb5e620f480000000u,
0xe35fa931a0000000u, 0x8e1bc9bf04000000u, 0xb1a2bc2ec5000000u, 0xde0b6b3a76400000u, 0x8ac7230489e80000u,
0xad78ebc5ac620000u, 0xd8d726b7177a8000u, 0x878678326eac9000u, 0xa968163f0a57b400u, 0xd3c21bcecceda100u,
0x84595161401484a0u, 0xa56fa5b99019a5c8u, 0xcecb8f27f4200f3au
}
};
return table;
}
/// (high, low) pairs
inline const std::array<std::uint64_t, 50>& pow10_major() noexcept
{
static const std::array<std::uint64_t, 50> table =
{
{
0xaddcb9e83c6b1793u, 0xdf4abe242a1bbf3eu, 0xaf8e5410288e1b6fu, 0x07ecf0ae5ee44ddau, 0xb1442798f49ffb4au, 0x99cd11cfdf41779du,
0xb2fe3f0b8599ef07u, 0x861fa7e6dcb4aa15u, 0xb4bca50b065abe63u, 0x0fed077a756b53aau, 0xb67f6455292cbf08u, 0x1a3bc84c17b1d543u,
0xb84687c269ef3bfbu, 0x3d5d514f40eea742u, 0xba121a4650e4ddebu, 0x92f34d62616ce413u, 0xbbe226efb628afeau, 0x890489f70a55368cu,
0xbdb6b8e905cb600fu, 0x5400e987bbc1c921u, 0xbf8fdb78849a5f96u, 0xde98520472bdd034u, 0xc16d9a0095928a27u, 0x75b7053c0f178294u,
0xc350000000000000u, 0x0000000000000000u, 0xc5371912364ce305u, 0x6c28000000000000u, 0xc722f0ef9d80aad6u, 0x424d3ad2b7b97ef6u,
0xc913936dd571c84cu, 0x03bc3a19cd1e38eau, 0xcb090c8001ab551cu, 0x5cadf5bfd3072cc6u, 0xcd036837130890a1u, 0x36dba887c37a8c10u,
0xcf02b2c21207ef2eu, 0x94f967e45e03f4bcu, 0xd106f86e69d785c7u, 0xe13336d701beba52u, 0xd31045a8341ca07cu, 0x1ede48111209a051u,
0xd51ea6fa85785631u, 0x552a74227f3ea566u, 0xd732290fbacaf133u, 0xa97c177947ad4096u, 0xd94ad8b1c7380874u, 0x18375281ae7822bdu,
0xdb68c2ca82ed2a05u, 0xa67398db9f6820e1u
}
};
return table;
}
/// one bit per power: whether the computed value is one unit too large
inline const std::array<std::uint32_t, 21>& pow10_fixups() noexcept
{
static const std::array<std::uint32_t, 21> table =
{
{
0x8d8fc810u, 0x06100293u, 0x19000000u, 0x00100000u, 0x00000908u, 0x00000000u, 0x04e00300u, 0x3807e0b2u, 0x3d83d793u, 0x0006f5ccu,
0x00000000u, 0xffff0000u, 0x8076337du, 0x4ff45ba0u, 0x09405033u, 0x034376d9u, 0x09000000u, 0x4e100501u, 0x076d14dcu, 0xf964f45eu,
0x0000003du
}
};
return table;
}
/// the 128-bit significand of 10^k, rounded down, for k in [-307, 341]
/// (compute_pow10 of Zmij)
inline uint128_parts compute_pow10(int k) noexcept
{
const auto i = static_cast<unsigned>(k + 307);
const std::uint64_t m = pow10_minor()[(i + 24) % 28];
const std::size_t j = 2 * static_cast<std::size_t>((i + 24) / 28);
const std::uint64_t h_hi = pow10_major()[j];
const std::uint64_t h_lo = pow10_major()[j + 1];
const std::uint64_t h1 = full_multiplication(h_lo, m).high;
const std::uint64_t c0 = h_lo * m;
const std::uint64_t c1 = h1 + (h_hi * m);
const std::uint64_t c2 = (c1 < h1 ? 1u : 0u) + full_multiplication(h_hi, m).high;
uint128_parts r{};
if ((c2 >> 63u) != 0)
{
r.high = c2;
r.low = c1;
}
else
{
r.high = (c2 << 1u) | (c1 >> 63u);
r.low = (c1 << 1u) | (c0 >> 63u);
}
r.low -= (pow10_fixups()[i >> 5u] >> (i & 31u)) & 1u;
return r;
}
/// The 128-bit significand of 10^k, rounded down, for k in [-342, 341].
/// Up to 10^308, the table for number parsing holds the same significands
/// (those of 5^k), except for k in [-27, -1], where it holds them one unit
/// larger (as the Eisel-Lemire algorithm needs them).
inline uint128_parts pow10(int k) noexcept
{
if (k > pow5_128_largest_power)
{
return compute_pow10(k); // (only for the smallest doubles)
}
const auto i = 2 * static_cast<std::size_t>(k - pow5_128_smallest_power);
uint128_parts r{pow5_128()[i + 1], pow5_128()[i]};
const std::uint64_t adjust = static_cast<unsigned>(k + 27) < 27u ? 1u : 0u;
r.high -= r.low < adjust ? 1u : 0u;
r.low -= adjust;
return r;
}
/// (x_hi * 2^64 + x_lo) * y >> 64, as 128 bits
inline uint128_parts umul192_hi128(std::uint64_t x_hi, std::uint64_t x_lo, std::uint64_t y) noexcept
{
const uint128_parts p = full_multiplication(x_hi, y);
uint128_parts r{};
r.low = p.low + full_multiplication(x_lo, y).high;
r.high = p.high + (r.low < p.low ? 1u : 0u);
return r;
}
/// (x * y + c) >> 64
inline std::uint64_t umul128_add_hi64(std::uint64_t x, std::uint64_t y, std::uint64_t c) noexcept
{
const uint128_parts p = full_multiplication(x, y);
return p.high + (p.low + c < p.low ? 1u : 0u);
}
/// the result of Zmij: the shorter candidate and, if that is outside the
/// rounding interval, the digit after it (16 bytes: returned in registers)
struct shortest_decimal
{
std::uint64_t integral; ///< the shorter candidate (15 or 16 digits for normal doubles)
int exponent; ///< the decimal exponent of the digit after it
unsigned char digit; ///< the digit after it (if has_digit)
bool has_digit; ///< whether the shortest decimal is integral * 10 + digit
};
/// The shortest decimal in the rounding interval of a positive finite double
/// given by its bits, the closest one if there are several (to_decimal of
/// Zmij, which keeps the last digit apart: the 15 or 16 digits before it can be
/// converted without a multiplication by 10 first). Always inlined: GCC
/// otherwise calls it, and its result goes through memory.
JSON_HEDLEY_ALWAYS_INLINE shortest_decimal to_shortest(std::uint64_t bits) noexcept
{
constexpr int extra_shift = 9;
const auto raw_exp = static_cast<int>((bits >> 52u) & 0x7FFu);
std::uint64_t bin_sig = bits & ((std::uint64_t{1} << 52u) - 1);
// a power of two has a narrower interval below (except the smallest normal)
const bool regular = bin_sig != 0 || raw_exp <= 1;
const int bin_exp = (raw_exp == 0 ? 1 : raw_exp) - 1075;
if (raw_exp != 0)
{
bin_sig |= std::uint64_t{1} << 52u;
}
// floor(log10(2^bin_exp)), or floor(log10(3/4 * 2^bin_exp)) for the irregular case
const int dec_exp = ((bin_exp * 315653) - (regular ? 0 : 131072)) >> 20;
// scaled by 10^(-dec_exp - 1): the integral part is the shorter candidate
const int shift = bin_exp + ((-(dec_exp + 1) * 217707) >> 16) + 1 + extra_shift;
const uint128_parts p10 = pow10(-dec_exp - 1);
const uint128_parts p = umul192_hi128(p10.high, p10.low, bin_sig << static_cast<unsigned>(shift));
std::uint64_t integral = p.high >> static_cast<unsigned>(extra_shift);
const std::uint64_t fractional = (p.high << static_cast<unsigned>(64 - extra_shift)) | (p.low >> static_cast<unsigned>(extra_shift));
std::uint64_t digit = 0;
bool round_up = false;
bool round_down = false;
if (JSON_HEDLEY_LIKELY(regular))
{
const std::uint64_t half_ulp = (p10.high >> static_cast<unsigned>(extra_shift + 1 - shift)) + (1 - (bin_sig & 1u));
round_up = fractional + half_ulp < fractional;
round_down = half_ulp > fractional;
// the last digit of the longer candidate, rounded to nearest
digit = umul128_add_hi64(fractional, 10, (std::uint64_t{1} << 63u) + 6);
if (fractional == (std::uint64_t{1} << 62u))
{
digit = 2; // 2.5 rounds to 2
}
}
else
{
const std::uint64_t half_ulp = p10.high >> static_cast<unsigned>(extra_shift + 1 - shift);
round_up = half_ulp > ~std::uint64_t{0} - fractional;
round_down = (half_ulp >> 1u) > fractional;
digit = umul128_add_hi64(fractional, 10, (std::uint64_t{1} << 63u) - 1);
const std::uint64_t lowest = umul128_add_hi64(fractional - (half_ulp >> 1u), 10, ~std::uint64_t{0});
digit = digit < lowest ? lowest : digit;
}
integral += round_up ? 1u : 0u;
// if the shorter candidate is outside the rounding interval: one digit more
return shortest_decimal{integral, dec_exp, static_cast<unsigned char>(digit), !round_up && !round_down};
}
/// The shortest decimal in the rounding interval of a positive finite double
/// given by its bits, as one number. The significand can end in zeros.
inline decimal to_decimal(std::uint64_t bits) noexcept
{
const shortest_decimal d = to_shortest(bits);
if (d.has_digit)
{
return decimal{(d.integral * 10) + d.digit, d.exponent};
}
return decimal{d.integral, d.exponent + 1};
}
} // namespace zmij
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -249,8 +249,9 @@ template<typename FloatType>
using native_float_t = typename std::conditional<std::numeric_limits<FloatType>::digits == 24, float, double>::type; using native_float_t = typename std::conditional<std::numeric_limits<FloatType>::digits == 24, float, double>::type;
/// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three /// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three
/// multiplications instead of eight (after simdjson and fast_float) /// multiplications instead of eight (after simdjson and fast_float); always
inline std::uint32_t parse_eight_digits(std::uint64_t v) noexcept /// inlined, as GCC otherwise calls it in the number loops
JSON_HEDLEY_ALWAYS_INLINE std::uint32_t parse_eight_digits(std::uint64_t v) noexcept
{ {
v = ((v & 0x0F0F0F0F0F0F0F0Fu) * 2561u) >> 8u; v = ((v & 0x0F0F0F0F0F0F0F0Fu) * 2561u) >> 8u;
v = ((v & 0x00FF00FF00FF00FFu) * 6553601u) >> 16u; v = ((v & 0x00FF00FF00FF00FFu) * 6553601u) >> 16u;
@@ -21,6 +21,8 @@
#undef JSON_NO_UNIQUE_ADDRESS #undef JSON_NO_UNIQUE_ADDRESS
#undef JSON_DISABLE_ENUM_SERIALIZATION #undef JSON_DISABLE_ENUM_SERIALIZATION
#undef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION #undef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
#undef JSON_DTOA_SSE2
#undef JSON_DTOA_NEON
#ifndef JSON_TEST_KEEP_MACROS #ifndef JSON_TEST_KEEP_MACROS
#undef JSON_CATCH #undef JSON_CATCH
+45 -16
View File
@@ -1366,8 +1366,9 @@ class serializer
/*! /*!
@brief dump an integer @brief dump an integer
Dump a given integer, appending it to @ref write_buffer. Works internally with Dump a given integer, appending it to @ref write_buffer (directly: copying
@a number_buffer. the digits from another buffer right after writing them waits until the
stores are done).
@param[in] x integer number (signed or unsigned) to dump @param[in] x integer number (signed or unsigned) to dump
@tparam NumberType either @a number_integer_t or @a number_unsigned_t @tparam NumberType either @a number_integer_t or @a number_unsigned_t
@@ -1402,33 +1403,57 @@ class serializer
return; return;
} }
// use a pointer to fill the buffer // use a pointer to fill the buffer (room for as much as number_buffer holds)
auto buffer_ptr = number_buffer.begin(); // NOLINT(llvm-qualified-auto,readability-qualified-auto) if (JSON_HEDLEY_UNLIKELY(write_buffer_pos + number_buffer.size() > write_buffer.size()))
{
flush();
}
auto* buffer_ptr = write_buffer.data() + write_buffer_pos;
number_unsigned_t abs_value; number_unsigned_t abs_value;
unsigned int n_chars{}; // one byte for the minus sign
unsigned int n_chars = 0;
if (is_negative_number(x)) if (is_negative_number(x))
{ {
*buffer_ptr = '-'; *buffer_ptr = '-';
abs_value = remove_sign(static_cast<number_integer_t>(x)); abs_value = remove_sign(static_cast<number_integer_t>(x));
n_chars = 1;
// account one more byte for the minus sign
n_chars = 1 + count_digits(abs_value);
} }
else else
{ {
abs_value = static_cast<number_unsigned_t>(x); abs_value = static_cast<number_unsigned_t>(x);
n_chars = count_digits(abs_value);
} }
// up to 16 digits: eight at a time (as the digits of floats), written
// without leading zeros
if (abs_value < 10000000000000000u)
{
const std::uint64_t value = abs_value;
const std::uint64_t upper = value / 100000000u;
const std::uint64_t first = dtoa_impl::eight_digit_bytes(upper != 0 ? upper : value);
const auto leading = static_cast<unsigned>(count_leading_zeros(first) / 8); // (first is not 0)
char* const p = buffer_ptr + n_chars;
dtoa_impl::store_msb_first(p, (first << (8 * leading)) + 0x3030303030303030u);
n_chars += 8 - leading;
if (upper != 0)
{
dtoa_impl::store_msb_first(p + 8 - leading, dtoa_impl::eight_digit_bytes(value - (upper * 100000000u)) + 0x3030303030303030u);
n_chars += 8;
}
write_buffer_pos += n_chars;
return;
}
n_chars += count_digits(abs_value);
// spare 1 byte for '\0' // spare 1 byte for '\0'
JSON_ASSERT(n_chars < number_buffer.size() - 1); JSON_ASSERT(n_chars < number_buffer.size() - 1);
// jump to the end to generate the string from backward, // jump to the end to generate the string from backward,
// so we later avoid reversing the result // so we later avoid reversing the result
buffer_ptr += static_cast<typename decltype(number_buffer)::difference_type>(n_chars); buffer_ptr += n_chars;
// Fast int2ascii implementation inspired by "Fastware" talk by Andrei Alexandrescu // Fast int2ascii implementation inspired by "Fastware" talk by Andrei Alexandrescu
// See: https://www.youtube.com/watch?v=o4-CwDo2zpg // See: https://www.youtube.com/watch?v=o4-CwDo2zpg
@@ -1451,14 +1476,13 @@ class serializer
*(--buffer_ptr) = static_cast<char>('0' + abs_value); *(--buffer_ptr) = static_cast<char>('0' + abs_value);
} }
put_buffer(number_buffer, n_chars); write_buffer_pos += n_chars;
} }
/*! /*!
@brief dump a floating-point number @brief dump a floating-point number
Dump a given floating-point number, appending it to @ref write_buffer. Works internally Dump a given floating-point number, appending it to @ref write_buffer.
with @a number_buffer.
@param[in] x floating-point number to dump @param[in] x floating-point number to dump
*/ */
@@ -1485,10 +1509,15 @@ class serializer
void dump_float(number_float_t x, std::true_type /*is_ieee_single_or_double*/) void dump_float(number_float_t x, std::true_type /*is_ieee_single_or_double*/)
{ {
auto* begin = number_buffer.data(); // directly into the write buffer: copying the text from number_buffer
// right after to_chars() wrote it waits until its stores are done
if (JSON_HEDLEY_UNLIKELY(write_buffer_pos + number_buffer.size() > write_buffer.size()))
{
flush();
}
auto* begin = write_buffer.data() + write_buffer_pos;
auto* end = ::nlohmann::detail::to_chars(begin, begin + number_buffer.size(), x); auto* end = ::nlohmann::detail::to_chars(begin, begin + number_buffer.size(), x);
write_buffer_pos += static_cast<std::size_t>(end - begin);
put_buffer(number_buffer, static_cast<std::size_t>(end - begin));
} }
JSON_HEDLEY_NON_NULL(1) JSON_HEDLEY_NON_NULL(1)
+18 -11
View File
@@ -828,14 +828,19 @@ indent_done:
const auto idx = static_cast<std::uint32_t>(emit(k, 0, 0, static_cast<std::size_t>(p - b), 0) - base); const auto idx = static_cast<std::uint32_t>(emit(k, 0, 0, static_cast<std::size_t>(p - b), 0) - base);
if (depth != 0) if (depth != 0)
{ {
const frame f = {cur_idx, cur_count, cur_is_object};
if (NLOHMANN_VIEW_LIKELY(depth <= 64)) if (NLOHMANN_VIEW_LIKELY(depth <= 64))
{ {
cold.shallow[depth - 1] = f; // field by field: a frame put together on the stack and
// copied would be read back wider than it was written,
// and that load waits until the stores are done
frame& f = cold.shallow[depth - 1];
f.idx = cur_idx;
f.count = cur_count;
f.is_object = cur_is_object;
} }
else else
{ {
cold.deep.push_back(f); cold.deep.push_back(frame{cur_idx, cur_count, cur_is_object});
} }
} }
++depth; ++depth;
@@ -851,20 +856,22 @@ indent_done:
n.next = static_cast<std::uint32_t>(out - base) - cur_idx; n.next = static_cast<std::uint32_t>(out - base) - cur_idx;
if (--depth != 0) if (--depth != 0)
{ {
frame f{};
if (NLOHMANN_VIEW_LIKELY(depth <= 64)) if (NLOHMANN_VIEW_LIKELY(depth <= 64))
{ {
f = cold.shallow[depth - 1]; const frame& f = cold.shallow[depth - 1];
}
else
{
f = cold.deep.back();
cold.deep.pop_back();
}
cur_idx = f.idx; cur_idx = f.idx;
cur_count = f.count; cur_count = f.count;
cur_is_object = f.is_object; cur_is_object = f.is_object;
} }
else
{
const frame f = cold.deep.back();
cold.deep.pop_back();
cur_idx = f.idx;
cur_count = f.count;
cur_is_object = f.is_object;
}
}
} }
NLOHMANN_VIEW_ALWAYS_INLINE bool literal(const char* text, std::size_t n, value_t k, std::uint8_t flags) NLOHMANN_VIEW_ALWAYS_INLINE bool literal(const char* text, std::size_t n, value_t k, std::uint8_t flags)
@@ -10,7 +10,7 @@
#include <array> // array #include <array> // array
#include <cstddef> // size_t #include <cstddef> // size_t
#include <cstdint> // uint32_t #include <cstdint> // uint8_t, uint32_t
#include <cstring> // memcpy #include <cstring> // memcpy
#include <functional> // less #include <functional> // less
#include <map> // map #include <map> // map
@@ -41,7 +41,9 @@ struct document_data
node* inline_tape = nullptr; ///< node array allocated together with this header node* inline_tape = nullptr; ///< node array allocated together with this header
std::size_t inline_cap = 0; std::size_t inline_cap = 0;
std::string arena{}; ///< decoded strings that contained escapes // NOLINT(readability-redundant-member-init) std::string arena{}; ///< decoded strings that contained escapes // NOLINT(readability-redundant-member-init)
std::size_t arena_size = 0; ///< bytes of decoded strings at base[1] (the arena, or those of a loaded image)
std::string owned{}; ///< owned copy of the input, if any // NOLINT(readability-redundant-member-init) std::string owned{}; ///< owned copy of the input, if any // NOLINT(readability-redundant-member-init)
std::vector<std::uint8_t> owned_image{}; ///< a loaded image the document owns (the text and the decoded strings point into it) // NOLINT(readability-redundant-member-init)
// hash indexes of large objects (see object_index.hpp) // hash indexes of large objects (see object_index.hpp)
static constexpr std::uint32_t index_min_members = 128; static constexpr std::uint32_t index_min_members = 128;
+56
View File
@@ -227,6 +227,62 @@ class editor
return View(&m_doc, slot); return View(&m_doc, slot);
} }
/// insert into an array before position idx (idx <= size()); returns a
/// view of the new element
template<typename V>
View insert(const View& array, std::size_t idx, V&& value)
{
node* const a = own(array);
if (a->kind != static_cast<std::uint8_t>(value_t::array))
{
throw_type_error(309, "cannot use insert() with ", array.type_name());
}
check_index(idx, a->len + 1);
const encoded e = encode(std::forward<V>(value));
node* const slot = new_slot(e);
node* const h = block_of(m_doc, a, 1);
std::memmove(h + 2 + idx, h + 1 + idx, (h->next - 1 - idx) * sizeof(node));
make_link(h[1 + idx], slot);
++h->next;
++h->len;
++a->len;
return View(&m_doc, slot);
}
/// remove all members with this key; returns their number
std::size_t erase(const View& object, string_view_t key)
{
node* const o = own(object);
if (o->kind != static_cast<std::uint8_t>(value_t::object))
{
throw_type_error(307, "cannot use erase() with ", object.type_name());
}
for (const node* k = nav::first(m_doc, o), *end = nav::end(m_doc, o); k != end; k = document_data::after(k + 1))
{
if (key_equals(*k, key))
{
return erase_members(o, key, false);
}
}
return 0;
}
/// remove an array element
void erase(const View& array, std::size_t idx)
{
node* const a = own(array);
if (a->kind != static_cast<std::uint8_t>(value_t::array))
{
throw_type_error(307, "cannot use erase() with ", array.type_name());
}
check_index(idx, a->len);
node* const h = block_of(m_doc, a, 0);
std::memmove(h + 1 + idx, h + 2 + idx, (h->next - 2 - idx) * sizeof(node));
--h->next;
--h->len;
--a->len;
}
private: private:
/// an encoded value: a scalar node, or the root of a new array/object /// an encoded value: a scalar node, or the root of a new array/object
struct encoded struct encoded
+602
View File
@@ -0,0 +1,602 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint8_t, uint16_t, uint32_t, uint64_t
#include <cstring> // memcmp, memcpy
#include <limits> // numeric_limits
#include <string> // string
#include <vector> // vector
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/errors.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
#include <nlohmann/detail/view/number.hpp>
#include <nlohmann/detail/view/object_index.hpp>
#include <nlohmann/detail/view/scan.hpp>
// Images: a document stored so that loading it needs no parsing.
//
// Layout (little-endian): a 64-byte header, the nodes, the text (the source,
// followed by the number tokens written by edits), a NUL, the decoded strings
// (followed by the strings written by edits), a NUL. The idea is that of
// zero-copy formats such as FlatBuffers (https://github.com/google/flatbuffers)
// and YaFF (https://github.com/yandex/yaff); no code is taken from them.
// check_image follows the idea of FlatBuffers' Verifier (bounds and
// structure) and also checks what the parser guarantees about strings and
// numbers, so that reading and serializing a checked image is safe and yields
// valid JSON.
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
/// how load() checks an image
enum class image_check
{
/// everything the parser guarantees: structure and bounds, strings (valid
/// UTF-8; source strings without quotes, backslashes, and control
/// characters), and numbers (well-formed, matching the stored values)
full,
/// structure and bounds only: reading and serializing are safe, but a
/// crafted image can yield invalid UTF-8, strings that serialize to
/// invalid JSON, or numbers that differ from their text
bounds,
/// none: for images from a trusted source only (a damaged image is
/// undefined behavior)
none,
};
struct image_header
{
std::array<char, 4> magic; ///< "NJVI"
std::uint32_t version; ///< 1
std::uint64_t node_count;
std::uint64_t text_size;
std::uint64_t arena_size;
std::array<std::uint64_t, 4> reserved; ///< zero (for later versions)
};
static_assert(sizeof(image_header) == 64, "the image header must be 64 bytes");
constexpr std::uint32_t image_version = 1;
/// the largest node count and text or string size of an image (as for parsed
/// documents, offsets and counts must fit 32 bits)
constexpr std::uint64_t image_limit = 0xFFFFFFF0u;
/// Copy the current structure of an edited document into nodes in document
/// order, as the parser would have written them. Text written by edits is
/// appended to text_tail (number tokens) and arena_tail (strings); floats that
/// are not finite become null, as dump() writes them.
inline void compact_nodes(const document_data& d, std::size_t arena_size, std::vector<node>& out, std::string& text_tail, std::string& arena_tail)
{
struct frame
{
const node* cur;
const node* end;
std::size_t index; ///< the container's node in out
std::uint32_t count;
bool object;
};
std::vector<frame> stack;
const auto string_node = [&](const node & s)
{
node r = s;
r.extra = 0;
r.flags = static_cast<std::uint8_t>(s.flags & node_flags::storage);
if (r.flags == node_flags::edited)
{
r.off = static_cast<std::uint32_t>(arena_size + arena_tail.size());
arena_tail.append(d.str(s), s.len);
r.flags = node_flags::escaped;
}
return r;
};
const auto emit = [&](const node * v)
{
node r = *v;
switch (static_cast<value_t>(v->kind))
{
case value_t::object:
case value_t::array:
r.flags = 0;
r.extra = 0;
r.off = (v->flags & (node_flags::moved | node_flags::is_new)) != 0 ? 0 : v->off;
r.len = 0; // counted below
r.next = 0; // set when the container is complete
stack.push_back(frame{d.first_child_edited(v), d.child_end_edited(v), out.size(), 0, v->kind == static_cast<std::uint8_t>(value_t::object)});
break;
case value_t::string:
r = string_node(*v);
break;
case value_t::number_integer:
case value_t::number_unsigned:
if ((v->flags & node_flags::storage) == node_flags::edited)
{
r.off = static_cast<std::uint32_t>(d.size + text_tail.size());
text_tail.append(d.str(*v), number_length(*v));
}
r.flags = 0;
break;
case value_t::number_float:
if ((v->flags & node_flags::storage) == node_flags::edited)
{
const char* const t = d.str(*v);
if (t[0] == 'n' || t[0] == 'i' || (v->len > 1 && t[1] == 'i'))
{
r = node{}; // nan and infinity: null, as dump() writes them
r.kind = static_cast<std::uint8_t>(value_t::null);
break;
}
r.off = static_cast<std::uint32_t>(d.size + text_tail.size());
text_tail.append(t, v->len);
r.extra = 0xFFFFu; // the digit layout is not recorded
}
r.flags = 0;
break;
case value_t::boolean:
r.flags = static_cast<std::uint8_t>(v->flags & node_flags::is_true);
break;
case value_t::null:
case value_t::binary:
case value_t::discarded:
default:
r.flags = 0;
break;
}
out.push_back(r);
};
emit(d.tape);
while (!stack.empty())
{
frame& top = stack.back();
if (top.cur == top.end)
{
node& c = out[top.index];
c.len = top.count;
c.next = static_cast<std::uint32_t>(out.size() - top.index);
stack.pop_back();
continue;
}
++top.count;
const node* v = nullptr;
if (top.object)
{
out.push_back(string_node(*top.cur));
v = document_data::deref(top.cur + 1);
top.cur = document_data::after(top.cur + 1);
}
else
{
v = document_data::deref(top.cur);
top.cur = document_data::after(top.cur);
}
emit(v); // may grow the stack (top is not used afterwards)
}
}
/// the document as an image
inline std::vector<std::uint8_t> save_image(const document_data& d)
{
#if !NLOHMANN_VIEW_LITTLE_ENDIAN
throw_type_error(320, "json_document images need a little-endian target"); // LCOV_EXCL_LINE
#endif
const std::size_t arena_size = d.arena_size;
const node* nodes = d.tape;
std::size_t count = d.tape_size;
std::vector<node> compacted;
std::string text_tail;
std::string arena_tail;
if (d.edits)
{
compact_nodes(d, arena_size, compacted, text_tail, arena_tail);
nodes = compacted.data();
count = compacted.size();
}
const std::size_t text_size = d.size + text_tail.size();
const std::size_t total_arena = arena_size + arena_tail.size();
if (NLOHMANN_VIEW_UNLIKELY(text_size >= image_limit || total_arena >= image_limit || count >= image_limit))
{
// LCOV_EXCL_START (4 GiB)
throw_out_of_range(416, "images of 4 GiB or more are not supported by json_document");
// LCOV_EXCL_STOP
}
image_header h{};
h.magic = {{'N', 'J', 'V', 'I'}};
h.version = image_version;
h.node_count = count;
h.text_size = text_size;
h.arena_size = total_arena;
std::vector<std::uint8_t> image(sizeof(h) + (count * sizeof(node)) + text_size + 1 + total_arena + 1);
std::uint8_t* o = image.data();
std::memcpy(o, &h, sizeof(h));
o += sizeof(h);
std::memcpy(o, nodes, count * sizeof(node));
// the hash indexes are rebuilt by load()
for (std::size_t i = 0; i < count; ++i)
{
if (nodes[i].kind == static_cast<std::uint8_t>(value_t::object) && nodes[i].extra != 0)
{
node n = nodes[i];
n.extra = 0;
std::memcpy(o + (i * sizeof(node)), &n, sizeof(node));
}
}
o += count * sizeof(node);
const auto append = [&o](const char* s, std::size_t n)
{
if (n != 0)
{
std::memcpy(o, s, n);
o += n;
}
};
append(d.src, d.size);
append(text_tail.data(), text_tail.size());
*o++ = 0;
append(d.base[1], arena_size);
append(arena_tail.data(), arena_tail.size());
*o = 0;
return image;
}
/// whether a number node matches its token the way the parser records it
/// (after the bounds check)
inline bool check_number(const node& n, const unsigned char* text)
{
const std::size_t len = number_length(n);
const unsigned char* const s = text + n.off;
const unsigned char* const e = s + len;
const unsigned char* p = s;
const bool negative = *p == '-';
p += negative ? 1 : 0;
const unsigned char* const int_start = p;
if (p == e)
{
return false;
}
if (*p == '0')
{
++p;
}
else if (*p >= '1' && *p <= '9')
{
while (p != e && is_digit(*p))
{
++p;
}
}
else
{
return false;
}
const auto int_digits = static_cast<std::size_t>(p - int_start);
std::size_t frac_digits = 0;
bool is_float = false;
if (p != e && *p == '.')
{
const unsigned char* const f0 = ++p;
while (p != e && is_digit(*p))
{
++p;
}
if (p == f0)
{
return false;
}
frac_digits = static_cast<std::size_t>(p - f0);
is_float = true;
}
std::int64_t exponent = 0;
if (p != e && (*p | 0x20u) == 'e')
{
++p;
const bool exp_negative = p != e && *p == '-';
p += (p != e && (*p == '+' || *p == '-')) ? 1 : 0;
if (p == e || !is_digit(*p))
{
return false;
}
while (p != e && is_digit(*p))
{
exponent = exponent < 100000 ? (exponent * 10) + (*p - '0') : exponent;
++p;
}
exponent = exp_negative ? -exponent : exponent;
is_float = true;
}
if (p != e)
{
return false;
}
if (n.kind == static_cast<std::uint8_t>(value_t::number_float))
{
// the digit layout the parser records (or "many", as compaction
// writes it), and a finite value
const auto layout = static_cast<std::uint16_t>((int_digits < 255 ? int_digits : 255) | ((frac_digits < 255 ? frac_digits : 255) << 8u));
if (n.extra != layout && n.extra != 0xFFFFu)
{
return false;
}
// parse() rejects floats that overflow; as there, only a number whose
// magnitude could reach 1e308 needs the conversion
if (static_cast<std::int64_t>(int_digits) + exponent > 300)
{
const auto v = float_value<double>(reinterpret_cast<const char*>(s), n); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
return v <= (std::numeric_limits<double>::max)() && v >= -(std::numeric_limits<double>::max)();
}
return true;
}
// integers: the token's value is the stored one; number_integer nodes of
// edits can be non-negative (as basic_json keeps the type of a value)
const bool integer = n.kind == static_cast<std::uint8_t>(value_t::number_integer);
if (is_float || int_digits > 20 || (negative && !integer))
{
return false;
}
// (at most 19 digits cannot overflow; 20 digits are compared with 2^64 - 1)
if (int_digits == 20 && std::memcmp(int_start, "18446744073709551615", 20) > 0)
{
return false;
}
std::uint64_t m = 0;
for (const unsigned char* d = int_start; d != int_start + int_digits; ++d)
{
m = (m * 10) + static_cast<std::uint64_t>(*d - '0');
}
if (integer && m > (negative ? std::uint64_t{1} << 63u : (std::uint64_t{1} << 63u) - 1))
{
return false;
}
return integer_bits(n) == (negative ? 0 - m : m);
}
/// Check the nodes of a loaded image against its text and decoded strings:
/// kinds, flags, and `extra`; extents and element counts of arrays and
/// objects; keys; bounds; string contents (source strings as the parser
/// leaves them: no quotes, backslashes, or control characters; all strings
/// valid UTF-8); and number tokens.
inline bool check_image(const node* nodes, std::size_t count, const unsigned char* text, std::size_t text_size,
const unsigned char* arena, std::size_t arena_size, bool full)
{
struct frame
{
std::size_t end;
std::uint32_t len;
std::uint32_t seen;
bool object;
bool expect_key;
};
std::vector<frame> stack;
const auto check_string = [&](const node & n) -> bool
{
if ((n.flags & ~node_flags::escaped) != 0 || n.extra != 0)
{
return false;
}
const bool decoded = (n.flags & node_flags::escaped) != 0;
const unsigned char* const base = decoded ? arena : text;
const std::size_t limit = decoded ? arena_size : text_size;
if (n.off > limit || n.len > limit - n.off)
{
return false;
}
if (!full)
{
return true;
}
const unsigned char* const b = base + n.off;
return decoded ? valid_utf8_prefix(b, n.len) == n.len : scan_string_run(b, b + n.len) == b + n.len;
};
// bounds of a number token; the recorded digit layout must lie within it
const auto number_in_bounds = [&](const node & n) -> bool
{
const std::size_t len = number_length(n);
if (len == 0 || n.off > text_size || len > text_size - n.off)
{
return false;
}
if (n.kind != static_cast<std::uint8_t>(value_t::number_float))
{
return (n.extra >> 8u) == 0;
}
// float_value() reads the sign, the integer digits, and the point and
// fraction digits the layout records (a layout of more than 19 digits
// means the general conversion, which stays within the token)
const std::size_t int_digits = n.extra & 0xFFu;
const std::size_t frac_digits = n.extra >> 8u;
const std::size_t need = (text[n.off] == '-' ? 1u : 0u) + int_digits + (frac_digits != 0 ? frac_digits + 1 : 0);
return int_digits + frac_digits > 19 || need <= len;
};
std::size_t i = 0;
for (;;)
{
// close finished arrays and objects
while (!stack.empty() && i == stack.back().end)
{
const frame f = stack.back();
if (f.seen != f.len || (f.object && !f.expect_key))
{
return false;
}
stack.pop_back();
if (!stack.empty())
{
++stack.back().seen;
stack.back().expect_key = true;
}
}
if (i == count)
{
return stack.empty();
}
if (i != 0 && stack.empty())
{
return false; // nodes after the root
}
const node& n = nodes[i];
if (!stack.empty() && stack.back().object && stack.back().expect_key)
{
if (n.kind != static_cast<std::uint8_t>(value_t::string) || !check_string(n))
{
return false;
}
stack.back().expect_key = false;
++i;
continue;
}
bool complete = true;
switch (static_cast<value_t>(n.kind))
{
case value_t::null:
// (the offset of a literal is read to size the output of dump())
if (n.flags != 0 || n.extra != 0 || n.off > text_size)
{
return false;
}
break;
case value_t::boolean:
if ((n.flags & ~node_flags::is_true) != 0 || n.extra != 0 || n.off > text_size)
{
return false;
}
break;
case value_t::string:
if (!check_string(n))
{
return false;
}
break;
case value_t::number_integer:
case value_t::number_unsigned:
case value_t::number_float:
if (n.flags != 0 || !number_in_bounds(n) || (full && !check_number(n, text)))
{
return false;
}
break;
case value_t::array:
case value_t::object:
{
const std::size_t limit = stack.empty() ? count : stack.back().end;
if (n.flags != 0 || n.extra != 0 || n.next == 0 || n.next > limit - i || n.off > text_size)
{
return false;
}
stack.push_back(frame{i + n.next, n.len, 0, n.kind == static_cast<std::uint8_t>(value_t::object), true});
complete = false;
break;
}
case value_t::binary:
case value_t::discarded:
default:
return false;
}
++i;
if (complete && !stack.empty())
{
++stack.back().seen;
stack.back().expect_key = true;
}
}
}
[[noreturn]] NLOHMANN_VIEW_NOINLINE inline void throw_invalid_image(const char* what)
{
throw_parse_error(116, concat("invalid json_document image: ", what));
}
/// Read an image into d. The text and the decoded strings stay in the image;
/// the nodes are copied (so that they are aligned, and edits can change them).
inline void load_image(document_data& d, const std::uint8_t* image, std::size_t size, image_check check)
{
#if !NLOHMANN_VIEW_LITTLE_ENDIAN
throw_type_error(320, "json_document images need a little-endian target"); // LCOV_EXCL_LINE
#endif
if (image == nullptr || size < sizeof(image_header))
{
throw_invalid_image("too short");
}
image_header h{};
std::memcpy(&h, image, sizeof(h));
// (the reserved fields are for later versions)
if (std::memcmp(h.magic.data(), "NJVI", 4) != 0 || h.version != image_version
|| (h.reserved[0] | h.reserved[1] | h.reserved[2] | h.reserved[3]) != 0)
{
throw_invalid_image("unknown format");
}
const std::size_t room = size - sizeof(h);
if (h.node_count == 0 || h.node_count > room / sizeof(node) || h.node_count >= image_limit || h.text_size >= image_limit || h.arena_size >= image_limit)
{
throw_invalid_image("sizes out of range");
}
const auto count = static_cast<std::size_t>(h.node_count);
const auto text_size = static_cast<std::size_t>(h.text_size);
const auto arena_size = static_cast<std::size_t>(h.arena_size);
const std::size_t text_at = sizeof(h) + (count * sizeof(node));
// the text, a NUL, the decoded strings, a NUL, and nothing after them
if (size - text_at < 2 || text_size > size - text_at - 2 || arena_size != size - text_at - text_size - 2
|| image[text_at + text_size] != 0 || image[size - 1] != 0)
{
throw_invalid_image("sizes out of range");
}
d.discarded = true;
d.edits.reset();
d.base[2] = nullptr;
d.owned.clear();
if (d.owned_image.empty() || image != d.owned_image.data())
{
d.owned_image.clear();
}
d.arena.clear();
d.indexes.clear();
d.index_slots.clear();
d.large_objects.clear();
d.tape_size = 0;
d.reserve(count);
std::memcpy(d.tape, image + sizeof(h), count * sizeof(node));
d.tape_size = count;
const std::uint8_t* const text = image + text_at;
const std::uint8_t* const arena = text + text_size + 1;
d.src = reinterpret_cast<const char*>(text); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
d.size = text_size;
d.base[0] = d.src;
d.base[1] = reinterpret_cast<const char*>(arena); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
d.arena_size = arena_size;
if (check != image_check::none && !check_image(d.tape, count, text, text_size, arena, arena_size, check == image_check::full))
{
throw_invalid_image("the check failed");
}
// the hash indexes of large objects, as after parsing
for (std::size_t i = 0; i < count; ++i)
{
node& n = d.tape[i];
if (n.kind == static_cast<std::uint8_t>(value_t::object))
{
n.extra = 0;
if (n.len >= document_data::index_min_members)
{
d.large_objects.push_back(static_cast<std::uint32_t>(i));
}
}
}
build_object_indexes(d);
d.discarded = false;
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -22,5 +22,7 @@
#undef NLOHMANN_VIEW_NEON #undef NLOHMANN_VIEW_NEON
#undef NLOHMANN_VIEW_SSE2 #undef NLOHMANN_VIEW_SSE2
#undef NLOHMANN_VIEW_SSSE3 #undef NLOHMANN_VIEW_SSSE3
#undef NLOHMANN_VIEW_SSSE3_DISPATCH
#undef NLOHMANN_VIEW_SSSE3_TARGET
#undef NLOHMANN_VIEW_VECTOR #undef NLOHMANN_VIEW_VECTOR
#undef NLOHMANN_VIEW_VECTOR_UTF8 #undef NLOHMANN_VIEW_VECTOR_UTF8
+67 -30
View File
@@ -26,43 +26,76 @@ namespace detail
namespace view namespace view
{ {
/*!
@brief locate the decimal point and the end of the mantissa of a float token
Also checks that the token is a JSON number. Tokens of the parser and of edits
always are; an image loaded with image_check::bounds can hold any bytes, which
must not reach the conversion (it expects a well-formed token).
*/
inline bool float_token_layout(const char* first, const char* last, std::size_t& dot, std::size_t& mantissa_end) noexcept
{
const auto digit = [last](const char* q)
{
return q != last && is_digit(static_cast<unsigned char>(*q));
};
const char* p = first;
p += (p != last && *p == '-') ? 1 : 0;
if (!digit(p) || (*p == '0' && digit(p + 1)))
{
return false;
}
while (digit(p))
{
++p;
}
dot = std::string::npos;
if (p != last && *p == '.')
{
dot = static_cast<std::size_t>(p - first);
if (!digit(++p))
{
return false;
}
while (digit(p))
{
++p;
}
}
mantissa_end = static_cast<std::size_t>(p - first);
if (p != last && (*p == 'e' || *p == 'E'))
{
++p;
p += (p != last && (*p == '+' || *p == '-')) ? 1 : 0;
if (!digit(p))
{
return false;
}
while (digit(p))
{
++p;
}
}
return p == last;
}
/*! /*!
@brief the value of the float token of a node, as parse() converts it @brief the value of the float token of a node, as parse() converts it
Uses the lexer's conversion (detail::convert_float), so that the values are Uses the lexer's conversion (detail::convert_float), so that the values are
bit-identical to parse(): float and double are converted without allocation bit-identical to parse(): float and double are converted without allocation
and independent of the locale. The digit layout recorded while parsing locates and independent of the locale. A token that is not a JSON number (only in a
the decimal point and the exponent without scanning the token. damaged image loaded with image_check::bounds) yields 0.
*/ */
template<typename FloatType> template<typename FloatType>
NLOHMANN_VIEW_NOINLINE FloatType float_value(const char* first, const node& n) NLOHMANN_VIEW_NOINLINE FloatType float_value(const char* first, const node& n)
{ {
const char* const last = first + n.len; const char* const last = first + n.len;
const std::size_t neg = first[0] == '-' ? 1 : 0; std::size_t dot = 0;
const std::size_t int_digits = n.extra & 0xFFu; std::size_t mantissa_end = 0;
const std::size_t frac_digits = n.extra >> 8u; if (NLOHMANN_VIEW_UNLIKELY(!float_token_layout(first, last, dot, mantissa_end)))
std::size_t dot = std::string::npos;
std::size_t mantissa_end = n.len;
if (int_digits != 255 && frac_digits != 255)
{ {
dot = frac_digits != 0 ? neg + int_digits : std::string::npos; return FloatType{};
mantissa_end = neg + int_digits + (frac_digits != 0 ? 1 + frac_digits : 0);
}
else
{
// more digits than the layout records: locate them
for (std::size_t i = 0; i < n.len; ++i)
{
if (first[i] == '.')
{
dot = i;
}
else if (first[i] == 'e' || first[i] == 'E')
{
mantissa_end = i;
break;
}
}
} }
return convert_float<FloatType>(first, last, dot, mantissa_end); return convert_float<FloatType>(first, last, dot, mantissa_end);
} }
@@ -93,16 +126,20 @@ NLOHMANN_VIEW_ALWAYS_INLINE float_significand layout_decimal(const unsigned char
} }
if (p != e) if (p != e)
{ {
// [eE][+-]digits; huge exponents saturate (the parser rejected overflow) // [eE][+-]digits; huge exponents saturate (the parser rejected
// overflow). The token is not read beyond e, and the digits are taken
// as unsigned, so that a token that is not well-formed (a damaged
// image loaded with image_check::bounds) yields a wrong value, but no
// overflow.
++p; ++p;
const bool exp_negative = *p == '-'; const bool exp_negative = p != e && *p == '-';
p += (*p == '-' || *p == '+') ? 1 : 0; p += (p != e && (*p == '-' || *p == '+')) ? 1 : 0;
std::int64_t exp_value = 0; std::int64_t exp_value = 0;
for (; p != e; ++p) for (; p != e; ++p)
{ {
if (exp_value < 0x10000000) if (exp_value < 0x10000000)
{ {
exp_value = (exp_value * 10) + (*p - '0'); exp_value = (exp_value * 10) + static_cast<unsigned char>(*p - '0');
} }
} }
q += exp_negative ? -exp_value : exp_value; q += exp_negative ? -exp_value : exp_value;
+17 -2
View File
@@ -67,13 +67,19 @@ NLOHMANN_VIEW_ALWAYS_INLINE std::uint16_t load16(const unsigned char* p) noexcep
/// in predicted branches: 16 for keys, whose lengths repeat from record to /// in predicted branches: 16 for keys, whose lengths repeat from record to
/// record, and 8 for string values (Value) where a vector loop follows, as /// record, and 8 for string values (Value) where a vector loop follows, as
/// their lengths vary more. Longer runs continue 16 bytes at a time with NEON /// their lengths vary more. Longer runs continue 16 bytes at a time with NEON
/// or SSE2, else eight bytes at a time. /// or SSE2, else eight bytes at a time. With SSE2, the run is checked 16 bytes
/// at a time from its first byte instead: on x86-64, one compare that finds
/// the end of most keys and short values is faster than a branch per byte (on
/// AArch64, where a NEON mask costs more and branches predict well, slower).
template<bool Value = false> template<bool Value = false>
NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned char* p, const unsigned char* e) noexcept NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned char* p, const unsigned char* e) noexcept
{ {
const std::uint8_t* plain = string_plain(); const std::uint8_t* plain = string_plain();
for (;;) for (;;)
{ {
#if NLOHMANN_VIEW_SSE2
p = vector_plain_run(p, e);
#else
if (e - p >= 16) if (e - p >= 16)
{ {
#define NLOHMANN_VIEW_STEP(i) if (NLOHMANN_VIEW_LIKELY(plain[p[i]] != 0)) {} else { p += (i); goto stop; } #define NLOHMANN_VIEW_STEP(i) if (NLOHMANN_VIEW_LIKELY(plain[p[i]] != 0)) {} else { p += (i); goto stop; }
@@ -107,6 +113,7 @@ NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned
#endif #endif
continue; continue;
} }
#endif
while (p != e && plain[*p] != 0) while (p != e && plain[*p] != 0)
{ {
++p; ++p;
@@ -115,15 +122,23 @@ NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned
{ {
return p; return p;
} }
#if !NLOHMANN_VIEW_SSE2
stop: stop:
#endif
if (*p < 0x80) if (*p < 0x80)
{ {
return p; // quote, backslash, or control character return p; // quote, backslash, or control character
} }
#if NLOHMANN_VIEW_VECTOR_UTF8 #if NLOHMANN_VIEW_VECTOR_UTF8
#if NLOHMANN_VIEW_SSSE3_DISPATCH
if (NLOHMANN_VIEW_LIKELY(cpu_has_ssse3()))
#endif
{
// non-ASCII: the vector check, out of line // non-ASCII: the vector check, out of line
return scan_string_vector(p, e, plain); return scan_string_vector(p, e, plain);
#else }
#endif
#if !NLOHMANN_VIEW_VECTOR_UTF8 || NLOHMANN_VIEW_SSSE3_DISPATCH
// non-ASCII: a run of well-formed sequences (the library's check, so // non-ASCII: a run of well-formed sequences (the library's check, so
// that exactly what json::parse accepts is accepted) // that exactly what json::parse accepts is accepted)
do do
+507 -7
View File
@@ -75,6 +75,23 @@ class output_buffer
m_pos += n; m_pos += n;
} }
/// the write position and the end of the writable space, for a writer
/// that keeps the position in a local variable (set_cursor() hands it back)
char* cursor() const noexcept
{
return m_pos;
}
char* limit() const noexcept
{
return m_end;
}
void set_cursor(char* p) noexcept
{
m_pos = p;
}
private: private:
static StringType& sized(StringType& out, std::size_t estimate) static StringType& sized(StringType& out, std::size_t estimate)
{ {
@@ -95,6 +112,105 @@ class output_buffer
char* m_end; char* m_end;
}; };
/// The length of the run at s that dump() writes unchanged without
/// ensure_ascii: all bytes but quotes, backslashes, and control characters.
/// Unlike detail::string_bulk_run(), non-ASCII bytes are not validated: the
/// strings of a document are valid UTF-8 (a damaged image loaded with
/// image_check::bounds can have others, which are then written unchanged).
inline std::size_t plain_output_run(const unsigned char* s, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
const std::uint64_t v = read_eight_bytes(s + i);
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"'
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\'
const std::uint64_t stop = (((q - ones) & ~q) | ((b - ones) & ~b) | ((v - 0x2020202020202020ull) & ~v)) & high;
if (stop != 0)
{
// the lowest flagged byte is the first stop: borrows only flag bytes above a true one
return i + (static_cast<std::size_t>(count_trailing_zeros(stop)) / 8);
}
}
for (; i < n; ++i)
{
if (s[i] == '"' || s[i] == '\\' || s[i] < 0x20)
{
return i;
}
}
return n;
}
/// A stack that starts in a buffer of the caller (a local array) and moves to
/// the heap (a vector of the caller) only when that is full, so that dumps of
/// shallow documents need no allocation. The top is a pointer, as in
/// std::vector. The address of the stack never escapes (the growth gets the
/// vector and returns the new storage), so its pointers stay in registers.
template<typename T>
class small_stack
{
public:
small_stack(T* buffer, std::size_t capacity, std::vector<T>& heap) noexcept
: m_begin(buffer), m_top(buffer), m_end(buffer + capacity), m_heap(&heap)
{}
small_stack(const small_stack&) = delete;
small_stack(small_stack&&) = delete;
small_stack& operator=(const small_stack&) = delete;
small_stack& operator=(small_stack&&) = delete;
~small_stack() = default;
NLOHMANN_VIEW_ALWAYS_INLINE void push_back(const T& x)
{
if (NLOHMANN_VIEW_UNLIKELY(m_top == m_end))
{
const std::size_t used = size();
const std::size_t capacity = 2 * static_cast<std::size_t>(m_end - m_begin);
m_begin = grow(*m_heap, m_begin, used, capacity);
m_top = m_begin + used;
m_end = m_begin + capacity;
}
*m_top++ = x;
}
NLOHMANN_VIEW_ALWAYS_INLINE T& back() noexcept
{
return m_top[-1];
}
NLOHMANN_VIEW_ALWAYS_INLINE void pop_back() noexcept
{
--m_top;
}
NLOHMANN_VIEW_ALWAYS_INLINE bool empty() const noexcept
{
return m_top == m_begin;
}
NLOHMANN_VIEW_ALWAYS_INLINE std::size_t size() const noexcept
{
return static_cast<std::size_t>(m_top - m_begin);
}
private:
/// the used entries moved to heap storage of the given capacity
NLOHMANN_VIEW_NOINLINE static T* grow(std::vector<T>& heap, const T* begin, std::size_t used, std::size_t capacity)
{
std::vector<T> bigger(capacity);
std::copy(begin, begin + used, bigger.begin());
heap.swap(bigger);
return heap.data();
}
T* m_begin;
T* m_top;
T* m_end;
std::vector<T>* m_heap;
};
/// how the view's dump() writes a value /// how the view's dump() writes a value
struct dump_style struct dump_style
{ {
@@ -129,6 +245,18 @@ class view_serializer
void dump(const node* root) void dump(const node* root)
{ {
if (!m_style.pretty && !m_style.ensure_ascii)
{
if (m_style.source_numbers)
{
dump_compact<true>(root);
}
else
{
dump_compact<false>(root);
}
return;
}
struct frame struct frame
{ {
const node* pos; ///< next element, or key of the next member const node* pos; ///< next element, or key of the next member
@@ -136,7 +264,9 @@ class view_serializer
bool object; bool object;
bool first; ///< nothing written yet bool first; ///< nothing written yet
}; };
std::vector<frame> stack; std::array<frame, 32> buffer; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
std::vector<frame> heap;
small_stack<frame> stack(buffer.data(), buffer.size(), heap);
const node* n = root; const node* n = root;
for (;;) for (;;)
{ {
@@ -207,6 +337,289 @@ class view_serializer
} }
private: private:
/*!
@brief the compact output without ensure_ascii (the default dump())
The same walk as dump(), with the write position in a local variable
(stores through char pointers would otherwise force a reload of the
buffer's members after each one), and with strings and number tokens of
the source copied by fixed-size moves of 32 bytes where the source has
that many bytes left, instead of a library call per token. The buffer
keeps 64 bytes of slack for the overshoot.
*/
/// a string that is not a plain string of the source (decoded, or written
/// by an edit), without ensure_ascii: runs without characters to escape
/// are copied
NLOHMANN_VIEW_NOINLINE void write_decoded(const node& n)
{
const auto* const s = reinterpret_cast<const unsigned char*>(m_doc.str(n)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
m_out.put('"');
for (std::size_t i = 0; i < n.len;)
{
const std::size_t run = plain_output_run(s + i, n.len - i);
if (run != 0)
{
m_out.put(reinterpret_cast<const char*>(s + i), run); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
i += run;
continue;
}
write_codepoint<false>(s[i], s + i, 1); // a quote, a backslash, or a control character
++i;
}
m_out.put('"');
}
/// the copies of dump_compact() that are not fixed-size moves (long
/// strings, or near the end of the source); out of line, so that the
/// compiler does not merge the fixed-size moves into this call
NLOHMANN_VIEW_NOINLINE static void copy_long(char* to, const char* from, std::size_t n) noexcept
{
std::memcpy(to, from, n);
}
template<bool SourceNumbers>
void dump_compact(const node* root)
{
struct frame
{
const node* pos; ///< (editable documents) next element, or key of the next member
const node* end;
bool object;
};
std::array<frame, 32> buffer; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
std::vector<frame> heap;
small_stack<frame> stack(buffer.data(), buffer.size(), heap);
const char* const src = m_doc.src;
const char* const src_end = src + m_doc.size;
char* w = m_out.cursor();
char* lim = m_out.limit();
// room for n bytes and the slack
const auto room = [&](std::size_t n)
{
if (NLOHMANN_VIEW_UNLIKELY(static_cast<std::size_t>(lim - w) < n + 64))
{
m_out.set_cursor(w);
m_out.reserve(n + 64);
w = m_out.cursor();
lim = m_out.limit();
}
};
// copy n bytes of the source (after room(n))
const auto copy = [&](const char* from, std::size_t n)
{
if (n <= 32 && src_end - from >= 32)
{
std::memcpy(w, from, 32);
}
else if (n <= 256 && src_end - from >= static_cast<std::ptrdiff_t>(n) + 32)
{
for (std::size_t i = 0; i < n; i += 32)
{
std::memcpy(w + i, from + i, 32);
}
}
else
{
copy_long(w, from, n);
}
w += n;
};
// a literal of n bytes (after room(n))
const auto literal = [&](const char* text, std::size_t n)
{
std::memcpy(w, text, n);
w += n;
};
// a string that is not a plain string of the source (out of line, so
// that the cursor stays in a register here)
const auto escaped = [&](const node & n)
{
m_out.set_cursor(w);
write_decoded(n);
w = m_out.cursor();
lim = m_out.limit();
};
// Read-only documents: the elements of a container follow it in the
// node array, so the walk goes through the array in order, and a
// frame only needs the end of its container. Editable documents: the
// elements of a moved container live elsewhere, so a frame keeps the
// position of the next element (see navigation).
// The innermost open container is kept in registers (cur; end ==
// nullptr: none), the stack holds the ones around it.
frame cur{nullptr, nullptr, false};
const node* n = root;
for (;;)
{
// write the value at n (read-only documents: and advance n)
bool opened = false;
switch (static_cast<value_t>(n->kind))
{
case value_t::string:
if ((n->flags & node_flags::storage) == 0)
{
room(n->len + 2);
*w++ = '"';
copy(src + n->off, n->len);
*w++ = '"';
}
else
{
escaped(*n);
}
break;
case value_t::number_integer:
case value_t::number_unsigned:
{
const std::uint32_t len = number_length(*n);
room(len);
if (Editable && (n->flags & node_flags::storage) != 0)
{
copy_long(w, m_doc.str(*n), len); // a canonical token written by an edit
w += len;
break;
}
const char* const token = src + n->off;
if (!SourceNumbers && NLOHMANN_VIEW_UNLIKELY(len == 2 && token[0] == '-' && token[1] == '0'))
{
*w++ = '0'; // parse() reads -0 as the integer 0
}
else
{
copy(token, len);
}
break;
}
case value_t::number_float:
if (SourceNumbers && (n->flags & node_flags::storage) != node_flags::edited)
{
room(n->len);
copy(src + n->off, n->len);
}
else if (std::is_same<number_float_t, double>::value)
{
room(64);
w = write_double_at(w, *n);
}
else
{
m_out.set_cursor(w);
write_float_node(*n);
w = m_out.cursor();
lim = m_out.limit();
}
break;
case value_t::boolean:
room(8);
if ((n->flags & node_flags::is_true) != 0)
{
literal("true", 4);
}
else
{
literal("false", 5);
}
break;
case value_t::object:
case value_t::array:
{
const bool object = n->kind == static_cast<std::uint8_t>(value_t::object);
room(8);
if (n->len == 0)
{
literal(object ? "{}" : "[]", 2);
}
else
{
*w++ = object ? '{' : '[';
stack.push_back(cur);
if (Editable)
{
cur = frame{nav::first(m_doc, n), nav::end(m_doc, n), object};
}
else
{
cur = frame{nullptr, n + n->next, object};
}
opened = true;
}
break;
}
case value_t::null:
room(8);
literal("null", 4);
break;
case value_t::binary: // LCOV_EXCL_LINE (not in a document)
case value_t::discarded: // LCOV_EXCL_LINE
default: // LCOV_EXCL_LINE
break; // LCOV_EXCL_LINE
}
if (!Editable)
{
++n; // the next node: the first element of an opened container, or the node after a scalar
}
// go to the next value: close finished containers, then separate
// (a container just opened has an element)
if (!opened)
{
for (;;)
{
if (cur.end == nullptr)
{
m_out.set_cursor(w);
m_out.finish();
return;
}
if ((Editable ? cur.pos : n) != cur.end)
{
break;
}
room(1);
*w++ = cur.object ? '}' : ']';
cur = stack.back();
stack.pop_back();
}
room(1);
*w++ = ',';
}
const node* const at = Editable ? cur.pos : n;
if (cur.object)
{
const node& key = *at;
if ((key.flags & node_flags::storage) == 0)
{
room(key.len + 3);
*w++ = '"';
copy(src + key.off, key.len);
w[0] = '"';
w[1] = ':';
w += 2;
}
else
{
escaped(key);
room(1);
*w++ = ':';
}
if (Editable)
{
n = nav::value(at + 1);
cur.pos = document_data::after(at + 1);
}
else
{
++n;
}
}
else if (Editable)
{
n = nav::value(at);
cur.pos = document_data::after(at);
}
}
}
void newline(std::size_t level) void newline(std::size_t level)
{ {
if (m_style.pretty) if (m_style.pretty)
@@ -258,7 +671,7 @@ class view_serializer
} }
else else
{ {
write_float(float_value<number_float_t>(m_doc, n)); write_float_node(n);
} }
break; break;
case value_t::object: // LCOV_EXCL_LINE (containers are written by dump()) case value_t::object: // LCOV_EXCL_LINE (containers are written by dump())
@@ -270,6 +683,83 @@ class view_serializer
} }
} }
/// a float node as dump() writes it
void write_float_node(const node& n)
{
write_float_node(n, std::is_same<number_float_t, double> {});
}
void write_float_node(const node& n, std::false_type /*other*/)
{
write_float(float_value<number_float_t>(m_doc, n));
}
void write_float_node(const node& n, std::true_type /*double*/)
{
m_out.reserve(64);
m_out.set_cursor(write_double_at(m_out.cursor(), n));
}
/*!
@brief (doubles) the float at n as dump() writes it, at w (64 bytes of room)
A token of at most 15 significant digits is written from its digits,
without a conversion: two decimals of at most 15 digits are farther
apart than the rounding interval of a (normal) double (the argument
behind DBL_DIG), so the token's digits are the shortest ones of its
double, which the library's conversion writes (Zmij). Other tokens are
converted from the digits already read.
*/
char* write_double_at(char* w, const node& n)
{
const unsigned int_digits = n.extra & 0xFFu;
const unsigned frac_digits = n.extra >> 8u;
if ((n.flags & node_flags::storage) != node_flags::edited && int_digits + frac_digits <= 19)
{
const auto* const first = reinterpret_cast<const unsigned char*>(m_doc.src + n.off); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
const float_significand d = layout_decimal(first, first + n.len, int_digits, frac_digits, reinterpret_cast<const unsigned char*>(m_doc.src + m_doc.size)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
// (the exponent keeps the value far from subnormals and overflow)
if (d.w != 0 && d.w < 1000000000000000u && d.exponent >= -290 && d.exponent <= 290)
{
*w = '-';
w += d.negative ? 1 : 0;
// (without leading zeros, all digits of the token count)
const unsigned char lead = first[d.negative ? 1 : 0];
return lead != '0' ? ::nlohmann::detail::dtoa_impl::write_short_decimal(w, d.w, static_cast<int>(int_digits + frac_digits), static_cast<int>(d.exponent))
: ::nlohmann::detail::dtoa_impl::write_short_decimal(w, d.w, static_cast<int>(d.exponent));
}
return write_double_value_at(w, decimal_to_float<double>(d)); // (without reading the token again)
}
return write_double_value_at(w, static_cast<double>(float_value<number_float_t>(m_doc, n)));
}
/// n bytes of text at w
static char* write_text_at(char* w, const char* text, std::size_t n) noexcept
{
std::memcpy(w, text, n);
return w + n;
}
/// a double as dump() writes it, at w (64 bytes of room)
static char* write_double_value_at(char* w, double x)
{
// (from the bits: without the checks of to_chars())
std::uint64_t bits = 0;
std::memcpy(&bits, &x, sizeof(bits));
if (NLOHMANN_VIEW_UNLIKELY((bits & 0x7FF0000000000000u) == 0x7FF0000000000000u))
{
return write_text_at(w, "null", 4);
}
*w = '-';
w += bits >> 63u;
bits &= ~(std::uint64_t{1} << 63u);
if (bits == 0)
{
return write_text_at(w, "0.0", 3);
}
return ::nlohmann::detail::dtoa_impl::write_shortest(w, ::nlohmann::detail::zmij::to_shortest(bits));
}
/// as serializer::dump_float() /// as serializer::dump_float()
void write_float(number_float_t x) void write_float(number_float_t x)
{ {
@@ -317,7 +807,9 @@ class view_serializer
m_out.put('"'); m_out.put('"');
} }
/// as serializer::dump_escaped() for valid UTF-8 (the view has no other) /// as serializer::dump_escaped(); strings of a document are valid UTF-8,
/// except in a damaged image loaded with image_check::bounds, for which
/// this throws what basic_json::dump() throws for the string
template<bool EnsureAscii> template<bool EnsureAscii>
void write_escaped(const unsigned char* s, std::size_t n) void write_escaped(const unsigned char* s, std::size_t n)
{ {
@@ -341,12 +833,13 @@ class view_serializer
} }
std::uint32_t codepoint = s[i]; std::uint32_t codepoint = s[i];
std::size_t len = 1; std::size_t len = 1;
if (codepoint >= 0xC0) if (codepoint >= 0x80)
{ {
len = 2; len = validate_one_utf8(s + i, n - i);
if (codepoint >= 0xE0) if (NLOHMANN_VIEW_UNLIKELY(len == 0))
{ {
len = codepoint >= 0xF0 ? 4 : 3; invalid_utf8(s, n);
return;
} }
codepoint &= 0xFFu >> (len + 1); codepoint &= 0xFFu >> (len + 1);
for (std::size_t k = 1; k < len; ++k) for (std::size_t k = 1; k < len; ++k)
@@ -359,6 +852,13 @@ class view_serializer
} }
} }
/// throw what basic_json::dump() throws for a string that is not valid UTF-8
NLOHMANN_VIEW_NOINLINE static void invalid_utf8(const unsigned char* s, std::size_t n)
{
const string_t dumped = BasicJsonType(string_t(reinterpret_cast<const char*>(s), n)).dump(); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
static_cast<void>(dumped);
}
template<bool EnsureAscii> template<bool EnsureAscii>
void write_codepoint(std::uint32_t codepoint, const unsigned char* bytes, std::size_t len) void write_codepoint(std::uint32_t codepoint, const unsigned char* bytes, std::size_t len)
{ {
+61 -7
View File
@@ -10,6 +10,7 @@
#pragma once #pragma once
#include <array> // array #include <array> // array
#include <atomic> // atomic
#include <cstddef> // size_t #include <cstddef> // size_t
#include <cstdint> // uint8_t, uint64_t #include <cstdint> // uint8_t, uint64_t
@@ -18,11 +19,13 @@
// Vector code for long runs of string bytes. NEON (AArch64) and SSE2 (x86-64) // Vector code for long runs of string bytes. NEON (AArch64) and SSE2 (x86-64)
// belong to the baseline instruction sets and are used by default. The vector // belong to the baseline instruction sets and are used by default. The vector
// UTF-8 check needs NEON, or SSSE3 if JSON_VIEW_USE_SSSE3 is defined: SSSE3 is // UTF-8 check needs NEON or SSSE3. SSSE3 is not part of x86-64, and the code
// not part of x86-64, so it must not depend on the flags of a translation unit // must not depend on the flags of a translation unit (two translation units
// (two translation units with different flags would have different // with different flags would have different definitions of the same inline
// definitions of the same inline functions). JSON_VIEW_NO_SIMD selects the // functions): the check is compiled for SSSE3 with a function attribute and
// portable code. // used where the CPU has SSSE3 (all x86-64 CPUs since about 2011), else the
// portable check. JSON_VIEW_USE_SSSE3 skips the CPU check (for code compiled
// for SSSE3 anyway); JSON_VIEW_NO_SIMD selects the portable code.
#if !defined(JSON_VIEW_NO_SIMD) && defined(__aarch64__) && (defined(__GNUC__) || defined(__clang__)) && NLOHMANN_VIEW_LITTLE_ENDIAN #if !defined(JSON_VIEW_NO_SIMD) && defined(__aarch64__) && (defined(__GNUC__) || defined(__clang__)) && NLOHMANN_VIEW_LITTLE_ENDIAN
#include <arm_neon.h> #include <arm_neon.h>
#define NLOHMANN_VIEW_NEON 1 #define NLOHMANN_VIEW_NEON 1
@@ -41,8 +44,24 @@
#else #else
#define NLOHMANN_VIEW_SSSE3 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum) #define NLOHMANN_VIEW_SSSE3 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#endif #endif
#if NLOHMANN_VIEW_SSE2 && !NLOHMANN_VIEW_SSSE3 && ((defined(__clang__) && __clang_major__ >= 4) || (defined(__GNUC__) && !defined(__clang__) && (__GNUC__ > 4 || (__GNUC__ == 4 && __GNUC_MINOR__ >= 9))))
// (GCC before 4.9 has no SSSE3 intrinsics without -mssse3)
#include <cpuid.h>
#include <tmmintrin.h>
#define NLOHMANN_VIEW_SSSE3_DISPATCH 1 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET __attribute__((target("ssse3")))
#elif NLOHMANN_VIEW_SSE2 && !NLOHMANN_VIEW_SSSE3 && defined(_MSC_VER)
// (MSVC compiles intrinsics of any instruction set)
#include <intrin.h>
#include <tmmintrin.h>
#define NLOHMANN_VIEW_SSSE3_DISPATCH 1 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET
#else
#define NLOHMANN_VIEW_SSSE3_DISPATCH 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#define NLOHMANN_VIEW_SSSE3_TARGET
#endif
#define NLOHMANN_VIEW_VECTOR (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSE2) #define NLOHMANN_VIEW_VECTOR (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSE2)
#define NLOHMANN_VIEW_VECTOR_UTF8 (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSSE3) #define NLOHMANN_VIEW_VECTOR_UTF8 (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSSE3 || NLOHMANN_VIEW_SSSE3_DISPATCH)
NLOHMANN_JSON_NAMESPACE_BEGIN NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail namespace detail
@@ -90,6 +109,40 @@ NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* vector_plain_run(const unsigned
} }
#endif #endif
#if NLOHMANN_VIEW_SSSE3_DISPATCH
/// whether the CPU has SSSE3 (CPUID leaf 1, ECX bit 9)
inline bool cpu_ssse3() noexcept
{
#if defined(_MSC_VER) && !defined(__clang__)
std::array<int, 4> regs {{}};
__cpuid(regs.data(), 1);
return (static_cast<unsigned>(regs[2]) & (1u << 9u)) != 0;
#else
unsigned eax = 0;
unsigned ebx = 0;
unsigned ecx = 0;
unsigned edx = 0;
return __get_cpuid(1, &eax, &ebx, &ecx, &edx) != 0 && (ecx & (1u << 9u)) != 0;
#endif
}
/// whether the CPU has SSSE3, asked once: the answer is kept in an atomic
/// that is initialized at compile time, so that neither a guard of a local
/// static nor a global constructor is needed (threads that ask at the same
/// time all store the same answer)
NLOHMANN_VIEW_ALWAYS_INLINE bool cpu_has_ssse3() noexcept
{
static std::atomic<int> known{0}; // 0: not asked yet, 1: no, 2: yes
int state = known.load(std::memory_order_relaxed);
if (NLOHMANN_VIEW_UNLIKELY(state == 0))
{
state = cpu_ssse3() ? 2 : 1;
known.store(state, std::memory_order_relaxed);
}
return state == 2;
}
#endif
#if NLOHMANN_VIEW_VECTOR_UTF8 #if NLOHMANN_VIEW_VECTOR_UTF8
/// Tables of the UTF-8 check of J. Keiser and D. Lemire, "Validating UTF-8 In /// Tables of the UTF-8 check of J. Keiser and D. Lemire, "Validating UTF-8 In
/// Less Than One Instruction Per Byte" (2021), as in simdjson ("lookup4"): each /// Less Than One Instruction Per Byte" (2021), as in simdjson ("lookup4"): each
@@ -192,8 +245,9 @@ compares, and the UTF-8 check covers the bytes up to it. Returns where the
string scan stops, like scan_string_run: before ill-formed UTF-8 and for the string scan stops, like scan_string_run: before ill-formed UTF-8 and for the
last bytes of the input, the bytes are checked one sequence at a time. Out of last bytes of the input, the bytes are checked one sequence at a time. Out of
line, so that no constants of the check occupy registers in the parse loop. line, so that no constants of the check occupy registers in the parse loop.
On x86-64, it is compiled for SSSE3 (see cpu_has_ssse3()).
*/ */
NLOHMANN_VIEW_NOINLINE inline const unsigned char* scan_string_vector(const unsigned char* p, const unsigned char* e, const std::uint8_t* plain) noexcept NLOHMANN_VIEW_SSSE3_TARGET NLOHMANN_VIEW_NOINLINE inline const unsigned char* scan_string_vector(const unsigned char* p, const unsigned char* e, const std::uint8_t* plain) noexcept
{ {
using lookup = utf8_lookup4<>; using lookup = utf8_lookup4<>;
const unsigned char* block = p; const unsigned char* block = p;
+105 -7
View File
@@ -25,7 +25,7 @@
#define INCLUDE_NLOHMANN_JSON_VIEW_HPP_ #define INCLUDE_NLOHMANN_JSON_VIEW_HPP_
#include <cstddef> // size_t #include <cstddef> // size_t
#include <cstdint> // uint32_t #include <cstdint> // uint8_t, uint32_t
#include <cstring> // memcpy, strlen #include <cstring> // memcpy, strlen
#include <iterator> // distance, input_iterator_tag, iterator_traits #include <iterator> // distance, input_iterator_tag, iterator_traits
#include <map> // map #include <map> // map
@@ -53,6 +53,7 @@
#include <nlohmann/detail/view/edit.hpp> #include <nlohmann/detail/view/edit.hpp>
#include <nlohmann/detail/view/edit_storage.hpp> #include <nlohmann/detail/view/edit_storage.hpp>
#include <nlohmann/detail/view/errors.hpp> #include <nlohmann/detail/view/errors.hpp>
#include <nlohmann/detail/view/image.hpp>
#include <nlohmann/detail/view/input.hpp> #include <nlohmann/detail/view/input.hpp>
#include <nlohmann/detail/view/iterator.hpp> #include <nlohmann/detail/view/iterator.hpp>
#include <nlohmann/detail/view/lookup.hpp> #include <nlohmann/detail/view/lookup.hpp>
@@ -591,8 +592,10 @@ class basic_json_view
style.indent_char = indent_char; style.indent_char = indent_char;
style.ensure_ascii = ensure_ascii; style.ensure_ascii = ensure_ascii;
style.source_numbers = numbers == number_format::source; style.source_numbers = numbers == number_format::source;
// the compact text is about as long as the source text of the value // the compact text is about as long as the source text of the value;
const std::size_t estimate = source_extent() + (style.pretty ? source_extent() / 2 : 0) + 64; // the compact writer keeps 64 bytes of slack, so that it does not grow
// the buffer just before the end
const std::size_t estimate = source_extent() + (style.pretty ? source_extent() / 2 : 0) + 160;
detail::view::view_serializer<BasicJsonType, Editable>(*m_doc, out, estimate, style).dump(m_node); detail::view::view_serializer<BasicJsonType, Editable>(*m_doc, out, estimate, style).dump(m_node);
return out; return out;
} }
@@ -934,7 +937,7 @@ class basic_json_document
/// whether the document holds its own copy of the text /// whether the document holds its own copy of the text
bool owns_source() const noexcept bool owns_source() const noexcept
{ {
return m_data && !m_data->owned.empty() && m_data->src == m_data->owned.data(); return m_data && ((!m_data->owned.empty() && m_data->src == m_data->owned.data()) || !m_data->owned_image.empty());
} }
/// number of index nodes (values plus object keys) /// number of index nodes (values plus object keys)
@@ -943,7 +946,7 @@ class basic_json_document
return m_data ? m_data->tape_size : 0; return m_data ? m_data->tape_size : 0;
} }
/// bytes held by the document (index, decoded strings, owned text) /// bytes held by the document (index, decoded strings, owned text or image)
std::size_t memory_usage() const noexcept std::size_t memory_usage() const noexcept
{ {
if (!m_data) if (!m_data)
@@ -952,7 +955,7 @@ class basic_json_document
} }
return sizeof(document_data) + (m_data->inline_cap * sizeof(detail::view::node)) return sizeof(document_data) + (m_data->inline_cap * sizeof(detail::view::node))
+ (m_data->tape != m_data->inline_tape ? m_data->tape_cap * sizeof(detail::view::node) : 0) + (m_data->tape != m_data->inline_tape ? m_data->tape_cap * sizeof(detail::view::node) : 0)
+ m_data->arena.capacity() + m_data->owned.capacity() + m_data->arena.capacity() + m_data->owned.capacity() + m_data->owned_image.capacity()
+ (m_data->indexes.capacity() * sizeof(document_data::object_index)) + (m_data->index_slots.capacity() * sizeof(std::uint32_t)) + (m_data->indexes.capacity() * sizeof(document_data::object_index)) + (m_data->index_slots.capacity() * sizeof(std::uint32_t))
+ (m_data->large_objects.capacity() * sizeof(std::uint32_t)) + (m_data->large_objects.capacity() * sizeof(std::uint32_t))
+ (m_data->edits != nullptr ? m_data->edits->bytes : 0); + (m_data->edits != nullptr ? m_data->edits->bytes : 0);
@@ -972,8 +975,10 @@ class basic_json_document
// allocate everything first, so that an exception leaves the document // allocate everything first, so that an exception leaves the document
// unchanged // unchanged
// (the decoded strings of a loaded image stay in the image)
const bool arena_in_use = d.base[1] == d.arena.data();
const bool shrink_arena = d.arena.capacity() > d.arena.size(); const bool shrink_arena = d.arena.capacity() > d.arena.size();
std::string arena(shrink_arena ? d.arena : std::string()); std::string arena(shrink_arena && arena_in_use ? d.arena : std::string());
// (edits link to the nodes of the index, which then stays in place) // (edits link to the nodes of the index, which then stays in place)
const bool shrink_tape = d.tape != d.inline_tape && d.tape_size != d.tape_cap && d.edits == nullptr; const bool shrink_tape = d.tape != d.inline_tape && d.tape_size != d.tape_cap && d.edits == nullptr;
const bool into_header = d.tape_size <= d.inline_cap; const bool into_header = d.tape_size <= d.inline_cap;
@@ -989,9 +994,61 @@ class basic_json_document
if (shrink_arena) if (shrink_arena)
{ {
d.arena.swap(arena); d.arena.swap(arena);
if (arena_in_use)
{
d.base[1] = d.arena.data(); d.base[1] = d.arena.data();
} }
} }
}
////////////
// images //
////////////
/// how load() checks an image (full, bounds, or none)
using image_check = detail::view::image_check;
/// The document as an image that load() reads without parsing: the node
/// index, the text, and the decoded strings. An edited document is
/// written in its current state (floats that are not finite become null,
/// as in dump()).
std::vector<std::uint8_t> save() const
{
if (NLOHMANN_VIEW_UNLIKELY(!m_data || m_data->discarded))
{
detail::view::throw_type_error(320, "cannot save a discarded json_document");
}
return detail::view::save_image(*m_data);
}
/// Read an image written by save(). The image is borrowed: it must stay
/// alive and unchanged while the document is used.
NLOHMANN_VIEW_NODISCARD
static basic_json_document load(const std::uint8_t* image, std::size_t size, const image_check check = image_check::full)
{
basic_json_document d;
d.ensure_data(nullptr, 0);
detail::view::load_image(*d.m_data, image, size, check);
return d;
}
/// read an image (borrowed)
NLOHMANN_VIEW_NODISCARD
static basic_json_document load(const std::vector<std::uint8_t>& image, const image_check check = image_check::full)
{
return load(image.data(), image.size(), check);
}
/// read an image and keep it (no copy)
NLOHMANN_VIEW_NODISCARD
static basic_json_document load(std::vector<std::uint8_t>&& image, const image_check check = image_check::full)
{
basic_json_document d;
d.ensure_data(nullptr, 0);
d.m_data->owned_image = std::move(image);
detail::view::load_image(*d.m_data, d.m_data->owned_image.data(), d.m_data->owned_image.size(), check);
return d;
}
/////////// ///////////
// edits // // edits //
@@ -1062,6 +1119,45 @@ class basic_json_document
return editor().push_back(array, std::forward<V>(value)); return editor().push_back(array, std::forward<V>(value));
} }
/// insert into an array before position idx (idx <= size()); returns a
/// view of the new element
template < typename I, typename V, typename std::enable_if < std::is_integral<I>::value && !std::is_same<I, bool>::value, int >::type = 0 >
view_type insert(view_type array, I idx, V && value)
{
return editor().insert(array, index(idx), std::forward<V>(value));
}
/// remove all members with this key; returns their number
std::size_t erase(view_type object, string_view_t key)
{
return editor().erase(object, key);
}
/// remove an array element
template < typename I, typename std::enable_if < std::is_integral<I>::value && !std::is_same<I, bool>::value, int >::type = 0 >
void erase(view_type array, I idx)
{
editor().erase(array, index(idx));
}
/// remove the value at a JSON pointer; returns the number of removed
/// values
std::size_t erase(const json_pointer& ptr)
{
if (ptr.empty())
{
detail::view::throw_out_of_range(405, "JSON pointer has no parent");
}
const view_type parent = root().at(ptr.parent_pointer());
const auto& token = ptr.back();
if (parent.is_array())
{
erase(parent, pointer_index(token));
return 1;
}
return erase(parent, string_view_t(token.data(), token.size()));
}
private: private:
using input_kind = detail::view::input_kind; using input_kind = detail::view::input_kind;
@@ -1137,6 +1233,7 @@ class basic_json_document
{ {
d.owned.clear(); d.owned.clear();
} }
d.owned_image.clear();
d.src = src; d.src = src;
d.size = size; d.size = size;
d.tape_size = 0; d.tape_size = 0;
@@ -1161,6 +1258,7 @@ class basic_json_document
{ {
d.base[0] = d.src; d.base[0] = d.src;
d.base[1] = d.arena.data(); d.base[1] = d.arena.data();
d.arena_size = d.arena.size();
detail::view::build_object_indexes(d); detail::view::build_object_indexes(d);
d.discarded = false; d.discarded = false;
return; return;
+769 -44
View File
@@ -8860,8 +8860,9 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc
} }
/// eight bytes as a little-endian word (compilers fold this into one load on /// eight bytes as a little-endian word (compilers fold this into one load on
/// little-endian targets) /// little-endian targets; always inlined, as GCC otherwise calls it in the
inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept /// number loops)
JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
{ {
return static_cast<std::uint64_t>(b[0]) | (static_cast<std::uint64_t>(b[1]) << 8u) return static_cast<std::uint64_t>(b[0]) | (static_cast<std::uint64_t>(b[1]) << 8u)
| (static_cast<std::uint64_t>(b[2]) << 16u) | (static_cast<std::uint64_t>(b[3]) << 24u) | (static_cast<std::uint64_t>(b[2]) << 16u) | (static_cast<std::uint64_t>(b[3]) << 24u)
@@ -8870,7 +8871,7 @@ inline std::uint64_t read_eight_bytes(const unsigned char* b) noexcept
} }
/// eight bytes as a little-endian word /// eight bytes as a little-endian word
inline std::uint64_t read_eight_bytes(const char* p) noexcept JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const char* p) noexcept
{ {
return read_eight_bytes(reinterpret_cast<const unsigned char*>(p)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast) return read_eight_bytes(reinterpret_cast<const unsigned char*>(p)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
} }
@@ -9479,8 +9480,9 @@ template<typename FloatType>
using native_float_t = typename std::conditional<std::numeric_limits<FloatType>::digits == 24, float, double>::type; using native_float_t = typename std::conditional<std::numeric_limits<FloatType>::digits == 24, float, double>::type;
/// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three /// the value of the eight ASCII digits in @a v (see read_eight_bytes()), three
/// multiplications instead of eight (after simdjson and fast_float) /// multiplications instead of eight (after simdjson and fast_float); always
inline std::uint32_t parse_eight_digits(std::uint64_t v) noexcept /// inlined, as GCC otherwise calls it in the number loops
JSON_HEDLEY_ALWAYS_INLINE std::uint32_t parse_eight_digits(std::uint64_t v) noexcept
{ {
v = ((v & 0x0F0F0F0F0F0F0F0Fu) * 2561u) >> 8u; v = ((v & 0x0F0F0F0F0F0F0F0Fu) * 2561u) >> 8u;
v = ((v & 0x00FF00FF00FF00FFu) * 6553601u) >> 16u; v = ((v & 0x00FF00FF00FF00FFu) * 6553601u) >> 16u;
@@ -24416,11 +24418,275 @@ NLOHMANN_JSON_NAMESPACE_END
#include <array> // array #include <array> // array
#include <cmath> // signbit, isfinite #include <cmath> // signbit, isfinite
#include <cstddef> // size_t
#include <cstdint> // intN_t, uintN_t #include <cstdint> // intN_t, uintN_t
#include <cstring> // memcpy, memmove #include <cstring> // memcpy, memmove
#include <limits> // numeric_limits #include <limits> // numeric_limits
#include <type_traits> // conditional #include <type_traits> // conditional
#ifdef _MSC_VER
#include <cstdlib> // _byteswap_uint64
#endif
// SSE2 (every x86-64 CPU) and NEON (every 64-bit Arm CPU) convert the 16
// digits of a double at once
#if defined(__x86_64__) || (defined(_M_X64) && !defined(_M_ARM64EC))
#include <emmintrin.h>
#define JSON_DTOA_SSE2 1
#define JSON_DTOA_NEON 0
#elif (defined(__aarch64__) || defined(_M_ARM64)) && !defined(_M_ARM64EC) && !defined(__ARM_BIG_ENDIAN)
#include <arm_neon.h>
#define JSON_DTOA_SSE2 0
#define JSON_DTOA_NEON 1
#else
#define JSON_DTOA_SSE2 0
#define JSON_DTOA_NEON 0
#endif
// #include <nlohmann/detail/conversions/zmij.hpp>
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2025 Victor Zverovich <https://github.com/vitaut/zmij>
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint32_t, uint64_t
// #include <nlohmann/detail/abi_macros.hpp>
// #include <nlohmann/detail/bit_ops.hpp>
// #include <nlohmann/detail/input/pow5_table.hpp>
// #include <nlohmann/detail/macro_scope.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief the shortest decimal representation of a double
A C++11 port of the conversion of Zmij by Victor Zverovich
(https://github.com/vitaut/zmij, MIT license): the shortest decimal in the
rounding interval of a double, the closest one if there are several. Zmij
credits Xiang JunBo (producing the shorter candidate without a division) and
Dougall Johnson (the compressed powers of ten). The powers of ten are taken
from the table for number parsing (pow5_table.hpp) where it holds them, and
computed from the compressed tables of Zmij beyond it.
*/
namespace zmij
{
/// significand * 10^exponent
struct decimal
{
std::uint64_t significand;
int exponent;
};
/// the compressed powers of ten of Zmij
inline const std::array<std::uint64_t, 28>& pow10_minor() noexcept
{
static const std::array<std::uint64_t, 28> table =
{
{
0x8000000000000000u, 0xa000000000000000u, 0xc800000000000000u, 0xfa00000000000000u, 0x9c40000000000000u,
0xc350000000000000u, 0xf424000000000000u, 0x9896800000000000u, 0xbebc200000000000u, 0xee6b280000000000u,
0x9502f90000000000u, 0xba43b74000000000u, 0xe8d4a51000000000u, 0x9184e72a00000000u, 0xb5e620f480000000u,
0xe35fa931a0000000u, 0x8e1bc9bf04000000u, 0xb1a2bc2ec5000000u, 0xde0b6b3a76400000u, 0x8ac7230489e80000u,
0xad78ebc5ac620000u, 0xd8d726b7177a8000u, 0x878678326eac9000u, 0xa968163f0a57b400u, 0xd3c21bcecceda100u,
0x84595161401484a0u, 0xa56fa5b99019a5c8u, 0xcecb8f27f4200f3au
}
};
return table;
}
/// (high, low) pairs
inline const std::array<std::uint64_t, 50>& pow10_major() noexcept
{
static const std::array<std::uint64_t, 50> table =
{
{
0xaddcb9e83c6b1793u, 0xdf4abe242a1bbf3eu, 0xaf8e5410288e1b6fu, 0x07ecf0ae5ee44ddau, 0xb1442798f49ffb4au, 0x99cd11cfdf41779du,
0xb2fe3f0b8599ef07u, 0x861fa7e6dcb4aa15u, 0xb4bca50b065abe63u, 0x0fed077a756b53aau, 0xb67f6455292cbf08u, 0x1a3bc84c17b1d543u,
0xb84687c269ef3bfbu, 0x3d5d514f40eea742u, 0xba121a4650e4ddebu, 0x92f34d62616ce413u, 0xbbe226efb628afeau, 0x890489f70a55368cu,
0xbdb6b8e905cb600fu, 0x5400e987bbc1c921u, 0xbf8fdb78849a5f96u, 0xde98520472bdd034u, 0xc16d9a0095928a27u, 0x75b7053c0f178294u,
0xc350000000000000u, 0x0000000000000000u, 0xc5371912364ce305u, 0x6c28000000000000u, 0xc722f0ef9d80aad6u, 0x424d3ad2b7b97ef6u,
0xc913936dd571c84cu, 0x03bc3a19cd1e38eau, 0xcb090c8001ab551cu, 0x5cadf5bfd3072cc6u, 0xcd036837130890a1u, 0x36dba887c37a8c10u,
0xcf02b2c21207ef2eu, 0x94f967e45e03f4bcu, 0xd106f86e69d785c7u, 0xe13336d701beba52u, 0xd31045a8341ca07cu, 0x1ede48111209a051u,
0xd51ea6fa85785631u, 0x552a74227f3ea566u, 0xd732290fbacaf133u, 0xa97c177947ad4096u, 0xd94ad8b1c7380874u, 0x18375281ae7822bdu,
0xdb68c2ca82ed2a05u, 0xa67398db9f6820e1u
}
};
return table;
}
/// one bit per power: whether the computed value is one unit too large
inline const std::array<std::uint32_t, 21>& pow10_fixups() noexcept
{
static const std::array<std::uint32_t, 21> table =
{
{
0x8d8fc810u, 0x06100293u, 0x19000000u, 0x00100000u, 0x00000908u, 0x00000000u, 0x04e00300u, 0x3807e0b2u, 0x3d83d793u, 0x0006f5ccu,
0x00000000u, 0xffff0000u, 0x8076337du, 0x4ff45ba0u, 0x09405033u, 0x034376d9u, 0x09000000u, 0x4e100501u, 0x076d14dcu, 0xf964f45eu,
0x0000003du
}
};
return table;
}
/// the 128-bit significand of 10^k, rounded down, for k in [-307, 341]
/// (compute_pow10 of Zmij)
inline uint128_parts compute_pow10(int k) noexcept
{
const auto i = static_cast<unsigned>(k + 307);
const std::uint64_t m = pow10_minor()[(i + 24) % 28];
const std::size_t j = 2 * static_cast<std::size_t>((i + 24) / 28);
const std::uint64_t h_hi = pow10_major()[j];
const std::uint64_t h_lo = pow10_major()[j + 1];
const std::uint64_t h1 = full_multiplication(h_lo, m).high;
const std::uint64_t c0 = h_lo * m;
const std::uint64_t c1 = h1 + (h_hi * m);
const std::uint64_t c2 = (c1 < h1 ? 1u : 0u) + full_multiplication(h_hi, m).high;
uint128_parts r{};
if ((c2 >> 63u) != 0)
{
r.high = c2;
r.low = c1;
}
else
{
r.high = (c2 << 1u) | (c1 >> 63u);
r.low = (c1 << 1u) | (c0 >> 63u);
}
r.low -= (pow10_fixups()[i >> 5u] >> (i & 31u)) & 1u;
return r;
}
/// The 128-bit significand of 10^k, rounded down, for k in [-342, 341].
/// Up to 10^308, the table for number parsing holds the same significands
/// (those of 5^k), except for k in [-27, -1], where it holds them one unit
/// larger (as the Eisel-Lemire algorithm needs them).
inline uint128_parts pow10(int k) noexcept
{
if (k > pow5_128_largest_power)
{
return compute_pow10(k); // (only for the smallest doubles)
}
const auto i = 2 * static_cast<std::size_t>(k - pow5_128_smallest_power);
uint128_parts r{pow5_128()[i + 1], pow5_128()[i]};
const std::uint64_t adjust = static_cast<unsigned>(k + 27) < 27u ? 1u : 0u;
r.high -= r.low < adjust ? 1u : 0u;
r.low -= adjust;
return r;
}
/// (x_hi * 2^64 + x_lo) * y >> 64, as 128 bits
inline uint128_parts umul192_hi128(std::uint64_t x_hi, std::uint64_t x_lo, std::uint64_t y) noexcept
{
const uint128_parts p = full_multiplication(x_hi, y);
uint128_parts r{};
r.low = p.low + full_multiplication(x_lo, y).high;
r.high = p.high + (r.low < p.low ? 1u : 0u);
return r;
}
/// (x * y + c) >> 64
inline std::uint64_t umul128_add_hi64(std::uint64_t x, std::uint64_t y, std::uint64_t c) noexcept
{
const uint128_parts p = full_multiplication(x, y);
return p.high + (p.low + c < p.low ? 1u : 0u);
}
/// the result of Zmij: the shorter candidate and, if that is outside the
/// rounding interval, the digit after it (16 bytes: returned in registers)
struct shortest_decimal
{
std::uint64_t integral; ///< the shorter candidate (15 or 16 digits for normal doubles)
int exponent; ///< the decimal exponent of the digit after it
unsigned char digit; ///< the digit after it (if has_digit)
bool has_digit; ///< whether the shortest decimal is integral * 10 + digit
};
/// The shortest decimal in the rounding interval of a positive finite double
/// given by its bits, the closest one if there are several (to_decimal of
/// Zmij, which keeps the last digit apart: the 15 or 16 digits before it can be
/// converted without a multiplication by 10 first). Always inlined: GCC
/// otherwise calls it, and its result goes through memory.
JSON_HEDLEY_ALWAYS_INLINE shortest_decimal to_shortest(std::uint64_t bits) noexcept
{
constexpr int extra_shift = 9;
const auto raw_exp = static_cast<int>((bits >> 52u) & 0x7FFu);
std::uint64_t bin_sig = bits & ((std::uint64_t{1} << 52u) - 1);
// a power of two has a narrower interval below (except the smallest normal)
const bool regular = bin_sig != 0 || raw_exp <= 1;
const int bin_exp = (raw_exp == 0 ? 1 : raw_exp) - 1075;
if (raw_exp != 0)
{
bin_sig |= std::uint64_t{1} << 52u;
}
// floor(log10(2^bin_exp)), or floor(log10(3/4 * 2^bin_exp)) for the irregular case
const int dec_exp = ((bin_exp * 315653) - (regular ? 0 : 131072)) >> 20;
// scaled by 10^(-dec_exp - 1): the integral part is the shorter candidate
const int shift = bin_exp + ((-(dec_exp + 1) * 217707) >> 16) + 1 + extra_shift;
const uint128_parts p10 = pow10(-dec_exp - 1);
const uint128_parts p = umul192_hi128(p10.high, p10.low, bin_sig << static_cast<unsigned>(shift));
std::uint64_t integral = p.high >> static_cast<unsigned>(extra_shift);
const std::uint64_t fractional = (p.high << static_cast<unsigned>(64 - extra_shift)) | (p.low >> static_cast<unsigned>(extra_shift));
std::uint64_t digit = 0;
bool round_up = false;
bool round_down = false;
if (JSON_HEDLEY_LIKELY(regular))
{
const std::uint64_t half_ulp = (p10.high >> static_cast<unsigned>(extra_shift + 1 - shift)) + (1 - (bin_sig & 1u));
round_up = fractional + half_ulp < fractional;
round_down = half_ulp > fractional;
// the last digit of the longer candidate, rounded to nearest
digit = umul128_add_hi64(fractional, 10, (std::uint64_t{1} << 63u) + 6);
if (fractional == (std::uint64_t{1} << 62u))
{
digit = 2; // 2.5 rounds to 2
}
}
else
{
const std::uint64_t half_ulp = p10.high >> static_cast<unsigned>(extra_shift + 1 - shift);
round_up = half_ulp > ~std::uint64_t{0} - fractional;
round_down = (half_ulp >> 1u) > fractional;
digit = umul128_add_hi64(fractional, 10, (std::uint64_t{1} << 63u) - 1);
const std::uint64_t lowest = umul128_add_hi64(fractional - (half_ulp >> 1u), 10, ~std::uint64_t{0});
digit = digit < lowest ? lowest : digit;
}
integral += round_up ? 1u : 0u;
// if the shorter candidate is outside the rounding interval: one digit more
return shortest_decimal{integral, dec_exp, static_cast<unsigned char>(digit), !round_up && !round_down};
}
/// The shortest decimal in the rounding interval of a positive finite double
/// given by its bits, as one number. The significand can end in zeros.
inline decimal to_decimal(std::uint64_t bits) noexcept
{
const shortest_decimal d = to_shortest(bits);
if (d.has_digit)
{
return decimal{(d.integral * 10) + d.digit, d.exponent};
}
return decimal{d.integral, d.exponent + 1};
}
} // namespace zmij
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
// #include <nlohmann/detail/macro_scope.hpp> // #include <nlohmann/detail/macro_scope.hpp>
@@ -25324,6 +25590,88 @@ void grisu2(char* buf, int& len, int& decimal_exponent, FloatType value)
grisu2(buf, len, decimal_exponent, w.minus, w.w, w.plus); grisu2(buf, len, decimal_exponent, w.minus, w.w, w.plus);
} }
/*!
@brief the shortest digits of a positive finite float (other than double): Grisu2
*/
template<typename FloatType>
JSON_HEDLEY_NON_NULL(1)
void shortest_digits(char* buf, int& len, int& decimal_exponent, FloatType value)
{
grisu2(buf, len, decimal_exponent, value);
}
/*!
@brief the shortest digits of a positive finite double: the conversion of
Zmij (see zmij.hpp), which always finds the shortest digits that read back as
the same value (Grisu2 does not for about one double in a thousand), and the
closest of them if there are several
v = buf * 10^decimal_exponent, as for grisu2()
*/
JSON_HEDLEY_NON_NULL(1)
inline void shortest_digits(char* buf, int& len, int& decimal_exponent, double value)
{
static_assert(std::numeric_limits<double>::is_iec559 && std::numeric_limits<double>::digits == 53,
"internal error: the conversion of Zmij needs IEEE 754 binary64 doubles");
JSON_ASSERT(std::isfinite(value));
JSON_ASSERT(value > 0);
std::uint64_t bits = 0;
std::memcpy(&bits, &value, sizeof(bits));
zmij::decimal d = zmij::to_decimal(bits);
// without trailing zeros (up to 16): 8, 4, 2, 1 at a time
while (d.significand % 100000000 == 0)
{
d.significand /= 100000000;
d.exponent += 8;
}
if (d.significand % 10000 == 0)
{
d.significand /= 10000;
d.exponent += 4;
}
if (d.significand % 100 == 0)
{
d.significand /= 100;
d.exponent += 2;
}
if (d.significand % 10 == 0)
{
d.significand /= 10;
d.exponent += 1;
}
// at most 17 digits, written from the back two at a time
static constexpr const char* pairs =
"00010203040506070809101112131415161718192021222324252627282930313233343536373839"
"40414243444546474849505152535455565758596061626364656667686970717273747576777879"
"8081828384858687888990919293949596979899";
std::array<char, 20> digits{};
std::size_t n = digits.size();
while (d.significand >= 100)
{
const std::uint64_t two_digits = d.significand % 100; // a variable: GCC calls a cast of the remainder useless where std::uint64_t is std::size_t
const auto i = static_cast<std::size_t>(two_digits) * 2;
d.significand /= 100;
n -= 2;
digits[n] = pairs[i];
digits[n + 1] = pairs[i + 1];
}
if (d.significand >= 10)
{
const auto i = static_cast<std::size_t>(d.significand) * 2;
n -= 2;
digits[n] = pairs[i];
digits[n + 1] = pairs[i + 1];
}
else
{
digits[--n] = static_cast<char>('0' + d.significand);
}
len = static_cast<int>(digits.size() - n);
std::memcpy(buf, digits.data() + n, static_cast<std::size_t>(len));
decimal_exponent = d.exponent;
}
/*! /*!
@brief appends a decimal representation of e to buf @brief appends a decimal representation of e to buf
@return a pointer to the element following the exponent. @return a pointer to the element following the exponent.
@@ -25453,6 +25801,374 @@ inline char* format_buffer(char* buf, int len, int decimal_exponent,
return append_exponent(buf, n - 1); return append_exponent(buf, n - 1);
} }
/// eight decimal digits (a value below 10^8) as bytes 0..9, the first digit
/// in the most significant byte: three steps that divide all lanes at once
/// by a multiplication (the conversion of Xiang JunBo, as in Zmij)
inline std::uint64_t eight_digit_bytes(std::uint64_t abcdefgh) noexcept
{
const std::uint64_t abcd_efgh = abcdefgh + (((std::uint64_t{1} << 32u) - 10000u) * ((abcdefgh * (((std::uint64_t{1} << 40u) / 10000u) + 1u)) >> 40u));
const std::uint64_t ab_cd_ef_gh = abcd_efgh + (((std::uint64_t{1} << 16u) - 100u) * (((abcd_efgh * (((std::uint64_t{1} << 19u) / 100u) + 1u)) >> 19u) & 0x7F0000007Fu));
return ab_cd_ef_gh + (((std::uint64_t{1} << 8u) - 10u) * (((ab_cd_ef_gh * (((std::uint64_t{1} << 10u) / 10u) + 1u)) >> 10u) & 0x000F000F000F000Fu));
}
/// store the bytes of v, the most significant one first (one byte swap and
/// one store where the byte order is known: compilers do not reliably merge
/// the byte stores once this is inlined)
inline void store_msb_first(char* p, std::uint64_t v) noexcept
{
#if defined(__BYTE_ORDER__) && defined(__ORDER_LITTLE_ENDIAN__) && __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__
v = __builtin_bswap64(v);
std::memcpy(p, &v, sizeof(v));
#elif defined(__BYTE_ORDER__) && defined(__ORDER_BIG_ENDIAN__) && __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
std::memcpy(p, &v, sizeof(v));
#elif defined(_MSC_VER) // (little-endian on all its targets)
v = _byteswap_uint64(v);
std::memcpy(p, &v, sizeof(v));
#else
for (unsigned i = 0; i < 8; ++i)
{
p[i] = static_cast<char>(v >> (56u - (8u * i)));
}
#endif
}
/*!
@brief digits * 10^exp for a double, in the layout of format_buffer()
The layout is that of format_buffer() with min_exp -4 and max_exp 15 (the
digits10 of double). The digits are converted eight at a time and placed
with fixed-size moves instead of per-digit loops and moves of the buffer.
@param[in] digits the digits (not 0, at most 17 digits; trailing zeros allowed)
@param[in] exp the decimal exponent of the last digit
@return a pointer past the text; up to 41 bytes at @a first are written
(some beyond the returned end)
*/
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_decimal(char* first, std::uint64_t digits, int exp) noexcept
{
JSON_ASSERT(digits != 0 && digits < 100000000000000000u);
const std::uint64_t upper = digits / 100000000u;
const std::uint64_t b0 = upper / 100000000u; // (one digit: it is its own byte)
const std::uint64_t b1 = eight_digit_bytes(upper % 100000000u);
const std::uint64_t b2 = eight_digit_bytes(digits % 100000000u);
// leading and trailing zero digits: zero bytes, counted without division
int leading = 16;
int zeros = 16;
if (b0 != 0)
{
leading = count_leading_zeros(b0) / 8;
}
else if (b1 != 0)
{
leading = 8 + (count_leading_zeros(b1) / 8);
}
else
{
leading += count_leading_zeros(b2) / 8;
}
if (b2 != 0)
{
zeros = count_trailing_zeros(b2) / 8;
}
else if (b1 != 0)
{
zeros = 8 + (count_trailing_zeros(b1) / 8);
}
// (else: 16, b0 is the one digit that is not 0)
// the digits as text at text + leading, then '0's, so that fixed-size
// moves need not check how many digits there are
std::array<char, 64> text; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
store_msb_first(text.data(), b0 + 0x3030303030303030u);
store_msb_first(text.data() + 8, b1 + 0x3030303030303030u);
store_msb_first(text.data() + 16, b2 + 0x3030303030303030u);
std::memset(text.data() + 24, '0', 40);
const int k = 24 - leading - zeros; // significant digits
const int n = k + exp + zeros; // position of the decimal point after the first digit
const char* const s0 = text.data() + leading;
if (-4 < n && n <= 15)
{
// "0.[000]digits" (n <= 0) is the digits after 1 - n leading '0's
// with the point after the first; "digits[000].0" (n >= k) and
// "dig.its" put the point after n characters
const int pad = n <= 0 ? 1 - n : 0;
const char* const s = s0 - pad;
const int len = k + pad;
const int point = n + pad;
std::memcpy(first, s, 16);
std::memcpy(first + point + 1, s + point, 24);
first[point] = '.';
return first + (point >= len ? point + 2 : len + 1);
}
// d.igitse+XX, with at least two exponent digits (as append_exponent())
std::memcpy(first, s0, 16);
std::memcpy(first + 2, s0 + 1, 16);
first[1] = '.';
char* const end = first + (k == 1 ? 1 : k + 1);
const int e = n - 1;
const auto ea = static_cast<unsigned>(e < 0 ? -e : e);
const bool three = ea >= 100;
end[0] = 'e';
end[1] = e < 0 ? '-' : '+';
end[2] = static_cast<char>('0' + (three ? ea / 100 : (ea / 10) % 10));
end[3] = static_cast<char>('0' + (three ? (ea / 10) % 10 : ea % 10));
end[4] = static_cast<char>('0' + (ea % 10));
return end + (three ? 5 : 4);
}
/*!
@brief the shortest decimal of a positive double (Zmij), as write_decimal()
writes it
For a normal double, the shorter candidate has 15 or 16 digits: they are
converted at once (two halves of eight digits) and followed by the digit
after them, if there is one, without the multiplication and division by 10
that counting the digits of one number would take. The fixed layouts move
the digits after the point by one byte.
@return a pointer past the text; up to 41 bytes at @a first are written
(some beyond the returned end)
*/
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_shortest(char* first, const zmij::shortest_decimal d) noexcept
{
const std::uint64_t sig = d.integral;
if (JSON_HEDLEY_UNLIKELY(sig < 100000000000000u || sig >= 10000000000000000u))
{
// (subnormals)
return d.has_digit ? write_decimal(first, (sig * 10) + d.digit, d.exponent) : write_decimal(first, sig, d.exponent + 1);
}
const bool sixteen = sig >= 1000000000000000u; // (else 15 digits)
const int last = d.has_digit ? d.digit : 0;
const std::uint64_t upper = sig / 100000000u;
#if JSON_DTOA_SSE2
// NOLINTBEGIN(portability-simd-intrinsics)
// the two halves in the 64-bit lanes, each as abcd * 2^32 + efgh, then as
// bytes (as eight_digit_bytes(), one lane each)
const __m128i x = _mm_set_epi64x(static_cast<long long>(sig - (upper * 100000000u)), static_cast<long long>(upper));
const __m128i abcd = _mm_srli_epi64(_mm_mul_epu32(x, _mm_set1_epi64x(109951163)), 40); // 2^40 / 10000 + 1
const __m128i abcd_efgh = _mm_add_epi64(x, _mm_mul_epu32(abcd, _mm_set1_epi64x(4294957296))); // 2^32 - 10000
// 32-bit lanes in the order of the text: abcd, efgh of both halves
const __m128i fours = _mm_shuffle_epi32(abcd_efgh, _MM_SHUFFLE(2, 3, 0, 1));
const __m128i ab = _mm_srli_epi16(_mm_mulhi_epu16(fours, _mm_set1_epi32(5243)), 3);
const __m128i ab_cd = _mm_or_si128(_mm_slli_epi32(_mm_sub_epi16(fours, _mm_mullo_epi16(ab, _mm_set1_epi32(100))), 16), ab);
// 16-bit lanes ab (< 100) -> bytes a, b: 256 * ab - 2559 * (ab / 10)
const __m128i bytes = _mm_sub_epi16(_mm_slli_epi16(ab_cd, 8), _mm_mullo_epi16(_mm_set1_epi16(2559), _mm_mulhi_epu16(ab_cd, _mm_set1_epi16(6554))));
// the last digit that is not 0 (sig is not 0)
const auto nonzero = static_cast<std::uint64_t>(_mm_movemask_epi8(_mm_cmpgt_epi8(bytes, _mm_setzero_si128())));
const int digits = 63 - count_leading_zeros(nonzero) + (sixteen ? 1 : 0); // without trailing zeros
const __m128i chars = _mm_add_epi8(bytes, _mm_set1_epi8('0'));
// the 16 characters from the first digit
const __m128i s = sixteen ? chars : _mm_or_si128(_mm_srli_si128(chars, 1), _mm_slli_si128(_mm_cvtsi32_si128('0' + last), 15));
const char s16 = static_cast<char>(sixteen ? '0' + last : '0'); // the 17th
const auto store_16 = [&s](char* p) noexcept
{
std::memcpy(p, &s, 16);
};
const char first_digit = static_cast<char>(_mm_cvtsi128_si32(s));
// NOLINTEND(portability-simd-intrinsics)
#elif JSON_DTOA_NEON
// as with SSE2: the halves in 32-bit lanes, then abcd, efgh of both
const uint32x2_t halves = vcreate_u32(upper | ((sig - (upper * 100000000u)) << 32u));
const uint32x2_t abcd = vmovn_u64(vshrq_n_u64(vmull_n_u32(halves, static_cast<std::uint32_t>(((std::uint64_t{1} << 40u) / 10000u) + 1u)), 40));
const uint32x2_t efgh = vmls_n_u32(halves, abcd, 10000u);
const uint32x4_t fours = vcombine_u32(vzip1_u32(abcd, efgh), vzip2_u32(abcd, efgh));
const uint32x4_t ab = vshrq_n_u32(vmulq_n_u32(fours, 5243u), 19);
const uint16x8_t ab_cd = vreinterpretq_u16_u32(vorrq_u32(ab, vshlq_n_u32(vmlsq_n_u32(fours, ab, 100u), 16)));
const uint16x8_t tens = vshrq_n_u16(vmulq_n_u16(ab_cd, 103u), 10);
const uint8x16_t bytes = vreinterpretq_u8_u16(vorrq_u16(tens, vshlq_n_u16(vmlsq_n_u16(ab_cd, tens, 10u), 8)));
// the last digit that is not 0 (sig is not 0): a nibble per byte
const std::uint64_t nonzero = vget_lane_u64(vreinterpret_u64_u8(vshrn_n_u16(vreinterpretq_u16_u8(vtstq_u8(bytes, bytes)), 4)), 0);
const int digits = ((63 - count_leading_zeros(nonzero)) / 4) + (sixteen ? 1 : 0); // without trailing zeros
const uint8x16_t chars = vaddq_u8(bytes, vdupq_n_u8('0'));
// the 16 characters from the first digit
const uint8x16_t s = sixteen ? chars : vextq_u8(chars, vdupq_n_u8(static_cast<std::uint8_t>('0' + last)), 1);
const char s16 = static_cast<char>(sixteen ? '0' + last : '0'); // the 17th
const auto store_16 = [&s](char* p) noexcept
{
vst1q_u8(reinterpret_cast<std::uint8_t*>(p), s); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
};
const auto first_digit = static_cast<char>(vgetq_lane_u8(s, 0));
#else
const std::uint64_t hi = eight_digit_bytes(upper);
const std::uint64_t lo = eight_digit_bytes(sig - (upper * 100000000u));
// trailing zero digits: zero bytes (sig is not 0)
const int zeros = lo != 0 ? count_trailing_zeros(lo) / 8 : 8 + (count_trailing_zeros(hi) / 8);
const int digits = 15 - zeros + (sixteen ? 1 : 0); // without trailing zeros
// the 16 characters from the first digit
const std::uint64_t s_hi = (sixteen ? hi : (hi << 8u) | (lo >> 56u)) + 0x3030303030303030u;
const std::uint64_t s_lo = (sixteen ? lo : (lo << 8u) | static_cast<std::uint64_t>(last)) + 0x3030303030303030u;
const char s16 = static_cast<char>(sixteen ? '0' + last : '0'); // the 17th
const auto store_16 = [s_hi, s_lo](char* p) noexcept
{
store_msb_first(p, s_hi);
store_msb_first(p + 8, s_lo);
};
const auto first_digit = static_cast<char>(s_hi >> 56u);
#endif
const int len = d.has_digit ? 16 + (sixteen ? 1 : 0) : digits; // significant digits
const int n = 16 + (sixteen ? 1 : 0) + d.exponent; // digits before the point
if (JSON_HEDLEY_LIKELY(n >= 1 && n <= 15))
{
// "dig.its" and "digits[000].0": the digits after the point move by
// one byte ('0's follow the digits)
#if JSON_DTOA_SSE2
// NOLINTBEGIN(portability-simd-intrinsics)
// (in the register: reading the digits back from memory right after
// storing them waits until the stores are done)
const __m128i at = _mm_set1_epi8(static_cast<char>(n));
const __m128i index = _mm_setr_epi8(0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15);
const __m128i before = _mm_cmpgt_epi8(at, index);
const __m128i after = _mm_cmpgt_epi8(index, at);
const __m128i text = _mm_or_si128(_mm_or_si128(_mm_and_si128(s, before), _mm_and_si128(_mm_slli_si128(s, 1), after)),
_mm_andnot_si128(_mm_or_si128(before, after), _mm_set1_epi8('.')));
std::memcpy(first, &text, 16);
first[16] = static_cast<char>(_mm_extract_epi16(s, 7) >> 8);
first[17] = s16;
// NOLINTEND(portability-simd-intrinsics)
#elif JSON_DTOA_NEON
const uint8x16_t index = vcombine_u8(vcreate_u8(0x0706050403020100u), vcreate_u8(0x0F0E0D0C0B0A0908u));
const uint8x16_t at = vdupq_n_u8(static_cast<std::uint8_t>(n));
const uint8x16_t after_point = vbslq_u8(vcgtq_u8(index, at), vextq_u8(vdupq_n_u8(0), s, 15), vdupq_n_u8('.'));
vst1q_u8(reinterpret_cast<std::uint8_t*>(first), vbslq_u8(vcltq_u8(index, at), s, after_point)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
first[16] = static_cast<char>(vgetq_lane_u8(s, 15));
first[17] = s16;
#else
store_16(first);
first[16] = s16;
std::uint64_t after_point[2]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays,cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
std::memcpy(after_point, first + n, 16);
std::memcpy(first + n + 1, after_point, 16);
first[n] = '.';
#endif
return first + (n >= len ? n + 2 : len + 1);
}
if (n <= 0 && n > -4)
{
// "0.[000]digits"
std::memset(first, '0', 8);
first[1] = '.';
store_16(first + 2 - n);
first[18 - n] = s16;
return first + 2 - n + len;
}
// d.igitse+XX, with at least two exponent digits (as append_exponent())
store_16(first + 1);
first[17] = s16;
first[0] = first_digit;
first[1] = '.';
char* const end = first + (len == 1 ? 1 : len + 1);
const int e = n - 1;
const auto ea = static_cast<unsigned>(e < 0 ? -e : e);
const bool three = ea >= 100;
end[0] = 'e';
end[1] = e < 0 ? '-' : '+';
end[2] = static_cast<char>('0' + (three ? ea / 100 : (ea / 10) % 10));
end[3] = static_cast<char>('0' + (three ? (ea / 10) % 10 : ea % 10));
end[4] = static_cast<char>('0' + (ea % 10));
return end + (three ? 5 : 4);
}
/// the powers of ten up to 10^16
inline const std::array<std::uint64_t, 17>& powers_of_ten_16() noexcept
{
static const std::array<std::uint64_t, 17> powers =
{
{
1u, 10u, 100u, 1000u, 10000u, 100000u, 1000000u, 10000000u, 100000000u, 1000000000u, 10000000000u,
100000000000u, 1000000000000u, 10000000000000u, 100000000000000u, 1000000000000000u, 10000000000000000u
}
};
return powers;
}
/*!
@brief digits * 10^exp, as write_decimal() writes it, for the digits of a
double that need no conversion (count digits, at most 15, the first not 0;
trailing zeros allowed): extended to 16 digits and written by write_shortest()
@return a pointer past the text; up to 41 bytes at @a first are written
(some beyond the returned end)
*/
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_short_decimal(char* first, std::uint64_t digits, int count, int exp) noexcept
{
JSON_ASSERT(digits >= powers_of_ten_16()[static_cast<std::size_t>(count - 1)] && count <= 15);
const int scale = 16 - count;
return write_shortest(first, zmij::shortest_decimal{digits * powers_of_ten_16()[static_cast<std::size_t>(scale)], exp - scale - 1, 0, false});
}
/// as write_short_decimal(), counting the digits (not 0, less than 10^15)
JSON_HEDLEY_NON_NULL(1)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_short_decimal(char* first, std::uint64_t digits, int exp) noexcept
{
JSON_ASSERT(digits != 0 && digits < 1000000000000000u);
// floor(log10(2^bits)) + 1 digits, or one less
const int log2_bound = ((64 - count_leading_zeros(digits)) * 1233) >> 12;
const int count = log2_bound + (digits >= powers_of_ten_16()[static_cast<std::size_t>(log2_bound)] ? 1 : 0);
return write_short_decimal(first, digits, count, exp);
}
/// a positive finite float (other than double): Grisu2 and format_buffer()
template<typename FloatType>
JSON_HEDLEY_NON_NULL(1, 2)
JSON_HEDLEY_RETURNS_NON_NULL
char* write_positive(char* first, const char* last, FloatType value)
{
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10);
// Compute v = buffer * 10^decimal_exponent.
// The decimal digits are stored in the buffer, which needs to be interpreted
// as an unsigned decimal integer.
// len is the length of the buffer, i.e., the number of decimal digits.
int len = 0;
int decimal_exponent = 0;
shortest_digits(first, len, decimal_exponent, value);
JSON_ASSERT(len <= std::numeric_limits<FloatType>::max_digits10);
// Format the buffer like printf("%.*g", prec, value)
constexpr int kMinExp = -4;
// Use digits10 here to increase compatibility with version 2.
constexpr int kMaxExp = std::numeric_limits<FloatType>::digits10;
JSON_ASSERT(last - first >= kMaxExp + 2);
JSON_ASSERT(last - first >= 2 + (-kMinExp - 1) + std::numeric_limits<FloatType>::max_digits10);
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10 + 6);
return format_buffer(first, len, decimal_exponent, kMinExp, kMaxExp);
}
/// a positive finite double: the shortest digits (Zmij), laid out by
/// write_shortest() (through a local buffer if [first, last) is shorter than
/// the 41 bytes it may write)
JSON_HEDLEY_NON_NULL(1, 2)
JSON_HEDLEY_RETURNS_NON_NULL
inline char* write_positive(char* first, const char* last, double value)
{
static_assert(std::numeric_limits<double>::is_iec559 && std::numeric_limits<double>::digits == 53,
"internal error: the conversion of Zmij needs IEEE 754 binary64 doubles");
std::uint64_t bits = 0;
std::memcpy(&bits, &value, sizeof(bits));
const zmij::shortest_decimal d = zmij::to_shortest(bits);
if (JSON_HEDLEY_LIKELY(last - first >= 41))
{
return write_shortest(first, d);
}
std::array<char, 64> buf; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
const auto len = static_cast<std::size_t>(write_shortest(buf.data(), d) - buf.data());
JSON_ASSERT(static_cast<std::size_t>(last - first) >= len);
std::memcpy(first, buf.data(), len);
return first + len;
}
} // namespace dtoa_impl } // namespace dtoa_impl
/*! /*!
@@ -25470,7 +26186,6 @@ JSON_HEDLEY_NON_NULL(1, 2)
JSON_HEDLEY_RETURNS_NON_NULL JSON_HEDLEY_RETURNS_NON_NULL
char* to_chars(char* first, const char* last, FloatType value) char* to_chars(char* first, const char* last, FloatType value)
{ {
static_cast<void>(last); // maybe unused - fix warning
JSON_ASSERT(std::isfinite(value)); JSON_ASSERT(std::isfinite(value));
// Use signbit(value) instead of (value < 0) since signbit works for -0. // Use signbit(value) instead of (value < 0) since signbit works for -0.
@@ -25496,28 +26211,7 @@ char* to_chars(char* first, const char* last, FloatType value)
JSON_HEDLEY_DIAGNOSTIC_POP JSON_HEDLEY_DIAGNOSTIC_POP
#endif #endif
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10); return dtoa_impl::write_positive(first, last, value);
// Compute v = buffer * 10^decimal_exponent.
// The decimal digits are stored in the buffer, which needs to be interpreted
// as an unsigned decimal integer.
// len is the length of the buffer, i.e., the number of decimal digits.
int len = 0;
int decimal_exponent = 0;
dtoa_impl::grisu2(first, len, decimal_exponent, value);
JSON_ASSERT(len <= std::numeric_limits<FloatType>::max_digits10);
// Format the buffer like printf("%.*g", prec, value)
constexpr int kMinExp = -4;
// Use digits10 here to increase compatibility with version 2.
constexpr int kMaxExp = std::numeric_limits<FloatType>::digits10;
JSON_ASSERT(last - first >= kMaxExp + 2);
JSON_ASSERT(last - first >= 2 + (-kMinExp - 1) + std::numeric_limits<FloatType>::max_digits10);
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10 + 6);
return dtoa_impl::format_buffer(first, len, decimal_exponent, kMinExp, kMaxExp);
} }
} // namespace detail } // namespace detail
@@ -26876,8 +27570,9 @@ class serializer
/*! /*!
@brief dump an integer @brief dump an integer
Dump a given integer, appending it to @ref write_buffer. Works internally with Dump a given integer, appending it to @ref write_buffer (directly: copying
@a number_buffer. the digits from another buffer right after writing them waits until the
stores are done).
@param[in] x integer number (signed or unsigned) to dump @param[in] x integer number (signed or unsigned) to dump
@tparam NumberType either @a number_integer_t or @a number_unsigned_t @tparam NumberType either @a number_integer_t or @a number_unsigned_t
@@ -26912,33 +27607,57 @@ class serializer
return; return;
} }
// use a pointer to fill the buffer // use a pointer to fill the buffer (room for as much as number_buffer holds)
auto buffer_ptr = number_buffer.begin(); // NOLINT(llvm-qualified-auto,readability-qualified-auto) if (JSON_HEDLEY_UNLIKELY(write_buffer_pos + number_buffer.size() > write_buffer.size()))
{
flush();
}
auto* buffer_ptr = write_buffer.data() + write_buffer_pos;
number_unsigned_t abs_value; number_unsigned_t abs_value;
unsigned int n_chars{}; // one byte for the minus sign
unsigned int n_chars = 0;
if (is_negative_number(x)) if (is_negative_number(x))
{ {
*buffer_ptr = '-'; *buffer_ptr = '-';
abs_value = remove_sign(static_cast<number_integer_t>(x)); abs_value = remove_sign(static_cast<number_integer_t>(x));
n_chars = 1;
// account one more byte for the minus sign
n_chars = 1 + count_digits(abs_value);
} }
else else
{ {
abs_value = static_cast<number_unsigned_t>(x); abs_value = static_cast<number_unsigned_t>(x);
n_chars = count_digits(abs_value);
} }
// up to 16 digits: eight at a time (as the digits of floats), written
// without leading zeros
if (abs_value < 10000000000000000u)
{
const std::uint64_t value = abs_value;
const std::uint64_t upper = value / 100000000u;
const std::uint64_t first = dtoa_impl::eight_digit_bytes(upper != 0 ? upper : value);
const auto leading = static_cast<unsigned>(count_leading_zeros(first) / 8); // (first is not 0)
char* const p = buffer_ptr + n_chars;
dtoa_impl::store_msb_first(p, (first << (8 * leading)) + 0x3030303030303030u);
n_chars += 8 - leading;
if (upper != 0)
{
dtoa_impl::store_msb_first(p + 8 - leading, dtoa_impl::eight_digit_bytes(value - (upper * 100000000u)) + 0x3030303030303030u);
n_chars += 8;
}
write_buffer_pos += n_chars;
return;
}
n_chars += count_digits(abs_value);
// spare 1 byte for '\0' // spare 1 byte for '\0'
JSON_ASSERT(n_chars < number_buffer.size() - 1); JSON_ASSERT(n_chars < number_buffer.size() - 1);
// jump to the end to generate the string from backward, // jump to the end to generate the string from backward,
// so we later avoid reversing the result // so we later avoid reversing the result
buffer_ptr += static_cast<typename decltype(number_buffer)::difference_type>(n_chars); buffer_ptr += n_chars;
// Fast int2ascii implementation inspired by "Fastware" talk by Andrei Alexandrescu // Fast int2ascii implementation inspired by "Fastware" talk by Andrei Alexandrescu
// See: https://www.youtube.com/watch?v=o4-CwDo2zpg // See: https://www.youtube.com/watch?v=o4-CwDo2zpg
@@ -26961,14 +27680,13 @@ class serializer
*(--buffer_ptr) = static_cast<char>('0' + abs_value); *(--buffer_ptr) = static_cast<char>('0' + abs_value);
} }
put_buffer(number_buffer, n_chars); write_buffer_pos += n_chars;
} }
/*! /*!
@brief dump a floating-point number @brief dump a floating-point number
Dump a given floating-point number, appending it to @ref write_buffer. Works internally Dump a given floating-point number, appending it to @ref write_buffer.
with @a number_buffer.
@param[in] x floating-point number to dump @param[in] x floating-point number to dump
*/ */
@@ -26995,10 +27713,15 @@ class serializer
void dump_float(number_float_t x, std::true_type /*is_ieee_single_or_double*/) void dump_float(number_float_t x, std::true_type /*is_ieee_single_or_double*/)
{ {
auto* begin = number_buffer.data(); // directly into the write buffer: copying the text from number_buffer
// right after to_chars() wrote it waits until its stores are done
if (JSON_HEDLEY_UNLIKELY(write_buffer_pos + number_buffer.size() > write_buffer.size()))
{
flush();
}
auto* begin = write_buffer.data() + write_buffer_pos;
auto* end = ::nlohmann::detail::to_chars(begin, begin + number_buffer.size(), x); auto* end = ::nlohmann::detail::to_chars(begin, begin + number_buffer.size(), x);
write_buffer_pos += static_cast<std::size_t>(end - begin);
put_buffer(number_buffer, static_cast<std::size_t>(end - begin));
} }
JSON_HEDLEY_NON_NULL(1) JSON_HEDLEY_NON_NULL(1)
@@ -34681,6 +35404,8 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
#undef JSON_NO_UNIQUE_ADDRESS #undef JSON_NO_UNIQUE_ADDRESS
#undef JSON_DISABLE_ENUM_SERIALIZATION #undef JSON_DISABLE_ENUM_SERIALIZATION
#undef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION #undef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
#undef JSON_DTOA_SSE2
#undef JSON_DTOA_NEON
#ifndef JSON_TEST_KEEP_MACROS #ifndef JSON_TEST_KEEP_MACROS
#undef JSON_CATCH #undef JSON_CATCH
File diff suppressed because it is too large Load Diff
+4 -1
View File
@@ -10,7 +10,7 @@ CXXFLAGS += -std=c++11
CPPFLAGS += -I ../single_include CPPFLAGS += -I ../single_include
FUZZER_ENGINE = src/fuzzer-driver_afl.cpp FUZZER_ENGINE = src/fuzzer-driver_afl.cpp
FUZZERS = parse_afl_fuzzer parse_bson_fuzzer parse_cbor_fuzzer parse_msgpack_fuzzer parse_ubjson_fuzzer parse_bjdata_fuzzer parse_bon8_fuzzer parse_json_view_fuzzer FUZZERS = parse_afl_fuzzer parse_bson_fuzzer parse_cbor_fuzzer parse_msgpack_fuzzer parse_ubjson_fuzzer parse_bjdata_fuzzer parse_bon8_fuzzer parse_json_view_fuzzer json_view_image_fuzzer
fuzzers: $(FUZZERS) fuzzers: $(FUZZERS)
parse_afl_fuzzer: parse_afl_fuzzer:
@@ -19,6 +19,9 @@ parse_afl_fuzzer:
parse_json_view_fuzzer: parse_json_view_fuzzer:
$(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-parse_json_view.cpp -o $@ $(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-parse_json_view.cpp -o $@
json_view_image_fuzzer:
$(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-json_view_image.cpp -o $@
parse_bson_fuzzer: parse_bson_fuzzer:
$(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-parse_bson.cpp -o $@ $(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-parse_bson.cpp -o $@
+19 -3
View File
@@ -7,7 +7,7 @@ users ask when they pick a library. They are not built by CMake or run by CI.
## Reproducing the numbers ## Reproducing the numbers
`compare.py` builds both programs against `include/` of this checkout, runs them, and writes the results together with `compare.py` builds the programs against `include/` of this checkout, runs them, and writes the results together with
everything needed to reproduce them to `results/<date>-<host>.md` (and `.csv`): the date, the commit, the CPU, the everything needed to reproduce them to `results/<date>-<host>.md` (and `.csv`): the date, the commit, the CPU, the
OS, the compiler, the flags, and the versions of all libraries. OS, the compiler, the flags, and the versions of all libraries.
@@ -52,9 +52,18 @@ JSON-RPC request (`rpc`):
`bench_corpus.cpp` runs parse, traverse, and dump on any list of files, so that no library is tuned to a handful of `bench_corpus.cpp` runs parse, traverse, and dump on any list of files, so that no library is tuned to a handful of
documents. documents.
`bench_edit.cpp` measures read-modify-write: parse, apply the same logical edits with each library's own API, and
serialize (compact). Workloads: `patch` (a handful of edits at fixed places) and `update` (edits in every record).
An editable `json_document` edits in place; yyjson copies its immutable document into a mutable one first
(`yyjson_doc_mut_copy`); Boost.JSON and `json::parse` build mutable DOMs; simdjson cannot edit a document. All
outputs are checked to describe the same value.
Before anything is timed, all engines must accept each document and agree on the traversal: the number of values, the Before anything is timed, all engines must accept each document and agree on the traversal: the number of values, the
bytes of all strings and keys, and the sum of all numbers. All engines run interleaved in every round, and the best bytes of all strings and keys, and the sum of all numbers. All engines run interleaved in every round, and the best
round is reported, as time and as a factor of the `json_view` time (below 1 means faster than `json_view`). round is reported, as time and as a factor of the `json_view` time (below 1 means faster than `json_view`). Each timed
call follows an untimed call of the same engine: otherwise the engine after `json::parse` pays for the allocator
cleaning up the tens of thousands of nodes `json::parse` just freed (with glibc, this made `json_view` look 1.7 times
slower on citm_catalog traverse).
The engines do not all offer the same features, which the numbers should be read with: The engines do not all offer the same features, which the numbers should be read with:
@@ -62,11 +71,18 @@ The engines do not all offer the same features, which the numbers should be read
|---|---|---|---|---| |---|---|---|---|---|
| `json_view` | immutable index into the text | yes | no | a fresh document per parse; "reused" parses into the same document | | `json_view` | immutable index into the text | yes | no | a fresh document per parse; "reused" parses into the same document |
| yyjson | immutable (`yyjson_read`) | yes | via a mutable copy | | | yyjson | immutable (`yyjson_read`) | yes | via a mutable copy | |
| simdjson DOM | immutable, parser reused | yes | no | | | simdjson DOM | immutable, parser reused | yes | no | "fresh" uses a new parser per parse |
| simdjson On-Demand | none: forward-only, lazy | no | no | only traverse and select | | simdjson On-Demand | none: forward-only, lazy | no | no | only traverse and select |
| Boost.JSON | owning, mutable DOM | yes | yes | monotonic resource | | Boost.JSON | owning, mutable DOM | yes | yes | monotonic resource |
| `json::parse` | owning, mutable DOM | yes | yes | | | `json::parse` | owning, mutable DOM | yes | yes | |
Reusing memory matters as much as the parser. simdjson DOM reuses its parser, so it writes into memory it already
touched; a fresh `json_view` document or yyjson document gets new memory for every parse. On Linux, glibc returns large
blocks to the system when they are freed, so every fresh parse of a large document pays a page fault per 4 KiB page:
on x86-64 Linux, a fresh `json_view` parse of jeopardy took about twice as long as a reused one. On macOS on Apple
silicon, with 16 KiB pages, the difference is much smaller. Compare "json_view (reused)" with "simdjson DOM", and the
fresh `json_view` with "simdjson DOM (fresh)" and yyjson.
## Published results ## Published results
Results are only published with the file `compare.py` wrote, which names the machine and the versions; see Results are only published with the file `compare.py` wrote, which names the machine and the versions; see
@@ -29,6 +29,7 @@
#include <chrono> #include <chrono>
#include <cmath> #include <cmath>
#include <cstdio> #include <cstdio>
#include <cstdlib>
#include <cstring> #include <cstring>
#include <fstream> #include <fstream>
#include <functional> #include <functional>
@@ -195,6 +196,11 @@ static void walk(const boost::json::value& v, stats& st)
static std::string slurp(const std::string& p) static std::string slurp(const std::string& p)
{ {
std::ifstream f(p, std::ios::binary); std::ifstream f(p, std::ios::binary);
if (!f)
{
std::fprintf(stderr, "cannot open %s\n", p.c_str());
std::exit(1);
}
std::stringstream ss; std::stringstream ss;
ss << f.rdbuf(); ss << f.rdbuf();
return ss.str(); return ss.str();
@@ -222,6 +228,7 @@ int main(int argc, char** argv)
} }
std::FILE* csv = std::fopen("bench_corpus.csv", "w"); std::FILE* csv = std::fopen("bench_corpus.csv", "w");
std::fprintf(csv, "file,bytes,workload,engine,ns\n"); std::fprintf(csv, "file,bytes,workload,engine,ns\n");
json_document reused;
simdjson::dom::parser sj; simdjson::dom::parser sj;
for (const auto& path : files) for (const auto& path : files)
{ {
@@ -272,8 +279,10 @@ int main(int argc, char** argv)
{ {
"parse", { "parse", {
{"json_view", [&] { auto x = json_document::parse(s); g_sink = static_cast<double>(x.node_count()); }}, {"json_view", [&] { auto x = json_document::parse(s); g_sink = static_cast<double>(x.node_count()); }},
{"json_view (reused)", [&] { reused.read(s); g_sink = static_cast<double>(reused.node_count()); }},
{"yyjson", [&] { yyjson_doc* x = yyjson_read(s.data(), s.size(), 0); g_sink = static_cast<double>(yyjson_doc_get_val_count(x)); yyjson_doc_free(x); }}, {"yyjson", [&] { yyjson_doc* x = yyjson_read(s.data(), s.size(), 0); g_sink = static_cast<double>(yyjson_doc_get_val_count(x)); yyjson_doc_free(x); }},
{"simdjson DOM", [&] { auto e = sj.parse(ps).value_unsafe(); g_sink = e.is_object(); }}, {"simdjson DOM", [&] { auto e = sj.parse(ps).value_unsafe(); g_sink = e.is_object(); }},
{"simdjson DOM (fresh)", [&] { simdjson::dom::parser p; auto e = p.parse(ps).value_unsafe(); g_sink = e.is_object(); }},
#if JSON_VIEW_BENCH_BOOST #if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); g_sink = v.is_object(); }}, {"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); g_sink = v.is_object(); }},
#endif #endif
@@ -306,6 +315,9 @@ int main(int argc, char** argv)
{ {
for (std::size_t k = 0; k < wl.second.size(); ++k) for (std::size_t k = 0; k < wl.second.size(); ++k)
{ {
// an untimed call first: whatever the previous engine left to the allocator
// (e.g. thousands of freed json nodes) is cleaned up here, not in the timing
wl.second[k].fn();
const auto t0 = std::chrono::steady_clock::now(); const auto t0 = std::chrono::steady_clock::now();
wl.second[k].fn(); wl.second[k].fn();
best[k] = std::min(best[k], std::chrono::duration<double, std::nano>(std::chrono::steady_clock::now() - t0).count()); best[k] = std::min(best[k], std::chrono::duration<double, std::nano>(std::chrono::steady_clock::now() - t0).count());
+583
View File
@@ -0,0 +1,583 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Read-modify-write benchmark: parse a document, apply the same logical edits
// with each library's own API, and serialize it (compact).
//
// json_view json_editable_document: edits in place, unchanged values stay in the index
// yyjson yyjson_read + yyjson_doc_mut_copy (the way to edit a parsed document)
// Boost.JSON parse into a mutable DOM (monotonic resource, precise numbers), serialize
// json::parse nlohmann::json today
// simdjson has no mutable document and is not part of this comparison.
//
// Workloads:
// patch a handful of edits at fixed places (scalars, a new member, a new array element)
// update edits in every record (twitter: 100 statuses, citm: 243 performances,
// canada: 480 rings, jeopardy: 216,930 questions): set scalars, erase a
// member, add a member (canada: replace the first point of every ring)
//
// Build: see README.md (same flags as bench_view.cpp).
#include <nlohmann/json_view.hpp>
#if JSON_VIEW_BENCH_BOOST
#include <boost/json.hpp>
#include <boost/json/src.hpp>
#endif
#include <yyjson.h>
#include <algorithm>
#include <chrono>
#include <cstdio>
#include <cstdlib>
#include <fstream>
#include <functional>
#include <sstream>
using nlohmann::json;
using nlohmann::json_editable_document;
using nlohmann::json_editable_view;
#if JSON_VIEW_BENCH_BOOST
namespace bj = boost::json;
#endif
static volatile std::size_t g_sink;
// ---------------- json_view ----------------
static std::string edit_view(const std::string& name, const std::string& s, bool update)
{
json_editable_document d = json_editable_document::parse(s);
const json_editable_view r = d.root();
if (name == "twitter")
{
if (update)
{
std::int64_t i = 0;
for (const json_editable_view st : r["statuses"])
{
d.set(st, "retweet_count", i++);
d.set(st, "favorited", true);
d.set(st, "text", "redacted");
d.erase(st, "entities");
d.set(st, "edited", true);
}
}
else
{
d.set(r["search_metadata"], "count", 200);
d.set(r["statuses"][0], "text", "patched");
d.set(r["statuses"][0]["user"], "followers_count", 1);
d.set(r["statuses"][99], "favorited", true);
d.set(r, "patched", true);
}
}
else if (name == "citm_catalog")
{
if (update)
{
for (const json_editable_view p : r["performances"])
{
d.set(p, "name", "performance");
d.set(p, "start", p["start"].get<std::int64_t>() + 1);
d.erase(p, "seatMapImage");
d.set(p, "edited", true);
}
}
else
{
d.set(r["events"]["138586341"], "name", "patched");
d.set(r["performances"][0], "start", 0);
d.set(r["venueNames"], "PLEYEL_PLEYEL", "Salle");
d.set(r, "patched", true);
}
}
else if (name == "canada")
{
const json_editable_view coords = r["features"][0]["geometry"]["coordinates"];
if (update)
{
for (const json_editable_view ring : coords)
{
d.set(ring, 0, json::array({0.5, 0.5}));
}
}
else
{
d.set(r["features"][0]["properties"], "name", "patched");
d.set(r, "type", "FeatureCollection2");
d.set(coords[0], 0, json::array({0.0, 0.0}));
}
}
else if (name == "jeopardy")
{
if (update)
{
for (const json_editable_view q : r)
{
d.set(q, "value", "$1");
d.erase(q, "air_date");
}
}
else
{
d.set(r[0], "value", "$0");
d.set(r[100000], "answer", "patched");
d.set(r[216929], "round", "x");
d.push_back(r, json::object({{"category", "NEW"}, {"value", "$5"}}));
}
}
else if (name == "status")
{
d.set(r, "retweet_count", 1);
d.set(r["user"], "name", "x");
}
else if (name == "rpc")
{
d.set(r, "id", 4);
d.set(r["params"], "subtrahend", 24);
}
return r.dump();
}
// ---------------- nlohmann::json ----------------
static std::string edit_json(const std::string& name, const std::string& s, bool update)
{
json r = json::parse(s);
if (name == "twitter")
{
if (update)
{
std::int64_t i = 0;
for (auto& st : r["statuses"])
{
st["retweet_count"] = i++;
st["favorited"] = true;
st["text"] = "redacted";
st.erase("entities");
st["edited"] = true;
}
}
else
{
r["search_metadata"]["count"] = 200;
r["statuses"][0]["text"] = "patched";
r["statuses"][0]["user"]["followers_count"] = 1;
r["statuses"][99]["favorited"] = true;
r["patched"] = true;
}
}
else if (name == "citm_catalog")
{
if (update)
{
for (auto& p : r["performances"])
{
p["name"] = "performance";
p["start"] = p["start"].get<std::int64_t>() + 1;
p.erase("seatMapImage");
p["edited"] = true;
}
}
else
{
r["events"]["138586341"]["name"] = "patched";
r["performances"][0]["start"] = 0;
r["venueNames"]["PLEYEL_PLEYEL"] = "Salle";
r["patched"] = true;
}
}
else if (name == "canada")
{
json& coords = r["features"][0]["geometry"]["coordinates"];
if (update)
{
for (auto& ring : coords)
{
ring[0] = json::array({0.5, 0.5});
}
}
else
{
r["features"][0]["properties"]["name"] = "patched";
r["type"] = "FeatureCollection2";
coords[0][0] = json::array({0.0, 0.0});
}
}
else if (name == "jeopardy")
{
if (update)
{
for (auto& q : r)
{
q["value"] = "$1";
q.erase("air_date");
}
}
else
{
r[0]["value"] = "$0";
r[100000]["answer"] = "patched";
r[216929]["round"] = "x";
r.push_back(json::object({{"category", "NEW"}, {"value", "$5"}}));
}
}
else if (name == "status")
{
r["retweet_count"] = 1;
r["user"]["name"] = "x";
}
else if (name == "rpc")
{
r["id"] = 4;
r["params"]["subtrahend"] = 24;
}
return r.dump();
}
// ---------------- yyjson ----------------
static std::string edit_yyjson(const std::string& name, const std::string& s, bool update)
{
yyjson_doc* idoc = yyjson_read(s.data(), s.size(), 0);
yyjson_mut_doc* d = yyjson_doc_mut_copy(idoc, nullptr);
yyjson_doc_free(idoc);
yyjson_mut_val* r = yyjson_mut_doc_get_root(d);
auto get = [](yyjson_mut_val * o, const char* k)
{
return yyjson_mut_obj_get(o, k);
};
if (name == "twitter")
{
yyjson_mut_val* sts = get(r, "statuses");
if (update)
{
std::size_t idx, max;
yyjson_mut_val* st;
std::int64_t i = 0;
yyjson_mut_arr_foreach(sts, idx, max, st)
{
yyjson_mut_set_sint(get(st, "retweet_count"), i++);
yyjson_mut_set_bool(get(st, "favorited"), true);
yyjson_mut_set_str(get(st, "text"), "redacted");
yyjson_mut_obj_remove_key(st, "entities");
yyjson_mut_obj_add_bool(d, st, "edited", true);
}
}
else
{
yyjson_mut_set_sint(get(get(r, "search_metadata"), "count"), 200);
yyjson_mut_val* s0 = yyjson_mut_arr_get(sts, 0);
yyjson_mut_set_str(get(s0, "text"), "patched");
yyjson_mut_set_sint(get(get(s0, "user"), "followers_count"), 1);
yyjson_mut_set_bool(get(yyjson_mut_arr_get(sts, 99), "favorited"), true);
yyjson_mut_obj_add_bool(d, r, "patched", true);
}
}
else if (name == "citm_catalog")
{
if (update)
{
std::size_t idx, max;
yyjson_mut_val* p;
yyjson_mut_arr_foreach(get(r, "performances"), idx, max, p)
{
yyjson_mut_set_str(get(p, "name"), "performance");
yyjson_mut_val* start = get(p, "start");
yyjson_mut_set_sint(start, yyjson_mut_get_sint(start) + 1);
yyjson_mut_obj_remove_key(p, "seatMapImage");
yyjson_mut_obj_add_bool(d, p, "edited", true);
}
}
else
{
yyjson_mut_set_str(get(get(get(r, "events"), "138586341"), "name"), "patched");
yyjson_mut_set_sint(get(yyjson_mut_arr_get(get(r, "performances"), 0), "start"), 0);
yyjson_mut_set_str(get(get(r, "venueNames"), "PLEYEL_PLEYEL"), "Salle");
yyjson_mut_obj_add_bool(d, r, "patched", true);
}
}
else if (name == "canada")
{
yyjson_mut_val* f0 = yyjson_mut_arr_get(get(r, "features"), 0);
yyjson_mut_val* coords = get(get(f0, "geometry"), "coordinates");
if (update)
{
static const double half[2] = {0.5, 0.5};
std::size_t idx, max;
yyjson_mut_val* ring;
yyjson_mut_arr_foreach(coords, idx, max, ring)
{
yyjson_mut_arr_replace(ring, 0, yyjson_mut_arr_with_real(d, half, 2));
}
}
else
{
static const double zero[2] = {0.0, 0.0};
yyjson_mut_set_str(get(get(f0, "properties"), "name"), "patched");
yyjson_mut_set_str(get(r, "type"), "FeatureCollection2");
yyjson_mut_arr_replace(yyjson_mut_arr_get(coords, 0), 0, yyjson_mut_arr_with_real(d, zero, 2));
}
}
else if (name == "jeopardy")
{
if (update)
{
std::size_t idx, max;
yyjson_mut_val* q;
yyjson_mut_arr_foreach(r, idx, max, q)
{
yyjson_mut_set_str(get(q, "value"), "$1");
yyjson_mut_obj_remove_key(q, "air_date");
}
}
else
{
yyjson_mut_set_str(get(yyjson_mut_arr_get(r, 0), "value"), "$0");
yyjson_mut_set_str(get(yyjson_mut_arr_get(r, 100000), "answer"), "patched");
yyjson_mut_set_str(get(yyjson_mut_arr_get(r, 216929), "round"), "x");
yyjson_mut_val* o = yyjson_mut_obj(d);
yyjson_mut_obj_add_str(d, o, "category", "NEW");
yyjson_mut_obj_add_str(d, o, "value", "$5");
yyjson_mut_arr_append(r, o);
}
}
else if (name == "status")
{
yyjson_mut_set_sint(get(r, "retweet_count"), 1);
yyjson_mut_set_str(get(get(r, "user"), "name"), "x");
}
else if (name == "rpc")
{
yyjson_mut_set_sint(get(r, "id"), 4);
yyjson_mut_set_sint(get(get(r, "params"), "subtrahend"), 24);
}
std::size_t n = 0;
char* out = yyjson_mut_write(d, 0, &n);
std::string result(out, n);
std::free(out);
yyjson_mut_doc_free(d);
return result;
}
#if JSON_VIEW_BENCH_BOOST
// ---------------- Boost.JSON ----------------
static std::string edit_boost(const std::string& name, const std::string& s, bool update)
{
bj::monotonic_resource mr;
bj::parse_options opt;
opt.numbers = bj::number_precision::precise; // correctly rounded, like the others
bj::value v = bj::parse(s, &mr, opt);
bj::object* const obj = v.if_object(); // nullptr for jeopardy (an array)
if (name == "twitter")
{
bj::array& sts = (*obj)["statuses"].as_array();
if (update)
{
std::int64_t i = 0;
for (auto& e : sts)
{
bj::object& st = e.as_object();
st["retweet_count"] = i++;
st["favorited"] = true;
st["text"] = "redacted";
st.erase("entities");
st["edited"] = true;
}
}
else
{
(*obj)["search_metadata"].as_object()["count"] = 200;
bj::object& s0 = sts[0].as_object();
s0["text"] = "patched";
s0["user"].as_object()["followers_count"] = 1;
sts[99].as_object()["favorited"] = true;
(*obj)["patched"] = true;
}
}
else if (name == "citm_catalog")
{
if (update)
{
for (auto& e : (*obj)["performances"].as_array())
{
bj::object& p = e.as_object();
p["name"] = "performance";
p["start"] = p["start"].as_int64() + 1;
p.erase("seatMapImage");
p["edited"] = true;
}
}
else
{
(*obj)["events"].as_object()["138586341"].as_object()["name"] = "patched";
(*obj)["performances"].as_array()[0].as_object()["start"] = 0;
(*obj)["venueNames"].as_object()["PLEYEL_PLEYEL"] = "Salle";
(*obj)["patched"] = true;
}
}
else if (name == "canada")
{
bj::object& f0 = (*obj)["features"].as_array()[0].as_object();
bj::array& coords = f0["geometry"].as_object()["coordinates"].as_array();
if (update)
{
for (auto& ring : coords)
{
ring.as_array()[0] = bj::array({0.5, 0.5});
}
}
else
{
f0["properties"].as_object()["name"] = "patched";
(*obj)["type"] = "FeatureCollection2";
coords[0].as_array()[0] = bj::array({0.0, 0.0});
}
}
else if (name == "jeopardy")
{
bj::array& a = v.as_array();
if (update)
{
for (auto& e : a)
{
bj::object& q = e.as_object();
q["value"] = "$1";
q.erase("air_date");
}
}
else
{
a[0].as_object()["value"] = "$0";
a[100000].as_object()["answer"] = "patched";
a[216929].as_object()["round"] = "x";
a.push_back(bj::object({{"category", "NEW"}, {"value", "$5"}}));
}
}
else if (name == "status")
{
(*obj)["retweet_count"] = 1;
(*obj)["user"].as_object()["name"] = "x";
}
else if (name == "rpc")
{
(*obj)["id"] = 4;
(*obj)["params"].as_object()["subtrahend"] = 24;
}
return bj::serialize(v);
}
#endif
// ---------------- harness ----------------
static std::string slurp(const std::string& p)
{
std::ifstream f(p, std::ios::binary);
if (!f)
{
std::fprintf(stderr, "cannot open %s\n", p.c_str());
std::exit(1);
}
std::stringstream ss;
ss << f.rdbuf();
return ss.str();
}
int main(int argc, char** argv)
{
if (argc < 2)
{
std::fprintf(stderr, "usage: %s <json_test_data directory> [rounds] [document]\n", argv[0]);
return 1;
}
const std::string T = std::string(argv[1]) + "/";
const int rounds = argc > 2 ? std::atoi(argv[2]) : 20;
const std::string only = argc > 3 ? argv[3] : "";
struct doc
{
std::string name, text;
int batch;
};
std::vector<doc> docs;
for (const char* f :
{"nativejson-benchmark/twitter.json", "nativejson-benchmark/citm_catalog.json", "nativejson-benchmark/canada.json", "jeopardy/jeopardy.json"
})
{
std::string n = std::string(f).substr(std::string(f).find('/') + 1);
docs.push_back({n.substr(0, n.size() - 5), slurp(T + f), 1});
}
docs.push_back({"status", json::parse(docs[0].text)["statuses"][0].dump(), 200});
docs.push_back({"rpc", R"({"jsonrpc": "2.0", "method": "subtract", "params": {"minuend": 42, "subtrahend": 23}, "id": 3})", 5000});
using fn = std::string (*)(const std::string&, const std::string&, bool);
const std::vector<std::pair<std::string, fn>> engines =
{
{"json_view", edit_view}, {"yyjson", edit_yyjson},
#if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", edit_boost},
#endif
{"json::parse", edit_json}
};
std::FILE* csv = std::fopen("bench_edit.csv", "w");
std::fprintf(csv, "doc,bytes,workload,engine,ns\n");
for (const auto& dc : docs)
{
if (!only.empty() && dc.name != only)
{
continue;
}
for (const bool update :
{
false, true
})
{
if (update && (dc.name == "status" || dc.name == "rpc"))
{
continue;
}
// all engines must produce the same value
const json expected = json::parse(edit_json(dc.name, dc.text, update));
bool ok = true;
for (const auto& e : engines)
{
ok = ok && json::parse(e.second(dc.name, dc.text, update)) == expected;
}
std::vector<double> best(engines.size(), 1e300);
const int r = dc.text.size() > 10000000 ? std::max(3, rounds / 4) : rounds;
for (int i = 0; i < r; ++i)
{
for (std::size_t k = 0; k < engines.size(); ++k)
{
// an untimed call first: whatever the previous engine left to the allocator
// (e.g. thousands of freed json nodes) is cleaned up here, not in the timing
g_sink = engines[k].second(dc.name, dc.text, update).size();
const auto t0 = std::chrono::steady_clock::now();
for (int b = 0; b < dc.batch; ++b)
{
g_sink = engines[k].second(dc.name, dc.text, update).size();
}
const double ns = std::chrono::duration<double, std::nano>(std::chrono::steady_clock::now() - t0).count() / dc.batch;
best[k] = std::min(best[k], ns);
}
}
const char* wl = update ? "update" : "patch";
std::printf("%-13s %-7s %s", dc.name.c_str(), wl, ok ? "" : "[OUTPUT MISMATCH] ");
for (std::size_t k = 0; k < engines.size(); ++k)
{
const double us = best[k] / 1e3;
std::printf(" %s %.*fus (%.2fx)", engines[k].first.c_str(), us < 10 ? 3 : (us < 1000 ? 1 : 0), us, best[k] / best[0]);
std::fprintf(csv, "%s,%zu,%s,%s,%.1f\n", dc.name.c_str(), dc.text.size(), wl, engines[k].first.c_str(), best[k]);
}
std::printf("\n");
std::fflush(stdout);
}
}
std::fclose(csv);
}
+16 -2
View File
@@ -10,7 +10,8 @@
// //
// json_view nlohmann/json_view.hpp (fresh document per parse / reused) // json_view nlohmann/json_view.hpp (fresh document per parse / reused)
// yyjson yyjson_read(): immutable document, random access // yyjson yyjson_read(): immutable document, random access
// simdjson DOM dom::parser (reused, as recommended): immutable, random access // simdjson DOM dom::parser (reused, as recommended; "fresh": a new parser
// per parse): immutable, random access
// references (different feature sets): // references (different feature sets):
// simdjson OD On-Demand: forward-only, lazy // simdjson OD On-Demand: forward-only, lazy
// Boost.JSON owning, mutable DOM (monotonic resource) // Boost.JSON owning, mutable DOM (monotonic resource)
@@ -19,7 +20,8 @@
// Workloads: parse (build + free), traverse (visit everything, convert every // Workloads: parse (build + free), traverse (visit everything, convert every
// number, touch every string and key), select (a few fields per document), // number, touch every string and key), select (a few fields per document),
// dump (compact serialization of the parsed document). // dump (compact serialization of the parsed document).
// All engines run interleaved in every round; the best round is reported. // All engines run interleaved in every round, each timed call after an untimed
// one of the same engine; the best round is reported.
#include <nlohmann/json_view.hpp> #include <nlohmann/json_view.hpp>
#if JSON_VIEW_BENCH_BOOST #if JSON_VIEW_BENCH_BOOST
@@ -33,6 +35,7 @@
#include <chrono> #include <chrono>
#include <cmath> #include <cmath>
#include <cstdio> #include <cstdio>
#include <cstdlib>
#include <fstream> #include <fstream>
#include <functional> #include <functional>
#include <map> #include <map>
@@ -581,6 +584,11 @@ static double pick_od(const std::string& name, simdjson::ondemand::document& d)
static std::string slurp(const std::string& p) static std::string slurp(const std::string& p)
{ {
std::ifstream f(p, std::ios::binary); std::ifstream f(p, std::ios::binary);
if (!f)
{
std::fprintf(stderr, "cannot open %s\n", p.c_str());
std::exit(1);
}
std::stringstream ss; std::stringstream ss;
ss << f.rdbuf(); ss << f.rdbuf();
return ss.str(); return ss.str();
@@ -657,6 +665,7 @@ int main(int argc, char** argv)
{"json_view (reused)", [&] { reused.read(s); g_sink = static_cast<double>(reused.node_count()); }}, {"json_view (reused)", [&] { reused.read(s); g_sink = static_cast<double>(reused.node_count()); }},
{"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); g_sink = static_cast<double>(yyjson_doc_get_val_count(d)); yyjson_doc_free(d); }}, {"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); g_sink = static_cast<double>(yyjson_doc_get_val_count(d)); yyjson_doc_free(d); }},
{"simdjson DOM", [&] { auto e = sj.parse(ps).value_unsafe(); g_sink = e.is_object(); }}, {"simdjson DOM", [&] { auto e = sj.parse(ps).value_unsafe(); g_sink = e.is_object(); }},
{"simdjson DOM (fresh)", [&] { simdjson::dom::parser p; auto e = p.parse(ps).value_unsafe(); g_sink = e.is_object(); }},
#if JSON_VIEW_BENCH_BOOST #if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); g_sink = v.is_object(); }}, {"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); g_sink = v.is_object(); }},
#endif #endif
@@ -664,6 +673,7 @@ int main(int argc, char** argv)
}}); }});
workloads.push_back({"traverse", { workloads.push_back({"traverse", {
{"json_view", [&] { auto d = json_document::parse(s); stats st; walk(d.root(), st); g_sink = st.num; }}, {"json_view", [&] { auto d = json_document::parse(s); stats st; walk(d.root(), st); g_sink = st.num; }},
{"json_view (reused)", [&] { reused.read(s); stats st; walk(reused.root(), st); g_sink = st.num; }},
{"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); stats st; walk(yyjson_doc_get_root(d), st); g_sink = st.num; yyjson_doc_free(d); }}, {"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); stats st; walk(yyjson_doc_get_root(d), st); g_sink = st.num; yyjson_doc_free(d); }},
{"simdjson DOM", [&] { stats st; walk(sj.parse(ps).value_unsafe(), st); g_sink = st.num; }}, {"simdjson DOM", [&] { stats st; walk(sj.parse(ps).value_unsafe(), st); g_sink = st.num; }},
{"simdjson OD", [&] { auto d = od.iterate(ps).value_unsafe(); stats st; walk_od(d.get_value().value_unsafe(), st); g_sink = st.num; }}, {"simdjson OD", [&] { auto d = od.iterate(ps).value_unsafe(); stats st; walk_od(d.get_value().value_unsafe(), st); g_sink = st.num; }},
@@ -674,6 +684,7 @@ int main(int argc, char** argv)
}}); }});
workloads.push_back({"select", { workloads.push_back({"select", {
{"json_view", [&] { auto d = json_document::parse(s); g_sink = pick(name, d.root()); }}, {"json_view", [&] { auto d = json_document::parse(s); g_sink = pick(name, d.root()); }},
{"json_view (reused)", [&] { reused.read(s); g_sink = pick(name, reused.root()); }},
{"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); g_sink = pick(name, yyjson_doc_get_root(d)); yyjson_doc_free(d); }}, {"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); g_sink = pick(name, yyjson_doc_get_root(d)); yyjson_doc_free(d); }},
{"simdjson DOM", [&] { g_sink = pick(name, sj.parse(ps).value_unsafe()); }}, {"simdjson DOM", [&] { g_sink = pick(name, sj.parse(ps).value_unsafe()); }},
{"simdjson OD", [&] { auto d = od.iterate(ps).value_unsafe(); g_sink = pick_od(name, d); }}, {"simdjson OD", [&] { auto d = od.iterate(ps).value_unsafe(); g_sink = pick_od(name, d); }},
@@ -710,6 +721,9 @@ int main(int argc, char** argv)
{ {
for (std::size_t k = 0; k < wl.second.size(); ++k) for (std::size_t k = 0; k < wl.second.size(); ++k)
{ {
// an untimed call first: whatever the previous engine left to the allocator
// (e.g. thousands of freed json nodes) is cleaned up here, not in the timing
wl.second[k].fn();
const auto t0 = std::chrono::steady_clock::now(); const auto t0 = std::chrono::steady_clock::now();
for (int b = 0; b < dc.batch; ++b) for (int b = 0; b < dc.batch; ++b)
{ {
+14 -5
View File
@@ -9,7 +9,7 @@
"""Compare json_view with yyjson, simdjson, Boost.JSON, and json::parse. """Compare json_view with yyjson, simdjson, Boost.JSON, and json::parse.
Builds bench_view.cpp and bench_corpus.cpp against the include/ directory of Builds bench_view.cpp, bench_corpus.cpp, and bench_edit.cpp against the include/ directory of
this checkout, runs them, and writes the results with everything needed to this checkout, runs them, and writes the results with everything needed to
reproduce them (date, commit, CPU, OS, compiler, library versions, flags) to reproduce them (date, commit, CPU, OS, compiler, library versions, flags) to
results/<date>-<host>.md and .csv next to this script. results/<date>-<host>.md and .csv next to this script.
@@ -154,11 +154,14 @@ def download_library(name, work):
if not os.path.isfile(archive): if not os.path.isfile(archive):
print(f'downloading {pin["url"]}', flush=True) print(f'downloading {pin["url"]}', flush=True)
# the URLs are the https constants in PINNED, and the SHA-256 is checked below # the URLs are the https constants in PINNED, and the SHA-256 is checked below
urllib.request.urlretrieve(pin['url'], archive) # nosec B310 # (into a .part file first, so that an interrupted download is not kept)
urllib.request.urlretrieve(pin['url'], archive + '.part') # nosec B310
os.replace(archive + '.part', archive)
with open(archive, 'rb') as f: with open(archive, 'rb') as f:
digest = hashlib.sha256(f.read()).hexdigest() digest = hashlib.sha256(f.read()).hexdigest()
if digest != pin['sha256']: if digest != pin['sha256']:
sys.exit(f'error: SHA-256 of {archive} is {digest}, expected {pin["sha256"]}') os.remove(archive) # downloaded again by the next run
sys.exit(f'error: SHA-256 of {archive} is {digest}, expected {pin["sha256"]} (removed)')
src = os.path.join(work, 'download', pin['dir']) src = os.path.join(work, 'download', pin['dir'])
if not os.path.isdir(src): if not os.path.isdir(src):
with tarfile.open(archive) as t: with tarfile.open(archive) as t:
@@ -214,6 +217,10 @@ def main():
ap.add_argument('--corpus', nargs='*', default=[], help='more files for bench_corpus') ap.add_argument('--corpus', nargs='*', default=[], help='more files for bench_corpus')
ap.add_argument('--build-dir', default=os.path.join(HERE, 'build'), help='where to build (default: build/ next to this script)') ap.add_argument('--build-dir', default=os.path.join(HERE, 'build'), help='where to build (default: build/ next to this script)')
args = ap.parse_args() args = ap.parse_args()
# the benchmarks run in the build directory: make the paths absolute
args.data = os.path.abspath(args.data)
args.corpus = [os.path.abspath(f) for f in args.corpus]
args.build_dir = os.path.abspath(args.build_dir)
cxx = os.environ.get('CXX', 'c++') cxx = os.environ.get('CXX', 'c++')
cc = os.environ.get('CC', 'cc') cc = os.environ.get('CC', 'cc')
@@ -249,7 +256,7 @@ def main():
objects.append(obj) objects.append(obj)
binaries = {} binaries = {}
for bench in ['bench_view', 'bench_corpus']: for bench in ['bench_view', 'bench_corpus', 'bench_edit']:
exe = os.path.join(args.build_dir, bench) exe = os.path.join(args.build_dir, bench)
run([cxx] + flags + include + [os.path.join(HERE, bench + '.cpp')] + objects + link + ['-o', exe]) run([cxx] + flags + include + [os.path.join(HERE, bench + '.cpp')] + objects + link + ['-o', exe])
binaries[bench] = exe binaries[bench] = exe
@@ -261,6 +268,8 @@ def main():
capture_output=True, text=True).stdout capture_output=True, text=True).stdout
outputs['bench_corpus'] = run([binaries['bench_corpus']] + corpus, cwd=args.build_dir, outputs['bench_corpus'] = run([binaries['bench_corpus']] + corpus, cwd=args.build_dir,
capture_output=True, text=True).stdout capture_output=True, text=True).stdout
outputs['bench_edit'] = run([binaries['bench_edit'], args.data, str(max(1, args.rounds // 2))], cwd=args.build_dir,
capture_output=True, text=True).stdout
for name, text in outputs.items(): for name, text in outputs.items():
print(text) print(text)
@@ -292,7 +301,7 @@ def main():
f.write(f'\n## {name}\n\n```\n{text.rstrip()}\n```\n') f.write(f'\n## {name}\n\n```\n{text.rstrip()}\n```\n')
with open(stem + '.csv', 'w', encoding='utf-8') as out: with open(stem + '.csv', 'w', encoding='utf-8') as out:
out.write(''.join(f'# {key}: {value}\n' for key, value in meta)) out.write(''.join(f'# {key}: {value}\n' for key, value in meta))
for name in ['bench_view', 'bench_corpus']: for name in ['bench_view', 'bench_corpus', 'bench_edit']:
path = os.path.join(args.build_dir, name + '.csv') path = os.path.join(args.build_dir, name + '.csv')
if os.path.isfile(path): if os.path.isfile(path):
with open(path, encoding='utf-8') as f: with open(path, encoding='utf-8') as f:
+6
View File
@@ -9,6 +9,12 @@ Additionally, `parse_json_view_fuzzer` (`tests/src/fuzzer-parse_json_view.cpp`)
produces, and that a rejected input makes both parsers throw with an identical `what()`. It takes plain JSON text, so it produces, and that a rejected input makes both parsers throw with an identical `what()`. It takes plain JSON text, so it
reuses the `corpus_json` corpus rather than a format of its own. reuses the `corpus_json` corpus rather than a format of its own.
`json_view_image_fuzzer` (`tests/src/fuzzer-json_view_image.cpp`) tests the images of `json_document` (`save()` and
`load()`). It uses each input twice: as an image, which `load()` must either reject with `parse_error.116` or read
safely (with `image_check::full`, the document must also serialize to the JSON it reads as), and as a JSON text, whose
image must load and serialize to the same text. A corpus of images can be made from JSON files with a small program
that calls `json_document::parse(text).save()`; plain JSON files work as well.
## Corpus creation ## Corpus creation
For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a
+95
View File
@@ -0,0 +1,95 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
/*
This file implements a test of json_document images suitable for fuzz
testing. The input is used twice:
- as an image: json_document::load() with image_check::full must either throw
a parse_error or yield a document that serializes to the JSON text it reads
as; with image_check::bounds, reading and serializing must be safe (checked
by the sanitizers), and serializing may only throw type_error.316
- as a JSON text: if json_document::parse() accepts it, the image of the
document must load (with every check) and serialize to the same text
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
#include <cassert>
#include <cstdint>
#include <string>
#include <vector>
#include <nlohmann/json.hpp>
#include <nlohmann/json_view.hpp>
// the checks below are assertions; NDEBUG would compile them away
#ifdef NDEBUG
#error "the fuzzer drivers must be built without NDEBUG"
#endif
using json = nlohmann::json;
using json_document = nlohmann::json_document;
using image_check = json_document::image_check;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// the input as an image
for (const image_check check :
{
image_check::full, image_check::bounds
})
{
json_document d;
try
{
d = json_document::load(data, size, check);
}
catch (const json::parse_error& e)
{
assert(e.id == 116);
continue;
}
std::string dumped;
try
{
dumped = d.root().dump();
}
catch (const json::type_error& e)
{
// invalid UTF-8 can only pass the bounds check
assert(check == image_check::bounds && e.id == 316);
continue;
}
const json j = d.root().materialize();
if (check == image_check::full)
{
assert(json::parse(dumped) == j);
// an image of the loaded document is the input
assert(d.save() == std::vector<std::uint8_t>(data, data + size));
}
}
// the input as a JSON text
const std::string text(reinterpret_cast<const char*>(data), size); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
const json_document parsed = json_document::parse(text, false);
if (!parsed.is_discarded())
{
const std::vector<std::uint8_t> image = parsed.save();
for (const image_check check :
{
image_check::full, image_check::bounds, image_check::none
})
{
const json_document loaded = json_document::load(image, check);
assert(loaded.root().dump() == parsed.root().dump());
}
}
return 0;
}
+11 -11
View File
@@ -33,7 +33,7 @@ TEST_CASE("Binary Formats" * doctest::skip())
const auto ubjson_2_size = json::to_ubjson(j, true).size(); const auto ubjson_2_size = json::to_ubjson(j, true).size();
const auto ubjson_3_size = json::to_ubjson(j, true, true).size(); const auto ubjson_3_size = json::to_ubjson(j, true, true).size();
CHECK(json_size == 2090303); CHECK(json_size == 2090234);
CHECK(bjdata_1_size == 1112030); CHECK(bjdata_1_size == 1112030);
CHECK(bjdata_2_size == 1224148); CHECK(bjdata_2_size == 1224148);
CHECK(bjdata_3_size == 1224148); CHECK(bjdata_3_size == 1224148);
@@ -46,16 +46,16 @@ TEST_CASE("Binary Formats" * doctest::skip())
CHECK(ubjson_3_size == 1169069); CHECK(ubjson_3_size == 1169069);
CHECK((100.0 * double(json_size) / double(json_size)) == Approx(100.0)); CHECK((100.0 * double(json_size) / double(json_size)) == Approx(100.0));
CHECK((100.0 * double(bjdata_1_size) / double(json_size)) == Approx(53.199)); CHECK((100.0 * double(bjdata_1_size) / double(json_size)) == Approx(53.201));
CHECK((100.0 * double(bjdata_2_size) / double(json_size)) == Approx(58.563)); CHECK((100.0 * double(bjdata_2_size) / double(json_size)) == Approx(58.565));
CHECK((100.0 * double(bjdata_3_size) / double(json_size)) == Approx(58.563)); CHECK((100.0 * double(bjdata_3_size) / double(json_size)) == Approx(58.565));
CHECK((100.0 * double(bon8_size) / double(json_size)) == Approx(50.509)); CHECK((100.0 * double(bon8_size) / double(json_size)) == Approx(50.511));
CHECK((100.0 * double(bson_size) / double(json_size)) == Approx(85.849)); CHECK((100.0 * double(bson_size) / double(json_size)) == Approx(85.853));
CHECK((100.0 * double(cbor_size) / double(json_size)) == Approx(50.497)); CHECK((100.0 * double(cbor_size) / double(json_size)) == Approx(50.499));
CHECK((100.0 * double(msgpack_size) / double(json_size)) == Approx(50.526)); CHECK((100.0 * double(msgpack_size) / double(json_size)) == Approx(50.528));
CHECK((100.0 * double(ubjson_1_size) / double(json_size)) == Approx(53.199)); CHECK((100.0 * double(ubjson_1_size) / double(json_size)) == Approx(53.201));
CHECK((100.0 * double(ubjson_2_size) / double(json_size)) == Approx(58.563)); CHECK((100.0 * double(ubjson_2_size) / double(json_size)) == Approx(58.565));
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(55.928)); CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(55.930));
} }
SECTION("twitter.json") SECTION("twitter.json")
+51
View File
@@ -1147,6 +1147,57 @@ TEST_CASE("json_view dump")
CHECK(d.root().dump() == json::parse(text).dump()); CHECK(d.root().dump() == json::parse(text).dump());
CHECK(d.root().dump() == "[1.5,100.0,0,-0.0,1.2345678901234568e+29,18446744073709551615,-9223372036854775808,0.1,1e-07,5e-324]"); CHECK(d.root().dump() == "[1.5,100.0,0,-0.0,1.2345678901234568e+29,18446744073709551615,-9223372036854775808,0.1,1e-07,5e-324]");
CHECK(d.root().dump(-1, ' ', false, json_view::number_format::source) == "[1.50,1E2,-0,-0.0,123456789012345678901234567890,18446744073709551615,-9223372036854775808,0.1,1e-7,5e-324]"); CHECK(d.root().dump(-1, ' ', false, json_view::number_format::source) == "[1.50,1E2,-0,-0.0,123456789012345678901234567890,18446744073709551615,-9223372036854775808,0.1,1e-7,5e-324]");
// also indented, and with ensure_ascii
CHECK(d.root().dump(0, ' ', false, json_view::number_format::source) == "[\n1.50,\n1E2,\n-0,\n-0.0,\n123456789012345678901234567890,\n18446744073709551615,\n-9223372036854775808,\n0.1,\n1e-7,\n5e-324\n]");
CHECK(d.root().dump(-1, ' ', true, json_view::number_format::source) == "[1.50,1E2,-0,-0.0,123456789012345678901234567890,18446744073709551615,-9223372036854775808,0.1,1e-7,5e-324]");
// float tokens of up to 17 significant digits in every spelling: those
// of at most 15 digits are written from their digits, the others
// through the conversion; both as dump() writes them
{
std::mt19937_64 tokens(1170); // NOLINT(cert-msc32-c,cert-msc51-cpp,bugprone-random-generator-seed)
// a number below n; the remainder is a std::uint64_t, which is
// std::size_t on some platforms and wider on others
const auto draw = [&tokens](std::size_t n)
{
const std::uint64_t r = tokens() % n;
return static_cast<std::size_t>(r);
};
std::string many_tokens = "[";
for (int i = 0; i < 20000; ++i)
{
const std::size_t length = 1 + draw(17);
std::string digits(1, static_cast<char>('1' + draw(9)));
for (std::size_t k = 1; k < length; ++k)
{
digits += static_cast<char>('0' + draw(10));
}
digits += std::string(draw(4), '0'); // trailing zeros
std::string token = draw(3) == 0 ? "-" : "";
const std::size_t point = draw(digits.size() + 1);
if (point == 0)
{
token += "0." + std::string(draw(5), '0') + digits;
}
else
{
token += digits.substr(0, point) + (point < digits.size() ? "." + digits.substr(point) : "");
}
// an exponent that keeps the value between about 1e-320 and 1e300
const int exponent = static_cast<int>(draw(600)) - 300 - static_cast<int>(point);
if (draw(4) != 0)
{
token += (draw(2) == 0 ? "e" : "E") + std::string(exponent >= 0 && draw(2) == 0 ? "+" : "") + std::to_string(exponent);
}
else if (point == digits.size())
{
token += ".0"; // (a float, not an integer)
}
many_tokens += (i != 0 ? "," : "") + token;
}
many_tokens += ']';
CHECK(json_document::parse(many_tokens).root().dump() == json::parse(many_tokens).dump());
}
// random doubles, written as parse() and dump() would // random doubles, written as parse() and dump() would
std::mt19937_64 rng(1170); // NOLINT(cert-msc32-c,cert-msc51-cpp,bugprone-random-generator-seed) std::mt19937_64 rng(1170); // NOLINT(cert-msc32-c,cert-msc51-cpp,bugprone-random-generator-seed)
+85 -1
View File
@@ -26,6 +26,7 @@ using ptr_t = ordered_json::json_pointer;
#include <functional> #include <functional>
#include <iterator> #include <iterator>
#include <limits> #include <limits>
#include <map>
#include <random> #include <random>
#include <string> #include <string>
#include <vector> #include <vector>
@@ -278,12 +279,45 @@ TEST_CASE("json_view edits: differential")
d.set(tv, key, v); d.set(tv, key, v);
j[p][key] = v; j[p][key] = v;
} }
else if (op == 6 && target.is_object() && !target.empty()) // erase a member
{
const std::string key = std::next(target.begin(), r(static_cast<int>(target.size()))).key();
if (r(2) == 0)
{
d.erase(tv, key);
}
else
{
d.erase(p / key);
}
j[p].erase(key);
}
else if (op == 7 && (target.is_array() || target.is_null())) // push_back else if (op == 7 && (target.is_array() || target.is_null())) // push_back
{ {
const ordered_json v = random_value(2); const ordered_json v = random_value(2);
d.push_back(tv, v); d.push_back(tv, v);
j[p].push_back(v); j[p].push_back(v);
} }
else if (op == 8 && target.is_array()) // insert
{
const auto i = static_cast<std::size_t>(r(static_cast<int>(target.size()) + 1));
const ordered_json v = random_value(2);
d.insert(tv, i, v);
j[p].insert(j[p].begin() + static_cast<std::ptrdiff_t>(i), v);
}
else if (op == 9 && target.is_array() && !target.empty()) // erase an element
{
const auto i = static_cast<std::size_t>(r(static_cast<int>(target.size())));
if (r(2) == 0)
{
d.erase(tv, i);
}
else
{
d.erase(p / i);
}
j[p].erase(i);
}
else if (op == 10 && target.is_array() && !target.empty()) // assign an element else if (op == 10 && target.is_array() && !target.empty()) // assign an element
{ {
const auto i = static_cast<std::size_t>(r(static_cast<int>(target.size()))); const auto i = static_cast<std::size_t>(r(static_cast<int>(target.size())));
@@ -350,6 +384,13 @@ TEST_CASE("json_view edits: errors")
CHECK_THROWS_WITH_AS(d.set(root["a"], 2, 1), "[json.exception.out_of_range.401] array index 2 is out of range", json::out_of_range&); CHECK_THROWS_WITH_AS(d.set(root["a"], 2, 1), "[json.exception.out_of_range.401] array index 2 is out of range", json::out_of_range&);
CHECK_THROWS_WITH_AS(d.set(root["a"], -1, 1), "[json.exception.out_of_range.401] array index -1 is out of range", json::out_of_range&); CHECK_THROWS_WITH_AS(d.set(root["a"], -1, 1), "[json.exception.out_of_range.401] array index -1 is out of range", json::out_of_range&);
CHECK_THROWS_WITH_AS(d.push_back(root["o"], 1), "[json.exception.type_error.308] cannot use push_back() with object", json::type_error&); CHECK_THROWS_WITH_AS(d.push_back(root["o"], 1), "[json.exception.type_error.308] cannot use push_back() with object", json::type_error&);
CHECK_THROWS_WITH_AS(d.insert(root["n"], 0, 1), "[json.exception.type_error.309] cannot use insert() with number", json::type_error&);
CHECK_THROWS_WITH_AS(d.insert(root["a"], 3, 1), "[json.exception.out_of_range.401] array index 3 is out of range", json::out_of_range&);
CHECK_THROWS_WITH_AS(d.erase(root["n"], "k"), "[json.exception.type_error.307] cannot use erase() with number", json::type_error&);
CHECK_THROWS_WITH_AS(d.erase(root["o"], 0), "[json.exception.type_error.307] cannot use erase() with object", json::type_error&);
CHECK_THROWS_WITH_AS(d.erase(root["a"], 2), "[json.exception.out_of_range.401] array index 2 is out of range", json::out_of_range&);
CHECK_THROWS_WITH_AS(d.erase(json::json_pointer("")), "[json.exception.out_of_range.405] JSON pointer has no parent", json::out_of_range&);
CHECK_THROWS_WITH_AS(d.erase(json::json_pointer("/missing/x")), "[json.exception.out_of_range.403] key 'missing' not found", json::out_of_range&);
CHECK_THROWS_WITH_AS(d.set(json::json_pointer("/a/01"), 1), "[json.exception.parse_error.106] parse error: array index '01' must not begin with '0'", json::parse_error&); CHECK_THROWS_WITH_AS(d.set(json::json_pointer("/a/01"), 1), "[json.exception.parse_error.106] parse error: array index '01' must not begin with '0'", json::parse_error&);
CHECK_THROWS_WITH_AS(d.set(root, json_editable_view()), "[json.exception.type_error.302] type must be a value, but is discarded", json::type_error&); CHECK_THROWS_WITH_AS(d.set(root, json_editable_view()), "[json.exception.type_error.302] type must be a value, but is discarded", json::type_error&);
CHECK_THROWS_WITH_AS(d.set(root, json::binary({1, 2})), "[json.exception.type_error.319] cannot store a binary value in a json_document", json::type_error&); CHECK_THROWS_WITH_AS(d.set(root, json::binary({1, 2})), "[json.exception.type_error.319] cannot store a binary value in a json_document", json::type_error&);
@@ -390,6 +431,27 @@ TEST_CASE("json_view edits: views and values")
CHECK(inner.get<int>() == 5); CHECK(inner.get<int>() == 5);
} }
SECTION("views keep referring to their value")
{
json_editable_document d = json_editable_document::parse(R"({"a": [10, 20, 30], "b": {"c": "text"}})");
const json_editable_view a = d.root()["a"];
const json_editable_view twenty = a[1];
const json_editable_view c = d.root()["b"]["c"];
d.insert(a, 0, 5);
d.push_back(a, 40);
CHECK(twenty.get<int>() == 20);
CHECK(a[2].get<int>() == 20);
d.erase(a, 2);
CHECK(twenty.get<int>() == 20); // an erased value keeps its last value
d.set(c, 7);
CHECK(c.get<int>() == 7); // a held view sees an assignment
d.set(d.root()["b"], json::array({1, 2}));
CHECK(d.root()["b"].dump() == "[1,2]");
CHECK(d.root().dump() == R"({"a":[5,10,30,40],"b":[1,2]})");
CHECK(d.root()["a"][0].source_offset() == static_cast<std::size_t>(-1)); // a new value
CHECK(d.root()["a"][1].source_offset() != static_cast<std::size_t>(-1));
}
SECTION("strings stay valid while more edits come") SECTION("strings stay valid while more edits come")
{ {
json_editable_document d = json_editable_document::parse("[]"); json_editable_document d = json_editable_document::parse("[]");
@@ -420,6 +482,19 @@ TEST_CASE("json_view edits: views and values")
CHECK(d.root().materialize().dump() == json::parse(R"([1.5, 100.0, 0.1, null, null, 18446744073709551615, -9223372036854775808])").dump()); CHECK(d.root().materialize().dump() == json::parse(R"([1.5, 100.0, 0.1, null, null, 18446744073709551615, -9223372036854775808])").dump());
} }
SECTION("numbers of other float types")
{
// doubles have their own path to the output; other float types are
// written as basic_json writes them, non-finite values as null
using json_float = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
using document_float = nlohmann::basic_json_document<json_float, true>;
document_float d = document_float::parse("[1.5]");
d.push_back(d.root(), std::numeric_limits<float>::quiet_NaN());
d.push_back(d.root(), -std::numeric_limits<float>::infinity());
CHECK(d.root().dump() == "[1.5,null,null]");
CHECK(d.root().dump(2) == json_float::parse("[1.5, null, null]").dump(2));
}
SECTION("nulls become containers, and the root can be replaced") SECTION("nulls become containers, and the root can be replaced")
{ {
json_editable_document d = json_editable_document::parse("[null, null]"); json_editable_document d = json_editable_document::parse("[null, null]");
@@ -433,6 +508,10 @@ TEST_CASE("json_view edits: views and values")
d.set(json::json_pointer("/x/3"), 4); // the size of the array appends too d.set(json::json_pointer("/x/3"), 4); // the size of the array appends too
d.set(json::json_pointer("/y"), false); d.set(json::json_pointer("/y"), false);
CHECK(d.root().dump() == R"({"x":[1,2,3,4],"y":false})"); CHECK(d.root().dump() == R"({"x":[1,2,3,4],"y":false})");
CHECK(d.erase(json::json_pointer("/x/0")) == 1);
CHECK(d.erase(json::json_pointer("/y")) == 1);
CHECK(d.erase(json::json_pointer("/nothing")) == 0);
CHECK(d.root().dump() == R"({"x":[2,3,4]})");
} }
SECTION("duplicate keys") SECTION("duplicate keys")
@@ -440,6 +519,9 @@ TEST_CASE("json_view edits: views and values")
json_editable_document d = json_editable_document::parse(R"({"a": 1, "b": 2, "a": 3})"); json_editable_document d = json_editable_document::parse(R"({"a": 1, "b": 2, "a": 3})");
d.set(d.root(), "a", 4); // the first member is assigned, the others dropped d.set(d.root(), "a", 4); // the first member is assigned, the others dropped
CHECK(d.root().dump() == R"({"a":4,"b":2})"); CHECK(d.root().dump() == R"({"a":4,"b":2})");
d = json_editable_document::parse(R"({"a": 1, "b": 2, "a": 3})");
CHECK(d.erase(d.root(), "a") == 2);
CHECK(d.root().dump() == R"({"b":2})");
} }
SECTION("values from other documents") SECTION("values from other documents")
@@ -472,7 +554,9 @@ TEST_CASE("json_view edits: views and values")
d.set(d.root(), "new", 1); // appended: the members move, the lookup is linear d.set(d.root(), "new", 1); // appended: the members move, the lookup is linear
CHECK(d.root()["new"].get<int>() == 1); CHECK(d.root()["new"].get<int>() == 1);
CHECK(d.root()["k199"].get<int>() == 199); CHECK(d.root()["k199"].get<int>() == 199);
CHECK(d.root().size() == 201); d.erase(d.root(), "k0");
CHECK(!d.root().contains("k0"));
CHECK(d.root().size() == 200);
} }
SECTION("reuse and memory") SECTION("reuse and memory")
+805
View File
@@ -0,0 +1,805 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json_view.hpp>
using nlohmann::json;
using nlohmann::ordered_json;
using nlohmann::json_document;
using nlohmann::json_editable_document;
using nlohmann::ordered_json_document;
using nlohmann::ordered_json_editable_document;
using image_check = json_document::image_check;
using nlohmann::detail::view::node;
#include <array>
#include <cstdint>
#include <cstring>
#include <fstream>
#include <functional>
#include <limits>
#include <random>
#include <sstream>
#include <string>
#include <utility>
#include <vector>
#include <test_data.hpp>
#if !(defined(__BYTE_ORDER__) && defined(__ORDER_BIG_ENDIAN__) && __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__)
namespace
{
#if !defined(JSON_NOEXCEPTION)
std::string exception_of(const std::function<void()>& f)
{
try
{
f();
}
catch (const json::exception& e)
{
return e.what();
}
return "";
}
const char* const check_failed = "[json.exception.parse_error.116] parse error: invalid json_document image: the check failed";
#endif
std::string read_file(const std::string& name)
{
std::ifstream f(std::string(TEST_DATA_DIRECTORY) + name, std::ios::binary);
std::stringstream ss;
ss << f.rdbuf();
return ss.str();
}
// the offsets of the parts of an image
constexpr std::size_t header_size = 64;
std::uint64_t header_field(const std::vector<std::uint8_t>& image, std::size_t offset)
{
std::uint64_t v = 0;
std::memcpy(&v, image.data() + offset, sizeof(v));
return v;
}
void set_header_field(std::vector<std::uint8_t>& image, std::size_t offset, std::uint64_t v)
{
std::memcpy(image.data() + offset, &v, sizeof(v));
}
std::size_t node_count(const std::vector<std::uint8_t>& image)
{
const std::uint64_t count = header_field(image, 8);
return static_cast<std::size_t>(count);
}
std::size_t text_at(const std::vector<std::uint8_t>& image)
{
return header_size + (node_count(image) * sizeof(node));
}
node node_at(const std::vector<std::uint8_t>& image, std::size_t i)
{
node n{};
std::memcpy(&n, image.data() + header_size + (i * sizeof(node)), sizeof(node));
return n;
}
void set_node(std::vector<std::uint8_t>& image, std::size_t i, const node& n)
{
std::memcpy(image.data() + header_size + (i * sizeof(node)), &n, sizeof(node));
}
#if !defined(JSON_NOEXCEPTION)
/// the result of loading an image with a check: "" or the exception message
std::string load_result(const std::vector<std::uint8_t>& image, image_check check)
{
return exception_of([&]
{
const json_document d = json_document::load(image, check);
static_cast<void>(d);
});
}
/// a copy of the image with node i changed by f
template<typename F>
std::vector<std::uint8_t> corrupted(const std::vector<std::uint8_t>& image, std::size_t i, F f)
{
std::vector<std::uint8_t> b = image;
node n = node_at(b, i);
f(n);
set_node(b, i, n);
return b;
}
#endif
/// a document and the documents loaded from its image must be equal
template<typename Document>
void check_round_trip(const Document& d)
{
const std::vector<std::uint8_t> image = d.save();
for (const image_check check :
{
image_check::full, image_check::bounds, image_check::none
})
{
const json_document l = json_document::load(image, check);
CHECK(l.root().dump() == d.root().dump());
CHECK(l.root().dump(2) == d.root().dump(2));
CHECK(l.root().materialize() == json(d.root().materialize()));
// an image of a loaded document is the same image
CHECK(l.save() == image);
}
// an editable document can be loaded, too
const ordered_json_editable_document e = ordered_json_editable_document::load(image);
CHECK(e.root().dump() == d.root().dump());
}
std::uint32_t rng()
{
static std::mt19937 generator(5295); // NOLINT(cert-msc32-c,cert-msc51-cpp,bugprone-random-generator-seed): reproducible
// result_type is std::uint_fast32_t, which may be wider than 32 bits
const std::mt19937::result_type value = generator();
return static_cast<std::uint32_t>(value);
}
} // namespace
TEST_CASE("json_view images: round trips")
{
SECTION("small documents")
{
for (const char* text :
{
"null", "true", "false", "0", "-0", "42", "-42", "18446744073709551615", "-9223372036854775808",
"123456789012345678901234567890", "1.5", "-1.25e-300", "1E308", "0.1000000000000000000000000001",
"\"\"", "\"text\"", R"("esc\"aped\n\u00e9\ud83d\ude00")", "\"\xc3\xa9\xe3\x81\x82\"",
"[]", "{}", "[[]]", "[{}]", "{\"\":{}}",
R"({"a": [1, 2.5, "x\ty", true, null, {"b": []}], "c": {"d": -3, "eA": "f"}})",
R"({"k": 1, "k": 2, "l": [], "k": 3})",
" [1 , 2 ] "
})
{
CAPTURE(text);
// false positive: parse() returns a document with a root
// @infer-ignore NULLPTR_DEREFERENCE
check_round_trip(json_document::parse(text));
// false positive: parse() returns a document with a root
// @infer-ignore NULLPTR_DEREFERENCE
check_round_trip(ordered_json_document::parse(text));
}
}
SECTION("files")
{
for (const char* name :
{
"/json_testsuite/sample.json", "/nativejson-benchmark/canada.json", "/nativejson-benchmark/citm_catalog.json",
"/nativejson-benchmark/twitter.json", "/json_tests/pass1.json", "/json_tests/pass2.json", "/json_tests/pass3.json"
})
{
CAPTURE(name);
const std::string text = read_file(name);
const json_document d = json_document::parse(text);
check_round_trip(d);
// what a loaded document reads is what parse() produces
CHECK(json_document::load(d.save()).root().materialize() == json::parse(text));
}
}
SECTION("images are deterministic")
{
const std::string text = R"({"b": [1, 2, {"c": "\u00e9"}], "a": 1.5})";
const json_document d = json_document::parse(text);
CHECK(d.save() == json_document::parse(text).save());
CHECK(d.save() == json_editable_document::parse(text).save());
CHECK(d.save() == ordered_json_document::parse(text).save());
const json_document copy = json_document::parse_copy(text);
CHECK(copy.save() == d.save());
}
SECTION("large objects get their hash index again")
{
std::string text = "{";
for (int i = 0; i < 1000; ++i)
{
text += (i != 0 ? ",\"k" : "\"k") + std::to_string(i) + "\":" + std::to_string(i);
}
text += R"(,"k7":"a duplicate","inner":{)";
for (int i = 0; i < 200; ++i)
{
text += (i != 0 ? ",\"m" : "\"m") + std::to_string(i) + "\":" + std::to_string(-i);
}
text += "}}";
const json_document d = json_document::parse(text);
const std::vector<std::uint8_t> image = d.save();
for (const image_check check :
{
image_check::full, image_check::none
})
{
const json_document l = json_document::load(image, check);
for (int i = 0; i < 1000; ++i)
{
CHECK(l.root()["k" + std::to_string(i)] == d.root()["k" + std::to_string(i)]);
}
CHECK(l.root()["k7"].get<int>() == 7); // the first of duplicate keys
CHECK(l.root()["inner"]["m199"].get<int>() == -199);
CHECK(!l.root().contains("k1000"));
// the index is not part of the image
CHECK(l.save() == image);
}
// the nodes of objects in the image do not carry the number of an index
CHECK(node_at(image, 0).extra == 0);
}
}
TEST_CASE("json_view images: edited documents")
{
const std::string text = R"({"name": "x", "n": 1, "f": 2.5, "list": [1, 2, 3], "obj": {"a": "\u00e9", "b": [true]}, "s": "a\"b"})";
SECTION("every kind of edit")
{
ordered_json_editable_document d = ordered_json_editable_document::parse(text);
d.set(d.root()["name"], "a new \"name\""); // string in the edit arena
d.set(d.root()["n"], -17); // negative integer
d.set(d.root(), "p", 5); // non-negative number_integer
d.set(d.root(), "u", 18446744073709551615u); // unsigned
d.set(d.root()["f"], 0.1); // float token
d.set(d.root(), "nan", std::numeric_limits<double>::quiet_NaN());
d.set(d.root(), "inf", -std::numeric_limits<double>::infinity());
d.push_back(d.root()["list"], "pushed"); // moved array
d.insert(d.root()["list"], 0, ordered_json::object({{"new", {1, 2}}}));
d.erase(d.root()["list"], 2);
d.erase(d.root(), "s");
d.set(d.root()["obj"], "c", ordered_json::array({1, "two", 3.5, nullptr, false})); // new object member with a new array
d.set(d.root(), "copy", d.root()["obj"]); // a copy of a subtree
d.set(d.root(), "key \xc3\xa9", true); // a key in the edit arena
const std::vector<std::uint8_t> image = d.save();
const ordered_json expected = ordered_json::parse(d.root().dump());
for (const image_check check :
{
image_check::full, image_check::bounds, image_check::none
})
{
const ordered_json_document l = ordered_json_document::load(image, check);
CHECK(l.root().dump() == d.root().dump());
CHECK(l.root().materialize() == expected);
CHECK(l.root()["nan"].is_null());
CHECK(l.root()["inf"].is_null());
CHECK(l.root()["p"].is_number_integer());
CHECK(l.root()["p"].get<int>() == 5);
CHECK(l.root()["f"].get<double>() == 0.1);
CHECK(l.root()["u"].get<std::uint64_t>() == 18446744073709551615u);
}
// the node index is in document order again: an image of the loaded
// document is the same image
CHECK(ordered_json_document::load(image).save() == image);
// number tokens of edits follow the source; the text is the source's
// prefix
const ordered_json_document l = ordered_json_document::load(image);
REQUIRE(l.source().size() > text.size());
CHECK(std::string(l.source().data(), text.size()) == text);
}
SECTION("a loaded document can be edited and saved again")
{
const std::vector<std::uint8_t> first = json_editable_document::parse(text).save();
json_editable_document d = json_editable_document::load(first);
d.set(d.root()["obj"]["a"], "changed");
d.push_back(d.root()["list"], 4);
d.set(d.root(), "z", json::array({json::object()}));
const std::vector<std::uint8_t> second = d.save();
const json_document l = json_document::load(second);
CHECK(l.root().dump() == d.root().dump());
CHECK(l.root()["obj"]["a"] == "changed");
CHECK(l.root()["list"].size() == 4);
}
SECTION("the root replaced")
{
json_editable_document d = json_editable_document::parse(text);
d.set(d.root(), json::array({1, "x"}));
check_round_trip(d);
d.set(d.root(), 3.5);
check_round_trip(d);
d.set(d.root(), "text");
check_round_trip(d);
}
}
TEST_CASE("json_view images: ownership")
{
const std::string text = R"({"a": "esc\u00e9aped", "b": [1, 2]})";
const std::vector<std::uint8_t> image = json_document::parse(text).save();
SECTION("borrowed")
{
const json_document d = json_document::load(image);
CHECK(!d.owns_source());
CHECK(d.root()["a"] == "esc\xc3\xa9" "aped");
const json_document p = json_document::load(image.data(), image.size());
CHECK(!p.owns_source());
CHECK(p.root() == d.root());
// the text is the image's
CHECK(d.source().data() == reinterpret_cast<const char*>(image.data() + text_at(image)));
}
SECTION("owned")
{
std::vector<std::uint8_t> copy = image;
const std::uint8_t* const data = copy.data();
json_document d = json_document::load(std::move(copy));
CHECK(d.owns_source());
CHECK(d.source().data() == reinterpret_cast<const char*>(data + text_at(image)));
CHECK(d.memory_usage() >= image.size());
CHECK(d.root()["b"][1] == 2);
// read() replaces the image
d.read(std::string("[1]"));
CHECK(d.owns_source());
CHECK(d.root().dump() == "[1]");
const std::string borrowed = "[2]";
d.read(borrowed);
CHECK(!d.owns_source());
}
SECTION("shrink_to_fit keeps the decoded strings of the image")
{
json_document d = json_document::parse(R"(["\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9\u00e9"])");
d.shrink_to_fit();
const std::vector<std::uint8_t> img = d.save();
json_document l = json_document::load(img);
l.shrink_to_fit();
CHECK(l.root().dump() == d.root().dump());
CHECK(l.save() == img);
}
}
// the remaining tests are about the exceptions of load() and save()
#if !defined(JSON_NOEXCEPTION)
TEST_CASE("json_view images: errors")
{
SECTION("a literal as the root: dump() after loading")
{
for (const char* text :
{
"null", "true", "false"
})
{
std::vector<std::uint8_t> image = json_document::parse(text).save();
node n = node_at(image, 0);
n.off = static_cast<std::uint32_t>(image.size());
set_node(image, 0, n);
CHECK(load_result(image, image_check::full) == check_failed);
CHECK(load_result(image, image_check::bounds) == check_failed);
}
}
SECTION("saving a discarded document")
{
const json_document empty{};
CHECK(exception_of([&] { static_cast<void>(empty.save()); }) == "[json.exception.type_error.320] cannot save a discarded json_document");
const json_document failed = json_document::parse("[1,", false);
CHECK(exception_of([&] { static_cast<void>(failed.save()); }) == "[json.exception.type_error.320] cannot save a discarded json_document");
}
const std::vector<std::uint8_t> image = json_document::parse(R"({"a": [1, "\u00e9"]})").save();
const std::string prefix = "[json.exception.parse_error.116] parse error: invalid json_document image: ";
SECTION("header and sizes")
{
CHECK(exception_of([]
{
const json_document d = json_document::load(nullptr, 0);
static_cast<void>(d);
}) == prefix + "too short");
CHECK(exception_of([&]
{
const json_document d = json_document::load(image.data(), 63);
static_cast<void>(d);
}) == prefix + "too short");
std::vector<std::uint8_t> bad = image;
bad[0] = 'X';
CHECK(load_result(bad, image_check::full) == prefix + "unknown format");
bad = image;
bad[4] = 2; // version
CHECK(load_result(bad, image_check::full) == prefix + "unknown format");
for (std::size_t reserved = 32; reserved < 64; reserved += 8)
{
bad = image;
bad[reserved + 3] = 1;
CHECK(load_result(bad, image_check::none) == prefix + "unknown format");
}
const auto sizes = [&](std::size_t offset, std::uint64_t v)
{
std::vector<std::uint8_t> b = image;
set_header_field(b, offset, v);
return load_result(b, image_check::none);
};
CHECK(sizes(8, 0) == prefix + "sizes out of range"); // no nodes
CHECK(sizes(8, 1000) == prefix + "sizes out of range"); // more nodes than bytes
CHECK(sizes(8, 0xFFFFFFF0u) == prefix + "sizes out of range");
CHECK(sizes(16, 0xFFFFFFF0u) == prefix + "sizes out of range"); // text size
CHECK(sizes(16, header_field(image, 16) + 1) == prefix + "sizes out of range");
CHECK(sizes(16, image.size()) == prefix + "sizes out of range");
CHECK(sizes(24, 0xFFFFFFF0u) == prefix + "sizes out of range"); // decoded string size
CHECK(sizes(24, header_field(image, 24) - 1) == prefix + "sizes out of range");
// the NULs after the text and the decoded strings
bad = image;
bad[text_at(image) + header_field(image, 16)] = 'x';
CHECK(load_result(bad, image_check::none) == prefix + "sizes out of range");
bad = image;
bad.back() = 'x';
CHECK(load_result(bad, image_check::none) == prefix + "sizes out of range");
// nothing after the image
bad = image;
bad.push_back(0);
CHECK(load_result(bad, image_check::none) == prefix + "sizes out of range");
// nodes, but not even room for the NULs
bad.assign(image.begin(), image.begin() + static_cast<std::ptrdiff_t>(text_at(image)));
CHECK(load_result(bad, image_check::none) == prefix + "sizes out of range");
}
}
TEST_CASE("json_view images: check")
{
// nodes: 0 { 1 "s" 2 "x\"y" (escaped) 3 "i" 4 -12 5 "u" 6 7 7 "f" 8 1.5e300 9 "b" 10 true 11 "n" 12 null
// 13 "a" 14 [ 15 "t" 16 {} ]
const std::string text = R"({"s":"x\"y","i":-12,"u":7,"f":1.5e300,"b":true,"n":null,"a":["t",{}]})";
const std::vector<std::uint8_t> image = json_document::parse(text).save();
REQUIRE(load_result(image, image_check::full).empty());
REQUIRE(node_count(image) == 17);
// bounds: rejected by both checks; content: only by the full one
const auto rejected = [&](const std::vector<std::uint8_t>& b, bool bounds)
{
CHECK(load_result(b, image_check::full) == check_failed);
CHECK(load_result(b, image_check::bounds) == (bounds ? check_failed : ""));
};
SECTION("kinds")
{
const std::array<std::uint8_t, 4> kinds = {{8, 9, 10, 200}}; // binary, discarded, link, unknown
for (const std::uint8_t kind : kinds)
{
rejected(corrupted(image, 12, [&](node & n)
{
n.kind = kind;
}), true);
}
// a key that is not a string
rejected(corrupted(image, 1, [](node & n)
{
n.kind = 0;
n.len = 0;
n.off = 0;
}), true);
}
SECTION("flags and extra")
{
rejected(corrupted(image, 12, [](node & n)
{
n.flags = 4;
}), true);
rejected(corrupted(image, 12, [](node & n)
{
n.extra = 1;
}), true);
rejected(corrupted(image, 10, [](node & n)
{
n.flags = 5;
}), true);
rejected(corrupted(image, 10, [](node & n)
{
n.extra = 1;
}), true);
rejected(corrupted(image, 1, [](node & n)
{
n.flags = 2; // a string in the edit arena
}), true);
rejected(corrupted(image, 1, [](node & n)
{
n.extra = 3;
}), true);
rejected(corrupted(image, 4, [](node & n)
{
n.flags = 2;
}), true);
rejected(corrupted(image, 4, [](node & n)
{
n.extra = static_cast<std::uint16_t>(n.extra | 0x100u); // an integer with fraction digits
}), true);
rejected(corrupted(image, 0, [](node & n)
{
n.flags = 8; // moved
}), true);
rejected(corrupted(image, 0, [](node & n)
{
n.extra = 1; // a hash index
}), true);
}
SECTION("bounds")
{
const std::size_t text_size = header_field(image, 16);
const std::size_t arena_size = header_field(image, 24);
rejected(corrupted(image, 1, [&](node & n)
{
n.off = static_cast<std::uint32_t>(text_size + 1);
}), true);
rejected(corrupted(image, 1, [&](node & n)
{
n.len = static_cast<std::uint32_t>(text_size);
}), true);
rejected(corrupted(image, 2, [&](node & n)
{
n.len = static_cast<std::uint32_t>(arena_size + 1);
}), true);
rejected(corrupted(image, 6, [&](node & n)
{
n.off = static_cast<std::uint32_t>(text_size);
}), true);
rejected(corrupted(image, 6, [&](node & n)
{
n.off = static_cast<std::uint32_t>(text_size + 5);
}), true);
rejected(corrupted(image, 6, [](node & n)
{
n.extra = 0; // no digits
}), true);
rejected(corrupted(image, 8, [&](node & n)
{
n.len = static_cast<std::uint32_t>(text_size);
}), true);
rejected(corrupted(image, 8, [](node & n)
{
n.len = 2; // shorter than the recorded digits
}), true);
rejected(corrupted(image, 14, [&](node & n)
{
n.off = static_cast<std::uint32_t>(text_size + 1);
}), true);
// literals: their offset sizes the output of dump()
rejected(corrupted(image, 10, [&](node & n)
{
n.off = static_cast<std::uint32_t>(text_size + 1);
}), true);
rejected(corrupted(image, 12, [&](node & n)
{
n.off = 0xFFFFFFFFu;
}), true);
}
SECTION("structure")
{
rejected(corrupted(image, 0, [](node & n)
{
n.next = 0;
}), true);
rejected(corrupted(image, 0, [](node & n)
{
n.next = 18; // beyond the image
}), true);
rejected(corrupted(image, 14, [](node & n)
{
n.next = 4; // beyond the enclosing object
}), true);
rejected(corrupted(image, 0, [](node & n)
{
n.len = 6; // member count
}), true);
rejected(corrupted(image, 14, [](node & n)
{
n.len = 3; // element count
}), true);
rejected(corrupted(image, 0, [](node & n)
{
n.next = 14; // the object ends after the key "a"
n.len = 7;
}), true);
rejected(corrupted(image, 0, [](node & n)
{
n.next = 13; // nodes after the root
n.len = 6;
}), true);
rejected(corrupted(image, 0, [](node & n)
{
n.kind = 2; // an array: the "keys" are values, and the counts do not match
}), true);
const std::vector<std::uint8_t> as_array = corrupted(image, 16, [](node & n)
{
n.kind = 2; // {} as []: fine
});
CHECK(load_result(as_array, image_check::full).empty());
CHECK(json_document::load(as_array).root().dump() == R"({"s":"x\"y","i":-12,"u":7,"f":1.5e+300,"b":true,"n":null,"a":["t",[]]})");
}
SECTION("strings")
{
// a quote in a source string (the full check only)
std::vector<std::uint8_t> b = image;
const std::size_t t = text_at(image);
const node t15 = node_at(image, 15);
b[t + t15.off] = '"';
rejected(b, false);
// a control character
b[t + t15.off] = '\n';
rejected(b, false);
// invalid UTF-8 in a decoded string
b = image;
const node s2 = node_at(image, 2);
b[t + header_field(image, 16) + 1 + s2.off] = 0xFF;
rejected(b, false);
}
SECTION("numbers")
{
const std::size_t t = text_at(image);
const node i4 = node_at(image, 4);
const node u6 = node_at(image, 6);
const node f8 = node_at(image, 8);
const auto at_token = [&](const node & n, std::size_t k, std::uint8_t c)
{
std::vector<std::uint8_t> b = image;
b[t + n.off + k] = c;
return b;
};
rejected(at_token(i4, 1, 'x'), false); // -x2
rejected(at_token(i4, 1, '0'), false); // -02
rejected(at_token(i4, 0, '1'), false); // 112 != -12
rejected(at_token(f8, 1, 'x'), false); // 1x5e300
rejected(at_token(f8, 2, 'e'), false); // 1.ee300
rejected(at_token(f8, 4, 'x'), false); // 1.5ex00
rejected(at_token(f8, 3, '0'), false); // 1.50300: another layout
rejected(at_token(f8, 4, '9'), false); // 1.5e900: overflow
rejected(at_token(f8, 0, 'x'), false);
rejected(at_token(u6, 0, '8'), false); // 8 != 7
rejected(corrupted(image, 4, [](node & n)
{
n.kind = 6; // "-12" as unsigned: a sign
n.extra = 3;
}), false);
// a non-negative number_integer (as edits write it): fine
const std::vector<std::uint8_t> positive = corrupted(image, 6, [](node & n)
{
n.kind = 5;
n.extra = 0;
});
CHECK(load_result(positive, image_check::full).empty());
CHECK(json_document::load(positive).root()["u"].is_number_integer());
rejected(corrupted(image, 8, [](node & n)
{
n.kind = 6; // a float token as integer
n.extra = 7;
}), false);
}
SECTION("float tokens of an image checked for bounds only")
{
// A float node whose layout records "many" digits is converted from
// its token alone; a token that is not a JSON number reads as 0.
const std::vector<std::uint8_t> img = json_document::parse("[1.5e300,2]").save();
const std::size_t t = text_at(img);
for (const char* token :
{
"x.5e300", "01.5e30", "1.xe300", "1.5ex00", "1.5e+x0", "1.5e30x", "-.5e300", "1.5E300"
})
{
CAPTURE(token);
std::vector<std::uint8_t> b = img;
node n = node_at(b, 1);
n.extra = 0xFFFFu;
set_node(b, 1, n);
std::memcpy(b.data() + t + n.off, token, n.len);
const json_document d = json_document::load(b, image_check::bounds);
const auto v = d.root()[0].get<double>();
CHECK(v == (std::string(token) == "1.5E300" ? 1.5e300 : 0.0));
CHECK(load_result(b, image_check::full) == (std::string(token) == "1.5E300" ? "" : check_failed));
}
}
SECTION("integer ranges")
{
// tokens of many digits, which the parser stores as floats
const std::string big = R"([123456789012345678901234, 99999999999999999999, 9223372036854775808])";
const std::vector<std::uint8_t> img = json_document::parse(big).save();
const auto as_integer = [&](std::size_t i, std::uint8_t kind, std::uint16_t extra)
{
std::vector<std::uint8_t> b = img;
node n = node_at(b, i);
n.kind = kind;
n.extra = extra;
set_node(b, i, n);
return load_result(b, image_check::full);
};
CHECK(as_integer(1, 6, 24) == check_failed); // more than 20 digits
CHECK(as_integer(2, 6, 20) == check_failed); // more than 2^64 - 1
CHECK(as_integer(3, 5, 18) == check_failed); // more than 2^63 - 1 as number_integer
}
}
TEST_CASE("json_view images: damaged images")
{
// A damaged image must be rejected, or read safely; with the full check,
// it also serializes to the JSON it reads as.
const std::vector<std::string> texts =
{
R"({"a": [1, -2, 3.25, "x\u00e9y", true, null], "b": {"c": "\"q\"", "d": 1e10}, "e": ""})",
R"([[[[]]], {"k": {"k": {"k": 12345678901234567890}}}, "\ud83d\ude00", -0.0, 0])",
};
for (const std::string& text : texts)
{
const std::vector<std::uint8_t> image = json_document::parse(text).save();
for (int round = 0; round < 3000; ++round)
{
std::vector<std::uint8_t> b = image;
const std::uint32_t flips = 1 + (rng() % 3);
for (std::uint32_t k = 0; k < flips; ++k)
{
// mostly the nodes, where the damage matters most
const std::size_t at = rng() % 4 != 0 ? header_size + (rng() % (b.size() - header_size)) : rng() % b.size();
b[at] = static_cast<std::uint8_t>(rng() % 3 == 0 ? rng() : b[at] ^ (1u << (rng() % 8)));
}
for (const image_check check :
{
image_check::full, image_check::bounds
})
{
json_document d;
try
{
d = json_document::load(b, check);
}
catch (const json::parse_error& e)
{
CHECK(e.id == 116);
continue;
}
std::string dumped;
std::string dumped_ascii;
try
{
dumped = d.root().dump();
dumped_ascii = d.root().dump(-1, ' ', true);
}
catch (const json::type_error& e)
{
// invalid UTF-8 (the bounds check only)
CHECK(check == image_check::bounds);
CHECK(e.id == 316);
continue;
}
const json j = d.root().materialize();
if (check == image_check::full)
{
CHECK(json::parse(dumped) == j);
CHECK(json::parse(dumped_ascii) == j);
}
}
}
}
}
#endif
#else
TEST_CASE("json_view images: big-endian targets")
{
const json_document d = json_document::parse("[1]");
CHECK_THROWS_WITH_AS(d.save(), "[json.exception.type_error.320] json_document images need a little-endian target", json::type_error&);
}
#endif
+260 -1
View File
@@ -15,6 +15,23 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::detail::dtoa_impl::reinterpret_bits; using nlohmann::detail::dtoa_impl::reinterpret_bits;
#include <array>
#include <cmath>
#include <cstdint>
#include <cstdio>
#include <cstdlib>
#include <iomanip>
#include <limits>
#include <locale>
#include <random>
#include <sstream>
#include <string>
#include <utility>
#include <vector>
#if defined(JSON_HAS_CPP_17)
#include <charconv>
#endif
namespace namespace
{ {
float make_float(uint32_t sign_bit, uint32_t biased_exponent, uint32_t significand) float make_float(uint32_t sign_bit, uint32_t biased_exponent, uint32_t significand)
@@ -450,7 +467,7 @@ TEST_CASE("formatting")
check_double( 1.2345e+18, "1.2345e+18" ); // 1.2345e+18 1.2345e+18 1.2345e18 check_double( 1.2345e+18, "1.2345e+18" ); // 1.2345e+18 1.2345e+18 1.2345e18
check_double( 1.2345e+19, "1.2345e+19" ); // 1.2345e+19 1.2345e+19 1.2345e19 check_double( 1.2345e+19, "1.2345e+19" ); // 1.2345e+19 1.2345e+19 1.2345e19
check_double( 1.2345e+20, "1.2345e+20" ); // 1.2345e+20 1.2345e+20 1.2345e20 check_double( 1.2345e+20, "1.2345e+20" ); // 1.2345e+20 1.2345e+20 1.2345e20
check_double( 1.2345e+21, "1.2344999999999999e+21" ); // 1.2345e+21 1.2344999999999999e+21 1.2345e21 check_double( 1.2345e+21, "1.2345e+21" ); // 1.2345e+21 1.2344999999999999e+21 1.2345e21
check_double( 1.2345e+22, "1.2345e+22" ); // 1.2345e+22 1.2345e+22 1.2345e22 check_double( 1.2345e+22, "1.2345e+22" ); // 1.2345e+22 1.2345e+22 1.2345e22
} }
@@ -514,3 +531,245 @@ TEST_CASE("formatting")
check_integer(1000000000000000000LL, "1000000000000000000"); check_integer(1000000000000000000LL, "1000000000000000000");
} }
} }
namespace
{
// a small unsigned big integer (32-bit limbs, least significant first), to
// recompute the powers of ten of the shortest double conversion
using big = std::vector<std::uint32_t>;
void big_mul_small(big& x, std::uint32_t m)
{
std::uint64_t carry = 0;
for (auto& limb : x)
{
const std::uint64_t v = (static_cast<std::uint64_t>(limb) * m) + carry;
limb = static_cast<std::uint32_t>(v);
carry = v >> 32u;
}
if (carry != 0)
{
x.push_back(static_cast<std::uint32_t>(carry));
}
}
void big_div_small(big& x, std::uint32_t d)
{
std::uint64_t rest = 0;
for (std::size_t i = x.size(); i-- > 0;)
{
const std::uint64_t v = (rest << 32u) | x[i];
x[i] = static_cast<std::uint32_t>(v / d);
rest = v % d;
}
while (!x.empty() && x.back() == 0)
{
x.pop_back();
}
}
std::size_t big_bit_length(const big& x)
{
std::size_t n = 32 * x.size();
for (std::uint32_t top = x.back(); (top & 0x80000000u) == 0; top <<= 1u)
{
--n;
}
return n;
}
bool big_bit(const big& x, std::size_t i)
{
return ((x[i / 32] >> (i % 32)) & 1u) != 0;
}
/// the 128 most significant bits of x (floor), shifted left if x has fewer bits
std::pair<std::uint64_t, std::uint64_t> big_top128(const big& x)
{
const std::size_t n = big_bit_length(x);
std::uint64_t high = 0;
std::uint64_t low = 0;
for (std::size_t k = 0; k < 128; ++k)
{
const bool bit = k < n && big_bit(x, n - 1 - k);
if (k < 64)
{
high = (high << 1u) | (bit ? 1u : 0u);
}
else
{
low = (low << 1u) | (bit ? 1u : 0u);
}
}
return {high, low};
}
/// the digits (without trailing zeros) and the decimal exponent of a
/// representation "[-]d[.ddd][e[+-]x]"
std::pair<std::string, int> digits_and_exponent(const std::string& s)
{
std::string digits;
int point = -1;
int exponent = 0;
for (std::size_t i = 0; i < s.size(); ++i)
{
const char c = s[i];
if (c >= '0' && c <= '9')
{
digits += c;
}
else if (c == '.')
{
point = static_cast<int>(digits.size());
}
else if (c == 'e' || c == 'E')
{
exponent = std::stoi(s.substr(i + 1));
break;
}
}
int e = exponent + (point < 0 ? static_cast<int>(digits.size()) : point) - static_cast<int>(digits.size());
const std::size_t first = digits.find_first_not_of('0');
digits = first == std::string::npos ? "0" : digits.substr(first);
while (digits.size() > 1 && digits.back() == '0')
{
digits.pop_back();
++e;
}
return {digits, e};
}
/// whether the decimal digits * 10^e reads back as v
bool reads_back(const std::string& digits, int e, double v)
{
const std::string text = digits + "e" + std::to_string(e);
return std::strtod(text.c_str(), nullptr) == v;
}
/// Check the representation of a positive finite double: it reads back as
/// the same value, and no representation with fewer digits does.
void check_shortest(double v)
{
std::array<char, 33> buf{};
char* end = nlohmann::detail::to_chars(buf.data(), buf.data() + 32, v);
const std::string text(buf.data(), end);
CAPTURE(text);
CHECK(std::strtod(text.c_str(), nullptr) == v);
// the layout is that of format_buffer() for the same digits
std::array<char, 64> reference{};
int len = 0;
int exponent = 0;
nlohmann::detail::dtoa_impl::shortest_digits(reference.data(), len, exponent, v);
const char* const reference_end = nlohmann::detail::dtoa_impl::format_buffer(reference.data(), len, exponent, -4, 15);
CHECK(text == std::string(reference.data(), static_cast<std::size_t>(reference_end - reference.data())));
const auto de = digits_and_exponent(text);
const std::string& digits = de.first;
if (digits.size() > 1)
{
// the decimals of one digit fewer next to the value
// (a stream rather than snprintf("%.*e"), whose output GCC cannot bound)
std::ostringstream shorter;
shorter.imbue(std::locale::classic());
shorter << std::scientific << std::setprecision(static_cast<int>(digits.size()) - 2) << v;
const auto near = digits_and_exponent(shorter.str());
// as an integer with digits.size() - 1 digits
std::string m = near.first;
int e = near.second;
while (m.size() < digits.size() - 1)
{
m += '0';
--e;
}
const std::uint64_t mid = std::stoull(m);
for (const std::uint64_t candidate :
{
mid - 1, mid, mid + 1
})
{
CAPTURE(candidate);
CHECK(!reads_back(std::to_string(candidate), e, v));
}
}
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
// the closest of the shortest representations, as std::to_chars finds it
std::array<char, 64> std_text{};
const auto r = std::to_chars(std_text.data(), std_text.data() + std_text.size(), v, std::chars_format::scientific);
CHECK(digits_and_exponent(std::string(std_text.data(), r.ptr)) == de);
#endif
}
} // namespace
TEST_CASE("shortest digits of doubles")
{
SECTION("powers of ten")
{
// the 128-bit significands of 10^k, rounded down, recomputed
for (int k = -342; k <= 341; ++k)
{
CAPTURE(k);
big x{1};
if (k >= 0)
{
for (int i = 0; i < k; ++i)
{
big_mul_small(x, 10);
}
}
else
{
// floor(2^b / 10^-k) for a b that leaves more than 128 bits
const int b = 128 + 64 + (4 * -k);
x.assign(static_cast<std::size_t>(b / 32) + 1, 0);
x.back() = 1u << (b % 32);
for (int i = 0; i < -k; ++i)
{
big_div_small(x, 10);
}
}
const auto expected = big_top128(x);
const auto actual = nlohmann::detail::zmij::pow10(k);
CHECK(actual.high == expected.first);
CHECK(actual.low == expected.second);
}
}
SECTION("boundary values")
{
for (const double v :
{
std::numeric_limits<double>::min(), std::numeric_limits<double>::max(), std::numeric_limits<double>::denorm_min(),
std::nextafter(std::numeric_limits<double>::min(), 0.0), 1.0, 2.0, 0.1, 0.3, 1e21, 1e22, 1e23, 5e-324, 9007199254740993.0,
1.2345e+21, 2.2250738585072014e-308, 1.7976931348623157e308, 4.9406564584124654e-324, 123456789012345680.0
})
{
check_shortest(v);
}
// all powers of two (their rounding interval is narrower below)
for (int e = -1074; e <= 1023; ++e)
{
check_shortest(std::ldexp(1.0, e));
}
// powers of ten and their neighbors
for (int e = -323; e <= 308; ++e)
{
const double p = std::strtod(("1e" + std::to_string(e)).c_str(), nullptr);
check_shortest(p);
check_shortest(std::nextafter(p, 0.0));
check_shortest(std::nextafter(p, std::numeric_limits<double>::infinity()));
}
}
SECTION("random doubles")
{
std::mt19937_64 rng(5295); // NOLINT(cert-msc32-c,cert-msc51-cpp,bugprone-random-generator-seed): reproducible
for (int i = 0; i < 100000; ++i)
{
const std::uint64_t bits = rng() & 0x7FFFFFFFFFFFFFFFu;
const auto v = reinterpret_bits<double>(bits);
if (std::isfinite(v) && v != 0)
{
check_shortest(v);
}
}
}
}