Commit Graph
457 Commits
Author SHA1 Message Date
Niels Lohmann c006db93d4 Reject integer lengths in json_document::parse
parse(ptr, len) compiled: len converted to allow_exceptions, and ptr was
read as a C string, past the end of a buffer without a terminating NUL.
Delete the overloads of parse, parse_copy, accept, and read that take an
integer other than bool where the flags are expected.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 16:10:11 +02:00
Niels Lohmann 86feea50ab Select the Zmij conversion by a trait, not by the double overload
The non-template overloads for double asserted binary64 doubles wherever json.hpp was included, so the library no longer compiled where double is not IEEE 754 binary64 (AVR, -fshort-double). A trait now picks Zmij for any binary64 type, including a long double of that format (MSVC, Apple Arm), and Grisu2 for the others. Remove the unused write_short_decimal, powers_of_ten_16, zmij::decimal, zmij::to_decimal, and shortest_digits(double).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 16:07:43 +02:00
Niels Lohmann 0e729be261 Load string scan words with memcpy and use MSVC bit-scan intrinsics
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 16:01:48 +02:00
Niels Lohmann 1da4b4987d Make the downloads of json_view's compare.py robust
Hash the archives in chunks instead of reading them into memory, download
with a timeout, extract into a temporary directory that is renamed into
place only after success (a half-extracted directory was trusted forever),
and split CXX and CC into arguments so that values like 'ccache g++' work.
Note at the pins that a SHA-256 mismatch of a GitHub tag archive means
that GitHub regenerated it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 15:59:36 +02:00
Niels Lohmann d1b1668668 Merge branch 'json-view/22-view-dump-fast' into json-view/15-view-bench
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:53 +02:00
Niels Lohmann cfcda6dbbe Merge branch 'json-view/21-images' into json-view/22-view-dump-fast
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:51 +02:00
Niels Lohmann 8018ac65de Merge branch 'json-view/19-edit-set' into json-view/21-images
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:50 +02:00
Niels Lohmann 75ef044426 Merge branch 'json-view/16-view-simd' into json-view/19-edit-set
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:45 +02:00
Niels Lohmann 1ec2d710f7 Merge branch 'json-view/13-view-dump' into json-view/16-view-simd
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:42 +02:00
Niels Lohmann 8860bf6f3a Merge branch 'json-view/11-view-access' into json-view/13-view-dump
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:40 +02:00
Niels Lohmann 7852da2bc2 Merge branch 'json-view/08-view-builder' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:37 +02:00
Niels Lohmann 0012f65f60 Merge branch 'json-view/23-zmij' into json-view/08-view-builder
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:35 +02:00
Niels Lohmann 4c43e03d40 Merge CI fixes into json-view/23-zmij
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:39:25 +02:00
Niels Lohmann 78bcd76a36 Merge branch 'json-view/22-view-dump-fast' into json-view/15-view-bench
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:37:51 +02:00
Niels Lohmann 8d27a4afbd Merge branch 'json-view/21-images' into json-view/22-view-dump-fast
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:37:49 +02:00
Niels Lohmann 92bf5fbfb1 Merge branch 'json-view/19-edit-set' into json-view/21-images
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/view/document_data.hpp
#	single_include/nlohmann/json_view.hpp
2026-10-08 16:37:29 +02:00
Niels Lohmann 9962b3cf43 Merge branch 'json-view/16-view-simd' into json-view/19-edit-set
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/view/document_data.hpp
#	single_include/nlohmann/json_view.hpp
2026-10-08 16:33:54 +02:00
Niels Lohmann d22193a8b6 Fix CI findings in the Zmij writer and its tests
- bit_ops.hpp: include macro_scope.hpp (JSON_HEDLEY_ALWAYS_INLINE); fixes IWYU
- to_chars.hpp: C4100 for the unused parameter in release builds, clang-tidy
  sign comparison, cpplint runtime/int, GCC -Wstrict-overflow (unsigned abs)
- unit-to_chars.cpp: parse with the library instead of strtod (MinGW's strtod
  rounds some 16/17 digit inputs wrongly); no floating-point std::to_chars
  with icpc

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:33:32 +02:00
Niels Lohmann 4f1c34e44d Merge branch 'json-view/13-view-dump' into json-view/16-view-simd
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/view/document_data.hpp
#	single_include/nlohmann/json_view.hpp
2026-10-08 16:29:14 +02:00
Niels Lohmann 5db2fa633c Merge branch 'json-view/11-view-access' into json-view/13-view-dump
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:27:26 +02:00
Niels Lohmann acde46421b Merge branch 'json-view/08-view-builder' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:27:24 +02:00
Niels Lohmann 63396cbe00 Merge branch 'json-view/23-zmij' into json-view/08-view-builder
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	Makefile
#	meson.build
2026-10-08 16:26:33 +02:00
Niels Lohmann 22b82b487b Merge branch 'json-view/02b-float-parser' into json-view/23-zmij
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:25:42 +02:00
Niels Lohmann 818ec541d9 Fix MSVC errors in image test and add image_check doc anchor
Move the raw string literal out of the CHECK macro (MSVC preprocessor),
cast 64-bit header fields to std::size_t (C4244 on 32-bit), and give the
image_check section in load.md a real anchor for the mkdocs strict build.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:22:58 +02:00
Niels Lohmann a7de414ce8 Avoid raw string with escapes inside CHECK macro (MSVC C2017)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:19:11 +02:00
Niels Lohmann 43c75bef51 Merge branch 'develop' into json-view/02b-float-parser
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	tests/src/unit-class_lexer.cpp
2026-10-08 16:19:06 +02:00
Niels Lohmann 3d7f554927 Use the with_*_t aliases in tests, examples, and docs (#5787)
* Use the with_*_t aliases in tests, examples, and docs

Replace spelled-out basic_json<...> instantiations that only change one
or two template parameters with nlohmann::json::with_*_t (or
ordered_json::with_*_t when the object type is ordered_map). Types that
change all three number types chain with_integers_t and with_float_t.

The raw basic_json<...> spelling stays where the template parameter
list itself is the subject: the alias tests in unit-udt.cpp, explicit
instantiations, and the ordered_json/compile-time docs.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix unit-large_json for clang and JSON_DIAGNOSTICS

Two test problems from #5781 broke CI on develop: CAPTURE(depth); trips
clang's -Wextra-semi-stmt, and the type_error.321 messages did not
account for the diagnostics path prefix.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use static_cast in unit-hash for clang-tidy

#5772 added functional casts that clang-tidy reports as C-style casts
(google-readability-casting). Also append a char instead of a
one-character string in unit-large_json.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Declare the expected message prefix const in unit-large_json

Without JSON_DIAGNOSTICS the prefix was never modified, which
clang-tidy reports (misc-const-correctness).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-08 16:11:16 +02:00
Suyog Verma a269794db7 Use MSVC intrinsics for full multiplication (#5782)
* Use MSVC intrinsics for full multiplication

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

* Fix formatting in unit-class_lexer

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

* Address review feedback

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

---------

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>
2026-10-08 08:49:32 +02:00
Niels Lohmann c26aed8d51 Merge branch 'json-view/22-view-dump-fast' into json-view/15-view-bench
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:27:00 +02:00
Niels Lohmann 8c5d30b330 Merge branch 'json-view/21-images' into json-view/22-view-dump-fast
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:58 +02:00
Niels Lohmann c7138cc9b4 Merge branch 'json-view/19-edit-set' into json-view/21-images
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:55 +02:00
Niels Lohmann 0a4c6d1a8f Merge branch 'json-view/16-view-simd' into json-view/19-edit-set
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:52 +02:00
Niels Lohmann ad715372bf Merge branch 'json-view/13-view-dump' into json-view/16-view-simd
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:50 +02:00
Niels Lohmann c39680779c Merge branch 'json-view/11-view-access' into json-view/13-view-dump
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:47 +02:00
Niels Lohmann dfdd69d234 Merge branch 'json-view/08-view-builder' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:44 +02:00
Niels Lohmann 33b7b30d06 Merge branch 'json-view/23-zmij' into json-view/08-view-builder
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:42 +02:00
Niels Lohmann cb3c0177ed Merge branch 'json-view/02b-float-parser' into json-view/23-zmij
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:39 +02:00
Niels Lohmann 39d34f30ba Merge branch 'develop' into json-view/02b-float-parser
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:13 +02:00
Niels Lohmann 88ddacb84b Fix CI warnings in own float parser
- pow5_table.hpp: pow5_128_largest_power was unused in this branch's
  own code (GCC -Werror=unused-const-variable); tie it to the table
  size with a static_assert instead of removing it, since a later
  branch in the stack (json-view/23-zmij) uses it.
- number_parse.hpp: rename the local variable `copy` to `buffer` to
  satisfy cpplint's build/include_what_you_use check.
- unit-class_lexer.cpp: extend the NOLINT list on the seeded mt19937
  with bugprone-random-generator-seed, and parenthesize
  `8 * sizeof(Bits) - 1` for clang-tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 20:26:12 +02:00
Niels LohmannandAfonso Januário a5f5d3059b Make std::hash<basic_json> consistent with operator== for numbers (#5772)
* Make std::hash<basic_json> consistent with operator== for numbers

operator== converts between number_integer, number_unsigned, and
number_float before comparing, so json(0), json(0U), and json(0.0)
all compare equal. hash() folded the specific value_t into the
result for each of the three numeric cases, giving each a distinct
hash and breaking the standard Hash requirement that a == b implies
hash(a) == hash(b). A std::unordered_set could therefore hold all
three as separate elements even though they compare equal.

hash() now treats all three numeric variants the same way: it
converts the value to number_float_t and combines it with a single
shared type tag, so any two numbers operator== considers equal hash
identically regardless of which internal type actually holds them.

Updated the accompanying test to check this consistency directly
(including via an actual unordered_set) instead of asserting that 0,
0U, and 0.0 hash differently, since that assumption was the bug.
Also corrected the function's own doc comment and the std::hash API
docs, which described the old behavior as intended.

Fixes #5400

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove now-unused number_integer_t/number_unsigned_t typedefs in hash()

Merging the three numeric branches into one that only reads
number_float_t left these two aliases unused, which several CI
configurations treat as a build error under -Wunused-local-typedefs.

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark the unordered_set in the hash regression test const

clang-tidy's misc-const-correctness check flagged it: the set is
never mutated after construction, only read via size().

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Normalize -0.0 in number hashes and test range ends

operator== compares numbers exactly since #5459, so equal numbers
share one value and convert to the same number_float_t. Update the
comment accordingly, map -0.0 to 0.0 before hashing (std::hash need
not do that), and test -0.0 and the ends of the integer ranges. Show
hash(0.0) in the docs example and note the change in the version
history.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Clarify hash documentation after review

- Say "may hash differently" for null, false, and numbers, since a
  collision across types is possible.
- Name the storage types (signed integer, unsigned integer,
  floating-point number) instead of example literals.
- Explain that the hash survives converting an integer to
  number_float_t but not the lossy conversion back, and that unequal
  numbers may share a hash.
- State that the example hash values are illustrative only and vary by
  platform, compiler, compiler version, and library version.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Afonso Januário <afonso-januario@hotmail.com>
2026-10-07 19:18:51 +02:00
43fe8928e1 Stop binary writers overflowing the stack on deep values (#5781)
* Stop binary writers overflowing the stack on deep values

to_cbor, to_msgpack, and to_ubjson recurse once per nesting level.
The parser is iterative, so a value the library accepts can crash on
the way back out.

Keep the existing recursive path for the first 128 levels and finish
anything deeper on a heap stack. Output is unchanged. BSON is left
alone because its extra size walk is a separate change.

Rebased onto the value-type output sink. The heap frames now initialize
every member, which is what -Weffc++ was rejecting.

See #5392.

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
(cherry picked from commit cf65ac438f)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Redesign the iterative binary writers around a shared recursion depth limit

Address the open review on the non-recursive CBOR/MessagePack/UBJSON/BJData
writers (#5518):

- Delete the CBOR array/object prefix helpers; both the recursive and
  iterative paths call write_cbor_head(), which already existed on develop.
- MessagePack: share one write_msgpack_array_prefix()/write_msgpack_object_prefix()
  helper per container kind between the recursive and iterative paths, both
  going through to_msgpack_length() so an over-long container throws
  out_of_range.412 identically either way.
- Reuse detail::recursion_depth_limit() instead of a separate constant, the
  same bound serializer::dump() and write_bson_document() already use.
- Redesign the frames after bson_frame/dump_frame: only a container with
  elements is ever pushed, its header is written at the point it is pushed,
  and the iterator is set in the frame's constructor instead of a
  default-then-assign two-step with a since-removed "started" flag. The
  UBJSON frame keeps only the value pointer, the per-element prefix_required
  flag, and the iterator; write_closer and is_object are no longer stored,
  since the former is always !use_count (use_count is constant for the whole
  document) and the latter follows from value->is_object().
- Factor the BJData ND-array shape check into is_bjdata_ndarray(), used by
  both the recursive object case and the iterative pushing logic.
- Give the frame classes the GCC -Weffc++ treatment already used for
  diff_frame: a noexcept converting constructor plus the five special members
  defaulted with no explicit noexcept.
- Fix two @ref self-references in write_cbor/write_msgpack/write_ubjson's own
  doc comments to point at the public to_cbor/to_msgpack/to_ubjson/to_bjdata
  API instead.
- The iterative object-key write for CBOR/MessagePack now runs the same
  strict-mode check_utf8() against the parent object as diagnostics context
  that the recursive path already ran, so the two paths raise identical
  diagnostics across the switch-over.
- Rewrite the tests: round trips instead of a bare size check, byte-exact
  comparisons against the recursive output at depths around the bound, a
  deep object and a BJData ND-array past the bound, a deep discarded value
  (type_error.321), and the OSS-Fuzz 566583014 CBOR/MessagePack regression.

BSON is unaffected by this change; it already walks its documents
iteratively and is covered separately by #5553.

Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: ayush-singh-0601 <singhayush062006@gmail.com>
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
2026-10-07 19:18:20 +02:00
Niels Lohmann 3c465beb61 Add tests for error_handler_t::keep in dump() (#4555)
Squashed onto develop from:
- Add error_handler_t::keep to copy invalid UTF-8 bytes unchanged
- Mention error_handler_t::keep in the README

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 19:17:57 +02:00
Niels Lohmann 069ace74af Round-trip BJData ND-array annotations exactly (single precision, key order) (#5707)
Squashed onto develop from:
- Round-trip BJData ND-array annotations exactly (single precision, key order)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 19:17:18 +02:00
Niels Lohmann eb34240b90 Compare json_view with yyjson, simdjson, and Boost.JSON
tests/benchmarks/json_view/ holds a comparison with other
libraries, answering the question users ask when they pick one.
It is not built by CMake and not run by CI.

- bench_view.cpp: parse, traverse, select and dump of twitter,
  citm_catalog, canada, jeopardy, a tweet and an RPC request,
  against json_view, yyjson, simdjson, Boost.JSON and json::parse,
  all agreeing and run interleaved.
- bench_corpus.cpp: parse, traverse and dump of any file list.
- bench_edit.cpp: parses, edits through each library's own API,
  and serializes; all outputs must describe the same value.
- compare.py: builds both programs with system or pinned,
  SHA-256-checked downloads, and writes results with the commit,
  CPU, OS, compiler and library versions.
- README.md and a workflow that runs on demand or via a PR label.

Fairness fixes folded in: each engine runs once untimed before
each timed round, so the next one no longer pays for the
previous one's cleanup; dump is also compared with source numbers
against yyjson's raw-number mode; a "simdjson DOM (fresh)" column
and a reused-document column for json_view were added;
compare.py's paths are absolute; downloads use a .part file and a
failed checksum removes the archive.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:57 +02:00
Niels Lohmann accba4d3b6 Write json_view's dump() without a call or conversion per token
The default dump() (no indentation, no ensure_ascii) gets its own
writer that makes the same walk and produces the same output:

- the write position stays in a local variable instead of a
  member, so the compiler keeps it in a register across stores
  through aliasing char pointers;
- strings and number tokens are copied with fixed-size 32-byte
  moves wherever enough source bytes remain, instead of one
  memcpy call per token;
- the innermost open container lives in local variables; a stack
  that starts as a local array of 32 entries holds the rest;
- unedited documents are walked through the node array in order,
  and integer tokens are read from the source directly.

On top of that, float tokens of at most 15 significant digits are
written straight from their digits via zmij::to_shortest() and
write_shortest(), without converting to a double and back: such
decimals are farther apart than a double's rounding interval, so
the token's digits are the double's shortest digits. Tokens of
16+ digits, or edited values, still go through decimal_to_float().
The view's own NEON write_decimal() is removed in favor of the
shared writer, and the dump output now grows in 64 KiB steps
instead of being resized to its estimate at once.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:54 +02:00
Niels Lohmann 73ceb71c95 Add images of json_documents: save() and load()
An image is a document stored so that loading it needs no
parsing: save() writes the node index, the text and the decoded
strings; the static load() reads an image written by save().

load() takes a pointer and size, a borrowed vector, or an owned
rvalue vector; the nodes are copied so they are aligned and can
be edited, while the text and decoded strings stay in the image.

image_check controls how much load() trusts the input: full
checks structure, bounds, strings and numbers, the parser's own
guarantees; bounds checks structure and bounds only; none skips
all checks, for images from a trusted source.

Layout is little-endian only ("NJVI" header, nodes, text, decoded
strings), following the idea of zero-copy formats such as
FlatBuffers and YaFF; the check follows FlatBuffers' Verifier.

New errors: parse_error.116 for a malformed image or a failed
check, type_error.320 for a discarded document or a big-endian
target.

A dedicated fuzzer and 6,000 seeded corruptions, checked under
ASan/UBSan, found and fixed two gaps: unchecked reserved header
fields, and unbounded null/boolean offsets that could make
dump() throw std::length_error.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:51 +02:00
Niels Lohmann c6ee5a64be Add editable json_documents: set, push_back, insert, and erase
basic_json_document gets a second template parameter, Editable
(false by default), plus the aliases json_editable_document,
json_editable_view, ordered_json_editable_document and
ordered_json_editable_view.

Editable documents can change values and structure without
rewriting the source text: set()/push_back() on values, keys,
array indices and JSON pointers; insert() before an array
element; erase() of an object key, array index or JSON pointer.

New values and element sequences go into edit storage that the
document owns and never moves, so views keep referring to their
value across edits and a parsed node never moves. Read-only
documents walk the plain node array and are unaffected.

Strings are checked for UTF-8 on entry, so dump() of an editable
document never throws type_error.316. Binary values cannot be
stored (type_error.319).

A seeded differential test applies random edits to an editable
document and to the equivalent ordered_json and compares both
after every step.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:38 +02:00
Niels Lohmann 8daec2b596 Scan json_view strings with SIMD and index large objects
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.

Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.

On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.

Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.

simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:32 +02:00
Niels Lohmann da1ca7f9d7 Add dump() and comparisons to json_view
Add basic_json_view::dump() and the comparison operators, and read
floats from the parser's digit layout instead of rescanning the
token.

dump(indent, indent_char, ensure_ascii, number_format) writes a
value the way ordered_json::parse(text).dump() writes it for the
same arguments: members in document order, all of them should a
key occur more than once; strings escaped by the same rules, using
the library's scanning kernels; floats written with the library's
to_chars conversion, so the output equals basic_json's byte for
byte; integers copied from the source, where they are already
canonical, except -0, which parse() reads as 0. There is no
error_handler argument, because the view only holds valid UTF-8.
number_format::source copies numbers exactly as they appear in the
source (e.g. "1.50", "1E2", "-0"), which basic_json cannot provide.
operator<< takes the indentation from the stream width, as for
basic_json. The writer walks iteratively, so nesting depth is
limited by memory only.

operator== and operator!= compare two views, or a view and a
basic_json value in either order, by the rules basic_json's
operator== uses: numbers compare by value across their types,
objects compare by their members with duplicate keys resolved as
parse() resolves them, member order matters only where the object
type keeps one, and discarded views compare as discarded basic_json
values do, including under JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON.
Nothing is materialized except single scalars.

While parsing, the view now records where the integer digits, the
fraction digits, and the exponent of a float token are, so floats
and doubles with at most 19 digits are read from that layout with
the library's decimal_to_float() instead of rescanning the token.
Both round correctly, so the values are those of parse(). get<double>(),
materialize(), dump(), and the comparisons all use it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:29 +02:00
Niels Lohmann 77acd4563c Add element access, iteration, values, and JSON pointers to json_view
Give basic_json_view the read-only access functions of basic_json:
operator[] and at() with keys and indices, front()/back(), find(),
contains(), count(), begin()/end() and cbegin()/cend(), items()
with structured bindings from C++17 on, and type_name().

Exceptions have the ids and messages of the const functions of
basic_json. Where basic_json has undefined behavior the view
answers safely: operator[] with a missing key or an out-of-range
index returns a discarded view, and front()/back() of an empty
container throw invalid_iterator.214. Objects are iterated in
document order, and all members are visited; duplicate-key lookups
find the first member (as yyjson and simdjson do), while parse(),
materialize(), and the map conversions keep the last value, as
parse() does. Keys of up to 16 bytes are compared with two
overlapping loads.

Add value conversions: get<T>()/get_to() for arithmetic types,
bool, nullptr_t, strings (std::basic_string copied,
string_view_t without a copy), BasicJsonType, views, std::vector,
and maps with string keys; get_string() for the string without a
copy; number_token() for the number exactly as written in the
source; value() with keys and JSON pointers; and operator[]/at()/
contains() with JSON pointers. Everything else, including types
with from_json(), goes through materialize() of that subtree.
get<T>() of arithmetic types is inlined down to the conversion, so
reading an integer needs no call.

detail::json_pointer_access exposes a pointer's reference tokens
to code outside basic_json.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:26 +02:00