Commit Graph
421 Commits
Author SHA1 Message Date
Niels Lohmann 44683394f5 Check strings copied from another document for valid UTF-8
An editable document only holds valid UTF-8, but copy_scalar() copied the
strings and keys of a view of another document unchecked. A document of
a weaker check (or a borrowed text that changed after parsing) could
therefore bring ill-formed UTF-8 into it. The copy is checked now, with
the error that dump() reports for the string; copies within the same
document stay unchecked.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:33 +02:00
Niels Lohmann 1d644340d4 Copy nested values into a document without recursion
Counting and copying the nodes of a view or a basic_json value into an
editable document recursed once per nesting level, so that a deeply
nested value overflowed the stack. The four functions now walk the value
with an explicit stack, as materialize() does.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:31 +02:00
Niels Lohmann 671b589d71 Add editable json_documents: set, push_back, insert, and erase
basic_json_document gets a second template parameter, Editable
(false by default), plus the aliases json_editable_document,
json_editable_view, ordered_json_editable_document and
ordered_json_editable_view.

Editable documents can change values and structure without
rewriting the source text: set()/push_back() on values, keys,
array indices and JSON pointers; insert() before an array
element; erase() of an object key, array index or JSON pointer.

New values and element sequences go into edit storage that the
document owns and never moves, so views keep referring to their
value across edits and a parsed node never moves. Read-only
documents walk the plain node array and are unaffected.

Strings are checked for UTF-8 on entry, so dump() of an editable
document never throws type_error.316. Binary values cannot be
stored (type_error.319).

A seeded differential test applies random edits to an editable
document and to the equivalent ordered_json and compares both
after every step.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:30 +02:00
Niels Lohmann 3c2a32ae74 Move the dump, comparison, and large-object tests to unit-json_view_dump.cpp
The dump and comparison tests moved to their own file in the dump pull
request; the hash index tests of this branch join them there.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:28 +02:00
Niels Lohmann 9ed32daab8 Index strings with size_t in the colliding-keys test (MSVC C4244 on Win32)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:27 +02:00
Niels Lohmann d3137f4161 Avoid GCC useless-cast and strict-overflow warnings in unit-json_view.cpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:26 +02:00
Niels Lohmann 9a5c0ef4a7 Fix clang-tidy 22 findings in unit-json_view.cpp and unit-json_view_builder.cpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:25 +02:00
Niels Lohmann a60bff1d7a Keep the first of duplicate keys in the object hash index again
Lookups in objects with 128 members or more return the first member of a
repeated key again, like the linear search of smaller objects. This undoes
the code change of a69542046; its test now expects the first member from
lookups (with and without a table) and the last value from materialize()
and basic_json::parse().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:23 +02:00
Niels Lohmann 594d65f6f1 Keep the last of duplicate keys in the object hash index
Lookups in objects with 128 members or more now return the last member of
a repeated key, like the linear search of smaller objects and like
materialize() and basic_json::parse(). build_object_index let the first
occurrence win, so the same text gave different results depending on the
size of the object.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:22 +02:00
Niels Lohmann 1f0efd425a Release the spare capacity of the hash index in shrink_to_fit
shrink_to_fit() now trims the tables of large objects like the node array and the decoded strings, and the list of large objects is released as soon as the tables are built.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:20 +02:00
Niels Lohmann 745d0eb328 Bound the probe length of the large-object hash index
The key hash is not seeded, so keys chosen to collide made building the table quadratic (20,000 colliding keys took 470 ms to parse). A key may now sit at most 64 slots from its home slot; if a key would sit further away, the table is dropped and the object is searched linearly. Lookups stop after the same distance.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:19 +02:00
Niels Lohmann e277ffb647 Scan json_view strings with SIMD and index large objects
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.

Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.

On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.

Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.

simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:17 +02:00
Niels Lohmann b3086f7e49 Split the dump and comparison tests off unit-json_view.cpp
The MinGW linker of the Windows clang jobs cannot link object files with more
than 32767 sections ("relocation truncated to fit: IMAGE_REL_AMD64_REL32
against `.rdata'"). unit-json_view.cpp reaches that limit as the stack
grows, so its "json_view dump" and "json_view comparison" test cases move
into unit-json_view_dump.cpp. The test generator and has_duplicate_keys()
that both files use move into json_view_test_helpers.hpp.

The new file mentions JSON_HAS_CPP_17, so it is built for C++17 like the file
it was split from.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:14 +02:00
Niels Lohmann 0e2d8975a4 Fix the clang-tidy 22 findings in json_view and its tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:13 +02:00
Niels Lohmann 9d43acd174 Document and test first-wins lookups next to dump() and ==
dump() writes every member and == resolves duplicate keys as parse()
does, whereas lookups find the first member of a duplicate key.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:13 +02:00
Niels Lohmann 8fe8c2c8d5 Hold the documents of the dump and comparison tests in names
root() of a temporary document is deleted: its views would dangle.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:11 +02:00
Niels Lohmann 90626e7899 Size the dump buffer of a small value by its own extent
The source extent of a value was read from the next node, falling back to the rest of the document when that node held a decoded string. dump() of a small value could thus allocate a buffer as large as the document. Skip a few such nodes, cap the fallback estimate, and shrink a buffer that is much larger than its output.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:10 +02:00
Niels Lohmann cbdb1fe520 Avoid raw string with escapes inside CHECK macro (MSVC C2017)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:09 +02:00
Niels Lohmann 135f0f37bd Add dump() and comparisons to json_view
Add basic_json_view::dump() and the comparison operators, and read
floats from the parser's digit layout instead of rescanning the
token.

dump(indent, indent_char, ensure_ascii, number_format) writes a
value the way ordered_json::parse(text).dump() writes it for the
same arguments: members in document order, all of them should a
key occur more than once; strings escaped by the same rules, using
the library's scanning kernels; floats written with the library's
to_chars conversion, so the output equals basic_json's byte for
byte; integers copied from the source, where they are already
canonical, except -0, which parse() reads as 0. There is no
error_handler argument, because the view only holds valid UTF-8.
number_format::source copies numbers exactly as they appear in the
source (e.g. "1.50", "1E2", "-0"), which basic_json cannot provide.
operator<< takes the indentation from the stream width, as for
basic_json. The writer walks iteratively, so nesting depth is
limited by memory only.

operator== and operator!= compare two views, or a view and a
basic_json value in either order, by the rules basic_json's
operator== uses: numbers compare by value across their types,
objects compare by their members with duplicate keys resolved as
parse() resolves them, member order matters only where the object
type keeps one, and discarded views compare as discarded basic_json
values do, including under JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON.
Nothing is materialized except single scalars.

While parsing, the view now records where the integer digits, the
fraction digits, and the exponent of a float token are, so floats
and doubles with at most 19 digits are read from that layout with
the library's decimal_to_float() instead of rescanning the token.
Both round correctly, so the values are those of parse(). get<double>(),
materialize(), dump(), and the comparisons all use it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:08 +02:00
Niels Lohmann 794fdc2b2e Fix the clang-tidy 22 findings in json_view and its tests
The same changes as on the dump branch, where they were first made, so
that this branch passes clang-tidy on its own.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:07 +02:00
Niels Lohmann f7c1b3495a Avoid GCC useless-cast warnings in the integer index tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:05 +02:00
Niels Lohmann decf6de9c7 Return the first member of duplicate keys from lookups again
Looking up the last member of a duplicate key cannot stop at a match, so
every lookup scanned the whole object (1.6 to 3.4 times slower for small
objects). operator[](key), at, find, value, contains, count, and JSON
pointer resolution return the first member again, as yyjson and simdjson
do; materialize() and get<map>() keep the last value, as parse().

The documentation says so in the feature page and on each lookup page,
and explains how to get the value parse() would give. The integer index
templates and the discarded chaining of operator[] stay.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:03 +02:00
Niels Lohmann f85c4bf8b6 Use named documents in the json_view tests that called root() on a temporary
root() of an rvalue document is deleted, because the views would dangle.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:01 +02:00
Niels Lohmann 5badf2eac5 Replace the removed operator bool of basic_json_view in the header and its tests
A view is tested with is_discarded(); the lookups in at(), value() and
contains(json_pointer) and the unit tests no longer rely on the explicit
conversion to bool.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:00 +02:00
Niels Lohmann 4ebac0263e Make view lookups last-wins and chained access safe
* Lookups (operator[], at, find, contains, count, value, JSON pointers)
  return the last member of a duplicate key, as materialize() and
  parse() keep it.
* operator[] on a discarded view returns a discarded view instead of
  throwing, so v["a"]["b"] is safe for a missing "a".
* operator[] and at() take any integer type (not only int and size_t),
  fixing ambiguous calls with unsigned, long, std::int64_t, ...
* Fix the operator[] documentation, which claimed a discarded view for
  a type mismatch where type_error.305 is thrown.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:59 +02:00
Niels Lohmann 041a561e56 Add element access, iteration, values, and JSON pointers to json_view
Give basic_json_view the read-only access functions of basic_json:
operator[] and at() with keys and indices, front()/back(), find(),
contains(), count(), begin()/end() and cbegin()/cend(), items()
with structured bindings from C++17 on, and type_name().

Exceptions have the ids and messages of the const functions of
basic_json. Where basic_json has undefined behavior the view
answers safely: operator[] with a missing key or an out-of-range
index returns a discarded view, and front()/back() of an empty
container throw invalid_iterator.214. Objects are iterated in
document order, and all members are visited; duplicate-key lookups
find the first member (as yyjson and simdjson do), while parse(),
materialize(), and the map conversions keep the last value, as
parse() does. Keys of up to 16 bytes are compared with two
overlapping loads.

Add value conversions: get<T>()/get_to() for arithmetic types,
bool, nullptr_t, strings (std::basic_string copied,
string_view_t without a copy), BasicJsonType, views, std::vector,
and maps with string keys; get_string() for the string without a
copy; number_token() for the number exactly as written in the
source; value() with keys and JSON pointers; and operator[]/at()/
contains() with JSON pointers. Everything else, including types
with from_json(), goes through materialize() of that subtree.
get<T>() of arithmetic types is inlined down to the conversion, so
reading an integer needs no call.

detail::json_pointer_access exposes a pointer's reference tokens
to code outside basic_json.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:58 +02:00
Niels Lohmann 41fd507e8f Silence a clang-tidy use-after-move finding in the json_view test
const_text is const, so std::move does not move from it; the second call
is part of the const rvalue test.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:57 +02:00
Niels Lohmann 4a4aa8cfc0 Make max_input_size a function (GCC -Wunused-const-variable)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:54 +02:00
Niels Lohmann 1d948b31e9 Check the steps of the deep-nesting test separately
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:54 +02:00
Niels Lohmann bf0a485c9d Annotate intentional patterns in unit-json_view.cpp for clang-tidy
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:52 +02:00
Niels Lohmann 5b5e0c2750 Fix clang-tidy findings in unit-json_view_builder.cpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:51 +02:00
Niels Lohmann 938e0d16ff Fuzz json_document with exact-size buffers and parse options
The fuzzer only parsed a std::string with default options. Also parse an
exact-size byte vector, which has no NUL after its last byte and takes the
bounds-checked path, and derive ignore_comments and ignore_trailing_commas
from the first input byte, comparing against json::parse and json::accept
with the same options.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:50 +02:00
Niels Lohmann a532412bb2 State the exact input size limit of json_document
The check rejects inputs of 0xFFFFFFF0 bytes or more, but the exception
message and the documentation said 4 GiB. Name the limit once
(max_input_size), and state 4 GiB minus 16 bytes in the message and the
documentation. Test the limit with a container that only claims the size.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:47 +02:00
Niels Lohmann b009db69b3 Check the node array size for overflow
reserve(n) and the growth of the index computed n * sizeof(node) without a
check, which wraps around on 32-bit targets for inputs of about 1 GiB and
allocates a too small array. Throw std::bad_alloc for a count beyond the
address space and clamp the growth step to it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:46 +02:00
Niels Lohmann 95704857ba Store node integers by halves on big-endian targets
The non-little-endian path of the node writer wrote the integer's native
word over len and next, so len got the high half there. Compose and split
the value explicitly (len is the low half, next the high half); the
little-endian path stays a plain memcpy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:46 +02:00
Niels Lohmann 6661390b23 Delete json_document::root() on temporary documents
auto v = json_document::parse(text).root() compiled and left the view
dangling. Delete the overload for rvalue documents, take a named document
in the tests, and document the lifetime rule.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:45 +02:00
Niels Lohmann 240dae052f Remove the conversion to bool from json_view
explicit operator bool meant "refers to a value", which silently differs
from what a basic_json converts to. Use !v.is_discarded() instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:44 +02:00
Niels Lohmann 4b1eddb286 Copy const rvalue strings and document the accepted inputs
json_document::parse(std::move(const_string)) failed to compile with
"no matching read_kind": only a non-const rvalue std::string can be moved
from. Treat a const rvalue as a copied byte container.

parse() does not accept everything BasicJsonType::parse() does: a FILE*
and pointers to or arrays of wide characters are rejected at compile time.
List the supported inputs instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:43 +02:00
Niels Lohmann 45ce9b51f5 Reject integer lengths in json_document::parse
parse(ptr, len) compiled: len converted to allow_exceptions, and ptr was
read as a C string, past the end of a buffer without a terminating NUL.
Delete the overloads of parse, parse_copy, accept, and read that take an
integer other than bool where the flags are expected.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:42 +02:00
Niels Lohmann 1d59a0f4a6 Add json_document and json_view: node index, parser, and document
Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).

A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.

The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.

basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.

detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.

A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:40 +02:00
Niels Lohmann 9783b2db58 Silence MSVC C4127 for a platform-constant check in the to_chars test
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:38 +02:00
Niels Lohmann 26383cfcb8 Select the Zmij conversion by a trait, not by the double overload
The non-template overloads for double asserted binary64 doubles wherever json.hpp was included, so the library no longer compiled where double is not IEEE 754 binary64 (AVR, -fshort-double). A trait now picks Zmij for any binary64 type, including a long double of that format (MSVC, Apple Arm), and Grisu2 for the others. Remove the unused write_short_decimal, powers_of_ten_16, zmij::decimal, zmij::to_decimal, and shortest_digits(double).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:37 +02:00
Niels Lohmann bf59f687c3 Fix CI findings in the Zmij writer and its tests
- bit_ops.hpp: include macro_scope.hpp (JSON_HEDLEY_ALWAYS_INLINE); fixes IWYU
- to_chars.hpp: C4100 for the unused parameter in release builds, clang-tidy
  sign comparison, cpplint runtime/int, GCC -Wstrict-overflow (unsigned abs)
- unit-to_chars.cpp: parse with the library instead of strtod (MinGW's strtod
  rounds some 16/17 digit inputs wrongly); no floating-point std::to_chars
  with icpc

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:36 +02:00
Niels Lohmann 2dc6471968 Write doubles with the shortest digits (Żmij), digits in registers
Write doubles with the conversion of Zmij by Victor Zverovich (MIT),
ported to C++11 (detail/conversions/zmij.hpp). It finds the
shortest decimal that reads back as the same double, and the
closest one if there are several. Grisu2, used until now, is fast
but not always shortest: it sometimes writes a 17th digit where 16
suffice, or a last digit that is not the closest. The layout is
unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2.

Digits are converted eight at a time with the BCD conversion of
Xiang JunBo, as in Zmij, and written with one byte swap per eight
digits and fixed-size moves instead of per-digit loops. Leading and
trailing zeros are counted from those bytes. to_chars() uses a
local buffer when the caller's is shorter than the 41 bytes this
may write. The powers of ten come from the number-parsing table,
adjusted where it holds values rounded up, and extended with Zmij's
compressed tables beyond 10^308.

write_shortest() converts its 16 digits in one vector register
(SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the
decimal point inside the register, avoiding a store-forwarding
stall that cost about 25% of the time to write a double. dump()
writes floats and integers straight into the serializer's write
buffer instead of copying them from a member buffer, and small
integers eight digits at a time. read_eight_bytes() and
parse_eight_digits() are marked always-inline, which GCC had been
calling out of line in the number-parsing loops.

Of one million random doubles, about 0.14% are now written with
different digits, always to a value that still reads back as the
same double.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:35 +02:00
Niels Lohmann 80bfa0ac73 Fix clang-tidy 22 findings in unit-class_lexer.cpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:32 +02:00
Niels Lohmann ba6f158378 Load string scan words with memcpy and use MSVC bit-scan intrinsics
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:31 +02:00
Niels Lohmann 6c82661a3d Fix CI warnings in own float parser
- pow5_table.hpp: pow5_128_largest_power was unused in this branch's
  own code (GCC -Werror=unused-const-variable); tie it to the table
  size with a static_assert instead of removing it, since a later
  branch in the stack (json-view/23-zmij) uses it.
- number_parse.hpp: rename the local variable `copy` to `buffer` to
  satisfy cpplint's build/include_what_you_use check.
- unit-class_lexer.cpp: extend the NOLINT list on the seeded mt19937
  with bugprone-random-generator-seed, and parenthesize
  `8 * sizeof(Bits) - 1` for clang-tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:30 +02:00
Niels Lohmann 422684995c Speed up the lexer: own float parser, string scan, and \u table
Give the library its own correctly rounded float converter for
binary32 and binary64 (IEEE 754), and speed up the lexer's string
and escape scanning.

The converter splits a number token into sign, significand, and
decimal exponent, then tries Clinger's fast path, then a templated
Eisel-Lemire step, and falls back to an exact big-integer digit
comparison for tokens with more than 19 significant digits whose two
candidate values round differently. This replaces std::from_chars
and strtod/strtof for both formats, so parsed values no longer
depend on the C/C++ library or the current locale. The strtold
fallback kept for other long double formats (x87, binary128) now
also copies a multi-byte decimal point correctly, fixing #5660.
eisel_lemire() and decimal_to_float() are always inlined so callers
keep the whole conversion in their hot loop.

The string-scanning kernels in string_scan.hpp find a stop byte with
the trailing-zero count of the SWAR mask instead of a byte loop, and
scalar_string_bulk_run() validates a run of multi-byte UTF-8
sequences one after another instead of re-searching after each one.

get_codepoint() decodes a contiguous \uXXXX escape with one table
lookup per byte instead of four range-checked get() calls; the
streaming path and all error positions are unchanged.

Adds 508 generated hard float-parsing cases with expected binary32
and binary64 bits, and kernel-comparison tests for the string scans
and the escape table against byte-by-byte references.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:29 +02:00
Niels Lohmann c59ad95252 Expect type_error.321 for deep discarded values in the BON8 writer test (#5807)
#5806 added the test while #5802 made to_bon8 reject discarded values; both
merged, so the test failed on develop.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:16 +02:00
Niels LohmannandClaude Opus 5.5 b38a2c9177 Stop the BON8 writer overflowing the stack on deep values (#5806)
* Stop the BON8 writer overflowing the stack on deep values

to_bon8() recursed once per nesting level, so a value the iterative
BON8 reader accepts (e.g. ~24k nested one-element arrays) crashed on
the way back out. #5781 bounded the CBOR, MessagePack, and UBJSON/BJData
writers, but BON8 was merged before it and was not covered.

Apply the same scheme: recurse for the first recursion_depth_limit()
levels, then finish the value with write_bon8_iterative, which keeps
the open containers on a heap stack (reusing binary_container_frame)
and writes the 0xFE closer when it leaves a container with more than
four elements. The output is byte-for-byte unchanged.

Fixes https://issues.oss-fuzz.com/issues/572238015

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge the array and object branches of the BON8 iterative writer

write_bon8_iterative and write_bon8_value_or_push handled arrays and
objects in separate branches that repeated the end-of-container check,
the 0xFE closer and the marker computation. Share those parts and branch
only where arrays and objects really differ. The output is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hz7VJi1FTKr6gpseLErbbS
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-11 08:24:35 +02:00