Commit Graph
1161 Commits
Author SHA1 Message Date
Niels Lohmann 5badf2eac5 Replace the removed operator bool of basic_json_view in the header and its tests
A view is tested with is_discarded(); the lookups in at(), value() and
contains(json_pointer) and the unit tests no longer rely on the explicit
conversion to bool.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:28:00 +02:00
Niels Lohmann 4ebac0263e Make view lookups last-wins and chained access safe
* Lookups (operator[], at, find, contains, count, value, JSON pointers)
  return the last member of a duplicate key, as materialize() and
  parse() keep it.
* operator[] on a discarded view returns a discarded view instead of
  throwing, so v["a"]["b"] is safe for a missing "a".
* operator[] and at() take any integer type (not only int and size_t),
  fixing ambiguous calls with unsigned, long, std::int64_t, ...
* Fix the operator[] documentation, which claimed a discarded view for
  a type mismatch where type_error.305 is thrown.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:59 +02:00
Niels Lohmann 041a561e56 Add element access, iteration, values, and JSON pointers to json_view
Give basic_json_view the read-only access functions of basic_json:
operator[] and at() with keys and indices, front()/back(), find(),
contains(), count(), begin()/end() and cbegin()/cend(), items()
with structured bindings from C++17 on, and type_name().

Exceptions have the ids and messages of the const functions of
basic_json. Where basic_json has undefined behavior the view
answers safely: operator[] with a missing key or an out-of-range
index returns a discarded view, and front()/back() of an empty
container throw invalid_iterator.214. Objects are iterated in
document order, and all members are visited; duplicate-key lookups
find the first member (as yyjson and simdjson do), while parse(),
materialize(), and the map conversions keep the last value, as
parse() does. Keys of up to 16 bytes are compared with two
overlapping loads.

Add value conversions: get<T>()/get_to() for arithmetic types,
bool, nullptr_t, strings (std::basic_string copied,
string_view_t without a copy), BasicJsonType, views, std::vector,
and maps with string keys; get_string() for the string without a
copy; number_token() for the number exactly as written in the
source; value() with keys and JSON pointers; and operator[]/at()/
contains() with JSON pointers. Everything else, including types
with from_json(), goes through materialize() of that subtree.
get<T>() of arithmetic types is inlined down to the conversion, so
reading an integer needs no call.

detail::json_pointer_access exposes a pointer's reference tokens
to code outside basic_json.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:58 +02:00
Niels Lohmann 10d36d3af2 Read the node array base after emit() in the view builder's open()
emit() moves the node array when it grows, and the subtraction read base
in the same expression, so the order was unspecified. MSVC Release builds
without forced inlining read the old base; the container index then
pointed outside the array and close() wrote out of bounds.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:56 +02:00
Niels Lohmann eeae3e9121 Make json_view.hpp pass include-what-you-use
ci_single_binaries runs IWYU with --error on every header. For json_view.hpp
it suggested adding <array> (std::array is used), the headers that json.hpp
already provides (abi_config, abi_macros, input_adapters, json_pointer,
cpp_future, string_concat, value_t, json_fwd), and <version> for
std::nullptr_t, and removing <cstddef>.

- include <array>
- keep <cstddef> (nullptr_t, size_t; IWYU attributes them to <version> and <cstring>)
- tell IWYU not to suggest the headers that json.hpp provides: the amalgamated
  json_view.hpp only includes json.hpp, so including them here would duplicate
  their definitions
- export json.hpp, keep macro_unscope.hpp, and drop the unused forward declaration of the document class

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:55 +02:00
Niels Lohmann 4a4aa8cfc0 Make max_input_size a function (GCC -Wunused-const-variable)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:54 +02:00
Niels Lohmann c0dbdec1ca Suppress Flawfinder's read() finding on the deleted overload
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:53 +02:00
Niels Lohmann 0c30a9b6de Give document_data a user-provided constructor
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:50 +02:00
Niels Lohmann a5a6ff6b7c Align the banner comment of json_view.hpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:49 +02:00
Niels Lohmann f0168aaa5f Remove the unused legacy_discarded_value_comparison from abi_config
Nothing reads it: the view does not compare values.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:48 +02:00
Niels Lohmann a532412bb2 State the exact input size limit of json_document
The check rejects inputs of 0xFFFFFFF0 bytes or more, but the exception
message and the documentation said 4 GiB. Name the limit once
(max_input_size), and state 4 GiB minus 16 bytes in the message and the
documentation. Test the limit with a container that only claims the size.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:47 +02:00
Niels Lohmann b009db69b3 Check the node array size for overflow
reserve(n) and the growth of the index computed n * sizeof(node) without a
check, which wraps around on 32-bit targets for inputs of about 1 GiB and
allocates a too small array. Throw std::bad_alloc for a count beyond the
address space and clamp the growth step to it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:46 +02:00
Niels Lohmann 95704857ba Store node integers by halves on big-endian targets
The non-little-endian path of the node writer wrote the integer's native
word over len and next, so len got the high half there. Compose and split
the value explicitly (len is the low half, next the high half); the
little-endian path stays a plain memcpy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:46 +02:00
Niels Lohmann 6661390b23 Delete json_document::root() on temporary documents
auto v = json_document::parse(text).root() compiled and left the view
dangling. Delete the overload for rvalue documents, take a named document
in the tests, and document the lifetime rule.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:45 +02:00
Niels Lohmann 240dae052f Remove the conversion to bool from json_view
explicit operator bool meant "refers to a value", which silently differs
from what a basic_json converts to. Use !v.is_discarded() instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:44 +02:00
Niels Lohmann 4b1eddb286 Copy const rvalue strings and document the accepted inputs
json_document::parse(std::move(const_string)) failed to compile with
"no matching read_kind": only a non-const rvalue std::string can be moved
from. Treat a const rvalue as a copied byte container.

parse() does not accept everything BasicJsonType::parse() does: a FILE*
and pointers to or arrays of wide characters are rejected at compile time.
List the supported inputs instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:43 +02:00
Niels Lohmann 45ce9b51f5 Reject integer lengths in json_document::parse
parse(ptr, len) compiled: len converted to allow_exceptions, and ptr was
read as a C string, past the end of a buffer without a terminating NUL.
Delete the overloads of parse, parse_copy, accept, and read that take an
integer other than bool where the flags are expected.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:42 +02:00
Niels Lohmann 7bb286df21 Fix document_data for old Clang (3.4-3.6)
Drop the empty braced NSDMIs of the std::string members: old Clang
rejects the defaulted constructor when it is used by a member
initializer before the end of the class definition.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:41 +02:00
Niels Lohmann 1d59a0f4a6 Add json_document and json_view: node index, parser, and document
Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).

A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.

The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.

basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.

detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.

A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:40 +02:00
Niels Lohmann 1f5c61ca7d Credit Zmij's authors in to_chars.hpp, the README, and the license page
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:37 +02:00
Niels Lohmann 26383cfcb8 Select the Zmij conversion by a trait, not by the double overload
The non-template overloads for double asserted binary64 doubles wherever json.hpp was included, so the library no longer compiled where double is not IEEE 754 binary64 (AVR, -fshort-double). A trait now picks Zmij for any binary64 type, including a long double of that format (MSVC, Apple Arm), and Grisu2 for the others. Remove the unused write_short_decimal, powers_of_ten_16, zmij::decimal, zmij::to_decimal, and shortest_digits(double).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:37 +02:00
Niels Lohmann bf59f687c3 Fix CI findings in the Zmij writer and its tests
- bit_ops.hpp: include macro_scope.hpp (JSON_HEDLEY_ALWAYS_INLINE); fixes IWYU
- to_chars.hpp: C4100 for the unused parameter in release builds, clang-tidy
  sign comparison, cpplint runtime/int, GCC -Wstrict-overflow (unsigned abs)
- unit-to_chars.cpp: parse with the library instead of strtod (MinGW's strtod
  rounds some 16/17 digit inputs wrongly); no floating-point std::to_chars
  with icpc

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:36 +02:00
Niels Lohmann 2dc6471968 Write doubles with the shortest digits (Żmij), digits in registers
Write doubles with the conversion of Zmij by Victor Zverovich (MIT),
ported to C++11 (detail/conversions/zmij.hpp). It finds the
shortest decimal that reads back as the same double, and the
closest one if there are several. Grisu2, used until now, is fast
but not always shortest: it sometimes writes a 17th digit where 16
suffice, or a last digit that is not the closest. The layout is
unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2.

Digits are converted eight at a time with the BCD conversion of
Xiang JunBo, as in Zmij, and written with one byte swap per eight
digits and fixed-size moves instead of per-digit loops. Leading and
trailing zeros are counted from those bytes. to_chars() uses a
local buffer when the caller's is shorter than the 41 bytes this
may write. The powers of ten come from the number-parsing table,
adjusted where it holds values rounded up, and extended with Zmij's
compressed tables beyond 10^308.

write_shortest() converts its 16 digits in one vector register
(SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the
decimal point inside the register, avoiding a store-forwarding
stall that cost about 25% of the time to write a double. dump()
writes floats and integers straight into the serializer's write
buffer instead of copying them from a member buffer, and small
integers eight digits at a time. read_eight_bytes() and
parse_eight_digits() are marked always-inline, which GCC had been
calling out of line in the number-parsing loops.

Of one million random doubles, about 0.14% are now written with
different digits, always to a value that still reads back as the
same double.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:35 +02:00
Niels Lohmann d99a485ce7 Mark the unsigned long of the MSVC bit scans for cpplint
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:33 +02:00
Niels Lohmann ba6f158378 Load string scan words with memcpy and use MSVC bit-scan intrinsics
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:31 +02:00
Niels Lohmann 6c82661a3d Fix CI warnings in own float parser
- pow5_table.hpp: pow5_128_largest_power was unused in this branch's
  own code (GCC -Werror=unused-const-variable); tie it to the table
  size with a static_assert instead of removing it, since a later
  branch in the stack (json-view/23-zmij) uses it.
- number_parse.hpp: rename the local variable `copy` to `buffer` to
  satisfy cpplint's build/include_what_you_use check.
- unit-class_lexer.cpp: extend the NOLINT list on the seeded mt19937
  with bugprone-random-generator-seed, and parenthesize
  `8 * sizeof(Bits) - 1` for clang-tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:30 +02:00
Niels Lohmann 422684995c Speed up the lexer: own float parser, string scan, and \u table
Give the library its own correctly rounded float converter for
binary32 and binary64 (IEEE 754), and speed up the lexer's string
and escape scanning.

The converter splits a number token into sign, significand, and
decimal exponent, then tries Clinger's fast path, then a templated
Eisel-Lemire step, and falls back to an exact big-integer digit
comparison for tokens with more than 19 significant digits whose two
candidate values round differently. This replaces std::from_chars
and strtod/strtof for both formats, so parsed values no longer
depend on the C/C++ library or the current locale. The strtold
fallback kept for other long double formats (x87, binary128) now
also copies a multi-byte decimal point correctly, fixing #5660.
eisel_lemire() and decimal_to_float() are always inlined so callers
keep the whole conversion in their hot loop.

The string-scanning kernels in string_scan.hpp find a stop byte with
the trailing-zero count of the SWAR mask instead of a byte loop, and
scalar_string_bulk_run() validates a run of multi-byte UTF-8
sequences one after another instead of re-searching after each one.

get_codepoint() decodes a contiguous \uXXXX escape with one table
lookup per byte instead of four range-checked get() calls; the
streaming path and all error positions are unchanged.

Adds 508 generated hard float-parsing cases with expected binary32
and binary64 bits, and kernel-comparison tests for the string scans
and the escape table against byte-by-byte references.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 10:27:29 +02:00
Niels LohmannandClaude Opus 5.5 b38a2c9177 Stop the BON8 writer overflowing the stack on deep values (#5806)
* Stop the BON8 writer overflowing the stack on deep values

to_bon8() recursed once per nesting level, so a value the iterative
BON8 reader accepts (e.g. ~24k nested one-element arrays) crashed on
the way back out. #5781 bounded the CBOR, MessagePack, and UBJSON/BJData
writers, but BON8 was merged before it and was not covered.

Apply the same scheme: recurse for the first recursion_depth_limit()
levels, then finish the value with write_bon8_iterative, which keeps
the open containers on a heap stack (reusing binary_container_frame)
and writes the 0xFE closer when it leaves a container with more than
four elements. The output is byte-for-byte unchanged.

Fixes https://issues.oss-fuzz.com/issues/572238015

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge the array and object branches of the BON8 iterative writer

write_bon8_iterative and write_bon8_value_or_push handled arrays and
objects in separate branches that repeated the end-of-container check,
the 0xFE closer and the marker computation. Share those parts and branch
only where arrays and objects really differ. The output is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hz7VJi1FTKr6gpseLErbbS
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-11 08:24:35 +02:00
Niels Lohmann eb8899a29c Restore v3.12.0 support for custom object key types (#5795)
* Restore v3.12.0 support for custom object key types

Custom object_t types whose key_type is not string_t compiled with
v3.12.0 for several APIs that unreleased changes broke:

- to_bson failed for every custom key type (#5553 kept a const string_t*
  to the key); the nested entry's header is now written where the entry
  is found.
- Copying deep values (and parse, merge_patch, update, insert) required
  operator== on keys (#5389); keys without one are now paired via find().
- to_cbor/to_msgpack required an implicit conversion to string_t (#5746,
  #5328); keys without one go through a temporary basic_json again.
- at() required a conversion to string_t for its error message (#5727);
  other keys are passed to concat() unchanged again.

The new unit-custom-object-key-type.cpp covers five key types with
different capabilities.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use the with_object_t alias for the custom object key test types

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid floating-point equality in custom key type test

GCC with -Werror=float-equal rejects comparing the double value with ==.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Write the head of nested BSON elements in one helper

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Suppress bugprone-return-const-ref-from-parameter in key_for_message

The reference is only passed to concat() within the full-expression that
holds the key, like the similar helpers in binary_writer.hpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the value of the custom test key types private

clang-tidy (cppcoreguidelines-non-private-member-variables-in-classes)
rejects the protected member; the derived key types use a protected
accessor instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Pass keys with data() and size() unchanged into the at() miss message

key_for_message() converted every key that string_t can be constructed
from, so a miss on a string_t or string_view key copied it before
concat() copied it again. Keys that concat() can append through data()
and size() are now passed through; only other keys (string literals,
key types that just convert to string_t) are converted.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 08:24:31 +02:00
Niels Lohmann 74a03dec84 Throw type_error.321 for discarded values in to_bon8 (#5802)
write_bon8_value silently skipped discarded values, but the array/object
count marker still counted them, so [1, discarded] produced 82 91: a marker
announcing two elements followed by one. Throw type_error.321 like the other
binary writers do, at any nesting level.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 08:22:00 +02:00
Niels Lohmann 2cca04ae6f Report one-character non-numeric array indices like longer ones (#5800)
* Report one-character non-numeric array indices like longer ones

A JSON pointer reference token that is not a number but has only one
character (e.g. "/a/x") was reported as out_of_range.404 ("unresolved
reference token"), because the "is not a number" check only ran for
tokens longer than one character; "/a/xy" got parse_error.109. Both now
throw parse_error.109. "-" and the empty token are still reported as
out_of_range.404. As a consequence, value(json_pointer, default) on an
array now throws for "/x" as it already did for "/xy".

Also document why ordered_map::erase's destroy/placement-new loop on
pair<const Key, T> is kept despite [basic.life]/8 before C++20.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document the one-character array index change

Add 3.13.0 version-history entries to at, operator[], value, patch,
patch_inplace, and unflatten, and describe in exceptions.md which array
indices throw parse_error.109 and which out_of_range.404.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 08:21:41 +02:00
Niels Lohmann 1649d9eda4 Use and extend the detail helpers to remove duplicated code (#5783)
* Use and extend the detail helpers to remove duplicated code

Library:
- binary_reader: format bytes with hex_byte() instead of snprintf
- add throw_type_must_be() for the 20 copies of type_error.302
- binary_reader: add last_byte_error()/unexpected_byte() for the 32
  "parse error at the last read byte" sites (replaces bon8_error)
- json_sax: add check_container_size() for out_of_range.408 and
  diagnostic_positions::set_container_start/_end()
- json_pointer: add throw_no_parent() (405) and throw_unresolved() (404)

Tests:
- unit-class_parser uses the shared utils::SaxCountdown
- move SaxEventLogger (and its ExitAfter* variants) from unit-class_parser
  and unit-deserialization into the new tests/src/test_sax.hpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Factor out more repeated error paths and test boilerplate

Library:
- add throw_cannot_use_with() for the 34 copies of type_error.304-312
  "cannot use X with Y"
- iter_impl: add throw_cannot_get_value() (invalid_iterator.214)
- parser: add syntax_error() for the 12 parse_error.101 sites
- ordered_map: share the four at() bodies via at_impl()
- json_sax_dom_callback_parser: add pop_container() for end_object()
  and end_array()
- binary_reader: build the two UBJSON/BJData length-type messages with
  concat() and last_byte_error()

Tests:
- move same_value(), the NDEBUG guard, and step 0 (parse without
  exceptions) of the seven fuzzer drivers into tests/src/fuzzer_common.hpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the exception ids and address review comments

Add detail::exception_id, a scoped enum with one named enumerator per
documented exception id, and use it for every id in the library. The
create() functions get an overload for it; the int overloads stay for
user code.

Following the review of #5783: add binary_reader::invalid_byte() and
length_type_error(), basic_json::throw_subscript_wrong_type(), move the
fuzzer includes and the using-declaration into fuzzer_common.hpp, and
rename test_sax.hpp to sax_event_loggers.hpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Silence -Wweak-vtables for the SAX event loggers

The loggers moved from anonymous namespaces in the test files into
sax_event_loggers.hpp, so clang now warns that their vtables are
emitted in every translation unit.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Silence MSVC 2015 C4100 in parser::syntax_error for static SAX::parse_error

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Inline invalid_byte() into its call sites

The wrapper only fixed the message string of unexpected_byte(), which is
the same kind of per-argument helper that was declined for
throw_type_must_be(). Call unexpected_byte("invalid byte", ...) directly.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-11 08:21:13 +02:00
Bas van Dijkandbvandijk 84ee12a67c Make UTF8_ACCEPT/UTF8_REJECT inline variables to avoid TU-local exposure in modules (#5797)
Signed-off-by: bvandijk <bas.van.dijk@cern.ch>
Co-authored-by: bvandijk <bas.van.dijk@cern.ch>
2026-10-09 23:38:42 +02:00
Niels Lohmann 2913e96433 Unflatten in time and memory linear in the pointer depth (#5793)
* Unflatten in time and memory linear in the pointer depth

#5443 made unflatten() decide between arrays and objects independently
of the iteration order by collecting the pointer prefixes that have a
reference token 0 below them in a std::set<std::vector<string_t>>.
Every such prefix was stored as a copy of all its reference tokens, and
get_and_create() compared whole prefix vectors at every step, so
unflattening a pointer of depth d took time and memory quadratic in d:
a 10,000-level array pointer took 18 s and 1.3 GB, a 100,000-level one
did not finish.

The prefixes are now numbered nodes of a tree, so each is stored once
and get_and_create() follows the tree token by token. The result is
unchanged, including its independence of the iteration order.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Initialize prefix_tree members to satisfy -Weffc++

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move prefix_tree setup and child insertion into member functions

The constructor now creates the root node, add_child() inserts a
reference token below a prefix and returns the child's number, and
find_child() looks one up for get_and_create().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 17:57:31 +02:00
Niels Lohmann 44e8597701 Flatten deeply nested values without recursing per nesting level (#5792)
* Flatten deeply nested values without recursing per nesting level

json_pointer::flatten() called itself once per nesting level, so
flatten() on a value nested deeply enough exhausted the call stack.
#5547 and #5548 fixed merge_patch() and diff() from #5393, but flatten()
was left out.

flatten() now walks the value with an explicit stack and keeps the path
in one buffer that grows and shrinks with it. It has a single code path
and no depth limit: the old version built a new path string per child,
so the iterative one is no slower on shallow values and much faster on
deep ones. The output, including the order of an ordered_json result,
is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Construct flatten frames in place

Give the frame a constructor so both call sites can use emplace_back, as
suggested in the review; index starts at 0 for every frame.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 17:06:41 +02:00
Niels Lohmann 374dfe4f0f Keep converted object keys alive while writing UBJSON and BJData (#5791)
* Keep converted object keys alive while writing UBJSON and BJData

Since #5746, write_ubjson and write_ubjson_iterative pass each object key
to sanitize_utf8_for_write and keep the returned reference. When
object_t::key_type is not string_t but converts to it, the argument is a
temporary that is destroyed at the end of the statement, and the
function returns a reference to it in every case but a sanitized copy,
so the key bytes are read from a dead object (AddressSanitizer:
stack-use-after-scope). Default json and ordered_json are unaffected.

Bind the key to a named object_key_string_t first: a reference when
key_type is string_t, so no copy is added there, and a converted copy
otherwise. A deleted overload of sanitize_utf8_for_write for anything
other than string_t turns a recurrence into a compile error.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Suppress -Wunused-member-function for the converting_key test type

converting_key::data() is only called when JSON_DIAGNOSTICS is enabled.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy findings in the UBJSON/BJData converted-key fix

Suppress hicpp/modernize-use-equals-delete on the deleted
sanitize_utf8_for_write overload: it guards a private helper and must stay
private. Replace the C-style array in the new test with std::array.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-09 17:05:36 +02:00
Alex Prabhat Bara 69a0c1b82c Avoid allocating temporary basic_json for cbor and msgpack object keys (#5328)
* avoid allocating temporary basic_json for CBOR and MessagePack object keys

Signed-off-by: alexprabhat99 <alexpbara@gmail.com>

* add size() to the custom object key test type

UBJSON and BJData access object keys through size() and c_str()
directly, so the key type now provides both and the comment says why.

Signed-off-by: alexprabhat99 <alexpbara@gmail.com>

* address review: drop key size()/c_str(), test keys below the depth limit

Nothing in the library calls size() or c_str() on an object key, so the
test key type only keeps data(), which JSON_DIAGNOSTICS needs.

The CBOR and MessagePack custom key tests now also nest objects deeper
than detail::recursion_depth_limit(), so keys written by
write_cbor_iterative and write_msgpack_iterative are covered as well.

Signed-off-by: alexprabhat99 <alexpbara@gmail.com>

---------

Signed-off-by: alexprabhat99 <alexpbara@gmail.com>
2026-10-09 13:22:07 +02:00
Suyog Verma a269794db7 Use MSVC intrinsics for full multiplication (#5782)
* Use MSVC intrinsics for full multiplication

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

* Fix formatting in unit-class_lexer

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

* Address review feedback

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

---------

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>
2026-10-08 08:49:32 +02:00
Niels LohmannandAfonso Januário a5f5d3059b Make std::hash<basic_json> consistent with operator== for numbers (#5772)
* Make std::hash<basic_json> consistent with operator== for numbers

operator== converts between number_integer, number_unsigned, and
number_float before comparing, so json(0), json(0U), and json(0.0)
all compare equal. hash() folded the specific value_t into the
result for each of the three numeric cases, giving each a distinct
hash and breaking the standard Hash requirement that a == b implies
hash(a) == hash(b). A std::unordered_set could therefore hold all
three as separate elements even though they compare equal.

hash() now treats all three numeric variants the same way: it
converts the value to number_float_t and combines it with a single
shared type tag, so any two numbers operator== considers equal hash
identically regardless of which internal type actually holds them.

Updated the accompanying test to check this consistency directly
(including via an actual unordered_set) instead of asserting that 0,
0U, and 0.0 hash differently, since that assumption was the bug.
Also corrected the function's own doc comment and the std::hash API
docs, which described the old behavior as intended.

Fixes #5400

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove now-unused number_integer_t/number_unsigned_t typedefs in hash()

Merging the three numeric branches into one that only reads
number_float_t left these two aliases unused, which several CI
configurations treat as a build error under -Wunused-local-typedefs.

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark the unordered_set in the hash regression test const

clang-tidy's misc-const-correctness check flagged it: the set is
never mutated after construction, only read via size().

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Normalize -0.0 in number hashes and test range ends

operator== compares numbers exactly since #5459, so equal numbers
share one value and convert to the same number_float_t. Update the
comment accordingly, map -0.0 to 0.0 before hashing (std::hash need
not do that), and test -0.0 and the ends of the integer ranges. Show
hash(0.0) in the docs example and note the change in the version
history.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Clarify hash documentation after review

- Say "may hash differently" for null, false, and numbers, since a
  collision across types is possible.
- Name the storage types (signed integer, unsigned integer,
  floating-point number) instead of example literals.
- Explain that the hash survives converting an integer to
  number_float_t but not the lossy conversion back, and that unequal
  numbers may share a hash.
- State that the example hash values are illustrative only and vary by
  platform, compiler, compiler version, and library version.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Afonso Januário <afonso-januario@hotmail.com>
2026-10-07 19:18:51 +02:00
43fe8928e1 Stop binary writers overflowing the stack on deep values (#5781)
* Stop binary writers overflowing the stack on deep values

to_cbor, to_msgpack, and to_ubjson recurse once per nesting level.
The parser is iterative, so a value the library accepts can crash on
the way back out.

Keep the existing recursive path for the first 128 levels and finish
anything deeper on a heap stack. Output is unchanged. BSON is left
alone because its extra size walk is a separate change.

Rebased onto the value-type output sink. The heap frames now initialize
every member, which is what -Weffc++ was rejecting.

See #5392.

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
(cherry picked from commit cf65ac438f)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Redesign the iterative binary writers around a shared recursion depth limit

Address the open review on the non-recursive CBOR/MessagePack/UBJSON/BJData
writers (#5518):

- Delete the CBOR array/object prefix helpers; both the recursive and
  iterative paths call write_cbor_head(), which already existed on develop.
- MessagePack: share one write_msgpack_array_prefix()/write_msgpack_object_prefix()
  helper per container kind between the recursive and iterative paths, both
  going through to_msgpack_length() so an over-long container throws
  out_of_range.412 identically either way.
- Reuse detail::recursion_depth_limit() instead of a separate constant, the
  same bound serializer::dump() and write_bson_document() already use.
- Redesign the frames after bson_frame/dump_frame: only a container with
  elements is ever pushed, its header is written at the point it is pushed,
  and the iterator is set in the frame's constructor instead of a
  default-then-assign two-step with a since-removed "started" flag. The
  UBJSON frame keeps only the value pointer, the per-element prefix_required
  flag, and the iterator; write_closer and is_object are no longer stored,
  since the former is always !use_count (use_count is constant for the whole
  document) and the latter follows from value->is_object().
- Factor the BJData ND-array shape check into is_bjdata_ndarray(), used by
  both the recursive object case and the iterative pushing logic.
- Give the frame classes the GCC -Weffc++ treatment already used for
  diff_frame: a noexcept converting constructor plus the five special members
  defaulted with no explicit noexcept.
- Fix two @ref self-references in write_cbor/write_msgpack/write_ubjson's own
  doc comments to point at the public to_cbor/to_msgpack/to_ubjson/to_bjdata
  API instead.
- The iterative object-key write for CBOR/MessagePack now runs the same
  strict-mode check_utf8() against the parent object as diagnostics context
  that the recursive path already ran, so the two paths raise identical
  diagnostics across the switch-over.
- Rewrite the tests: round trips instead of a bare size check, byte-exact
  comparisons against the recursive output at depths around the bound, a
  deep object and a BJData ND-array past the bound, a deep discarded value
  (type_error.321), and the OSS-Fuzz 566583014 CBOR/MessagePack regression.

BSON is unaffected by this change; it already walks its documents
iteratively and is covered separately by #5553.

Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: ayush-singh-0601 <singhayush062006@gmail.com>
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
2026-10-07 19:18:20 +02:00
Niels Lohmann 069ace74af Round-trip BJData ND-array annotations exactly (single precision, key order) (#5707)
Squashed onto develop from:
- Round-trip BJData ND-array annotations exactly (single precision, key order)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 19:17:18 +02:00
Niels Lohmann 367336c83d Fix CI on develop after #5585 (#5779)
* Fix CI on develop after #5585

- test-diagnostics-optimized: -O3 makes GCC's -Winline and
  -Wsuggest-attribute=pure/const warnings fire with the ci_test_gcc flag
  set; turn them off for this test.
- test-diagnostics-optimized: suppress Clang's -Wexit-time-destructors for
  the static table in to_json.
- Infer: raise pulse-max-disjuncts from 20 to 40. With the default,
  Pulse loses the stored type in basic_json::replace_value() and reports
  false null dereferences of get_ptr() results in unit-pointer_access.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Ignore Infer's false USE_AFTER_DELETE in ordered_map::erase

Infer's std::string model keeps the buffer of a moved-from string, so the
destroy-and-reconstruct loop in erase(first, last) looks like it destroys a
buffer twice.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark throw_on_discarded()'s parameters as used without exceptions

With JSON_NOEXCEPTION, JSON_THROW expands to std::abort(), so Clang's
-Wunused-parameter breaks test-disabled_exceptions (since #5761).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Skip the span_input_adapter sax_parse checks with deleted deprecated functions

The #5676 regression test (#5740) calls the deprecated
sax_parse(span_input_adapter&&, ...), which JSON_DELETE_DEPRECATED_FUNCTIONS
deletes, so ci_test_delete_deprecated_functions failed to build.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fall back to the first entry in test-diagnostics-optimized's to_json

clang-tidy (clang-analyzer-security.ArrayBound) flagged it->second for a
value not in the table. Use the same fallback as
NLOHMANN_JSON_SERIALIZE_ENUM; the test still fails with -Werror=array-bounds
on the headers from before #5585.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy findings in tests from #5762 and #5774

- unit-regression2.cpp (#5762): const/auto for the destroy() test values;
  NOLINT the intended copy in check_destroy_edge_case().
- unit-serialization.cpp (#5774): build the expected strings with += instead
  of chained operator+ (performance-inefficient-string-concatenation).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:37:19 +02:00
Niels Lohmann 30e542e52a Declare the namespace-scope constants in to_chars.hpp inline (#5776)
MSVC warns with C5260 that kAlpha and kGamma have internal linkage when
the header is used as a header unit or through a module. JSON_INLINE_VARIABLE
makes them inline variables from C++17 on.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 09:03:18 +02:00
Niels Lohmann 1c754cfe31 Copy a pair-shaped array value under JSON_BRACE_INIT_COPY_SEMANTICS (#5701)
With JSON_BRACE_INIT_COPY_SEMANTICS enabled, single-element brace
initialization from a JSON value decided whether to copy the value or
build an object by inspecting the value's runtime shape: a two-element
array whose first element is a string, such as ["key", 42], was turned
into an object instead of being copied. This made the behavior depend
on the element's content, and it did not distinguish an existing value
of this shape from a nested braced pair written in the source, such as
the inner {"key", "value"} of {{"key", "value"}}.

json_ref now records whether it was constructed from a braced list
(true only for the std::initializer_list<json_ref> constructor used
for nested braced lists) or from a value. The initializer-list
constructor uses this to copy or move a single non-braced-list element
before deciding whether the list describes an object, so a JSON value
is always copied regardless of its shape, while a braced pair written
in the source still creates an object.

Fixes #5662.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 07:39:05 +02:00
Niels LohmannandJoseph.Demarest ff6f3d7d4a Reject nested indefinite-length CBOR string chunks (#5766)
* fix(cbor): reject nested indefinite string chunks

Signed-off-by: Joseph.Demarest <joseph@demarest.dev>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments on nested indefinite-length CBOR strings

Rename is_chunk to inside_indefinite, update the stale test section
names, and use lowercase comments like the surrounding code.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Joseph.Demarest <joseph@demarest.dev>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Joseph.Demarest <joseph@demarest.dev>
2026-10-07 07:38:52 +02:00
Mohd Quamar TyagiandNiels Lohmann 675e519966 Allow SAX parsing of tagged CBOR (#5740)
* Allow SAX parsing of tagged CBOR

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* Consolidate sax_parse overloads with default tag_handler parameter

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* Add version history entry for tag_handler in sax_parse documentation

---------

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 07:38:24 +02:00
Niels Lohmann 21a69230bd Create a value before giving it its type (#5585)
* Create a value before giving it its type

Squashed onto develop from:
- Create a value before giving it its type
- Skip the failed-allocation test when exceptions are disabled
- Keep the created pointer rather than an uninitialized json_value
- Skip the vector<bool> failed-allocation check for VS 2015 with iterator debugging
- Test the remaining to_json overloads with a failing allocation
- Skip the to_json allocation-failure section on VS 2015 Debug
- Fix false GCC -Warray-bounds error with JSON_DIAGNOSTICS at -O3 (#5744)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Store the new value with a helper in all to_json constructors

Every external_constructor<>::construct now creates the new value first and
hands it to basic_json::replace_value(), which destroys the old value, stores
the new one before setting its type (as elsewhere in this PR), sets the
parents, and checks the invariant.

The std::vector<bool> and std::valarray overloads use array_t's range
constructor, which the other array overloads already rely on. Range views
keep their loop, as begin() and end() of a view may have different types.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-06 22:59:34 +02:00
162e13b86f Add with_*_t alias templates to create basic_json types with changed template parameters (#5758)
* Add helper types to make it easier to create a basic_json type with modified template parameters

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename with_changed_*_t aliases to with_*_t and merge integer/unsigned aliases

Per review discussion on #3898 between gregmarr and nlohmann:
- rename with_changed_X_t to with_X_t for brevity
- replace the separate with_changed_integer_t/with_changed_unsigned_t
  aliases with a single with_integers_t<NumberIntegerType2, NumberUnsignedType2>
- add @sa doc comment links for the upcoming documentation page

Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add documentation for the with_*_t member alias templates

Add docs/mkdocs/docs/api/basic_json/with_t.md documenting with_object_t,
with_array_t, with_string_t, with_boolean_t, with_integers_t, with_float_t,
with_allocator_t, with_json_serializer_t, with_binary_t and with_base_class_t,
with an accompanying example, and link the page from the basic_json member
types list and the mkdocs navigation.

Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add tests for the with_*_t member alias templates

Check with std::is_same that each with_*_t alias produces the expected
basic_json type, and that with_string_t keeps nlohmann::ordered_map as
the object type when used on ordered_json.

Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix with_t nav entry and document chaining of the with_*_t aliases

Indent the with_t entry in mkdocs.yml so it is listed under basic_json,
explain that the aliases can be chained and work on ordered_json, and
test both, including json::with_object_t<ordered_map> == ordered_json.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add docset entry for basic_json::with_t

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: barcode <barcode@example.com>
Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
2026-10-06 22:33:29 +02:00
fhgffy 69874e4544 Fix NUL bytes in UBJSON/BJData high-precision numbers (#5760)
* Fix NUL-terminated UBJSON and BJData high-precision payloads

Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>

* Use English comments for the high-precision NUL fix

Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>

* Track NUL bytes while reading high-precision payloads

Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>

* Drop dates from code comments.

Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>

* Report a NUL in a high-precision number where it is read

Return the parse error from the read loop so the byte offset points at the NUL, and drop the separate check after lexing.

Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>

* Use fixed text for the high-precision NUL error

2026-10-06: Use the known zero byte directly instead of formatting and concatenating it. Preserve the SAX token, error message, and byte offset.
Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>

---------

Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
2026-10-06 22:30:52 +02:00
fhgffy 9b179cee1e Handle all value_t enumerators in has_no_children() (#5770)
2026-10-07: add the scalar enum labels already handled by the default
branch so -Werror=switch-enum builds succeed. Keep the existing behavior
and synchronize the generated single header.

Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
2026-10-06 22:29:22 +02:00