Commit Graph
1058 Commits
Author SHA1 Message Date
Niels Lohmann 0b6a55ea6b Merge remote-tracking branch 'origin/json-view/11-view-access' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:20:19 +02:00
Niels Lohmann 5286c2843b Merge remote-tracking branch 'origin/json-view/10-view-document' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:20:14 +02:00
Niels Lohmann 3ba495ddf9 Merge remote-tracking branch 'origin/json-view/08-view-builder' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:20:09 +02:00
Niels Lohmann 2739c0b0af Merge remote-tracking branch 'origin/json-view/04-unicode-escapes' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:20:04 +02:00
Niels Lohmann 471b35c5e8 Merge remote-tracking branch 'origin/json-view/03-string-scan' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:19:58 +02:00
Niels Lohmann e87813232f Merge remote-tracking branch 'origin/json-view/02b-float-parser' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:19:20 +02:00
Niels Lohmann 15b0cc0561 Merge remote-tracking branch 'origin/develop' into HEAD
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:18:51 +02:00
Niels Lohmann d8d47be4a5 Fix CI on develop after merging the ready-to-merge PRs (#5754)
* Keep the serializer conversion for objects whose keys cannot be converted

#5591 added a test converting nlohmann::json into a basic_json whose
string type cannot be constructed from std::string. That instantiates
convert_iteratively(), whose members.emplace_back(next.key(), ...) needs
exactly that key conversion, and broke the build of unit-alt-string.
Dispatch on the key's constructibility and leave such conversions to the
serializers, as the levels above the nesting bound already do (#3425).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the remaining CI failures on develop

- unit-wstring: with a 16-bit wchar_t (Windows), a lone surrogate is
  reported as the ill-formed byte 0xFF since #5704; the std::wstring
  expectations still had the previous <U+0000>.
- ci_single_binaries: json_literals.hpp (#5610) and json.hpp include each
  other on purpose, and IWYU, not following the cycle, asks to replace
  json.hpp with json_fwd.hpp. Report its findings without failing the
  build, as already done for json.hpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the library warnings and noexcept specifications from the merged PRs

- binary_reader: rename the error_handler constructor parameter, which
  shadowed the member (-Wshadow, -Wshadow-field-in-constructor; #5746)
- basic_json(copy_construct_tag, ...): declare it noexcept when copying
  the base class is (GCC 16 -Wnoexcept; #5690)
- the scalar-on-left legacy comparison operators: noexcept only when
  converting the scalar is, like their member counterparts (#5682, #5751)
- compare_leaves: use std::is_eq/is_lt/is_gt instead of comparing a
  std::partial_ordering with 0 (-Wzero-as-null-pointer-constant; #5686)
- serializer: silence MSVC C4127 for the EnsureAscii template parameter
  (#5741, #5746)
- clang-tidy: return the sanitized reference in binary_writer, take the
  key of ordered_map::find_impl by const reference (#5727), and mark the
  switches over parse_array_index (#5728)
- ordered_map: keep <memory> for std::allocator (IWYU)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Split unit-conversions.cpp so MinGW can link it

clang 18 with the MinGW linker failed to link test-conversions_cpp17
("relocation truncated to fit: IMAGE_REL_AMD64_REL32"). As windows.yml
recommends, keep the objects small by splitting the test file.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the tests added by the merged PRs for all CI configurations

- discard the results of dump() and from_*() in CHECK_THROWS with
  utils::ignore_return_value (GCC -Werror=unused-result)
- give unit-bson's huge_string_t a default constructor (MSVC C2512,
  GCC 5, clang 3.5)
- unit-disabled_exceptions: use the literals namespace when the global
  UDLs are off (ci_test_noglobaludls; #5700)
- unit-binary_utf8_strict: expect the JSON pointer prefix with
  JSON_DIAGNOSTICS (#5741)
- skip the tests that rely on exceptions under JSON_NOEXCEPTION
  (#5678, #5732)
- clang-tidy and clang -Werror: static test data, CAPTURE(...);,
  const-correctness, use-after-move alias, unused conversion operator,
  a missing <iterator> include

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Title the macro examples and add JSON_STRICT_BINARY_UTF8 to the docset

The documentation style check requires "Example: ..." titles on pages with several examples (#5741, #5591) and a docset entry for every macro page.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Regenerate BUILD.bazel and nlohmann_json.natvis

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5746 added detail/output/error_handler.hpp and #5741 the json_abi_sbu8 ABI tag.

* Install libidn11 for the CMake 3.5.0 binary in ci_cmake_flags

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5733 moved ci_cmake_options from ubuntu:focal to ubuntu:24.04, which no longer ships libidn.so.11; the CMake 3.5.0 release binary links against it, so every ci_cmake_flags run has failed since. Install focal's libidn11 package for that matrix entry only.

* Suppress Infer's false STACK_VARIABLE_ADDRESS_ESCAPE in get_impl

get_impl() returns its local by value. A test added by the merged PRs instantiates it with a type Infer misreads, so ci_infer reported the 2021 code for the first time.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 17:18:19 +02:00
Niels Lohmann 73e9eae3c1 Add an error_handler parameter for UTF-8 to the binary readers and writers (#5746)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 12:13:49 +02:00
Niels Lohmann f56b418c56 Follow each binary format's UTF-8 rule: strict writers (CBOR/UBJSON/BJData/BSON), lenient readers (#5741)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 12:13:48 +02:00
Niels LohmannandRaphael Grimm 40021f38fb Add basic_json::as_base_class and document name conflicts with custom base classes (#5589)
* Add basic_json::as_base_class and document name conflicts with custom base classes

Members of basic_json hide members of a custom base class with the same
name, and future releases may add members that hide ones accessible
today. Document this in json_base_class_t and add as_base_class() to
reach hidden members without spelling out the cast.

Also make json_base_class_t a public member type. It was documented
since 3.12.0, but declared private, so users could not name it.

Supersedes #3899.

Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add as_base_class to the docset search index

New public members get an entry in docs/docset/docSet.sql (as done for
to_bon8/from_bon8 in #2998). Without it, the Dash/Zeal docset built from
the documentation cannot find basic_json::as_base_class.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Silence clang-tidy for the hidden type_name() in the base class test

ci_clang_tidy failed with readability-convert-member-functions-to-static
on base_class_with_hidden_members::type_name(). It must stay a
non-static member: the test shows that it is hidden by the non-static
basic_json::type_name() and reachable through as_base_class().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
2026-10-04 12:00:31 +02:00
Niels Lohmann c5650eaa3c Reject malformed UTF-16/UTF-32 units in wide-string input (#5704)
The wide-string input adapter (used for std::u16string, std::u32string,
std::wstring, and iterators over 2- or 4-byte character types) passed
some malformed code units on to the lexer as values that are neither a
byte (0x00..0xFF) nor char_traits<char>::eof(). As a result:

- A lone UTF-16 surrogate inside true/false/null was accepted if its low
  byte matched the expected letter, or ended the input silently if it
  was the last unit.
- A high surrogate followed by a unit that is not its low surrogate
  swallowed that unit; if the swallowed unit was the newline ending a
  // comment, the comment silently extended over the next line.
- Where wint_t is a signed int (macOS, the BSDs), a negative wchar_t
  collided with char_traits<char>::eof() (ending the input early) or was
  truncated to its low byte, depending on its value.

The UTF-32 helper now converts the code unit to std::uint32_t before the
range checks, so a negative unit reaches the same "emit 0xFF" branch
already used for code points above U+10FFFF. The UTF-16 helper now
peeks at the next unit before consuming it, and emits 0xFF instead of
the raw surrogate when no valid pair is found, matching how ill-formed
UTF-8 bytes are rejected elsewhere in the lexer.

Fixes #5645.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:59:03 +02:00
Niels Lohmann 38a2db260c Remove dead metaprogramming and duplicated code in traits and pointers (#5728)
* Remove unused is_sax and is_detected_convertible

detail::is_sax had no user: the parser and the binary reader only use
is_sax_static_asserts, so is_sax was a second, unchecked copy of the
SAX event list. is_sax_static_asserts asserted boolean(bool) twice in
a row, and detail::is_detected_convertible was never used anywhere.

Remove all three and include <cstddef> for size_t instead of <cstdint>.
Only names in nlohmann::detail are removed; behavior, public API and ABI
are unchanged. The diagnostics for an incomplete SAX handler are the
same, apart from the duplicated boolean() message.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Replace meta/logic.hpp with a disjunction trait

meta/logic.hpp added a second set of type-level boolean helpers
(cxpr_and, cxpr_or, cxpr_not, ...) next to the existing conjunction
and negation in type_traits.hpp. It was used only by one static_assert
in from_json_tuple_impl, two of its templates were never used, and it
was the only header without the license banner and relied on
transitive includes for <type_traits>.

Add the missing disjunction next to conjunction and negation, use the
three in the static_assert, and delete logic.hpp together with its
BUILD.bazel entry. same_sign now uses disjunction as well, which
resolves the 2022 TODO waiting for such a trait.

The static_assert accepts and rejects the same types as before. Only
names in nlohmann::detail change; behavior, public API and ABI are
unchanged.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove unused would_call_std_* from NLOHMANN_CAN_CALL_STD_FUNC_IMPL

Besides detail::result_of_begin/end, which is_range and iterator_t use,
the macro defined a namespace detail2 with a tag type, a catch-all
overload and would_call_std_begin/end, plus would_call_std_begin/end
structs directly in namespace nlohmann. Nothing has used them since
they were added in #3020.

Reduce the macro to its detail part. Without the trailing struct the
';' after the two invocations would be an empty declaration that
-Wextra-semi flags, so drop it. macro_scope.hpp included
meta/detected.hpp only for this macro; all users of detected.hpp
include it (or type_traits.hpp) themselves, so remove the include.

Behavior and ABI are unchanged. The undocumented, untested and unused
names nlohmann::would_call_std_begin, nlohmann::would_call_std_end and
namespace nlohmann::detail2 are no longer declared.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Simplify is_ordered_map to reuse has_capacity

is_ordered_map re-detected capacity() with a C++03 sizeof/vararg
trick right after has_capacity did the same detection through
is_detected. For ordered_map, the old trick took the address of
std::vector::capacity, which [namespace.std]/6 makes unspecified.
Reuse has_capacity instead, which removes the unspecified-behavior
pointer-to-std-member and two NOLINT suppressions.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove duplicate const overload of json_pointer::get_checked

The const and non-const get_checked() overloads had byte-identical
50-line bodies, differing only in the signature. The remaining
template deduces a const-qualified BasicJsonType for const callers,
so at(), the out_of_range::create() calls and the bounds check all
still work.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix tautological clause in iter_impl's iterator category assertion

The static_assert meant to check the LegacyBidirectionalIterator
named requirement had a first clause comparing
std::bidirectional_iterator_tag to itself, which is always true and
checks nothing; only array_t::iterator was actually being checked,
despite the message claiming object iterators were checked too.
Drop the tautological clause, reword the message to describe what
is actually checked, and note that object_t may use a forward-only
iterator as long as reverse iteration and operator-- are unused.
The check is intentionally not extended to object_t::iterator, since
that would reject object types with forward-only iterators that
compile and work correctly today.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix misplaced and stale comments in JSON_HAS_RANGES and conversions

The JSON_HAS_RANGES feature-detection block had its libc++ comment
sitting above the clang+libstdc++ branch it does not describe,
leaving the libc++ branch uncommented and the clang+libstdc++ branch
without its own rationale. Move each comment to sit under its own
branch, and give the clang+libstdc++ branch (added in issue 5161) its
own one-line reason referencing that issue instead of reusing the
libc++ branch's comment. Also fix a duplicated-word typo ("in large
in large cpp files") in from_json.hpp, drop two unanswered 2017
design questions left as comments in type_traits.hpp and
from_json.hpp that no longer reflect open questions, and correct
NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's @since tag from 3.12.0 to
3.13.0, the release it was actually introduced in.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Support any-rank C arrays in from_json, not just rank 1-4

from_json() for C arrays had four hand-unrolled overloads (rank 1-4,
added incrementally in #4262), each with its own nested loops. to_json()
already handles any rank recursively, so a rank-5+ C array could be
serialized but not read back with get_to()/get<>().

Replace the four overloads with one from_json() SFINAE-constrained on
get<remove_all_extents<T>::type>() existing, forwarding to a pair of
mutually recursive from_json_c_array_element() helpers: one assigns a
non-array element via get<T>(), the other loops over a array element and
recurses one dimension at a time. Each dimension still goes through at(),
so type_error.304/out_of_range.401 stay unchanged; ranks 1-4 keep their
existing behavior and semantics.

Adds rank-5 round-trip and mismatched-shape tests to unit-conversions.cpp.

Public API: additive only (rank 5+ C arrays become readable).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 1

* Move templated_json_throw into nlohmann::detail

templated_json_throw() was defined in macro_scope.hpp, which is included
outside NLOHMANN_JSON_NAMESPACE_BEGIN, so the helper leaked into the
global namespace as ::templated_json_throw with no ABI tag. Unqualified
lookup in NLOHMANN_JSON_SERIALIZE_ENUM_STRICT could then bind to a
same-named function declared in the user's own namespace instead, which
fails to compile with Clang ("does not name a template").

Move the helper next to the exception classes in exceptions.hpp, inside
nlohmann::detail, and call it qualified as
::nlohmann::detail::templated_json_throw<...>(...) from both macro
expansion sites. Rewrite the doc comment to give the real reason for the
helper (JSON_THROW may expand to code that discards its argument, e.g.
when exceptions are disabled) and fix the "supress" typo.

templated_json_throw was never released (added by #5151 after v3.12.0),
so it can be moved freely.

Adds a regression test that expands NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
inside a namespace declaring its own templated_json_throw.

Public API: no change (::templated_json_throw was an unreleased,
unintentional global-namespace leak with no callers relying on its
location).

Overlaps #5698, which rewrites the same two macro call lines; the
overlapping hunks are small and should be trivial to reconcile on
rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 2

* Factor the repeated JSON_HAS_RANGES/MinGW guard into one macro

The std::ranges view conversion (excluded on MinGW because of its
incomplete C++20 ranges support, #4916) was gated by the same
#if JSON_HAS_RANGES && !defined(__MINGW32__) condition at seven
independent sites in to_json.hpp and type_traits.hpp, with the MinGW
rationale duplicated in two of them and missing from the rest. Since the
sites come in matching pairs (one enables is_compatible_range_view and a
view-based overload, the other adds the exclusion to the
plain-array-type overload), a drift between any pair would produce an
ambiguous or missing overload on exactly one platform.

Add JSON_HAS_RANGE_VIEW_CONVERSION next to JSON_HAS_RANGES in
macro_scope.hpp, combining both conditions with the #4916 reasoning in
one place, #undef it in macro_unscope.hpp, and use it at all seven
sites. This does not fold the MinGW check into JSON_HAS_RANGES itself:
JSON_HAS_RANGES is user-overridable and also gates the
enable_borrowed_range specialization in iteration_proxy.hpp, which is
not excluded on MinGW.

No behavior or public API change: JSON_HAS_RANGE_VIEW_CONVERSION expands
to exactly the condition that was previously written out at each site.

Overlaps #5585, #5600 and #3575, which touch the same to_json.hpp and
type_traits.hpp lines; the change here is a mechanical
search-and-replace of the guard condition and should rebase cleanly.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 11

* De-duplicate from_json.hpp's map and array-fallback bodies

Several from_json() overload pairs in from_json.hpp were copies of each
other, so a fix has to be applied twice (as #5681 already does):

- from_json(..., std::map&) and from_json(..., std::unordered_map&) for
  non-string keys had identical 16-line bodies: array check, m.clear(),
  pair check loop, m.emplace(...). Route both through a new
  from_json_pair_array_to_map(j, m) helper.
- The from_json_array_impl priority_tag<1> and priority_tag<0> fallbacks
  ran the same std::transform/std::inserter loop, differing only in
  ret.reserve(j.size()). Merge them into one body and, modeled on the
  existing from_json_object_reserve, add a from_json_array_reserve pair
  so the reserve() call is only made for ConstructibleArrayType that
  support it.

Error ids (type_error.302), messages, diagnostic paths ((at(0)/at(1))
and behavior for types with/without reserve() are unchanged; only the
duplication is removed.

Public API: no change.

Overlaps #5681, which changes the "&j" to "&p" line in both map bodies;
the shared helper here should make that a one-line change instead of two
on rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 5

* Unify json_pointer's three array-index parsers

array_index(), contains() and get_checked_or_null() each re-implemented
the RFC 6901 array-index rules and the size_type range check: array_index()
does the canonical parse and throws; contains() (which must not throw,
#5395) re-validates every digit by hand and runs its own strtoull/ERANGE
check before calling array_index() anyway, parsing every array token
twice; get_checked_or_null() wraps array_index() in JSON_TRY/
JSON_INTERNAL_CATCH (detail::out_of_range&) to turn an unrepresentable
index into "not found".

Add a single private, noexcept parse_array_index(s, idx) returning an
array_index_status (ok / leading_zero / not_a_number / unresolved /
exceeds_size_type). array_index() becomes a thin wrapper mapping each
status to the existing parse_error.106/109 or out_of_range.404/410;
contains() and get_checked_or_null() switch on the status directly. This
removes contains()'s digit-validation loop and its second strtoull call,
and get_checked_or_null()'s JSON_TRY/JSON_INTERNAL_CATCH.

Bugfix as a consequence: get_checked_or_null()'s JSON_TRY/
JSON_INTERNAL_CATCH was dead code under JSON_NOEXCEPTION (JSON_TRY
expands to "if(true)" and the catch to "if(false)", so JSON_THROW's
std::abort() ran unconditionally), meaning value() and contains() would
abort instead of returning the default/false for an out-of-range-sized
or oversized array index when exceptions are disabled (#5672). Switching
on parse_array_index()'s return value instead of relying on an actual
throw/catch fixes this: get_checked_or_null() now returns nullptr for
array_index_status::unresolved/exceeds_size_type in every build
configuration, and still calls JSON_THROW (aborting under
JSON_NOEXCEPTION, as before) only for a malformed index
(leading_zero/not_a_number), matching its documented @throw list.

All existing error ids, messages and diagnostic paths are unchanged; a
few reference tokens that used to fail contains()'s manual per-character
validation (e.g. "1a") now fail via array_index_status::unresolved
instead, with no observable difference since contains() only returns
bool.

Adds regression tests to unit-element_access2.cpp's "access on array
type" section covering value() with an index that exceeds size_type and
one with a trailing non-digit, both of which must yield the default
value rather than abort/throw.

Public API: no change.

Overlaps #5700, #5614 and #5692, which touch the contains() and
get_checked_or_null() array hunks; this change replaces those hunks with
calls into the new shared parser, so a rebase will need to re-apply
their token-handling changes (e.g. the empty-token case) on top of the
switch statements here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 4

* Regenerate single_include after merging develop

The merge commit kept develop's single_include/nlohmann/json.hpp because
make amalgamate saw it as up to date.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review: switch in array_index, drop redundant inline

- json_pointer::array_index() dispatches on array_index_status with a
  switch, matching the other parse_array_index() caller
- drop `inline` from the function templates this PR adds or moves in
  from_json.hpp
- reword a comment that described the change rather than the code

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:50:42 +02:00
Niels Lohmann 1df1e8a845 Make cross-string-type basic_json conversion explicit without implicit conversions (#5591)
* Make cross-string-type basic_json conversion explicit without implicit conversions

The converting constructor from another basic_json specialization was
always implicit, so a value with a different string_t (std::wstring, a
string with a custom allocator, ...) silently converted into a temporary,
e.g. when passed to a function taking const nlohmann::json&. Such
conversions do not produce correct values (#3425), and
JSON_USE_IMPLICIT_CONVERSIONS=0 did not catch them.

When JSON_USE_IMPLICIT_CONVERSIONS is 0, the constructor is now explicit
if the string types differ. Specializations sharing a string type (json
and ordered_json, different serializers or object maps) stay implicitly
convertible, so the NLOHMANN_DEFINE_TYPE_* macros keep working with
nested json members. get<BasicJsonType>() constructs explicitly and
works in both modes.

Fixes #2649.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Construct explicitly in get_to() and to_json(std::optional)

With JSON_USE_IMPLICIT_CONVERSIONS=0 the conversion from a basic_json
with a different string type is now explicit, but two library paths
still assigned such a value implicitly and failed to compile inside the
library:

- get_to() with a basic_json target (the #2175 overload) did
  `v = *this`, so json(42).get_to(alt_json&) broke although
  get<alt_json>() works.
- to_json(BasicJsonType&, const std::optional<T>&) is constrained on
  std::is_constructible (which accepts the explicit constructor) but did
  `j = *opt`, so converting a std::optional<alt_json> into a json broke.

Both now construct the value explicitly, as get_impl() already does.

Also replace static_cast<bool>(JSON_USE_IMPLICIT_CONVERSIONS) with a
comparison: clang-tidy's modernize-use-bool-literals rejected the cast
of the integer literal the macro expands to, failing ci_clang_tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:52 +02:00
Niels Lohmann a212d3b2e4 Preserve the object comparator's state in a deep copy past the nesting bound (#5722)
* Preserve the object comparator's state in a deep copy past the nesting bound

copy_object_level(), used by the copy constructor and copy assignment once a
value is nested deeper than the iterative deep copy's bound (128 levels, or
every copy under JSON_NO_THREAD_LOCAL), built each object's copy with the
object type's plain range constructor. That default-constructs the object's
comparator instead of copying the original's. For an object type whose
comparator carries state, such as a std::map that compares keys
case-sensitively only when constructed that way, the copy then ordered - and
could even deduplicate - its keys differently from the original.

Add detail::is_comparator_constructible_object_type, a detection trait for
object types that provide a key_comp() and a constructor taking a range and a
comparator, the way std::map does. copy_object_level now dispatches on it: an
object type that qualifies gets its copy built with src_object.key_comp()
passed along; other object types, such as nlohmann::ordered_map (which has a
key_compare for its std::map-like interface, but no key_comp()), keep using
the plain range constructor exactly as before.

merge_patch and update() were checked for the same pattern; neither is
affected, since both only ever add members one at a time to an object that
already has its own comparator (or start a brand new default-constructed one),
rather than rebuilding an object_t from a range copied out of an existing,
possibly custom-comparator object.

Fixes #5649.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep astyle from padding the create_object_with_comparator templates

Spell the negated condition as detail::negation<...> instead of a leading
'!', which made astyle spread the template header out.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:47 +02:00
Niels Lohmann 1edf0ef041 Fix value(json_pointer, default) aborting under JSON_NOEXCEPTION (#5700)
With exceptions disabled (JSON_NOEXCEPTION or -fno-exceptions),
value(const json_pointer&, default) called std::abort() for array
reference tokens that array_index() rejects with out_of_range.404/410:
indices too large to fit size_type, the empty token ("/"), and tokens
like "/1a". With exceptions enabled, the same tokens correctly yielded
the default value, because get_checked_or_null() relied on
JSON_TRY/JSON_INTERNAL_CATCH (detail::out_of_range&) to turn the
exception into nullptr; under JSON_NOEXCEPTION, JSON_THROW aborts
before that catch is ever reached.

get_checked_or_null() now detects those out-of-range tokens itself,
the same way contains(json_pointer) already does (#5495), and only
calls array_index() for tokens that must still raise parse_error.106
or parse_error.109 (e.g. "/01", "/+1"), matching the documented
behavior of value().

Added regression tests to tests/src/unit-disabled_exceptions.cpp
(built with JSON_NOEXCEPTION and -fno-exceptions) and the matching
checks to tests/src/unit-element_access2.cpp for normal exception
mode.

Fixes #5672.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:43 +02:00
Niels Lohmann 756b28c2b8 Keep a NUL byte ending a // comment as the end of input (#5696)
With the default NUL handling (JSON_STRICT_NUL_HANDLING not set), a NUL
byte in the input is treated as the real end of input everywhere -
except when it immediately ends a `//` comment: scan_comment() matched
'\0' as a comment terminator like '\n', so the NUL was consumed as
part of the comment and scan() never saw it as end of input; the next
get() then kept reading past it. Multi-line comments and
JSON_STRICT_NUL_HANDLING=1 were unaffected, since there the NUL is
just part of the comment text.

Fix scan_comment() to leave the NUL unconsumed (unget()) instead of
returning it as part of the comment, so the following scan() reports
it as end of input, exactly as for a NUL anywhere else.

Fixes #5659.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:39 +02:00
Niels Lohmann 49cd427196 Copy-construct the base class of a deep copy's elements, not assign it (#5690)
The bounded-descent copy added by #5389 built the elements of a deep copy
(nested past the 128-level bound) by default-constructing them and then
having copy_metadata() assign their base class afterwards. That assignment
is only instantiated for values nested past the bound, but being called
from copy_structured() at all meant it was compiled for every copy, so a
CustomBaseClass that is copy-constructible but not move-assignable (for
example one with a const data member) no longer let its basic_json be
copy-constructed, at any depth.

copy_array_level() and copy_object_level() now build each element with a
private-tag-selected constructor that copy-constructs the base class (and,
under JSON_DIAGNOSTIC_POSITIONS, copies the positions) directly, the same
way the copy constructor already builds elements within the 128-level
bound. Copying a basic_json is therefore back to requiring only a
copy-constructible base class, as documented and as it was before #5389;
copy assignment is unchanged and still requires an assignable one.

Fixes #5674.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:36 +02:00
Niels Lohmann b730946432 Fix key types convertible to std::string_view breaking lookups (#5689)
Since #4958, a key type implicitly convertible to std::string_view was
accepted by is_usable_as_basic_json_key_type without checking that the
object's comparator can actually compare object_t::key_type with that
key type. The key was then forwarded unchanged to the underlying map,
so const operator[], at, find, count, contains, erase and value failed
to compile (a hard error inside <map>) for a key convertible only to
std::string_view, and value() rejected such keys outright. For keys
convertible to both std::string and std::string_view, the KeyType&&
templates now won overload resolution over the object_t::key_type
overloads and then failed the same way, a regression from 3.12.0. Only
the non-const operator[] worked, because it uses emplace(), which
constructs a std::string from the key explicitly. ordered_json was not
affected, since ordered_map checks comparability itself.

Add a trait, is_string_view_convertible_key_type, that recognizes a key
type that is convertible to std::string_view but not directly
comparable with the object's key type, provided std::string_view itself
is comparable with it. at(), operator[], find(), count(), contains(),
erase() and value() now route such keys through a new lookup_key()
helper that converts them to std::string_view before they reach the
object, matching how the object's transparent comparator already
supports std::string_view lookups.

Fixes #5663.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:32 +02:00
Niels Lohmann 0d01d6ae90 Classify leaves with operator<=> itself past the nesting bound (#5686)
* Classify leaves with operator<=> itself past the nesting bound

In C++20, an ordered comparison past the nesting bound classified a pair of
leaves by asking == first and then order_leaves(), which calls < and > -
both derived from <=>. For a pair of binary values with the same bytes but a
different subtype, == reports them unequal, while <=> (through
std::vector<std::uint8_t>::operator<=>) reports them equivalent, so the pair
ended the comparison as unordered instead of letting the next element
decide - unlike an array or object within the bound, which compares such a
pair with its own operator<=> and gets equivalent. So operator<=>, and the
<, <=, >, >= derived from it, could give a different result for the same two
values depending on how deeply the values were nested, or unordered at every
depth with JSON_NO_THREAD_LOCAL defined.

compare_leaves() now classifies such a pair in C++20 with operator<=> itself
instead, matching how a value within the bound is compared; the equality-only
and pre-C++20 ordered cases are unchanged. Which of the three runs is chosen
by overloading on std::integral_constant<bool, Ordered>, the same tag
dispatch order_leaves() already uses, rather than a runtime "if (Ordered)" on
a template parameter, which MSVC would flag as a constant condition (C4127).

Added a regression test to unit-comparison.cpp that nests such a pair 0, 127,
128 and 200 levels deep (127 stays within the 128-level bound, 128 and 200
do not) and checks that operator<=> and operator< agree at every depth.

Fixes #5654.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop the version history note for a bug that was never released

The regression came from #5390, which is not in any release. Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:29 +02:00
Niels Lohmann 6f2048cd5d Accept lvalues in ordered_map::emplace's value parameter (#5685)
* Accept lvalues in ordered_map::emplace's value parameter

ordered_map::emplace(key, value) took the mapped value only by T&&, an
rvalue reference rather than a forwarding reference, so
ordered_json::emplace("a", value) failed to compile whenever value was
an lvalue or a const lvalue, even though the same call compiles for
json (whose object_t is std::map, with a variadic emplace). Turn the
value parameter into a separately-deduced forwarding reference,
constrained with std::is_constructible so the overloads still only
accept something convertible to the mapped type. std::map-compatible
semantics are unchanged: emplace still does nothing if the key already
exists.

Open PR #5609 also touches ordered_map.hpp (moving values on vector
growth); this change only touches the two emplace() overloads and
should not conflict.

Fixes #5673.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid astyle's padding in ordered_map::emplace's template headers

Use detail::conjunction instead of && and drop the redundant V&& in detail::is_constructible, so astyle keeps the usual template formatting. Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:25 +02:00
Niels Lohmann df27cc3d4c Add scalar-on-left overloads for legacy discarded comparisons in C++20 (#5682)
With JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1 and C++20, a scalar on
the left-hand side of <= or >= (e.g., `1 <= discarded`) yielded false
instead of the documented true. The C++20 legacy block only had member
operators, which are only candidates when the basic_json is the left
operand; for a scalar on the left, overload resolution picked the
candidate rewritten from operator<=>, which does not emulate the legacy
behavior. The C++17 branch already has scalar-on-the-left friend
overloads for <= and >=; add the equivalent pair to the C++20 legacy
block.

Added a regression test to tests/src/unit-comparison.cpp covering all
four operand orders for both operators.

Fixes #5665.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:21 +02:00
Niels Lohmann d11e89471b Assert on missing array indices in const operator[] and document the JSON pointer case (#5606)
The const operator[] overloads are unchecked by design, and a missing key
or index is undefined behavior. The key overload guards this with a
runtime assertion, but the index overload did not, although the element
access documentation says an assertion fires in both cases. The const
JSON pointer overload inherits both through json_pointer::get_unchecked(),
so a pointer to a missing array index read out of bounds even in debug
builds, and its documentation promised out_of_range.404 for any pointer
that cannot be resolved.

Add JSON_ASSERT(idx < size()) to const operator[](size_type), which also
covers the index leg of the const JSON pointer overload. Document the
undefined behavior for the const JSON pointer overload in operator[].md
and in the runtime assertions page. Release builds are unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:46:18 +02:00
Niels Lohmann 6d0867a85f Make comparisons with scalars noexcept only when the conversion is (#5751)
The comparison operators taking a scalar (==, !=, <, <=, >, >=, and
C++20's <=>) convert the scalar to a basic_json and compare, but were
unconditionally noexcept. When that conversion throws, the program
called std::terminate instead of propagating the exception, e.g. when
comparing a json with a string literal under memory pressure
(std::bad_alloc) or with an enum value not mapped by
NLOHMANN_JSON_SERIALIZE_ENUM_STRICT (out_of_range.410). clang-tidy
22.1 reports the latter as bugprone-exception-escape.

Declare the 16 scalar overloads
noexcept(std::is_nothrow_constructible<basic_json, ScalarType>::value):
they stay noexcept for numbers, Booleans, nullptr, and plain enums,
and are noexcept(false) for strings and enums whose to_json may throw.
The comparisons of two basic_json values are unchanged.

Restore the strict-enum comparisons removed from unit-conversions.cpp
in the previous PR, check that comparing an unmapped strict enum now
throws, and pin the new exception specifications in unit-noexcept.cpp.
Document the exception safety of overload (2) on all seven operator
pages. Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 17:25:17 +02:00
Niels Lohmann f17207377b Mark an Infer false positive in get_impl()
The tests added here instantiate basic_json::get() with a type for which Infer 1.3.0 reports STACK_VARIABLE_ADDRESS_ESCAPE on "return ret;", although ret is returned by value. Suppress it on that line, as develop does for its own Infer false positives (#5750).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:43:28 +02:00
Niels Lohmann 3b30446d1b Merge branch 'json-view/11-view-access' into json-view/12-view-values
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:37:21 +02:00
Niels Lohmann 960d3435a1 Merge branch 'json-view/10-view-document' into json-view/11-view-access
Conflicts in See also lists (docs/mkdocs/docs/api/basic_json/begin.md,docs/mkdocs/docs/api/basic_json/cbegin.md,docs/mkdocs/docs/api/basic_json/cend.md,docs/mkdocs/docs/api/basic_json/end.md,docs/mkdocs/docs/api/basic_json/type_name.md), where develop (#5638) and this branch both edited: kept develop's entries and added this branch's basic_json_view links.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:37:16 +02:00
Niels Lohmann ceb1cdb774 Merge branch 'json-view/08-view-builder' into json-view/10-view-document
Conflicts in the See also lists of nine basic_json pages, is_discarded.md, and features/index.md, where develop (#5638) and this branch both added entries: kept both. Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:36:41 +02:00
Niels Lohmann 80761cf255 Merge branch 'json-view/04-unicode-escapes' into json-view/08-view-builder
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:36:09 +02:00
Niels Lohmann eed512f1d5 Merge branch 'json-view/03-string-scan' into json-view/04-unicode-escapes
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:36:05 +02:00
Niels Lohmann 3f672f036d Merge branch 'json-view/02b-float-parser' into json-view/03-string-scan
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:36:00 +02:00
Niels Lohmann 83302ff69d Merge branch 'develop' into json-view/02b-float-parser
Conflicts:
- number_parse.hpp: kept this branch's float parser, which replaces the
  Eisel-Lemire code that develop's side changed (#5750 made its digit
  counter unsigned; this parser has no such counter, and it compiles
  cleanly with GCC's -Wstrict-overflow=5).
- number_handling.md, template_parameters.md: kept this branch's
  description of the conversion and added develop's "Before version
  3.13.0" sentence.

Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:35:40 +02:00
Niels Lohmann 63c10a51fc Review and extend the documentation, and check it in CI (#5638)
* Review and extend the documentation, and check it in CI

A review of all documentation pages found factual errors, dead links,
missing cross-references, and gaps in examples. This fixes them and adds
checks so the same problems are caught automatically.

Fixes:
- wrong signatures and version histories (operator!= C++20 member,
  binary() subtype type, get<PointerType>(), JSON_NO_THREAD_LOCAL, ...)
- stale descriptions (number parsing since #5283, UBJSON table, SAX
  example that no longer compiled, tsl::ordered_map advice)
- dead internal and external links; repology.org badges (the domain is
  suspended) replaced by badges that query the registries directly
- deprecation notes link the migration guide; the guide itself fixed

Additions:
- "See also" sections, cross-references, 25 runnable examples, 12
  Mermaid diagrams, new API pages for json_pointer::operator<=> and
  byte_container_with_subtype::operator==/!=
- landing page, guides for untrusted input and performance
- "unreleased" badge after versions newer than the latest release

Checks:
- strict documentation build (broken links/anchors fail it); CI and
  the publish workflow fetch the full history the build needs
- weekly external link check, Mermaid syntax check in CI
- check_structure.py: example titles, heading levels, alt texts,
  header links, docset index coverage; its unused-example check works
  again
- all examples produce the same output on every platform

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep the customer links that could not be fixed

A dead link on the customers page is still the evidence of where the
use of the library was documented. Keep the original URLs of the entries
without a working replacement (Marne, Cisco Webex Desk Camera, Philips
Hue, CyberArk) and exclude exactly these URLs from the link check.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the duplicate-key recipe's claim about SAX positions

The SAX interface's key() receives no position either; only parse_error()
does. Also note that the recipe does not report the path to the repeated
key (see discussion #5085).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Say the library is available as a single header and mention json_fwd.hpp

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct documentation errors found while hunting for bugs

- patch/patch_inplace: list the JSON pointer errors parse_error.106-109
  and out_of_range.402/404, and quote the actual parse_error.105 message.
- unflatten: list parse_error.106/107/108 and out_of_range.404.
- to_bson: list out_of_range.415 (binary subtype above 255) and note
  that 412 and 415 are new in 3.13.0.
- to_string: state that string_t must be convertible to std::string, also
  in the StringType requirements table.
- JSON Lines: a `while (input >> j)` loop also throws after the last value
  for concatenated JSON values; show a loop that works for both.
- BON8: a string gets 0xFF only if nothing follows it in the message; a
  string at the end of an array or object is ended by 0xFE.
- custom_string_type.hpp: add operator+=(char), which the "Always
  required" list asks for (json_pointer::to_string, flatten, unflatten,
  and diff did not compile), and an ADL int_to_string for diff and items.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cache the release headers with functools.lru_cache

Codacy (Pylint) flagged the mutable default argument that header() used
as its cache. functools.lru_cache keeps the same memoization without it.
The script's output is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:32:15 +02:00
Niels Lohmann 4d46bce4e1 Fix CI on develop (#5750)
* Fix CI configuration broken by recent merges and tool updates

- gcc_flags.cmake: drop -Wexperimental-fmv-target, which GCC 16 accepts
  only on aarch64; amd64 rejects it, so every GCC job failed while
  checking the compiler.
- ci_get_cmake: add VERBATIM so the checksum pipeline is passed to the
  shell intact (the unescaped `$'` broke the generated Makefile and
  build.ninja, failing ci_cmake_flags and ci_module_cpp20); match the
  SHA-256 entry case-insensitively, as CMake 3.5.0 lists the archive as
  "Linux-x86_64"; and unpack with --strip-components, as that archive's
  top-level directory is spelled "Linux" too.
- ci_single_binaries: compile json.hpp's TU without IWYU's --error, as
  the comment above the gate already intends.
- tests: restore -Wno-deprecated-declarations for all non-MSVC compilers
  (#5737 kept it for GCC only), as several tests call deprecated
  functions on purpose; include thirdparty/fifo_map as SYSTEM.
- .clang-tidy: set misc-use-internal-linkage.AnalyzeTypes to false;
  clang-tidy 22.1 extended the check to classes and enums and flagged
  100 test helper types.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix library warnings and a JSON_DIAGNOSTICS parent bug

- binary_reader: pass integers to sax->number_integer() through
  conditional_static_cast<number_integer_t>, making the existing
  narrowing for a narrow number_integer_t explicit (MSVC C4244 and GCC
  -Wconversion/-Warith-conversion with the int16_t test from #5694);
  mark two Infer DEAD_STORE false positives with @infer-ignore.
- to_json: set the parents of an array built from a C++20 range view
  after all elements are in place; a reallocating push_back moved the
  earlier elements and left their parent pointers stale, failing the
  JSON_DIAGNOSTICS invariant assertion.
- json.hpp: suppress MSVC C4127 for the new is_ordered_map check in
  diff(), like the three existing ones; spell out std::formatter::parse's
  return and iterator types for clang-tidy 22.1.
- number_parse: make the Eisel-Lemire digit counter unsigned
  (GCC -Wstrict-overflow).
- string_utils: take encode_utf8's callable by const reference
  (cppcoreguidelines-missing-std-forward) and drop a \u from its doc
  comment (-Wdocumentation-unknown-command).
- ordered_map: include <memory> for std::allocator (cpplint).

Ran make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix tests failing in CI on develop

- unit-allocator: skip the #5640 test under MSVC STL iterator debugging,
  where containers allocate a debug proxy in noexcept move constructors
  and a failing allocation terminates; move a decrement out of an if
  condition (bugprone-inc-dec-in-conditions).
- unit-conversions: expect the "(/0)" path with JSON_DIAGNOSTICS;
  compare strict enums via get<>() rather than through the noexcept
  operator==(ScalarType, json), which bugprone-exception-escape flags.
- unit-alt-string: suppress -Wexit-time-destructors for the strict enum
  macro and misc-use-internal-linkage for its enum.
- unit-bjdata: call the static lookup functions through the type and
  pass unsigned char (-Wsign-conversion on amd64).
- Mark Infer false positives with @infer-ignore in unit-diagnostics,
  unit-pointer_access, unit-udt, and unit-conversions.
- Smaller clang-tidy 22.1 findings in unit-class_parser,
  unit-constructor2, unit-custom-base-class, unit-locale-cpp, and
  unit-noexcept.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:23:55 +02:00
Niels Lohmann 6aec1a085f Report BON8 input that ends after a UTF-8 lead byte as truncated (#5677)
A lead byte (0xC2..0xF7) inside a string begins either another character
(if a continuation byte follows) or an integer (otherwise). When the input
ended right after the lead byte, the reader took the missing byte as "not
a continuation byte", ended the string before the lead byte, and treated
the lead byte as the start of the next value. With strict=false, a message
cut off there was therefore read as a shorter value: the 11 bytes of
"😀😀é" cut after 9 bytes gave "😀😀", and ["aé"] cut after 3 of its 5
bytes gave ["a"]. With strict=true, the input was rejected with a
misleading message ("expected end of input"), or, for a key, with
parse_error.112 instead of 110.

Either reading of the lead byte leaves the message incomplete: a string at
the end of a message must be terminated by 0xFF, so the lead byte cannot
belong to a following message. Report parse_error.110 (unexpected end of
input) for strings and keys, as the comment on get_bon8_string() already
requires and as the reference decoder (HikoGUI) does.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 07:47:50 +02:00
Niels Lohmann b1595c1b40 Merge branch 'json-view/11-view-access' into json-view/12-view-values
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:13 +02:00
Niels Lohmann 3d6d610fdc Merge branch 'json-view/10-view-document' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:03 +02:00
Niels Lohmann 134b2f0efe Mark json_view.hpp's read() and strlen as Flawfinder false positives
json_document::read is a member function, not POSIX read(), and the C
string overload requires null-terminated input like json::parse.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:27:01 +02:00
Niels Lohmann 05b6cd0892 Merge branch 'json-view/11-view-access' into json-view/12-view-values
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:19:53 +02:00
Niels Lohmann 95d10dab70 Merge branch 'json-view/10-view-document' into json-view/11-view-access
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:19:51 +02:00
Niels Lohmann 371a8a3d9f Merge branch 'json-view/08-view-builder' into json-view/10-view-document
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	Makefile
#	cmake/ci.cmake
2026-10-01 10:19:49 +02:00
Niels Lohmann ae01d57694 Merge branch 'json-view/03-string-scan' into json-view/04-unicode-escapes
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:18:33 +02:00
Niels Lohmann 6a757ca675 Merge branch 'json-view/02b-float-parser' into json-view/03-string-scan
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:18:33 +02:00
Niels Lohmann cb51f80e34 Merge branch 'json-view/04-unicode-escapes' into json-view/08-view-builder
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 10:18:33 +02:00
Niels Lohmann 9d44e3f359 Merge branch 'develop' into json-view/02b-float-parser
Conflicted only in tests/src/unit-class_lexer.cpp, where develop's #5737
lint fix (CAPTURE(x); -> CAPTURE(x)) collided with this PR's rewrite of
the Eisel-Lemire float tests; kept the PR's new tests and applied the
lint-fixed CAPTURE style. single_include regenerated via make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:53:02 +02:00
Niels Lohmann e400780533 Re-amalgamate single_include (#5745)
* Re-amalgamate single_include

#5737 changed 13 headers under include/ but merged without the matching
single_include/nlohmann/json.hpp update, so the amalgamated header still
had, among others, the GCC C++20 -Wignored-attributes pragma block and
the clang -Wdocumentation push/pop that #5737 removed, the forwarding
from_json tuple/array helpers it replaced with const references, and
lacked the output_adapter char_traits changes it added.

Regenerated with `make amalgamate` (astyle 3.4.13). The diff is exactly
`git diff b54ed188e e5a89d671 -- include/` (164+/95-); json_fwd.hpp and
json_literals.hpp were already up to date.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Run the amalgamation check on pushes to develop

#5737 was merged 39 seconds after its last push, while its own
check_amalgamation run was still queued (earlier runs had been
cancelled by the concurrency group), so the stale single_include
reached develop without any failing check.

Also run the check on pushes to develop, without cancelling in-progress
develop runs. The "save" job (PR number/author for the comment
workflow) only runs for pull requests, the checkout falls back to
github.sha, and comment_check_amalgamation.yml only comments for
PR-triggered runs, since push runs have no PR and no "pr" artifact.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:09:34 +02:00
Niels Lohmann cff0a61369 Share DOM SAX position handling; fix stale parser and lexer comments (#5731)
* Share the diagnostic-position setter of the DOM SAX parsers

json_sax_dom_parser and json_sax_dom_callback_parser each had a private
copy of handle_diagnostic_positions_for_json_value(), identical except
for comments. Move the body into one static member function,
detail::diagnostic_positions::set_from_lexer(value, lexer), which both
classes call with their lexer pointer. basic_json befriends the new
struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position
members are private.

The discarded case is reached through the callback parser, so the
LCOV_EXCL markers that only the dom parser's copy had are gone. The
NOLINT on the unreachable default case loses the stray
"-warnings-as-errors", which is not a check name.

The start-position setup in start_object()/start_array() is left alone,
as #5706 is editing the callback parser's versions.

Behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the parser comments on recursion and skip_to_state_evaluation

The class documentation called the parser a recursive descent parser,
but sax_parse_internal() is a loop that keeps the open containers on an
explicit stack. The comment at the end of an array and of an object
said the flag is set to false while the code below it sets it to true.
Describe what the code does instead.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update the discard_number_values comments to the current number path

The comments explaining the accept() shortcut in convert_number() and
the member documentation still argued in terms of strtoull()/strtoll()
and errno, which #5283 replaced with convert_integer(), and pointed at
scan_number() instead of convert_number(). They also did not say that
scan_number_bulk_contiguous() converts integers itself, so the shortcut
is only reached for input without bulk access, with
JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back.

Rewrite both comments to describe the digit-count check in front of
convert_integer(), keeping the 18-digit bound and the json_sax_acceptor
argument. The stale <cstdlib> comment is left for after #5616, which
edits that include block.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* List the UTF-8 validators instead of calling the DFA the only one

The documentation of decode() called the Hoehrmann DFA the single
source of truth for UTF-8 validation. It is used only by the serializer
and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text
strings). The lexer's scan_string() switch, validate_one_utf8() /
valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer)
and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges
on their own.

Replace the sentence with a list of the four validators, what each is
used for, and a note that they must accept the same sequences. Sharing
code between them was considered and dropped: it would save a few lines
in a validator that is entangled with BON8 pushback, and #5677 is
editing the BON8 byte path.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale doc comments and include lists in the input headers

input_adapters.hpp included <memory> and <numeric> for the removed
shared_ptr-based adapter design but used neither; it called
(std::min) without including <algorithm>. json_sax.hpp used
std::numeric_limits without including <limits>. Also corrected
comments that no longer matched the code: input_stream_adapter does
not skip the input's BOM (the lexer's skip_bom() does), the
span_input_adapter comment named the no-longer-existing
input_buffer_adapter type, lexer::get_string() does not reset the
token, binary_reader's get_number() doc opened with /* instead of
/*! (so Doxygen skipped it) and omitted BON8 from its endianness
note, and the UBJSON-binary-types note did not mention that BJData
'B' arrays are read as binary.

Left out: the lgtm suppression on lexer.hpp's scan_number() (in
#5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's
start_object/start_array docs (in draft #5267's hunk), per the
verdict's conflict list.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the strict-EOF/release_lookahead/error block in parser::parse()

json_sax_dom_callback_parser and json_sax_dom_parser branches of
parser::parse() ran the same ~25 lines after sax_parse_internal():
the strict-mode EOF check (raising parse_error.101 through the SAX
parser), release_lookahead() in non-strict mode, and mapping an
errored SAX parser to a discarded result. The two copies had already
drifted apart in formatting and in the second copy's "see above"
comment.

Add a private parse_dom(DomSax&, strict) member that runs this shared
sequence once and returns whether the SAX parser did not error; both
branches of parse() now only construct their DOM SAX parser, call
parse_dom(), and (for the callback parser) map a discarded top-level
value to null. sax_parse() is left untouched, since it only runs the
EOF check and release_lookahead() when sax_parse_internal() succeeded,
unlike parse(), which runs them unconditionally.

Behavior-preserving: same operations in the same order for both SAX
parser kinds. Verified with unit-class_parser (strict/non-strict,
callback and non-callback), unit-deserialization and
unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean
make amalgamate / make check-amalgamation diff.

Overlaps #5601, which touches the same lines.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5712 item 2

* Share the code point to UTF-8 encoding between the wide-string helpers and the lexer

The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three
times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32
code point, in the UTF-16 helper for both a BMP code unit and a valid
surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The
copies had drifted: the UTF-32 helper masked the leading bits of each
byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift
alone, even though both give the same result for a code point that is
already known to be in range.

Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder
that invokes a callable once per output byte, most significant byte
first. Use it in the three valid-code-point branches (UTF-32 code
points up to U+10FFFF, UTF-16 code units outside the surrogate range,
and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where
out forwards to add(). The UTF-16 helper's deliberate pass-through of
malformed surrogate units and the UTF-32 helper's 0xFF sentinel for
code points above U+10FFFF are untouched, since neither reaches the new
helper.

Behavior-preserving: same bytes in the same order for every valid code
point, verified with unit-class_lexer, unit-class_parser,
unit-deserialization, unit-wstring and the non-test-data parts of
unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse
microbenchmark that shows no change (about 73 ms either way, median of
3, 1M escape sequences). single_include/ regenerated with make
amalgamate; make check-amalgamation leaves a clean tree.

Overlaps #5704, which rewrites the wide_string_input_helper
specializations touched here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5712 item 6

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:32:32 +02:00
Niels Lohmann 3926fcaac3 Deduplicate serializer dump code; fix stale includes, docs, and lint (#5729)
* Share scalar serialization between dump_internal and dump_value

dump_value()'s cases for string, binary, boolean, number_integer,
number_unsigned, number_float, discarded and null were a byte-for-byte
copy of dump_internal()'s (added together in #5285 for the iterative
fallback path). Any future change to scalar output had to be made in
both places, or the recursive and depth-limited paths would silently
start producing different bytes.

Extract the shared cases into a private dump_scalar() and have both
dump_internal() and dump_value() call it. Output is unchanged: dump(),
dump(4), dump(-1,' ',true) and the replace/ignore error_handler_t
variants are byte-identical over the json_test_data corpus before and
after, and dump() throughput on a scalar-heavy document is unaffected.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop serializer.hpp's dependency on binary_writer.hpp

The only use of binary_writer in serializer.hpp was
binary_writer<BasicJsonType, char>::to_char_type() to write the
U+FFFD replacement character's three bytes. With CharType=char this
is an identity conversion, so the include of binary_writer.hpp (and
transitively binary_reader.hpp) pulled in a large, unrelated header
for a no-op call.

Write the three bytes directly instead. serializer.hpp compiles
standalone with -Wall -Wextra -Werror, with and without
-funsigned-char, and unit-serialization's error_handler_t::replace
cases (with and without ensure_ascii) still pass. Moving
binary_writer's to_char_type/to_msgpack_length to its private section
is left as an optional follow-up.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale #include lines in the output headers

output_adapters.hpp included <algorithm> and <iterator> for std::copy
and std::back_inserter, which have not been used there since #3569
(2022). serializer.hpp included <algorithm> for std::reverse (also
unused), <cmath> for labs/isnan/signbit (only std::isfinite is used)
and <utility> for std::move (nothing from <utility> is used there),
while using std::next without including <iterator> at all, relying on
getting it transitively through output_adapters.hpp's own stale
<iterator>.

Drop the unused includes, add <iterator> for std::next, and correct
the remaining include comments. Both headers still compile standalone
with -Wall -Wextra -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove JSON_HEDLEY_NON_NULL(2) from write_characters() overrides

output_vector_adapter, output_stream_adapter and output_string_adapter
declared their write_characters(const CharType*, std::size_t) override
JSON_HEDLEY_NON_NULL(2), but binary_writer legitimately calls it with
a null pointer and length 0 for an empty string or binary value; the
type-erased call path only stayed silent under UBSan because the
static callee at those call sites is the unattributed virtual base.
A nonnull attribute on a definition lets GCC and Clang assume the
parameter is non-null inside the function body even when the call is
virtual, so this was latent undefined behavior, not just style.

Drop the attribute from the three overrides and document the
(nullptr, 0) contract on output_adapter_protocol::write_characters.
unit-cbor, unit-msgpack, unit-bson and unit-bon8 (which all exercise
empty binary/string payloads through the stream and vector/string
adapters) pass under -fsanitize=address,undefined,nonnull-attribute.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update stale serializer doc comments to match the current implementation

dump_internal()'s doc block still described the pre-#5285/#5449
implementation: an escape_string() function that does not exist
(the function is dump_escaped), integer conversion "implicitly via
operator<<" (dump_integer actually uses a digit-pair lookup table),
and floating-point conversion via "%g" (IEEE-754 types go through
to_chars, others through snprintf). dump_value()'s comment said
elements are pushed for dump_internal to walk, but it is
dump_iteratively() that walks the stack. dump_escaped(), dump_integer()
and dump_float() each said they write "to output stream @a o", which
has not been true since the writer moved to write_buffer.

Doc-only change; no behavior, API or ABI impact.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge duplicate byte-to-hex helper and drop stale '| 0' promotions

serializer::hex_bytes() and binary_writer::hex_byte() had identical
bodies. Keep one, detail::hex_byte() in string_utils.hpp, and use it
from both. Also drop the `| 0` at the two serializer call sites
(hex_bytes(byte | 0) and hex_bytes(s.back() | 0)): #3088 (7440786b8)
added it so that `ss << std::hex << (byte | 0)` printed a number
rather than a char with the old stringstream writer; the int result
just narrows back to uint8_t now, so it was a no-op.

Behavior is unchanged: unit-serialization, unit-bon8 (whose
type_error.316 messages exercise this code) and unit-diagnostics
pass, and both headers still compile standalone with -Werror.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Trim two stale lint suppressions in serializer.hpp

dump_integer()'s `auto buffer_ptr = number_buffer.begin();` carried
NOLINT entries for cppcoreguidelines-pro-type-vararg and hicpp-vararg,
left over from the snprintf-based implementation (#3088); there is no
variadic call on that line, so keep only the qualified-auto
suppressions it actually needs. remove_sign()'s assert checked
`x < 0 && x < (std::numeric_limits<number_integer_t>::max)()) `with a
NOLINT(misc-redundant-expression) to hide it; the second conjunct is
always true once x < 0, and has been since 6ce2f35ba (2019), so
reduce the assert to `x < 0` and drop the suppression instead of
masking it.

Both are documentation-only changes to assertions/suppressions, not
behavior. The to_chars.hpp `#if 0` branch this item also flagged is
left alone, next to draft PR #5634's pending hunk.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the Hoehrmann SPDX copyright line to string_utils.hpp

serializer.hpp carried the SPDX-FileCopyrightText line for Björn
Hoehrmann's UTF-8 decoder, but the decoder (decode() and the utf8d
table) has lived in string_utils.hpp since #5185 (d19f7f5dc);
serializer.hpp now only calls decode(). Move the copyright line to
where the code it covers actually is.

Part of #5709

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop calling std::localeconv() on every dump()

The serializer constructor snapshotted std::localeconv() into a
locale_chars member on every dump(), even though the only reader is
dump_float(number_float_t, std::false_type)'s snprintf path, taken
only for a number_float_t that is neither IEEE single nor double.
localeconv() is not required to be thread-safe with setlocale(), so
every dump() paid for and raced on a lookup that almost never mattered.

Remove locale_chars and the locale member. Right before the
thousands-separator/decimal-point fixups in the snprintf path, read
std::localeconv() into local thousands_sep/decimal_point variables
(null-checked, first byte only, as before) - the same way
lexer::get_decimal_point() already does since #5597. Output is
unchanged unless the locale changes during a single dump(); in that
case the fixups now match what snprintf just produced, instead of a
value snapshotted before the call.

Overlaps draft PR #5608, which touches the same constructor and
dump_float() lines to move this code into a new
dump_float_snprintf(); this lands the lookup change now as #5709 asks,
and #5608 can do the lookup inside dump_float_snprintf() when it
rebases.

Verification: the full json_test_data corpus (742 files, dump(),
dump(4) and dump(-1,' ',true)) is byte-identical to before the change
under the C locale. Added a test pinning the new per-conversion
lookup: it switches LC_NUMERIC mid-dump() (via a streambuf that
switches on its first write, after the serializer's write buffer has
been flushed once but before a later float is converted) and checks
the decimal point is still normalized using the locale active at
conversion time. On a platform where long double is IEEE-754 double
(e.g. 64-bit Arm), dump_float() takes the locale-independent
to_chars() path and the test is a no-op there; it is meaningful on a
platform where long double is extended precision (most x86 targets).

Part of #5709 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:31:42 +02:00
Niels Lohmann 18dd5663b0 Diff deeply nested values without recursing per nesting level (#5548)
* Diff deeply nested values without recursing per nesting level

diff() descended into both values once per nesting level, and compared
them with operator== on every level on the way, which recurses as well.
Values nested deeply enough - 25,000 levels on an 8 MiB stack - exhausted
the call stack and terminated the process, although parse() accepts
them without complaint. On such a chain the per-level comparisons and
path strings also made diff() quadratic in time and memory.

Both the recursion and operator== only descend as far as the source is
nested. So diff() first checks, recursing at most diff_depth_limit()
(128) levels, whether the source is nested more deeply than that. If not
- all but a vanishing minority of values - the recursive algorithm
diffs it exactly as before, now as diff_recursively(). Otherwise
diff_iteratively() walks the two values on an explicit stack, emitting
the same operations in the same order. It does not compare arrays and
objects with operator== up front (equal ones yield no operations
anyway), keeps the path in one buffer instead of a new string per
level, and hands every subtree that is not nested too deeply back to
diff_recursively(), so equal parts are still skipped quickly.

The check costs one pass over the source. On a 3,000-object document
that is about 30% of diffing two equal values (which is just an
operator== call), about 10% of diffing values that differ in a few
places, and noise when arrays change length. Once operator== no longer
recurses (#5390), the check can go.

Tests check that the patch reproduces the target at every depth up to
300, for json and ordered_json, including reordered members. They also
check the exact operation for a difference deep inside, and diff values
nested 100,000 levels deep.

Fixes #5393 for diff().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make diff_frame a member struct that declares its special members

GCC's -Weffc++ (an error in CI) asks a class with pointer members, a
user constructor and a non-trivial destructor to declare its copy
constructor and copy assignment; diff_frame's vector and basic_json
members make its destructor non-trivial. Declare all five as defaulted,
which also satisfies clang-tidy's special-member-functions check. Leave
their exception specifications implicit: GCC 4.8 rejects an explicit
one that differs from the implicit one, as it does for flatten_task in
#5517.

The converting constructor cannot throw, and is now declared noexcept
for GCC's -Wnoexcept, which flags the emplace_back() under C++26
otherwise. The struct also moves from diff_iteratively() into the class,
like dump_frame in the serializer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use the shared recursion limit in diff()

diff_depth_limit() is gone in favor of detail::recursion_depth_limit().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Diff fewer nesting depths so the test does not time out under Valgrind

Checking every depth up to 300 made test-json_patch exceed the 1500 s ctest
timeout in ci_test_valgrind. Check the depths up to 16, those around the
recursion limit of 128, and 300 instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark the diff frame's value-initialized members for clang-tidy

The braces are kept for GCC's -Weffc++, as in json_sax.hpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Bound diff()'s descent with a depth count instead of scanning the source

Now that operator== no longer recurses (#5390), diff() can keep its per-level
equality shortcut all the way down. It diffs recursively for the first
detail::recursion_depth_limit() levels, as merge_patch() does, and hands
anything deeper to diff_iteratively(). The nesting_exceeds() scan, which
cost about 30% on equal documents, is gone, and diff() is on par with
develop again.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Note that the diff frame reference is invalidated by pop_back() too

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep diff()'s recursive levels small and its result elided

diff_recursively built every patch operation in place from initializer
lists. Unoptimized builds give each of those temporaries its own stack
slot, so every level of the bounded descent cost kilobytes of stack
(about 6 KB with clang -O0), and the 128 recursive levels overflowed the
1 MB stack of MSVC Debug in the "deeply nested values" test. The
operations and the key comparison of two objects are now built by
separate functions, which diff_iteratively shares, and both diff
functions append to one result instead of returning a patch per level
that the caller copies. With clang -O0, diffing values nested 300 levels
deep now peaks at about 190 KB of stack instead of 880 KB.

Since diff() now owns the only returned value, clang's -Wnrvo no longer
reports the returns of diff_recursively, which alternated between the
local patch and diff_iteratively's result.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Copy the diff frame's members instead of holding a reference to it

The loop in diff_iteratively held a reference to the top frame, which
enter() invalidates when it pushes and the end of the loop invalidates
when it pops. Nothing used it afterwards, but a later change could. As in
the other iterative walks, the members the loop reads are now copied out
as constants and the ones it advances are changed through stack.back().
The frame as a whole is not copied: it holds the common keys and the
"add" operations of an object.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:29:32 +02:00
Niels Lohmann a32f61eb98 Fix update() and merge_patch() when the argument is *this or one of its members (#5678) 2026-09-30 23:04:57 +02:00