Commit Graph
5212 Commits
Author SHA1 Message Date
Niels Lohmann 7e8e8e219b Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/input/binary_reader.hpp
#	include/nlohmann/json.hpp
#	single_include/nlohmann/json.hpp
2026-09-30 22:53:03 +02:00
Niels Lohmann c261578431 Deduplicate binary reader/writer helpers and fix stale comments (#5730)
* Fix stale and missing comments in binary_writer

The doc block of write_number() ended up above the byte_swap() helpers
added in #5286, about 80 lines from the function. It was also a plain
comment that Doxygen skips, said "write a number to output input", and
left BON8 out of the big-endian formats. Move it back onto
write_number() as a /*! block and fix the text.

write_bson() documented "@pre j.type() == value_t::object", but it
throws type_error.317 for every other type, and to_bson() relies on
that. Document the exception instead.

Explain why the CBOR binary subtype is always written with a 0xD8..0xDB
head and never in the one-byte tag form: binary_reader with
cbor_tag_handler_t::store only keeps those heads as a subtype, so
switching to write_cbor_head() would break round trips for subtypes
0..23.

Also fix the grammar of the to_char_type comment. Comments only; no
change in behavior, API or ABI.

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge the duplicated UBJSON/BJData integer marker ladders

write_number_with_ubjson_prefix() (unsigned and signed overloads) and
ubjson_prefix() (number_integer and number_unsigned cases) each picked
the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their
own independent if/else ladder, and the values beyond 64 bits were
handled by a second, tag-dispatched pair of ladders. An optimized
container announces the marker of its first element via ubjson_prefix()
and then writes every element through write_number_with_ubjson_prefix(),
so the two had to be kept in lockstep by hand across four call sites.

Replace all of that with one ubjson_integer_prefix() built on
value_in_range_of<T>, and one write_ubjson_integer_payload() that
writes the value (or, for 'H', the decimal digits) for a given marker.
write_number_with_ubjson_prefix() and ubjson_prefix() keep their
signatures and now just call these two helpers.

Behavior, the public API and the ABI are unchanged. Verified with a
new regression test covering scalars and $-optimized arrays/objects at
every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for
to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing
to_ubjson/to_bjdata output before and after over the json_test_data
corpus (bit-identical).

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove dead get_char parameters in binary_reader

The non-recursive rewrite of the binary readers (#5505, #5506, #5507)
left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char
parameters dead: parse_cbor_internal() has one caller and it always
passes true, and parse_ubjson_internal() has one caller and it always
uses the true default. Both parameters, and the @param docs describing
the "reuse the last character" mode they used to select, no longer
correspond to anything.

Drop both parameters, initialise fetch/prefix unconditionally, and
update the two call sites in sax_parse(). parse_cbor_value()'s and
get_ubjson_string()'s own get_char parameters are unrelated and are
left alone; both still have a false caller.

Also delete a stray `@return whether a valid MessagePack value was
passed to the SAX parser` doxygen block that sits directly above
parse_msgpack_value()'s real doc comment, a leftover of the same
rewrite.

Behavior, the public API and the ABI are unchanged; these are private
members of detail::binary_reader. Verified by compiling with
-Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and
unit-msgpack (offline, against the stubbed test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Share the IEEE half-precision decoder between CBOR and BJData

binary_reader had two ~45-line copies of the IEEE 754 half-precision
decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is
normalised, the two blocks were identical except for the byte order
used to assemble the 16-bit half (CBOR is big endian, BJData is little
endian). Any future change to half-float decoding had to be made and
kept in sync in both places.

Add one get_half_float(format, little_endian) helper that does the two
get()/unexpect_eof() reads, assembles the half in the requested byte
order, decodes it per RFC 8949 Appendix D, and calls sax->number_float.
Both cases now just call it with their byte order; the BJData case
keeps its bjdata-only guard.

Behavior, the public API and the ABI are unchanged. Verified with a
scratch probe comparing the old and new decoders bit-for-bit (NaN by
isnan()) over all 65536 wire byte pairs, in both formats, and by
running unit-cbor and unit-bjdata (offline, against the stubbed
test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the MessagePack unsigned-integer writer ladder

The number_integer (non-negative branch) and number_unsigned cases in
write_msgpack() each held their own copy of the fixint/uint8/16/32/64
ladder, kept in lockstep only by a comment ("we used the code from the
value_t::number_unsigned case here"). Both copies mixed union members:
the signed copy compared number_unsigned but wrote number_integer, and
vice versa.

Extract write_msgpack_unsigned(std::uint64_t), mirroring how
write_cbor_head() already avoids the same duplication for CBOR, and
call it from both cases. Each case now reads only its own active
union member. Output bytes are unchanged for the default 64-bit
number types.

#5710 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Unify float marker selection and fix the long double compile error

Four formats picked between a float32 and float64 marker through four
different helper styles: dummy-argument overloads for CBOR and
MessagePack, an std::is_same template for BON8, and a runtime if-chain
on input_format_t for write_compact_float(). With number_float_t set
to long double, to_cbor, to_msgpack and to_ubjson failed inside the
library with "call to 'get_cbor_float_prefix' is ambiguous", while
to_bson kept working because write_bson_double() takes a plain double.

Change write_compact_float() to take the two marker bytes directly
(each of its three callers already knows them at compile time) instead
of an input_format_t it only forwarded, and delete the now-unused
get_cbor_float_prefix(), get_msgpack_float_prefix(),
get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn
the two get_ubjson_float_prefix() overloads into one template. Both
write_compact_float() and get_ubjson_float_prefix() now report an
unsupported number_float_t with a static_assert naming the requirement,
rather than an ambiguous-overload error; the assert lives in the
function body, not the class scope, so to_bson with long double is
unaffected.

Verified with a probe basic_json<..., long double>: to_bson still
compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail
to compile with the new static_assert message.

This changes the text of an existing compile error for users with an
unsupported number_float_t (documented as a public-API-visible change
in #5710).

#5710 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the BJData ndarray writer's dtype dispatch and drop <map>

write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on
every call just to translate the _ArrayType_ name to a dtype marker
(the only reason binary_writer.hpp included <map>), then mapped dtype
to C++ type twice more: once as a switch for the range-check pass and
once as a separate if/else chain for the write pass, with nothing
checking that the two agreed. The caller also ran three at() lookups,
and the callee called value.at(key) about ten more times for the same
three members.

Replace the map with bjdata_ndarray_type_marker(), a plain string
comparison chain (a C++11 constexpr function cannot contain a switch,
so this mirrors binary_reader's own static table style). Replace the
switch/if-chain pair with one write_bjdata_ndarray_elements() that
switches on dtype once and calls a per-type helper -
write_bjdata_ndarray_element<T>() for the eight integer dtypes and
write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag
selecting the range check or the actual write, so the two passes can
no longer disagree on the type. _ArrayType_, _ArraySize_ and
_ArrayData_ are now looked up once into references, and the four
header marker bytes ('[', '$', '#') are written through to_char_type()
like the rest of the UBJSON/BJData writer.

The 'd' (single-precision) rule is left exactly as before, since #5707
is expected to change it separately.

Verified byte-for-byte identical output before/after for every dtype
(including the Draft 2/Draft 3 'byte' fallback and the use_count/
use_type combinations) via a standalone probe, plus round-tripping
through from_bjdata().

Overlaps #5707, which is expected to touch the 'd' dtype case, and
#5518, which is expected to move the write_bjdata_ndarray() call site.

#5710 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Assert that write_bson_document() consumes every calc_bson_sizes() entry

calc_bson_sizes() and write_bson_document() are a hand-synchronized
pair of passes over the same object/array tree, introduced by #5553:
the size pass appends to nested_sizes in visiting order, and the write
pass consumes the table by position with nested_sizes[next_size++].
Nothing checked that the write pass consumed the whole table. If a
future change touched only one of the two passes - for example to skip
or reject an entry - every later size prefix in the document would be
silently wrong.

Add JSON_ASSERT(next_size == nested_sizes.size()) where
write_bson_document() returns, so such a future drift between the two
passes is caught immediately (JSON_ASSERT expands to nothing in
release builds using assert(), and the fuzzers/tests already build
with it enabled). The two passes agree today, so this changes nothing
observable; it only guards against the risk described in #5710 item 5.

Extracting a shared stepper for the two passes (the second half of the
proposed change) is left for a follow-up: it only saves ~30 lines and
the issue asks for it only if the result reads clearly, which needs
more room to get right than a mechanical cleanup pass allows.

#5710 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the BJData lookup tables static functions instead of members

binary_reader held bjd_optimized_type_markers and bjd_types_map as
non-static const members (12 string_t objects for the type-name table),
built and destroyed on every from_cbor/from_msgpack/from_bson/
from_ubjson/from_bon8/from_bjdata call even though only from_bjdata
ever reads them. They also needed the #define/decltype/#undef
workaround from #3637 and two NOLINTNEXTLINE suppressions, and
binary_writer already carries the same two lists in another form
(is_bjdata_excluded_type_marker() and a local std::map in
write_bjdata_ndarray(), the latter removed by the item-4 commit), so
the excluded-marker lists could drift apart.

Replace bjd_optimized_type_markers with static constexpr
is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain
as binary_writer's is_bjdata_excluded_type_marker(). Replace
bjd_types_map with a non-constexpr static bjd_type_name(char_int_type)
switch returning nullptr for an unknown marker (a C++11 constexpr
function cannot contain a switch). Delete both
JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the
NOLINTNEXTLINE suppressions, detail::make_array() (no longer used
anywhere), and the now-unused <algorithm> and <array> includes.

Update the two call sites (the ND-array excluded-type check and the
_ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's
"LUT arrays are sorted" section, which only checked the two tables'
internal ordering, with a check of all 12 type names and all 8
excluded markers against both new functions.

#5711 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read CBOR's 1/2/4/8-byte argument through one helper

parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte
big-endian unsigned integer" ladder four times over:
- twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore
  branch and once, nearly identically, in the ::store branch (~90
  lines to read one integer);
- twice more for container lengths, once for array heads 0x98-0x9B and
  once for map heads 0xB8-0xBB, where the 1/2-byte forms called
  enter_array()/enter_object() directly and the 4/8-byte forms
  additionally went through get_cbor_container_size().

Add get_cbor_argument(std::uint64_t&), reading the width selected by
current & 0x1F via the same get_number() calls as before (so EOF is
reported exactly as before), and route all four sites through it:
- 0xD8-0xDB now read the argument once per branch instead of switching
  on `current` a second time; behavior split cleanly from embedded tags
  0xC0-0xD7 (tag value in the head, no argument to read), which is now
  its own case block that no longer has to fall into the ::store
  switch's "default" case to reach the same tag_pending = true; return
  true; outcome.
- 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always
  going through get_cbor_container_size() (harmless for 1/2-byte
  lengths, which already always fit).

Verified byte-for-byte identical behavior before/after with a
standalone probe covering embedded and multi-byte tags under all three
tag_handler_t settings, a tag over a byte string (subtype path),
truncated tag/length arguments of every width, and array/map lengths
of every width, including the out_of_range.408 "excessive size" case:
same exceptions, same messages, same chars_read, same successful
results.

Left the string/byte-string length ladders in get_cbor_string()/
get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is
expected to touch them separately.

Overlaps #5601 (adds a branch right above the embedded-tag case) and
#5607 (touches the integer cases 0x18-0x1B, which share this ladder's
shape in separate hunks).

#5711 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add leave_container() to match enter_container()

Every container is opened through enter_container(), whose docs
promise that a check placed there runs before every start event. The
close side had no equivalent: the same
"container_stack.pop_back(); dispatch to end_object() or end_array()"
sequence was written out separately in BSON, CBOR, MessagePack,
UBJSON/BJData and BON8, each copying the pattern of keeping an
is_object flag around the pop_back() that would otherwise invalidate
a reference to it. A check needed on close would have had to be added
in five places, and a sixth copy could go unnoticed.

Add leave_container() next to enter_container(), doing the same
pop-then-dispatch, and replace the five sites with it. Each site keeps
its own surrounding logic (BSON's check_bson_document_size() call
before popping, MessagePack's is_object copy used again below,
UBJSON/BJData's remaining-container handling after popping, BON8's
top used again below); only the repeated pop/dispatch line pair is
now shared.

Verified all six binary-format unit suites and unit-regression2's
deep-nesting tests (dependent count/reuse count and the bjdata ndarray
depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan.

Overlaps #5601, which is expected to add a sixth close site in its own
skip loop; that site can route through leave_container() too once it
lands.

#5711 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop passing the input format to sax_parse() when the reader already has it

binary_reader's constructor stores the format in the input_format
member, and sax_parse(format, sax_, strict, tag_handler) took the same
value again purely to dispatch on it. Every in-tree caller passed the
same value both times (all 16 from_cbor/from_msgpack/from_ubjson/
from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three
public basic_json::sax_parse() overloads), so nothing was broken
today, but a caller of the detail class directly (only reachable via
JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could
pass a mismatched pair - say bjdata to the constructor and ubjson to
sax_parse - and dispatch on one format while applying the other
format's rules; the default-constructed input_format_t::json reader
would additionally hit JSON_ASSERT(false) in exception_message() on
its first error.

Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the
existing overload with the stored input_format, and switch every
caller to it: the 16 from_*() sites (keeping their
`// cppcheck-suppress[accessMoved]` comments) and the three
basic_json::sax_parse() overloads, all of which already had the format
available from their own `format` parameter. The four-argument overload
is kept for anyone still calling it, now with
JSON_ASSERT(format == input_format) so a mismatch fails immediately
in a debug build (assert-enabled binaries, including the fuzzers and
test suite) instead of misbehaving; verified with a probe that
constructs a reader for one format and calls the explicit overload
with another, which aborts on that assertion as expected.

Removing or asserting against the constructor's input_format_t::json
default, which would affect direct detail users, is left as a separate
decision per #5711 item 5.

Overlaps #5601, which is expected to add an AllowRecovery template
parameter to sax_parse() and touch these same call sites in json.hpp.

#5711 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate UBJSON/BJData signed-count handling, drop dead ndarray checks

get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently
sized signed integer and then repeated the same "reject negative with
error 113" check; only 'L' additionally checked value_in_range_of for
the out_of_range.408 case. Any change to that error path had to be
made four times.

Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read,
the negative check and the range check once, and route all four
markers through it. The range check is a no-op for 'i'/'I'/'l' (their
values always fit std::size_t) and only live for 'L' on a 32-bit
std::size_t target, matching today's behavior exactly.

In the ndarray dimension-product loop, the preceding loop already
returns early on any zero dimension and result starts at 1, so `i > 0`
in the pre-multiplication overflow check was always true, and
`result == 0` in the post-multiplication check could not be reached
either: two positive factors whose product does not overflow (as the
pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&`
and narrow the post-check to `result == npos`, the one case the
pre-check cannot rule out (an exact, non-overflowing match with the
sentinel reserved for unknown-size containers), with a comment
explaining why.

Verified byte-for-byte identical behavior before/after with a
standalone probe covering negative counts for every marker, a matching
positive count, and ndarray inputs, plus the full unit-ubjson and
unit-bjdata suites (same assertion counts as before this change).

Overlaps #5601 (rewrites the four parse_error calls and the overflow
checks touched here) and #5607/#5707 (touch neighboring lines in the
same functions).

#5711 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop redundant format parameter and dummy float argument (review)

binary_reader::sax_parse(format, ...) only ever had to equal the format
given to the constructor, which it asserted. With every caller already
on the format-less overload, remove the four-argument overload and
dispatch on the stored input_format directly. binary_reader is a
detail class, so this is not a public API change.

get_ubjson_float_prefix() took a value only to deduce its type; make
the type an explicit template argument instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:48:32 +02:00
Niels Lohmann 35802e78d6 Remove dead Makefile targets; document macro_builder; tidy serve_header (#5735)
* Remove dead doctest help entry and pretty_format target from Makefile

The top-level Makefile still carried three leftovers:

- The help text listed a "doctest" target that was removed in #4560,
  so "make doctest" fails with "No rule to make target". The example
  check now runs as "make check_output -C docs".
- "pretty_format" ran clang-format on all sources, but .clang-format
  was deleted in #4573, so the target reformatted everything in the
  default LLVM style, against the Artistic Style formatting that
  "make pretty" applies and CI enforces.
- "clean" removed benchmarks/files/numbers/*.json, a directory that no
  longer exists since the benchmarks moved to tests/benchmarks (#3462).

Only maintainer tooling changes; the library is not affected.

Part of #5717

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document tools/macro_builder and tidy up serve_header.py

tools/macro_builder generates the NLOHMANN_JSON_EXPAND,
NLOHMANN_JSON_GET_MACRO and NLOHMANN_JSON_PASTE* macros in
macro_scope.hpp, but nothing referred to it. Add a README that explains
what it generates, how to run it and where the output goes, and which
dependent tables (NLOHMANN_JSON_DOUBLE_PASTE, NLOHMANN_JSON_TYPE_BODY)
are maintained by hand. Point to it from a comment above
NLOHMANN_JSON_EXPAND. The generator itself is unchanged; following the
README reproduces the header byte for byte.

In serve_header.py, drop the LGTM suppression (LGTM.com shut down in
2022), replace the """.""" placeholder docstrings with real ones, and
import socket and ssl at module level. DualStackServer.server_bind uses
socket, which was only imported under __main__; when the module was
imported instead, the NameError was swallowed and IPV6_V6ONLY was not
cleared.

The header change is a comment only; behavior, API and ABI are
unchanged.

Part of #5717

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Hash every release_files artifact, not a hardcoded subset

The `release` target signed and copied json_fwd.hpp into release_files
alongside json.hpp, but the shasum line that writes hashes.txt only
listed json.hpp, include.zip and json.tar.xz. Users could not verify
the published json_fwd.hpp against hashes.txt.

Hash every file in release_files except the .asc signatures instead
of naming files by hand, so a newly shipped header (such as the
json_literals.hpp that #5610 adds to this target) cannot be missed
again.

Only affects the generated hashes.txt release artifact; the library
itself is unaffected.

Overlaps #5610, which touches the same lines to add json_literals.hpp
to the release target.

#5717 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the broken fuzz_testing* Makefile targets

fuzz_testing and fuzz_testing_{bon8,bson,cbor,msgpack,ubjson} seeded
fuzz-testing/testcases from tests/data, which was removed in dbf1a1f41
(2020) when the test data moved to the external json_test_data repo.
The find command found nothing, but the pipeline's exit status was
that of xargs, so the recipe still reported success with an empty
corpus, and the printed afl-fuzz command would refuse to start.

The recipes were also six near-identical copies with unquoted -name
patterns, used the legacy CXX=afl-clang++, and fuzzing-start/stop were
missing from both the help output and .PHONY. tests/fuzzing.md already
documents the working flow (download json_test_data, then
`make -C tests fuzzers`), so replace the six broken targets and their
help lines with a single pointer to that document instead of trying
to keep six copies of a fragile shell pipeline in sync.

This does not affect OSS-Fuzz, which builds through tests/Makefile.

Overlaps #5621, which adds a seventh copy of the same broken line for
fuzz_testing_json_view.

#5717 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove stale Travis comment above check-amalgamation

check-amalgamation carried "Note: this target is called by Travis",
left over from before the project switched off Travis CI. The prior
Makefile cleanup commit removed the other stale Travis-era leftovers
(the doctest help entry, pretty_format, and the benchmarks/ path in
clean) but missed this comment.

#5717 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale install/usage instructions in the vendored amalgamate README

tools/amalgamate/README.md is the unmodified upstream text and no
longer matches how the tool is used here:

- It named a Bitbucket origin that no longer exists; CHANGES.md
  already tracks the GitHub mirror commit this copy is based on.
- It asked for Python 2.7, but CI and the Makefile run the script
  with python3.
- It told readers to run ./test.sh (not vendored) and install to
  /usr/local/bin; in this repository the tool runs through
  `make amalgamate`.
- Its usage synopsis showed `-v` taking no argument, but the script's
  own argparser requires `choices=["yes", "no"]`, so that form fails
  with "argument -v/--verbose: expected one argument". The Makefile
  calls it as `--verbose=yes`.
- It pointed at test/source.c.json and test/include.h.json, which are
  not vendored; the configs actually used are config_json.json and
  config_json_fwd.json.

Rewrote only the Installing and Using sections to match; left the
"Here be dragons" caveats and the rest of the vendored code untouched
to avoid diverging further from upstream.

Overlaps #5615, which edits amalgamate.py, this README and CHANGES.md.

#5717 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Derive generate_natvis.py's ABI tag list and version from abi_macros.hpp

generate_natvis.py hard-coded abi_tags = ['_diag', '_ldvcmp', '_dp',
'_bics', '_psp', '_snul'] and required --version on the command line.
The source of truth is include/nlohmann/detail/abi_macros.hpp: the
NLOHMANN_JSON_ABI_TAG_* defines, the argument order of
NLOHMANN_JSON_ABI_TAGS_CONCAT, and NLOHMANN_JSON_VERSION_MAJOR/MINOR/
PATCH. Nothing checked that the copies stayed in sync, and they have
drifted apart before: _dp was added in #4517 but missed here until
#5544, and 3.11.3 shipped json_abi_v3_11_2 namespaces (#4340).

Parse the tag list (in NLOHMANN_JSON_ABI_TAGS_CONCAT order) and the
version from abi_macros.hpp instead of hard-coding them. Make
--version optional (falling back to the parsed version) and default
the output directory to the repository root the script lives in.

Add a "natvis" Makefile target that runs the script, and extend
check-amalgamation to regenerate nlohmann_json.natvis and fail on a
diff, the same way it already does for the amalgamated headers and
BUILD.bazel. Wire the same regeneration into check_amalgamation.yml,
using the tool copy checked out from develop (as the workflow already
does for amalgamate.py) and installing jinja2 from
tools/generate_natvis/requirements.txt. Update the tool's README to
say it must be re-run after adding an ABI tag or bumping the version.

Verified: a run against develop produces no diff (with either the
default or an explicit --version 3.12.0); adding a dummy
NLOHMANN_JSON_ABI_TAG_* without a matching #define makes the script
fail loudly instead of silently omitting the tag; xmllint --noout
passes on the regenerated file; and running the script from a
directory other than the one being checked (simulating the workflow's
separate tool checkout) against this repository root also produces no
diff.

Overlaps #5600, which added _ekmo to the same hand-written abi_tags
line and regenerated the file.

#5717 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Strip the leading "./" find(1) prefix from release hashes.txt entries

bc6e7db72 (#5717 item 1) switched the release target's shasum line from
naming files by hand to $$(find . -type f -not -name '*.asc' | sort),
so a newly shipped header is hashed automatically. Run from inside
release_files, that find prints paths as "./json.hpp" instead of
"json.hpp", so hashes.txt lists "./json.hpp" etc. instead of the plain
filenames it always used. shasum -c still verifies "./json.hpp" fine,
but it is a needless cosmetic regression for anyone reading the file
or matching it against release notes.

Strip the "./" prefix with sed before sorting, keeping the filenames
exactly as before while still hashing every artifact automatically.

Review fix for #5717 item 1 (PR #5735).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix check_amalgamation.yml: pass --version to generate_natvis.py

bff45f111 (#5717 item 3) made --version optional in
tools/generate_natvis/generate_natvis.py and wired the workflow's new
"Regenerate nlohmann_json.natvis" step to call it without --version,
relying on the script deriving the version from abi_macros.hpp itself.

But NATVIS_TOOL_DIR is checked out from develop, the same way TOOL_DIR
already is for amalgamate.py, precisely so an in-flight PR's tooling
changes cannot mark themselves clean. Until this PR (or an equivalent)
merges to develop, that checkout is the old generate_natvis.py, whose
--version argument is still required=True. The new step's invocation
of "generate_natvis.py $MAIN_DIR" (no --version) then fails argparse
on this PR's own CI run with "the following arguments are required:
--version", before the check ever gets to compare output.

Extract the version from $MAIN_DIR's own abi_macros.hpp in the
workflow and always pass it as --version. That satisfies the old
script's required argument and is accepted as an explicit override by
the new one, so the step behaves the same whether NATVIS_TOOL_DIR holds
the pre- or post-merge tool, and stays correct for later PRs that bump
the version.

Verified by running the workflow step's shell logic locally against
both the pre-#5717 generate_natvis.py (checked out at 633de8e44) and
the new one: both produce the identical nlohmann_json.natvis as the
committed file.

Review fix for #5717 item 3 (PR #5735).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Extend tools/macro_builder to also generate DOUBLE_PASTE and TYPE_BODY

tools/macro_builder only emitted NLOHMANN_JSON_EXPAND..PASTE64, so two
other tables that scale with the same max_args stayed hand-maintained
with nothing checking them: NLOHMANN_JSON_DOUBLE_PASTE (added by hand
in #4563 for the *_WITH_NAMES macros) and the 64-slot dispatch table of
NLOHMANN_JSON_TYPE_BODY (the #4041 zero-member/one-or-more-member
switch). Both tables pass one macro name per slot to the same
NLOHMANN_JSON_GET_MACRO dispatch as PASTE, so they can drift out of
sync with max_args exactly the way _dp did in the ABI tag list fixed
by #5544.

Extend main.cpp with build_double_paste_code() (same recursive-doubling
shape as build_paste_code(), but DOUBLE_PASTE consumes two arguments
per member, so an even slot index falls back to the next lower odd
DOUBLE_PASTE<N>) and build_type_body_table() (max_args - 1 MEMBERS
slots and one trailing EMPTY slot, 8 per line, matching how it is
written by hand today). Add a "type_body" argument that selects the
TYPE_BODY block, since it lives at a separate location in
macro_scope.hpp from the EXPAND..DOUBLE_PASTE63 block; plain invocation
is unchanged apart from covering the extended range. No longer emit
the tool's old trailing blank line, so its output is directly diffable
without post-processing.

Verified with c++ -std=c++11: running the tool (with and without
"type_body") and piping the raw output through the pinned astyle
reproduces both blocks of the current macro_scope.hpp byte for byte.
tests/src/unit-udt_macro.cpp (all NLOHMANN_DEFINE_TYPE_*/_WITH_NAMES/
zero-member variants) passes unchanged under -std=c++11 and -std=c++17
with -fsanitize=address,undefined.

Add a "macro_builder_check" Makefile target that builds main.cpp,
regenerates both blocks into a scratch directory inside the repository
(astyle's --project lookup needs the target files under the same tree
as .astylerc, unlike an external /tmp directory), and diffs them
against the corresponding ranges of macro_scope.hpp; wire it into
check-amalgamation next to the natvis check. Wire the same regeneration
into check_amalgamation.yml, splicing the (still unindented) generated
blocks back into the PR's own macro_scope.hpp before the existing
astyle/amalgamation step runs, so that step's own tree-wide astyle
pass both indents them and folds any drift into the amalgamation
patch/diff the workflow already produces.

Unlike amalgamate.py and generate_natvis.py, this step builds
tools/macro_builder/main.cpp from the pull request's own checkout
($MAIN_DIR) rather than a separate checkout of tools/ at develop: this
tool has no independent source of truth to regenerate against (its
README documents that it must reproduce macro_scope.hpp byte for
byte), so a develop-pinned copy would only reproduce the
generate_natvis.py trap fixed in a previous commit on this branch,
where a PR that teaches the tool to cover more of the file fails its
own CI until that PR merges and updates the develop copy.

Add tools/macro_builder/README.md documentation for both new tables
and the two-invocation usage, and a short pointer comment above
NLOHMANN_JSON_TYPE_BODY (the EXPAND pointer already covered the first
block; extended its wording to include DOUBLE_PASTE63).

Closes #5717 item 5 in full, completing what the documentation-only
"Document tools/macro_builder..." commit already on this branch left
open (that commit's README/pointer-comment half stands; it also covers
item 7).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:21:27 +02:00
Niels Lohmann 791cd88dfc Fix raw-TeX formulas and other documentation infrastructure debt (#5736)
* Fix math formulas rendering as raw TeX in the published docs

The privacy plugin self-hosts MathJax 2.7.0 but drops its
?config=TeX-MML-AM_CHTML query string, so the rehosted script loads no
input jax and the 10 formulas across 6 pages render as raw TeX to
readers. Remove pymdownx.arithmatex and the MathJax extra_javascript
entry, and rewrite the formulas in plain HTML (<sup>, <i>) instead.
This also drops a nine-year-old third-party script from every page.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix publish_documentation triggers and persist unneeded git credentials

publish_documentation.yml only triggered on docs/mkdocs/** pushes, but
the site also embeds .github/CODE_OF_CONDUCT.md, CONTRIBUTING.md,
SECURITY.md, cmake/{clang,gcc}_flags.cmake, .clang-tidy,
tools/astyle/.astylerc and tests/fmt_formatter/project/main.cpp via
pymdownx.snippets, so changes to those files never republished the
site. Extend the path filter to cover them, and switch runs-on from
the long-pinned ubuntu-22.04 to ubuntu-latest to match
ci_test_documentation.

Also add persist-credentials: false to the checkouts in ci_icpx and
ci_nvhpc (ubuntu.yml) and msvc-vs2026/msvc-arm64 (windows.yml), none
of which pushes with git, matching every other checkout in these
workflows.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix example build: broken debug echo, deprecations hidden for everything

docs/Makefile's debug echo used a space instead of a comma in
$(call cxx_standard ...), so it always printed an empty standard.
Every example was also compiled with -Wno-deprecated-declarations,
which would silently hide an accidental deprecated-API call in any of
them. Factor the duplicated compile flags into EXAMPLE_CPPFLAGS/
EXAMPLE_WARNFLAGS, build with -Werror=deprecated-declarations by
default, and only allow the three examples that intentionally
document deprecated API (the DEPRECATED_EXAMPLES list) to suppress it.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix check_structure.py NOLINT parsing and an off-by-one report line

`line.strip("<!-- NOLINT")` followed by `.strip(" -->")` strips any of
the characters in those sets from both ends, not a literal prefix; it
only happened to work for "Examples". A NOLINT'd section name
starting with N, O, L, I or T (e.g. "Notes", "Template parameters",
"Iterator invalidation", "Literals") was silently mangled, so the
suppression did not apply and the checker could report a spurious
missing/misordered section. Parse the comment with a regex instead.
The same fragile strip() pattern was used for heading text; replace it
with a plain prefix slice. Also fix the admonition_title report, which
used the 0-based line index while every other report in the file uses
lineno+1.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix docset README's non-existent make target and stale fallback URL

The README told readers to run `make nlohmann_json.docset`, but the
Makefile's targets are `all`, `JSON_for_Modern_C++.docset` and
`install_docset_zeal`; the documented command has failed with "No
rule to make target" since #2967 (2021). Point the README at the real
target and folder name. Info.plist's DashDocSetFallbackURL also still
pointed at the old nlohmann.github.io/json/ URL instead of the
canonical https://json.nlohmann.me/ from mkdocs.yml's site_url.

Leave list_missing_pages/list_removed_paths alone: they may become
redundant once #5638's check_docset() lands, which is a follow-up.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove two leftover Doxygen-era .link files from the examples directory

parse__iterator_pair.link and parse__pointers.link each held only a
Wandbox "online" permalink from the old Doxygen docs. #3071 deleted
every other .link file in 2021; these two came in through a parallel
PR (#3100) and were never referenced by any page, script or config.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix MacPorts CMake example to include its own snippet files

The MacPorts "Example: CMake" block included
integration/homebrew/example.cpp and
integration/homebrew/CMakeLists.txt instead of the MacPorts files
right next to it, a copy-paste slip from the Homebrew section. Nothing
referenced integration/macports/CMakeLists.txt as a result. The page
rendered correctly only because the homebrew, macports and
vcpkg/CMakeLists.txt snippets are byte-identical, so a future edit to
the MacPorts files would not have shown up on the page.

Part of #5718 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not persist git credentials in publish_documentation's checkout

The checkout step in publish_documentation.yml left the default
persist-credentials: true, so GITHUB_TOKEN stayed writable in
.git/config for the rest of the job (zizmor's artipacked finding). The
Deploy documentation step authenticates through its own github_token
input to peaceiris/actions-gh-pages and does not push with the
checked-out credentials, so persist-credentials: false is safe here,
matching every other checkout in the workflow set.

Overlaps #5638, which edits this same checkout step (adds
fetch-depth: 0); expect a rebase conflict there.

Part of #5718 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Replace list_missing_pages/list_removed_paths with a comm(1)-based diff

The docset Makefile's list_missing_pages ran one sqlite3 query per
mkdocs page, and list_removed_paths nested a loop over all mkdocs pages
inside a loop over all docset index paths (O(n*m) shell iteration).
Issue #5718 item 5 suggested removing or reducing these targets once
#5638's check_docset() lands, but that PR is still open and covers only
API pages and macros, not the full page set these targets check.
Replace the loops with two sorted path lists (DOCSET_PAGE_PATHS from
mkdocs' markdown sources, DOCSET_INDEX_PATHS from the built docset
index) compared with a single comm(1) call each, verified to produce
output identical to the old loops against the current docSet.dsidx.

The sed expression used '#' as its delimiter, which GNU Make reads as
a comment character even inside a variable assignment, truncating the
line and orphaning the closing paren of $(shell ...) ("unterminated
call to function 'shell': missing ')'"). Use '@' as the delimiter
instead.

Part of #5718 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:15:20 +02:00
Niels Lohmann 7d7055ec50 Fix stack overflow converting deep values between specializations (#5723)
* Fix stack overflow converting deep values between specializations

Constructing a basic_json from another specialization (json to
ordered_json or back, also via get<ordered_json>()) converted every
container with its range constructor, which calls the converting
constructor for each element. The call stack therefore grew with every
nesting level, and a value nested some 30,000 levels deep overflowed it.

The conversion now bounds its descent the way the copy constructor does
since #5387: the first 128 levels are converted exactly as before, and
below that convert_iteratively() finishes the value with an explicit
stack. It builds each container bottom-up from its converted elements
with the container's range constructor, so member order and keys that
become equal are handled as before, and it gives a value its type only
once its container exists, so an exception leaves nothing behind that
cannot be destroyed. Parents (JSON_DIAGNOSTICS) and positions
(JSON_DIAGNOSTIC_POSITIONS) are set for every value.

Converting a null value no longer resets its positions: the constructor
assigned null to a value that already was null, which swapped in the
positions of the temporary.

Fixes #5650.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Explain why converting null keeps positions and why next is a reference

Review feedback on #5723 (gregmarr): clarify in comments that the
converting constructor has already copied the positions of val, which
the null case keeps like every other case, and that next must be a
reference into pending so that ++next advances the stored iterator.

Comments only; no code change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Refer to recursion_depth_limit() in the convert_structured() docs

The comment still named nesting_depth_limit, which #5637 removed on
develop in favor of detail::recursion_depth_limit().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Advance the pending iterator through pending.back() and shorten the null comment

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:36:04 +02:00
Niels Lohmann 3b6ae43c53 Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions (#5598)
* Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions

basic_json can be constructed from std::tuple<json&>, which it turns into
a one-element array. Because of this, std::tuple picks its converting
constructor that converts the whole source tuple instead of the
element-wise one. As a result, std::tuple<const json&> built from
std::forward_as_tuple(j) binds to a temporary (a compile error with libc++,
a dangling reference with other standard libraries), and std::tuple<json>
built the same way holds [j] instead of a copy of j.

The new opt-in macro JSON_DISABLE_TUPLE_REFERENCE_CONVERSION (CMake option
JSON_DisableTupleReferenceConversion) removes the conversion from a
one-element tuple holding a reference to the same basic_json type, so
std::tuple converts element-wise. It is off by default, so existing
behavior is unchanged. It does not change any function body and therefore
is not part of the ABI tag.

Fixes #2226

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert one-element tuples to arrays on every compiler

to_json for std::tuple assigns a braced list, j = { std::get<Idx>(t)... }.
With a single element that is itself a basic_json, Apple clang 15 and 16
treat j = {x} as a copy of x, so std::tuple<json>{true} became true
instead of [true]. The macOS jobs (Xcode 15.1, 16.1) failed the new
checks in unit-disable-tuple-reference-conversion and unit-regression2.

The one-element overload that already handles
JSON_BRACE_INIT_COPY_SEMANTICS builds the array (or object, for a
[string, value] element) explicitly, the same way the initializer-list
constructor does. Use it unconditionally. The output is unchanged on
compilers that already wrapped the element.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Skip json reference tuple tests on clang < 4 and GCC < 5

ci_test_compilers_gcc_old (4.8) and ci_test_compilers_clang (3.4) could
not compile the new tuple tests. Creating a std::tuple of basic_json
references, e.g. std::forward_as_tuple(j), makes these compilers
instantiate basic_json's conversion operator for libstdc++'s internal
tuple bases, which fails hard. This happens with and without
JSON_DISABLE_TUPLE_REFERENCE_CONVERSION, so it is a limitation of these
compilers, not of the new option.

Tested with the CI images: clang 3.4 to 3.9 and GCC 4.8 and 4.9 fail,
clang 4, 5, and 6 and GCC 5 and 6 compile all cases. Skip only the
checks that create such tuples; the is_constructible checks still run.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:53:56 +02:00
Niels Lohmann 6a073dbae4 Give operator>> a strong exception-safety guarantee (#5695)
operator>> parsed directly into its basic_json& target, so a parse
error left the target holding whatever was parsed before the error
instead of its previous value. With JSON_DIAGNOSTICS=1, that partial
value also violated the class invariant, because the parent pointers
of an array or object's elements are only set when the container is
closed, which a failed parse never reaches; copying such a value then
aborted in assert_invariant().

Fix it the way basic_json::parse() already handles this: parse into a
temporary and move it into the target only once parsing succeeds, so
the target is left unchanged if an exception is thrown.

Fixes #5652.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:34 +02:00
Niels Lohmann fdcc569eee Use only documented StringType members in json_pointer (#5692)
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.

Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.

Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).

Fixes #5666.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:31 +02:00
Niels Lohmann 5f1727cef2 Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	tests/src/unit-alt-string.cpp
2026-09-30 20:11:28 +02:00
Niels Lohmann 6fc0d501f3 Move the user-defined string literals to <nlohmann/json_literals.hpp> and add JSON_NO_AUTOMATIC_UDLS (#5610)
* Add JSON_NO_UDLS to leave out the user-defined string literals

The bodies of operator""_json and operator""_json_pointer call the
parser, so every translation unit including the library instantiates it,
even if it never parses anything. Defining JSON_NO_UDLS leaves the
literals out entirely, which saves 15-35% compile time for such
translation units (#5294). Nothing changes if the macro is not defined.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mention JSON_NO_UDLS in the list of exported module symbols

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the user-defined string literals to <nlohmann/json_literals.hpp>

Following the review in #5294, the literals now live in their own header
instead of being removed entirely: <nlohmann/json.hpp> includes it at the
end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined,
so a project can opt out globally and include the header only where the
literals are used.

The header only uses public and standard macros, because the library's
internal macros are undefined at the end of json.hpp and the amalgamation
inlines macro_scope.hpp only once. For the same reason, the library no
longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition
is still visible to the header. The single-header copy is identical to
the multi-header one, as it only includes <nlohmann/json.hpp>. The module
always exports the literals.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: include cycle, GCC 4.8 literal operator spacing, and global UDLs off in the JSON_NO_AUTOMATIC_UDLS test

- Suppress clang-tidy misc-header-include-cycle on the intentional mutual
  include of json.hpp and json_literals.hpp.
- Use operator"" _json with a space for GCC 4.8 in the test's detection
  aliases, as the header does.
- Only test the global literal operators when JSON_USE_GLOBAL_UDLS is on.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Declare the literal operators through a local macro

The GCC 4.8 spacing condition was repeated for both operator definitions
and the global using-declarations. NLOHMANN_JSON_LITERAL_OPERATOR(suffix)
now selects operator""##suffix or operator"" suffix in one place and is
undefined at the end of json_literals.hpp.

Suggested by gregmarr in review.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:08:17 +02:00
Niels Lohmann 67435c9c7e Give a deep copy its type only after its container exists (#5721)
When copying a value nested deeper than 128 levels, and an allocation
fails while an inner array or object is being copied, the partially
built copy ended up with an element typed array/object but holding a
null pointer. That element was already a fully constructed member of
its parent's container, so destroying the parent during stack
unwinding dereferenced the null pointer (release builds) or failed
assert_invariant() (debug builds), instead of letting std::bad_alloc
reach the caller.

copy_iteratively() set a pending worklist element's type right after
popping it, before the next loop iteration created its container in
copy_array_level()/copy_object_level(). Move that type assignment into
those two functions, right after the container is successfully
created, and drop the premature one in copy_iteratively(), so a
half-built element stays a null value - as copy_shallow()'s comment
already promised - until it can safely hold one.

Add a regression test to tests/src/unit-allocator.cpp that copies a
value nested 130 levels deep (both arrays and objects, with a
std::map- and an ordered_map-backed object_t) and fails every
allocation of the copy in turn: each attempt must throw std::bad_alloc
without crashing, and the source must stay unchanged.

Fixes #5640.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:08:00 +02:00
Niels Lohmann 2ea6d8c127 Require the found key to equal the looked-up key when comparing objects (#5720)
For an object type whose comparator treats unequal keys as equivalent
(for example a std::map with a case-insensitive comparator),
compare_iteratively() looked up a mismatched left key in the right
object with find(), which uses the object's own comparator, and
accepted whatever entry it found without checking that the keys are
actually equal. A case-insensitive comparator then found "KEY" for
"key", so two objects nested past the recursion bound (or at every
depth with JSON_NO_THREAD_LOCAL) could compare equal even though the
object type's own operator== - and basic_json itself, below the bound
- consider them different.

Accept the found entry only if its key equals (not just compares
equivalent to) the looked-up key.

Fixes #5655.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:56 +02:00
Niels Lohmann 66877675b1 Check the iterator range for binary values in basic_json(first, last) (#5719)
basic_json(first, last) treated value_t::binary like the structured
types (array, object) in the range check, so it always copied the
whole binary value regardless of the iterators, even for an empty
range such as (b.end(), b.end()). The other primitive types (number,
boolean, string) already reject such a range with
invalid_iterator.204, and erase(first, last) already does the same
for binary values, so this made the constructor inconsistent with
both. Move case value_t::binary into the group of checked primitive
types.

Also update the two matching passages in basic_json.md that describe
overload 7, and add a version-history note.

Fixes #5670.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:52 +02:00
Niels Lohmann 7fd6895788 Hide a discarded container's content from the parser callback (#5706)
When a parser callback rejects an object's or array's start event,
json_sax_dom_callback_parser kept calling it for everything inside
that container anyway: nested keys, values, and the start/end events
of containers below it. This contradicts parser_callback_t's own
documentation, which promises that discarding a container at its
start event also hides its content from the callback.

The same code path also kept a full copy of every key inside such a
discarded container in key_stack until the whole parse finished,
because the early return for values that are not stored skipped the
matching pop. Filtering out a large subtree is the main reason to use
a callback, so this made peak memory during the parse scale with the
size of the very subtree the callback was trying to skip.

Fix start_object(), start_array(), and key() so that a container
whose own start event was discarded, or that is nested inside one, is
never handed to the callback, and no longer pushes onto the key
stacks. A container whose start event was accepted but whose key was
rejected still gets its content reported, as documented ("the
callback is still called for the associated value, but its return
value has no further effect"); only its own bookkeeping is skipped
since it will not be stored.

Fixes #5643.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:48 +02:00
Niels Lohmann 6ae17630a4 Reject integral keys for contains(), find(), and count() at compile time (#5705)
j.contains(0), j.find(0), and j.count(0) used to compile: the literal 0 is
a null pointer constant, so it converts to a null const char*, and the
overloads taking const typename object_t::key_type& accepted it by
constructing a std::string from that null pointer, which is undefined
behavior (a crash with both libc++ and libstdc++). value(0, default_value)
had the same problem in C++11, where the object comparator is not
transparent.

Add deleted overloads for integral arguments to contains(), find()
(const and non-const), count(), and value() so that these calls are
compile errors in every supported language mode instead of crashing.
Calls with string, string_view, json_pointer, and size-typed element
access (at(), operator[](), erase()) are unaffected.

Fixes #5657.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:45 +02:00
Niels Lohmann 5bd766aa50 Move to_bson's binary subtype check into calc_bson_sizes (#5703)
to_bson() rejected a binary value's subtype above 255 (out_of_range.415)
in write_bson_binary(), which only has the binary_t, not the basic_json
value that holds it, so the exception was created with no JSON_DIAGNOSTICS
context even though the equivalent to_msgpack() check names the value's
path. The check also ran after the document size, all preceding elements,
and this element's header and length had already reached the output
adapter, so a caller-provided std::vector or std::string ended up holding
a truncated document.

calc_bson_sizes() already walks every value before anything is written,
to size embedded documents and arrays and to reject invalid keys
(out_of_range.409) up front. The subtype check now runs there instead,
in calc_bson_binary_size(), which is given the basic_json value so the
exception can use it as context. The now-redundant check in
write_bson_binary() is removed, since calc_bson_sizes() always throws
first if any binary value in the document has an oversized subtype.

Fixes #5675.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:41 +02:00
Niels Lohmann 4bb1b14b06 Trim the compiler-appended NUL from wide/UTF string literals too (#5702)
With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16,
UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]"))
failed with parse_error.101 at the terminating NUL of the literal,
and accept() returned false. The array overload of input_adapter()
only dropped the compiler-added trailing '\0' for arrays of char,
so for wchar_t, char16_t, char32_t, and char8_t arrays that
terminator was passed to the parser as data, which the macro then
rejected.

Broaden the trimming to every character type that a string literal
can use (char, wchar_t, char16_t, char32_t, and, since C++20,
char8_t). Arrays of any other element type (unsigned char,
std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a
trailing zero byte there is still read as genuine data.

Fixes #5658.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:37 +02:00
Niels LohmannandClaude Sonnet 5 bfe0f32d71 Fix std::terminate and null pointer access in input_stream_adapter (#5699)
Parsing from a std::istream crashed in two unusual but valid stream
states, both in input_stream_adapter:

- With eofbit in the stream's exceptions() mask, get_character() sets
  eofbit via is->clear(), which throws std::ios_base::failure. While
  that exception unwinds, ~input_stream_adapter() called clear() again
  to reset eofbit, which is still set and still in the exception mask,
  so it throws a second time out of the (implicitly noexcept)
  destructor and std::terminate() is called. The destructor now only
  calls clear() if a bit other than eofbit remains set, so the first
  exception can propagate normally.
- For an std::istream without a stream buffer (rdbuf() == nullptr,
  e.g. std::istream(nullptr)), the constructor stored the null
  pointer without checking it, and get_character() dereferenced it.
  input_adapter(std::istream&) now throws parse_error.101 for such a
  stream, the same as it already does for a null FILE* or char*.

Added regression tests to unit-deserialization.cpp and, for the
JSON_PRECISE_STREAM_POSITION variant of get_character(), to
unit-precise-stream-position.cpp; both crashed before this fix.
Documented the two exceptions in parse.md and operator_gtgt.md.

Fixes #5646.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-30 20:07:33 +02:00
Niels Lohmann 44a88d85be Fix NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's from_json message (#5698)
from_json built its out_of_range.410 message with "..." + j.dump(). If the
unmatched value is (or contains) a string with invalid UTF-8, that dump()
itself throws type_error.316, so the caller got type_error.316 instead of
the documented out_of_range.410; such strings can reach get<Enum>()
unvalidated, e.g. from from_cbor()/from_msgpack(). With a custom string_t,
j.dump() returns that type, and "const char*" + string_t does not compile
unless the type happens to provide operator+, so the macro failed to
compile for such types.

Build the message with detail::concat(), which appends any type exposing
data()/size() and always yields a std::string, and dump with
error_handler_t::replace so building the message itself cannot throw.

Added regression tests: an invalid-UTF-8 case in the existing strict-enum
test in unit-conversions.cpp, and a strict-enum use with alt_string (the
custom string_t from unit-alt-string.cpp) to cover the compile failure.

Fixes #5667.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:29 +02:00
Niels Lohmann 5b11a0282c Exchange the CustomBaseClass subobject in basic_json::swap() (#5697)
basic_json::swap() (and the friend swap() and the pre-C++20 std::swap
overload that forward to it) only exchanged m_data.m_type/m_data.m_value,
leaving each value's json_base_class_t subobject in place. This is
inconsistent with the copy and move constructors and copy assignment,
which all carry the base class along with the value, so after
a.swap(b) any metadata stored in a CustomBaseClass ended up attached to
the wrong value. Algorithms that mix swap() with moves, such as
std::sort, scrambled the metadata across the whole container.

Fix the member swap() to also exchange the json_base_class_t subobject
and extend the noexcept specifications of swap() and the friend swap()
accordingly.

Fixes #5653.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:24 +02:00
Niels Lohmann bfea6f36d3 Fix to_msgpack() reading the inactive number union member (#5694)
* Fix to_msgpack() reading the inactive number union member

basic_json stores number_integer and number_unsigned in a union, and
number_unsigned_t only has to be at least as wide as number_integer_t
(with the default types, both are 64-bit and have the same
representation). When number_integer_t is narrower, write_msgpack()
read the wrong union member in two places:

- The number_unsigned case wrote number_integer's bits instead of
  number_unsigned's, silently writing the wrong value whenever it
  did not fit in number_integer_t.
- The number_integer case (non-negative branch) picked the encoded
  width by comparing number_unsigned's bits, which is undefined
  behavior, though the value written was still number_integer's, so
  at worst a too-wide encoding was chosen.

Read the active member in both cases, like the other binary writers
(CBOR, UBJSON, BJData, BSON, BON8) already do.

Fixes #5644.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cast number_integer to number_unsigned_t only once in to_msgpack()

Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:19 +02:00
Niels Lohmann e444a66276 Copy values before inserting an initializer list into an array (#5693)
* Copy values before inserting an initializer list into an array

insert(pos, {...}) inserted wrong values when the initializer list
contained const references to elements of the array being inserted
into. json_ref stores only a pointer for a const lvalue, so the
initializer_list_t range passed straight to the array's range insert
aliased the array's own storage; std::vector::insert(pos, first, last)
may move or shift elements before copying from that range, so the
source elements were already stale by the time they were read
(different wrong results on libc++ and libstdc++).

Copy the referenced values into a temporary array_t first, then move
that temporary into place, so the source range never aliases the
array being modified.

Fixes #5656.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use the reserve_array helper in the initializer_list insert fix

The previous commit called array_t::reserve() directly on the
temporary buffer used to copy an ilist's values before inserting.
std::deque, a documented ArrayType (tests/src/unit-custom-array-type.cpp),
has no reserve(), so insert(pos, initializer_list) no longer compiled
for it. Use the existing detail::reserve_array() SFINAE helper (already
used by the SAX DOM parser) instead, which leaves array types without
reserve() untouched.

Added a regression check that deque_json::insert(pos, {...}) compiles
and handles the aliasing case from #5656.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:16 +02:00
Niels Lohmann 1151826508 Take diff()'s fast path unless the object type reorders members (#5691)
* Take diff()'s fast path unless the object type reorders members

For every object type except an insertion-ordered one like ordered_map,
diff() no longer produced a member-by-member patch when target had a key
that sorts before a key the two objects share: it fell through to the
slow path, which removes every member of source and re-adds every member
of target, instead of just adding the new key.

#5465 added an order check to require the fast path to also reproduce
target's member order, needed because ordered_map's patch()-driven "add"
appends a new member at the end. The check compared the common keys'
order between source and target and also required that every added key
come after every common key in target's order ("new_keys_form_suffix").
The comment above it argued this check is always true for std::map, and
that reasoning is correct for the order of the common keys themselves,
but not for new_keys_form_suffix: a std::map iterates in sorted key
order, so a new key that sorts before an existing common key is
enumerated between common keys, making new_keys_form_suffix false even
though std::map's own key order does not need reordering at all - it
places every member itself, regardless of insertion history, so a
member-by-member diff already reproduces target's iteration order.

Only require the order check for an object type that keeps insertion
order, using the same detail::is_ordered_map trait the library already
uses to recognize such an object type in set_parent(). Every other
object type - std::map in key order, a hash map in an order its
operator== ignores - always takes the fast path.

Added a regression test to unit-json_patch.cpp: the issue's example now
yields a single "add" op for json, while ordered_json still takes the
slow path to reproduce target's member order.

Fixes #5639.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Only track the target key order in diff() for insertion-ordered objects

common_keys_target_order and new_keys_form_suffix are only read when object_t keeps its members in insertion order; skip building them otherwise. Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:10 +02:00
Niels Lohmann 6218aa212b Throw std::length_error for operator[](SIZE_MAX) instead of corrupting the array (#5687)
For idx == SIZE_MAX, the non-const array operator[] computed the new size as
idx + 1, which wraps to 0. resize(0) then emptied the array, and the
subsequent operator[](idx) on the now-empty vector wrote one element before
its buffer. Every other too-large index (e.g. SIZE_MAX - 1) already went
through resize(), which throws std::length_error and leaves the array
unchanged; SIZE_MAX was the one value for which the overflow bypassed that
safety net.

Add a guard that throws std::length_error before computing idx + 1 when idx
is the largest representable size_type value, so the array is left
unchanged, matching the exception vector::resize() already throws for
smaller (but still too large) indices.

Fixes #5647.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:06 +02:00
Niels Lohmann 42f89e3130 Let to_json(std::optional<T>) propagate exceptions from T's to_json (#5684)
* Let to_json(std::optional<T>) propagate exceptions from T's to_json

The overload was marked noexcept even though its body assigns *opt to
the JSON value, which calls T's to_json (or allocates for std::string,
std::vector, or json). Any exception from there -- a user-defined
to_json reporting an error, or std::bad_alloc -- called std::terminate()
instead of propagating. The noexcept also made basic_json's converting
constructor noexcept(true) for std::optional<T>, so json j = opt; could
not report the error either.

Fixes #5642.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: make to_json(std::optional<T>) conditionally noexcept

GCC's -Wnoexcept (an error in ci_test_gcc) fired at to_json_fn's
noexcept(noexcept(to_json(j, val))): after dropping the unconditional
noexcept, to_json(std::optional<int>) had no exception specification
although GCC could prove its body cannot throw. It also made
json(std::optional<int>) lose its noexcept.

Declare the overload noexcept exactly when assigning the contained
value to the JSON value is (std::is_nothrow_assignable<BasicJsonType&,
const T&>), which is what the body does. std::optional<int> is noexcept
again; a T whose to_json may throw still propagates the exception.
Static assertions in the test check both cases.

The test's throwing to_json triggered -Wmissing-prototypes and
-Wmissing-noreturn (clang) and -Wmissing-declarations and
-Wsuggest-attribute=noreturn (GCC). Move the type and its to_json into
an anonymous namespace, mark the function [[noreturn]], and compile
them only without JSON_NOEXCEPTION, like the test that uses them.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:07:02 +02:00
Niels Lohmann 6d7845d207 Fix deprecated json_pointer/string operator== warning in value() (#5683)
value(KeyType&&, default) is constrained on is_comparable_with_object_key,
which passes KeyType as a reference. is_comparable's dispatch on
is_json_pointer_of<A, B> only matches a json_pointer as a plain type or a
plain reference, so a const-qualified reference (as produced when KeyType
is deduced from a json_pointer argument) fell through to
is_comparable_no_json_pointer, which instantiates the deprecated
json_pointer/string comparison operators. This made ordered_json's
transparent comparator (and any transparent comparator on a custom string
type) warn under -Wdeprecated-declarations when calling
value(json_pointer, default), even though no such comparison is ever
performed. at() was already fixed for this in #5289, which does not use
is_comparable_with_object_key.

Strip references and cv-qualifiers with uncvref_t before the
is_json_pointer_of dispatch, so any reference-to-json_pointer is
recognized regardless of qualifiers.

Fixes #5664.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:59 +02:00
Niels Lohmann 678fd3017b Fix element path for map/unordered_map JSON_DIAGNOSTICS errors (#5681)
When converting a JSON array to std::map or std::unordered_map with a
non-string key, each element must itself be a [key, value] array. If an
element is not an array, from_json() threw type_error 302 with the outer
array's value (&j) as the exception context, so with JSON_DIAGNOSTICS
enabled the message pointed at the whole array instead of the offending
element (e.g. "(/outer/m)" instead of "(/outer/m/2)"), even though the
message text already described the element's type.

Both from_json() overloads now pass the element (&p) as the context, so
the reported JSON Pointer matches the type named in the message, the
same way std::vector<std::vector<T>> and similar conversions already do.

Fixes #5668.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:55 +02:00
Niels Lohmann fc4c9c3446 Fix clear() to also reset the subtype of a binary value (#5680)
* Fix clear() to also reset the subtype of a binary value

clear() on a binary value cleared the bytes but left the subtype
untouched, so the result was not equal to a default-constructed
binary value even though the documentation says clear() has the same
effect as *this = basic_json(type()). The fix calls
byte_container_with_subtype::clear_subtype() alongside the existing
clear() call.

Extended the "filled binary" clear() test in unit-modifiers.cpp with
a case that uses a subtype, since the existing cases only covered
binary values without one.

Fixes #5669.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the table alignment in clear.md

Addresses review comment by @gregmarr.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:51 +02:00
Niels Lohmann 68beba727c Fix from_json() for enums with underlying type bool (#5679)
get_arithmetic_value() rejects boolean_t, so the default from_json()
for enums failed to compile for an enum whose underlying type is
bool (e.g. enum class Flag : bool { off, on }), even though the
matching to_json() serializes such enums as an unsigned number.

Read the underlying value through number_unsigned_t in that case,
matching what to_json() writes, then cast back to the underlying
type before constructing the enum.

Fixes #5671.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:46 +02:00
Niels Lohmann 0050d0f7a4 Share one nesting depth limit between all bounded descents (#5637)
Copying and comparing stopped their descent at
basic_json::nesting_depth_limit(), while serializing, hashing and
merging used detail::recursion_depth_limit(). Both were 128, but
nothing kept them equal. The thread-local count now tests against
detail::recursion_depth_limit() as well, and a static_assert keeps
the limit small enough for the byte that holds the count.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:43 +02:00
Niels Lohmann 6a8a7735ed Do not throw in contains() for an empty array reference token (#5614)
* Do not throw in contains() for an empty array reference token

json_pointer::contains() rejected malformed array indices, but an empty
reference token (e.g. "/a/" where "a" is an array, or "/" on an array)
passed every check and reached array_index(), which throws
out_of_range.404. contains() must not throw (cf. #5395), so it now
returns false for an empty token. at() still throws out_of_range.404.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: bind j_nested_const by reference in the json_pointer test

clang-tidy (ci_clang_tidy) flagged the new test with
performance-unnecessary-copy-initialization: the local copy
j_nested_const of j_nested is never modified. Bind it as a const
reference instead; it still exercises the const overloads of at() and
contains().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:39 +02:00
Niels Lohmann 65af260117 Move instead of deep-copy ordered_json values when an object grows (#5609)
* Move instead of deep-copy ordered_json values when an object grows

ordered_map keeps its elements in a std::vector<std::pair<const Key, T>>.
With a std::string key, that pair is not nothrow move constructible (the
const key has to be copied), so std::vector copies every element when it
reallocates. For ordered_json, this deep-copies every member value an
object already holds, including whole nested subtrees, on each growth
step.

Grow the storage in ordered_map instead, copying the keys and moving the
values. This happens in two phases, so the strong exception guarantee is
kept without try/catch. The first phase may throw, but only touches a
temporary buffer: it copies the keys, value-initializes the values, and
constructs the new element. The second phase moves the values (noexcept)
and swaps the buffers. Because the new element is constructed before any
value is moved, arguments that refer to elements of the container stay
valid, as with std::vector. Types that cannot take this path keep the
std::vector behavior.

Parsing into ordered_json (ParseStringOrdered, Apple M1 Max, clang -O3):
twitter 3.20 -> 1.70 ms, citm_catalog 7.73 -> 3.67 ms, jeopardy 219 ->
177 ms, canada unchanged. The number of allocations for twitter and
citm_catalog drops by two thirds.

Also add ParseStringOrdered rows to the benchmarks, and document the
growth behavior and the exception safety of ordered_map.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: skip std::pair noexcept assumptions on EDG-based compilers

ci_icpc and ci_nvhpc failed to compile unit-ordered_map.cpp: the static
assertion that std::pair<const std::string, ordered_json> is not nothrow
move-constructible fails there. The EDG front end (Intel icpc 2021.10,
NVIDIA nvc++ 25.5) considers the defaulted move constructor of
std::pair<const Key, T> noexcept even if copying Key can throw. With these
compilers, std::vector already moves such elements itself when it grows,
and ordered_map correctly leaves growing to it.

The same misjudgement makes std::vector call std::terminate when a key copy
throws during growth, so the exception-safety test with throwing_key would
abort on these compilers as well.

Skip the static assertion and the exception-safety section when __EDG__ is
defined. Verified with icpc 2021.10 (-std=gnu++11) and nvc++ 25.5 (C++11 and
C++17) on Compiler Explorer: unit-ordered_map and unit-disabled_exceptions
build and pass.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:34 +02:00
Niels Lohmann d268eaa693 Convert floats with Eisel-Lemire when std::from_chars is unavailable (#5617)
* Convert floats with Eisel-Lemire when std::from_chars is unavailable

Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit
coordinates of canada.json) went to strtod unless std::from_chars was
available. It is not used in C++11/14, and not with libc++, which does not
define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's
compute_float) now converts them with integer arithmetic, correctly rounded
for any token with at most 19 significant digits. Longer tokens are
truncated; the result is used if w and w + 1 round alike, else strtod
decides as before. Overflow still yields infinity (out_of_range.406).

The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit
test recomputes every entry with big-integer arithmetic. Further tests:
known values generated with Python (whose float() is correctly rounded),
200,000 round trips through to_chars, and the 128-bit multiplication and
leading-zero count against big-integer references (both with and without a
128-bit type). Checked against strtod on 6.5 million tokens, among them
60,000 exact halfway cases: no difference.

json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang);
other files unchanged. Compile time of a TU including json.hpp: +0.7%.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI: unused parse result and find() == npos in the Eisel-Lemire tests

GCC (-Werror=unused-result) rejected CHECK_THROWS_WITH_AS(json::parse(...))
because parse() is [[nodiscard]]; assign the result to a dummy json as the
other tests do. clang-tidy flagged longer.find('.') == npos with
abseil-string-find-str-contains; store the position in a variable first.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:26 +02:00
Niels Lohmann 5b9ffae8bb Move the float conversion chain out of the lexer (#5616)
lexer::convert_number() converted float tokens with std::from_chars (when
available), Clinger's fast path, and the locale-aware strtod fallback, all
as lexer members. They are now free functions in number_parse.hpp:

- convert_float_fast(): std::from_chars, then Clinger's fast path, skipped
  when the mantissa has too many significant digits
- convert_float_locale_aware(): strtof/strtod/strtold with the decimal point
  of the current locale, retried when the locale changed (#5198)

so that other code converting JSON number tokens gets the same values. No
change in behavior; the lexer no longer includes <clocale> and <cstdlib>.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:06:25 +02:00
Niels Lohmann 04f6ddd227 Keep external headers as #include when amalgamating (#5615)
A header that builds on json.hpp (such as the planned json_view.hpp) must
not inline json.hpp: its single-header version would contain a second copy
of the library, and that copy would change with every library change.

The optional config key "external" lists include paths that are kept as
#include directives. Only the first directive per path is kept; repeated
ones are commented out, as the tool already does for inlined headers.
config_json_view.json uses it for json_view.hpp; the existing configs do
not set it, and json.hpp and json_fwd.hpp regenerate byte-identically.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:05:51 +02:00
Niels Lohmann c16dd7e4f5 Fix CI: do not hide base class methods in the #3989 test SAX parsers
clang-tidy's bugprone-derived-method-shadowing-base-method rejected the
test SAX parsers that derive from json_sax_dom_parser (or the test's
SaxEventLogger) and redefine their non-virtual event functions. The
recovering DOM parsers of unit-class_parser.cpp and unit-regression2.cpp
now hold a json_sax_dom_parser and forward to it, SaxEventLogger gets a
flag to recover from errors instead of a derived class, and the one
function unit-alt-string.cpp redefines is marked as intended.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 04:50:11 +02:00
Niels Lohmann 792853d725 Fix CI: clang-tidy and JSON_NOEXCEPTION in the #3989 changes
ci_clang_tidy flagged nested conditional operators in the BON8 skip
code and in remove_incomplete_utf8_sequence()
(readability-avoid-nested-conditional-operator), and two branches with
the same body in recover_string() (bugprone-branch-clone). Use if
chains instead of the nested conditionals and merge the two branches
into one condition; the short-circuit order is unchanged.

ci_test_noexceptions aborted in the #3989 regression test: it compares
the first recovered error with the message of the exception that
from_cbor() and friends throw, and under JSON_NOEXCEPTION that call
aborts instead of throwing. Guard binary_error_message() and its uses
with #if !defined(JSON_NOEXCEPTION), as other tests in the suite do.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 16:42:57 +02:00
Niels Lohmann 4bdf1b7e74 Fix CI: no global std::string in the #3989 regression test
The U+FFFD constant was a global std::string, which the clang CI build
rejects with -Wglobal-constructors. It is now a function.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 09:27:46 +02:00
Niels Lohmann b1e9d98e41 Merge remote-tracking branch 'origin/claude/fix-issue-3989-db7e45' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:31:07 +02:00
Niels Lohmann f855d257df Fix CI: keep raw strings with backslashes out of test macros for MSVC
MSVC stringizes the arguments of doctest's CHECK() so that a raw string
literal becomes an ordinary one, and then reported the "\q" in the new
alt_string recovery test as warning C4129, an error with /WX. The input
is now a variable. The same pattern with "\u0000" in the parser's
recovery test is replaced by a JSON value built from a std::string.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:30:39 +02:00
Niels Lohmann 3e683e9c04 Merge branch 'develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:22:25 +02:00
Niels Lohmann 633de8e44b Fix CI: clang-tidy and GCC -Wnoexcept in the locale test (#5613)
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:

- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
  calls whose result was discarded. Check the result, like the other
  resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
  callbacks, which cannot throw but were not declared noexcept.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:20:43 +02:00
Niels Lohmann d1d84ed9af Fix Flawfinder: format the BSON element type without snprintf
Moving the report of an unsupported BSON element type into
skip_unsupported_bson_element() moved its snprintf() call, which
Flawfinder then reported as a new CWE-134 finding. The two hexadecimal
digits are now computed directly; the message is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 20:36:28 +02:00
Niels Lohmann fc03b9912e Look up the locale decimal point at conversion time, not lexer construction (#5597)
* Look up the locale decimal point at conversion time, not lexer construction

The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.

token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.

Fixes #5198

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the strtod retry loop when the decimal point is unchanged

convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.

Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Weffc++ errors in the #5198 locale test

GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:56:11 +02:00
Niels Lohmann de8529f99b Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/input/binary_reader.hpp
#	single_include/nlohmann/json.hpp
2026-09-28 17:53:51 +02:00
Niels Lohmann 9e1a09eec0 Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys

CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:

  syntax error while parsing MessagePack object key: only string keys
  are supported, but found nil; last byte: 0xC0

The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.

Refs #2766, #3381

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Point the MessagePack key note to the spec's profile section

The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:51:07 +02:00
Niels Lohmann 677794f076 Fix CI: recover numbers with the string operations every string_t has
recover_number() used back(), pop_back(), front(), and
find_first_of(const char*), which the minimal alt_string of
unit-alt-string.cpp does not provide. Since the binary readers recover
UBJSON/BJData high-precision numbers with it, from_ubjson() instantiated
it, too, and the test no longer compiled. It now uses only size(),
operator[], and resize(), and unit-alt-string.cpp recovers from errors
in JSON text, UBJSON, and CBOR.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 12:17:11 +02:00
Niels Lohmann 437a95cfdb Repair complete items in binary formats when parse_error() returns true (#3989)
When the SAX parser asks to recover, the binary readers now repair an
item whose end is known and read on after it, as RFC 8949, Section 5.3
describes for CBOR:

- CBOR: tags are ignored, and simple values other than false, true, and
  null become null (RFC 8949, Section 6.1); a negative integer below the
  range of number_integer_t becomes the nearest floating-point number.
- Strings that are not valid UTF-8 get U+FFFD for each ill-formed
  sequence, as in JSON text; so does a UBJSON/BJData char above 0x7F.
- UBJSON/BJData high-precision numbers keep their longest valid
  beginning (via the lexer's recover_token()), or become infinity.
- Members whose key is not a string are skipped (CBOR, MessagePack,
  BON8), like members without a key in JSON text.
- BSON elements of types the library does not read (ObjectId, datetime,
  decimal128, ...) become null; a string without its terminator and a
  document whose size does not match are kept.

Where the end of an item is unknown, reading stops as before, except
that BSON skips to the end of the document, whose size it knows.

The value read before such an error is now completed by the reader from
its container stack, as the JSON parser does, instead of by a proxy SAX
parser, which is removed. Like the parser, binary_reader gets an
AllowRecovery template parameter, so that from_*() compile without the
new code.

Tests: a table of repairs, numbers out of range, errors that stop, and
a sweep over changed and removed bytes of eight encodings that checks
balanced events and that the first error is the one from_*() reports.
All fuzzers now run a recovering checker; the binary ones also check
that it reports an error exactly when from_*() fails.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 21:40:59 +02:00
Niels Lohmann c8735246d0 Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 21:02:38 +02:00
Niels Lohmann 373005f7ac Fix MSVC: avoid reserving by faked size in MessagePack size tests (#5604)
The "Size above uint32" tests for arrays and objects fake a container
size of 2^32 and expect to_msgpack() to throw out_of_range.412. But
to_msgpack(j) first reserves binary_reserve_hint(j) bytes, which is
size + 1 for arrays and 2 * size + 1 for objects, i.e. 4 or 8 GiB.
Linux and macOS overcommit, so the reservation succeeds; on Windows it
throws std::bad_alloc before the size check is reached (seen with
msvc-vs2026 Debug x64 on the object test).

Write into a caller-owned vector instead, so nothing is reserved.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 20:57:16 +02:00