Compare commits

...
Author SHA1 Message Date
Niels Lohmann b87a618151 Pin Google Benchmark to release 1.9.5
The benchmarks fetched Google Benchmark's main branch, so two builds on
different days could measure with different library code, and CMake 3.30
and later warn that the single-argument FetchContent_Populate() is
deprecated. Fetch the 1.9.5 release archive, verified by its SHA-256,
with FetchContent_MakeAvailable() instead. That needs CMake 3.14; Google
Benchmark itself already needed 3.13.

Its -Werror is switched off, so a newer compiler's new warnings cannot
break the pinned release, and its install rules are no longer added.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-24 09:48:18 +02:00
Niels Lohmann 67ae3fd85b Point ci_benchmarks at tests/benchmarks
The target has configured ${PROJECT_SOURCE_DIR}/benchmarks since it was
added in #2561, but the benchmarks live in tests/benchmarks.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-24 07:10:37 +02:00
Niels Lohmann 71fe9f0b96 Document the benchmarks, and make them build and compare versions again
The benchmark project hasn't configured since #4793: download_test_data.cmake
compiles cmake/detect_libcpp_version.cpp relative to CMAKE_SOURCE_DIR, which
is tests/benchmarks when that is the top-level project, so try_run fails and
so does `make run_benchmarks`. The path is now relative to the module itself,
which is the same file for the main build.

The Dump benchmark discarded dump()'s result, which is [[nodiscard]] by now;
it warned, and left the optimizer free to shorten the loop. The result is
now kept with benchmark::DoNotOptimize.

A new cache variable, JSON_BENCHMARK_INCLUDE_DIR, names the directory holding
the nlohmann/json.hpp to benchmark (single_include by default, as before),
so the same benchmarks can be built against two versions and compared.

tests/benchmarks/README.md documents what is measured, how to build and run
the benchmarks, how to read the output, and how to compare two versions with
Google Benchmark's compare.py; it recommends doing so by hand before a
release rather than in CI.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-24 07:07:59 +02:00
Niels LohmannandClaude Sonnet 5 ed513715a8 Document that a NUL byte in the input is treated as end of input (#5534)
* docs: document that a NUL byte in the input is treated as end of input

A NUL byte anywhere in the input - trailing, or embedded ahead of more
otherwise well-formed JSON - is currently treated the same as genuine
end of input, so parsing silently stops there instead of raising the
parse_error.101 any other unexpected byte triggers. This mirrors the
NUL-terminated-C-string convention already used when no explicit input
length is given (json::parse(const char*) already stops at strlen()),
just applied uniformly rather than only when a length is genuinely
unavailable.

This behavior predates this change and is not being altered here -
changing it would be an observable, backwards-incompatible behavior
change for any caller that (knowingly or not) depends on it, which is
not something to do silently in a patch. Documenting the current,
verified behavior as a new FAQ entry instead, so it's an intentional
and discoverable part of the contract rather than a surprise.

Fixes #5530.

Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY

* Add JSON_STRICT_NUL_HANDLING opt-in macro for issue #5530

A NUL byte anywhere in the input is currently treated the same as real
end of input, rather than raising parse_error.101 like any other
unexpected byte (documented in the previous commit's FAQ entry). A full
unconditional fix was tried in PR #5532 but rejected as too risky to
ship by default: any caller could depend on the current behavior, even
unknowingly (e.g. a zero-padded buffer). On PR #5534, gregmarr proposed
a compile-time opt-in flag instead, and the maintainer agreed, wanting
it available now and defaulting to the corrected behavior in 4.0.0.

This mirrors the existing JSON_BRACE_INIT_COPY_SEMANTICS precedent as
closely as sensible:
- JSON_STRICT_NUL_HANDLING defaults to 0 (off); the three lexer sites
  that treat '\0' as EOF/comment-terminator are gated with
  `#if !JSON_STRICT_NUL_HANDLING` so the default-off behavior is
  byte-for-byte identical to today's.
- input_adapters.hpp's `T (&array)[N]` overload additionally trims a
  single trailing '\0' from a `char` array (e.g. a string literal like
  `json::parse("123")`) when the macro is on, so that case keeps
  working; every other element type (unsigned char, std::uint8_t, ...)
  always keeps its full extent. This intentionally does *not* reuse the
  existing strlen()-based pointer overload via SFINAE-excluding `char`
  from the array overload, as originally sketched for this change: that
  approach is ambiguous against the newer generic container overload
  added since PR #5532, and even where it compiles, strlen()-scanning a
  `char` array that is not NUL-terminated within its bounds reads past
  the end of the array (confirmed with AddressSanitizer). Trimming only
  a single trailing byte, without scanning, avoids both problems.
- Documented via docs/mkdocs/docs/api/macros/json_strict_nul_handling.md,
  linked from the macros index/nav/features page, the FAQ entry, and
  the parse/accept/operator>> reference pages.
- Tested in unit-class_parser.cpp and unit-deserialization.cpp, default
  state unguarded and opt-in state guarded. Since the library itself
  #undefs the macro at the end of json.hpp (as JSON_BRACE_INIT_COPY_SEMANTICS
  already does), a plain `#if defined(JSON_STRICT_NUL_HANDLING)` guard
  after the include never actually triggers; the tests instead capture
  the command-line value into a test-local macro before including the
  header. A few pre-existing fixtures elsewhere (std::array<uint8_t, N>
  sized one larger than their literal, relying on value-initialization
  to silently add a trailing zero byte) needed the same one-byte
  adjustment to keep passing under the opt-in behavior.

Unlike the precedent, this adds a proper `JSON_StrictNulHandling` CMake
option (rather than a raw -DCMAKE_CXX_FLAGS injection) and wires its
ci_test_strict_nul_handling target into the ci_cmake_options job matrix
in .github/workflows/ubuntu.yml, so the opt-in build is actually
exercised in CI -- closing the one gap in the precedent's own CI setup
(ci_test_brace_init_copy_semantics is defined but never referenced by
any workflow, so it has never actually run).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Clarify where JSON_STRICT_NUL_HANDLING does not reject NUL bytes

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <niels.lohmann@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 06:55:21 +02:00
Niels Lohmann f751547a81 Add missing contributors to the README thanks list (#5543)
Add 60 contributors whose work was not yet credited and update seven
links that pointed to renamed or reassigned GitHub accounts.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-23 20:43:40 +02:00
dependabot[bot] 06b0452189 Bump mkdocs-git-revision-date-localized-plugin in /docs/mkdocs (#5537)
Bumps [mkdocs-git-revision-date-localized-plugin](https://github.com/timvink/mkdocs-git-revision-date-localized-plugin) from 1.5.4 to 1.6.0.
- [Release notes](https://github.com/timvink/mkdocs-git-revision-date-localized-plugin/releases)
- [Commits](https://github.com/timvink/mkdocs-git-revision-date-localized-plugin/compare/v1.5.4...v1.6.0)

---
updated-dependencies:
- dependency-name: mkdocs-git-revision-date-localized-plugin
  dependency-version: 1.6.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-23 20:32:31 +02:00
Niels LohmannandClaude 1054b2097e Speed up binary writing: value-type output sink + byte-swap number encoding (#5286)
* Devirtualize binary_writer via a value-type output sink

to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson wrote every byte through
output_adapter_t, a shared_ptr<output_adapter_protocol> whose
write_character/write_characters are virtual. Unlike the lexer (templated
on a concrete InputAdapterType), the binary writer never got that
treatment, so binary output paid a vtable lookup per byte and a
make_shared per call.

Template binary_writer on an OutputSinkType and give it two concrete,
non-virtual sinks:

- output_vector_sink: appends straight into a std::vector (push_back /
  insert), used by the vector-returning to_* convenience functions. No
  vtable, no shared_ptr; the writes inline.
- output_adapter_sink: forwards to a type-erased output_adapter_t, so the
  existing to_*(j, output_adapter) overloads (streams, strings, custom
  adapters) keep working exactly as before -- one virtual call each,
  unchanged.

binary_writer keeps a convenience constructor taking output_adapter_t
(building the default output_adapter_sink), so the adapter overloads are
untouched; only the convenience functions switch to the vector sink. The
friend declaration and the basic_json binary_writer alias gain the new
(defaulted) template parameter.

Output is byte-for-byte identical: verified across ~3000 randomized
values plus curated edge cases (all scalar widths, strings with invalid
UTF-8, binary, nested arrays/objects) for CBOR, MessagePack, UBJSON (both
size/type settings), BJData, and BSON, plus the output_adapter path, in
C++11/17/20. Warning-clean under clang -Weverything and the gcc pedantic
set; clang-tidy clean on the changed headers; make check-amalgamation
clean.

Throughput (g++ -O3, vs develop): scalar-dense binary output such as
integer arrays ~1.4x; many small to_cbor calls ~1.04x (DOM traversal
bound); string/blob-heavy output unchanged (already bulk-bound). No
workload regressed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix CI failures from binary_writer output-sink change

Four CI jobs failed on the initial commit; all are addressed here without
changing any output (binary encodings remain byte-for-byte identical to
develop across the differential corpus):

1. ci_test_gcc / cuda (-Werror=duplicated-branches): for number_float_t ==
   float, static_cast<float>(n) is the identity, so write_compact_float's
   two branches are intentionally identical. Once the concrete vector sink
   is inlined, GCC constant-folds and diagnoses this (the type-erased path
   hid it behind a non-inlined virtual call). Silence -Wduplicated-branches
   for GCC (clang has no such warning) alongside the existing -Wfloat-equal
   pragma.

2. ci_static_analysis_clang (UBSan nonnull-attribute): binary_writer passes
   a null pointer with length 0 for empty strings/binary. output_vector_sink
   / output_adapter_sink declared write_characters JSON_HEDLEY_NON_NULL, so
   the sanitizer flagged the (harmless) zero-length call once the sink was
   called directly rather than through the attribute-free virtual base. Drop
   the attribute from both sinks, matching the pre-existing behavior.

3. ci_cpplint (build/include_what_you_use): output_adapter_sink uses
   std::move; add #include <utility>.

4. ci_cuda_example (nvcc 11.8): NVCC's front end rejects the default
   template argument on the binary_writer alias template. Revert the alias
   to its original single-parameter form (relying on binary_writer's own
   defaulted OutputSinkType) and spell out the full type in the vector-sink
   convenience functions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Encode big-endian numbers with a byte swap instead of std::reverse

write_number() reordered multi-byte numbers for the big-endian formats
(CBOR/MessagePack/UBJSON) with std::reverse over the byte array. GCC
lowered only some sizes to a bswap; clang kept a scalar byte shuffle
(0 bswap instructions in the CBOR number path). Replace the reverse with
size-dispatched __builtin_bswap16/32/64 helpers (portable shift fallback
for other compilers; std::reverse retained for exotic sizes such as a
long double number_float_t).

Codegen: the CBOR number path now emits bswap on both compilers
(gcc 2 -> 16, clang 0 -> 4). Output is byte-for-byte identical to the
previous implementation across the binary differential corpus.

Throughput (isolated vs the std::reverse version, best of 9):
  CBOR int64 array   gcc +7%   clang +10%
  CBOR uint16 array  gcc +27%  clang flat

Modest but consistent on number-dense encodings; negligible on
string/blob-heavy output, as expected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Reserve output capacity up front for binary serialization

The vector-returning to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson grew
the output buffer purely by geometric reallocation. Reserving an estimate
up front avoids the early reallocations, which is the dominant per-byte
cost for array/object-heavy output.

The estimate (binary_reserve_hint) is deliberately conservative and safe
against untrusted input: it consults only the top-level element count
(O(1), no walk of the DOM), guards the multiplication against overflow,
and clamps the result to a fixed 1 MiB ceiling, so a large or hostile DOM
can never force an oversized allocation here. The buffer still grows
geometrically past the hint, so an underestimate only costs a few later
reallocations; scalars/strings/binary are written in one shot and get no
hint. Reserving capacity does not change the bytes produced.

Throughput (g++/clang -O3, vs the previous commit):
  cbor int array     +10% / +13%
  cbor object array  +20% / +38%

Output is byte-for-byte identical to develop across the binary
differential corpus.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review findings on the binary writer output sinks

- binary_reserve_hint(): the 4-bytes-per-element estimate over-reserved by up
  to 4x for arrays of small scalars (CBOR encodes 0..23 in one byte), and the
  returned vector kept that capacity. Make the hint a strict lower bound on the
  encoded size instead, which also removes the 1 MiB clamp whose branch no test
  could reach (the largest container in the suite has 65793 elements).

- Guard the -Wduplicated-branches pragma with __GNUC__ >= 7. The warning does
  not exist before GCC 7, so naming it made GCC 4.8/4.9/5/6 - which the CI
  matrix still builds - warn under -Wpragmas on every including translation
  unit, breaking downstream -Werror builds.

- Constrain the adapter constructor of binary_writer with the enable_if its
  documentation already claimed, so a writer over some other sink type is no
  longer advertised as constructible from an output adapter.

- Let output_vector_adapter wrap output_vector_sink rather than duplicating the
  append logic, so the type-erased and templated paths share one implementation.

- Collapse the three copies of the memcpy/byte_swap/memcpy dance into a single
  byte_swap_buffer() helper, and add the MSVC _byteswap_* intrinsics so MSVC no
  longer falls back to the scalar shuffle this change exists to eliminate.

- Add a vector_writer() helper for the five vector-returning to_* overloads
  instead of spelling out the writer type at each call site, and drop a dead
  default member initializer on output_adapter_sink.

- New tests: the vector sink and the adapter sink must produce identical bytes
  for every format (the two to_* overloads no longer delegate to each other and
  could otherwise drift), and binary_reserve_hint() must never exceed the size
  actually written.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Route the -Wduplicated-branches pragma through Hedley

Match #5485, which moved the binary writer's hand-rolled diagnostic
pragmas onto JSON_HEDLEY_PRAGMA (merged into develop while this branch
was open). The devirtualization's -Wduplicated-branches suppression in
write_compact_float was the one raw '#pragma GCC diagnostic' left; it
now uses JSON_HEDLEY_PRAGMA like the adjacent -Wfloat-equal line, still
guarded to GCC >= 7 and non-clang (the warning exists only there).

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-23 08:59:46 +02:00
Niels Lohmann f92024b317 De-duplicate the swap() diagnostic-positions characterization test (#5540) 2026-09-23 07:48:21 +02:00
Niels Lohmann 0b20b7e622 Reject MessagePack/BSON binary subtypes that don't fit their wire format (#5469)
* Reject MessagePack/BSON binary subtypes that don't fit their wire format

Both formats store byte_container_with_subtype's subtype (a uint64_t)
in a single byte. The writers cast to std::int8_t/std::uint8_t without
a range check, so subtypes above 255 were silently truncated modulo
256 instead of raising an error. Throw out_of_range.413 instead when
the subtype exceeds the representable range of 0-255.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the new binary-subtype regression test out of unit-regression2.cpp

unit-regression2.cpp is already at the edge of what the MinGW linker
can relocate; adding this test's ~26 lines tips test-regression2_cpp20
(clang, Windows) over into "relocation truncated to fit:
IMAGE_REL_AMD64_REL32 against `.rdata'" (see 8ce64b9c1 / b82717c8a for
the same failure mode). Split the test along format lines instead:
MessagePack assertions move to unit-msgpack.cpp, BSON assertions to
unit-bson.cpp. The CBOR round-trip guard is dropped as redundant --
unit-cbor.cpp's "Tagged values" section already round-trips subtypes
up to 8589934590, far past the 70000 checked here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-22 21:52:32 +02:00
Niels Lohmann d2c1a6a272 Reject array insert(pos, first, last) iterators not pointing into an array (#5468)
The array-range insert() overload checked that pos fits the current
value and that first/last share the same owning value, but never
verified that value is itself an array. Passing iterators from an
object, a primitive, or null handed value-initialized (singular)
std::vector iterators straight to array_t::insert(), which is
undefined behavior. Add the missing is_array() check, mirroring the
equivalent check already present in the object-range insert()
overload.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-22 21:52:31 +02:00
Niels Lohmann a2b19d6158 Honor allow_exceptions=false for excessive array/object size (out_of_range.408) (#5467)
* Honor allow_exceptions=false for excessive array/object size (out_of_range.408)

The SAX DOM parsers' start_object()/start_array() threw out_of_range.408
directly via JSON_THROW when a binary format (CBOR/UBJSON/BJData) declared
a container size exceeding max_size(), bypassing the allow_exceptions flag
that every other malformed-input error path in these classes honors via
parse_error(). This meant that json::from_cbor(data, true, false) etc.
could still throw (or abort under JSON_NOEXCEPTION) instead of returning a
discarded value, contrary to the allow_exceptions=false contract.

Route all four call sites (two in json_sax_dom_parser, two in
json_sax_dom_callback_parser) through parse_error() instead, matching the
existing error-handling pattern used elsewhere in this file. Behavior is
unchanged when allow_exceptions is true (the default); the exception
message and type are identical.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop a non-portable exact exception message check in the 408 test

The allow_exceptions=false regression test checked the exact message
text produced when allow_exceptions=true (the default). On platforms
where std::size_t is 32-bit (e.g. mingw x86, MSVC Win32 builds), a
declared CBOR length of 2^63 is intercepted earlier, by
get_cbor_container_size()'s own (pre-existing, already correct)
length-narrowing check, with different wording than this fix's
start_array()/start_object() size check -- same error code, same
"still throws when allow_exceptions=true" guarantee, different text.

CHECK_THROWS_AS already verifies the behavior this test cares about
(still throws json::out_of_range, unchanged); drop the exact-message
assertion since it isn't portable across size_t widths and doesn't
add coverage of this fix specifically.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Werror=unused-result on json::from_cbor() in the 408 regression test

from_cbor() is [[nodiscard]]; CHECK_THROWS_AS() otherwise discards its
result, which GCC flags under -Werror. Assign to a throwaway json, as
the rest of the suite already does for from_cbor()/from_msgpack().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-22 21:52:31 +02:00
Niels Lohmann e485441123 Restore a duplicate key's prior value when the callback rejects its new value (#5466)
* Restore a duplicate key's prior value when the callback rejects its new value

json_sax_dom_callback_parser::key() unconditionally overwrote the object
slot for a key with a `discarded` placeholder as soon as the key was
accepted by the parser callback. For a duplicate key (legal JSON), this
destroyed the pre-existing value from an earlier occurrence of the same
key before the new value was even parsed. If the new value was then
rejected by the callback, remove_discarded_value() erased the member
entirely instead of leaving the original value in place, contradicting
the documented behavior that a discarded value behaves as if it was
never read.

Add a small stash of (slot pointer, previous value) pairs so that when
key() overwrites an existing member with the discarded placeholder, the
previous value can be restored later if the corresponding value (scalar,
object, or array) is rejected, instead of being erased. The stash entry
is dropped without restoring once the new value is definitively
accepted (in handle_value() for scalars, end_object()/end_array() for
containers), so a duplicate key whose new value is accepted still keeps
the last value as before. Non-duplicate keys are unaffected: rejecting
their value still removes the member entirely, since there is nothing
to restore.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark parser-callback test lambdas noexcept to fix GCC -Wnoexcept -Werror

GCC's libstdc++ std::function move assignment evaluates a noexcept
check that invokes a wrapped callable in an unevaluated context; a
non-noexcept parser_callback_t lambda then trips -Wnoexcept ("noexcept-
expression evaluates to 'false'"), which CI's ci_test_gcc job builds
with -Werror. The pre-existing parser_callback_t test lambdas in this
file already work around this by declaring themselves noexcept; apply
the same fix to the three added lambdas that didn't.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-22 21:52:30 +02:00
Niels Lohmann e564136c22 Make diff() account for member order in ordered_json objects (#5465)
* Make diff() account for member order in ordered_json objects

diff() compared source/target objects purely by key set, ignoring
relative member order. For ordered_json (insertion-ordered, vector-
backed object_t), two objects that differ only in member order are
unequal via operator==, but diff() never emitted any patch operation
to fix the order, so source.patch(diff(source, target)) == target
could fail to hold.

Fix by detecting when common keys appear in a different relative
order in source vs. target (or when a new key would need to land
somewhere other than the end), and in that case removing and
re-adding the affected keys in target's order, which relies on
patch()'s "add" op appending new keys at the end of an ordered_map.
For plain json (std::map-backed, always key-sorted iteration) this
is a no-op and the original minimal per-key diff path is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid redundant lookups in diff()'s object-order tracking

The previous fix for ordered_json member order re-derived common-key
order and suffix information with extra target.find()/source.find()
calls layered on top of the pre-existing removed/added-key passes,
instead of reusing those same passes. This roughly tripled the number
of map lookups per diff() call for every object, including plain
`json`, where the reordering path is never taken.

Piggyback the order tracking (and the "add" op construction for new
keys) onto the two passes the algorithm already needs to detect
removed/added keys, and walk the fast path's recursion in lockstep
with the precomputed common-key list instead of re-querying `target`.
This restores diff() to its pre-existing lookup count; benchmarked at
n=1000 keys, ordered_json::diff() was roughly 2x slower than baseline
before this change and is back within noise of baseline after it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Preserve diff()'s original op ordering and fix a slow-path deletion gap

Splitting removed-key detection and common-key recursion into separate
passes (for the earlier lookup-count fix) changed the emitted patch's
op order: all "remove" ops now came before all recursive per-key diffs,
instead of interleaved in source's iteration order as the original
implementation did. This broke docs/mkdocs/docs/examples/diff.output's
exact-match CI check (ci_test_examples) even though the patch was still
semantically correct.

Defer "remove" emission into the same walk that does the recursive
diffs, so common keys and deleted keys are interleaved in source order
again, matching historical output.

While restructuring that walk, the reordering ("slow path") branch was
only emitting "remove" for keys common to both objects, never for keys
present in source but genuinely absent from target -- a key deleted
alongside an actual reorder would silently survive the patch. Fixed by
removing every source key in the slow path (both deleted and common
keys need removing there; common keys are then re-added in target's
order). Verified with a targeted reorder+deletion case and a fresh
20,000-case round-trip fuzz run (0 failures).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-22 21:52:30 +02:00
Niels LohmannandClaude Sonnet 5 663013ce64 Fix incorrect diagnostic-positions test assertion for swap() (#5539)
The characterization test added in #5482 asserted that basic_json::swap()
does NOT exchange start_position/end_position, based on a misreading of
the code cited for #5420. In fact swap() (json.hpp, around line 3637)
does swap start_position/end_position along with the value, consistent
with copy-assignment. The test's assumption was backwards, so it failed
on every CI job across every branch/PR since the commit landed. Correct
the assertions to match the actual (and correct) behavior: positions are
exchanged together with values.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-22 21:35:00 +02:00
Niels Lohmann c41152e620 Fall back to plain-object encoding when to_bjdata()'s _ArrayType_ annotation is not a string (#5494)
* Fall back to plain-object encoding when _ArrayType_ is not a string

write_bjdata_ndarray() looked up _ArrayType_ by calling get<string_t>()
directly, which throws type_error.302 when the annotation is not a
string (e.g. a number, null, boolean, array, or object). Per the
documented BJData ndarray contract, an object only qualifies for the
compact ndarray encoding if _ArrayType_ names a known type; anything
else must fall back to plain-object encoding, the same way an unknown
type-name string already does.

Add an is_string() check before the get<string_t>() call so a
non-string _ArrayType_ takes the existing "unrecognized type name"
fallback path instead of throwing.

Fixes #5398.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Relax the BJData fuzzer's round-trip check from byte-exact to value-exact

Fixing #5398 lets to_bjdata() proceed past the object it used to reject,
which exposed a pre-existing, unrelated round-trip quirk to the fuzzer:
a binary_t value serialized through the non-optimized ("$U#"-less)
array encoding is parsed back as a plain array of numbers, since
from_bjdata() has no way to tell "array of uint8 numbers" apart from
"array of bytes" without that optimized header. Re-serializing that
plain array then goes through the generic smallest-type writer, which
- unrelated to this PR, and long predating it - prefers the 'i' (int8)
marker over 'U' (uint8) for values that fit both, so the re-encoded
bytes can differ from the original even though both decode to the same
value.

This is not introduced by the #5398 fix; the same divergence reproduces
from a bare json::binary_t value with no _ArrayType_ annotation
involved at all, on the commit immediately preceding it. A general fix
would mean changing the shared UBJSON/BJData smallest-type selection
that hundreds of existing tests pin to 'i' for small positive
integers, which is out of scope and too risky for this PR.

Update fuzzer-parse_bjdata.cpp's round-trip assertions to check that
re-serializing is value-stable (from_bjdata(to_bjdata(j)) == j) rather
than byte-exact, matching the guarantee BJData actually provides, and
add a regression test in unit-bjdata.cpp using the exact OSS-Fuzz input
that documents the behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Compare dump()s instead of json values in the BJData fuzzer's round-trip check

The value-stability assertion added to fix the earlier OSS-Fuzz crash
(json::from_bjdata(to_bjdata(j2)) == j2) itself broke on a NaN payload:
IEEE 754 NaN is never equal to itself, so operator== reports two
structurally-identical trees containing a non-finite double as
different -- not a round-trip bug, just NaN's ordinary
non-reflexivity. dump() serializes any non-finite double the same
deterministic way (as JSON null, since JSON cannot represent NaN or
Infinity), so comparing dumps is stable under exactly the values that
break operator==.

Verified against both the original OSS-Fuzz crash input and the new
one (0x68 0x68 0x7c, which decodes to a NaN), plus a local 2.5M-case
random-input sweep with no failures.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:20:22 +02:00
Niels Lohmann 502e9d66f6 Fix to_bjdata() emitting the Draft-3-only 'B' marker in default Draft-2 mode (#5479)
* Fix to_bjdata() emitting the Draft-3-only 'B' marker in default Draft-2 mode

_ArrayType_ = "byte" mapped unconditionally to the BJData type marker
'B', regardless of the requested bjdata_version. 'B' is defined only by
BJData Draft 3; with the default version (draft2), this produced a
stream that is invalid for Draft 2 and, unlike every other
_ArrayType_, round-tripped back as a binary value instead of the
original annotated object.

Only accept "byte" / emit 'B' when bjdata_version selects Draft 3.
Under Draft 2, fall back to the same plain-object encoding used
elsewhere in this function for other invalid-annotation cases, so the
value round-trips correctly.

Fixes #5404.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Future-proof the Draft-3-only 'B' marker gate

@gregmarr pointed out that dtype == 'B' && bjdata_version != draft3
only future-proofs by accident, since bjdata_version_t currently has
exactly two values. Compare with < instead, so a later draft that
keeps the 'B' marker valid does not need this gate revisited.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:20:21 +02:00
Niels Lohmann e8e1ba0db9 Fix dead ill-formed-fourth-byte UTF-8 test sections (byte3/byte4 typo) (#5499)
* Fix dead ill-formed-fourth-byte UTF-8 test sections (byte3/byte4 typo)

The "ill-formed: wrong fourth byte" SECTIONs in unit-unicode3.cpp,
unit-unicode4.cpp, and unit-unicode5.cpp guarded their loop with a check
on byte3 instead of byte4. Since the enclosing loop already restricts
byte3 to its valid range, the guard was always true and the section's
"continue" fired unconditionally, so check_utf8string()/check_utf8dump()
were never actually invoked for a malformed fourth byte.

Fixing the guard naively (byte3 -> byte4) would also have swept the full
byte2 x byte3 combinatorics for every byte4 value, adding millions of
redundant iterations: the lexer validates continuation bytes strictly in
sequence with early exit (see next_byte_in_range() in lexer.hpp), so once
byte2/byte3 are within their valid range, the byte4 outcome does not
depend on which valid byte2/byte3 values were chosen. Instead, byte2 and
byte3 are now held to a small hedge of representative valid prefixes
(range corners plus a midpoint) while byte4 is still swept exhaustively
over its full 0x00-0xFF range, since that is the actual property under
test. Also fixed the garbled "skip fourth second byte" comment in
unit-unicode3.cpp.

Verified offline: before the fix, the "wrong fourth byte" subcase
executes 0 assertions in all three files (proving it was dead code);
after the fix, it executes 11520 (unicode3), 34560 (unicode4), and 11520
(unicode5) assertions, and a deliberately reintroduced bug in the
lexer's byte4 range check causes it to fail (proving it is now
meaningful). Total per-file assertion counts grow by the same small
amounts, not by millions, and all other sections in these files still
pass unchanged.

Fixes #5416

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep full byte2 x byte3 combinatorics in the wrong-fourth-byte sections

The maintainer wants exhaustive coverage of every byte combination here
rather than the representative-prefix reduction, matching the style of
the sibling "wrong second/third byte" sections in the same files.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:16:43 +02:00
Niels Lohmann 29ba5973b6 Add missing diagnostic-positions test coverage (lifetime, input adapters, SAX) (#5482)
* Add missing diagnostic-positions test coverage (lifetime, input adapters, SAX)

Building on the merged unit-class_parser.cpp from #5417, add
characterization tests (regression protection for existing behavior, not a
behavior change) for JSON_DIAGNOSTIC_POSITIONS:

- value lifetime: copy ctor copies positions recursively, move ctor resets
  the moved-from value to npos, and mutating a parsed document (operator[],
  push_back, erase) leaves the parent's stale span and siblings' positions
  untouched while new values get npos.
- input adapters: wide-string input positions count transcoded UTF-8 bytes
  (not wide characters), BOM-prefixed input's start_pos() reflects the
  skipped 3-byte BOM, istringstream/ifstream/iterator-pair inputs report
  consistent (non-npos) positions, and binary formats (CBOR, MessagePack,
  UBJSON, BSON) always report npos.
- a user-constructed json_sax_dom_parser with no lexer (as used when driving
  json::sax_parse() directly) reports npos for every value, since it has no
  m_lexer_ref to source positions from.

While characterizing swap(), found that basic_json::swap() (and the friend
swap() that forwards to it) does not swap start_position/end_position,
unlike copy-assignment's operator=(basic_json), which does as part of its
copy-and-swap implementation. This looks like a real inconsistency/bug, but
per the scope of this test-only change it is only pinned (not fixed) here;
see the comment at the "swap() does NOT exchange positions" section.

Fixes #5420

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix MSVC source-encoding portability in the wide-string position test

Use é escapes instead of a literal UTF-8-encoded 'é' inside the L""
literal, so the wide string's content does not depend on the compiler's
assumed source character set (MSVC without /utf-8 decodes raw non-ASCII
source bytes using the system code page rather than as UTF-8, which was
producing a wstring of unexpected length/content and failing the
ws.size()/end_pos() assertions on Windows CI).

Also reworded a comment that unintentionally embedded the literal
substring "TODO check", which clang-tidy's google-readability-todo check
flags regardless of quoting context.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:16:40 +02:00
Niels Lohmann 75efd6b1c3 Test suite: cover untested macro configs, std::formatter branches, patch_inplace, and fix a duplicate TEST_CASE name (#5492)
* test: cover JSON_NO_IO, JSON_THROW/TRY/CATCH_USER, JSON_SKIP_LIBRARY_VERSION_CHECK, and JSON_DisableEnumSerialization in CI (#5423)

These four supported configuration macros were never actually compiled
anywhere in the test matrix:

- JSON_NO_IO and the JSON_THROW_USER/JSON_TRY_USER/JSON_CATCH_USER trio
  are exercised together in a new tests/src/unit-no_io_and_user_exceptions.cpp,
  which is automatically picked up by the existing unit-*.cpp test glob and
  thus built across the whole standard test matrix.
- JSON_SKIP_LIBRARY_VERSION_CHECK is exercised by a new, dedicated
  tests/src/skip_library_version_check.cpp, compiled directly by the new
  ci_test_skiplibraryversioncheck target in cmake/ci.cmake: the scenario it
  simulates (mixing two differently-versioned inclusions of the library)
  unavoidably triggers the compiler's own "macro redefined" warning, which
  would fail under the library's own -Weverything/-Werror unit test matrix
  for a reason unrelated to the macro under test.
- JSON_DisableEnumSerialization already had #if-guarded tests in several
  unit-*.cpp files (from #4384), but no CMake target ever actually set the
  JSON_DisableEnumSerialization CMake option, so that guarded code was never
  compiled. Add ci_test_disableenumserialization, mirroring the existing
  ci_test_noimplicitconversions/ci_test_noglobaludls targets. Building the
  full test suite with this option on surfaced one real, narrow gap: get<T>()
  on std::vector<std::byte> (used by unit-regression2.cpp's custom BinaryType
  tests) relies on std::byte being handled via enum serialization, so add the
  same #if-guard convention to the two affected SECTIONs there.

Both new CI targets are added to the ci_cmake_options matrix in
.github/workflows/ubuntu.yml, alongside the existing ci_test_* targets.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: cover multi-digit widths and bare alignment in std::formatter<json> (#5423)

Every existing std::formatter spec with a width used a single digit (e.g.
"{:2}"), so the width-parsing loop's accumulation of a second/third digit was
never exercised; add multi-digit width cases. Likewise, every existing spec
with an alignment character also had an explicit fill character, so the
bare-alignment branch (e.g. "{:<}", with no fill) was never exercised; add
cases asserting it keeps the default space indent character.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: add coverage for patch_inplace() (#5423)

patch_inplace() had no unit test at all. Add a happy-path case mirroring an
existing patch() example, and -- more importantly -- pin its distinguishing
contract versus patch(): when a multi-operation JSON Patch fails partway
through, patch_inplace() (which mutates the document directly, operation by
operation) leaves whatever operations already succeeded applied, whereas
patch() (which applies the patch to an internal copy that is discarded on
exception) leaves the original completely untouched either way. Verified
empirically against the current implementation before writing the assertions.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: fix duplicate TEST_CASE name in unit-no-mem-leak-on-adl-serialize.cpp (#5423)

Two distinct TEST_CASEs were both named "check_for_mem_leak_on_adl_to_json-2".
doctest allows duplicate names, so both still ran, but it makes
--test-case=<name> filtering and reporting ambiguous. Rename the second one
to "-3", continuing the existing "-1"/"-2" sequence.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: add direct coverage for the std::u8string to_json overload (#5423)

The ADL to_json overload for std::basic_string<char8_t, ...> was only ever
reached indirectly, via std::filesystem::path::u8string(). Add a test that
constructs a json value directly from a std::u8string, gated the same way as
the overload itself (include/nlohmann/detail/conversions/to_json.hpp): behind
both the std::filesystem::path feature guard and __cpp_lib_char8_t, since the
overload only exists when both are satisfied.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* test: verify move semantics of byte_container_with_subtype's rvalue constructors (#5423)

The two rvalue-reference constructors were never distinguished from their
const-lvalue-reference twins by any test. Add a "move semantics" section that
constructs from an rvalue std::vector, checks the resulting container keeps
the exact same buffer address as the source (a stronger check than just
observing the source ended up empty, since a copy-then-clear could do that
too), and confirms the source vector was left empty.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard patch_inplace() partial-application test against JSON_NOEXCEPTION

The "distinguishing contract vs patch(): partial application on
failure" test relies on doc.patch_inplace(patch) actually throwing so
the partially-applied state can be observed right after the throw
point. Under ci_test_noexceptions, JSON_THROW() calls std::abort()
instead of throwing, and doctest's --no-throw test filter (which that
CI job passes) makes CHECK_THROWS_AS() a no-op that never even
evaluates its expression -- so patch_inplace() is never called and the
follow-up assertions fail against the untouched original document.

Guard the whole SECTION with #if !defined(JSON_NOEXCEPTION), following
the same convention already used elsewhere in the test suite (e.g.
unit-class_parser.cpp) for exception-dependent tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix MSVC C2220 in the std::u8string conversion test

MSVC's C5321 ("nonstandard extension used: encoding '\xNN' as a
multi-byte utf-8 character") is promoted to a hard error by our MSVC CI
configs. It fires because the test composed a non-ASCII UTF-8 sequence
inside a u8"" literal using raw \x byte escapes; MSVC treats that as
nonstandard and suggests using \u universal-character-names instead,
which every compiler agrees on and which compiles down to the exact
same encoded bytes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard the JSON_THROW_USER test against JSON_NOEXCEPTION and GCC's -Wunused-result

Two independent CI configurations failed to build/run this new test:

- ci_test_noexceptions runs the whole suite with -DJSON_NOEXCEPTION and
  doctest's "--no-throw" filter, which compiles CHECK_THROWS_AS() down
  to a no-op that never even invokes the guarded expression. Since this
  test's whole point is to observe json_throw_user_call_count after
  json::parse()/at() actually throw, it can't be meaningfully run under
  that filter (our JSON_THROW_USER override still throws real
  exceptions regardless of JSON_NOEXCEPTION, but the assertion never
  gets a chance to run). Guard the TEST_CASE with
  #if !defined(JSON_NOEXCEPTION), mirroring the existing precedent in
  unit-json_patch.cpp.

- ci_test_gcc and ci_test_standards_gcc(11) failed with
  -Werror=unused-result on the discarded json::parse() return value.
  json::parse() is marked warn_unused_result, and unlike a real
  [[nodiscard]] attribute, GCC does not consider that satisfied by
  doctest's (void)-cast around the expression in C++11 mode. Assign the
  result to a discarded local instead, matching the established
  `json _ = json::parse(...)` idiom already used throughout
  unit-class_parser.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Suppress a clang-tidy false positive on an intentional defensive copy

performance-unnecessary-copy-initialization suggests copy_for_patch
could be a reference since it's never modified -- but the copy is the
point: it guards against a hypothetical regression where patch()
mutates its receiver, which a reference could never catch (the
follow-up assertion would just compare `original` to itself).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy findings in the JSON_NO_IO/JSON_THROW_USER test

- bugprone-macro-parentheses: wrap the JSON_THROW_USER macro argument in
  parentheses at the throw site.
- modernize-raw-string-literal: switch two escaped JSON string literals to
  raw string literals.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:15:46 +02:00
Niels Lohmann f58db1c9e8 Add test coverage for ordered_json/alt_json across binary formats and patch/diff/flatten APIs (#5480)
* Add test coverage for ordered_json/alt_json across binary formats and patch/diff/flatten APIs

Closes a test-coverage gap from #5421: ordered_json (and the alt_string-based
basic_json specialization from unit-alt-string.cpp) were never round-tripped
through the binary formats (CBOR/MessagePack/UBJSON/BSON/BJData), nor through
flatten()/unflatten(), diff()/patch()/patch_inplace(), or merge_patch(). Also
adds a std::formatter<ordered_json> spot-check, mirroring the precedent set
by the format_as() ADL-deduction test.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Pass alt_string's std::string constructor argument by value (clang-tidy modernize-pass-by-value)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:15:20 +02:00
Niels Lohmann ff65f688f7 Cover binary values in the indentation regression tests (#5533)
#5285's indentation regression test didn't exercise json::binary,
which serializes as an object but always writes its byte array
compactly (dump_byte()). #5186 had covered this case before it was
closed as superseded; port just that coverage here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-16 20:11:19 +02:00
dependabot[bot] af91eee2cc Bump the codeql-action group with 4 updates (#5536)
Bumps the codeql-action group with 4 updates: [github/codeql-action/init](https://github.com/github/codeql-action), [github/codeql-action/autobuild](https://github.com/github/codeql-action), [github/codeql-action/analyze](https://github.com/github/codeql-action) and [github/codeql-action/upload-sarif](https://github.com/github/codeql-action).


Updates `github/codeql-action/init` from 4.37.9 to 4.38.0
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/cdf488f595d80d6e07e03d4674febd5ab45fa938...b96794f015dfd88f77b49b1c93e0fa7110f94c63)

Updates `github/codeql-action/autobuild` from 4.37.9 to 4.38.0
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/cdf488f595d80d6e07e03d4674febd5ab45fa938...b96794f015dfd88f77b49b1c93e0fa7110f94c63)

Updates `github/codeql-action/analyze` from 4.37.9 to 4.38.0
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/cdf488f595d80d6e07e03d4674febd5ab45fa938...b96794f015dfd88f77b49b1c93e0fa7110f94c63)

Updates `github/codeql-action/upload-sarif` from 4.37.9 to 4.38.0
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/cdf488f595d80d6e07e03d4674febd5ab45fa938...b96794f015dfd88f77b49b1c93e0fa7110f94c63)

---
updated-dependencies:
- dependency-name: github/codeql-action/init
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: codeql-action
- dependency-name: github/codeql-action/autobuild
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: codeql-action
- dependency-name: github/codeql-action/analyze
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: codeql-action
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.38.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: codeql-action
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-16 20:11:06 +02:00
Niels LohmannandClaude Opus 5 c72f37a40d Support custom object/array types and improve template parameter handling (#5443)
* docs: document the implicit requirements on basic_json's template parameters

The requirements that basic_json places on its eleven template parameters
were only implied by how the library uses the resulting object_t, array_t,
string_t, etc. Consumers had to discover them by trial and error.

Add "Template Parameter Requirements" collecting them, split into what is
always required and what is only required when a particular part of the API
is instantiated. Notable findings that were previously undocumented:

- ObjectType must provide a key_compare member type (actual_object_comparator
  names object_t::key_compare in both arms of a std::conditional), and its
  third template parameter is used as a comparator, so std::unordered_map
  cannot be used without a wrapper.
- ArrayType must provide capacity() -- push_back(), emplace_back(),
  operator+=(), and operator[](size_type) call it unconditionally -- and
  needs random-access iterators, so std::deque and std::list do not work.
- StringType needs contiguous, null-terminated data(), a one-byte value_type,
  and either assignability from std::to_string or an ADL int_to_string().
- NumberFloatType must be float, double, or long double for parsing and
  serialization; the integer types must satisfy std::is_integral.
- AllocatorType must be stateless, support incomplete types, and use plain
  pointers.
- BooleanType and the number types are union members and must be trivial.

Link the new page from the basic_json overview, the types feature page, and
the individual type alias pages, and correct the container examples given for
ObjectType (std::unordered_map) and ArrayType (std::list), which do not work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix object_comparator_t for object types without key_compare

detail::actual_object_comparator selected between object_t::key_compare and
default_object_comparator_t with std::conditional. Both type arguments of
std::conditional are named eagerly, so object_t::key_compare had to exist
regardless of the condition, and the has_key_compare guard added in 3.11.0
never took effect: any ObjectType without a key_compare member type failed to
compile while instantiating basic_json itself.

Use detected_or_t instead, which resolves through a SFINAE partial
specialization and only names object_t::key_compare when it exists. The
selected type is unchanged for every object type that compiled before, so
object_comparator_t -- a public member type -- keeps its meaning and ABI.

has_key_compare had no other users and is removed.

Add a regression test using an adapter around std::unordered_map, which has no
key_compare; it fails to compile without this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: list the types that are known to work for each template parameter

Follow up on the template parameter requirements page: state, for every
template parameter, which concrete types work and where they stop working.
Each entry was verified by compiling and running a common workload (DOM
access, dump, parse, CBOR/MessagePack round-trip, flatten, hash) against that
instantiation.

Findings worth calling out:

- ObjectType no longer needs a key_compare member type, so the std::unordered_map
  adapter only has to restore the template argument order. A hash-ordered
  ObjectType works everywhere except unflatten(), which reconstructs an array
  only when it meets the reference token 0 before the other indices.
- ArrayType: std::deque works when wrapped to add capacity(); std::list does not.
- StringType: std::pmr::string and std::basic_string with a custom allocator
  compile for the DOM, dump, and parse, but not for the binary readers, flatten,
  or diff, because the library assigns std::string values to string_t and
  int_to_string cannot be overloaded for a type in namespace std.
- NumberFloatType: long double works for dump and parse but not for the binary
  formats, which have no encoding for it.
- BinaryType: std::vector<std::byte> supports assignment, get, and the binary
  formats, but neither dump nor std::hash<basic_json>.

Also record the object_comparator_t fix in its version history.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix unflatten and binary dumping for non-default configurations

unflatten() decided between array and object by looking at the first reference
token it happened to see for a node: it started an array only when that token
was 0. With a sorted object type the token 0 always arrives first, so the
result was correct by accident; with an object type whose iteration order is
unspecified, {"/c/2":3,"/c/1":2,"/c/0":1} unflattened to an object with the
keys "0", "1", and "2" instead of an array.

Collect the pointer prefixes that have a reference token 0 among their children
before building the result, and let get_and_create() consult that set. The
outcome is now independent of the iteration order and matches, for every input,
what a sorted object type produced before: a value is restored as an array if
and only if one of its keys is 0. Iterating the flattened object in a different
order would have been simpler, but it would have changed the key order of the
result for insertion-ordered object types.

The serializer, std::hash, and the UBJSON writer converted the elements of a
binary value to an integer implicitly, which does not compile for a BinaryType
whose value type is std::byte, and which made dump() write the bytes of a
signed value type as negative numbers. Convert to std::uint8_t explicitly in
all three places, so every byte type dumps as 0..255. The default
std::vector<std::uint8_t> configuration is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: note which Abseil containers can be used as template arguments

Checked against Abseil release 20250127.0 with the same workload as the other
entries on the page (DOM access, dump, parse, CBOR/MessagePack/UBJSON
round-trip, flatten, hash), with and without JSON_DIAGNOSTICS.

absl::flat_hash_map and absl::node_hash_map work as ObjectType through an
adapter that restores the template argument order and makes erase(iterator)
return the following iterator, which Abseil's returns as void. The page now
carries that adapter, and notes that absl::flat_hash_map does not keep
references to the mapped values valid across insertions while
absl::node_hash_map does. Both have a capacity() member, so JSON_DIAGNOSTICS
already refreshes the parent pointers conservatively for them.

absl::btree_map and absl::InlinedVector cannot be used at all: object_t and
array_t are formed while basic_json is still incomplete, and both inspect
their value type at class scope. std::map and std::vector are required by the
standard to tolerate this, third-party containers generally are not, so the
page states the constraint on its own rather than only per container.

absl::InlinedVector does work as BinaryType, where it is instantiated with a
complete type. absl::FixedArray and absl::Cord are not usable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Relax the ArrayType and ObjectType requirements

Two requirements forced users of otherwise suitable containers to write a
wrapper, and neither was load-bearing.

array_t::capacity() was read in push_back(), emplace_back(), operator+=(), and
operator[](size_type), but set_parent() only looks at the value under
JSON_DIAGNOSTICS; without diagnostics it was computed and discarded. Read it
through array_capacity(), which reports unknown_size() when diagnostics are off
or when the array type has no capacity() at all, and treat an unknown capacity
as "the elements may have moved" so the parent pointers are refreshed
conservatively. std::deque now works as ArrayType, in both builds, and
capacity() is no longer named at all in a default build. Since the capacity is
now only meaningful for array insertions, it moves out of set_parent() into
set_parent_after_array_insert().

basic_json::erase(iterator) assigned the object's erase() return value, which
requires the container to return the following iterator. Abseil's hash maps
return void to avoid computing a successor the caller may not need. Detect that
and compute the successor before erasing; containers that return an iterator,
including the vector-backed ordered_map where a precomputed successor would be
wrong, keep the existing path.

Together these leave an Abseil hash map needing only an alias that restores the
template argument order, and no adapter at all for std::deque.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not require string_t to be convertible from std::string

Three places built a std::string and handed it to something expecting a
string_t: the UBJSON high-precision number reader, which every binary reader
instantiates, and the BSON writer's array element size calculation and write.
That silently required string_t to be implicitly convertible from std::string,
which std::string itself and types with a string_view conversion satisfy, but
many string types do not.

Construct the string_t explicitly from the data and size, which the
requirements already cover. This makes boost::container::string, eastl::string,
std::pmr::string, and std::basic_string with a custom allocator work as
StringType, none of which could previously be used with any binary format.

Add binary format coverage to the alt_string test, which had none, including a
UBJSON high-precision number -- the case that goes through the reader path.
BSON stays uncovered there: it additionally needs string_t::find(value_type),
which alt_string does not provide.

Also record which containers from Boost, Abseil, and EASTL work for each
template parameter, and correct two claims: std::pmr::string is usable after
this change, and tsl::ordered_map is not usable at all, because its iterators
expose the mapped value as const while basic_json modifies it in place.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: record compatibility for the common header-only hash maps

ankerl::unordered_dense (map and segmented_map), phmap (flat_hash_map and
node_hash_map), and robin_hood::unordered_flat_map all work as ObjectType
through the same adapter as Abseil's and Boost's hash maps, which only has to
restore the template argument order.

phmap::btree_map and robin_hood::unordered_node_map do not: like the other
btree containers they require a complete value type.

Note that none of these hash maps defines key_compare, so every one of them
depends on object_comparator_t falling back to default_object_comparator_t.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: record Folly and the remaining vector replacements

Folly works, with the caveat that its headers need C++20: folly::fbstring as
StringType, folly::fbvector and folly::small_vector as ArrayType,
folly::fbvector<std::uint8_t> as BinaryType, and folly::F14NodeMap as
ObjectType through the usual argument-order adapter. folly::F14FastMap is the
exception and requires a complete value type.

For ArrayType, boost::container::devector, boost::container::static_vector
(within its fixed capacity), and std::pmr::vector work as well.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: cover fifo_map, gtl, folly::sorted_vector_map, and Qt

nlohmann::fifo_map works through the adapter that has always been documented
for it, and preserves the insertion order. Restore its mention in the object
order page, which was dropped together with the tsl::ordered_map one: unlike
ordered_map it keeps a lookup index, so it is the insertion-ordered option
without the quadratic cost.

gtl::flat_hash_map and folly::sorted_vector_map work as well, the latter
through an alias that drops the allocator, whose value type it disagrees on.
gtl::btree_map does not, for the same reason as the other btree containers.

None of the Qt containers can be used, each for its own reason: QMap has no
value_type, QHash iterators yield the mapped value rather than a pair, QList
has no max_size(), QByteArray spells empty() as isEmpty(), and QString is
UTF-16.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: qualify the std::pmr::string support claim

Listing std::pmr::string as fully supported was an overclaim: it was only ever
checked with the default memory resource, which is not what PMR is for.

basic_json cannot be given an allocator or a memory resource, so a pmr string
inside a value always allocates from std::pmr::get_default_resource(), and
assigning an arena-backed string into a value silently drops its resource,
because polymorphic_allocator does not propagate on copy construction. Passing
polymorphic_allocator as AllocatorType does not compile either. Only the
process-global set_default_resource() redirects these allocations.

Say so, and separate the row from std::basic_string with a custom stateless
allocator, which is unaffected.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: remove a duplicated StringType compatibility section

The StringType section carried two 'Compatible types' tables and two copies of
the reference-implementation tip. The second table was a stale copy from before
the binary format string fixes and still listed std::pmr::string and
std::basic_string with a custom allocator as unusable, contradicting the
corrected table a few lines above it, and it dragged along the old explanation
that blamed int_to_string.

Drop the stale copy and put the surviving table before the notes, so the
'see below' in the std::pmr::string row points forwards.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: correct the template parameter requirements after independent verification

Every claim on the page was re-checked by compiling and running it, including
the rows that say a type cannot be used, which were checked to fail for the
documented reason and not merely to fail. Twenty-four claims were wrong.

The most consequential: the incomplete-type constraint applies to ObjectType
only. object_t is instantiated inside the class definition, because it is
probed for key_compare; array_t is only named there and is not instantiated
until basic_json is complete. So eastl::vector, QList and QVector are not
excluded by incomplete types at all -- they simply have no max_size() -- and
absl::InlinedVector is excluded for a subtler reason of its own.

Further corrections: ObjectType does not need erase(key), which has a fallback,
but does need at(key) for UBJSON output; only == and < are used, or == and <=>
under C++20, not all six; the documented adapter does not fit ankerl or
robin_hood. ArrayType needs no initializer-list insert, and value_type, the
(count, value) constructor and swappability are per-function, not always.
BinaryType needs a range insert for CBOR indefinite-length byte strings and
does not need push_back. StringType needs append(const StringType&)
unconditionally, and does not need operator!= or operator== against const
char*; empty(), resize(n) and reserve(n) are per-subsystem; int_to_string is
needed by diff, items and std::hash rather than by JSON Pointer or flatten.
BooleanType must be implicitly convertible from bool, and JSONSerializer's
second parameter need not carry a default.

std::pmr::string was wrong in the other direction this time: a moved-in string
does keep its memory resource, and later growth allocates from it. Only copies
land on the default resource.

Five requirement violations are not caught at compile time rather than the two
the page claimed; they are now listed together up front. Split every
compatibility table into what works and what does not, as the reasons in the
second half are the useful part.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Reduce the string_t and array_t members the library requires

Several members were required only because of how the library happened to
be written, not because the functionality needs them. Dropping them widens
the set of usable string and array types, and one of them was also a
performance problem.

string_t:

- c_str() is gone. Every call site already knew the length and passed it
  along, so data() is enough. The one place that did not, the diagnostics
  path in exceptions.hpp, now builds the token from data() and size(),
  which also stops it from truncating keys that contain a null byte.
- back() is gone; the serializer indexes the last character instead.
- find(str, pos), replace(), and substr() are gone. escape() and
  unescape() rebuilt the string with one replace() per escaped character,
  which moves the tail every time: escaping a string of n characters that
  all need escaping cost O(n^2). Both now scan with find_first_of() -- a
  member the pointer parser already required -- and append whole runs, so
  the common case is one search and one copy. Escaping 64000 tildes drops
  from 717 ms to 20 ms; a string with nothing to escape gets faster too
  (8.4 ms to 5.8 ms), because the scan is still a single memchr per pass.
  json_pointer::split() takes its reference tokens with the
  (const char*, size_type) constructor rather than substr().
- json_pointer::to_string() accumulates with concat<string_t> instead of
  letting concat default to std::string and converting afterwards, so
  streaming a json_pointer no longer requires string_t to be assignable
  from a std::string.

array_t:

- at(size_type) is gone. basic_json::at(size_type) checked the index by
  calling array_t::at() and translating std::out_of_range, which also
  required the array type to throw that exact exception. It now compares
  against size() and uses operator[]. The thrown exception, its message,
  and the behaviour under JSON_NOEXCEPTION are unchanged.

The BSON writer wrote the terminating null byte out of the string's own
buffer (size() + 1). It now writes the byte itself, so string_t::data()
need not be null-terminated for to_bson().

The tests pin the reduced API: alt_string loses the five dropped members
and gains coverage of the escaping paths, and a std::vector whose at() is
hidden is used as an ArrayType.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: record the reduced string_t and array_t requirements

Drop c_str(), back(), find(str, pos), replace(), and substr() from the
StringType requirements and at(size_type) from the ArrayType ones, and
note the string assignment the JSON pointer code performs. Streaming a
json_pointer no longer needs assignability from a std::string.

Add the non-null-terminated data() to the list of violations that are not
diagnosed at compile time -- it was described in the StringType section
but missing from the summary at the top -- and correct the QString row,
which no longer fails for the c_str() it lacks.

JSON_CATCH_USER no longer wraps a catch of std::out_of_range: the last one
went away with array_t::at(). Describe what the library actually catches.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use character literals for the signed BinaryType test

MSVC rejects char(0xFF) with C4310 (cast truncates constant value),
which the Windows workflow treats as an error. The character literals
carry the same byte values without a narrowing cast.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not instantiate a hash map with an incomplete basic_json in the tests

object_t is probed for key_compare inside the definition of basic_json, so
it is instantiated while basic_json is still incomplete. Whether a hash map
survives that depends on the standard library: libstdc++ 9 needs the size of
the mapped type to instantiate std::unordered_map's node type and rejects
the adapter, which broke the GCC 9 builds.

The test now derives its no-key_compare object type from std::map -- which
does cope -- and shadows the inherited key_compare member type with an
entity that is not a type, so the library's probe finds none, exactly as for
a hash map. The unflatten() order-independence checks in unit-json_pointer
already cover the behaviour that the unordered object type was there for.
The limitation is documented for std::unordered_map.

Also address two Clang-Tidy findings the earlier commits introduced:
erase_from_object() declares its iterator with auto, and at(size_type) checks
the type first and then falls through to the return instead of throwing from
an else branch.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep diagnostic key paths null-terminated

Building the token from data() and size() kept an embedded null byte in the
key, and since what() hands out a C string, that truncated the whole message
rather than just the key: to_bson() on a key containing U+0000 reported
"[json.exception.out_of_range.409] (/en" instead of the full explanation.
This broke test-bson under JSON_DIAGNOSTICS.

Constructing from data() alone stops at the first null byte, which is what
c_str() did before, so the message is unchanged -- without requiring
string_t to provide c_str().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not parse the value in the array-at() test

JSON_DIAGNOSTIC_POSITIONS adds the byte range of the value to the exception
message, which a parsed value has and an in-memory one does not, so the two
message checks failed in that configuration. Build the array in memory
instead of parsing it; the test is about at(size_type) not needing
array_t::at(), and the byte range is beside the point.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the custom BinaryType tests into their own translation unit

The two sections added to unit-regression2.cpp brought a third full
basic_json instantiation into a translation unit that was already large.
With Clang on MinGW that pushed the object over the reach of a 32-bit
relocation and test-regression2_cpp20.exe failed to link:

    relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'

unit-regression2.cpp is restored to exactly what it was before, and the
coverage moves to unit-custom-binary-type.cpp, next to the object and array
type tests it belongs with. The signed value type is now also covered in
C++11, where std::byte is not available.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not require the container iterators to be nothrow move constructible

iter_impl declared its defaulted move operations noexcept. The exception
specification a defaulted function gets implicitly follows from its members,
here internal_iterator, which holds the object and array iterators. libstdc++
gives std::deque's iterator a user-provided copy constructor without noexcept
before version 11, so the implicit specification is noexcept(false) and does
not match the declared one. That deletes the function -- and with g++ 4.8,
which predates CWG 1778, it is an error outright:

    error: function 'iter_impl<basic_json<std::map, std::deque> >::iter_impl(
    iter_impl&&)' defaulted on its first declaration with an
    exception-specification that differs from the implicit declaration

So std::deque, which this branch documents as a usable array type, could not
be used with an older standard library. Leaving the specification to be
computed cannot mismatch; iteration_proxy_value already spells out the same
condition next door.

The default configuration is unaffected: json::iterator, json::const_iterator
and ordered_json::iterator stay nothrow move constructible and move
assignable, which the test now checks so it cannot regress unnoticed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address two Clang-Tidy findings the custom container tests exposed

Both come from instantiating basic_json with containers other than the
default ones, and neither shows up with the Clang-Tidy version available
outside CI:

- insert(const_iterator, basic_json&&) forwards its by-value iterator to
  the const-reference overload. performance-unnecessary-value-param asks
  for the copy to be a move; it only fires for an iterator that is not
  trivially copyable, as std::deque's is not. The NOLINT on the function
  does not cover it, because the finding is reported where the parameter
  is used rather than where it is declared. Move it, which is what the
  check asks for and is a (very small) improvement in its own right.

- cppcoreguidelines-use-enum-class rejects the unnamed enum that shadowed
  the inherited key_compare member type. An enum class would not do, since
  it declares a type of that name and the probe would find it again; a
  member function declaration hides the name just as well.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Assert the iterators' exception specification relative to the container

The test pinned that nlohmann::json's iterators stay nothrow movable after
iter_impl's defaulted move operations lost their declared noexcept. That is
not a property of the library, though: the exception specification is now
computed from the container iterators, so it holds only for standard library
implementations whose iterators are themselves nothrow movable.

MSVC's checked iterators before VS2017 are not -- _Iterator_base12 registers
the iterator with the container's debug proxy in a copy constructor that
carries no noexcept -- so the assertions fail on a Visual Studio 2015 debug
build, which is the one debug configuration in the AppVeyor matrix and has no
counterpart in the GitHub Actions matrix.

Assert what the change actually guarantees instead: the iterators are nothrow
movable exactly when the object and array iterators they are built from are.
That still pins the default configuration against a silent regression, and it
is true whatever the standard library provides.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Detect a void-returning erase() through a named trait

erase_from_object() distinguished its two overloads with a decltype of a
member call written inline in a default template argument. Every other
detection in the library goes through the detector machinery in detected.hpp
instead -- has_erase_with_key_type is the same question about the same member
function -- and the inline form is the one shape older compilers are least
reliable about.

Express it the same way: detect_erase_with_iterator plus is_detected_exact,
both of which the library already relies on elsewhere. No behaviour changes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Give the custom container types only the constructors the library uses

The three container types in the new tests inherited every constructor of
their base with using Base::Base. That asks for more than the test needs: the
library builds an object or an array by default construction, by copy or move,
and -- when converting between two basic_json types or from an initializer
list -- from an iterator range. Declaring those directly makes the requirement
visible in the test, and keeps object types out of a corner where a compiler
has to declare std::map's whole constructor set for a derived class while
basic_json is still incomplete.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Temporarily disable the new custom container tests

AppVeyor is the only CI that builds MSVC 2015 and 2017, and it has now
rejected three heads of this branch. Its build log is not reachable from
where this is being worked on, so the verdict is a single bit and the cause
has to be narrowed down by bisection.

Everything else stays: the library changes, the reduced alt_string, and the
unflatten() tests. If AppVeyor passes with these three translation units
disabled, the cause is one of the six basic_json instantiations they add; if
it fails, it is in the library. Either way this commit is reverted.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard the disabled tests with a macro rather than #if 0

Clang-Tidy's readability-avoid-unconditional-preprocessor-if rejects a literal
#if 0. Use a macro that is never defined instead, which the check does not
look at. Still temporary, and reverted together with the previous commit.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Re-enable the array and binary container tests

AppVeyor passed with all three new translation units disabled, so the library
changes, the reduced alt_string, and the unflatten() tests are fine on MSVC
2015 and 2017; the cause is one of the six basic_json instantiations the new
tests add.

Bring back two of the three. If AppVeyor passes again, the cause is in
unit-custom-object-type.cpp, which is the one still disabled; if it fails, it
is in one of these two and needs one more split. Still temporary.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Diagnose two silently violated template parameter requirements

Both were on the list of requirements that are not caught at compile time and
corrupt values rather than failing, and both are a plain size comparison:

- A BinaryType whose value_type is wider than one byte, which the readers and
  writers reinterpret as raw bytes anyway.
- A NumberUnsignedType too narrow to hold the absolute value of every
  NumberIntegerType value, which makes basic_json(INT64_MIN).dump() yield -0
  for std::int64_t with std::uint32_t.

Neither static_assert rejects a configuration that worked before: both only
fire where the result was already wrong. Also add the two comments the review
asked for, in write_bson_string() and calc_bson_array_size(), matching the
ones their counterparts already carry.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Align the template parameter tables and record what is now diagnosed

Every table in the page is reformatted so each column is exactly as wide as
its widest cell, which is what the review asked for in a dozen places: the
separator rows that ran two dashes long, the stray spaces, and the columns
padded well past their content.

The row listing six containers that require a complete mapped type is split
in two so that one cell no longer sets the width of the whole table.

Content changes: NumberUnsignedType is described as any unsigned integer type
at least as wide as NumberIntegerType rather than any unsigned integer type;
the two requirements that are now static_asserts move out of the list of
violations that are not caught at compile time; and the two places that
require a non-const operator[] say why data() will not do (std::string has no
non-const data() before C++17).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Bisect the other way: only the object container tests

The previous head touched only docs/, which AppVeyor's only_commits filter
skips, so it produced no build and no status at all -- the pull request looked
green without ever having been built on MSVC 2015 or 2017.

Swap the guards instead of repeating that step: unit-custom-object-type.cpp is
enabled and the array and binary translation units are disabled. AppVeyor
already passed with all three disabled, so a failure here pins the cause on
no_key_compare_json or void_erase_json, and a pass pins it on the array or
binary file. Still temporary.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Split the two object types apart

AppVeyor failed with only unit-custom-object-type.cpp enabled and passed with
all three new translation units disabled, so the cause is one of the two
object types in this file and not the array or binary ones.

Guard out void_erase_map and leave no_key_compare_map, which separates the two
constructs under suspicion: shadowing the inherited key_compare member type
with an entity that is not a type, and hiding the inherited erase with a
void-returning overload. A failure here points at the first, a pass at the
second. Still temporary.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Build the no-key_compare object type by composition, not inheritance

The "object type without key_compare" test failed on AppVeyor's MSVC
2017 jobs (/std:c++17): its no_key_compare_map derived publicly from
std::map and shadowed the inherited key_compare type with a same-named
member function, relying on ordinary member hiding to make key_compare
unreachable as a type for the library's detection trait. MSVC 2017
does not honor that hiding for a typename-qualified lookup performed
from outside the class and still resolves key_compare to the base's
comparator type, so object_comparator_t incorrectly picked it up
instead of falling back to default_object_comparator_t.

Wrapping a std::map by composition instead removes the base class
entirely, so there is no key_compare to find under any lookup rule,
on any compiler. Also drops the now-unneeded JSON_BISECT_CUSTOM_CONTAINER_TESTS
guard left over from narrowing this down: the void_erase_map test in
the same file was never the cause and is re-enabled unconditionally.

Verified locally with clang++ and g++ under C++17 and C++20.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Re-enable the array and binary custom-container tests

unit-custom-array-type.cpp and unit-custom-binary-type.cpp were still
guarded behind JSON_BISECT_CUSTOM_CONTAINER_TESTS from bisecting the
AppVeyor failure fixed in 7c39f3227, which was unrelated to either
file. The macro was never defined, so none of these tests actually ran
in CI. Verified locally with clang++ and g++ under C++17 and C++20
before removing the guards.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix indentation of custom_object_type per astyle

The one-line function bodies in the composition-based no_key_compare_map
(7c39f3227) do not match the project's Allman brace style, which the
ci_test_amalgamation job enforces with astyle. Reformatted with the
pinned astyle 3.4.13; no functional change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* docs: add compiled reference implementations for the container template parameters

Each of ObjectType, ArrayType, StringType, and BinaryType now links to a
minimal, self-contained header (docs/mkdocs/docs/examples/custom_*_type.hpp)
that wraps the corresponding standard container by composition and satisfies
every "Always required" member listed on that page. Unlike the prose
requirement lists, these are real code: each header has a companion .cpp that
instantiates a basic_json specialization with it and is compiled and run by
the existing ci_test_examples check (docs/Makefile's check_output_portable),
so the reference implementations cannot silently drift from what the library
actually requires. The .output files were generated with that same target.

StringType's existing pointer to tests/src/unit-alt-string.cpp's alt_string
is kept alongside the new header as a more thorough, battle-tested example.

Verified locally: astyle (pinned 3.4.13, project .astylerc) on the new files;
clang++/g++ under C++11/17/20 for each example against the amalgamated
header; `make check_output_portable` in docs/; `mkdocs build --strict` and
scripts/check_structure.py for the page itself.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Declare no_key_compare_map's accessors noexcept

The GCC C++20 job builds with -Wnoexcept and -Werror, and the standard
library takes noexcept(c.begin()) and noexcept(c.end()) in ranges_base.h and
range_access.h. Forwarding to std::map without repeating its noexcept made
those expressions false, which the warning reports as an error:

  error: noexcept-expression evaluates to 'false' because of a call to
         no_key_compare_map<...>::begin()          [-Werror=noexcept]
  note:  but ... does not throw; perhaps it should be declared 'noexcept'

Give the accessors the exception specification of what they forward to.
std::map declares begin, end, cbegin, cend, empty, size, max_size, and clear
noexcept, so the wrapper does too. swap is left alone: std::map's is only
conditionally noexcept, and nothing asks for it.

void_erase_map is unaffected because it still derives from std::map and
inherits accessors that already carry the specification.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Write out what the defaulted constructor of no_key_compare_map implied

ci_test_gcc builds with -Weffc++, which asks for data to be initialized in a
member initialization list; a defaulted default constructor does not do that:

  error: 'no_key_compare_map<...>::data' should be initialized in the member
         initialization list                              [-Werror=effc++]

Writing the constructor out satisfies that but drops the exception
specification the defaulted one carried, which -Wnoexcept then objects to
where the standard library takes noexcept(construct(...)). Declare it the way
the defaulted constructor was: noexcept when the wrapped map's default
constructor is.

This is the cost of composition -- inheritance carried std::map's exception
specifications and initialization for free, and forwarding by hand has to
restate them.

Checked with the repository's own GCC warning set from cmake/gcc_flags.cmake,
all 346 flags, at C++11, C++17 and C++20: no diagnostics for this file, nor
for the two custom container translation units that were disabled while the
MSVC failure was narrowed down and are built again now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark no_key_compare_map::swap noexcept

Clang-Tidy rejects a swap that is not:

  error: swap functions should be marked noexcept
         [cppcoreguidelines-noexcept-swap,performance-noexcept-swap]

It was left unmarked on the grounds that std::map::swap is only
conditionally noexcept, so an unconditional promise would be wrong for a
comparator or allocator that can throw while swapping. Both concerns are met
by taking the specification from the wrapped map rather than asserting one:
noexcept(noexcept(data.swap(other.data))). Clang-Tidy accepts that, and no
NOLINT is needed.

Last in the series of specifications that inheritance used to supply and
composition has to write out by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 20:46:37 +02:00
Niels Lohmann 6487678bc5 Fix swap(array_t&)/swap(object_t&) to update parent pointers under JSON_DIAGNOSTICS (#5464)
Both overloads swapped the underlying container storage but never called
set_parents(), leaving elements moved into *this with stale m_parent
pointers (typically nullptr from the free-standing array_t/object_t).
This produced wrong JSON Pointer paths in diagnostic messages and could
trip assert_invariant() on subsequent copies. Mirrors the fix already
applied in swap(reference other).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-15 20:46:28 +02:00
Niels Lohmann 3bfe2b6da7 Stop the privacy plugin from downloading the repology.org badges (#5523)
repology.org refuses requests coming from GitHub Actions runners. The
Material privacy plugin logs the failed download as a warning, but then
raises FileNotFoundError when it reads back the cache entry it never
wrote, so ci_test_build_documentation aborts the whole documentation
build. develop and every open pull request are currently red because of
this.

The badges are fine for readers -- repology only rejects non-browser
clients -- so exclude them from the plugin's asset self-hosting and keep
them as external references. All other external assets (fonts, shields,
CDN files) are still downloaded and self-hosted as before.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-14 21:19:41 +02:00
Niels Lohmann 1da2f68992 Only reserve array capacity if the array type supports it (#5522) 2026-09-14 06:44:01 +02:00
Niels Lohmann c0b2878a44 Swap diagnostic positions in basic_json::swap() (#5493)
basic_json::swap() (and the friend swap() that forwards to it) only
exchanged m_data.m_type/m_data.m_value, leaving start_position/end_position
untouched under JSON_DIAGNOSTIC_POSITIONS. This is inconsistent with
copy-assignment's operator=(basic_json), which swaps positions as part of
its copy-and-swap implementation, so after swap(a, b) each value ended up
with the other value's content but its own original position.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-13 17:43:59 +02:00
Niels Lohmann a50c2537eb Reserve capped array capacity in json_sax_dom_parser::start_array() for definite-length binary arrays (#5476)
* Reserve capped array capacity for definite-length binary arrays

CBOR, MessagePack, and the optimized [$type#count UBJSON/BJData form all
pass an exact element count to sax->start_array(len), but
json_sax_dom_parser::start_array() (and the callback variant) only used
len for an overflow check against max_size() and never reserved the
underlying vector, so each element triggered a reallocation cascade via
emplace_back().

Reserve upfront, but cap the reservation at 16384 elements: max_size()
for a std::vector is far larger than any realistic input, so an
unbounded reserve(len) would let a crafted/truncated header (e.g. CBOR
0x9A + a huge uint32 count with no data) trigger a multi-gigabyte
allocation attempt instead of the normal graceful parse_error. With the
cap, a hostile length still fails fast with the existing parse_error,
while realistic arrays get a single up-front allocation.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the huge-claimed-length DoS regression tests portable across size_t widths

On a platform where size_t is narrower than 64 bits (e.g. 32-bit mingw/msvc
x86), the previously-hardcoded huge test lengths either collide with that
platform's unknown_size() sentinel (CBOR/MessagePack, both using exactly
SIZE_MAX) or exceed the platform's smaller vector<json>::max_size()
(UBJSON/BJData's 0x7FFFFFFF), so the header is now rejected outright
(out_of_range.408) instead of being accepted and only found short of data
(parse_error.110). Both are safe, bounded rejections of the hostile input;
the property under test -- no attempt to allocate space for billions of
elements -- holds either way. Accept both outcomes instead of pinning the
64-bit-only exact result.

Also fixed an unrelated clang-tidy finding (google-readability-casting) on
the functional-style std::size_t(...) casts in the neighboring "arrays of
various sizes" section.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix remaining CI failures in the huge-claimed-length DoS regression tests

- Apply the same google-readability-casting fix (std::size_t{N} instead
  of std::size_t(N)) to the "arrays of various sizes" section in
  unit-msgpack.cpp, unit-ubjson.cpp, and unit-bjdata.cpp; only
  unit-cbor.cpp had been fixed previously, since clang-tidy's build
  didn't get far enough to report the other three in the same pass.

- json_sax_dom_parser::start_array()'s max_size() check calls JSON_THROW
  directly rather than going through sax->parse_error(), so unlike the
  scanner's own "not enough data" parse_error it is not gated by
  allow_exceptions=false. On a platform where a header's claimed count
  exceeds max_size() (e.g. 32-bit, for UBJSON/BJData's 0x7FFFFFFF test
  value), from_ubjson/from_bjdata(input, true, false) can therefore still
  throw instead of returning a discarded value. Make that assertion
  tolerant of either outcome, same as the main exception-catching check
  above it.

- Guard all four "a huge claimed length..." SECTIONs with
  #if !defined(JSON_NOEXCEPTION), matching this test suite's existing
  convention for exception-dependent tests: under JSON_NOEXCEPTION,
  JSON_THROW never produces a catchable C++ exception at all (it aborts
  the process), so a section that relies on try/catch to distinguish
  between two acceptable outcomes cannot be expressed under that build
  configuration regardless of platform.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use (std::min)(len, reserve_cap) instead of a ternary in start_array()

Addresses review feedback from @gregmarr on PR #5476.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-12 21:09:33 +02:00
Niels Lohmann aa391dc0a5 Fix two develop CI regressions: dump() nodiscard warning and binary-reader const-correctness (#5520)
* Discard dump()'s [[nodiscard]] return value in an exception-only check

CHECK_THROWS_WITH_AS(j.dump(), ...) called dump() only to trigger and
catch the exception, but never used the return value. dump() is
warn_unused_result, so GCC's pedantic build (-Werror --all-warnings)
rejected it as -Werror=unused-result, breaking ci_test_gcc. Wrapped in
utils::ignore_return_value(), matching every other such call in this
file.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark container_frame top as const in CBOR/UBJSON readers

clang-tidy's misc-const-correctness flagged these on PR #5520's CI: the
BSON sibling copy was already const, but these two were left mutable
even though only container_stack.back().remaining is ever written.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 17:55:46 +02:00
Niels Lohmann d2514a46f7 Add benchmarks for the binary readers (#5510)
The benchmark suite covered parsing JSON text, dumping, and serializing to
CBOR, but only one binary read: FromMsgpack. Nothing measured from_cbor,
from_ubjson, from_bjdata or from_bson, so a change to binary_reader.hpp had
no baseline to be compared against.

Add read benchmarks for every format, in the two shapes that matter: from a
contiguous buffer, which is what most callers pass, and from a FILE*, which
is what FromMsgpack already measures and which compiles to different code.
FromMsgpack itself is left untouched so its numbers stay comparable across
releases. The input is derived at setup time by serializing a parsed test
file, because the test data repository ships JSON only.

The test files are wide and shallow, but the readers' cost is per container,
so add three value shapes they do not cover -- deeply nested containers, many
sibling containers, and one flat array of numbers -- plus an indefinite-length
CBOR string, a form the writer never emits and which therefore has to be
assembled by hand. UBJSON and BJData are also captured in their size- and
type-annotated form, which the readers handle in a separate code path.

BSON requires an object at the top level, so it cannot reuse the array-rooted
test files; it is captured on the object-rooted ones, and the shapes are
wrapped in an object so every format measures the same value. The setup marks
the benchmark as skipped rather than letting the exception escape if that
requirement is ever violated.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 13:15:09 +02:00
Niels Lohmann 0452641c18 Read BSON documents without recursing per nesting level (#5508)
* Read BSON documents without recursing per nesting level

An embedded document (record type 0x03) or array (0x04) was read by calling
back into the document reader, which read its element list, which called the
element reader again for the next embedded one. The native call stack
therefore grew with the nesting depth of the input, and about seven bytes buy
a level, so a document of a few hundred kilobytes crashes the process
(#5104). This is the last of the four binary formats to still do that.

Apply the same shape as the other three: open_bson_document() reads the size
prefix and opens the document, parse_bson_element_internal() calls it for both
record types instead of recursing, and parse_bson_internal() loops over the
element list of whichever document is innermost, closing it when its
terminator is reached and resuming the one below.

check_bson_document_size() is unchanged, and so is when it runs: a document is
still measured from the byte before its size prefix to the byte after its
terminator, and still reported before the end event. The frame carries those
two values, which is what a per-document check needs once the reads are
interleaved rather than nested. Nothing else about the element reader changes.

unit-bson passes unchanged. Round trips through to_bson of nested objects,
arrays, arrays of objects and mixed nesting are identical to the previous
commit, as are the errors for a truncated document, an unsupported record
type, a negative size and a size that does not match, including their byte
offsets. A 30,000-level document built by to_bson is now read to completion
where it used to crash.

Note for sequencing: #5185 changes parse_bson_internal(), the element list and
the array reader, which are the functions this commit restructures. It should
land first; this commit then keeps its checks and moves them onto the loop.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make parse_bson_internal's end-of-document top a copy, not a reference

Same issue as the CBOR and UBJSON/BJData readers: top aliased
container_stack.back() and was read (top.is_object) right after
container_stack.pop_back() ended its lifetime. A copy stays valid
regardless of what happens to the stack; nothing here mutates the live
entry, so no field needs to go through container_stack.back() directly.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:45 +02:00
Niels Lohmann a14619b354 Read UBJSON and BJData containers without recursing per nesting level (#5507)
* Read UBJSON and BJData containers without recursing per nesting level

get_ubjson_array() and get_ubjson_object() read their elements by calling
back into the value reader, which called them again for a nested container,
so the native call stack grew with the nesting depth of the input. '[' alone
opens a container, so half a million of them crashes the process before the
input runs out (#5104). The optimized forms reach the same path through a
size or type annotation, and in plain UBJSON '[' and '{' are permitted as the
type of an optimized container, so "[$[#i\x01" repeated nests just as deeply
at six bytes a level.

Both readers now only open their container, and parse_ubjson_internal() loops:
it closes the containers that have ended, claims the next element of the
innermost one, reads its key when it is an object, and works out the marker
of the value to read next. That last part is where the formats differ, and
the loop follows what the four element loops used to do:

  - a sized, typed container gives its elements no marker of their own
  - a sized, untyped container reads one for each element
  - a container that ends at a marker has the byte already, from the test
    against ']' or '}'; for an object it is the first byte of the key

The ND-array wrapper and the 'B' binary shortcut stay as they are. Both read
a complete value rather than opening a container, and their elements are
always scalars: BJData does not permit '[' or '{' as an optimized type, which
is also why only plain UBJSON needed the type-marker case above.

A container of no-ops keeps its behaviour of holding no elements while still
announcing its declared size to the SAX parser, by opening it and then
setting its count to zero.

unit-ubjson and unit-bjdata pass unchanged, 1.39 million assertions between
them, and a behaviour comparison against the previous commit over every
container form -- sized, unsized, typed, untyped, empty, no-op, ND-array,
binary, and the forms nested inside one another -- gives identical values,
error codes, messages and byte offsets. 500,000 levels of each vector now
report a parse error instead of crashing, and a well-formed 100,000-level
value is read to completion.

The driver costs about 3 % on parsing 60,000 small objects and one array of a
million integers, for the reason given in the previous commit; reading the
frame once per element rather than per branch halved what it cost before.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the UBJSON/BJData advance loop's top a copy, not a reference

Same issue as the CBOR reader: top aliased container_stack.back() and
was read (top.is_object) right after container_stack.pop_back() ended
its lifetime. A copy stays valid regardless of what happens to the
stack; the one place that mutates the live entry (--top.remaining) now
goes through container_stack.back() directly.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:44 +02:00
Niels Lohmann 91ab3e81f5 Read CBOR containers and tags without recursing per nesting level (#5506)
* Read CBOR containers and tags without recursing per nesting level

get_cbor_array() and get_cbor_object() read their elements by calling back
into the value reader, which called them again for a nested container, and a
tag was handled by reading the tagged value the same way. All three cost
native stack, and all three cost a single byte to encode: 0x9F opens an
indefinite-length array, 0x81 a one-element array, and 0xC2 is a tag. Half a
million of any of them crashes the process before the input runs out (#5104).

Apply the shape the MessagePack reader already uses: the open containers live
on the heap stack, parse_cbor_value() reads a single value and only opens a
container rather than reading it to its end, and parse_cbor_internal() loops,
resuming the innermost container after each element.

Two things are specific to CBOR. An indefinite-length container ends at a
break marker rather than at a count, and testing for that marker consumes a
byte which is the first byte of the next element when it is not one; the
frame's count is npos for those, and the driver tracks whether the next value
starts at a fresh byte. And a tag is not a value of its own: instead of
reading the tagged value by recursing, the value reader reports that a tag was
read and the driver reads on, so a chain of tags costs no stack at all.

The switch that decodes a value is unchanged apart from the twelve container
cases and the two tag sites. Verified against the previous commit over
definite and indefinite arrays and maps, all four counted forms, empty
containers, nesting of the forms inside each other, truncated inputs, and all
three tag handlers: identical values, error codes, messages and byte offsets.
500,000 levels of each of the three vectors now report parse_error.110 instead
of crashing, and a well-formed 200,000-level value is read to completion.

On performance: the driver does per element what a counted loop used to do
per container, and CBOR pays for it more than MessagePack because the value
reader also has to be told whether to fetch a byte. Parsing 60,000 small
objects and one array of a million integers is 3 to 4 % slower than the
recursive reader, measured over five alternating runs. Against develop the
same two inputs are about 44 % faster, because the entry point no longer
copies the value it parsed; the earlier commit in this series is what pays
for that.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make parse_cbor_internal's top a copy so it survives pop_back()

top aliased container_stack.back(), and was still read (top.is_object)
right after container_stack.pop_back() destroyed the element it aliased.
Nothing currently reorders those two lines, but the comment claiming the
reference's lifetime was already fine only accounted for reallocation
from a push, not this. A trivially-copyable container_frame makes top a
copy instead, so reads of it stay valid regardless of what happens to the
stack; the one place that mutates the live entry now does so through
container_stack.back() directly rather than through top.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:44 +02:00
Niels Lohmann 0dd8ca9023 Read MessagePack containers without recursing per nesting level (#5505)
get_msgpack_array() and get_msgpack_object() read their elements by calling
back into parse_msgpack_internal(), which calls them again for a nested
container. The native call stack therefore grew with the nesting depth of the
input, and each level costs only one byte to encode: 0x91 is a one-element
array, so a few hundred thousand of them crash the process before any of the
input is rejected (#5104).

Keep the open containers on a heap stack instead, the way
parser::sax_parse_internal() has always done for JSON text. A frame records
how many elements are left and whether to close with end_object() or
end_array(); parse_msgpack_value() reads a single value and, for a container,
only opens it; and parse_msgpack_internal() loops, resuming the innermost
container after each element and closing it when its count runs out. Whether
the value that was begun is complete is answered by the stack being empty, so
no separate bookkeeping is needed.

The switch that decodes a value is untouched apart from the six container
cases, which now call enter_container() rather than a reader that loops. That
keeps this diff to the control flow and leaves the decoding of every other
type byte-identical.

enter_container() is the only place a binary reader emits start_object() or
start_array(), so a check that rejects a container can be added there once and
is guaranteed to run before the start event. The frame type and the stack are
shared, ready for the other three formats.

Verified against develop over empty, nested, counted (array 16/32, map 16/32)
and truncated inputs: identical values, error codes, messages and byte
offsets. 300,000 levels now report parse_error.110 instead of crashing, and a
well-formed 300,000-level value is read to completion through the SAX
interface, where develop crashes.

Reading such a value into a basic_json needs the return-by-move change as
well, without which the recursive copy constructor overflows on the way out;
that is the parent commit, and the test for the value path covers the two
together. Timing is unchanged: parsing 60,000 small objects and one array of
a million integers is within run-to-run noise of develop either way.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:43 +02:00
Niels Lohmann d9c55eb225 Split unit-regression2.cpp so the MinGW linker can relocate it (#5511)
Linking test-regression2 with clang and MinGW fails with

    relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'

once the translation unit grows past a certain size: the code can no longer
reach the read-only data it references within the range of a 32-bit
relocation. The file is one of the largest in the test suite and had been
sitting just under that limit, so an unrelated change elsewhere in the
library is enough to tip it over. It is already the second such file --
unit-regression1.cpp was split for size before -- and windows.yml already
carries a workaround for the same limit hitting the debug sections of this
same target, where -g0 was enough because that relocation was against
`.debug_line'. This one is against `.rdata', which no compiler flag avoids.

Move the second half of the regression tests, and the helper types only they
use, into unit-regression3.cpp. The sections are independent -- every
statement in "regression tests 2" was already inside a SECTION -- so they
move unchanged, and the counts confirm nothing was lost: 168 assertions
before the split, 50 plus 118 after.

The result is that both files are comfortably smaller than the one that used
to link, measured with clang at -O1 for C++20:

                        read-only data        text     object
    before                      58,233   1,287,764  3,158,120
    unit-regression2.cpp        48,161   1,012,988  2,522,296
    unit-regression3.cpp        41,710     772,704  1,878,880

No CMake change is needed: tests/CMakeLists.txt globs src/unit-*.cpp, so the
new file is picked up and built for every standard like its siblings.

CONTRIBUTING.md pointed contributors at unit-regression2.cpp for new bug
tests; it now points at the smaller file and says why the two exist, so the
split does not quietly undo itself.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:42 +02:00
Niels Lohmann 44f8ec30e9 Bound UBJSON optimized arrays of a valueless type (#5504)
* Bound UBJSON optimized arrays of a valueless type

An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized UBJSON array of one of those has no
payload: reading an element consumes no input at all. Its declared count is
therefore the only thing that decides how much is allocated, and nothing
bounded it. "[$Z#l" and a four-byte count is nine bytes of input describing
two billion values; #2793 reports 35 GB and 150 seconds from ten bytes, and
OSS-Fuzz has an out-of-memory and a timeout report for the same shape.

Every other type costs at least one byte per element, so the end of the input
bounds it. 'N' (no-op) is already skipped rather than stored. Objects are not
affected either: each element is preceded by its key, which costs bytes. And
BJData already refuses these markers as an optimized type, so this is a plain
UBJSON matter.

Reject a count above 1,048,576 elements for those three types with
out_of_range.408, the code this reader already uses for a declared size it
will not honour. The check runs before the SAX start event, so no container
is opened and then abandoned.

Rejecting on the read side alone would break the guarantee that anything
to_ubjson() writes can be read back, and would trip the round-trip assertion
in fuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized
encoding, one byte per element, for arrays of these types above the same
limit. Its decision depends only on the array's size, which is identical for
a value and for anything parsed back from it, so the round trip is stable.

No existing test changes: the largest such count in the test suite is 65,793.
The excessive-size test that already used this shape still passes, now
rejected a little earlier than by the max_size() check it used to reach.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Note the 1,048,576 valueless-array limit as (1 << 20) in the docs

Addresses review feedback from @gregmarr on PR #5504: spell out the
binary/hex form next to the decimal count so it reads as the round
power-of-two it is, matching how include/nlohmann/detail/input/binary_reader.hpp
defines max_valueless_container_size. Applied in both docs/exceptions.md
and ubjson.md, as requested.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:42 +02:00
Niels Lohmann 4bb2de8cc5 Reject a nested BJData ndarray dimension vector where it is read (#5503)
get_ubjson_size_type() takes an inside_ndarray parameter saying whether it is
being called for an ndarray's dimension vector, where another ndarray is not
allowed. It then seeded the flag it passes down to get_ubjson_size_value()
with `false` rather than with that parameter, and only consulted
inside_ndarray afterwards, on the '$' branch.

So on the '#' branch nothing stopped the descent: every "#[" pair of an input
like "[" followed by "#[#[#[..." opened another dimension vector, several
native stack frames deeper each time, and the recursion was only reported on
the way back out. 100,000 pairs crash the process. This is #5104 again, in a
path that has nothing to do with containers.

Seed the flag with inside_ndarray, which is what get_ubjson_size_value()
documents it wants: "for input, `true` means already inside an ndarray vector
or ndarray dimension is not allowed". The nested '[' is then refused where it
is read, so the length of the chain no longer matters.

Both post-checks gain `&& !inside_ndarray`, because an ndarray was found
*here* only if the flag flipped -- get_ubjson_size_value() only ever returns
`true` when its initial value was `false`, as its documentation says. With
that, the "ndarray can not be recursive" branch is unreachable: a recursive
ndarray is now caught one level earlier, and reported as "ndarray dimensional
vector is not allowed" like every other nested dimension vector.

Three existing expectations move accordingly (vR2, vR4, vR6). All three now
fail earlier, and all three now report the same error that vR1, vR5 and vH
already reported for the same shape, which is the more consistent outcome.
Everything else is unchanged: valid 1D and 2D ndarrays, optimized containers
and plain arrays produce identical results, and unit-ubjson is untouched.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:41 +02:00
Niels Lohmann dd50f0eb16 Stop CBOR indefinite-length strings from recursing per chunk (#5502)
get_cbor_string() and get_cbor_binary() handled the indefinite-length forms
(0x7F and 0x5F) by calling themselves once per chunk. Each chunk therefore
cost a native stack frame, and since a chunk may itself be an indefinite-
length string, an input of repeated 0x7F bytes reached one frame per input
byte: 200,000 of them crash the process with SIGSEGV before a single byte is
rejected. This is the same defect as #5104, in a path the container-level
work does not touch.

Count the open levels instead of recursing through them. That is enough here
because every chunk is appended to the same result -- get_bytes() writes at
result.size() -- so there is no per-level state to keep. The temporary chunk
string and its copy into the result go away with the recursion.

The definite-length cases move to get_cbor_string_chunk() and
get_cbor_binary_chunk() unchanged, including their error messages, which
still name 0x7F and 0x5F because those are handled one level up.

Behaviour is unchanged. Comparing against develop over the interesting byte
sequences -- empty, single-chunk, nested, over-closed and truncated forms,
both strings and byte arrays, and an indefinite-length map key -- produces
identical values, error codes, messages and byte offsets. The 200,000-level
input now reports parse_error.110 at byte 200001 instead of crashing.

Note that nesting these is not valid CBOR: RFC 8949, Section 3.2.3 forbids
it. This does not change that either way -- it has always been accepted, and
rejecting it is a separate decision (#5317, #5325). Should it be rejected
later, that is now one condition on the level counter rather than a change to
the control flow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:34:41 +02:00
Niels Lohmann fae364db24 Move the scanned string into the value instead of copying it (#5458)
* Move the scanned string into the value instead of copying it

The SAX interface documents that the string handed to json_sax::string() may
be moved from, and the DOM handlers already move the one handed to binary().
string() did not, so every string value was copy-constructed out of the
lexer's token buffer, which then kept the buffer alive at its high-water mark
until the next token overwrote it.

Moving hands that buffer to the new value instead. The allocation count is
unchanged - the value needed one either way - but the copy is gone.

  jeopardy      247.3 ms -> 240.9 ms  (-2.6%)
  citm_catalog    4.61 ms ->   4.48 ms (-2.8%)
  40k 30-char strings 7.80 ms -> 7.64 ms (-2.1%)

Note this deliberately does not extend to the object key. Moving the key
hands the lexer's buffer - sized for the largest token seen so far - to a key
that is usually short, so the next value has to grow a fresh buffer. Measured,
that costs 11.9% on a document of many small keys with longer values.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Spell out the move rationale at every handle_value(std::move) site

The comment explaining why the value is moved sat only on
json_sax_dom_parser::string(), and the callback parser's string() pointed
at it with "see json_sax_dom_parser::string()". That reference cannot be
searched for - the function is declared as `bool string(string_t& val)`
inside the class, so the qualified name appears nowhere - and the two
binary() overloads, which have always moved, carried no explanation at all.

Put the same comment on all four sites and name json_sax, which is
greppable, instead of a member that is not.

Comment-only; no generated code changes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:22:51 +02:00
Niels Lohmann 1dc1d09fc6 Do not search a container for the value the callback rejected (#5457)
When a parser callback rejects a value, the placeholder stored for it has to
be removed from its parent again. remove_discarded_value() found it by
scanning the parent from the beginning, so filtering a container cost one
scan per rejected member - quadratic in the number of members of a single
container.

A rejected value can only ever be the one most recently added to its parent:
the last element of an array, or the placeholder key() stored under the
current key in an object. Record that key alongside the existing
key_keep_stack, and for a container record it again alongside ref_stack so
end_object()/end_array() can find it in the parent. Removal is then O(1) for
an array and O(log n) for an object, and finding nothing there means nothing
was stored, so there is nothing to remove.

The key for a container is read before handle_value() may consume it, so it
is also correct when the callback rejects the container at its start event
and it never reaches its parent at all.

Discarding half the members of one object, before -> after:

     members     value rejected    container rejected at start
      16 000    392 ms -> 3.7 ms      803 ms ->  8.2 ms
      64 000   6238 ms -> 14.6 ms   12651 ms -> 32.2 ms
     128 000  25339 ms -> 30.6 ms

Results are unchanged: 48 000 randomized documents parsed under 12 different
filtering callbacks - covering duplicate keys, empty keys, rejected keys and
containers rejected at both their start and end events - produce byte
identical output before and after.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:22:50 +02:00
Niels Lohmann 471584616f Benchmark parsing of pretty-printed JSON (#5456)
Every input in the benchmark corpus is minified or only lightly spaced, so
none of them exercise the lexer's whitespace handling. Real-world JSON is
frequently indented - configuration files, pretty-printed API responses,
anything kept under version control - where the insignificant whitespace can
outweigh the data itself.

Add a ParseIndented family that re-serializes each document with an
indentation and parses that. The content is identical to the matching
ParseString row, so the pair isolates the cost of the whitespace alone.
Measured here, parsing the indented form costs 16-34% more than the minified
form of the same document, which nothing in the suite currently reports.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:22:50 +02:00
Niels LohmannandClaude da42a627dc Stop dump() from heap-allocating its output adapter per call (#5449)
* Stop dump() from heap-allocating its output adapter per call

The serializer held its output sink as output_adapter_t<char>
(a std::shared_ptr<output_adapter_protocol<char>>), which dump() and
operator<< built via make_shared -- one heap allocation per call for a
sink that only wraps a reference to the caller's string or stream.

Hold the sink as a non-owning output_adapter_protocol<char>* instead and
construct the concrete adapter on the stack at the call site. The write
path (o->write_characters) is unchanged, so output is byte-for-byte
identical; a compact dump() of a small object drops from 2 heap
allocations to 1 (only the returned string remains), ~3% faster.

Completes the per-call allocation cleanup on this branch, which already
removed the indent_string buffer (both were reported in #5413).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L1oJ2ggRHS37zeVe94QTA1
Signed-off-by: Claude <noreply@anthropic.com>

* Take the output adapter by reference at the serializer ctor

Per review: the serializer still holds the adapter as a non-owning
pointer, but the constructor now takes output_adapter_protocol<char>&
and takes its address internally, so every call site passes a
reference. A reference cannot be null and reads as a borrow, which
makes the lifetime contract harder to get wrong than handing over a
raw pointer. The stored member and the write path are unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Claude <noreply@anthropic.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude <noreply@anthropic.com>
2026-09-11 08:22:50 +02:00
Niels LohmannandClaude Opus 4.8 ff80ed3295 Speed up dump(), and keep it from overflowing the stack (#5285)
* Add SWAR bulk fast path to string serialization (dump_escaped)

When ensure_ascii is false, dump_escaped previously ran every byte of
every string and object key through the UTF-8 DFA decoder, even for the
common case of ordinary text with nothing to escape. This mirrors the
per-byte cost the parser had before the contiguous fast paths.

At a character boundary, bulk-copy the longest run of bytes that need no
escaping using string_bulk_run() - the same SWAR scanner and UTF-8 bulk
validator the lexer's contiguous path uses - and only fall back to the
byte-at-a-time DFA loop for the first byte that needs individual handling
(a quote, backslash, control character, or ill-formed/truncated UTF-8).
Because every "hard" or invalid byte is still processed by the unchanged
byte path, escaping output and error handling (including strict-mode
error 316 position and message) are byte-identical to before.

The ensure_ascii=true path is unchanged: it must escape non-ASCII and
0x7F, which string_bulk_run does not stop on, so a separate predicate
would be needed for it.

Verified byte-for-byte identical dump output against the pre-change
implementation across ~20k randomized byte strings plus curated edge
cases (all escapes, control chars, valid multibyte, surrogates,
overlong, truncated sequences) for both ensure_ascii settings and all
three error handlers, in C++11/17/20 at -O2/-O3.

Throughput (g++ -O3, ensure_ascii=false, vs pre-change):
  long ASCII strings   4.2x
  twitter-like objects 2.3x
  dense CJK            1.4x  (further headroom with JSON_USE_SIMDUTF)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Buffer serializer output and add ensure_ascii string fast path

Two further serialization speedups on top of the ensure_ascii=false bulk
copy, both reusing the SWAR primitives in detail/input/string_scan.hpp.

1. Internal write buffer (devirtualization). Every structural character
   ('{', '"', ',', ...) previously went straight to the output adapter
   through a virtual call. Route all writes through put_char/put_chars
   into a 1 KiB buffer that flushes in bulk; the public dump() flushes
   once the top-level value is done (the recursive worker is split out as
   dump_internal). Runs larger than the buffer are written straight
   through, so large payloads are not copied twice. This is the dominant
   cost for object/array-heavy values.

2. ensure_ascii fast path. dump_escaped previously ran the UTF-8 DFA over
   every byte when escaping non-ASCII. Add find_ascii_copyable_run() (a
   SWAR scan stopping at '"', '\\', < 0x20, 0x7F, and >= 0x80) so runs of
   printable ASCII are bulk-copied, with the byte path handling each
   escape/non-ASCII byte exactly as before.

Behavior is unchanged: dump output is byte-for-byte identical to the
previous implementation across ~20k randomized byte strings plus curated
edge cases (all escapes, control chars, 0x7F, valid multibyte,
surrogates, overlong, truncated), for object/array/pretty output, both
ensure_ascii settings, and all three error handlers, in C++11/17/20 at
-O2/-O3. New unit tests cover the buffer flush boundaries, the escape and
0x7F handling, multibyte under both settings, and invalid-UTF-8 handling.

Throughput (g++ -O3, vs the ensure_ascii=false-only baseline):
  long ASCII, ensure_ascii=0   4.2x
  long ASCII, ensure_ascii=1   4.1x
  twitter-like objects         2.7x
  dense CJK                    1.8x

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Flush serializer buffer in dump_escaped unit test

test-convenience failed (macOS finished first; the failure is
platform-independent) because check_escaped() calls the internal
serializer::dump_escaped() directly and then reads the output stream.
Since dump_escaped() now writes into the serializer's internal write
buffer, the bytes were still buffered and the stream was empty.

Expose flush() under JSON_PRIVATE_UNLESS_TESTED (same visibility as
dump_escaped) and flush in check_escaped() before inspecting the output.
Per-string flushing inside dump_escaped() was rejected on purpose: it
would defeat the buffering that makes object/array-heavy dumps faster.
Library behavior is unchanged (flush()'s body is identical; only its
access label moved).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid deep recursion in serialization write-buffer test

The "many small structural writes exceed the write buffer" subcase built
a 1100-deep nested array and dumped it to force >1024 consecutive
single-character writes through put_char (exercising the write buffer's
flush-when-full branch). dump() recurses per nesting level, so on MSVC
debug builds (smaller default stack, larger frames) this overflowed the
stack and crashed test-serialization; Linux/macOS have enough headroom to
hide it.

Replace the nesting with a flat array of 500 empty strings. Each element
emits '"', '"', ',' via put_char, so the dump is a long run of
single-character writes (1501 bytes > the 1024-byte buffer) at nesting
depth two, hitting the same flush branch without deep recursion. Library
code is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Split the write-buffer helpers and write indentation directly

Follow-up to @gregmarr's review: put_chars() was doing four unrelated jobs, so
give the two that can be made safe their own entry points.

- put_literal(): takes the literal by reference and deduces the length from the
  array bound, so the 27 hand-counted lengths at the call sites can no longer
  drift from the literals they describe. A literal is checked at compile time to
  fit the buffer, so this path needs no write-through branch.

- put_buffer(): takes the fixed-size buffer itself rather than a bare pointer,
  so the length can be checked against the buffer's own bound.

- put_indent(): memsets the indentation into the write buffer, filling and
  flushing it as needed. This removes indent_string entirely, and with it both
  bugs of #5186: the indentation string was grown by doubling, which is not
  enough when indent_step more than doubles it (a heap over-read - dump(2000)
  read 2000 bytes out of a 1024-byte string), and the grown part was filled with
  a space instead of the configured indent_char. next_indent() keeps that PR's
  assertion against the unsigned indentation accumulation wrapping on deep
  nesting.

put_chars() keeps the two cases that are genuinely a pointer and a count: the
run-length copies out of the string being escaped, and to_chars() output.

Tests cover an indent_step wider than the write buffer, a non-space indentation
character past the old growth point, and nesting whose accumulated indentation
spans several buffer-fulls. All three fail against develop.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fill the indentation buffer once instead of once per flush

@gregmarr's point on the fill-and-flush loop: flushing does not disturb what
the write buffer holds, so an indentation spanning several buffer-fulls only
has to be written into the buffer once and can then be handed to the adapter
as many times as needed. The loop re-filled it every time, doing work it
already knew was there.

put_indent() now fills the room left in the buffer, and if anything remains,
flushes, fills the buffer once, and re-flushes that same content. It also
returns early for a zero-width indentation, which is what the closing brace of
every outermost value asks for.

Measured over a dump(), counting memset calls and bytes inside put_indent:

    indent       before              after
         4       1 call /     4 B    1 call /     4 B
      2000       2 calls /  2000 B   2 calls /  2046 B
    100000      98 calls / 100000 B  2 calls /  2046 B

The wide case is now constant work rather than proportional to the indentation
width; ordinary widths are unchanged. Tests extended to cover several whole
buffer-fulls and an exact multiple of the buffer size.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Tighten the write-buffer helpers after review

More of @gregmarr's review on the put_* split:

- Reattach the put_chars() doc comment, which the new helpers had been
  inserted in front of, leaving it describing put_indent().

- Compute the literal length once in put_literal() instead of spelling N - 1
  at each use.

- Add put_string(str, start, end), which keeps the pointer arithmetic and the
  bounds assertions inside the function instead of at the call site. With
  dump_float()'s to_chars() output moved onto put_buffer() as well, put_chars()
  now has no callers outside put_string()/put_buffer(): nothing passes a bare
  pointer and a count any more.

- Carry the indentation as std::size_t rather than unsigned int. It is a size,
  it is compared and combined with buffer sizes throughout, and the casts in
  put_indent() disappear. next_indent() keeps its assertion, which is far
  harder to trip on a 64-bit size_t but still reachable where that is 32 bits.

No output change: pretty and compact dumps, binary values included, are
byte-identical to develop.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Silence avoid-c-arrays on put_literal's array reference

clang-tidy flags the reference-to-array parameter under
cppcoreguidelines/hicpp/modernize-avoid-c-arrays, and the CI treats warnings as
errors. Binding to the array is the whole point here - it is what lets the
length be deduced from the literal instead of hand-written at the call site - so
suppress it the same way from_json(), to_json() and get_to() already suppress it
for their own T (&arr)[N] parameters.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Bound the descent of dump()

Serializing a container serializes its elements, so dump() descended into one
call per nesting level. A value nested deeply enough exhausted the call stack
and terminated the process with a segmentation fault - no exception, nothing
the caller could catch. Parsing such a value works, as the parser is
iterative, and so does destroying one, as #1436 made destruction iterative.

Bound how far the descent goes rather than take the call stack away from it.
The first 128 levels are written by exactly the code that always wrote them,
and only below that does dump_iteratively write out what is left, keeping the
containers it has entered on an explicit stack. Serializing can therefore no
longer exhaust the stack, however deeply a value is nested, while a value
nested less deeply than the bound pays only for one comparison per container.

Writing every value that way instead measured between 2% and 20% slower - 20%
on object-heavy documents - which is why the descent is kept for all but the
values that cannot afford it. The bound costs nothing measurable: between
-1.4% and +1.2% across compact and pretty output of number, integer, string,
object-heavy, wide-object and deeply nested documents.

The output is unchanged for every value. Both ways of writing a container
emit the separator in front of every element but the first, rather than
after every element but the last, which puts exactly one between each pair
and none at the end.

This fixes #5387 for dump(). The copy constructor is fixed in #5389.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fold ensure_ascii into the escaper and write bytes without dump_integer

Two hot spots that the write buffer and the bulk scanner left behind.

dump_escaped took ensure_ascii as a runtime flag and tested it inside the
loop, once per character run, although it cannot change while a string is
written. It is now a template parameter, dispatched once per string, which
folds the choice of scanner and lets each of the two be inlined into a loop
of its own. This is the hottest loop in the serializer: it runs over every
string and every object key.

A binary value's bytes went through dump_integer, which counts digits and
does 64-bit arithmetic for a number that is always in [0, 255]. dump_byte
writes the three digits it takes at most straight into the write buffer
instead. Any byte type that is not a plain unsigned byte is still left to
dump_integer, whose representation of it may differ.

Measured against the previous commit (medians of 9 interleaved runs, clang
-O3): binary values -33.8%, dense CJK with ensure_ascii -20.6%, key-heavy
objects -17.8%, deeply nested pretty output -17.9%, dense CJK without
ensure_ascii -11.8%, object-heavy documents -9.3% compact and -9.5% pretty,
a small value dumped in a loop -21.4%, wide objects -2.3%. Arrays of plain
ASCII strings measured 3.5% to 4.2% slower, the one shape that loses; number
and integer arrays are unchanged.

Also tried and dropped: leaving the write and string buffers uninitialized
rather than zeroing 1.5 KB per dump() call. It is worth -30% on small values,
but two nearly identical string workloads moved 18% apart in opposite
directions, so the measurements did not support it.

The output is unchanged for every value: the differential now also covers
every one of the 256 byte values, alone and together, in both binary layouts.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Write a byte without walking a pointer over the buffer

clang-tidy's misc-const-correctness reads the pointer dump_byte advanced over
the write buffer as one whose pointee could be const. Index the buffer
instead, which says the same thing without a raw pointer at all.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Parenthesize the reserve arithmetic in the deep-nesting test

clang-tidy's readability-math-missing-parentheses wants the multiplication
spelled out in reserve(6 * depth + 1), and CI treats its warnings as errors.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not scan for a copyable run that cannot exist

Under ensure_ascii, dump_escaped() calls find_ascii_copyable_run() at every
character boundary. When the text is dense non-ASCII - CJK, where every byte
is >= 0x80 - the scanner stops on its first byte and returns zero, so its SWAR
block runs once per character and buys nothing, on top of the escaping that
still has to happen afterwards.

A run can only be non-empty when the first byte is one the scanner may copy,
so test that single byte before calling it. Runs that do exist are found
exactly as before, so the bulk-copy win is unchanged; only the calls that were
always going to return zero are skipped.

Output is unchanged: the dump digest over canada/citm/twitter, in compact,
pretty and ensure_ascii form, matches develop byte for byte.

  dump(ensure_ascii=true)   develop    before     after
  CJK text                   3.54ms    4.25ms    3.36ms
  CJK, no ASCII at all       3.09ms    4.02ms    3.02ms
  Latin-1-ish text           4.39ms    3.04ms    2.93ms
  plain ASCII                3.92ms    0.80ms    0.79ms

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review of the write-buffer helpers

Three points from @gregmarr's review:

put_chars() is gone. It was the only entry point taking a bare pointer and
a count, and it existed only so put_string() and put_buffer() had something
to delegate to. Its body now lives in put_string(), and put_buffer() is
put_string(buffer, 0, length) - std::array already carries data() and
size(), so it satisfies the same interface a string does. Nothing appends
characters without a bound any more.

dump_escaped()'s documentation block was duplicated. The dispatcher was
inserted between the original comment and the function it described, and
the comment was copied rather than split. The worker now has its own short
comment saying why ensure_ascii is a template parameter.

The local in dump_byte() is deliberate, and is now documented as such:
writing through write_buffer[] is a char write, which may alias any object,
so with write_buffer_pos updated in place the compiler must reload and
store it around every digit. Measured on a dump of a 4 MiB binary value,
18.0 ms without the local against 7.4 ms with it.

Output is unchanged: byte-identical dumps across 77 files in compact,
pretty, ensure_ascii, pretty+ascii, indent 600 and tab-indent form.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review: drop unneeded backslash-escapes and duplicate scan loop

'"' does not need escaping in a char literal, unlike in a string literal.
find_ascii_copyable_run() also duplicated the byte-at-a-time search that
already exists as the loop's own scalar tail; break into it instead of
re-deriving the offset in a second, near-identical loop.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move pretty_print, ensure_ascii and indent_step into the serializer

None of these change over the life of a serializer, unlike current_indent
and depth, which do change on every recursive call. They are now captured
once in the constructor - matching indent_char and error_handler - instead
of being threaded through dump(), dump_internal(), dump_iteratively(),
dump_value() and dump_escaped() on every call.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the serializer from holding onto std::localeconv()'s pointer

loc was only ever read twice, immediately, to seed thousands_sep and
decimal_point; nothing else in the class used it. A local in the
constructor body serves the same purpose without keeping the pointer
around for the serializer's lifetime.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Keep thousands_sep/decimal_point const via a small locale_chars struct

const members can't be assigned in a constructor body, so seeding them
from std::localeconv() meant either dropping const or holding onto the
lconv* for longer than needed. A sub-object computes both from the
pointer in its own constructor and is itself initialized in serializer's
mem-initializer-list, so the two chars stay const, std::localeconv() is
still called exactly once, and nothing outlives the constructor.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-11 08:22:49 +02:00
Niels Lohmann 35cdd5f408 De-duplicate the diagnostic-positions test files via define-based recompilation (#5474)
tests/src/unit-class_parser_diagnostic_positions.cpp and
tests/src/unit-diagnostic-positions-only.cpp were maintained as near-copies
of unit-class_parser.cpp and unit-diagnostic-positions.cpp respectively, and
had drifted: trailing-comma handling, the #5342 filter-array/filter-value
sections, and the cross-input-adapter diagnostics test were never ported to
the positions-enabled copy.

Fold the position-specific assertions into the base files, guarded by
file a second time with the relevant macro set via CMake COMPILE_DEFINITIONS
(mirroring the existing test-comparison_legacy pattern) instead of
maintaining a separate source file. This removes the duplication and, as a
side effect, closes the coverage gaps above since the full test file now
compiles under JSON_DIAGNOSTIC_POSITIONS=1 as well.

Fixes #5417

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 08:18:43 +02:00
Niels Lohmann bfb07786cd Fix to_bjdata() silently truncating out-of-range _ArrayData_ elements (#5473)
write_bjdata_ndarray() validated that each _ArrayData_ element matched
the number kind (integer vs. float) named by _ArrayType_, but not its
range. An element that did not fit the target C++ type (e.g. 256 for
"uint8") was silently wrapped by the static_cast used to write it, or,
for "single", silently overflowed to infinity.

Range-check each element against the type named by _ArrayType_ before
writing it, reusing the existing fallback path that already encodes
the annotated object as a plain object for other invalid-annotation
cases in this function.

Fixes #5403.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-10 17:59:28 +02:00
Niels LohmannandClaude Opus 4.8 4a93aa4e2f Speed up parsing of contiguous input (numbers, strings, UTF-8) (#5283)
* Speed up number parsing in the lexer (fast paths from the fast_float/simdjson world)

The number scanner converted its already-validated digit buffer with
std::strtoull/std::strtoll/std::strtod. Those pull in locale and errno
machinery and dominate number-heavy parsing (strtod runs at ~6 M/s).

Replace them with dedicated parsers over the validated buffer:

- parse_integer_unsigned / parse_integer_signed: accumulate digits with
  overflow detection, falling back to the float path on overflow exactly
  as the strtoull/strtoll round-trip check did. Overflow behavior is
  unchanged for narrower or wider custom number types.

- parse_float_fast: Clinger's exact fast path for `double` (<=19 significant
  digits, |exp10| <= 22, significand < 2^53), where significand * 10^exp is
  exact under IEEE round-to-nearest. This is the same fast path used by
  fast_float/simdjson. It is bit-identical to strtod on this subset and
  declines (falling back to strtod) otherwise. Only `double` uses it; float
  and long double keep std::strtof/std::strtold via a templated overload.

Measured on representative data (g++ 13, -O3):
  - integers:  DOM parse +11%, SAX +25-34%
  - floats:    DOM parse +37%, SAX +70%  (clang: float DOM ~1.9x)

No dependencies added; header-only and C++11-clean. Existing parser,
lexer, conversion and deserialization unit tests pass unchanged; a
3M-value random-double fuzz matches strtod bit-for-bit.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add SWAR bulk string scanning for contiguous input (simdjson-style)

scan_string() read the input one character at a time through the input
adapter and classified every byte with a large switch. For contiguous
byte buffers we can instead scan 8 bytes at a time with a SWAR word test
that finds the first byte needing individual handling (the closing quote,
an escape, a control character, or a non-ASCII UTF-8 byte) and bulk-append
the ordinary run in one go.

- input adapters expose supports_bulk_scan / bulk_data / bulk_remaining /
  bulk_skip for provably-contiguous, same-type, 1-byte iterator ranges
  (raw pointers in every standard; std::string/std::vector/std::array and
  friends additionally in C++20 via std::contiguous_iterator).
- the lexer gains a bulk_scan capability (gated on lazy_token_string so
  bypassing the per-character capture cannot lose error diagnostics) and a
  scan_string_bulk() fast path; streaming/wide/user adapters are unchanged
  and keep the byte-at-a-time scanner.

The run contains no newline (all bytes < 0x20 are treated as special), so
position bookkeeping stays exact, and error tokens are still reconstructed
lazily from the consumed byte range. The SWAR special-byte test is pure
uint64_t arithmetic - no intrinsics, no runtime dispatch, C++11-clean.

Measured on representative data, pointer input, g++ 13 -O3
(string values discarded by accept() see the largest gains):

  long ASCII strings:  DOM +4.5x,  SAX +14x,   accept +17x  (to ~2 GB/s)
  short strings:       DOM +15%,   SAX +62%,   accept +85%
  escape-heavy:        DOM +31%,   SAX +26%,   accept +28%

Same-input parity verified: 200k randomized documents (escapes, multibyte
UTF-8, surrogate pairs) accept/parse identically via the contiguous SWAR
path and the streaming byte path; unit lexer/parser/diagnostic-position/
deserialization/conversions suites pass unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Validate UTF-8 in the bulk string scanner (portable ~2-3x on non-ASCII text)

The SWAR bulk string path stopped at the first non-ASCII byte and handed
every multibyte character to the byte-at-a-time scanner, whose per-byte
get()/next_byte_in_range()/add() machinery runs at roughly half the speed
of validating straight from the buffer. As a result, dense non-ASCII text
(CJK, emoji, accented Latin) parsed ~10-15x slower than ASCII.

Fold well-formed UTF-8 into the bulk run: scan_string_bulk() now, on a
non-ASCII lead byte, validates one sequence with validate_one_utf8() -
which mirrors scan_string()'s per-byte switch ranges exactly (rejecting
overlong forms, surrogates, and out-of-range code points) - and appends it
in place, continuing until the closing quote, an escape, a control byte,
or an ill-formed sequence. All error handling still defers to the byte
path, so error messages and positions are byte-for-byte unchanged.

Because only well-formed content is fast-pathed and every rejection falls
through to the existing scanner, behavior is identical; the win is purely
throughput. Measured on pointer input (accept, string values discarded):

  content        g++ 13         clang 18
  dense CJK      277 -> 648     ~605  MB/s   (~2.3x)
  dense emoji    299 -> 857     ~702  MB/s   (~2.6-2.9x)
  mixed 90% ASCII 246 -> 331    ~334  MB/s   (~1.35x)
  pure ASCII     unchanged (~3.2 / 4.1 GB/s)

Verified: 2,000,000 randomized documents built from arbitrary bytes
(overlong, surrogate, truncated, out-of-range sequences) accept/reject and
parse identically via the contiguous path and the streaming byte path;
lexer/parser/diagnostic-position/deserialization/conversions suites pass
unchanged. Pure C++11, no intrinsics.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Route contiguous byte containers through the pointer adapter (fast paths in C++11)

json::parse(std::string) - the most common entry point - did not benefit
from the contiguous fast paths (bulk string scanning, UTF-8 bulk
validation, memcpy for binary formats) in C++11..17: std::string::iterator
is a library wrapper, not a raw pointer, and pre-C++20 there is no portable
way to prove it contiguous, so supports_bulk_scan was false. Only raw
pointers, string literals, and C-arrays (and, in C++20, anything modelling
std::contiguous_iterator) took the fast path.

Detect contiguous single-byte containers (std::string, std::vector<char>,
std::vector<std::uint8_t>, std::string_view, ...) via is_contiguous_byte_
container and route them through an iterator_input_adapter built from
data()/data()+size(). The generic iterator-based container overload is
constrained to exclude these, so the two overloads are disjoint and there
is no ambiguity (a plain competing overload loses to the greedy
forwarding-reference container overload on reference binding, and a factory
partial-specialization is ambiguous - both were tried and rejected).

The pointer keeps the container's own element type, so char_type - and
therefore all parsing behavior - is byte-for-byte identical to the iterator
path (const char* for std::string, const std::uint8_t* for
std::vector<std::uint8_t>); only the raw pointer additionally turns on the
fast paths. Lifetimes are unchanged: the container outlives the adapter for
the full parse expression, exactly as the iterators it replaces did.

Measured, C++11, json::parse/accept(std::string), g++ 13:

  long ASCII strings:  accept 201 -> 3200 MB/s  (~16x),  parse 174 -> 1444
  dense CJK:           accept 263 ->  697 MB/s  (~2.6x)
  short strings:       accept 163 ->  243 MB/s  (~1.5x)

Verified: char_type preserved for std::string (char) and
std::vector<std::uint8_t> (uint8_t); CBOR/MsgPack round-trips from
std::vector<std::uint8_t> unchanged; 1,000,000 randomized documents accept
and parse identically via std::string and via std::istream;
deserialization/user-defined-input/parser/lexer/conversions/diagnostic-
position suites pass (20,480 assertions); warning-clean on g++ and clang in
C++11/17/20.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add optional simdutf backend for bulk UTF-8 validation (JSON_USE_SIMDUTF)

The bulk string scanner validates UTF-8 straight from a contiguous buffer.
The scalar validator caps at ~0.3-0.7 GB/s on non-ASCII text; a SIMD
validator reaches several GB/s. Rather than hand-rolling SIMD UTF-8
validation (easy to get subtly wrong - a from-scratch SSE attempt rejected
valid CJK), wire in the vetted simdutf library behind an opt-in switch.

simdutf is not header-only (it ships simdutf.cpp and uses runtime CPU
dispatch), so it is not vendored: defining JSON_USE_SIMDUTF includes
<simdutf.h> and routes the bulk validator through simdutf::validate_utf8;
the project supplies and links simdutf. Undefined (the default), nothing
external is included and the portable C++11 scalar path is used, so the
library stays header-only and its baseline behavior is unchanged.

Design keeps behavior identical either way:
- scan_string_bulk() now finds the run up to the next quote/escape/control
  byte (non-ASCII allowed) and validates it in one shot; on the rare
  validation failure it recomputes the exact valid prefix with the scalar
  helper, so ill-formed input still falls through to the byte path and is
  reported at the same position with the same message.
- the per-sequence scalar path is factored into scalar_string_bulk_run()
  and is the default backend; the refactor is behavior-preserving and does
  not change scalar throughput.

Verified: default and JSON_USE_SIMDUTF builds accept/reject/parse
identically across 2,000,000 arbitrary-byte documents and 1,000,000
mixed-escape/UTF-8 documents (differential fuzz vs the streaming byte
path); lexer/parser/diagnostic-position/deserialization suites pass under
both configurations (20,188 assertions with the backend enabled);
warning-clean on g++ and clang, C++11 and C++20, both configurations.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add a contiguous fast path for scanning numbers

scan_number() reads a number one character at a time through the input
adapter (get()) and appends each byte to token_buffer (add()) before
converting. For contiguous input, the per-character get()/add() overhead
dominates: it is roughly two thirds of the time spent on number-heavy
parsing, far more than the value conversion itself.

Add scan_number_bulk_contiguous(), which parses the whole number token
straight from the input buffer: it validates and classifies the extent
with the same grammar as scan_number()'s state machine, materializes
token_buffer in one copy (substituting the locale decimal point exactly as
scan_number() does), advances the adapter, and reuses the shared
convert_number() tail. On anything it does not recognize as a well-formed
number it makes no state change and returns token_type::uninitialized, so
the caller falls back to scan_number(), which then produces the exact
diagnostic. Errors and their positions are therefore unchanged.

The conversion tail is factored out of scan_number() into convert_number()
so both scanners share it; the fast path is selected by tag dispatch on the
existing bulk_scan capability, so streaming/wide/user adapters are
unaffected.

Measured on pointer input, g++ 13 -O3:
  - integers: parse +65%, accept +98%
  - floats:   parse +39%, accept +70%

Verified: 2,000,000 randomized number documents (including overflow-range
integers, long digit strings and %.17g doubles) parse identically via the
contiguous path and the streaming byte path, matching value, type and
round-trip text; the locale suite and existing parser/lexer/conversions/
deserialization tests pass; a new "lexer number fast path" test checks
contiguous-vs-streaming parity, token classification, and that malformed
numbers are rejected identically. Pure C++11, no intrinsics.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move byte-level scan/parse helpers out of lexer.hpp

lexer.hpp had grown by ~600 lines of byte-level helpers that have no
dependency on the lexer's template parameters and clutter the state
machine. Move them, unchanged, into two focused headers as free functions
in namespace detail:

- number_parse.hpp: parse_integer_unsigned/parse_integer_signed (now
  templated on the number type) and parse_float_fast (Clinger's exact
  double fast path, with the decimal point passed as an argument instead of
  read from a lexer member).
- string_scan.hpp: the SWAR string helpers (is_string_special,
  swar_string_special, find_string_special, validate_one_utf8,
  scalar_string_bulk_run) and the backend-dispatched string_bulk_run,
  including the optional simdutf include and find_string_delimiter.

lexer.hpp now includes these and calls the free functions; the methods that
touch lexer state (scan_string, scan_number, scan_string_bulk,
scan_number_bulk_contiguous, convert_number) stay put. This is a pure code
move with no behavior change: lexer.hpp drops from 2357 to 1934 lines, the
now-unused <cstdint>/<cstring>/<limits> includes are removed, and the
free-function form makes the SWAR helpers reusable elsewhere (e.g. the
serializer's string escaping).

Verified: default and JSON_USE_SIMDUTF builds compile; 2,000,000 number and
2,000,000 arbitrary-byte-string differential-fuzz documents parse
identically to before; lexer/parser/conversions/deserialization/locale/
diagnostic-position suites pass (20,576 assertions); warning-clean on g++
and clang in C++11/17/20; the amalgamation regenerates and passes
check-amalgamation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy findings and document JSON_USE_SIMDUTF in the nav

- number_parse.hpp: use std::array for the powers-of-ten table
  (avoid-c-arrays) and `auto` for the cast-initialized result
  (modernize-use-auto), matching the codebase style (cf. the serializer's
  utf8d table). Indexing casts keep the -Wsign-conversion build clean.
- add JSON_USE_SIMDUTF to the mkdocs navigation so the macro page is
  reachable.

No behavior change; clang-tidy is clean on the new headers and the
amalgamation is regenerated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix number fast path for custom string types without assign()

The contiguous number fast path materialized token_buffer with
token_buffer.assign(data, len), but string_t is only required to provide
the minimal interface the rest of the lexer uses (push_back, append,
clear, operator[], ...). Custom string types such as the test's alt_string
do not implement assign(), so scan_number_bulk_contiguous() failed to
compile for them (unit-alt-string), breaking the gcc/clang standards and
old-compiler CI jobs.

reset() already clears token_buffer, so fill it with append() - which
alt_string and std::string both provide and which the string fast path
already relies on - instead of assign().

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Satisfy clang-tidy: parenthesize math and drop unused forwarding reference

The CI clang-tidy (newer than the locally available version) reported two
additional checks on the new code:

- readability-math-missing-parentheses: parenthesize the (a * b) + c digit
  accumulations in number_parse.hpp.
- cppcoreguidelines-missing-std-forward: the contiguous-byte-container
  input_adapter overload took a forwarding reference but only reads
  data()/size() and never forwards it. It is already disjoint from the
  generic container overload via SFINAE, so a plain const& is correct and
  clearer (and keeps the container alive for the whole parse just as before).

No behavior change; char_type and routing are unchanged (std::string and
std::vector<std::uint8_t> still take the pointer adapter with char/uint8_t
char_type), CBOR/MsgPack round-trips and the 2M number fuzz still pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use std::from_chars (Eisel-Lemire) for float conversion when available

The Clinger fast path is exact only for the "easy" subset (<=19 significant
digits, |exp10| <= 22); high-precision and scientific floats fall through to
strtod, where the failed Clinger attempt actually makes parsing a net loss.
std::from_chars implements the Eisel-Lemire algorithm in modern standard
libraries: locale-independent, correctly rounded, and fast over the whole
value range.

convert_number() now tries parse_float_from_chars() first (guarded by
__cpp_lib_to_chars, so C++11 and libc++-without-float-support keep the
Clinger + strtod path unchanged), then Clinger, then strtof. from_chars is
used only when it consumes the entire token; a partial parse means a non-'.'
locale decimal point, and an under-/overflow (result_out_of_range) also
declines - in both cases the existing strtod fallback supplies the exact
value and the well-defined +/-inf/0 the parser expects, side-stepping the
P4168 divergence between implementations. float and long double now get the
fast path too (Clinger was double-only).

Measured, C++17, g++ 13 -O3, json::parse/accept:
  - canada-style floats: ~unchanged (Clinger already covered them)
  - high-precision (17 digits):  parse 2.1x, accept 2.5x
  - scientific (17 digits + exp): parse 3.6x, accept 4.1x

Verified: C++11 (Clinger/strtod) and C++17 (from_chars) parse every value -
including subnormals, boundary values, and 1e9999/1e-9999 over-/underflow -
to bit-identical results; 2M number-fuzz clean; conversions/deserialization/
locale/number-fast-path suites pass in both C++11 and C++17; clang-tidy
clean; warning-clean on g++ and clang in C++11/17/20.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard from_chars use on JSON_HAS_CPP_17, not just __cpp_lib_to_chars

libstdc++ 15 defines __cpp_lib_to_chars even in C++14 mode (via
bits/version.h pulled in by other headers), but <charconv> is only included
under JSON_HAS_CPP_17. That made parse_float_from_chars() reference
std::from_chars without the header in C++14 builds, breaking gcc-latest,
icpx, and the offline-testdata jobs.

Gate the use on JSON_HAS_CPP_17 && __cpp_lib_to_chars so it matches the
include condition exactly; C++11/14 always take the scalar fallback.
Verified by forcing __cpp_lib_to_chars in a C++14 build: the guard
suppresses std::from_chars and it compiles. C++17 behavior is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Disable Clinger float fast path under extended FP precision (x87)

The contiguous number fast path uses a Clinger-style exact algorithm
(significand * 10^scale in double arithmetic), which is only correctly
rounded when double operations are evaluated in true 53-bit precision.
On the x87 FPU used by 32-bit x86 (FLT_EVAL_METHOD == 2) the single
multiply/divide is computed in 80-bit and then double-rounded to double,
so a small fraction of values land 1 ULP off.

This surfaced as test-cbor_cpp11 and test-msgpack_cpp11 failing on the
mingw (x86) job for regression/floats.json: the C++17 builds pass because
they take the correctly-rounded std::from_chars path, while C++11 falls
back to parse_float_fast(). A 5M-sample check over shortest round-trip
decimals reproduces it: 0 divergences with 53-bit doubles, ~1 in 25 000
with 80-bit intermediates; declining to std::strtod fixes all of them.

Guard parse_float_fast() on FLT_EVAL_METHOD so it declines whenever the
platform evaluates doubles in extended precision, letting the caller use
the correctly-rounded std::from_chars / std::strtod path instead. On
mainstream x86-64/ARM64 (FLT_EVAL_METHOD == 0) the fast path is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard <charconv> include with __has_include for GCC 7

GCC 7 sets __cplusplus to the C++17 value under -std=gnu++1z, so
JSON_HAS_CPP_17 is defined, but its libstdc++ ships no <charconv> header
(added in GCC 8; floating-point from_chars in GCC 11). The unconditional
"#if defined(JSON_HAS_CPP_17) #include <charconv>" therefore failed to
compile there: "fatal error: charconv: No such file or directory" in the
ci_test_compilers_gcc (7) job.

Wrap the include in __has_include(<charconv>), mirroring the library's
existing handling of <version> and <filesystem> in macro_scope.hpp. When
the header is absent, __cpp_lib_to_chars stays undefined and
parse_float_from_chars() takes its scalar fallback, so the from_chars use
site (already gated on __cpp_lib_to_chars) is never reached. GCC 8-10,
which have <charconv> but no floating-point from_chars, are unaffected:
they include the header but still take the fallback. GCC 11+ is unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AXcDtEma2PjxgmPS9cQGzA
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Restore the column when ungetting a newline

The contiguous number fast path never reads the character that terminates
a number token, while scan_number() reads it and then ungets it. When that
character is a newline, get() has already cleared chars_read_current_line,
and unget() could only restore lines_read - leaving the column at 0. The
two paths therefore reported different columns for the same document:

    json::parse("[01\n]")      -> line 1, column 3
    json::parse(stringstream)  -> line 1, column 0

Remember the column the newline was read at so unget() can restore it.
Both paths now report the position the offending token actually starts at,
which also fixes the pre-existing column-0 artifact for streaming input.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Extend the bulk scan fast paths to sized sentinels

supports_bulk_scan required IteratorType and SentinelType to be the same
type, which excluded std::counted_iterator paired with std::default_sentinel_t
- the combination #5268 had already enabled for the memcpy fast path. Such
input fell back to the byte-at-a-time scanner even though it is contiguous
and its remaining length is computable in O(1).

Factor the "distance is computable in O(1)" test into sentinel_is_sized and
use it for iterator_is_contiguous, supports_seek, and supports_bulk_scan
alike, and share the std::ranges::distance/std::distance dispatch through a
remaining_count() helper.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document JSON_USE_SIMDUTF on the macro overview page

The macro was only listed in the API macro index; add it to the supported
macros overview alongside the other JSON_USE_* macros, and note that it
selects between two definitions of the same inline function and so must be
defined identically in every translation unit.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate source code

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cover the counted-iterator bulk scan paths

Sized sentinels newly reach the bulk string/number scanners and the
seek-based token reconstruction, so exercise both:

- diagnostics that quote the offending token, which are rebuilt from the
  consumed input via copy_consumed_range()
- inputs whose count ends before the underlying buffer does, including a
  closing quote that exists only behind the count, a cut inside an 8-byte
  SWAR stride, and a cut inside a UTF-8 sequence

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the SentinelType example on iterator_input_adapter

The comment offered "a C++20 sentinel or counted_iterator" as examples of a
SentinelType, but std::counted_iterator is the IteratorType - the sentinel it
pairs with is std::default_sentinel_t. #5268 corrected the same wording in the
API documentation and left the code comment behind.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not discard the parse result in the error position check

json::parse is declared warn_unused_result, and CHECK_THROWS_WITH_AS
evaluates its expression as a discarded statement, so the assertion broke
the -Werror builds (GCC -Werror=unused-result, MSVC C4834 under /WX).
Compare against the helper that already captures the message instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Lock the two number grammars together with a parity test

The JSON number grammar is encoded twice: as the scan_number() state machine
and as the contiguous fast path. The fast path declining on anything it does
not recognize keeps most divergence harmless, but if it ever accepted
something the state machine rejects the result would be a silent correctness
bug, and the existing test only pinned a hand-written list of numbers.

Enumerate every string of length 1..4 over "01.eE+-" (2800 tokens) and
require both paths to agree on the parsed value and on the exact error
message. Verified to fail if the fast path's grammar is perturbed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Skip token_buffer for integers on the contiguous path

An integer token does not need token_buffer: the number_integer and
number_unsigned SAX callbacks take only the value, and the overflow
diagnostic rebuilds the text from the input via get_token_string(). Convert
straight from the input buffer and materialize the token only for the
floating-point tail, which still needs a NUL-terminated buffer for strtod.

JSON_DIAGNOSTIC_POSITIONS derives a number's start position from
get_string().size(), so the copy is kept when that is enabled.

The integer dispatch is factored into convert_integer() and shared with
convert_number(), so both scanners keep using one implementation.

Integer-heavy input, 400k values, -O3:

              parse    accept
  gcc 16      +14%     +23%
  clang       +15%     +22%

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate source code

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Cover the bulk string and UTF-8 scanners

These paths had no dedicated tests and rested on differential fuzzing only.
Add three sections, all comparing the contiguous scanner against the
byte-at-a-time one on the parsed value and on the exact error message:

- every string of length 1..3 over an alphabet of ordinary ASCII, both
  specials, a control byte, escape characters, UTF-8 lead and continuation
  bytes, and a byte that is never valid - each at offset 0 and offset 9, so
  the bulk scanner sees them with and without a run behind them
- every kind of run-ending byte at each offset across two 8-byte SWAR words,
  so multibyte sequences also straddle the word boundary
- the boundaries of every range validate_one_utf8() recognizes: shortest and
  longest encodings, overlongs, both ends of the surrogate block, U+10FFFF
  and just past it, and truncated sequences

Verified to fail if the bulk validator accepts surrogates, and if the SWAR
word test stops detecting control characters.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Format the new string fast path test with astyle

The pinned astyle expands a braced-init-list used as a range-for range onto
several lines; hoist the two offsets into a named vector instead, which reads
better and leaves nothing for astyle to reformat.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix shadowed locals and guard the exception-dependent tests

Two problems in the tests added for the bulk scanners, both found by CI:

- the inner `const json j` in the counted-iterator diagnostics shadowed the
  one declared at test-case scope, which -Wshadow rejects on GCC and clang
  and C4456 rejects on MSVC under /WX; rename them
- the new parity checks parse deliberately invalid input, which calls
  std::abort() rather than throwing when JSON_NOEXCEPTION is defined, so
  they would have crashed the no-exception build; guard them the way the
  other tests do

json::accept() does not abort, so the UTF-8 range assertions stay compiled
without exceptions and keep covering validate_one_utf8() there.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Gate the C++20 iterator classification on JSON_HAS_RANGES

Making supports_bulk_scan depend on iterator_is_contiguous meant the trait
is now instantiated for every adapter, not only when get_elements() is
called. On standard libraries with an incomplete <ranges> that is fatal:
libstdc++ 10 evaluates std::contiguous_iterator<std::counted_iterator<T*>>
by calling std::to_address, which needs an operator-> its counted_iterator
does not have, so satisfaction checking is a hard error rather than false.
Reported by clang 14 + libstdc++ 10.

JSON_HAS_RANGES already encodes exactly this ("libstdc++ < 11 has incomplete
C++20 ranges", #4440), so require it for the C++20 branch. Affected
toolchains fall back to the pointer-only test and the byte-at-a-time
scanner, which parses identically, just without the bulk fast paths.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate source code

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Satisfy clang-tidy in the new bulk scanner tests

- give the helper lambdas an explicit std::string return type and return
  braced initializer lists (modernize-return-braced-init-list)
- replace the C-style array of test cases with a std::vector
  (modernize-avoid-c-arrays)
- silence pro-type-member-init on the two brace-initialized aggregates;
  default member initializers would stop them being aggregates in C++11

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid escaped literals in the counted-iterator diagnostics list

clang-tidy reads "[\"\\ud834\"]" as a literal better written raw, and the
two literals written next to each other in "[\"a\x01""b\"]" as a missing
comma. The concatenation was there to stop the hex escape swallowing the
following character; build those documents from explicit bytes instead and
use raw strings elsewhere. The byte sequences are unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Reattach convert_number's documentation

@gregmarr spotted that convert_integer() was inserted between convert_number()
and its doc block, leaving convert_integer() with two stacked blocks and
convert_number() with none. Comment only; no code change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments on the simdutf backend

string_bulk_run() had the same `return scalar_string_bulk_run(...)` in both
arms of the `#if`. The simdutf arm already falls through when validation
fails, so a single return after the `#endif` says the same thing.

The JSON_USE_SIMDUTF example showed `#include <simdutf.h>`, which
string_scan.hpp already does under the same guard; users only have to put the
header on the include path and link the library, not include it themselves.

Set the version history entry to 3.13.0, matching the other macro pages
documenting unreleased features.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Pin the number error position against a non-newline terminator

The comment claimed the reported column is the one the offending token starts
at. It is the column reached after the token's last character - which is the
actual point of the unget() change: a number terminated by a newline now
reports what the same number terminated by a space always did.

Assert that equality directly, and add a multi-character token where the start
and end columns differ, so the invariant cannot be read off a single-character
example.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not repeat the integer conversion that just failed

scan_number_bulk_contiguous() converts an integer token straight from the input
buffer. When the value does not fit, it materializes token_buffer and calls
convert_number(), which tried the very same integer conversion again before
falling back to floating point.

Recording the outcome in number_type skips the second attempt. The resulting
token type and value are unchanged: convert_number() reached the float tail
either way.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Require a container's value_type to match what data() points at

is_contiguous_byte_container accepted any type with a data() returning a
pointer to a single-byte integral plus a size(). That is duck typing: the two
members say nothing about size() counting the units data() points at.

A type where it does not - fixed-size records, say - was routed to the
pointer-based adapter and parsed as [data(), data() + size()) bytes, silently
truncating input the iterator-based adapter had read in full:

    struct record_buffer {
        using value_type = std::array<char, 4>;
        std::string bytes;
        const char* data() const;                        // raw bytes
        std::size_t size() const;                        // in records
        const char* begin() const; const char* end() const;
    };
    json::parse(record_buffer{"[1,2,3,4,5]"});           // parse error at column 3

Requiring the container's own value_type to be that same element type ties the
two together. Every contiguous standard container satisfies it, so std::string,
std::vector<char>, std::array<char, N> and std::string_view keep the fast path;
anything else falls back to the iterator-based adapter, which is always correct.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Only compile the simdutf backend from C++17 on

simdutf.h rejects anything below C++17 with an #error, so defining
JSON_USE_SIMDUTF in a C++11 or C++14 translation unit did not fail with a
message about simdutf being unavailable - it failed to compile at all, taking
the library's C++11 support with it. Nothing caught this because no build ever
compiled that path.

Gate the include and both uses on JSON_HAS_CPP_17, the same way number_parse.hpp
gates std::from_chars. Below C++17 the macro now has no effect and the scalar
validator runs; it accepts and rejects exactly the same input, so the macro is
safe to set project-wide even when some translation units use an older standard.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Test the simdutf backend in CI

JSON_USE_SIMDUTF was documented and shipped but never compiled by anything in
the repository, so nothing held the backend to the behavior the docs promise.

Add JSON_TestSimdutf (OFF by default), which fetches simdutf and defines
JSON_USE_SIMDUTF for every test target, and a ci_test_simdutf target that runs
the whole suite in that configuration. Because simdutf needs C++17, the suite is
built at C++11 as well, so one job covers both the scalar fallback with the
macro defined and simdutf itself.

The dependency hangs off test_main, whose usage requirements every test target
inherits. The library target and the installed CMake package are deliberately
untouched: making nlohmann_json link simdutf would put a find_dependency() in
the exported package, which is a separate decision.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Warn when JSON_TestSimdutf cannot reach the simdutf backend

simdutf needs C++17: without it the dependency does not even compile, and with
a C++17 compiler but no C++17-or-later standard under test it builds and then
goes unused. Either way the option silently did nothing useful, or broke the
configure step outright.

Resolve the tested standards first, then check them: when none of them can
reach simdutf, skip the dependency and say so, naming which of the two reasons
applies and how to fix it. The tests then run against the scalar validator,
which is what would have happened anyway.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXDi7NTMKAmoArUZSKMc4T
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use record_buffer::data() so clang does not flag it unneeded

The record_buffer test type declares data() and size() so the
is_contiguous_byte_container trait can see both and still reject the
type on its value_type. data() was never called, so clang's
-Wunneeded-member-function (under -Weverything -Werror) failed the
C++20 build. Assert that data() points at the underlying bytes: it
ODR-uses the member and documents the property the type is meant to
demonstrate.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Skip the float fast path when it cannot succeed

parse_float_fast() (Clinger) needs a significand below 2^53, so it always
declines once the mantissa has 17 or more significant digits. convert_number()
called it unconditionally, so those numbers were walked an extra time before
strtod had to run anyway. On streaming input, where scanning is byte-at-a-time
and there is no compensating win, that made canada.json about 6% slower than
develop.

Derive the significant-digit count from token_buffer indices - the digits are
not scanned again - and skip the call when it is guaranteed to decline. Both
scanners pass the offset where the mantissa ends; the count only has to be
corrected for a leading "0", which the JSON grammar admits nowhere else. The
integer path returns before the check, so integer-heavy input is unaffected.

Values are unchanged: this only avoids an attempt that would have failed.
Verified bit-exact against develop over every number in canada.json,
floats.json, signed_ints.json, unsigned_ints.json, small_signed_ints.json,
citm_catalog.json and twitter.json, for both the contiguous and the streaming
scanner.

  parse, streaming     develop    before     after
  canada.json           19.4ms    20.5ms    19.3ms
  floats.json          135.9ms   131.8ms   128.0ms

  parse, contiguous    develop    before     after
  canada.json           15.5ms    12.9ms    11.7ms
  floats.json           98.6ms    69.8ms    66.7ms

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-10 17:20:48 +02:00
Niels Lohmann d69fb8654d Return the parsed value by move from from_cbor() and friends (#5501)
The binary entry points end with

    return res ? result : basic_json(value_t::discarded);

The condition operator's second operand is an lvalue, so this is not a case
where the return value can be elided or implicitly moved from: every
successful from_cbor(), from_msgpack(), from_ubjson(), from_bjdata() and
from_bson() call deep-copies the value it just parsed, and then destroys the
original.

The copy is not cheap, and it is not incidental: basic_json's copy
constructor walks the whole value. Parsing a 2 MB CBOR document with 60,000
objects, median of 25 runs, clang 17 -O3:

    from_cbor      26.99 ms  ->  14.65 ms
    from_msgpack   26.82 ms  ->  14.82 ms

Moving instead of copying is the entire change; the parsed value is not used
again after the return expression is evaluated.

There is a second reason to prefer the move. The copy constructor recurses
once per nesting level, so the copy is also a stack-overflow path on the
return side, on a value the reader has already accepted. That is currently
masked because the readers themselves recurse and overflow first (#5104), but
it has to be fixed for making them iterative to have any effect.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-10 17:15:34 +02:00
Wu Shuwen 3e2de32226 Remove unused iomanip include (#5516)
Signed-off-by: dajiaohuang <mikewushuwen@outlook.com>
2026-09-10 08:27:05 +02:00
Niels Lohmann b17c272f46 Reserve capacity in from_json() object conversion when the target container supports it (#5472)
* Reserve capacity in from_json() object conversion when supported

The object-to-container from_json() overload filled the target
container one element at a time without reserving capacity, even
when the target type supports reserve() (e.g. std::unordered_map)
and the number of elements is already known. This caused unnecessary
rehashing while parsing large objects into such containers.

Add a reserve-detecting overload (from_json_object_impl), mirroring
the priority_tag-based SFINAE technique already used by the array
conversion path (from_json_array_impl), so that reserve(size()) is
called up front when available and the loop falls back unchanged
otherwise (e.g. for std::map).

Fixes #5406

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update doc example outputs for new object-conversion iteration order

Reserving capacity in from_json()'s object-conversion path before
inserting elements changes libstdc++'s std::unordered_map bucket
layout, which changes the iteration order used by
get__ValueType_const.cpp, get_to.cpp and operator__ValueType.cpp to
print the elements of a converted std::unordered_map<std::string,
json>. Verified against a clean develop checkout (built with the
same GCC/libstdc++ used in CI) that the old order was produced
without this PR's change and the new order is produced with it, and
that the three affected examples now match their updated expected
output byte-for-byte.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Factor out a reserve-dispatch helper instead of duplicating the object from_json loop

Addresses review feedback from @gregmarr on PR #5472: the emplace loop no
longer needs to exist twice for the reserve/no-reserve cases. A small
from_json_object_reserve() overload pair (SFINAE-dispatched on whether
reserve() exists, mirroring the priority_tag technique used elsewhere)
either calls reserve() or is a no-op; from_json_object_impl() calls it once.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Inline from_json_object_impl into from_json now that it is called only once

Addresses review feedback from @gregmarr on PR #5472: with the reserve
loop de-duplicated, from_json_object_impl no longer needs to be a
separate function that from_json immediately delegates to.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:54:12 +02:00
Niels Lohmann 8e26b9d4b3 Make contains(json_pointer) return false instead of throwing on unrepresentable array-index tokens (#5495)
contains(const json_pointer&) is documented to never throw, but a purely
numeric reference token that is syntactically a valid array index yet
numerically too large to be represented (exceeding size_type's max, or
exceeding ULLONG_MAX and causing strtoull() to set errno to ERANGE) made
it fall through to array_index(), which throws out_of_range.410/404.

Pre-check the token's magnitude the same way array_index() does, but
return false instead of throwing, mirroring how the surrounding code
already rejects other malformed tokens (leading zero, non-digit
characters, "-") without throwing. operator[]/at() are untouched and
keep throwing for these inputs.

Fixes #5395

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:53:25 +02:00
Niels Lohmann 8c8d14cb4e Reduce test-suite compile time: extract tiny per-standard test content; drop redundant legacy-comparison CI job (#5481)
* Extract C++17-only content from unit-items.cpp into its own test file

unit-items.cpp is a 1433-line file that was being compiled twice per
CI configuration (once for C++11, once for C++17) purely because it
contained a single, small JSON_HAS_CPP_17-gated SECTION ("structured
bindings", 14 lines). Move that SECTION into a new, dedicated file
(tests/src/unit-items-cpp17.cpp) so only that tiny file needs a
second build; unit-items.cpp itself now builds/tests only once. No
tests/CMakeLists.txt changes are needed since the existing
file(GLOB ... src/unit-*.cpp) plus json_test_add_test_for() already
auto-register and standard-gate any new unit-*.cpp file based on
whether it textually contains JSON_HAS_CPP_<N> (the same mechanism
already used for the existing unit-iterators3.cpp file, which follows
the identical pattern).

Verified with plain clang++ under -std=c++11/14/17/20 and via a local
CMake configure+build that:
- unit-items.cpp now only produces a test-items_cpp11 target (the
  former test-items_cpp17 target is gone) and its assertion/test-case
  counts are unchanged (2 test cases / 222 assertions) for every
  standard.
- The new unit-items-cpp17.cpp produces test-items-cpp17_cpp11 (an
  intentionally empty translation unit under C++11 that reports 0
  tests, 0 assertions, SUCCESS) and test-items-cpp17_cpp17 (1 test
  case / 1 assertion, identical to what "structured bindings" ran
  as before it was moved).

Separately, unit-regression1.cpp (1530 lines) was also being built
twice per CI configuration because it contained the substring
JSON_HAS_CPP_17 -- but on inspection this was dead code: an orphaned
"#ifdef JSON_HAS_CPP_17 / #include <variant> / #endif" left over from
when the actual std::variant-based regression test (issue #1292) was
relocated to unit-regression2.cpp. Nothing in unit-regression1.cpp
uses <variant>, so there is no SECTION/TEST_CASE to preserve here;
the dead include is simply removed. This was verified by grepping the
file for any other use of "variant" (none) and confirming issue #1292
is still covered by unit-regression2.cpp. Compiled and ran under
-std=c++11/14/17/20 and via CMake: unit-regression1.cpp now only
produces a test-regression1_cpp11 target (test-regression1_cpp17 is
gone) with an unchanged test-case count (3) under every standard.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix astyle indentation of #include inside #ifdef in unit-items-cpp17.cpp

This repo's astyle style keeps preprocessor directives at column 0
even inside #ifdef blocks. The new tests/src/unit-items-cpp17.cpp
had its #include <map>/#include <string> indented, which made the
'check' CI job's amalgamation/formatting diff non-empty and failed
the aggregate check.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:52:58 +02:00
Niels Lohmann d22c5c7641 Add test coverage for documented lenient BSON input handling (#5478)
* Add test coverage for documented lenient BSON input handling

Issue #5333 documented three intentionally-lenient behaviors of the BSON
reader (any non-zero byte accepted as a boolean `true`, BSON array element
keys not validated against the required decimal sequence, and the payload
of binary subtype 0x02 "old binary" returned as-is including its inner
length prefix), but none of them was pinned by a test, so a future change
could silently regress the documented behavior.

Also add coverage for the out_of_range.412 length-overflow check
(shared by binary, string, and (sub-)document BSON length fields) for
the string and document cases; only the binary case was previously
tested.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix 32-bit overflow in huge_string_t BSON length-overflow tests

huge_string_t doubles as basic_json's StringType, so it is used not only
for the JSON string value under test but also for object keys (e.g. "s",
"nested"). Making size() unconditionally lie about being huge therefore
inflated the keys' reported sizes as well, pushing the running totals
computed while walking the BSON document (calc_bson_object_size and
friends in binary_writer.hpp) past what a 32-bit std::size_t can hold.

On 64-bit platforms this happens to still produce a working (if
needlessly large) result, but on 32-bit platforms (e.g. the mingw x86 CI
job) the size_t arithmetic silently wraps around: for the "document" test
this merely surfaces the wrong number in the exception message, but for
the "string" test the wrapped total happens to fall back under
INT32_MAX, so the intended out_of_range.412 guard is skipped entirely and
the code goes on to actually write ~2 GiB worth of characters from the
key's real, tiny buffer - which is what raised the reported
"vector::_M_range_insert" exception instead of a controlled 412.

Make the fake-huge size opt-in via huge_string_t::as_huge() and only
apply it to the string value under test, leaving keys at their real
(small) size. This keeps every intermediate size well within 32-bit
size_t range on any platform, matching how huge_binary_t already avoids
the same trap (it is only ever used as the BSON value type, never as a
key). Expected out_of_range.412 messages are updated accordingly.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:52:23 +02:00
Niels Lohmann c9f1d2d646 Broaden JSON_HEDLEY_WARN_UNUSED_RESULT coverage to pure query functions (#5477)
* Broaden JSON_HEDLEY_WARN_UNUSED_RESULT coverage to pure query functions

Add JSON_HEDLEY_WARN_UNUSED_RESULT to the unambiguous, const,
side-effect-free observer functions whose return value is the entire
purpose of the call:

- dump()
- type(), type_name()
- all is_* predicates (is_primitive, is_structured, is_null,
  is_boolean, is_number, is_number_integer, is_number_unsigned,
  is_number_float, is_object, is_array, is_string, is_binary,
  is_discarded)
- empty(), size(), max_size()
- count(...) (both overloads) and contains(...) (all overloads,
  including the deprecated json_pointer<BasicJsonType> overload)

This mirrors the direction the standard library has taken with
[[nodiscard]] on the analogous std::vector/std::map members, and
catches real bugs such as `j.empty();` (meant `j.clear();`) or
`j.contains(k);` with the result thrown away.

Deliberately out of scope (left for a separate, later policy
decision, per the issue): at(), value(), get*(), flatten(),
unflatten(), patch(), merge_patch(), begin()/end(), comparison
operators, erase(), and emplace().

Compiling the full test suite (tests/src/unit-*.cpp) with
-Wunused-result -Werror uncovered one real hit: a regression test in
unit-regression2.cpp called dump() purely to check it does not throw,
discarding the result. Fixed by explicitly casting to void, since the
call is intentionally result-less there.

Fixes #5410

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix discarded nodiscard results across the test suite for GCC's warn_unused_result

A plain (void) cast on a call expression suppresses the C++17 [[nodiscard]]
warning but not GCC's warning for functions annotated via the GNU
__attribute__((warn_unused_result)) form -- which is what
JSON_HEDLEY_WARN_UNUSED_RESULT expands to on GCC. Several existing tests
that call a newly-annotated function (dump(), empty()) purely to check
that it throws/does not throw, discarding the result via (void), newly
warned (and failed -Werror builds) once the annotation was broadened.
Route those discards through a small ignore_return_value() helper
instead, which actually consumes the value and suppresses the warning
on both attribute forms.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use utils::ignore_return_value() for the issue #1445 dump() discard too

Addresses review feedback from @gregmarr on PR #5477: this call site was
still using the older "capture in a variable, then (void) it" pattern
from before this PR introduced utils::ignore_return_value(), instead of
the helper now used at every other discarded-nodiscard-result call site
this PR touches.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:50:39 +02:00
Niels Lohmann a316738cfa Add JSON_HEDLEY_WARN_UNUSED_RESULT to the current accept() overloads (#5471)
accept() is a pure query whose only effect is the returned bool; both
parse() overloads and the deprecated accept(span_input_adapter&&, ...)
overload already carry JSON_HEDLEY_WARN_UNUSED_RESULT, but the two
current, recommended accept() overloads were missing it. Add the
annotation to match, so discarding accept()'s result now warns under
-Wunused-result / [[nodiscard]], as it already does for parse().

Fixes #5407

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:50:39 +02:00
Niels Lohmann 8c38e270b5 Reject JSON Patch move when from is a proper prefix of path (#5497)
* Reject JSON Patch move when from is a proper prefix of path

RFC 6902 (section 4.4) forbids "from" from being a proper prefix of
"path" for a "move" operation: "a location cannot be moved into one
of its children." "move" is implemented as remove-then-add with no
check for this. For object targets, the subsequent "add" happened to
throw as a side effect of resolving through the now-removed parent,
but for array targets, removing the "from" element shifts subsequent
indices, so "path" silently re-resolves to a different element and
the operation "succeeds" with a silently corrupted document.

Add a check, before performing the remove/add, for whether "from" is
a proper prefix of "path" at the reference-token level. This compares
json_pointer's already-unescaped reference_tokens vectors (basic_json
is a friend of json_pointer) rather than the raw pointer strings, so
that tokens containing escaped '/' or '~' characters are compared
correctly, and a token that merely looks like a string prefix (e.g.
"/ab" vs "/abc/x") is not mistaken for a pointer-token prefix. When
"from" is a proper prefix of "path", throw out_of_range.414.

Fixes #5397.

Stacked on top of the fix for #5396 (branch
issue-5396-patch-remove-primitive-parent), since both touch the same
patch_inplace move/remove handling in include/nlohmann/json.hpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add root-pointer and array-append-token edge case tests for the move prefix check

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Replace std::equal with an explicitly-bounded loop in the move prefix check

The three-iterator std::equal(first1, last1, first2) form has no
explicit end iterator for the second range, which a static analyzer
(Flawfinder, CWE-126) flags as a potential over-read even though the
preceding size comparison already guarantees the second range is long
enough. Rather than argue the point, make the bound visible in the code
itself via an explicit loop -- every access to ptr.reference_tokens is
now guarded by the same index the loop condition bounds against
from_size.

(The C++14 four-iterator std::equal(first1, last1, first2, last2) form
was tried first as a more minimal fix, but this codebase targets C++11
and that overload is not safely usable under -std=c++11 with all
supported standard library implementations.)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Extract the move prefix check into a named helper lambda

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Account for JSON_DIAGNOSTIC_POSITIONS in the move-prefix-check error messages

out_of_range::create() includes a "(bytes X-Y)" position annotation when
JSON_DIAGNOSTIC_POSITIONS is enabled, which the ci_test_diagnostic_positions
CI job builds the whole suite with. The five new out_of_range.414
assertions only checked the annotation-free message. Confirmed
JSON_DIAGNOSTICS produces the same (annotation-free) message as the
default build for this particular throw site (its path-based annotation
is empty at the root, where &result always points here), so only two
message variants are needed, not three.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:49:12 +02:00
Niels Lohmann c9477ccf91 Throw when JSON Patch remove's target path resolves through a primitive or null parent (#5496)
RFC 6902 (section 4.2) requires the target location of a "remove"
operation to exist. operation_remove handled parent.is_object() and
parent.is_array(), but had no final else branch: when the resolved
parent was a primitive value or null, neither branch matched and the
operation silently did nothing instead of failing.

Add the missing else branch, throwing out_of_range.413 with wording
that matches the existing out_of_range.411 thrown by the analogous
"add" case (operation_add) for the same kind of invalid parent.

Fixes #5396.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:49:12 +02:00
Niels Lohmann d6660cf718 Route hand-rolled diagnostic pragmas through Hedley (#5485)
* Route hand-rolled diagnostic pragmas through Hedley

Several places in the library hand-roll compiler diagnostic suppression
with raw `#pragma`/`#ifdef __GNUC__`/`#ifdef __clang__` guards instead of
using the Hedley primitives already bundled and used elsewhere
(JSON_HEDLEY_DIAGNOSTIC_PUSH/POP, JSON_HEDLEY_PRAGMA, ...). Converted six
of the seven listed push/pop pairs to use those primitives instead of
raw `#pragma GCC diagnostic`/`#pragma clang diagnostic` text:

- include/nlohmann/json.hpp (~3770, ~3863): -Wfloat-equal
- include/nlohmann/detail/conversions/to_chars.hpp (~1078): -Wfloat-equal
- include/nlohmann/detail/output/binary_writer.hpp (~1844): -Wfloat-equal
- include/nlohmann/detail/iterators/iteration_proxy.hpp (~211): -Wmismatched-tags
- include/nlohmann/detail/exceptions.hpp (~36): -Wweak-vtables

iteration_proxy.hpp did not previously include macro_scope.hpp itself
(it only compiled because some other header included earlier in
json.hpp happened to pull macro_scope.hpp in first); it now includes it
directly like the other detail headers that use Hedley macros, so it is
self-contained.

Each push/pop pair now uses JSON_HEDLEY_DIAGNOSTIC_PUSH/POP
unconditionally (a no-op on compilers that don't need it) and wraps the
actual `#pragma ... diagnostic ignored` text in JSON_HEDLEY_PRAGMA so it
goes through Hedley's _Pragma()-based emission instead of a raw #pragma
line, while keeping the original `#ifdef __GNUC__` / `#if
defined(__clang__)` guard around the ignored-pragma itself.

Deviation from the issue's suggested transformation: the issue's example
replaces the `#ifdef __GNUC__` guard with `#if
JSON_HEDLEY_HAS_WARNING("-Wfloat-equal")`. JSON_HEDLEY_HAS_WARNING is
implemented purely via Clang's `__has_warning` builtin and evaluates to
0 on real GCC (`#define JSON_HEDLEY_HAS_WARNING(warning) (0)` when
`__has_warning` is not defined), so adopting it verbatim would silently
stop suppressing -Wfloat-equal on GCC -- a real regression, not just a
style change. The existing `#ifdef __GNUC__` / `#if defined(__clang__)`
guards were kept for the ignored-pragma to stay behavior-preserving, and
only the push/pop/pragma-emission mechanism was routed through Hedley.

Two of the seven locations from the issue (the -Wignored-attributes
push at the very top of json.hpp and its matching pop after
`#include <nlohmann/detail/macro_unscope.hpp>`) were intentionally left
unconverted:
- The push, at the very top of json.hpp, runs before
  `detail/macro_scope.hpp` (and therefore hedley.hpp) has been included
  anywhere in the translation unit, so JSON_HEDLEY_DIAGNOSTIC_PUSH is not
  yet defined at that point.
- The pop runs after `macro_unscope.hpp`, which -- via hedley_undef.hpp
  -- has already #undef'd every JSON_HEDLEY_* macro (by design, see
  #5408) precisely so they don't leak to users, so JSON_HEDLEY_DIAGNOSTIC_POP
  is no longer defined by the time the pop is reached either.
  Making this one pair work would require either hoisting the ~2000
  line vendored hedley.hpp to the very top of the amalgamated single
  header (a much bigger structural change to single_include than a pure
  mechanism swap) or special-casing this one pop ahead of the general
  macro cleanup. Both are riskier than the mechanical, behavior-preserving
  change requested, so this pair was left as-is.

## Validation

- Compiled include/nlohmann/json.hpp and single_include/nlohmann/json.hpp
  with `-Wall -Wextra -Wfloat-equal -Wmismatched-tags -Wweak-vtables`
  (clang, which self-identifies as __GNUC__ too): no warnings, same as
  before the change.
- Compiled and ran tests/src/unit-to_chars.cpp, unit-conversions.cpp,
  unit-iterators1.cpp, unit-iterators2.cpp, and unit-class_parser.cpp
  against the fixed include/: all pass.
- Compiled unit-msgpack.cpp, unit-bjdata.cpp, and unit-ubjson.cpp (which
  exercise binary_writer.hpp's write_compact_float extensively): all
  compile cleanly; the vast majority of assertions pass (the only
  failures are pre-existing environment issues unrelated to this change
  -- missing generated test-data files, not code correctness).
- Ran `make amalgamate`; the single_include diff is limited to exactly
  the lines touched in include/, with no unrelated reordering.
- No real (non-Apple) GCC was available in this environment to test
  directly; the `_Pragma("GCC diagnostic ...")` text emitted by
  JSON_HEDLEY_PRAGMA is byte-identical to the prior `#pragma GCC
  diagnostic ...` text, and the `#ifdef __GNUC__` guard is unchanged, so
  GCC's behavior is expected to be identical. CI covers the GCC matrix.

This PR is stacked on top of #5475 (issue-5408-hedley-undef-leak) since
both touch the same files; only the last commit here is new.

Fixes #5409.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Guard JSON_HEDLEY_DIAGNOSTIC_PUSH/POP with the same compiler check as the pragma they bracket

Addresses review feedback from @gregmarr on PR #5485: the push/pop calls
were unconditional, so compilers other than the one the ignored-pragma
targets (e.g. MSVC, or GCC where the pair only applies under __clang__)
now did a needless push/pop with nothing suppressed in between. Move the
existing #ifdef __GNUC__ / #if defined(__clang__) guard to also cover the
push/pop, restoring the original zero-overhead behavior on other compilers
while still emitting the pragma itself through Hedley.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:48:18 +02:00
Niels Lohmann c05347dd54 Undefine the four JSON_HEDLEY_* macros that leak after including json.hpp (#5475)
* Undefine the four JSON_HEDLEY_* macros that leak after including json.hpp

include/nlohmann/detail/macro_unscope.hpp includes hedley_undef.hpp to
#undef every JSON_HEDLEY_* macro so none of them leak into the including
translation unit. Four macros were missing from that list and therefore
stayed defined after #include <nlohmann/json.hpp>:

- JSON_HEDLEY_PRAGMA
- JSON_HEDLEY_PREDICT_TRUE
- JSON_HEDLEY_PREDICT_FALSE
- JSON_HEDLEY_CLANG_HAS_DECLSPEC_ATTRIBUTE

hedley_undef.hpp is generated (via `make update_hedley`) by grepping
hedley.hpp for its own internal `#undef JSON_HEDLEY_X` redefinition
guards. JSON_HEDLEY_PRAGMA/PREDICT_TRUE/PREDICT_FALSE have no such guard
in upstream Hedley, so they were never picked up. The guard for
JSON_HEDLEY_CLANG_HAS_DECLSPEC_ATTRIBUTE also has an upstream typo
(`JSON_HEDLEY_CLANG_HAS_DECLSPEC_DECLSPEC_ATTRIBUTE`), so hedley_undef.hpp
was undefining the wrong (never-defined) name.

Fixes:
- include/nlohmann/thirdparty/hedley/hedley_undef.hpp: corrected the
  DECLSPEC_ATTRIBUTE typo and added the three missing #undef lines,
  keeping the file's alphabetical ordering.
- Makefile (update_hedley target): changed hedley_undef.hpp generation to
  extract macro names directly from every `#define JSON_HEDLEY_...` in
  hedley.hpp instead of from existing `#undef` guards, so a future
  `make update_hedley` run undefines every macro Hedley actually defines,
  even ones without a pre-existing redefinition guard. This was not run
  in this PR (it would also pull in an unrelated upstream Hedley sync);
  hedley_undef.hpp was hand-patched instead and single_include was
  regenerated with `make amalgamate`.
- tests/src/unit-no-macro-leak.cpp: new regression test (picked up
  automatically by tests/CMakeLists.txt's existing unit-*.cpp glob) that
  includes json.hpp and then #ifdef/#error-checks every JSON_HEDLEY_*
  macro name, so any future leak of any of the 151 vendored macros fails
  the build, not just the four fixed here.

Fixes #5408.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Derive the JSON_HEDLEY_* leak-check test from hedley.hpp at build time

tests/src/unit-no-macro-leak.cpp previously hardcoded a static list of
~151 #ifdef/#error checks, one per JSON_HEDLEY_* macro name known at the
time it was written. That list would silently go stale the next time
`make update_hedley` pulls in a vendor update that adds, removes, or
renames a macro, since nothing would force it to be regenerated.

Add cmake/scripts/gen_hedley_undef_check.cmake, which derives the full
list of JSON_HEDLEY_* macro names directly from
include/nlohmann/thirdparty/hedley/hedley.hpp:

- tests/CMakeLists.txt uses it (MODE=checks) to (re)generate
  hedley_undef_checks.inc at configure and build time, and wires the
  generating custom target as a dependency of the test-no-macro-leak_cpp*
  targets so it can never build against a stale copy. unit-no-macro-leak.cpp
  now just #include-s the generated file inside its TEST_CASE instead of
  carrying the checks itself.
- The Makefile's `update_hedley` target now delegates hedley_undef.hpp
  generation to the same script (MODE=undef, new `update_hedley_undef`
  target), so the vendored header, the generated #undef list, and the
  generated test checks are all derived from the same extraction logic and
  cannot drift apart.

This mirrors the approach taken independently in #5415 for the same
issue (#5408), credited there to a self-regenerating mechanism that
"can never drift again" -- ported into this branch instead of the
static list originally proposed here.

Verified with a local CMake configure + build + ctest, both against
include/ (JSON_MultipleHeaders=ON) and against the amalgamated
single_include/nlohmann/json.hpp (JSON_MultipleHeaders=OFF), and by
temporarily deleting a #undef line from hedley_undef.hpp to confirm the
generated test actually fails on a real leak.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix REUSE compliance failure in gen_hedley_undef_check.cmake

The generated file's embedded banner contains the literal text
'SPDX-License-Identifier: MIT' as part of the *content* being written
to hedley_undef.hpp, not as this .cmake script's own REUSE header (it
is already covered by the blanket 'Files: *' rule in .reuse/dep5).

The reuse tool matched that embedded line as an SPDX tag for the
script itself and failed to parse the trailing 'MIT\n")' as a valid
SPDX License Expression, breaking ci_reuse_compliance. Wrap the
embedded banner in REUSE-IgnoreStart/REUSE-IgnoreEnd comments, as
recommended by the tool's own diagnostic output.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:48:17 +02:00
Niels Lohmann 6d86cc0f6b Speed up whitespace skipping in the lexer (#5490)
* Speed up whitespace skipping in the lexer

lexer::skip_whitespace() called get() for every whitespace byte, and
get() checks the (almost always false, once past the first character)
next_unget flag on every call. skip_whitespace() now reads its first
character with get() (needed to honor a pending unget() left over from
finishing the previous token, e.g. scan_number() always ungets the
character that terminated the number) and every further whitespace
character with a new get_ignoring_pending_unget() variant that skips
that branch, since nothing in the loop calls unget().

This is a narrower fix than the full contiguous-buffer bulk-skip
suggested in the issue (scan a run of whitespace directly in the
adapter's buffer and update position counters once per run). That
approach depends on bulk-scan adapter infrastructure
(supports_bulk_scan/bulk_data()/bulk_skip()) introduced by the open,
unmerged parser-performance PR #5283, which this change intentionally
does not depend on or replicate. Building new bulk-scan adapter
infrastructure from scratch was judged out of scope/riskier than
warranted here, so this change is limited to the safe, always-correct
improvement of removing redundant per-character bookkeeping from the
existing byte-at-a-time loop; full bulk-skipping is left as future
work once #5283 (or equivalent adapter support) lands.

Line/column/byte-offset bookkeeping is untouched and verified
bit-for-bit identical before and after this change, including for
pretty-printed (dump(4)) input with embedded newlines.

Fixes #5412

Stacked on top of the PR for #5411 (branch
issue-5411-lexer-skip-conversion).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix codegen regression in skip_whitespace() from #5490

Benchmarking found the get()/get_ignoring_pending_unget() split in
skip_whitespace() made long whitespace runs (e.g. indentation in
pretty-printed JSON) 1.75x-3.2x SLOWER instead of faster, reproducible
with both Apple Clang and GCC.

Root cause: rewriting the loop from a plain do-while into an initial
get() followed by a while-loop defeated the compiler's ability to keep
the input adapter's read/end pointers in registers across iterations;
both compilers instead reloaded them from memory on every character.
The function split itself was not the problem (it still fully
inlines); the loop's control-flow shape was.

The fix keeps the same two-function structure but restores a
do-while shape (guarded by an if for the "first char not whitespace"
case), which lets both compilers hoist the pointers back into
registers, matching or beating pre-#5490 performance.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Share the position-counter bump between get() and get_ignoring_pending_unget()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Extract current_is_whitespace() to deduplicate skip_whitespace()'s two whitespace checks

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use a raw string literal for the multi-line error-position test input

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix clang-tidy raw-string-literal finding and guard a new test against JSON_NOEXCEPTION

The issue #5412 whitespace-skipping test added a check_error() helper that
relies on catching json::parse_error to verify the exception message; under
JSON_NOEXCEPTION, JSON_THROW aborts instead of throwing, which crashed
ci_test_noexceptions (and cascaded into the other ci_cmake_options jobs).
Guard the whole section with #if !defined(JSON_NOEXCEPTION), matching the
existing pattern used by sibling tests in this file.

Also switch one escaped string literal to a raw string literal to satisfy
clang-tidy's modernize-raw-string-literal check.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:46:26 +02:00
Niels Lohmann faa35cc647 Skip integer conversion in accept()/SAX validation when the value is unused (#5484)
* Skip integer conversion in accept()/SAX validation when the value is unused

lexer::scan_number() always converted every numeric token with
strtoull()/strtoll() before returning, even though accept() (and any
consumer using json_sax_acceptor) immediately discards the converted
value. For value_unsigned/value_integer tokens whose digit count
already guarantees the value fits into 64 bits, the conversion cannot
change the accept/reject decision (such tokens are always finite and
unconditionally accepted), so scan_number() can skip strtoull()/
strtoll() entirely in that case when the caller signals it does not
need the value. Numbers with more digits keep using the exact,
unmodified conversion path, so overflow reclassification to
value_float (and the finiteness check on it) is unaffected.

parse() and value_float handling are completely unchanged.

Fixes #5411

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document why the digit-count fast path is safe regardless of number_unsigned_t/number_integer_t width

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Avoid temporary-string concatenation flagged by clang-tidy in the differential test

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 09:46:26 +02:00
dependabot[bot] 7b2d73cf2e Bump step-security/harden-runner from 2.21.0 to 2.21.1 (#5513)
Bumps [step-security/harden-runner](https://github.com/step-security/harden-runner) from 2.21.0 to 2.21.1.
- [Release notes](https://github.com/step-security/harden-runner/releases)
- [Commits](https://github.com/step-security/harden-runner/compare/05e31511f85b41b11d1cf0ef85d0992719546e2c...e14015d583714f6e62063499dc959a02595150a1)

---
updated-dependencies:
- dependency-name: step-security/harden-runner
  dependency-version: 2.21.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-07 21:37:08 +02:00
140 changed files with 17281 additions and 4811 deletions
+3 -1
View File
@@ -108,7 +108,9 @@ The tests are located in [`tests/src/unit-*.cpp`](https://github.com/nlohmann/js
are structured along the features of the library or the nature of the tests. Usually, it should be clear from the
context which existing file needs to be extended, and only very few cases require creating new test files.
When fixing a bug, edit `unit-regression2.cpp` and add a section referencing the fixed issue.
When fixing a bug, edit `unit-regression3.cpp` and add a section referencing the fixed issue.
`unit-regression2.cpp` holds the older tests; the two files exist because a single one grew large enough for the
MinGW linker to fail relocating it, so please keep adding to the smaller file rather than growing the larger one.
#### Exceptions
+2 -2
View File
@@ -11,7 +11,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -34,7 +34,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+1 -1
View File
@@ -9,7 +9,7 @@ jobs:
runs-on: ubuntu-22.04
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+4 -4
View File
@@ -27,7 +27,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -38,14 +38,14 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
uses: github/codeql-action/init@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4.38.0
with:
languages: c-cpp
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
uses: github/codeql-action/autobuild@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4.38.0
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
uses: github/codeql-action/analyze@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4.38.0
@@ -19,7 +19,7 @@ jobs:
pull-requests: write
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+1 -1
View File
@@ -17,7 +17,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+2 -2
View File
@@ -27,7 +27,7 @@ jobs:
security-events: write
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -43,6 +43,6 @@ jobs:
output: 'flawfinder_results.sarif'
- name: Upload analysis results to GitHub Security tab
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
uses: github/codeql-action/upload-sarif@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4.38.0
with:
sarif_file: ${{github.workspace}}/flawfinder_results.sarif
+1 -1
View File
@@ -17,7 +17,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+1 -1
View File
@@ -26,7 +26,7 @@ jobs:
runs-on: ubuntu-22.04
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+2 -2
View File
@@ -36,7 +36,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -76,6 +76,6 @@ jobs:
# Upload the results to GitHub's code scanning dashboard.
- name: "Upload to code-scanning"
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
uses: github/codeql-action/upload-sarif@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4.38.0
with:
sarif_file: results.sarif
+2 -2
View File
@@ -32,7 +32,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -61,7 +61,7 @@ jobs:
# Upload SARIF file generated in previous step
- name: Upload SARIF file
uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9
uses: github/codeql-action/upload-sarif@b96794f015dfd88f77b49b1c93e0fa7110f94c63 # v4.38.0
with:
sarif_file: semgrep.sarif
if: always()
+1 -1
View File
@@ -16,7 +16,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+6 -6
View File
@@ -35,7 +35,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -60,7 +60,7 @@ jobs:
target: [ci_test_amalgamation, ci_test_single_header, ci_cppcheck, ci_cpplint, ci_reproducible_tests, ci_non_git_tests, ci_offline_testdata, ci_reuse_compliance, ci_test_valgrind]
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -100,7 +100,7 @@ jobs:
container: ubuntu:focal
strategy:
matrix:
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls]
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_disableenumserialization, ci_test_skiplibraryversioncheck, ci_test_simdutf, ci_test_strict_nul_handling]
steps:
- name: Install build-essential
run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev
@@ -118,7 +118,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -369,7 +369,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
@@ -392,7 +392,7 @@ jobs:
target: [ci_test_examples, ci_test_build_documentation]
steps:
- name: Harden Runner
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
+6
View File
@@ -59,6 +59,7 @@ option(JSON_LegacyDiscardedValueComparison "Enable legacy discarded value compar
option(JSON_Install "Install CMake targets during install step." ${MAIN_PROJECT})
option(JSON_MultipleHeaders "Use non-amalgamated version of the library." ON)
option(JSON_SystemInclude "Include as system headers (skip for clang-tidy)." OFF)
option(JSON_StrictNulHandling "Build with strict NUL-byte handling enabled." OFF)
if (JSON_CI)
include(ci)
@@ -108,6 +109,10 @@ if (JSON_Diagnostics)
message(STATUS "Diagnostics enabled (JSON_DIAGNOSTICS=1)")
endif()
if (JSON_StrictNulHandling)
message(STATUS "Strict NUL-byte handling enabled (JSON_STRICT_NUL_HANDLING=1)")
endif()
if (JSON_Diagnostic_Positions)
message(STATUS "Diagnostic positions enabled (JSON_DIAGNOSTIC_POSITIONS=1)")
endif()
@@ -141,6 +146,7 @@ target_compile_definitions(
$<$<BOOL:${JSON_Diagnostics}>:JSON_DIAGNOSTICS=1>
$<$<BOOL:${JSON_Diagnostic_Positions}>:JSON_DIAGNOSTIC_POSITIONS=1>
$<$<BOOL:${JSON_LegacyDiscardedValueComparison}>:JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1>
$<$<BOOL:${JSON_StrictNulHandling}>:JSON_STRICT_NUL_HANDLING=1>
)
target_include_directories(
+18 -3
View File
@@ -1,4 +1,4 @@
.PHONY: pretty clean ChangeLog.md release
.PHONY: pretty clean ChangeLog.md release update_hedley update_hedley_undef
##########################################################################
# configuration
@@ -41,6 +41,8 @@ all:
@echo "fuzz_testing_ubjson - prepare fuzz testing of the UBJSON parser"
@echo "pretty - beautify code with Artistic Style"
@echo "run_benchmarks - build and run benchmarks"
@echo "update_hedley - download Hedley and regenerate hedley.hpp / hedley_undef.hpp"
@echo "update_hedley_undef - rebuild hedley_undef.hpp from the JSON_HEDLEY_* #define names in hedley.hpp"
##########################################################################
@@ -241,11 +243,24 @@ update_hedley:
rm -f include/nlohmann/thirdparty/hedley/hedley.hpp include/nlohmann/thirdparty/hedley/hedley_undef.hpp
curl https://raw.githubusercontent.com/nemequ/hedley/master/hedley.h -o include/nlohmann/thirdparty/hedley/hedley.hpp
$(SED) -i 's/HEDLEY_/JSON_HEDLEY_/g' include/nlohmann/thirdparty/hedley/hedley.hpp
grep "[[:blank:]]*#[[:blank:]]*undef" include/nlohmann/thirdparty/hedley/hedley.hpp | grep -v "__" | sort | uniq | $(SED) 's/ //g' | $(SED) 's/undef/undef /g' > include/nlohmann/thirdparty/hedley/hedley_undef.hpp
$(SED) -i '1s/^/#pragma once\n\n/' include/nlohmann/thirdparty/hedley/hedley.hpp
$(SED) -i '1s/^/#pragma once\n\n/' include/nlohmann/thirdparty/hedley/hedley_undef.hpp
$(MAKE) update_hedley_undef
$(MAKE) amalgamate
# Rebuild hedley_undef.hpp from every JSON_HEDLEY_* name that hedley.hpp
# #defines. Hedley does not #undef all of its public macros internally (see
# #5408), so grepping those #undef lines misses names such as
# JSON_HEDLEY_PRAGMA. cmake/scripts/gen_hedley_undef_check.cmake is the
# single source of truth for this extraction (tests/CMakeLists.txt uses the
# same script, in MODE=checks, to generate the matching leak-check test), so
# the vendored header, the generated #undef list, and the regression test
# cannot drift apart.
update_hedley_undef:
cmake -DHEDLEY_HPP=include/nlohmann/thirdparty/hedley/hedley.hpp \
-DOUTPUT=include/nlohmann/thirdparty/hedley/hedley_undef.hpp \
-DMODE=undef \
-P cmake/scripts/gen_hedley_undef_check.cmake
##########################################################################
# serve_header.py
##########################################################################
+67 -7
View File
@@ -1421,7 +1421,7 @@ I deeply appreciate the help of the following people.
6. [Joshua C. Randall](https://github.com/jrandall) fixed a bug in the floating-point serialization.
7. [Aaron Burghardt](https://github.com/aburgh) implemented code to parse streams incrementally. Furthermore, he greatly improved the parser class by allowing the definition of a filter function to discard undesired elements while parsing.
8. [Daniel Kopeček](https://github.com/dkopecek) fixed a bug in the compilation with GCC 5.0.
9. [Florian Weber](https://github.com/Florianjw) fixed a bug in and improved the performance of the comparison operators.
9. [Fiona Johanna Weber](https://github.com/Fiona-J-W) fixed a bug in and improved the performance of the comparison operators.
10. [Eric Cornelius](https://github.com/EricMCornelius) pointed out a bug in the handling with NaN and infinity values. He also improved the performance of the string escaping.
11. [易思龙](https://github.com/likebeta) implemented a conversion from anonymous enums.
12. [kepkin](https://github.com/kepkin) patiently pushed forward the support for Microsoft Visual Studio.
@@ -1523,14 +1523,14 @@ I deeply appreciate the help of the following people.
108. [Kevin Tonon](https://github.com/ktonon) overworked the C++11 compiler checks in CMake.
109. [Axel Huebl](https://github.com/ax3l) simplified a CMake check and added support for the [Spack package manager](https://spack.io).
110. [Carlos O'Ryan](https://github.com/coryan) fixed a typo.
111. [James Upjohn](https://github.com/jammehcow) fixed a version number in the compilers section.
111. [James Upjohn](https://github.com/jupjohn) fixed a version number in the compilers section.
112. [Chuck Atkins](https://github.com/chuckatkins) adjusted the CMake files to the CMake packaging guidelines and provided documentation for the CMake integration.
113. [Jan Schöppach](https://github.com/dns13) fixed a typo.
114. [martin-mfg](https://github.com/martin-mfg) fixed a typo.
115. [Matthias Möller](https://github.com/TinyTinni) removed the dependency from `std::stringstream`.
116. [agrianius](https://github.com/agrianius) added code to use alternative string implementations.
117. [Daniel599](https://github.com/Daniel599) allowed to use more algorithms with the `items()` function.
118. [Julius Rakow](https://github.com/jrakow) fixed the Meson include directory and fixed the links to [cppreference.com](https://cppreference.com).
118. [Julius Rakow](https://github.com/juliusrakow) fixed the Meson include directory and fixed the links to [cppreference.com](https://cppreference.com).
119. [Sonu Lohani](https://github.com/sonulohani) fixed the compilation with MSVC 2015 in debug mode.
120. [grembo](https://github.com/grembo) fixed the test suite and re-enabled several test cases.
121. [Hyeon Kim](https://github.com/simnalamburt) introduced the macro `JSON_INTERNAL_CATCH` to control the exception handling inside the library.
@@ -1581,7 +1581,7 @@ I deeply appreciate the help of the following people.
166. [Mark Beckwith](https://github.com/wythe) fixed a typo.
167. [yann-morin-1998](https://github.com/yann-morin-1998) helped to reduce the CMake requirement to version 3.1.
168. [Konstantin Podsvirov](https://github.com/podsvirov) maintains a package for the MSYS2 software distro.
169. [remyabel](https://github.com/remyabel) added GNUInstallDirs to the CMake files.
169. [remyabel](https://github.com/remyabel2) added GNUInstallDirs to the CMake files.
170. [Taylor Howard](https://github.com/taylorhoward92) fixed a unit test.
171. [Gabe Ron](https://github.com/Macr0Nerd) implemented the `to_string` method.
172. [Watal M. Iwasaki](https://github.com/heavywatal) fixed a Clang warning.
@@ -1608,7 +1608,7 @@ I deeply appreciate the help of the following people.
193. [Hubert Chathi](https://github.com/uhoreg) made CMake's version config file architecture-independent.
194. [OmnipotentEntity](https://github.com/OmnipotentEntity) implemented the binary values for CBOR, MessagePack, BSON, and UBJSON.
195. [ArtemSarmini](https://github.com/ArtemSarmini) fixed a compilation issue with GCC 10 and fixed a leak.
196. [Evgenii Sopov](https://github.com/sea-kg) integrated the library to the wsjcpp package manager.
196. [Evgenii Sopov](https://github.com/sea5kg) integrated the library to the wsjcpp package manager.
197. [Sergey Linev](https://github.com/linev) fixed a compiler warning.
198. [Miguel Magalhães](https://github.com/magamig) fixed the year in the copyright.
199. [Gareth Sylvester-Bradley](https://github.com/garethsb-sony) fixed a compilation issue with MSVC.
@@ -1702,7 +1702,7 @@ I deeply appreciate the help of the following people.
287. [NN](https://github.com/NN---) added the Visual Studio output directory to `.gitignore`.
288. [Romain Reignier](https://github.com/romainreignier) improved the performance of the vector output adapter.
289. [Mike](https://github.com/Mike-Leo-Smith) fixed the `std::iterator_traits`.
290. [Richard Hozák](https://github.com/zxey) added macro `JSON_NO_ENUM` to disable default enum conversions.
290. [Richard Hozák](https://github.com/richardhozak) added macro `JSON_NO_ENUM` to disable default enum conversions.
291. [vakokako](https://github.com/vakokako) fixed tests when compiling with C++20.
292. [Alexander “weej” Jones](https://github.com/alexweej) fixed an example in the README.
293. [Eli Schwartz](https://github.com/eli-schwartz) added more files to the `include.zip` archive.
@@ -1727,7 +1727,7 @@ I deeply appreciate the help of the following people.
312. [Gareth Sylvester-Bradley](https://github.com/garethsb) added `operator/=` and `operator/` to construct JSON pointers.
313. [Michael Macnair](https://github.com/mykter) added support for afl-fuzz testing.
314. [Berkus Decker](https://github.com/berkus) fixed a typo in the README.
315. [Illia Polishchuk](https://github.com/effolkronium) improved the CMake testing.
315. [Illia Polishchuk](https://github.com/ilqvya) improved the CMake testing.
316. [Ikko Ashimine](https://github.com/eltociear) fixed a typo.
317. [Raphael Grimm](https://github.com/barcode) added the possibility to define a custom base class.
318. [tocic](https://github.com/tocic) fixed typos in the documentation.
@@ -1797,6 +1797,66 @@ I deeply appreciate the help of the following people.
382. [bitFiedler](https://github.com/bitFiedler) made GDB pretty printer work with Python 3.8.
383. [Gianfranco Costamagna](https://github.com/LocutusOfBorg) fixed a compiler warning.
384. [risa2000](https://github.com/risa2000) made `std::filesystem::path` conversion to/from UTF-8 encoded string explicit.
385. [AM](https://github.com/maqnouch) fixed typos in the README.
386. [dmenendez-gruposantander](https://github.com/dmenendez-gruposantander) fixed typos in the comments of the examples.
387. [Mihai Stan](https://github.com/mstan-xx) fixed comparisons against the literal `0`.
388. [Matt Gumbel](https://github.com/intelmatt) fixed some `-Weffc++` warnings.
389. [vimpunk](https://github.com/vimpunk) moved a lambda out of an unevaluated context to support older compilers.
390. [Chris Harris](https://github.com/cjh1) fixed the compilation with GCC 4.8.
391. [Palmer Dabbelt](https://github.com/palmer-dabbelt) generated and installed a pkg-config file.
392. [Gus Pozuelo](https://github.com/ap-viavi) made `ordered_map` compatible with GCC 5.5, Clang 3.6, and Xcode 9.
393. [AK](https://github.com/Lioncky) fixed an MSVC build error caused by the `min`/`max` macros from `windows.h`.
394. [Sergiu Deitsch](https://github.com/sergiud) provided a fallback for missing `char8_t` support.
395. [Xiaochuan Ye](https://github.com/XueSongTap) fixed `from_msgpack` for `std::byte` input by specializing `std::char_traits`.
396. [Ville Vesilehto](https://github.com/thevilledev) fixed an overflow in the BJData size calculation and rejected overflowing negative integers in CBOR.
397. [NmPassTHFan](https://github.com/nmpassthf) replaced the deprecated `std::is_trivial` for C++26.
398. [Chris Ever](https://github.com/chirsz-ever) added the `ignore_trailing_commas` parser option.
399. [Kuan-Fu Wu](https://github.com/kfwu1999) fixed the example code for `json_pointer` initialization.
400. [David Kilzer](https://github.com/ddkilzer) added a missing header to the input adapters.
401. [Miko](https://github.com/mikomikotaishi) added proper C++20 module support, simplified the module API, and fixed missing exports.
402. [hitgirl](https://github.com/hitgil) fixed the CMake configuration when cross-compiling.
403. [Devon Thomas](https://github.com/ThomaDevOSU) mentioned the Artistic Style formatting in the contribution guidelines.
404. [Erik Hu](https://github.com/Erikhu1) made Coveralls upload errors non-fatal in the CI.
405. [co63oc](https://github.com/co63oc) fixed typos.
406. [DmitriBogdanov](https://github.com/DmitriBogdanov) fixed broken package manager links in the documentation.
407. [Bander](https://github.com/banderzhm) improved the MSVC compatibility of the C++ modules.
408. [Andy Choi](https://github.com/ccpong) removed an unnecessary `template` keyword before `get` in the README and the documentation.
409. [SamareshSingh](https://github.com/ssam18) fixed single-element brace initialization to copy/move instead of wrapping in an array, fixed the `WITH_DEFAULT` macros for `ordered_map`, and handled moved events in `serve_header.py`.
410. [Aditya](https://github.com/Lumowhisp) improved the documentation of the documentation generation.
411. [cheese1](https://github.com/cheese1) clarified the README.
412. [KhloodElhossiny](https://github.com/khloodelhossiny) enabled `std::string_view` keys in `operator[]`.
413. [Charles Cabergs](https://github.com/cacharle) fixed a `-Wtautological-constant-out-of-range-compare` warning.
414. [EALePain](https://github.com/EALePain) made the `std::tuple` conversion work with reference types such as `std::tie`.
415. [koala_oishi](https://github.com/chibi-dogs) fixed grammatical wording in the README.
416. [riccardoori11](https://github.com/riccardoori11) fixed a typo in the documentation.
417. [Swastik Bose](https://github.com/VasuBhakt) fixed the parent pointers after `update()` with `JSON_DIAGNOSTICS` and fixed the Doxygen autolinking of requirements.
418. [trdesilva](https://github.com/trdesilva) added `front`, `pop_front`, and `push_front` to `json_pointer`.
419. [Akhilesh Arora](https://github.com/akhilesharora) fixed an incomplete-type error with `ordered_json`.
420. [Hariom Phulre](https://github.com/hariomphulre) fixed the C++20 modules compilation with GCC.
421. [Kirill Lokotkov](https://github.com/RUSLoker) fixed printing `long double` values.
422. [George Sedov](https://github.com/radistmorse) added the `NLOHMANN_DEFINE_TYPE_*_WITH_NAMES` macros.
423. [Caillin Nugent](https://github.com/nugentcaillin) added the `NLOHMANN_JSON_SERIALIZE_ENUM_STRICT` macro.
424. [Cosmin D.](https://github.com/drcosmin) fixed `std::filesystem::path` conversions and added an MSVC workaround for `std::unique_ptr`.
425. [Paul Dreik](https://github.com/pauldreik) fixed a test relying on implementation-specific behavior.
426. [Daniel Falk](https://github.com/daniel-falk) added missing copyright notices to the SBOM.
427. [Federico Sfriso](https://github.com/federicosfriso05-dotcom) added support for constructing JSON values from C++20 range views.
428. [Luke Banicevic](https://github.com/banaboi) fixed corrupt BSON output for lengths exceeding `INT32_MAX`, cleaned up the BSON writer, and improved the documentation.
429. [Patrick Armstrong](https://github.com/Patrick10199) updated the CBOR references and the half-precision float assertions.
430. [Yash Bavadiya](https://github.com/xevrion) added checks to all BSON reads.
431. [hum4nBeing](https://github.com/hum4nBeing) fixed the overflow handling of high-precision numbers in UBJSON.
432. [tomatotomata](https://github.com/tomatotomata) added checks for reading CBOR tagged subtypes.
433. [YingqiDuan](https://github.com/YingqiDuan) documented the BSON interoperability.
434. [KBS](https://github.com/youdie006) documented the standards compliance and the strictness of `parse()` and `operator>>`.
435. [Petr Bělohlávek](https://github.com/petrbel) added Clang 21 and 22 to the CI.
436. [Dmitry Rantovov](https://github.com/darkdi) fixed the placement of a CBOR documentation block.
437. [ljcjclljc](https://github.com/ljcjclljc) fixed the comparison of large unsigned integers with signed integers.
438. [Sahil Kamate](https://github.com/sahilkamate03) fixed the handling of CBOR tags 0-5 and 21-23.
439. [Krishnanand G](https://github.com/Krishnanand-G) made the UBJSON writer reject `use_type` without `use_size`.
440. [whn](https://github.com/Whning0513) documented the lenient BSON input handling and corrected the complexity of `to_bson`.
441. [elix3r](https://github.com/22elix3r) fixed `update()` with `merge_objects` when merging a primitive into an object.
442. [Avionic Harshit](https://github.com/avionicharshit-byte) made `diff()` linear when an array shrinks.
443. [Qatadaha Bin Matloob](https://github.com/qatcod) fixed comparisons between integers and floats and fixed unparsable BJData output.
444. [Wu Shuwen](https://github.com/dajiaohuang) removed an unused include.
Thanks a lot for helping out! Please [let me know](mailto:mail@nlohmann.me) if I forgot someone.
+70 -1
View File
@@ -212,6 +212,24 @@ add_custom_target(ci_test_legacycomparison
COMMENT "Compile and test with legacy discarded value comparison enabled"
)
###############################################################################
# Validate UTF-8 with simdutf.
###############################################################################
add_custom_target(ci_test_simdutf
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON -DJSON_TestSimdutf=ON
# simdutf needs C++17, so the library falls back to its scalar validator
# below that: build the suite at C++11 to cover the fallback with the macro
# defined, and at C++17 to run every test against simdutf itself
"-DJSON_TestStandards=11\;17"
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_simdutf
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_simdutf
COMMAND cd ${PROJECT_BINARY_DIR}/build_simdutf && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure
COMMENT "Compile and test with simdutf UTF-8 validation enabled"
)
###############################################################################
# Enable brace-init copy semantics.
###############################################################################
@@ -227,6 +245,23 @@ add_custom_target(ci_test_brace_init_copy_semantics
COMMENT "Compile and test with brace-init copy semantics enabled"
)
###############################################################################
# Enable strict NUL-byte handling.
###############################################################################
add_custom_target(ci_test_strict_nul_handling
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON -DJSON_FastTests=ON -DJSON_StrictNulHandling=ON
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_strict_nul_handling
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_strict_nul_handling
# unit-testsuites contains a fixture (a "1e308" test value) that relies on the
# legacy NUL-as-end-of-input behavior this macro disables; exclude it here, as
# it is expected to fail under strict NUL handling and is out of scope for it
COMMAND cd ${PROJECT_BINARY_DIR}/build_strict_nul_handling && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure -E "test-testsuites"
COMMENT "Compile and test with strict NUL-byte handling enabled"
)
###############################################################################
# Disable global UDLs.
###############################################################################
@@ -242,6 +277,40 @@ add_custom_target(ci_test_noglobaludls
COMMENT "Compile and test with global UDLs disabled"
)
###############################################################################
# Disable enum serialization.
###############################################################################
add_custom_target(ci_test_disableenumserialization
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON -DJSON_FastTests=ON -DJSON_DisableEnumSerialization=ON
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_disableenumserialization
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_disableenumserialization
COMMAND cd ${PROJECT_BINARY_DIR}/build_disableenumserialization && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure
COMMENT "Compile and test with enum serialization disabled"
)
###############################################################################
# Skip the multiple-inclusion library version check.
###############################################################################
# tests/src/skip_library_version_check.cpp deliberately simulates a scenario
# (mixing two differently-versioned inclusions of the library in one
# translation unit) that unavoidably triggers the compiler's own "macro
# redefined" warning, so -- unlike the ci_test_* targets above -- it is
# compiled directly here, with a modest warning set, instead of being folded
# into the library's own -Weverything/-Werror unit test matrix.
add_custom_target(ci_test_skiplibraryversioncheck
COMMAND ${CMAKE_COMMAND} -E make_directory ${PROJECT_BINARY_DIR}/skip_library_version_check
COMMAND ${CMAKE_CXX_COMPILER} -std=c++11 -Wall -Wextra
-I${PROJECT_SOURCE_DIR}/include
${PROJECT_SOURCE_DIR}/tests/src/skip_library_version_check.cpp
-o ${PROJECT_BINARY_DIR}/skip_library_version_check/skip_library_version_check
COMMAND ${PROJECT_BINARY_DIR}/skip_library_version_check/skip_library_version_check
COMMENT "Compile and run a translation unit simulating a mismatched library version, with JSON_SKIP_LIBRARY_VERSION_CHECK defined"
)
###############################################################################
# Coverage.
###############################################################################
@@ -546,7 +615,7 @@ add_custom_target(ci_single_binaries
add_custom_target(ci_benchmarks
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Release -GNinja
-S${PROJECT_SOURCE_DIR}/benchmarks -B${PROJECT_BINARY_DIR}/build_benchmarks
-S${PROJECT_SOURCE_DIR}/tests/benchmarks -B${PROJECT_BINARY_DIR}/build_benchmarks
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_benchmarks --target json_benchmarks
COMMAND cd ${PROJECT_BINARY_DIR}/build_benchmarks && ./json_benchmarks
COMMENT "Run benchmarks"
+1 -1
View File
@@ -77,7 +77,7 @@ if(CMAKE_CROSSCOMPILING)
endif()
if(NOT DEFINED LIBCPP_VERSION_OUTPUT_CACHED)
try_run(RUN_RESULT_VAR COMPILE_RESULT_VAR
"${CMAKE_BINARY_DIR}" SOURCES "${CMAKE_SOURCE_DIR}/cmake/detect_libcpp_version.cpp"
"${CMAKE_BINARY_DIR}" SOURCES "${CMAKE_CURRENT_LIST_DIR}/detect_libcpp_version.cpp"
RUN_OUTPUT_VARIABLE LIBCPP_VERSION_OUTPUT
COMPILE_OUTPUT_VARIABLE LIBCPP_VERSION_COMPILE_OUTPUT
)
+112
View File
@@ -0,0 +1,112 @@
# Shared extractor for the JSON_HEDLEY_* macro names defined in hedley.hpp.
#
# Every macro that hedley.hpp #defines must be #undef-ed again once json.hpp
# has been fully processed (see include/nlohmann/detail/macro_unscope.hpp
# and https://github.com/nlohmann/json/issues/5408). Deriving the macro list
# straight from hedley.hpp here -- instead of hand-maintaining it in two
# places -- means hedley_undef.hpp and the regression test that checks for
# leaked macros can never drift apart, even after a future `make
# update_hedley` pulls in new macros from upstream Hedley.
#
# MODE=undef (default): write hedley_undef.hpp (SPDX header, #pragma once,
# one #undef per macro name) -- used by `make update_hedley_undef`
# MODE=checks: write one #ifdef/FAIL_CHECK/#endif per macro name,
# meant to be #include-d inside a TEST_CASE -- used by
# tests/CMakeLists.txt to (re)generate the include for
# tests/src/unit-no-macro-leak.cpp
#
# Required variables:
# HEDLEY_HPP path to include/nlohmann/thirdparty/hedley/hedley.hpp
# OUTPUT path of the file to (over)write
# Optional:
# MODE "undef" (default) or "checks"
if(NOT DEFINED HEDLEY_HPP OR NOT DEFINED OUTPUT)
message(FATAL_ERROR "HEDLEY_HPP and OUTPUT must be set")
endif()
if(NOT EXISTS "${HEDLEY_HPP}")
message(FATAL_ERROR "Hedley header not found: ${HEDLEY_HPP}")
endif()
if(NOT DEFINED MODE)
set(MODE undef)
endif()
if(NOT MODE STREQUAL "undef" AND NOT MODE STREQUAL "checks")
message(FATAL_ERROR "MODE must be undef or checks, got: ${MODE}")
endif()
# Line-anchored, like `grep -oE "^[[:blank:]]*#[[:blank:]]*define[[:blank:]]+JSON_HEDLEY_[A-Za-z0-9_]+"`.
# Unanchored matching would also pick up JSON_HEDLEY_* mentions inside
# comments or string literals elsewhere in the file, which must not turn
# into #undef lines.
file(STRINGS "${HEDLEY_HPP}" hedley_lines)
set(macro_names)
foreach(line IN LISTS hedley_lines)
if("${line}" MATCHES "^[ \t]*#[ \t]*define[ \t]+(JSON_HEDLEY_[A-Za-z0-9_]+)")
list(APPEND macro_names "${CMAKE_MATCH_1}")
endif()
endforeach()
if(NOT macro_names)
message(FATAL_ERROR "No JSON_HEDLEY_* macros found in ${HEDLEY_HPP}")
endif()
list(REMOVE_DUPLICATES macro_names)
# Lexicographic, locale-independent (ASCII-only names) -- matches `LC_ALL=C sort`.
list(SORT macro_names COMPARE STRING)
list(LENGTH macro_names macro_count)
set(generated "")
if(MODE STREQUAL "undef")
# Same banner `make update_hedley_undef` would stamp by hand, so the
# recipe is self-contained and its output is byte-stable across reruns.
# The embedded SPDX tags below are part of the *generated* file's
# content, not a REUSE header for this .cmake script itself (which is
# already covered by the blanket "Files: *" rule in .reuse/dep5) -- keep
# them wrapped in REUSE-IgnoreStart/End so `reuse lint` does not try to
# parse "MIT\n")" as this file's own SPDX-License-Identifier value.
# REUSE-IgnoreStart
string(APPEND generated "// __ _____ _____ _____\n")
string(APPEND generated "// __| | __| | | | JSON for Modern C++\n")
string(APPEND generated "// | | |__ | | | | | | version 3.12.0\n")
string(APPEND generated "// |_____|_____|_____|_|___| https://github.com/nlohmann/json\n")
string(APPEND generated "//\n")
string(APPEND generated "// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>\n")
string(APPEND generated "// SPDX-License-Identifier: MIT\n")
# REUSE-IgnoreEnd
string(APPEND generated "\n")
string(APPEND generated "#pragma once\n")
string(APPEND generated "\n")
foreach(name IN LISTS macro_names)
string(APPEND generated "#undef ${name}\n")
endforeach()
else()
string(APPEND generated "// This file is generated by cmake/scripts/gen_hedley_undef_check.cmake\n")
string(APPEND generated "// from include/nlohmann/thirdparty/hedley/hedley.hpp. Do not edit it by\n")
string(APPEND generated "// hand -- it is regenerated on every build. ${macro_count} macros checked.\n\n")
foreach(name IN LISTS macro_names)
string(APPEND generated "#ifdef ${name}\n")
string(APPEND generated " FAIL_CHECK(\"${name} leaked after including nlohmann/json.hpp\");\n")
string(APPEND generated "#endif\n")
endforeach()
endif()
get_filename_component(output_dir "${OUTPUT}" DIRECTORY)
if(output_dir)
file(MAKE_DIRECTORY "${output_dir}")
endif()
# Avoid rewriting the file (and busting downstream incremental rebuilds)
# when the content has not actually changed.
set(write_output TRUE)
if(EXISTS "${OUTPUT}")
file(READ "${OUTPUT}" existing_content)
if(existing_content STREQUAL generated)
set(write_output FALSE)
endif()
endif()
if(write_output)
file(WRITE "${OUTPUT}" "${generated}")
endif()
@@ -90,6 +90,10 @@ Linear in the length of the input. The parser is a predictive LL(1) parser.
A UTF-8 byte order mark is silently ignored.
By default, a `'\0'` (NUL) byte anywhere in the input is treated as end of input, rather than as an ordinary (and,
outside of a string, invalid) byte; see the [FAQ entry](../../home/faq.md#nul-bytes-in-the-input) for details and the
[`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) macro to opt into rejecting it instead.
## Examples
??? example
@@ -111,6 +115,8 @@ A UTF-8 byte order mark is silently ignored.
- [parse](parse.md) - deserialize from a compatible input
- [sax_parse](sax_parse.md) - parse input using the SAX interface
- [operator>>](../operator_gtgt.md) - deserialize from stream
- [`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input
instead of treating it as end of input
## Version history
@@ -120,6 +126,8 @@ A UTF-8 byte order mark is silently ignored.
- Added `ignore_trailing_commas` in version 3.13.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating
it as end of input; planned to become the default in version 4.0.0.
!!! warning "Deprecation"
+6 -1
View File
@@ -14,7 +14,11 @@ To store objects in C++, a type is defined by the template parameters explained
## Template parameters
`ArrayType`
: container type to store arrays (e.g., `std::vector` or `std::list`)
: container type to store arrays. It must be a vector-like container: the library uses `operator[]`, `at()`, and
`resize()`, and requires random-access iterators. `#!cpp std::vector` and `#!cpp std::deque` qualify;
`#!cpp std::list` does not. See
[Template Parameter Requirements](../../features/types/template_parameters.md#arraytype) for the full list of
requirements.
`AllocatorType`
: the allocator to use for objects (e.g., `std::allocator`)
@@ -66,3 +70,4 @@ Arrays are stored as pointers in a `basic_json` type. That is, for any access to
## Version history
- Added in version 1.0.0.
- Made `capacity()` optional, so that array types such as `#!cpp std::deque` can be used, in version 3.13.0.
+11 -1
View File
@@ -42,7 +42,9 @@ represent a byte array in modern C++.
`value_type` must additionally be exactly one byte wide (e.g., `std::uint8_t`/`char`/`std::byte`): the binary
serializers (CBOR, MessagePack, BSON, UBJSON) read and write the container's raw bytes via
`reinterpret_cast`, which is only correct for byte-sized elements -- a container like
`#!cpp std::vector<std::intptr_t>` will not work as `BinaryType`.
`#!cpp std::vector<std::intptr_t>` will not work as `BinaryType`. The elements must be stored contiguously, and
the binary readers additionally require `resize()` and `operator[]`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#binarytype) for the full list.
## Notes
@@ -50,6 +52,11 @@ represent a byte array in modern C++.
The default values for `BinaryType` is `#!cpp std::vector<std::uint8_t>`.
#### Supported byte types
`#!cpp std::vector<std::uint8_t>`, `#!cpp std::vector<char>`, and `#!cpp std::vector<std::byte>` are supported.
Regardless of which of them is configured, [`dump`](dump.md) writes the bytes as the numbers 0..255.
#### Custom BinaryType behavior
When a custom `BinaryType` is configured (other than the default `#!cpp std::vector<std::uint8_t>`), you can assign
@@ -126,3 +133,6 @@ type `#!cpp binary_t*` must be dereferenced.
## Version history
- Added in version 3.8.0. Changed the type of subtype to `std::uint64_t` in version 3.10.0.
- Fixed [`dump`](dump.md), [`std::hash`](std_hash.md), and [`to_ubjson`](to_ubjson.md) for byte types that are not
integers (e.g., `#!cpp std::byte`) in version 3.13.0. `dump` now writes the bytes of a signed byte type (e.g.,
`#!cpp char`) as 0..255 rather than as negative numbers.
@@ -11,6 +11,14 @@ literals `#!json true` and `#!json false`.
To store boolean values in C++, a type is defined by the template parameter `BooleanType` which chooses the type to use.
## Template parameters
`BooleanType`
: the type to store booleans. As it is stored directly inside a `basic_json` value (in a union), it must be a
trivially default-constructible, trivially copyable, and trivially destructible type that is convertible to and
from `#!cpp bool`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#booleantype).
## Notes
#### Default type
+4
View File
@@ -35,6 +35,10 @@ class basic_json;
| `BinaryType` | type for binary arrays | [`binary_t`](binary_t.md) |
| `CustomBaseClass` | extension point for user code | [`json_base_class_t`](json_base_class_t.md) |
The library imposes a number of requirements on these types that are not expressed as C++ concepts, such as the
container operations `object_t` and `array_t` must provide, or the fact that `StringType` must be `char`-based. They
are collected in [Template Parameter Requirements](../../features/types/template_parameters.md).
## Specializations
- [**json**](../json.md) - default specialization
@@ -88,6 +88,8 @@ Strong exception safety: if an exception occurs, the original value stays intact
do not belong to the same JSON value; example: `"iterators do not fit"`
- Throws [`invalid_iterator.211`](../../home/exceptions.md#jsonexceptioninvalid_iterator211) if `first` or `last`
are iterators into container for which insert is called; example: `"passed iterators may not belong to container"`
- Throws [`invalid_iterator.202`](../../home/exceptions.md#jsonexceptioninvalid_iterator202) if `first` or `last`
do not point to an array; example: `"iterators first and last must point to arrays"`
4. The function can throw the following exceptions:
- Throws [`type_error.309`](../../home/exceptions.md#jsonexceptiontype_error309) if called on JSON values other than
arrays; example: `"cannot use insert() with string"`
@@ -21,8 +21,11 @@ The default value for `CustomBaseClass` is `void`. In this case, an
#### Limitations
The type `CustomBaseClass` has to be a default-constructible class.
The type `CustomBaseClass` has to be a default-constructible, non-`final` class.
`basic_json` only supports copy/move construction/assignment if `CustomBaseClass` does so as well.
A `CustomBaseClass` with non-static data members forfeits `basic_json`'s
[standard layout](https://en.cppreference.com/w/cpp/named_req/StandardLayoutType) guarantee. See
[Template Parameter Requirements](../../features/types/template_parameters.md#custombaseclass).
## Examples
@@ -19,6 +19,12 @@ using json_serializer = JSONSerializer<T, SFINAE>;
The default values for `json_serializer` is [`adl_serializer`](../adl_serializer/index.md).
#### Requirements
A custom serializer must provide `#!cpp static void to_json(basic_json&, T)` for every type it serializes, and either
`#!cpp static void from_json(const basic_json&, T&)` or `#!cpp static T from_json(const basic_json&)` for every type it
deserializes. See [Template Parameter Requirements](../../features/types/template_parameters.md#jsonserializer).
## Examples
??? example
@@ -20,6 +20,16 @@ used.
To store floating-point numbers in C++, a type is defined by the template parameter `NumberFloatType` which chooses the
type to use.
## Template parameters
`NumberFloatType`
: the type to store floating-point numbers. Parsing and serialization are implemented in terms of
`#!cpp std::strtof`/`#!cpp std::strtod`/`#!cpp std::strtold` and `#!cpp std::snprintf`, so the type must be
`#!cpp float`, `#!cpp double`, or `#!cpp long double`. The
[binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`,
because they have no encoding for `#!cpp long double`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype).
## Notes
#### Default type
@@ -20,6 +20,13 @@ used.
To store integer numbers in C++, a type is defined by the template parameter `NumberIntegerType` which chooses the type
to use.
## Template parameters
`NumberIntegerType`
: the type to store signed integers. It must be a **signed integral** type (`#!cpp std::is_integral`) with a
`#!cpp std::numeric_limits` specialization, and it is stored directly inside a `basic_json` value. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberintegertype-and-numberunsignedtype).
## Notes
#### Default type
@@ -20,6 +20,14 @@ used.
To store unsigned integer numbers in C++, a type is defined by the template parameter `NumberUnsignedType` which chooses
the type to use.
## Template parameters
`NumberUnsignedType`
: the type to store unsigned integers. It must be an **unsigned integral** type (`#!cpp std::is_integral`) with a
`#!cpp std::numeric_limits` specialization, and it must be able to represent the absolute value of every
[`number_integer_t`](number_integer_t.md) value. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberintegertype-and-numberunsignedtype).
## Notes
#### Default type
@@ -30,3 +30,5 @@ and [`default_object_comparator_t`](default_object_comparator_t.md) otherwise.
- Added in version 3.0.0.
- Changed to be conditionally defined as `#!cpp typename object_t::key_compare` or `default_object_comparator_t` in
version 3.11.0.
- Fixed the fallback to `default_object_comparator_t`, which previously failed to compile for object types without a
`key_compare` member type, in version 3.13.0.
+6 -1
View File
@@ -18,7 +18,11 @@ To store objects in C++, a type is defined by the template parameters described
## Template parameters
`ObjectType`
: the container to store objects (e.g., `std::map` or `std::unordered_map`)
: the container to store objects. Its template parameters must have the same order and meaning as those of
`std::map`; in particular, the third parameter is a comparator. `#!cpp std::unordered_map`, whose third parameter
is a hash function, therefore needs an adapter -- see
[Template Parameter Requirements](../../features/types/template_parameters.md#objecttype) for the full list of
requirements, an adapter example, and the containers that are known to work.
`StringType`
: the type of the keys or names (e.g., `std::string`). The comparison function `std::less<StringType>` is used to
@@ -122,3 +126,4 @@ the object is silently converted as an array of key-value pairs, which is incorr
## Version history
- Added in version 1.0.0.
- Allowed object types whose `erase(iterator)` returns `#!cpp void` in version 3.13.0.
+8
View File
@@ -103,6 +103,10 @@ A UTF-8 byte order mark is silently ignored.
Invalid Unicode escapes and unpaired surrogates in the input are reported as
[`parse_error.101`](../../home/exceptions.md#jsonexceptionparse_error101) with a detailed message.
By default, a `'\0'` (NUL) byte anywhere in the input is treated as end of input, rather than as an ordinary (and,
outside of a string, invalid) byte; see the [FAQ entry](../../home/faq.md#nul-bytes-in-the-input) for details and the
[`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) macro to opt into rejecting it instead.
## Examples
??? example "Parsing from a character array"
@@ -236,6 +240,8 @@ Invalid Unicode escapes and unpaired surrogates in the input are reported as
- [accept](accept.md) - check if the input is valid JSON
- [sax_parse](sax_parse.md) - parse input using the SAX interface
- [operator>>](../operator_gtgt.md) - deserialize from stream
- [`JSON_STRICT_NUL_HANDLING`](../macros/json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input
instead of treating it as end of input
## Version history
@@ -246,6 +252,8 @@ Invalid Unicode escapes and unpaired surrogates in the input are reported as
- Added `ignore_trailing_commas` in version 3.13.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating
it as end of input; planned to become the default in version 4.0.0.
!!! warning "Deprecation"
+8
View File
@@ -34,6 +34,10 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
("add", "remove", "move")
- Throws [`out_of_range.411`](../../home/exceptions.md#jsonexceptionout_of_range411) if an "add" operation's target
location has a parent that is neither an object nor an array.
- Throws [`out_of_range.413`](../../home/exceptions.md#jsonexceptionout_of_range413) if a "remove" operation's target
location has a parent that is neither an object nor an array.
- Throws [`out_of_range.414`](../../home/exceptions.md#jsonexceptionout_of_range414) if a "move" operation's "from"
location is a proper prefix of its "path" location.
- Throws [`other_error.501`](../../home/exceptions.md#jsonexceptionother_error501) if "test" operation was
unsuccessful.
@@ -75,3 +79,7 @@ is thrown. In any case, the original value is not changed: the patch is applied
- Added in version 2.0.0.
- Added [`out_of_range.411`](../../home/exceptions.md#jsonexceptionout_of_range411) and stopped relying on an internal assertion when an "add" operation's
target location has a non-object/non-array parent in version 3.13.0.
- Added [`out_of_range.413`](../../home/exceptions.md#jsonexceptionout_of_range413) and stopped silently ignoring a "remove" operation whose target
location has a non-object/non-array parent in version 3.13.0.
- Added [`out_of_range.414`](../../home/exceptions.md#jsonexceptionout_of_range414) and rejected a "move" operation whose "from" location is a proper
prefix of its "path" location instead of silently producing a corrupted result in version 3.13.0.
@@ -30,6 +30,10 @@ No guarantees, value may be corrupted by an unsuccessful patch operation.
("add", "remove", "move")
- Throws [`out_of_range.411`](../../home/exceptions.md#jsonexceptionout_of_range411) if an "add" operation's target
location has a parent that is neither an object nor an array.
- Throws [`out_of_range.413`](../../home/exceptions.md#jsonexceptionout_of_range413) if a "remove" operation's target
location has a parent that is neither an object nor an array.
- Throws [`out_of_range.414`](../../home/exceptions.md#jsonexceptionout_of_range414) if a "move" operation's "from"
location is a proper prefix of its "path" location.
- Throws [`other_error.501`](../../home/exceptions.md#jsonexceptionother_error501) if "test" operation was
unsuccessful.
@@ -72,3 +76,7 @@ function throws an exception.
- Added in version 3.11.0.
- Added [`out_of_range.411`](../../home/exceptions.md#jsonexceptionout_of_range411) and stopped relying on an internal assertion when an "add" operation's
target location has a non-object/non-array parent in version 3.13.0.
- Added [`out_of_range.413`](../../home/exceptions.md#jsonexceptionout_of_range413) and stopped silently ignoring a "remove" operation whose target
location has a non-object/non-array parent in version 3.13.0.
- Added [`out_of_range.414`](../../home/exceptions.md#jsonexceptionout_of_range414) and rejected a "move" operation whose "from" location is a proper
prefix of its "path" location instead of silently producing a corrupted result in version 3.13.0.
@@ -23,6 +23,11 @@ JSON class into byte-sized characters during deserialization.
`StringType`. To work with wide-character data, convert it to/from UTF-8 at the boundary instead -- see the
FAQ's [wide string handling](../../home/faq.md#wide-string-handling) section for a conversion recipe.
Beyond the character type, the library expects a substantial part of the `#!cpp std::string` interface (contiguous
null-terminated `data()`, `substr()`, `find()`, `append()`, ...). See
[Template Parameter Requirements](../../features/types/template_parameters.md#stringtype) for the full list and
for the string types that are known to work.
## Notes
#### Default type
@@ -78,3 +83,5 @@ and an example.
## Version history
- Added in version 1.0.0.
- Removed the requirement that `string_t` be implicitly convertible from `#!cpp std::string`, which the BSON writer and
the UBJSON reader relied on, in version 3.13.0.
+6 -2
View File
@@ -34,10 +34,14 @@ void swap(typename binary_t::container_type& other);
```
1. Exchanges the contents of the JSON value with those of `other`. Does not invoke any move, copy, or swap operations on
individual elements. All iterators and references remain valid. The past-the-end iterator is invalidated.
individual elements. All iterators and references remain valid. The past-the-end iterator is invalidated. If macro
[`JSON_DIAGNOSTIC_POSITIONS`](../macros/json_diagnostic_positions.md) is defined to `#!cpp 1`, the
[`start_pos()`](start_pos.md)/[`end_pos()`](end_pos.md) diagnostic positions are exchanged along with the value.
2. Exchanges the contents of the JSON value from `left` with those of `right`. Does not invoke any move, copy, or swap
operations on individual elements. All iterators and references remain valid. The past-the-end iterator is
invalidated. Implemented as a friend function callable via ADL.
invalidated. Implemented as a friend function callable via ADL. If macro
[`JSON_DIAGNOSTIC_POSITIONS`](../macros/json_diagnostic_positions.md) is defined to `#!cpp 1`, the
[`start_pos()`](start_pos.md)/[`end_pos()`](end_pos.md) diagnostic positions are exchanged along with the value.
3. Exchanges the contents of a JSON array with those of `other`. Does not invoke any move, copy, or swap operations on
individual elements. All iterators and references remain valid. The past-the-end iterator is invalidated.
4. Exchanges the contents of a JSON object with those of `other`. Does not invoke any move, copy, or swap operations on
+9 -1
View File
@@ -37,7 +37,14 @@ Linear in the size of the JSON value.
## Notes
Empty objects and arrays are flattened by [`flatten()`](flatten.md) to `#!json null` values and cannot unflattened to
their original type. Apart from this example, for a JSON value `j`, the following is always true:
their original type.
A flattened array and a flattened object whose keys are array indices are indistinguishable, because both are
described by the same JSON pointers. A value is therefore restored as an array if and only if one of its keys is the
reference token `0`, and as an object otherwise: `#!json {"2": 1}` is restored unchanged, whereas `#!json {"0": 1}` is
restored as `#!json [1]`. This decision does not depend on the order in which the flattened object is iterated.
Apart from these two cases, for a JSON value `j`, the following is always true:
`#!cpp j == j.flatten().unflatten()`.
## Examples
@@ -63,3 +70,4 @@ their original type. Apart from this example, for a JSON value `j`, the followin
## Version history
- Added in version 2.0.0.
- Made the array/object decision independent of the object's iteration order in version 3.13.0.
+6
View File
@@ -14,6 +14,11 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_DIAGNOSTIC_POSITIONS**](json_diagnostic_positions.md) - access positions of elements
- [**JSON_NOEXCEPTION**](json_noexception.md) - switch off exceptions
## Parsing
- [**JSON_STRICT_NUL_HANDLING**](json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input instead of
treating it as end of input
## Language support
- [**JSON_HAS_CPP_11**<br>**JSON_HAS_CPP_14**<br>**JSON_HAS_CPP_17**<br>**JSON_HAS_CPP_20**](json_has_cpp_11.md) - set supported C++ standard
@@ -24,6 +29,7 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
- [**JSON_USE_SIMDUTF**](json_use_simdutf.md) - use the simdutf library to accelerate UTF-8 validation
## Library version
@@ -0,0 +1,126 @@
# JSON_STRICT_NUL_HANDLING
```cpp
#define JSON_STRICT_NUL_HANDLING /* value */
```
When defined to `1`, a `'\0'` (NUL) byte in JSON text input is rejected with `parse_error.101`, like any other
unexpected byte, instead of being silently treated as end of input.
The macro only affects the JSON text parser ([`parse`](../basic_json/parse.md), [`accept`](../basic_json/accept.md),
[`sax_parse`](../basic_json/sax_parse.md), and [`operator>>`](../operator_gtgt.md)). There are three cases where a NUL
byte is still not rejected:
- The binary formats ([`from_bjdata`](../basic_json/from_bjdata.md), [`from_bson`](../basic_json/from_bson.md),
[`from_cbor`](../basic_json/from_cbor.md), [`from_msgpack`](../basic_json/from_msgpack.md),
[`from_ubjson`](../basic_json/from_ubjson.md)) are never affected: there, `0x00` is ordinary data.
- A bare `const char*` pointer has no length of its own, so its length is still determined with `strlen()`. The first
NUL byte therefore still marks the end of the input, and nothing after it is read.
- One trailing `'\0'` at the end of a `char` array (e.g., a string literal) is trimmed; see the warning below.
## Default definition
The default value is `0` (disabled — existing behavior is preserved).
```cpp
#define JSON_STRICT_NUL_HANDLING 0
```
## Notes
!!! note "Background"
By default, a `'\0'` byte anywhere in the input is treated the same as the real end of the input, rather than as
an ordinary (and, outside of a string, invalid) byte. Everything from that byte onward is silently ignored,
without a parse error - including further, otherwise well-formed JSON:
```cpp
json::parse(std::string("123") + '\0'); // == 123, no error
json::parse(std::string("123") + '\0' + "true"); // == 123, the "true" is silently ignored too
```
This falls out of the same convention used when no explicit input length is given at all: parsing from a
`const char*` already stops at the first NUL byte via `strlen()`, since a bare pointer has no length of its own.
The library applies that same NUL-terminated-C-string convention uniformly, rather than only when a length is
genuinely unavailable - so a `std::string`, iterator range, or container whose content happens to include a NUL
byte is affected the same way a raw `const char*` would be (see the
[FAQ entry](../../home/faq.md#nul-bytes-in-the-input) for a fuller explanation).
This was not fixed unconditionally, because doing so is backwards-incompatible for any caller who happens to
depend on the current behavior - even unknowingly, for instance because their input already contains trailing
padding they never noticed was being discarded (see [#5530](https://github.com/nlohmann/json/issues/5530)).
This macro instead offers an opt-in path to the corrected behavior ahead of version 4.0.0, where it is planned to
become the default.
!!! warning "Opt-in only"
This macro must be defined **before** including `<nlohmann/json.hpp>`. Defining it after the include has no
effect.
Enabling it also changes how a `char` array (including a string literal, e.g. `json::parse("123")`) is read: such
an array normally carries a trailing `'\0'` contributed by the compiler, not by the source text. With this macro
enabled, that one trailing byte is trimmed if present so that parsing a string literal keeps working; every other
byte in the array - including any `'\0'` that is not the very last element - is read as real data and rejected
like any other unexpected byte. Arrays of any other element type (`unsigned char`, `std::uint8_t`, ...), as used
for CBOR or MessagePack, are never affected by this trimming; their full extent - including a genuine trailing
`0x00` - is always preserved, in both states of this macro.
!!! tip "Workaround without the macro"
To reject a NUL byte without enabling this macro, trim your input yourself before calling `parse()`:
```cpp
s.resize(s.find('\0')); // drop everything from the first NUL onward, if any
json::parse(s);
```
## Examples
??? example "Default behavior (macro not defined)"
Without the macro, a NUL byte silently ends parsing at that point:
```cpp
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
json j = json::parse(std::string("123") + '\0' + "true");
// j is 123 -- the '\0' and everything after it is silently ignored
}
```
??? example "Opt-in strict handling (macro defined to 1)"
With the macro, a NUL byte is rejected like any other unexpected byte:
```cpp
#define JSON_STRICT_NUL_HANDLING 1
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
json j = json::parse(std::string("123") + '\0' + "true");
// throws parse_error.101 -- the NUL byte is now invalid input,
// exactly like any other unexpected trailing byte
json ok = json::parse("123");
// ok is 123 -- parsing from a string literal still works
}
```
## See also
- [FAQ: NUL bytes in the input](../../home/faq.md#nul-bytes-in-the-input)
- [**parse**](../basic_json/parse.md) - deserialize from a compatible input
- [**accept**](../basic_json/accept.md) - check if the input is valid JSON
- [**operator>>**](../operator_gtgt.md) - deserialize from stream
## Version history
- Added in version 3.13.0.
- Planned to become the default (with the macro removed) in version 4.0.0.
@@ -12,9 +12,11 @@
Controls how exceptions are handled by the library.
1. This macro overrides [`#!cpp catch`](https://en.cppreference.com/w/cpp/language/try_catch) calls inside the library.
The argument is the type of the exception to catch. As of version 3.8.0, the library only catches `std::out_of_range`
exceptions internally to rethrow them as [`json::out_of_range`](../../home/exceptions.md#out-of-range) exceptions.
The macro is always followed by a scope.
The argument is the type of the exception to catch. The library uses it in a single place: to swallow any exception
escaping the parent-pointer check that [`JSON_DIAGNOSTICS`](json_diagnostics.md) adds to the class invariant. The
places where the library catches its own [`json::out_of_range`](../../home/exceptions.md#out-of-range) exceptions
use `JSON_INTERNAL_CATCH` instead, which `JSON_CATCH_USER` also overrides unless `JSON_INTERNAL_CATCH_USER` is
defined. The macro is always followed by a scope.
2. This macro overrides `#!cpp throw` calls inside the library. The argument is the exception to be thrown. Note that
`JSON_THROW_USER` should leave the current scope (e.g., by throwing or aborting), as continuing after it may yield
undefined behavior.
@@ -0,0 +1,71 @@
# JSON_USE_SIMDUTF
```cpp
#define JSON_USE_SIMDUTF
```
When defined, the parser validates the UTF-8 content of JSON strings that come from a **contiguous byte input**
(`std::string`, `std::vector<char>`/`<std::uint8_t>`, string literals, `const char*` ranges, …) using the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. On text with many
non-ASCII characters (e.g. CJK or emoji) this can validate several times faster.
This is an **opt-in external dependency**. The library itself remains header-only and its behavior is unchanged: the
same input is accepted or rejected either way, and every parse error is reported at the same position with the same
message (simdutf is only used to fast-path *valid* runs; anything it flags falls back to the scalar path so the exact
diagnostic is preserved). Streaming inputs (files, `std::istream`, wide strings, user-defined adapters) always use the
scalar path.
When `JSON_USE_SIMDUTF` is defined you must make the `simdutf.h` header available on the include path and link the
simdutf library. When it is not defined, no simdutf header is included and there is no dependency.
!!! note "Requires C++17"
simdutf requires C++17 and its header rejects older standards with an `#!cpp #error`. The backend is therefore only
compiled in from C++17 on. In C++11 and C++14 the macro has no effect and the scalar validator is used, which
accepts and rejects exactly the same input -- only throughput differs. Setting the macro project-wide is therefore
safe even when some translation units are built with an older standard.
!!! warning "Define consistently"
The macro selects between two definitions of the same inline validation function. It must therefore be defined
identically for **every** translation unit that includes the library; mixing translation units that define it with
ones that do not is an ODR violation. Prefer setting it as a compile definition on the target rather than with
`#!cpp #define` in individual source files.
## Default definition
By default, `#!cpp JSON_USE_SIMDUTF` is not defined and the portable C++11 scalar validator is used.
```cpp
#undef JSON_USE_SIMDUTF
```
## Examples
??? example
The code below enables the simdutf backend for UTF-8 validation.
```cpp
#define JSON_USE_SIMDUTF 1
#include <nlohmann/json.hpp>
...
```
The project must also link against simdutf, e.g. with CMake:
```cmake
target_compile_definitions(your_target PRIVATE JSON_USE_SIMDUTF)
target_link_libraries(your_target PRIVATE simdutf::simdutf)
```
!!! hint "Testing this configuration"
The unit tests can be built against the simdutf backend with the CMake option `JSON_TestSimdutf` (`OFF` by
default), which fetches simdutf and defines `JSON_USE_SIMDUTF` for every test target. The `ci_test_simdutf` target
runs the whole test suite in that configuration.
## Version history
- Added in version 3.13.0.
+11
View File
@@ -72,6 +72,13 @@ input >> j2; // parses the next value
Note that reading concatenated values does **not** work for [JSON Lines](../features/parsing/json_lines.md)
(newline-delimited JSON) input -- see that page for why and for the recommended alternative.
By default, a `'\0'` (NUL) byte encountered while reading a value is treated as end of input, rather than as an
ordinary (and, outside of a string, invalid) byte; see the [FAQ entry](../home/faq.md#nul-bytes-in-the-input) for
details and the [`JSON_STRICT_NUL_HANDLING`](macros/json_strict_nul_handling.md) macro to opt into rejecting it
instead. Because `operator>>` only parses a single value and does not require the rest of the stream to be consumed,
a NUL byte *after* a complete value has no effect on `operator>>` either way; it only matters while a value is still
being read.
!!! warning "Deprecation"
This function replaces function `#!cpp std::istream& operator<<(basic_json& j, std::istream& i)` which has
@@ -98,7 +105,11 @@ Note that reading concatenated values does **not** work for [JSON Lines](../feat
- [accept](basic_json/accept.md) - check if the input is valid JSON
- [parse](basic_json/parse.md) - deserialize from a compatible input
- [`JSON_STRICT_NUL_HANDLING`](macros/json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input
instead of treating it as end of input
## Version history
- Added in version 1.0.0.
- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating
it as end of input; planned to become the default in version 4.0.0.
@@ -0,0 +1,19 @@
#include <iostream>
#include <map>
#include <nlohmann/json.hpp>
#include "custom_array_type.hpp"
using custom_json = nlohmann::basic_json<std::map, custom_array_type>;
int main()
{
custom_json j = custom_json::array();
j.push_back(1);
j.push_back(2);
j.push_back(3);
std::cout << j.dump() << std::endl;
std::cout << std::boolalpha << (custom_json::parse(j.dump()) == j) << std::endl;
}
@@ -0,0 +1,152 @@
#pragma once
#include <memory>
#include <utility>
#include <vector>
// A minimal, self-contained ArrayType built around a private std::vector.
// See https://json.nlohmann.me/features/types/template_parameters/#arraytype
template<class T, class Allocator = std::allocator<T>>
class custom_array_type
{
using vector_t = std::vector<T, Allocator>;
vector_t data_;
public:
using value_type = typename vector_t::value_type;
using size_type = typename vector_t::size_type;
using iterator = typename vector_t::iterator;
using const_iterator = typename vector_t::const_iterator;
custom_array_type() = default;
custom_array_type(const custom_array_type&) = default;
custom_array_type(custom_array_type&&) = default;
custom_array_type& operator=(const custom_array_type&) = default;
custom_array_type& operator=(custom_array_type&&) = default;
template<class InputIt>
custom_array_type(InputIt first, InputIt last) : data_(first, last) {}
custom_array_type(size_type count, const T& value) : data_(count, value) {}
iterator begin()
{
return data_.begin();
}
iterator end()
{
return data_.end();
}
const_iterator begin() const
{
return data_.begin();
}
const_iterator end() const
{
return data_.end();
}
const_iterator cbegin() const
{
return data_.cbegin();
}
const_iterator cend() const
{
return data_.cend();
}
bool empty() const
{
return data_.empty();
}
size_type size() const
{
return data_.size();
}
size_type max_size() const
{
return data_.max_size();
}
void clear()
{
data_.clear();
}
void resize(size_type n)
{
data_.resize(n);
}
T& operator[](size_type pos)
{
return data_[pos];
}
const T& operator[](size_type pos) const
{
return data_[pos];
}
T& back()
{
return data_.back();
}
const T& back() const
{
return data_.back();
}
void push_back(const T& value)
{
data_.push_back(value);
}
void push_back(T&& value)
{
data_.push_back(std::move(value));
}
template<class... Args>
void emplace_back(Args&& ... args)
{
data_.emplace_back(std::forward<Args>(args)...);
}
void pop_back()
{
data_.pop_back();
}
iterator insert(const_iterator pos, const T& value)
{
return data_.insert(pos, value);
}
iterator insert(const_iterator pos, size_type count, const T& value)
{
return data_.insert(pos, count, value);
}
template<class InputIt>
iterator insert(const_iterator pos, InputIt first, InputIt last)
{
return data_.insert(pos, first, last);
}
iterator erase(const_iterator pos)
{
return data_.erase(pos);
}
iterator erase(const_iterator first, const_iterator last)
{
return data_.erase(first, last);
}
void swap(custom_array_type& other)
{
data_.swap(other.data_);
}
friend bool operator==(const custom_array_type& lhs, const custom_array_type& rhs)
{
return lhs.data_ == rhs.data_;
}
friend bool operator<(const custom_array_type& lhs, const custom_array_type& rhs)
{
return lhs.data_ < rhs.data_;
}
};
@@ -0,0 +1,2 @@
[1,2,3]
true
@@ -0,0 +1,21 @@
#include <cstdint>
#include <iostream>
#include <map>
#include <string>
#include <vector>
#include <nlohmann/json.hpp>
#include "custom_binary_type.hpp"
using custom_json = nlohmann::basic_json<std::map, std::vector, std::string, bool,
std::int64_t, std::uint64_t, double, std::allocator,
nlohmann::adl_serializer, custom_binary_type>;
int main()
{
const auto j = custom_json::binary({0x01, 0x02, 0x03});
std::cout << j.dump() << std::endl;
std::cout << std::boolalpha << (custom_json::from_cbor(custom_json::to_cbor(j)) == j) << std::endl;
}
@@ -0,0 +1,112 @@
#pragma once
#include <cstdint>
#include <initializer_list>
#include <vector>
// A minimal, self-contained BinaryType built around a private std::vector.
// See https://json.nlohmann.me/features/types/template_parameters/#binarytype
class custom_binary_type
{
using vector_t = std::vector<std::uint8_t>;
vector_t data_;
public:
using value_type = vector_t::value_type;
using size_type = vector_t::size_type;
using iterator = vector_t::iterator;
using const_iterator = vector_t::const_iterator;
custom_binary_type() = default;
custom_binary_type(const custom_binary_type&) = default;
custom_binary_type(custom_binary_type&&) = default;
custom_binary_type& operator=(const custom_binary_type&) = default;
custom_binary_type& operator=(custom_binary_type&&) = default;
template<class InputIt>
custom_binary_type(InputIt first, InputIt last) : data_(first, last) {}
// so basic_json::binary({0x01, 0x02}) can build one directly
custom_binary_type(std::initializer_list<std::uint8_t> init) : data_(init) {}
size_type size() const
{
return data_.size();
}
bool empty() const
{
return data_.empty();
}
void clear()
{
data_.clear();
}
void resize(size_type n)
{
data_.resize(n);
}
// read-only is enough: the writers only ever read from a binary value
const std::uint8_t* data() const
{
return data_.data();
}
std::uint8_t& operator[](size_type pos)
{
return data_[pos];
}
std::uint8_t operator[](size_type pos) const
{
return data_[pos];
}
std::uint8_t& back()
{
return data_.back();
}
std::uint8_t back() const
{
return data_.back();
}
iterator begin()
{
return data_.begin();
}
iterator end()
{
return data_.end();
}
const_iterator begin() const
{
return data_.begin();
}
const_iterator end() const
{
return data_.end();
}
const_iterator cbegin() const
{
return data_.cbegin();
}
const_iterator cend() const
{
return data_.cend();
}
template<class InputIt>
iterator insert(const_iterator pos, InputIt first, InputIt last)
{
return data_.insert(pos, first, last);
}
friend bool operator==(const custom_binary_type& lhs, const custom_binary_type& rhs)
{
return lhs.data_ == rhs.data_;
}
friend bool operator<(const custom_binary_type& lhs, const custom_binary_type& rhs)
{
return lhs.data_ < rhs.data_;
}
};
@@ -0,0 +1,2 @@
{"bytes":[1,2,3],"subtype":null}
true
@@ -0,0 +1,26 @@
#include <iostream>
#include <type_traits>
#include <vector>
#include <nlohmann/json.hpp>
#include "custom_object_type.hpp"
using custom_json = nlohmann::basic_json<custom_object_type, std::vector>;
int main()
{
custom_json j;
j["pi"] = 3.141;
j["happy"] = true;
j["list"] = {1, 2, 3};
std::cout << j.dump(2) << std::endl;
std::cout << std::boolalpha << (custom_json::parse(j.dump()) == j) << std::endl;
// custom_object_type has no key_compare member, so object_comparator_t
// falls back to its default
std::cout << std::boolalpha
<< std::is_same<custom_json::object_comparator_t, custom_json::default_object_comparator_t>::value
<< std::endl;
}
@@ -0,0 +1,144 @@
#pragma once
#include <map>
#include <utility>
// A minimal, self-contained ObjectType built around a private std::map.
// key_compare is deliberately not exposed: when an ObjectType has no
// key_compare member, the library falls back to its own default comparator.
// See https://json.nlohmann.me/features/types/template_parameters/#objecttype
template<class Key, class T, class Compare, class Allocator>
class custom_object_type
{
using map_t = std::map<Key, T, Compare, Allocator>;
map_t data_;
public:
using key_type = typename map_t::key_type;
using mapped_type = typename map_t::mapped_type;
using value_type = typename map_t::value_type;
using size_type = typename map_t::size_type;
using iterator = typename map_t::iterator;
using const_iterator = typename map_t::const_iterator;
custom_object_type() = default;
custom_object_type(const custom_object_type&) = default;
custom_object_type(custom_object_type&&) = default;
custom_object_type& operator=(const custom_object_type&) = default;
custom_object_type& operator=(custom_object_type&&) = default;
template<class InputIt>
custom_object_type(InputIt first, InputIt last) : data_(first, last) {}
iterator begin()
{
return data_.begin();
}
iterator end()
{
return data_.end();
}
const_iterator begin() const
{
return data_.begin();
}
const_iterator end() const
{
return data_.end();
}
const_iterator cbegin() const
{
return data_.cbegin();
}
const_iterator cend() const
{
return data_.cend();
}
bool empty() const
{
return data_.empty();
}
size_type size() const
{
return data_.size();
}
size_type max_size() const
{
return data_.max_size();
}
void clear()
{
data_.clear();
}
iterator find(const key_type& key)
{
return data_.find(key);
}
const_iterator find(const key_type& key) const
{
return data_.find(key);
}
size_type count(const key_type& key) const
{
return data_.count(key);
}
std::pair<iterator, bool> emplace(const key_type& key, const mapped_type& value)
{
return data_.emplace(key, value);
}
std::pair<iterator, bool> insert(const value_type& value)
{
return data_.insert(value);
}
template<class InputIt>
void insert(InputIt first, InputIt last)
{
data_.insert(first, last);
}
mapped_type& operator[](const key_type& key)
{
return data_[key];
}
mapped_type& at(const key_type& key)
{
return data_.at(key);
}
const mapped_type& at(const key_type& key) const
{
return data_.at(key);
}
iterator erase(iterator pos)
{
return data_.erase(pos);
}
iterator erase(iterator first, iterator last)
{
return data_.erase(first, last);
}
size_type erase(const key_type& key)
{
return data_.erase(key);
}
void swap(custom_object_type& other)
{
data_.swap(other.data_);
}
friend bool operator==(const custom_object_type& lhs, const custom_object_type& rhs)
{
return lhs.data_ == rhs.data_;
}
friend bool operator<(const custom_object_type& lhs, const custom_object_type& rhs)
{
return lhs.data_ < rhs.data_;
}
};
@@ -0,0 +1,11 @@
{
"happy": true,
"list": [
1,
2,
3
],
"pi": 3.141
}
true
true
@@ -0,0 +1,20 @@
#include <iostream>
#include <map>
#include <vector>
#include <nlohmann/json.hpp>
#include "custom_string_type.hpp"
using custom_json = nlohmann::basic_json<std::map, std::vector, custom_string_type>;
int main()
{
custom_json j;
j["pi"] = 3.141;
j["happy"] = true;
j["list"] = {1, 2, 3};
std::cout << j.dump(2) << std::endl;
std::cout << std::boolalpha << (custom_json::parse(j.dump()) == j) << std::endl;
}
@@ -0,0 +1,134 @@
#pragma once
#include <ostream>
#include <string>
// A minimal, self-contained StringType built around a private std::string.
// Wraps rather than inherits, so it exposes exactly what the library needs
// and nothing more of std::string's interface.
//
// Covers the "Always required" members, the extras needed for the binary
// formats, and the extras needed for JSON Pointer / flatten / unflatten /
// diff. Extending it further (e.g. for std::hash<basic_json> or to_bson) is
// a matter of adding the extra members listed in the "Required for other
// functionality" table.
//
// See https://json.nlohmann.me/features/types/template_parameters/#stringtype
class custom_string_type
{
std::string data_;
public:
using value_type = char;
using size_type = std::string::size_type;
using iterator = std::string::iterator;
using const_iterator = std::string::const_iterator;
static constexpr size_type npos = std::string::npos;
custom_string_type() = default;
custom_string_type(const custom_string_type&) = default;
custom_string_type(custom_string_type&&) = default;
custom_string_type& operator=(const custom_string_type&) = default;
custom_string_type& operator=(custom_string_type&&) = default;
// not explicit: the library relies on being able to hand it a string literal
custom_string_type(const char* s) : data_(s) {}
custom_string_type(const char* s, size_type count) : data_(s, count) {}
custom_string_type(size_type count, char ch) : data_(count, ch) {}
size_type size() const
{
return data_.size();
}
bool empty() const
{
return data_.empty();
}
void clear()
{
data_.clear();
}
void resize(size_type n)
{
data_.resize(n);
}
void resize(size_type n, char c)
{
data_.resize(n, c);
}
void reserve(size_type n)
{
data_.reserve(n);
}
// must stay null-terminated -- the parser hands this to std::strtoull &
// friends; std::string::data() has guaranteed that since C++11
const char* data() const
{
return data_.data();
}
void push_back(char c)
{
data_.push_back(c);
}
char& operator[](size_type pos)
{
return data_[pos];
}
char operator[](size_type pos) const
{
return data_[pos];
}
custom_string_type& append(const char* s, size_type count)
{
data_.append(s, count);
return *this;
}
custom_string_type& append(const custom_string_type& other)
{
data_.append(other.data_);
return *this;
}
size_type find_first_of(char c, size_type pos = 0) const
{
return data_.find_first_of(c, pos);
}
iterator begin()
{
return data_.begin();
}
iterator end()
{
return data_.end();
}
const_iterator begin() const
{
return data_.begin();
}
const_iterator end() const
{
return data_.end();
}
friend bool operator==(const custom_string_type& lhs, const custom_string_type& rhs)
{
return lhs.data_ == rhs.data_;
}
friend bool operator<(const custom_string_type& lhs, const custom_string_type& rhs)
{
return lhs.data_ < rhs.data_;
}
// not required by the library itself, but dump() returns a custom_string_type
// and this makes `std::cout << j.dump()` work as expected
friend std::ostream& operator<<(std::ostream& os, const custom_string_type& s)
{
return os << s.data_;
}
};
@@ -0,0 +1,10 @@
{
"happy": true,
"list": [
1,
2,
3
],
"pi": 3.141
}
true
@@ -4,8 +4,8 @@
Hello, world!
1 2 3 4 5
string: "Hello, world!"
number: {"floating-point":17.23,"integer":42}
null: null
string: "Hello, world!"
boolean: true
array: [1,2,3,4,5]
+1 -1
View File
@@ -4,8 +4,8 @@
Hello, world!
1 2 3 4 5
string: "Hello, world!"
number: {"floating-point":17.23,"integer":42}
null: null
string: "Hello, world!"
boolean: true
array: [1,2,3,4,5]
@@ -4,9 +4,9 @@
Hello, world!
1 2 3 4 5
string: "Hello, world!"
number: {"floating-point":17.23,"integer":42}
null: null
string: "Hello, world!"
boolean: true
array: [1,2,3,4,5]
[json.exception.type_error.302] type must be boolean, but is string
@@ -69,6 +69,13 @@ The library uses the following mapping from JSON values types to UBJSON types ac
Note that `use_size = true` alone may result in larger representations - the benefit of this parameter is that the
receiving side is immediately informed on the number of elements of the container.
An array whose type marker is `Z` (null), `T` (true) or `F` (false) stores no payload at all, because the marker
already is the value. Its declared count is therefore the only thing that decides how much memory the receiving side
allocates, and a handful of bytes can describe billions of elements. `from_ubjson` rejects such an array with
[`out_of_range.408`](../../home/exceptions.md#jsonexceptionout_of_range408) when the count exceeds 1,048,576
(`1 << 20`), and `to_ubjson` writes longer arrays of these types without the annotation, so any value it produces
can be read back.
!!! info "Binary values"
If the JSON data contains the binary type, the value stored is a list of integers, as suggested by the UBJSON
+21
View File
@@ -105,6 +105,19 @@ using the library with compilers that do not fully support C++11 and may only wo
See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](../api/macros/json_skip_unsupported_compiler_check.md).
## `JSON_STRICT_NUL_HANDLING`
When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other
unexpected byte, instead of being silently treated as end of input (see the
[FAQ entry](../home/faq.md#nul-bytes-in-the-input) for background). The default value is `0`, which preserves the
existing behavior; this is planned to become the default in version 4.0.0.
The strict handling can also be enabled with the CMake option
[`JSON_StrictNulHandling`](../integration/cmake.md#json_strictnulhandling) (`OFF` by default) which sets
`JSON_STRICT_NUL_HANDLING` accordingly.
See [full documentation of `JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md).
## `JSON_THROW_USER(exception)`
This macro overrides `#!cpp throw` calls inside the library. The argument is the exception to be thrown.
@@ -137,6 +150,14 @@ behavior is deprecated and switched off (`0`) by default.
See [full documentation of `JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`](../api/macros/json_use_legacy_discarded_value_comparison.md).
## `JSON_USE_SIMDUTF`
When defined, UTF-8 validation of JSON strings read from contiguous byte input is delegated to the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. This is an opt-in
external dependency and is not defined by default.
See [full documentation of `JSON_USE_SIMDUTF`](../api/macros/json_use_simdutf.md).
## `NLOHMANN_DEFINE_TYPE_*(...)`, `NLOHMANN_DEFINE_DERIVED_TYPE_*(...)`
The library defines 12 macros to simplify the serialization/deserialization of types. See the page on
+5 -1
View File
@@ -51,7 +51,11 @@ If you do want to preserve the **insertion order**, you can use the type [`nlohm
--8<-- "examples/ordered_json.output"
```
Alternatively, you can use a more sophisticated ordered map like [`tsl::ordered_map`](https://github.com/Tessil/ordered-map) ([integration](https://github.com/nlohmann/json/issues/546#issuecomment-304447518)) or [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) ([integration](https://github.com/nlohmann/json/issues/485#issuecomment-333652309)).
Alternatively, [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) also preserves the insertion order and, unlike [`ordered_map`](../api/ordered_map.md), keeps a lookup index, so it does not have the quadratic cost described below. It is used through a small adapter ([integration](https://github.com/nlohmann/json/issues/485#issuecomment-333652309)).
If the order does not matter and you only want faster lookup, `boost::unordered_flat_map`, `absl::flat_hash_map`, `absl::node_hash_map`, and several other hash maps work through an adapter that restores the template argument order `basic_json` expects; see [Template Parameter Requirements](types/template_parameters.md#objecttype). Note these are *unordered*, not insertion-ordered.
[`tsl::ordered_map`](https://github.com/Tessil/ordered-map) cannot be used: its iterators expose the mapped value as `const`, while `basic_json` needs to modify it in place.
The [`ordered_map`](../api/ordered_map.md) behind `nlohmann::ordered_json` is deliberately minimal and has no lookup
index, so every key access is a linear scan and building an object of `n` keys costs O(n²). This is unnoticeable at
+6 -1
View File
@@ -79,7 +79,8 @@ template<
class NumberFloatType = double,
template<typename U> class AllocatorType = std::allocator,
template<typename T, typename SFINAE = void> class JSONSerializer = adl_serializer,
class BinaryType = std::vector<std::uint8_t>
class BinaryType = std::vector<std::uint8_t>,
class CustomBaseClass = void
>
class basic_json;
```
@@ -106,6 +107,10 @@ using number_float_t = NumberFloatType;
using binary_t = nlohmann::byte_container_with_subtype<BinaryType>;
```
Not every type can be passed for these template arguments: the library uses the resulting types in ways that imply a
number of requirements, for instance that `StringType` is `char`-based or that `ArrayType` is vector-like. These
requirements are collected in [Template Parameter Requirements](template_parameters.md).
## Objects
@@ -0,0 +1,747 @@
# Template Parameter Requirements
Class [`basic_json`](../../api/basic_json/index.md) is configurable through eleven template parameters. The library
never formally states what a type passed for one of these parameters has to provide -- the requirements are implied by
the way the library uses the resulting [`object_t`](../../api/basic_json/object_t.md),
[`array_t`](../../api/basic_json/array_t.md), [`string_t`](../../api/basic_json/string_t.md), etc. This page collects
these requirements so they do not have to be discovered by trial and error. Each section lists the concrete types
that are known to work for that parameter and the ones that do not, checked against Boost 1.83, Abseil 20250127.0,
Folly, EASTL 3.21, `ankerl::unordered_dense`, `phmap`, `gtl`, `robin_hood`, `tsl::ordered_map`, and Qt 6.
## How to read this page
Requirements are split into two groups:
- **Always required** -- needed to instantiate `basic_json` at all, or needed by functions that virtually every program
uses (construction, element access, [`dump`](../../api/basic_json/dump.md)).
- **Required for ...** -- only needed when a particular part of the API is instantiated. Member function templates are
only instantiated when they are used, so a type may be perfectly usable even though it does not satisfy these
requirements, as long as the corresponding functions are never called.
!!! warning "Requirements are not checked"
Three requirements are checked with a `#!cpp static_assert`: the array iterator category, the width of
[`BinaryType`](#binarytype)'s `value_type`, and [`NumberUnsignedType`](#numberintegertype-and-numberunsignedtype)
being at least as wide as [`NumberIntegerType`](#numberintegertype-and-numberunsignedtype). The rest are not
diagnosed with dedicated error messages, and violating most of them results in a compiler error somewhere inside
the library. Four violations are not caught at compile time at all:
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers,
because the lexer hands the buffer to `#!cpp std::strtoull`/`#!cpp std::strtoll`/`#!cpp std::strtod`.
- A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation,
and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance.
- The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion
in a normal build, and only fail silently under `#!cpp NDEBUG`.
## Overview
| Template parameter | Default | Notable substitutes |
|-------------------------------------------------------------------|-----------------------------------|-----------------------------------------------------------------------|
| [`ObjectType`](#objecttype) | `std::map` | [`nlohmann::ordered_map`](../../api/ordered_map.md), Abseil hash maps |
| [`ArrayType`](#arraytype) | `std::vector` | `#!cpp std::deque` |
| [`StringType`](#stringtype) | `std::string` | `std::string`-like types over `char` |
| [`BooleanType`](#booleantype) | `bool` | none worth using |
| [`NumberIntegerType`](#numberintegertype-and-numberunsignedtype) | `std::int64_t` | any signed integer type |
| [`NumberUnsignedType`](#numberintegertype-and-numberunsignedtype) | `std::uint64_t` | any unsigned integer type at least as wide as `NumberIntegerType` |
| [`NumberFloatType`](#numberfloattype) | `double` | `float` (`long double`: no binary formats) |
| [`AllocatorType`](#allocatortype) | `std::allocator` | stateless allocators |
| [`JSONSerializer`](#jsonserializer) | `adl_serializer` | serializers with the same interface |
| [`BinaryType`](#binarytype) | `#!cpp std::vector<std::uint8_t>` | `#!cpp std::vector<char>` |
| [`CustomBaseClass`](#custombaseclass) | `void` | any default-constructible class |
!!! warning "Third-party containers and incomplete types"
`object_t` is instantiated inside the definition of `basic_json` -- it is probed for a `key_compare` member to
form [`object_comparator_t`](../../api/basic_json/object_comparator_t.md) -- i.e. while `basic_json` is still an
incomplete type. `#!cpp std::map` is required by the standard to support incomplete mapped types; most
third-party maps are not, and inspecting the mapped type at class scope (for instance with
`#!cpp std::is_trivially_move_assignable`) makes them unusable as `ObjectType`, no matter how their template
arguments are adapted. This rules out `absl::btree_map`, `phmap::btree_map`, `gtl::btree_map`,
`robin_hood::unordered_node_map`, `folly::F14FastMap`, and `eastl::hash_map`.
`array_t` is only *named* in the class definition and is not instantiated until `basic_json` is complete, so an
`ArrayType` that inspects its value type at class scope is generally fine -- `boost::container::small_vector` and
`static_vector` both reject incomplete value types yet work here. `absl::InlinedVector` is the exception: the
`#!cpp std::is_trivially_move_assignable<basic_json>` it evaluates while instantiating itself re-enters the
library's own trait machinery mid-instantiation.
!!! note "Folly requires C++20"
Folly's headers use `#!cpp consteval` and `#!cpp std::type_identity`, so any `basic_json` specialization that
names a Folly type has to be compiled as C++20 or later, whatever the rest of the library supports.
## `ObjectType`
`ObjectType` is instantiated as
```cpp
using object_t = ObjectType<StringType, // key_type
basic_json, // mapped_type
default_object_comparator_t, // key_compare
AllocatorType<std::pair<const StringType,
basic_json>>>; // allocator_type
```
i.e., the template arguments follow the order and meaning of `std::map`.
### Always required
- The template must be usable with **four** type arguments in the order shown above. The third argument is a
**comparator**; containers that expect something else in this position (e.g., a hash function) need an alias template
or wrapper -- see [Notes](#notes).
- An optional member type `key_compare`. If it is present it becomes
[`object_comparator_t`](../../api/basic_json/object_comparator_t.md); otherwise
[`default_object_comparator_t`](../../api/basic_json/default_object_comparator_t.md) is used.
- Member types `key_type`, `mapped_type`, `value_type`, and `iterator`.
- `value_type` must behave like `#!cpp std::pair<const key_type, mapped_type>`; the library accesses `.first` and
`.second` on it.
- `iterator` must be default-constructible and satisfy
[LegacyBidirectionalIterator](https://en.cppreference.com/w/cpp/named_req/BidirectionalIterator). The type returned
by `cbegin()`/`cend()` must satisfy the same requirements.
- Constructors: default, copy, move, and from an iterator range `(first, last)`.
- Member functions `begin()`, `end()`, `cbegin()`, `cend()`, `empty()`, `size()`, `max_size()`, `clear()`,
`find(key)`, `count(key)`, `emplace(key, value)`, `insert(value_type)`, `insert(first, last)`, `operator[](key)`,
`erase(iterator)`, and `erase(first, last)`. `erase(iterator)` may return the following iterator or `#!cpp void`;
in the latter case the library computes the successor itself, before erasing.
- `erase(key)` is **optional**: if the container does not provide one, the library falls back to `find(key)` followed
by `erase(iterator)`.
- `at(key)` is required only by [`to_ubjson`](../../api/basic_json/to_ubjson.md) and
[`to_bjdata`](../../api/basic_json/to_bjdata.md), but every container tried here provides it.
- `emplace` and `insert(value_type)` must return `#!cpp std::pair<iterator, bool>` and must have **unique-key**
semantics; multimaps cannot be used.
- The type must be swappable (via `std::swap` or an ADL `swap`).
- The comparison operators `==` and `<`; `!=`, `<=`, `>`, and `>=` are derived from them. Where the library uses
three-way comparison (C++20), `==` and `<=>` are required **instead** -- the six two-way operators do not satisfy
it. They implement [`basic_json`'s comparison operators](../../api/basic_json/operator_eq.md).
### Required for heterogeneous key lookup
The overloads of [`at`](../../api/basic_json/at.md), [`operator[]`](../../api/basic_json/operator%5B%5D.md),
[`find`](../../api/basic_json/find.md), [`contains`](../../api/basic_json/contains.md),
[`count`](../../api/basic_json/count.md), [`erase`](../../api/basic_json/erase.md), and
[`value`](../../api/basic_json/value.md) that accept a key type other than `object_t::key_type` require
- a **transparent** comparator, i.e. [`object_comparator_t`](../../api/basic_json/object_comparator_t.md) has a member
type `is_transparent` (this is why the default comparator is `#!cpp std::less<>` since C++14), and
- corresponding heterogeneous `find`, `count`, `erase`, and `operator[]` overloads on the container.
### Notes
#### `std::unordered_map` needs an adapter
`#!cpp std::unordered_map` cannot be passed directly: its third template parameter is a hash function, but
`basic_json` passes a comparator in that position. An alias template or wrapper that restores the expected argument
order makes it usable:
```cpp
template<class Key, class T, class IgnoredCompare, class Allocator>
struct unordered_map_object
: std::unordered_map<Key, T, std::hash<Key>, std::equal_to<Key>, Allocator>
{
using base_t = std::unordered_map<Key, T, std::hash<Key>, std::equal_to<Key>, Allocator>;
using base_t::base_t;
};
using unordered_json = nlohmann::basic_json<unordered_map_object>;
```
Whether `#!cpp std::unordered_map` can be instantiated at all depends on the standard library: `object_t` is formed
while `basic_json` is still incomplete (see the warning above), and libstdc++ 9 needs the size of the mapped type to
instantiate the hash map's node type, so the adapter does not compile there. Newer libstdc++ versions, and the hash
maps listed below, do not have that problem.
The adapter above works verbatim for Abseil's, Boost's, `phmap`'s and `gtl`'s hash maps, which all place the hash
function third and take a `#!cpp std::pair<const Key, T>` allocator fifth. Two need a different adapter:
- `ankerl::unordered_dense` expects an allocator over `#!cpp std::pair<Key, T>` (non-const key), so the allocator has
to be rebound to that or dropped.
- `robin_hood`'s fifth parameter is the non-type `MaxLoadFactor100`, so its adapter must drop the allocator entirely.
None of these hash maps defines `key_compare`, so all of them additionally rely on `object_comparator_t` falling back
to [`default_object_comparator_t`](../../api/basic_json/default_object_comparator_t.md); see
[`object_comparator_t`](../../api/basic_json/object_comparator_t.md).
#### Abseil hash maps
`absl::flat_hash_map` and `absl::node_hash_map` tolerate an incomplete value type, but they take a hash function as
their third template argument. The same adapter as for `#!cpp std::unordered_map` makes them usable:
```cpp
template<class Key, class T, class IgnoredCompare, class Allocator>
struct flat_hash_object
: absl::flat_hash_map<Key, T, absl::Hash<Key>, std::equal_to<Key>, Allocator>
{
using base_t = absl::flat_hash_map<Key, T, absl::Hash<Key>, std::equal_to<Key>, Allocator>;
using base_t::base_t;
};
using flat_hash_json = nlohmann::basic_json<flat_hash_object>;
```
`absl::node_hash_map` keeps references to the mapped values valid across insertions; `absl::flat_hash_map` does not,
which makes it behave like [`ordered_json`](../../api/ordered_json.md) with respect to
[iterator invalidation](../../api/basic_json/index.md#iterator-invalidation). Both expose a `capacity()` member
function, so [`JSON_DIAGNOSTICS`](../../api/macros/json_diagnostics.md) treats them conservatively and keeps the
parent pointers correct either way.
#### Iteration order
The library never relies on the container's iteration order for correctness; it does determine the order in which
object keys are serialized by [`dump`](../../api/basic_json/dump.md) and visited by
[`items`](../../api/basic_json/items.md). See [Object Order](../object_order.md).
#### `capacity()` marks a container as insertion-ordered
With [`JSON_DIAGNOSTICS`](../../api/macros/json_diagnostics.md) enabled, the library detects insertion-ordered maps by
probing for a `capacity()` member function (`nlohmann::ordered_map` inherits it from `std::vector`) and refreshes all
parent pointers after every insertion. An `ObjectType` that happens to have a `capacity()` member is therefore treated
conservatively -- this is correct, but slower.
#### Key order and duplicate keys
The library does not sort or de-duplicate keys itself; the behavior described in
[`object_t`](../../api/basic_json/object_t.md) is entirely the behavior of the chosen container.
!!! tip "Reference implementation"
`docs/mkdocs/docs/examples/custom_object_type.hpp` wraps a private `#!cpp std::map` and satisfies every
requirement above. It does not define `key_compare`, so `object_comparator_t` falls back to
[`default_object_comparator_t`](../../api/basic_json/default_object_comparator_t.md) -- a good starting point for
a custom `ObjectType`.
```cpp
--8<-- "examples/custom_object_type.hpp"
```
??? example "Compiling and using it"
```cpp
--8<-- "examples/custom_object_type.cpp"
```
Output:
```json
--8<-- "examples/custom_object_type.output"
```
### Compatible containers
| Container | Notes |
|----------------------------------------------------------------------------------|-------------------------------------------------------------------------------|
| `#!cpp std::map` (default) | |
| [`nlohmann::ordered_map`](../../api/ordered_map.md) | used by [`ordered_json`](../../api/ordered_json.md); keeps insertion order |
| [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) | keeps insertion order; adapter puts `fifo_map_compare` in the comparator slot |
| `boost::container::map`, `boost::container::flat_map` | no adapter needed |
| `#!cpp std::unordered_map` | through the adapter above; not with libstdc++ 9, see the note |
| `boost::unordered_map`, `boost::unordered_flat_map`, `boost::unordered_node_map` | through the adapter above |
| `absl::flat_hash_map`, `absl::node_hash_map` | through the adapter above; `flat_hash_map` moves mapped values on rehash |
| `phmap::flat_hash_map`, `phmap::node_hash_map`, `gtl::flat_hash_map` | through the adapter above |
| `ankerl::unordered_dense::map` and `segmented_map` | adapter must rebind or drop the allocator |
| `robin_hood::unordered_flat_map` | adapter must drop the allocator |
| `folly::F14NodeMap` | through the adapter above; requires C++20, see the note above |
| `folly::sorted_vector_map` | alias must drop the allocator, whose value type it disagrees on |
### Containers that cannot be used
| Container | Reason |
|--------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------|
| `absl::btree_map`, `phmap::btree_map`, `gtl::btree_map` | require a complete mapped type |
| `robin_hood::unordered_node_map`, `folly::F14FastMap`, `eastl::hash_map` | require a complete mapped type |
| `eastl::map` | EASTL iterators do not work with `#!cpp std::iterator_traits` |
| `tsl::ordered_map` | its iterators expose the mapped value as `#!cpp const` |
| `QMap` | no `value_type` member type |
| `QHash` | its `value_type` is the mapped type rather than a key/value pair, and its iterators dereference to the mapped value |
| `#!cpp std::multimap`, `#!cpp std::unordered_multimap` | `emplace` does not return `#!cpp std::pair<iterator, bool>` |
## `ArrayType`
`ArrayType` is instantiated as
```cpp
using array_t = ArrayType<basic_json, AllocatorType<basic_json>>;
```
### Always required
- The template must be usable with **two** type arguments (value type and allocator).
- Member types `value_type` and `iterator`.
- Constructors: default, copy, and move; and from an iterator range `(first, last)`.
- Member functions `begin()`, `end()`, `cbegin()`, `cend()`, `empty()`, `size()`, `max_size()`, `clear()`,
`operator[](size_type)`, `back()`, `push_back()`, `emplace_back()`, `pop_back()`, `resize()`,
`insert()` (single element, count, and range), `erase(pos)`, and `erase(first, last)`.
`basic_json::insert(pos, initializer_list)` goes through the range overload, so no initializer-list `insert` is
needed. `at(size_type)` is **not** required: [`basic_json::at(size_type)`](../../api/basic_json/at.md) checks the
index itself and then uses `operator[]`.
- `iterator` must be default-constructible, and it as well as the type returned by `cbegin()`/`cend()` must satisfy
[LegacyRandomAccessIterator](https://en.cppreference.com/w/cpp/named_req/RandomAccessIterator).
A `#!cpp static_assert` only checks for
[LegacyBidirectionalIterator](https://en.cppreference.com/w/cpp/named_req/BidirectionalIterator), but
[`dump`](../../api/basic_json/dump.md) (`cend() - 1`),
[`erase(idx)`](../../api/basic_json/erase.md) (`begin() + idx`), and the random-access operations of
[`basic_json::iterator`](../../api/basic_json/begin.md) require random access.
- The comparison operators, as for [`ObjectType`](#objecttype): `==` and `<`, or `==` and `<=>` under C++20.
### Required for individual functions
- A member type `value_type`, for [`to_bson`](../../api/basic_json/to_bson.md) of an array.
- A constructor from `(count, value)`, for
[`basic_json(size_type, const basic_json&)`](../../api/basic_json/basic_json.md).
- Swappability, via `#!cpp std::swap` or an ADL `swap`, for [`swap(array_t&)`](../../api/basic_json/swap.md).
!!! note "`capacity()` is optional"
With [`JSON_DIAGNOSTICS`](../../api/macros/json_diagnostics.md) enabled, the library reads `array_t::capacity()`
to find out whether adding an element reallocated the array and moved its elements, which would invalidate the
parent pointers. An array type without a `capacity()` member function is handled conservatively: the parent
pointers of all elements are refreshed after every insertion, which makes adding *n* elements cost O(*n*²). Only
diagnostics builds pay this; without them `capacity()` is never called.
!!! tip "Reference implementation"
`docs/mkdocs/docs/examples/custom_array_type.hpp` wraps a private `#!cpp std::vector` and satisfies every
requirement above -- a good starting point for a custom `ArrayType`.
```cpp
--8<-- "examples/custom_array_type.hpp"
```
??? example "Compiling and using it"
```cpp
--8<-- "examples/custom_array_type.cpp"
```
Output:
```json
--8<-- "examples/custom_array_type.output"
```
### Compatible containers
| Container | Notes |
|---------------------------------------------------------|-------------------------------------------------------------------------------------------|
| `#!cpp std::vector` (default) | |
| `#!cpp std::deque` | references survive appends, but not insertions elsewhere; see the `capacity()` note above |
| `#!cpp std::pmr::vector` | through an alias, as the allocator comes from `AllocatorType` instead |
| `boost::container::vector`, `deque`, `devector` | |
| `boost::container::stable_vector` | the only one tried that keeps references valid across *every* insertion |
| `boost::container::small_vector`, `folly::small_vector` | through an alias that fixes the inline capacity |
| `boost::container::static_vector` | through the same kind of alias, for arrays that stay within the fixed capacity |
| `folly::fbvector` | requires C++20, see the note above |
### Containers that cannot be used
| Container | Reason |
|-------------------------------------|-----------------------------------------------------------------------------------------------|
| `#!cpp std::list` | no `operator[]`, and no random-access iterators |
| `eastl::vector`, `QList`, `QVector` | no `max_size()`; they handle the incomplete value type fine |
| `absl::InlinedVector` | requires a complete value type, see the note above |
| `absl::FixedArray` | the size is fixed at construction, so `resize`, `push_back`, `insert` and `erase` are missing |
## `StringType`
`StringType` is used **both** for JSON string values and for the keys of JSON objects
(`string_t` and `object_t::key_type`).
### Always required
- A member type `value_type` that is one byte wide and `char`-compatible. The library stores and processes UTF-8
encoded `char` data and hands `data()` to `#!cpp std::strtoull`/`#!cpp std::strtoll`.
`#!cpp std::wstring`, `#!cpp std::u16string`, and `#!cpp std::u32string` are **not** valid choices; see the FAQ on
[wide string handling](../../home/faq.md#wide-string-handling).
- Constructors: default, copy, move, from `#!cpp const char*` (which must not be `#!cpp explicit`), from
`#!cpp (const char*, size_type)`, and from `#!cpp (size_type, char)`; and copy or move assignment.
- Member functions `size()`, `clear()`, `resize(n, c)`, `data()`, `push_back(char)`, and `operator[]`
(const and non-const, returning references). `c_str()` and `back()` are **not** required.
- `data()` must return a pointer to a contiguous, **null-terminated** buffer -- the parser hands it to
`#!cpp std::strtoull`. A type whose `data()` is not null-terminated does not fail to compile; it silently
misparses numbers.
- `append(const char*, size_type)`, used by [`dump`](../../api/basic_json/dump.md), and `append(const StringType&)`,
used by the CBOR reader for indefinite-length strings. The library's internal string concatenation additionally has
to append a `#!cpp char` and a `#!cpp const char*`; for each it selects between `append(arg)`, `#!cpp operator+=`,
`append(first, last)`, and `append(data, size)`.
- The comparison operator `==` against another `StringType`, and `<` for use as a key of the chosen
[`ObjectType`](#objecttype) (with the default comparator, `#!cpp std::less<>` must be able to compare two
`StringType` values, and a `StringType` with the key types used for lookup). `!=` is never applied to a
`StringType`, and `==` against `#!cpp const char*` is resolved by the implicit `#!cpp const char*` constructor.
### Required for the binary formats
- `resize(n)`, used by the readers to make room for a block of bytes.
- Non-const `operator[]`, into which the readers `#!cpp std::memcpy` those bytes. A non-`#!cpp const` `data()` would
serve just as well, but `#!cpp std::string` has only had one since C++17, and the library still supports C++11.
### Required for JSON Pointer, `flatten`, and `diff`
- A static member `npos` and the member function `find_first_of(char, size_type)` -- together with `data()`,
`reserve(n)`, and `append(const char*, size_type)` they implement the escaping and unescaping of reference tokens
described in RFC 6901. Neither `find(const StringType&, size_type)`, nor `substr(pos, count)`, nor
`replace(pos, count, const StringType&)` is required.
- `empty()`.
- `begin()` and `end()` -- used by
[`operator[](const json_pointer&)`](../../api/basic_json/operator%5B%5D.md) to decide whether a reference token
denotes an array index.
### Required for other functionality
| Functionality | Additional requirement |
|-----------------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| [`diff`](../../api/basic_json/diff.md), [`items`](../../api/basic_json/items.md), [`std::hash`](../../api/basic_json/std_hash.md) | conversion of a `#!cpp std::size_t` to `StringType`: either assignability from the result of `#!cpp std::to_string`, or an ADL overload `#!cpp void int_to_string(StringType&, std::size_t)` |
| [`std::hash<basic_json>`](../../api/basic_json/std_hash.md) | additionally a specialization of `#!cpp std::hash<StringType>` |
| [`to_bson`](../../api/basic_json/to_bson.md) | `find(value_type)` and `npos` |
| [`parse`](../../api/basic_json/parse.md) from a `string_t` | the input adapters must accept it; otherwise pass a character range |
| `#!cpp operator<<(std::ostream&, const json_pointer&)` | streamability to `#!cpp std::ostream` |
| exception messages | `data()` and `size()`, or `begin()` and `end()` |
### Compatible types
| Type | Notes |
|-----------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `#!cpp std::string` (default) | |
| `#!cpp std::basic_string` with a custom **stateless** allocator | |
| `#!cpp std::pmr::string` | see the warning below before relying on the memory resource |
| `boost::container::string` | needs a user-supplied `#!cpp std::hash` specialization (Boost provides `boost::hash` instead) |
| `folly::fbstring` | requires C++20, see the note above |
| `eastl::string` | needs a user-supplied `#!cpp std::hash` and an ADL `int_to_string` (it is not assignable from a `#!cpp std::string`); [`parse`](../../api/basic_json/parse.md) does not accept it directly -- pass a character range or a `#!cpp std::string` |
| a custom string class in a user-defined namespace | if the requirements above are met |
### Types that cannot be used
| Type | Reason |
|----------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------|
| `#!cpp std::wstring`, `#!cpp std::u16string`, `#!cpp std::u32string` | the character type is not one byte wide |
| `#!cpp std::u8string` | one byte wide, but `#!cpp char8_t` is not `#!cpp char`-compatible |
| `absl::Cord` | no `value_type`, and the storage is not contiguous |
| `QString` | no `append(const char*, size_type)`; its `QChar` is also two bytes wide, though that is never diagnosed |
!!! warning "A `std::pmr::string` mostly does not use the memory resource you choose"
`basic_json` cannot be given an allocator or a memory resource. `AllocatorType` is default-constructed at every
allocation and has to be stateless (see [`AllocatorType`](#allocatortype)), and string values the library creates
are constructed with their own default allocator. So:
- Every string the library itself produces -- from [`parse`](../../api/basic_json/parse.md), from
[`dump`](../../api/basic_json/dump.md), or by default construction -- allocates from
`#!cpp std::pmr::get_default_resource()`.
- **Copying** an arena-backed string into a value silently drops its memory resource: the copy lands on the
default resource, because `#!cpp std::pmr::polymorphic_allocator` does not propagate on copy construction.
Nothing warns about this.
- **Moving** one in does keep it, and later growth still allocates from that arena -- but it does not survive a
copy of the enclosing `basic_json`.
- Passing `#!cpp std::pmr::polymorphic_allocator` as `AllocatorType` does not work around any of this; it does
not compile.
Apart from moving a string in, the only way to redirect these allocations is the process-global
`#!cpp std::pmr::set_default_resource()`.
!!! tip "Reference implementation"
`docs/mkdocs/docs/examples/custom_string_type.hpp` wraps a private `#!cpp std::string` and satisfies every
requirement above -- a good starting point for a custom `StringType`. The unit test
`tests/src/unit-alt-string.cpp` contains a more thorough variant, `alt_string`, exercised against a larger part
of the API.
```cpp
--8<-- "examples/custom_string_type.hpp"
```
??? example "Compiling and using it"
```cpp
--8<-- "examples/custom_string_type.cpp"
```
Output:
```json
--8<-- "examples/custom_string_type.output"
```
## `BooleanType`
`boolean_t` is stored **directly** inside `basic_json`, as a member of an anonymous union.
### Always required
- A literal type that is trivially default-constructible, trivially copyable, and trivially destructible; otherwise the
union's special member functions are deleted.
- **Implicitly** convertible from `#!cpp bool` -- an `#!cpp explicit` constructor is not enough, because the
`to_json` overload for a custom `BooleanType` is constrained on `#!cpp std::is_convertible` -- and contextually
convertible to `#!cpp bool` (here an `#!cpp explicit operator bool` is fine).
- Comparison operators `==`, `!=`, `<`, `<=`, `>`, `>=` (or `<=>`).
- Convertible from and to `#!cpp bool` through the serializer, because
[`get<bool>()`](../../api/basic_json/get.md) is used internally.
There is little reason to use anything other than `#!cpp bool` here.
### Compatible types
`#!cpp bool` is the only usable choice. Another trivially copyable type that is implicitly convertible to and from
`#!cpp bool` -- `#!cpp std::uint8_t`, say -- does compile, and JSON booleans still round-trip, but the type then
serves as both `boolean_t` and an ordinary integer: `basic_json` can no longer be constructed or assigned from a
`#!cpp std::uint8_t` at all (the boolean and unsigned-integer `to_json` overloads become ambiguous), and
[`get<std::uint8_t>()`](../../api/basic_json/get.md) on a number throws
[`type_error.302`](../../home/exceptions.md#jsonexceptiontype_error302) instead of returning the value.
## `NumberIntegerType` and `NumberUnsignedType`
Both types are stored **directly** inside `basic_json`'s union.
### Always required
- `#!cpp std::is_integral` must be satisfied: `NumberIntegerType` must be a **signed** integer type,
`NumberUnsignedType` an **unsigned** integer type. Class types are not supported -- among others, the constructors
taking integer values are constrained on `#!cpp std::is_integral`.
- Trivially default-constructible, trivially copyable, and trivially destructible (union member).
- `#!cpp std::numeric_limits` must be specialized for both types.
- `NumberUnsignedType` must be able to represent the absolute value of every `NumberIntegerType` value; serialization
of negative numbers converts the value to `NumberUnsignedType`. A `#!cpp static_assert` requires it to be at least as
wide as `NumberIntegerType`, which is what that amounts to for the standard integer types.
- Both types must fit into the internal 64-character number buffer used by
[`dump`](../../api/basic_json/dump.md), which is the case for all standard integer types.
- [`std::hash<basic_json>`](../../api/basic_json/std_hash.md) additionally requires `#!cpp std::hash` specializations.
### Notes
The number types influence what the parser accepts: an integer literal that does not round-trip through the chosen type
is stored as [`number_float_t`](../../api/basic_json/number_float_t.md) instead. Choosing types narrower than 64 bits
therefore silently changes parse results rather than raising an error. See
[Number Handling](number_handling.md) for details.
### Compatible types
| Type pair | Support |
|----------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `#!cpp std::int64_t` / `#!cpp std::uint64_t` (default) | full |
| `#!cpp std::int32_t` / `#!cpp std::uint32_t`, `#!cpp long long` / `#!cpp unsigned long long` | full; narrower types change which literals the parser can represent |
| any other pair of standard signed/unsigned integer types | full |
| class types, enumerations | not usable; `#!cpp std::is_integral` must hold |
| `#!cpp bool`, or a type already used for another member of the union | not usable; `#!cpp std::is_integral<bool>` is in fact `#!cpp true`, but the `get_impl_ptr` overloads for `boolean_t`, `number_integer_t`, `number_unsigned_t` and `number_float_t` would collide |
## `NumberFloatType`
`number_float_t` is stored **directly** inside `basic_json`'s union.
### Always required
- Trivially default-constructible, trivially copyable, and trivially destructible (union member).
- `#!cpp std::numeric_limits` must be specialized; `max_digits10` is used to size the conversion.
- `#!cpp std::isfinite` must be applicable to the type.
### Required for parsing and serialization
`NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`:
- The [parser](../parsing/index.md) converts number literals with `#!cpp std::strtof`, `#!cpp std::strtod`, or
`#!cpp std::strtold`; the library provides overloads for exactly these three types.
- [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
(`#!cpp float` is promoted to `#!cpp double`).
If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary32 or binary64 number, `dump` uses the
Grisu2 algorithm, which produces the shortest representation that round-trips. Otherwise the `snprintf` fallback with
`max_digits10` digits is used.
### Required for the binary formats
`NumberFloatType` must be `#!cpp float` or `#!cpp double`. The writers for
[CBOR, MessagePack, UBJSON, BJData, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
binary32 or binary64 field and have no encoding for `#!cpp long double`.
### Compatible types
| Type | Support |
|--------------------------|-----------------------------------------------------------------------------------------------------------------------|
| `#!cpp double` (default) | full; short round-trip output through Grisu2 |
| `#!cpp float` | full; short round-trip output through Grisu2 |
| `#!cpp long double` | `dump` and `parse` only; the binary format writers do not compile, as they only handle IEEE 754 binary32 and binary64 |
| any other type | not usable |
## `AllocatorType`
`AllocatorType` is instantiated with **one** argument, for each of `object_t`, `array_t`, `string_t`, `binary_t`,
`basic_json`, and `#!cpp std::pair<const StringType, basic_json>`.
### Always required
- The template must be usable with exactly one type argument. The library instantiates `AllocatorType<T>` directly and
never uses `#!cpp std::allocator_traits<...>::rebind_alloc`.
- It must satisfy the [Allocator](https://en.cppreference.com/w/cpp/named_req/Allocator) named requirement so that
`#!cpp std::allocator_traits` can be used with it.
- It must be **default-constructible and stateless**. Objects are allocated with a default-constructed allocator and
deallocated with a *different* default-constructed allocator, and
[`get_allocator()`](../../api/basic_json/get_allocator.md) returns a default-constructed instance. Allocators
carrying state are not supported, so there is no way to tell a `basic_json` where to allocate from; see the note
under [`StringType`](#stringtype) for what that means in practice. A stateful allocator is **not diagnosed**: it
compiles and silently ignores the state.
- It must support **incomplete types**: `AllocatorType<basic_json>` is instantiated inside the definition of
`basic_json` itself.
- `#!cpp std::allocator_traits<AllocatorType<basic_json>>::pointer` becomes
[`basic_json::pointer`](../../api/basic_json/index.md#container-types), and iterators are constructed from raw
`#!cpp basic_json*` values. The `pointer` type must therefore be a plain pointer; fancy pointers are not supported.
### Compatible types
| Type | Support |
|-------------------------------------------------------------------|----------------------------------------|
| `#!cpp std::allocator` (default) | full |
| a custom stateless allocator template | full |
| stateful allocators, e.g. `#!cpp std::pmr::polymorphic_allocator` | not usable; see the requirements above |
## `JSONSerializer`
`JSONSerializer` is instantiated as `JSONSerializer<T, void>` and defaults to
[`adl_serializer`](../../api/adl_serializer/index.md).
### Always required
- The template must accept **two** type arguments. It does not have to give the second one a default -- `basic_json`
declares the parameter as `#!cpp template<typename T, typename SFINAE = void> class JSONSerializer`, so uses such as
`#!cpp JSONSerializer<T>` inside the library supply `#!cpp void` themselves. The second parameter exists so that
partial specializations can be constrained by SFINAE.
- For every type `T` that is converted **to** a JSON value, a static member function
`#!cpp static void to_json(basic_json&, T)` must exist.
- For every type `T` that is converted **from** a JSON value, either
`#!cpp static void from_json(const basic_json&, T&)` or `#!cpp static T from_json(const basic_json&)` must exist.
The latter form is required for types that are not default-constructible; see
[Arbitrary Types Conversions](../arbitrary_types.md).
- To support the [converting constructor](../../api/basic_json/basic_json.md) between different `basic_json`
specializations, `to_json` must be available for `boolean_t`, `number_integer_t`, `number_unsigned_t`,
`number_float_t`, `string_t`, `object_t`, `array_t`, and `binary_t` of the *source* specialization.
### Compatible types
| Type | Support |
|---------------------------------------------------------------------------|-------------------------------------------------------------------|
| [`nlohmann::adl_serializer`](../../api/adl_serializer/index.md) (default) | full |
| a class template deriving from `adl_serializer` | full; the usual way to change behavior while keeping the defaults |
| an unrelated template with the same interface | full, but it has to handle every type the library converts |
## `BinaryType`
`BinaryType` is not a JSON type; it is used for the byte strings of the
[binary formats](../binary_formats/index.md). It is wrapped as
```cpp
using binary_t = nlohmann::byte_container_with_subtype<BinaryType>;
```
### Always required
- A non-`final` class type -- [`byte_container_with_subtype`](../../api/byte_container_with_subtype/index.md) derives
from it publicly.
- A member type `value_type` that is **exactly one byte** wide (e.g., `#!cpp std::uint8_t`, `#!cpp char`, or
`#!cpp std::byte`). Readers and writers reinterpret the container's storage as raw bytes, so a wider `value_type` is
rejected with a `#!cpp static_assert`.
- Contiguous storage: the binary readers `#!cpp std::memcpy` into `#!cpp &binary[n]`, the writers `reinterpret_cast`
`data()`. `#!cpp data() + n` would do for the readers too, but they share one helper with
[`StringType`](#stringtype), whose non-`#!cpp const` `data()` is C++17 and later only.
- Default-constructible, copy-constructible, and move-constructible.
- Member functions `size()`, `empty()`, `data()`, `resize()`, `operator[]`, `back()`, `begin()`, `end()`, `cbegin()`,
and `cend()` with random-access iterators, and `insert(pos, first, last)`, which the CBOR reader uses to join the
chunks of an indefinite-length byte string. `push_back()` is **not** required.
- Comparison operators: `==` is used by
[`byte_container_with_subtype`](../../api/byte_container_with_subtype/index.md), the relational operators by
[`basic_json`'s comparison operators](../../api/basic_json/operator_le.md).
### Required for individual functions
- `clear()`, for [`basic_json::clear()`](../../api/basic_json/clear.md).
`max_size()`, `at()`, `reserve()`, `erase()`, `pop_back()`, and `emplace_back()` are **not** used at all.
See [`binary_t`](../../api/basic_json/binary_t.md) for how a non-default `BinaryType` changes the meaning of assigning
such a container to a `basic_json` value.
!!! tip "Reference implementation"
`docs/mkdocs/docs/examples/custom_binary_type.hpp` wraps a private `#!cpp std::vector<std::uint8_t>` and satisfies
every requirement above -- a good starting point for a custom `BinaryType`.
```cpp
--8<-- "examples/custom_binary_type.hpp"
```
??? example "Compiling and using it"
```cpp
--8<-- "examples/custom_binary_type.cpp"
```
Output:
```json
--8<-- "examples/custom_binary_type.output"
```
### Compatible containers
| Container | Notes |
|---------------------------------------------------------------------------------------------|---------------------------------------------------------------------------|
| `#!cpp std::vector<std::uint8_t>` (default) | |
| `#!cpp std::vector<char>`, `#!cpp std::vector<std::byte>` | `dump()` writes the bytes as 0..255 whichever is used |
| `boost::container::vector<std::uint8_t>`, `boost::container::small_vector<std::uint8_t, N>` | |
| `absl::InlinedVector<std::uint8_t, N>` | usable here, unlike as an `ArrayType`, because the value type is complete |
| `eastl::vector<std::uint8_t>` | usable here, unlike as an `ArrayType`, because `max_size()` is not needed |
| `folly::fbvector<std::uint8_t>` | requires C++20, see the note above |
### Containers that cannot be used
| Container | Reason |
|------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `QByteArray` | no `empty()` (it spells that `isEmpty()`); its `insert` takes an index rather than an iterator; and it converts to `string_t`, which makes `to_json` ambiguous between a string and a binary value |
| `#!cpp std::string` | `binary_t::container_type` and `string_t` would be the same type, so the two [`swap`](../../api/basic_json/swap.md) overloads collide and `basic_json` cannot be instantiated at all |
| `#!cpp std::deque<std::uint8_t>` | storage is not contiguous, so there is no `data()` |
| containers whose `value_type` is wider than one byte | see above -- accepted by the compiler, wrong at runtime |
## `CustomBaseClass`
`CustomBaseClass` is an extension point: unless it is `#!cpp void` (the default, which selects the empty
`nlohmann::json_default_base`), `basic_json` publicly derives from it.
### Always required
- A non-`final`, default-constructible class type.
- `basic_json` is copy-/move-constructible and copy-/move-assignable only if `CustomBaseClass` is.
### Notes
`basic_json` is documented to be a
[StandardLayoutType](https://en.cppreference.com/w/cpp/named_req/StandardLayoutType). Because `basic_json` has
non-static data members of its own, a `CustomBaseClass` with non-static data members forfeits this guarantee.
Note the namespace of `CustomBaseClass` becomes an associated namespace of `basic_json` for the purpose of
argument-dependent lookup.
See [`json_base_class_t`](../../api/basic_json/json_base_class_t.md) for an example.
### Compatible types
| Type | Support |
|----------------------------------------------|----------------------------------------------------------------------------|
| `#!cpp void` (default) | an empty base class is used; no effect on `basic_json` |
| any default-constructible, non-`final` class | full; see [`json_base_class_t`](../../api/basic_json/json_base_class_t.md) |
## Cross-specialization conversions
Converting a value from one `basic_json` specialization into another (see the
[converting constructor](../../api/basic_json/basic_json.md)) imposes two additional requirements that are not
diagnosed at compile time. With assertions enabled they abort on the `#!cpp JSON_ASSERT` at the end of the converting
constructor; under `#!cpp NDEBUG` they fail **silently** at runtime:
- The target `string_t` must be directly constructible from the source `string_t`. Otherwise the string is converted to
an array of character codes.
- The target `object_t::key_type` must be directly constructible from the source object's key type. Otherwise the
object is converted to an array of key/value pairs.
See [issue #3425](https://github.com/nlohmann/json/issues/3425), [`string_t`](../../api/basic_json/string_t.md), and
[`object_t`](../../api/basic_json/object_t.md).
## See also
- [Types](index.md) -- overview of how JSON values are stored
- [Number Handling](number_handling.md) -- how the number types affect parsing and serialization
- [Object Order](../object_order.md) -- using an insertion-ordered `ObjectType`
- [`basic_json`](../../api/basic_json/index.md) -- API documentation of the class template
+52
View File
@@ -868,6 +868,12 @@ The size of an array or object in a [binary format](../features/binary_formats/i
the size following `#` for [UBJSON](../features/binary_formats/ubjson.md)/[BJData](../features/binary_formats/bjdata.md),
or the encoded length for [CBOR](../features/binary_formats/cbor.md).
The exception is also thrown for a [UBJSON](../features/binary_formats/ubjson.md) array of a type that is encoded by its
marker alone (`Z`, `T` or `F`) whose declared count exceeds 1,048,576 (`1 << 20`). Such an array has no payload, so its
count alone decides how much memory is allocated, and a handful of bytes would otherwise describe billions of values.
[`to_ubjson`](../api/basic_json/to_ubjson.md) writes longer arrays of these types without the size and type annotation,
so any value it produces can still be read back.
!!! failure "Example messages"
```
@@ -879,6 +885,9 @@ or the encoded length for [CBOR](../features/binary_formats/cbor.md).
```
[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size
```
```
[json.exception.out_of_range.408] syntax error while parsing UBJSON size: excessive array size
```
### json.exception.out_of_range.409
@@ -933,6 +942,49 @@ BSON stores the length of documents, arrays, strings, and binary values in a sig
[`to_bson`](../api/basic_json/to_bson.md) produced documents with negative length prefixes that
[`from_bson`](../api/basic_json/from_bson.md) rejected.
### json.exception.out_of_range.413
A JSON Patch `remove` operation cannot be applied because the target location's parent is neither an object nor an array. Per [RFC 6902](https://datatracker.ietf.org/doc/html/rfc6902), a `remove` target must reference a member of an existing object or an element of an existing array; a primitive value (string, number, boolean, etc.) or `null` has no members or elements to remove.
!!! failure "Example message"
```
cannot remove value: the JSON Patch 'remove' target's parent is of type number, but must be an object or array
```
!!! note
This exception was added in version 3.13.0. Before that, this situation was silently ignored (the `remove` operation had no effect).
### json.exception.out_of_range.414
A JSON Patch `move` operation's `"from"` location is a proper prefix of its `"path"` location. Per [RFC 6902](https://datatracker.ietf.org/doc/html/rfc6902) (section 4.4), a location cannot be moved into one of its own children.
!!! failure "Example message"
```
cannot move value: 'from' path '/0' is a proper prefix of 'path' '/0/0'
```
!!! note
This exception was added in version 3.13.0. Before that, this situation could succeed with a corrupted result: for an array target, removing the "from" element before the "add" step shifted subsequent indices, so "path" silently re-resolved to a different element than intended.
### json.exception.out_of_range.415
MessagePack's ext type and BSON's binary subtype are each stored in a single byte. This exception is thrown when serializing a
[`byte_container_with_subtype`](../api/byte_container_with_subtype/index.md) whose subtype exceeds 255.
!!! failure "Example message"
```
[json.exception.out_of_range.415] subtype 70000 is too large for the MessagePack ext type (max 255)
```
!!! note
This exception was added in version 3.13.0. Before that, subtypes above 255 were silently truncated modulo 256 instead of raising an error.
## Further exceptions
This exception is thrown in case of errors that cannot be classified with the
+48
View File
@@ -90,6 +90,54 @@ The library supports **Unicode input** as follows:
In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding.
### NUL bytes in the input
!!! question "Questions"
- Why does `json::parse()` silently ignore part of my input?
- Why does a `std::string`/buffer with extra data after the JSON text parse without error, while a similar-looking string with extra text does not?
A `'\0'` (NUL) byte anywhere in the input is treated the same as the real end of the input, rather than as an ordinary (and, outside of a string, invalid) byte. Everything from that byte onward is silently ignored, without a parse error — including further, otherwise well-formed JSON:
```cpp
json::parse(std::string("123") + '\0'); // == 123, no error
json::parse(std::string("123") + '\0' + "true"); // == 123, the "true" is silently ignored too
```
This is different from any other unexpected trailing byte, which *does* raise [`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101):
```cpp
json::parse("123x"); // throws parse_error.101: unexpected additional data
```
This falls out of the same convention used when no explicit input length is given at all: `json::parse(const char*)` already stops at the first NUL byte via `strlen()`, since a bare pointer has no length of its own. The library applies that same NUL-terminated-C-string convention uniformly, rather than only when a length is genuinely unavailable — so a `std::string`, iterator range, or container whose content happens to include a NUL byte is affected the same way a raw `const char*` would be.
If your input may contain a trailing or embedded NUL that is **not** meant to signal the end of the JSON text — for instance, a fixed-size, zero-padded buffer — trim it yourself before calling `parse()`, since the library will otherwise silently stop there instead of raising an error:
```cpp
s.resize(s.find('\0')); // drop everything from the first NUL onward, if any
json::parse(s);
```
**Opt-in strict handling (since version 3.13.0)**
Manually trimming every input is easy to forget. If you define [`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md) to `1` before including the library, a `'\0'` byte is instead rejected like any other unexpected byte and raises `parse_error.101`, instead of being treated as end of input:
```cpp
#define JSON_STRICT_NUL_HANDLING 1
#include <nlohmann/json.hpp>
json::parse(std::string("123") + '\0'); // throws parse_error.101 instead of silently returning 123
```
This macro defaults to `0` (disabled, preserving the behavior described above) to avoid breaking existing code that may depend on it, even unknowingly; it is planned to become the default in version 4.0.0. See [its documentation](../api/macros/json_strict_nul_handling.md) for details, including how it also affects `char` arrays such as string literals.
Note that this is unrelated to an *unescaped* NUL byte occurring **inside** a quoted JSON string, which is a different, already-invalid case and is correctly rejected either way:
```cpp
json::parse(std::string("\"") + '\0' + "\""); // throws parse_error.101: control character U+0000 (NUL) must be escaped to \u0000
```
### Wide string handling
!!! question
+5
View File
@@ -198,6 +198,11 @@ Use the non-amalgamated version of the library. This option is `ON` by default.
Treat the library headers like system headers (i.e., adding `SYSTEM` to the [`target_include_directories`](https://cmake.org/cmake/help/latest/command/target_include_directories.html) call) to check for this library by tools like Clang-Tidy. This option is `OFF` by default.
### `JSON_StrictNulHandling`
Reject a `'\0'` (NUL) byte in the input instead of treating it as end of input, by defining the macro
[`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md). This option is `OFF` by default.
### `JSON_Valgrind`
Execute the test suite with [Valgrind](https://valgrind.org). This option is `OFF` by default. Depends on `JSON_BuildTests`.
+11 -1
View File
@@ -98,6 +98,7 @@ nav:
- Types:
- features/types/index.md
- features/types/number_handling.md
- features/types/template_parameters.md
- Integration:
- integration/index.md
- integration/migration_guide.md
@@ -293,9 +294,11 @@ nav:
- 'JSON_NO_IO': api/macros/json_no_io.md
- 'JSON_SKIP_LIBRARY_VERSION_CHECK': api/macros/json_skip_library_version_check.md
- 'JSON_SKIP_UNSUPPORTED_COMPILER_CHECK': api/macros/json_skip_unsupported_compiler_check.md
- 'JSON_STRICT_NUL_HANDLING': api/macros/json_strict_nul_handling.md
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
- 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md
- 'JSON_USE_SIMDUTF': api/macros/json_use_simdutf.md
- 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md
- 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md
- 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md
@@ -393,7 +396,14 @@ plugins:
- http://nlohmann.github.io/json/*
- https://nlohmann.github.io/json/*
- mailto:*
- privacy
- privacy:
# repology.org refuses requests from GitHub Actions runners, which made
# the privacy plugin abort the whole build when it could not download the
# package badges (the fetch fails, then reading the missing cache entry
# raises FileNotFoundError). Readers' browsers are served normally, so
# leave these badges as external references instead of self-hosting them.
assets_exclude:
- repology.org/*
- llmstxt:
markdown_description: >
JSON for Modern C++ is a C++11 header-only library implementing a JSON
+1 -1
View File
@@ -1,7 +1,7 @@
wheel==0.48.0
mkdocs==1.6.1 # documentation framework
mkdocs-git-revision-date-localized-plugin==1.5.4 # plugin "git-revision-date-localized"
mkdocs-git-revision-date-localized-plugin==1.6.0 # plugin "git-revision-date-localized"
mkdocs-material==9.7.7 # theme for mkdocs
mkdocs-material-extensions==1.3.1 # extensions
mkdocs-minify-plugin==0.8.0 # plugin "minify"
@@ -398,6 +398,17 @@ inline void from_json(const BasicJsonType& j, CompatibleArrayType& bin)
}
}
template<typename ConstructibleObjectType>
auto from_json_object_reserve(ConstructibleObjectType& obj, typename ConstructibleObjectType::size_type size, priority_tag<1> /*unused*/)
-> decltype(obj.reserve(size), void())
{
obj.reserve(size);
}
template<typename ConstructibleObjectType>
inline void from_json_object_reserve(ConstructibleObjectType& /*obj*/, std::size_t /*size*/, priority_tag<0> /*unused*/)
{}
template<typename BasicJsonType, typename ConstructibleObjectType,
enable_if_t<is_constructible_object_type<BasicJsonType, ConstructibleObjectType>::value, int> = 0>
inline void from_json(const BasicJsonType& j, ConstructibleObjectType& obj)
@@ -409,6 +420,7 @@ inline void from_json(const BasicJsonType& j, ConstructibleObjectType& obj)
ConstructibleObjectType ret;
const auto* inner_object = j.template get_ptr<const typename BasicJsonType::object_t*>();
from_json_object_reserve(ret, inner_object->size(), priority_tag<1> {});
for (const auto& p : *inner_object)
{
ret.emplace(p.first, p.second.template get<typename ConstructibleObjectType::mapped_type>());
@@ -1075,8 +1075,8 @@ char* to_chars(char* first, const char* last, FloatType value)
}
#ifdef __GNUC__
#pragma GCC diagnostic push
#pragma GCC diagnostic ignored "-Wfloat-equal"
JSON_HEDLEY_DIAGNOSTIC_PUSH
JSON_HEDLEY_PRAGMA(GCC diagnostic ignored "-Wfloat-equal")
#endif
if (value == 0) // +-0
{
@@ -1087,7 +1087,7 @@ char* to_chars(char* first, const char* last, FloatType value)
return first;
}
#ifdef __GNUC__
#pragma GCC diagnostic pop
JSON_HEDLEY_DIAGNOSTIC_POP
#endif
JSON_ASSERT(last - first >= std::numeric_limits<FloatType>::max_digits10);
+7 -4
View File
@@ -33,8 +33,8 @@
// code stumbling over this. See https://github.com/nlohmann/json/issues/4087
// for a discussion.
#if defined(__clang__)
#pragma clang diagnostic push
#pragma clang diagnostic ignored "-Wweak-vtables"
JSON_HEDLEY_DIAGNOSTIC_PUSH
JSON_HEDLEY_PRAGMA(clang diagnostic ignored "-Wweak-vtables")
#endif
NLOHMANN_JSON_NAMESPACE_BEGIN
@@ -101,7 +101,10 @@ class exception : public std::exception
{
if (&element.second == current)
{
tokens.emplace_back(element.first.c_str());
// data() is null-terminated, so a key containing
// a null byte is cut short here rather than
// truncating the whole message at what()
tokens.emplace_back(element.first.data());
break;
}
}
@@ -287,5 +290,5 @@ class other_error : public exception
NLOHMANN_JSON_NAMESPACE_END
#if defined(__clang__)
#pragma clang diagnostic pop
JSON_HEDLEY_DIAGNOSTIC_POP
#endif
+3 -1
View File
@@ -114,7 +114,9 @@ std::size_t hash(const BasicJsonType& j)
seed = combine(seed, static_cast<std::size_t>(j.get_binary().subtype()));
for (const auto byte : j.get_binary())
{
seed = combine(seed, std::hash<std::uint8_t> {}(byte));
// the cast is needed for binary types whose value type is not
// an integer (e.g., std::byte)
seed = combine(seed, std::hash<std::uint8_t> {}(static_cast<std::uint8_t>(byte)));
}
return seed;
}
File diff suppressed because it is too large Load Diff
+146 -21
View File
@@ -155,11 +155,31 @@ class input_stream_adapter
// General-purpose iterator-based adapter. It might not be as fast as
// theoretically possible for some containers, but it is extremely versatile.
// SentinelType defaults to IteratorType for backward compatibility, but may
// be a different type (e.g., a C++20 sentinel or counted_iterator).
// SentinelType defaults to IteratorType for backward compatibility, but may be
// a different type, e.g. a C++20 sentinel such as std::default_sentinel_t when
// IteratorType is a std::counted_iterator.
template<typename IteratorType, typename SentinelType = IteratorType>
class iterator_input_adapter
{
// Whether the number of elements between two positions can be computed in
// O(1): either the iterator and the sentinel have the same type (plain
// std::distance) or, in C++20, the sentinel is a sized sentinel for the
// iterator (std::ranges::distance), e.g. std::default_sentinel_t paired
// with std::counted_iterator.
//
// JSON_HAS_RANGES gates the C++20 branch: on standard libraries with an
// incomplete <ranges> (libstdc++ < 11, see #4440) evaluating
// std::contiguous_iterator on a std::counted_iterator is a hard error
// instead of yielding false, and these traits are instantiated for every
// adapter. Such toolchains fall back to the pointer-only test and simply
// use the byte-at-a-time scanner.
static constexpr bool sentinel_is_sized =
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>;
#else
std::is_same<IteratorType, SentinelType>::value;
#endif
public:
using char_type = typename std::iterator_traits<IteratorType>::value_type;
@@ -171,7 +191,7 @@ class iterator_input_adapter
// in wide_string_input_adapter, which does not expose this).
static constexpr bool supports_seek =
std::is_same<typename std::iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value
&& std::is_same<IteratorType, SentinelType>::value
&& sentinel_is_sized
&& sizeof(char_type) == 1;
iterator_input_adapter(IteratorType first, SentinelType last)
@@ -219,30 +239,60 @@ class iterator_input_adapter
private:
// whether IteratorType refers to a contiguous range and therefore supports
// a std::memcpy fast path (pointers always do; in C++20 we can also detect
// library iterators such as those of std::vector and std::string).
// Computing the available element count needs either same-type iterators
// (plain std::distance) or, in C++20, a sized sentinel (std::ranges::distance),
// e.g. std::counted_iterator paired with std::default_sentinel_t.
static constexpr bool iterator_is_contiguous =
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>)
&& (std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
// library iterators such as those of std::vector and std::string). The
// available element count must also be computable in O(1), hence
// sentinel_is_sized.
static constexpr bool iterator_is_contiguous = sentinel_is_sized &&
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
#else
std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
std::is_pointer<IteratorType>::value;
#endif
// number of unread elements in [current, end)
std::size_t remaining_count() const
{
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
return static_cast<std::size_t>(std::ranges::distance(current, end));
#else
return static_cast<std::size_t>(std::distance(current, end));
#endif
}
public:
// Whether the remaining input is a single contiguous block of 1-byte
// elements that the lexer can inspect directly (used for the SWAR string
// fast path).
static constexpr bool supports_bulk_scan =
iterator_is_contiguous && sizeof(char_type) == 1;
// Pointer to the next unread element; only valid when bulk_remaining() > 0.
const char_type* bulk_data() const
{
return &*current;
}
// Number of unread elements available as one contiguous block.
std::size_t bulk_remaining() const
{
return remaining_count();
}
// Consume @a n elements previously inspected via bulk_data().
void bulk_skip(std::size_t n)
{
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(n));
}
private:
// contiguous fast path: bulk copy the remaining range with std::memcpy
template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
{
const std::size_t wanted = count * sizeof(T);
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
const std::size_t available = static_cast<std::size_t>(std::ranges::distance(current, end)) * sizeof(char_type);
#else
const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type);
#endif
const std::size_t available = remaining_count() * sizeof(char_type);
const std::size_t copied = (std::min)(wanted, available);
if (JSON_HEDLEY_LIKELY(copied != 0))
{
@@ -570,6 +620,46 @@ typename iterator_input_adapter_factory<IteratorType, SentinelType>::adapter_typ
return factory_type::create(first, last);
}
// The element type a container's data() points at, cv-qualifiers removed.
// Ill-formed - and therefore SFINAE-friendly - for types without data().
template<typename ContainerType>
using container_data_t = typename std::remove_cv<typename std::remove_pointer <
decltype(std::declval<const ContainerType&>().data()) >::type >::type;
// The container's own element type, cv-qualifiers removed. It is looked up on
// the bare type so it is also found when ContainerType is deduced as a
// reference by the forwarding-reference overload below.
template<typename ContainerType>
using container_value_t = typename std::remove_cv <
typename std::remove_cv<typename std::remove_reference<ContainerType>::type>::type::value_type >::type;
// Detect a container that stores its elements contiguously as single bytes
// (std::string, std::vector<char/unsigned char>, std::array<char, N>,
// std::string_view, ...). Such inputs are wrapped in a pointer-based adapter so
// they benefit from the contiguous fast paths (bulk string scanning, memcpy for
// binary formats) in every C++ standard - not only in C++20, where the standard
// library iterators model std::contiguous_iterator and are detected directly.
//
// data() and size() on their own would be duck typing: they say nothing about
// size() counting the units data() points at, and reading [data(), data() +
// size()) as bytes would be wrong for a type where it does not. Requiring the
// container's own value_type to be that same single-byte element ties the two
// together; every contiguous standard container satisfies it. Anything else
// keeps the iterator-based adapter, which is always correct - only slower.
template<typename ContainerType, typename = void>
struct is_contiguous_byte_container : std::false_type {};
template<typename ContainerType>
struct is_contiguous_byte_container < ContainerType, void_t <
container_data_t<ContainerType>,
container_value_t<ContainerType>,
decltype(std::declval<const ContainerType&>().size()) >>
: std::integral_constant < bool,
std::is_pointer<decltype(std::declval<const ContainerType&>().data())>::value&&
std::is_integral<container_data_t<ContainerType>>::value&&
sizeof(container_data_t<ContainerType>) == 1 &&
std::is_same<container_data_t<ContainerType>, container_value_t<ContainerType>>::value > {};
// Convenience shorthand from container to iterator
// Enables ADL on begin(container) and end(container)
// Encloses the using declarations in namespace for not to leak them to outside scope
@@ -597,12 +687,32 @@ struct container_input_adapter_factory< ContainerType,
} // namespace container_input_adapter_factory_impl
template<typename ContainerType>
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType&& container)
// General container path (iterator-based). Contiguous single-byte containers
// are excluded here and routed through the pointer-based overload below.
template < typename ContainerType,
enable_if_t < !is_contiguous_byte_container<ContainerType>::value, int > = 0 >
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType && container)
{
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
}
// Contiguous single-byte containers (std::string, std::vector<char>, ...) are
// wrapped in a pointer-based adapter so the contiguous fast paths apply in every
// standard. The pointer keeps the container's own element type (const char* for
// std::string, const std::uint8_t* for std::vector<std::uint8_t>, ...), so the
// resulting char_type - and therefore the parsing behavior - is byte-for-byte
// identical to the iterator-based path; only the raw pointer additionally
// enables the bulk fast paths. The container outlives the adapter for the whole
// parse (temporaries live until the end of the full expression), exactly as the
// iterators it replaces did.
template < typename ContainerType,
enable_if_t < is_contiguous_byte_container<ContainerType>::value, int > = 0 >
auto input_adapter(const ContainerType& container)
-> decltype(input_adapter(container.data(), container.data() + container.size()))
{
return input_adapter(container.data(), container.data() + container.size());
}
// specialization for std::string
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
@@ -652,6 +762,21 @@ contiguous_bytes_input_adapter input_adapter(CharT b)
template<typename T, std::size_t N>
auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
{
#if JSON_STRICT_NUL_HANDLING
// A `char` array from string-literal initialization (e.g. json::parse("123"))
// carries a trailing '\0' contributed by the compiler, not by the source
// text; drop exactly that one byte so it is not mistaken for real trailing
// data. Every other element type (unsigned char, std::uint8_t, ...) keeps
// the full extent unconditionally, since a trailing zero byte there is
// genuine data (e.g. CBOR/MessagePack). This intentionally does not
// strlen()-scan the array (as the pointer overload above does for a
// null-delimited string): for a `char` array that is not NUL-terminated
// within its bounds, that would read past the end of the array.
if (std::is_same<typename std::remove_cv<T>::type, char>::value && N > 0 && array[N - 1] == 0)
{
return input_adapter(array, array + N - 1);
}
#endif
return input_adapter(array, array + N);
}
+217 -25
View File
@@ -8,15 +8,17 @@
#pragma once
#include <algorithm> // find_if, min
#include <cstddef>
#include <string> // string
#include <type_traits> // enable_if_t
#include <utility> // move
#include <utility> // move, pair
#include <vector> // vector
#include <nlohmann/detail/exceptions.hpp>
#include <nlohmann/detail/input/lexer.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/cpp_future.hpp>
#include <nlohmann/detail/string_concat.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
@@ -150,6 +152,29 @@ constexpr std::size_t unknown_size()
return (std::numeric_limits<std::size_t>::max)();
}
/*!
@brief reserve capacity for @a len elements in array @a arr
Reserving upfront avoids repeated reallocations while the elements are added,
but the reservation is capped so a bogus/hostile length (which is not bounded
by max_size(), unlike e.g. std::vector) cannot trigger an oversized allocation
for a small or truncated input.
The overload below is selected for array types without reserve() (e.g.,
std::deque), which are then left untouched.
*/
template<typename ArrayType>
auto reserve_array(ArrayType& arr, std::size_t len, priority_tag<1> /*unused*/)
-> decltype(arr.reserve(len), void())
{
constexpr std::size_t reserve_cap = 16384;
arr.reserve((std::min)(len, reserve_cap));
}
template<typename ArrayType>
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
{}
/*!
@brief SAX implementation to create a JSON value from SAX events
@@ -222,12 +247,16 @@ class json_sax_dom_parser
bool string(string_t& val)
{
handle_value(val);
// json_sax documents that the passed value may be moved from,
// so hand the buffer over instead of copying it
handle_value(std::move(val));
return true;
}
bool binary(binary_t& val)
{
// json_sax documents that the passed value may be moved from,
// so hand the buffer over instead of copying it
handle_value(std::move(val));
return true;
}
@@ -249,7 +278,7 @@ class json_sax_dom_parser
if (JSON_HEDLEY_UNLIKELY(len != detail::unknown_size() && len > ref_stack.back()->max_size()))
{
JSON_THROW(out_of_range::create(408, concat("excessive object size: ", std::to_string(len)), ref_stack.back()));
return parse_error(0, "", out_of_range::create(408, concat("excessive object size: ", std::to_string(len)), ref_stack.back()));
}
return true;
@@ -298,7 +327,12 @@ class json_sax_dom_parser
if (JSON_HEDLEY_UNLIKELY(len != detail::unknown_size() && len > ref_stack.back()->max_size()))
{
JSON_THROW(out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
return parse_error(0, "", out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
}
if (len != detail::unknown_size())
{
reserve_array(*ref_stack.back()->m_data.m_value.array, len, priority_tag<1> {});
}
return true;
@@ -532,12 +566,16 @@ class json_sax_dom_callback_parser
bool string(string_t& val)
{
handle_value(val);
// json_sax documents that the passed value may be moved from,
// so hand the buffer over instead of copying it
handle_value(std::move(val));
return true;
}
bool binary(binary_t& val)
{
// json_sax documents that the passed value may be moved from,
// so hand the buffer over instead of copying it
handle_value(std::move(val));
return true;
}
@@ -548,6 +586,11 @@ class json_sax_dom_callback_parser
const bool keep = callback(static_cast<int>(ref_stack.size()), parse_event_t::object_start, discarded);
keep_stack.push_back(keep);
// the key this object will be stored under, read before handle_value()
// may consume it; kept in lockstep with ref_stack so end_object() can
// find the object in its parent again
container_key_stack.push_back(current_key());
auto val = handle_value(BasicJsonType::value_t::object, true);
ref_stack.push_back(val.second);
@@ -568,7 +611,7 @@ class json_sax_dom_callback_parser
// check object limit
if (JSON_HEDLEY_UNLIKELY(len != detail::unknown_size() && len > ref_stack.back()->max_size()))
{
JSON_THROW(out_of_range::create(408, concat("excessive object size: ", std::to_string(len)), ref_stack.back()));
return parse_error(0, "", out_of_range::create(408, concat("excessive object size: ", std::to_string(len)), ref_stack.back()));
}
}
return true;
@@ -581,11 +624,24 @@ class json_sax_dom_callback_parser
// check callback for the key
const bool keep = callback(static_cast<int>(ref_stack.size()), parse_event_t::key, k);
key_keep_stack.push_back(keep);
// remember the key so a rejected value can be erased without searching
// the object for it (kept in lockstep with key_keep_stack)
key_stack.push_back(val);
// add discarded value at the given key and store the reference for later
if (keep && ref_stack.back())
{
object_element = &(ref_stack.back()->m_data.m_value.object->operator[](val) = discarded);
auto& obj = *ref_stack.back()->m_data.m_value.object;
const auto it = obj.find(val);
if (it != obj.end())
{
// this is a duplicate key (legal in JSON); remember its
// current value so it can be restored later if the new
// value is rejected by the callback, instead of being
// erased together with the discarded placeholder
duplicate_key_stash.emplace_back(&(it->second), it->second);
}
object_element = &(obj[val] = discarded);
}
return true;
@@ -597,13 +653,18 @@ class json_sax_dom_callback_parser
{
if (!callback(static_cast<int>(ref_stack.size()) - 1, parse_event_t::object_end, *ref_stack.back()))
{
// discard object
*ref_stack.back() = discarded;
// discard object, unless this slot holds a duplicate key's
// previous value pending restoration, in which case that
// value is restored instead of being discarded
if (!resolve_duplicate_key_stash(ref_stack.back(), true))
{
*ref_stack.back() = discarded;
#if JSON_DIAGNOSTIC_POSITIONS
// Set start/end positions for discarded object.
handle_diagnostic_positions_for_json_value(*ref_stack.back());
// Set start/end positions for discarded object.
handle_diagnostic_positions_for_json_value(*ref_stack.back());
#endif
}
}
else
{
@@ -617,18 +678,25 @@ class json_sax_dom_callback_parser
#endif
ref_stack.back()->set_parents();
// this object is finally, definitively kept; drop any
// pending duplicate-key stash entry for its slot since it
// can no longer be restored
resolve_duplicate_key_stash(ref_stack.back(), false);
}
}
JSON_ASSERT(!ref_stack.empty());
JSON_ASSERT(!keep_stack.empty());
JSON_ASSERT(!container_key_stack.empty());
ref_stack.pop_back();
keep_stack.pop_back();
const string_t object_key = std::move(container_key_stack.back());
container_key_stack.pop_back();
if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_structured())
{
// remove discarded value
remove_discarded_value(*ref_stack.back());
remove_discarded_value(*ref_stack.back(), object_key);
}
return true;
@@ -639,6 +707,9 @@ class json_sax_dom_callback_parser
const bool keep = callback(static_cast<int>(ref_stack.size()), parse_event_t::array_start, discarded);
keep_stack.push_back(keep);
// see start_object()
container_key_stack.push_back(current_key());
auto val = handle_value(BasicJsonType::value_t::array, true);
ref_stack.push_back(val.second);
@@ -659,7 +730,12 @@ class json_sax_dom_callback_parser
// check array limit
if (JSON_HEDLEY_UNLIKELY(len != detail::unknown_size() && len > ref_stack.back()->max_size()))
{
JSON_THROW(out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
return parse_error(0, "", out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
}
if (len != detail::unknown_size())
{
reserve_array(*ref_stack.back()->m_data.m_value.array, len, priority_tag<1> {});
}
}
@@ -686,23 +762,35 @@ class json_sax_dom_callback_parser
#endif
ref_stack.back()->set_parents();
// this array is finally, definitively kept; drop any
// pending duplicate-key stash entry for its slot since it
// can no longer be restored
resolve_duplicate_key_stash(ref_stack.back(), false);
}
else
{
// discard array
*ref_stack.back() = discarded;
// discard array, unless this slot holds a duplicate key's
// previous value pending restoration, in which case that
// value is restored instead of being discarded
if (!resolve_duplicate_key_stash(ref_stack.back(), true))
{
*ref_stack.back() = discarded;
#if JSON_DIAGNOSTIC_POSITIONS
// Set start/end positions for discarded array.
handle_diagnostic_positions_for_json_value(*ref_stack.back());
// Set start/end positions for discarded array.
handle_diagnostic_positions_for_json_value(*ref_stack.back());
#endif
}
}
}
JSON_ASSERT(!ref_stack.empty());
JSON_ASSERT(!keep_stack.empty());
JSON_ASSERT(!container_key_stack.empty());
ref_stack.pop_back();
keep_stack.pop_back();
const string_t object_key = std::move(container_key_stack.back());
container_key_stack.pop_back();
// remove discarded value
if (!ref_stack.empty() && ref_stack.back())
@@ -716,7 +804,7 @@ class json_sax_dom_callback_parser
// the array is either still stored under its key or was never
// stored, leaving the placeholder key() wrote; both show up as
// a discarded member of the parent object
remove_discarded_value(*ref_stack.back());
remove_discarded_value(*ref_stack.back(), object_key);
}
}
@@ -809,15 +897,92 @@ class json_sax_dom_callback_parser
}
#endif
/// remove the discarded value the callback rejected from its parent
static void remove_discarded_value(BasicJsonType& parent)
/// if there is a pending duplicate-key stash entry for this exact slot,
/// remove it from the stash; if restore_value is true, the stashed
/// previous value is moved back into the slot first (use this when the
/// new value at that slot was rejected); otherwise the stash entry is
/// simply dropped (use this when the new value was accepted, so it
/// correctly supersedes the old one and no restore should ever happen
/// for this slot again)
/// @return whether a matching stash entry was found (and processed)
bool resolve_duplicate_key_stash(BasicJsonType* slot, bool restore_value)
{
for (auto it = parent.begin(); it != parent.end(); ++it)
const auto it = std::find_if(duplicate_key_stash.begin(), duplicate_key_stash.end(),
[slot](const std::pair<BasicJsonType*, BasicJsonType>& entry)
{
if (it->is_discarded())
return entry.first == slot;
});
if (it == duplicate_key_stash.end())
{
return false;
}
if (restore_value)
{
*slot = std::move(it->second);
}
duplicate_key_stash.erase(it);
return true;
}
/*!
@brief the key the value now being handled will be stored under
Empty unless the enclosing container is an object, in which case it is the
key of the pending key() event. Read before handle_value() consumes that
key, so it is also correct when the value never reaches its parent.
*/
string_t current_key() const
{
if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_object()
&& !key_stack.empty())
{
return key_stack.back();
}
return string_t{};
}
/*!
@brief remove the discarded value the callback rejected from its parent,
unless it is a duplicate key's slot with a stashed previous value, in
which case that previous value is restored instead
A rejected value can only ever be the one most recently added to @a parent:
the last element of an array, or the placeholder key() stored under @a key
in an object. Looking there directly makes this O(1) resp. O(log n), where
searching @a parent for it made a filtering parse quadratic in the number of
members of a single container.
Finding no discarded value there means none was stored in the first place -
the callback rejected the value before it reached its parent - so there is
nothing to remove.
@param[in,out] parent the container to remove the rejected value from
@param[in] key the key the value was stored under; unused for arrays
*/
void remove_discarded_value(BasicJsonType& parent, const string_t& key)
{
if (parent.is_array())
{
auto& array = *parent.m_data.m_value.array;
if (!array.empty() && array.back().is_discarded())
{
parent.erase(it);
break;
array.pop_back();
}
}
else if (parent.is_object())
{
auto& object = *parent.m_data.m_value.object;
const auto it = object.find(key);
if (it != object.end() && it->second.is_discarded())
{
// a duplicate key's slot has a stashed previous value that
// must be restored instead of being erased
if (!resolve_duplicate_key_stash(&it->second, true))
{
object.erase(it);
}
}
}
}
@@ -867,11 +1032,14 @@ class json_sax_dom_callback_parser
if (!ref_stack.empty() && ref_stack.back() && ref_stack.back()->is_object())
{
JSON_ASSERT(!key_keep_stack.empty());
JSON_ASSERT(!key_stack.empty());
const bool placeholder_stored = key_keep_stack.back();
key_keep_stack.pop_back();
const string_t key = std::move(key_stack.back());
key_stack.pop_back();
if (placeholder_stored)
{
remove_discarded_value(*ref_stack.back());
remove_discarded_value(*ref_stack.back(), key);
}
}
return {false, nullptr};
@@ -904,8 +1072,10 @@ class json_sax_dom_callback_parser
JSON_ASSERT(ref_stack.back()->is_object());
// check if we should store an element for the current key
JSON_ASSERT(!key_keep_stack.empty());
JSON_ASSERT(!key_stack.empty());
const bool store_element = key_keep_stack.back();
key_keep_stack.pop_back();
key_stack.pop_back();
if (!store_element)
{
@@ -914,6 +1084,16 @@ class json_sax_dom_callback_parser
JSON_ASSERT(object_element);
*object_element = std::move(value);
if (!skip_callback)
{
// this scalar value finally, definitively replaces whatever was
// at this slot; drop any pending duplicate-key stash entry for
// it since it can no longer be restored (a container value at
// this slot is resolved later, in end_object()/end_array(),
// since skip_callback is true for the placeholder handling that
// happens here for those)
resolve_duplicate_key_stash(object_element, false);
}
return {true, object_element};
}
@@ -925,8 +1105,20 @@ class json_sax_dom_callback_parser
std::vector<bool> keep_stack {}; // NOLINT(readability-redundant-member-init)
/// stack to manage which object keys to keep
std::vector<bool> key_keep_stack {}; // NOLINT(readability-redundant-member-init)
/// the keys key() stored a placeholder for, in lockstep with key_keep_stack
std::vector<string_t> key_stack {}; // NOLINT(readability-redundant-member-init)
/// for each open container, the key it is stored under in its parent
/// object, in lockstep with ref_stack; unused where the parent is not an
/// object
std::vector<string_t> container_key_stack {}; // NOLINT(readability-redundant-member-init)
/// helper to hold the reference for the next object element
BasicJsonType* object_element = nullptr;
/// stash of (slot pointer, previous value) for object members that
/// already existed when key() was called again for the same key
/// (duplicate keys); used to restore the previous value if the new
/// value is later rejected by the callback, instead of erasing the
/// member entirely
std::vector<std::pair<BasicJsonType*, BasicJsonType>> duplicate_key_stash {};
/// whether a syntax error occurred
bool errored = false;
/// callback function
+510 -35
View File
@@ -19,7 +19,9 @@
#include <vector> // vector
#include <nlohmann/detail/input/input_adapters.hpp>
#include <nlohmann/detail/input/number_parse.hpp>
#include <nlohmann/detail/input/position_t.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
@@ -125,6 +127,25 @@ constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
return false;
}
// Detect whether an input adapter exposes a contiguous byte block that the
// lexer can scan directly (see iterator_input_adapter::supports_bulk_scan).
// Adapters without the flag - file, stream, wide-string, user-defined - fall
// back to the character-at-a-time string scanner.
template<typename InputAdapterType>
using detect_supports_bulk_scan = decltype(InputAdapterType::supports_bulk_scan);
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::true_type /*detected*/)
{
return InputAdapterType::supports_bulk_scan;
}
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::false_type /*detected*/)
{
return false;
}
/*!
@brief lexical analysis
@@ -146,13 +167,22 @@ class lexer : public lexer_base<BasicJsonType>
static constexpr bool lazy_token_string =
input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {});
/// whether string scanning may bulk-consume runs of ordinary characters
/// directly from a contiguous input buffer (SWAR fast path). This requires
/// the token to be reconstructible lazily (lazy_token_string), so bypassing
/// the per-character capture in get() cannot lose error diagnostics.
static constexpr bool bulk_scan =
lazy_token_string
&& input_adapter_supports_bulk_scan<InputAdapterType>(is_detected<detect_supports_bulk_scan, InputAdapterType> {});
public:
using token_type = typename lexer_base<BasicJsonType>::token_type;
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false) noexcept
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter))
, ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_)
{}
// deleted because of pointer members
@@ -265,6 +295,40 @@ class lexer : public lexer_base<BasicJsonType>
return true;
}
/// contiguous input: bulk-append the run of ordinary characters and complete
/// well-formed UTF-8 sequences starting at the current read position, leaving
/// the first byte that needs individual handling (the closing quote, an
/// escape, a control character, or an ill-formed UTF-8 byte) for get()
void scan_string_bulk(std::true_type /*bulk*/)
{
// a pending unget must be consumed through the normal path first
if (next_unget)
{
return;
}
const std::size_t remaining = ia.bulk_remaining();
if (remaining == 0)
{
return;
}
const auto* const data = reinterpret_cast<const unsigned char*>(ia.bulk_data());
const std::size_t pos = string_bulk_run(data, remaining);
if (pos == 0)
{
return;
}
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), pos);
ia.bulk_skip(pos);
// the run contains no newline (all bytes < 0x20 are treated as special),
// so only the flat character counters advance
position.chars_read_total += pos;
position.chars_read_current_line += pos;
}
/// streaming input: no bulk fast path
void scan_string_bulk(std::false_type /*bulk*/) const noexcept {}
/*!
@brief scan a string literal
@@ -290,6 +354,10 @@ class lexer : public lexer_base<BasicJsonType>
while (true)
{
// bulk-consume ordinary characters from contiguous input, then
// handle the next special byte through the switch below
scan_string_bulk(std::integral_constant<bool, bulk_scan> {});
// get the next character
switch (get())
{
@@ -884,7 +952,9 @@ class lexer : public lexer_base<BasicJsonType>
case '\n':
case '\r':
case char_traits<char_type>::eof():
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
return true;
default:
@@ -902,8 +972,10 @@ class lexer : public lexer_base<BasicJsonType>
{
switch (get())
{
case char_traits<char_type>::eof():
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
case char_traits<char_type>::eof():
{
error_message = "invalid comment; missing closing '*/'";
return false;
@@ -1008,6 +1080,12 @@ class lexer : public lexer_base<BasicJsonType>
// changed if minus sign, decimal point, or exponent is read
token_type number_type = token_type::value_unsigned;
// offset just past the last mantissa byte in token_buffer (i.e. the
// index of 'e'/'E', or the whole token when there is no exponent).
// convert_number() uses it to count significant digits; npos means
// "not seen an exponent yet" and is resolved at scan_number_done
std::size_t mantissa_end = std::string::npos;
// state (init): we just found out we need to scan a number
switch (current)
{
@@ -1193,6 +1271,9 @@ scan_number_decimal2:
scan_number_exponent:
// we just parsed an exponent
number_type = token_type::value_float;
// this label is reached only right after the 'e'/'E' was appended (from
// the zero, any1, and decimal2 states), so the mantissa ends before it
mantissa_end = token_buffer.size() - 1;
switch (get())
{
case '+':
@@ -1279,45 +1360,199 @@ scan_number_done:
// we are done scanning a number)
unget();
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
errno = 0;
// no exponent was scanned: the mantissa spans the whole token
if (mantissa_end == std::string::npos)
{
mantissa_end = token_buffer.size();
}
// try to parse integers first and fall back to floats
return convert_number(number_type, mantissa_end);
}
/*!
@brief convert an already-validated integer token to its value
The digit sequence in [first, last) has been validated by the caller, so a
dedicated parser can avoid the locale/errno overhead of std::strtoull.
@return the token type on success; token_type::uninitialized if @a
number_type is not an integer type or the value does not fit, in
which case the caller falls back to the floating-point conversion
(matching the previous std::strtoull/std::strtoll behavior)
*/
token_type convert_integer(token_type number_type, const char* first, const char* last)
{
if (number_type == token_type::value_unsigned)
{
const auto x = std::strtoull(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
if (parse_integer_unsigned(first, last, value_unsigned))
{
value_unsigned = static_cast<number_unsigned_t>(x);
if (value_unsigned == x)
{
return token_type::value_unsigned;
}
return token_type::value_unsigned;
}
}
else if (number_type == token_type::value_integer)
{
const auto x = std::strtoll(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
if (parse_integer_signed(first, last, value_integer))
{
value_integer = static_cast<number_integer_t>(x);
if (value_integer == x)
{
return token_type::value_integer;
}
return token_type::value_integer;
}
}
return token_type::uninitialized;
}
/*!
@brief check whether Clinger's fast path can still succeed for this token
parse_float_fast() needs a significand below 2^53. A mantissa with 17 or
more significant digits is at least 10^16 and therefore always exceeds it,
so calling the fast path would walk the token one extra time only to
decline before strtod has to run anyway.
Significant digits are the mantissa's digits from the first nonzero one on;
the sign, the decimal point, leading zeros, and the exponent do not count.
The answer is derived from indices - the digits are not scanned again - so
this stays off the hot path of the number scanners.
@param[in] mantissa_end offset just past the last mantissa byte in
token_buffer
@return false if parse_float_fast() is guaranteed to decline
*/
bool mantissa_fits_clinger(std::size_t mantissa_end) const
{
// 10^16 already exceeds 2^53, so 17 digits can never fit
constexpr std::size_t limit = 17;
const std::size_t neg = (!token_buffer.empty() && token_buffer[0] == '-') ? 1u : 0u;
const std::size_t has_dot = (decimal_point_position != std::string::npos) ? 1u : 0u;
// the JSON grammar restricts the integer part to "0" or [1-9][0-9]*, so
// a leading zero can only be a lone "0", which is not significant
const std::size_t lead_zero = (token_buffer[neg] == '0') ? 1u : 0u;
JSON_ASSERT(mantissa_end >= neg + has_dot + lead_zero);
std::size_t digits = mantissa_end - neg - has_dot - lead_zero;
if (JSON_HEDLEY_LIKELY(digits < limit))
{
return true;
}
// Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. Note
// token_buffer holds the locale's decimal point, so the fraction is
// located through decimal_point_position rather than by searching '.'.
if (lead_zero != 0)
{
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
for (std::size_t i = decimal_point_position + 1;
digits >= limit && i < mantissa_end && token_buffer[i] == '0'; ++i)
{
--digits;
}
}
return digits < limit;
}
/*!
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
@param[in] mantissa_end offset just past the last mantissa byte in
token_buffer (the index of 'e'/'E', or
token_buffer.size() when there is no exponent);
used to skip Clinger's fast path when it cannot
possibly succeed - see mantissa_fits_clinger()
*/
token_type convert_number(token_type number_type, std::size_t mantissa_end)
{
// If the caller does not need the converted value (only whether the
// input is syntactically valid; see json_sax_acceptor/accept()), an
// unsigned/integer token can be reported without calling
// strtoull()/strtoll() at all, *provided* we can already tell from
// the digit count alone that the conversion cannot overflow 64 bits.
// Such tokens are always finite and are accepted unconditionally by
// the parser regardless of their actual value (parser::sax_parse_internal()
// never checks finiteness for value_unsigned/value_integer), so the
// classification below is all that is needed.
//
// A decimal number with up to 18 digits is always representable in
// both std::uint64_t and std::int64_t (18 nines is ~1e18, well below
// both UINT64_MAX ~1.8e19 and INT64_MAX ~9.2e18), so strtoull()/strtoll()
// could not have set errno to ERANGE for it. Numbers with more digits
// (rare in practice) fall through to the exact code below, unchanged,
// so their handling -- including reclassification to value_float when
// the value overflows 64 bits, and rejection when it is not even
// finite as a double -- is bit-for-bit identical to before this
// optimization.
//
// Note this reasons about std::uint64_t/std::int64_t, not about
// number_unsigned_t/number_integer_t (BasicJsonType's own, possibly
// narrower, template parameters -- e.g. std::uint32_t). That is fine
// *only* because discard_number_values is exclusively set by
// accept() (see json.hpp), and accept() always parses through the
// library's own json_sax_acceptor -- never a user-supplied SAX
// consumer -- whose number_unsigned()/number_integer()/number_float()
// callbacks unconditionally discard their argument and return true.
// So for every caller that can reach this branch, neither the token
// classification below nor the eventual (possibly narrowed, and on
// this fast path left stale/unset) value_unsigned/value_integer is
// ever consulted -- an unsigned/integer token is accepted outright,
// and even a >18-digit token that this fast path deliberately falls
// through for is, once reclassified to value_float, still finite
// (and thus accepted) for any digit count that fits in number_unsigned_t
// or number_integer_t regardless of that type's width. If this
// function is ever taught to run with discard_number_values true for
// a caller that *does* read the converted value, this reasoning (and
// the fast path below) would need to be revisited.
if (discard_number_values)
{
constexpr std::size_t safe_digit_count = 18;
if (number_type == token_type::value_unsigned && token_buffer.size() <= safe_digit_count)
{
return token_type::value_unsigned;
}
if (number_type == token_type::value_integer && token_buffer.size() - 1 <= safe_digit_count)
{
return token_type::value_integer;
}
}
const char* const num_begin = token_buffer.data();
const char* const num_end = num_begin + token_buffer.size();
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, num_begin, num_end);
if (integer_result != token_type::uninitialized)
{
return integer_result;
}
}
// this code is reached if we parse a floating-point number or if an
// integer conversion above failed
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
}
// Skipping a fast path that cannot succeed is lossless and saves a full
// extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
{
return token_type::value_float;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
@@ -1326,6 +1561,158 @@ scan_number_done:
return token_type::value_float;
}
/*!
@brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
then produces the exact diagnostic. @a current is the first digit or the
leading minus (already read); the remaining bytes are taken from the adapter.
*/
token_type scan_number_bulk_contiguous()
{
// a pending unget offsets the buffer position from current; fall back
if (next_unget)
{
return token_type::uninitialized;
}
const std::size_t rem = ia.bulk_remaining();
if (rem == 0)
{
// the first digit is the last input byte; let scan_number() finish
return token_type::uninitialized;
}
// the byte before the next unread one is current (contiguous input)
const char* const data = reinterpret_cast<const char*>(ia.bulk_data()) - 1;
const std::size_t avail = rem + 1;
// validate + classify the number extent (mirrors scan_number()'s grammar)
std::size_t i = 0;
std::size_t dot_index = std::string::npos;
token_type number_type = token_type::value_unsigned;
if (data[0] == '-')
{
number_type = token_type::value_integer;
i = 1;
if (i >= avail)
{
return token_type::uninitialized;
}
}
if (data[i] == '0')
{
++i;
}
else if (data[i] >= '1' && data[i] <= '9')
{
++i;
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
else
{
return token_type::uninitialized;
}
if (i < avail && data[i] == '.')
{
number_type = token_type::value_float;
dot_index = i;
++i;
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
// the mantissa ends here, whether or not an exponent part follows
const std::size_t mantissa_end = i;
if (i < avail && (data[i] == 'e' || data[i] == 'E'))
{
number_type = token_type::value_float;
++i;
if (i < avail && (data[i] == '+' || data[i] == '-'))
{
++i;
}
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
const std::size_t len = i;
// reset() records where this token starts (for diagnostics), so it has
// to run before the input position advances below
reset();
// An integer token needs no token_buffer: the SAX callbacks for
// number_integer/number_unsigned take only the value, and the overflow
// diagnostic rebuilds the text from the input. Convert straight from the
// input buffer and leave token_buffer empty. (JSON_DIAGNOSTIC_POSITIONS
// derives a number's start position from get_string().size(), so there
// the token still has to be materialized.)
#if !JSON_DIAGNOSTIC_POSITIONS
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, data, data + len);
if (JSON_HEDLEY_LIKELY(integer_result != token_type::uninitialized))
{
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return integer_result;
}
// The value does not fit an integer, so this token converts as a
// float. Recording that here keeps convert_number() below from
// repeating the integer attempt that just failed.
number_type = token_type::value_float;
}
#endif
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return convert_number(number_type, mantissa_end);
}
/// contiguous input: try the number fast path, else the byte-path scanner
token_type scan_number_dispatch(std::true_type /*bulk*/)
{
const token_type t = scan_number_bulk_contiguous();
return (t != token_type::uninitialized) ? t : scan_number();
}
/// streaming input: always use the byte-path scanner
token_type scan_number_dispatch(std::false_type /*bulk*/)
{
return scan_number();
}
/*!
@param[in] literal_text the literal text to expect
@param[in] length the length of the passed literal text
@@ -1393,8 +1780,7 @@ scan_number_done:
*/
char_int_type get()
{
++position.chars_read_total;
++position.chars_read_current_line;
advance_position();
if (next_unget)
{
@@ -1406,6 +1792,23 @@ scan_number_done:
current = ia.get_character();
}
return track_after_read();
}
/// shared head of get() / get_ignoring_pending_unget(): bump the
/// per-character position counters (line-count-on-'\n' bookkeeping is
/// handled afterwards, in track_after_read(), once `current` is known)
void advance_position() noexcept
{
++position.chars_read_total;
++position.chars_read_current_line;
}
/// shared tail of get() / get_ignoring_pending_unget(): capture the
/// character for error messages (if needed) and update line/column
/// bookkeeping for the character now in `current`
char_int_type track_after_read()
{
// seekable adapters reconstruct the token lazily on error (see
// get_token_string), so the eager per-character copy is skipped
capture_char(std::integral_constant<bool, lazy_token_string> {});
@@ -1413,12 +1816,38 @@ scan_number_done:
if (current == '\n')
{
++position.lines_read;
// remember the column the newline was read at: chars_read_current_line
// is about to be cleared, and a matching unget() cannot reconstruct it
chars_read_before_newline = position.chars_read_current_line;
position.chars_read_current_line = 0;
}
return current;
}
/*!
@brief like get(), but for call sites that can prove no unget() is pending
get() has to check the `next_unget` flag on every call, because a
previous token may have ended with unget() (e.g. scan_number() always
ungets the character that terminated the number, so the next call to
scan() can see it again). skip_whitespace() reads that first,
possibly-ungotten character via a plain get(), but every further
character it reads is guaranteed to be a fresh read: nothing between
those calls invokes unget(). This variant skips the (otherwise always
false) next_unget branch for those calls; it is not a general
replacement for get().
*/
char_int_type get_ignoring_pending_unget()
{
JSON_ASSERT(!next_unget);
advance_position();
current = ia.get_character();
return track_after_read();
}
/// seekable adapter: nothing to capture, the token is rebuilt on error
void capture_char(std::true_type /*lazy*/) const noexcept {}
@@ -1446,12 +1875,20 @@ scan_number_done:
--position.chars_read_total;
// in case we "unget" a newline, we have to also decrement the lines_read
// and restore the column that get() cleared when it saw the newline;
// chars_read_current_line == 0 can only mean the last get() read one
if (position.chars_read_current_line == 0)
{
if (position.lines_read > 0)
{
--position.lines_read;
}
// chars_read_before_newline counts the newline itself, which is the
// character being ungotten, hence the -1
position.chars_read_current_line = (chars_read_before_newline > 0)
? chars_read_before_newline - 1
: 0;
}
else
{
@@ -1612,13 +2049,37 @@ scan_number_done:
return true;
}
/// whether `current` is one of the four JSON whitespace characters
bool current_is_whitespace() const noexcept
{
return current == ' ' || current == '\t' || current == '\n' || current == '\r';
}
void skip_whitespace()
{
// the first character may be a pending unget() left over from the
// previous token (see get_ignoring_pending_unget()); every
// subsequent character read by this loop is guaranteed fresh, since
// nothing below calls unget()
get();
if (!current_is_whitespace())
{
return;
}
// this is written as an if-guarded do-while (rather than a plain
// while loop) because that shape is what lets both GCC and Clang
// keep the input adapter's read pointer in a register across
// iterations; the equivalent while-loop measurably defeated that
// optimization in testing, turning long whitespace runs (e.g. the
// indentation of pretty-printed JSON) from a register-only loop
// into one that reloads the pointer from memory every character
do
{
get();
get_ignoring_pending_unget();
}
while (current == ' ' || current == '\t' || current == '\n' || current == '\r');
while (current_is_whitespace());
}
token_type scan()
@@ -1694,11 +2155,14 @@ scan_number_done:
case '7':
case '8':
case '9':
return scan_number();
return scan_number_dispatch(std::integral_constant<bool, bulk_scan> {});
// end of input (the null byte is needed when parsing from
// string literals)
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
// end of input; by default, a null byte is also treated as end of
// input for backwards compatibility (see JSON_STRICT_NUL_HANDLING
// to opt into rejecting a null byte in the input instead)
case char_traits<char_type>::eof():
return token_type::end_of_input;
@@ -1725,6 +2189,10 @@ scan_number_done:
/// the start position of the current token
position_t position {};
/// the value chars_read_current_line had when the last newline was read, so
/// that unget() can restore the column instead of leaving it at 0
std::size_t chars_read_before_newline = 0;
/// raw input token string for error messages; only populated for streaming
/// adapters (seekable adapters reconstruct it lazily via token_string_start)
std::vector<char_type> token_string {};
@@ -1754,6 +2222,13 @@ scan_number_done:
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
/// token classification and never looks at the converted numeric value;
/// when set, scan_number() may skip strtoull()/strtoll() for
/// value_unsigned/value_integer tokens whose digit count guarantees they
/// fit into 64 bits (see scan_number())
const bool discard_number_values = false;
};
} // namespace detail
@@ -0,0 +1,302 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint64_t
#include <limits> // numeric_limits
#include <nlohmann/detail/macro_scope.hpp>
// std::from_chars lives in <charconv>, but being in C++17 mode does not
// guarantee the header exists: GCC 7 sets __cplusplus to C++17 yet ships no
// <charconv> (added in GCC 8; floating-point support in GCC 11). Guard the
// include with __has_include so such toolchains fall back to the scalar path.
#if defined(JSON_HAS_CPP_17) && defined(__has_include)
#if __has_include(<charconv>)
#include <charconv> // from_chars (only used when __cpp_lib_to_chars is defined)
#include <system_error> // errc
#endif
#endif
// This file contains the value-conversion helpers used by the lexer to turn an
// already-validated number token into a value, without the locale/errno
// overhead of std::strtoull/std::strtod. They are free functions so the lexer
// stays focused on scanning; see lexer::convert_number().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief fast integer parser for an already-validated unsigned integer
The number scanner has already checked that [first, last) is a valid JSON
integer, so this only needs to accumulate the digits and detect overflow. This
avoids the locale/errno machinery of std::strtoull, which dominates
integer-heavy inputs.
@param[in] first pointer to the first character (a digit)
@param[in] last pointer past the last character
@param[out] value the parsed value on success
@return true if the value fit into @a NumberUnsignedType; false on overflow, in
which case the caller falls back to floating-point parsing (matching the
previous std::strtoull behavior)
*/
template<typename NumberUnsignedType>
bool parse_integer_unsigned(const char* first, const char* last, NumberUnsignedType& value) noexcept
{
// accumulate in the widest unsigned type used by the previous strtoull
// path so the overflow behavior is unchanged for custom number types
std::uint64_t x = 0;
constexpr std::uint64_t cutoff = (std::numeric_limits<std::uint64_t>::max)() / 10u;
constexpr std::uint64_t cutlim = (std::numeric_limits<std::uint64_t>::max)() % 10u;
for (const char* p = first; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(x > cutoff || (x == cutoff && digit > cutlim)))
{
return false;
}
x = (x * 10u) + digit;
}
value = static_cast<NumberUnsignedType>(x);
// reject values that do not round-trip into a narrower NumberUnsignedType
return static_cast<std::uint64_t>(value) == x;
}
/*!
@brief fast integer parser for an already-validated negative integer
@param[in] first pointer to the leading '-'
@param[in] last pointer past the last character
@param[out] value the parsed (negative) value on success
@return true on success; false on overflow (caller falls back to float)
*/
template<typename NumberIntegerType>
bool parse_integer_signed(const char* first, const char* last, NumberIntegerType& value) noexcept
{
// the state machine only reaches the signed path via a leading '-'
JSON_ASSERT(first != last && *first == '-');
std::uint64_t magnitude = 0;
// |INT64_MIN| == INT64_MAX + 1; this is the largest admissible magnitude
constexpr std::uint64_t limit = static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()) + 1u;
for (const char* p = first + 1; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(magnitude > (limit - digit) / 10u))
{
return false;
}
magnitude = (magnitude * 10u) + digit;
}
const std::int64_t x = (magnitude == limit)
? (std::numeric_limits<std::int64_t>::min)()
: -static_cast<std::int64_t>(magnitude);
value = static_cast<NumberIntegerType>(x);
// reject values that do not round-trip into a narrower NumberIntegerType
return static_cast<std::int64_t>(value) == x;
}
/*!
@brief exact fast path for parsing a `double` (Clinger's algorithm)
For the common case - at most 19 significant digits, a decimal exponent in
[-22, 22], and a significand below 2^53 - the value equals significand *
10^exp computed in IEEE-754 double arithmetic, which is exact under
round-to-nearest because both operands are exactly representable. This is the
same fast path used by fast_float/simdjson; the general cases are left to
std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
// true double precision. On platforms that keep intermediates in extended
// precision (e.g. the x87 FPU on 32-bit x86, where FLT_EVAL_METHOD == 2) the
// single significand * 10^scale step is double-rounded and can be 1 ULP off,
// so decline and let the caller fall back to the correctly-rounded
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
static const std::array<double, 23> powers_of_ten =
{
{
1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11,
1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22
}
};
const char* p = first;
bool negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
negative = (*p == '-');
++p;
}
std::uint64_t significand = 0;
int num_digits = 0;
int fractional_digits = 0;
bool seen_dot = false;
bool any_digit = false;
for (; p != last; ++p)
{
const char c = *p;
if (c >= '0' && c <= '9')
{
any_digit = true;
if (JSON_HEDLEY_UNLIKELY(num_digits >= 19))
{
return false; // significand may not fit into uint64_t
}
significand = (significand * 10u) + static_cast<std::uint64_t>(c - '0');
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (static_cast<DecimalPointType>(c) == decimal_point)
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
return false;
}
seen_dot = true;
}
else if (c == 'e' || c == 'E')
{
++p;
break;
}
else
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_digit))
{
return false;
}
int exponent = 0;
if (p != last) // an exponent part remains
{
bool exp_negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
exp_negative = (*p == '-');
++p;
}
bool any_exp_digit = false;
for (; p != last; ++p)
{
if (JSON_HEDLEY_UNLIKELY(*p < '0' || *p > '9'))
{
return false;
}
exponent = (exponent * 10) + (*p - '0');
any_exp_digit = true;
if (JSON_HEDLEY_UNLIKELY(exponent > 9999))
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_exp_digit))
{
return false;
}
if (exp_negative)
{
exponent = -exponent;
}
}
const int scale = exponent - fractional_digits;
if (JSON_HEDLEY_UNLIKELY(significand >= (static_cast<std::uint64_t>(1) << 53)))
{
return false; // significand not exactly representable as double
}
auto result = static_cast<double>(significand);
if (scale >= 0)
{
if (JSON_HEDLEY_UNLIKELY(scale > 22))
{
return false;
}
result *= powers_of_ten[static_cast<std::size_t>(scale)];
}
else
{
if (JSON_HEDLEY_UNLIKELY(-scale > 22))
{
return false;
}
result /= powers_of_ten[static_cast<std::size_t>(-scale)];
}
out = negative ? -result : result;
return true;
#endif
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
{
return false;
}
/*!
@brief parse a float with std::from_chars (Eisel-Lemire) when available
std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
@return true if the value was parsed exactly and fully; false to fall back
*/
template<typename FloatType>
bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept
{
// JSON_HAS_CPP_17 must gate the use as well as the <charconv> include above:
// some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even
// in C++14 mode, where <charconv> is not included.
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
const auto result = std::from_chars(first, last, out);
return result.ec == std::errc() && result.ptr == last;
#else
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(out);
return false;
#endif
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+3 -2
View File
@@ -72,9 +72,10 @@ class parser
parser_callback_t<BasicJsonType> cb = nullptr,
const bool allow_exceptions_ = true,
const bool ignore_comments = false,
const bool ignore_trailing_commas_ = false)
const bool ignore_trailing_commas_ = false,
const bool discard_number_values_ = false)
: callback(std::move(cb))
, m_lexer(std::move(adapter), ignore_comments)
, m_lexer(std::move(adapter), ignore_comments, discard_number_values_)
, allow_exceptions(allow_exceptions_)
, ignore_trailing_commas(ignore_trailing_commas_)
{
@@ -0,0 +1,287 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef> // size_t
#include <cstdint> // uint64_t
#include <cstring> // memcpy
#include <nlohmann/detail/macro_scope.hpp>
// Optional SIMD backend for bulk UTF-8 validation. This is an opt-in external
// dependency: nlohmann/json itself stays header-only and the C++11 scalar
// validator below is always available; defining JSON_USE_SIMDUTF additionally
// requires the simdutf headers on the include path and linking the simdutf
// library. See string_bulk_run().
//
// simdutf.h itself requires C++17 - it rejects older standards with an #error -
// so the backend is only compiled in from C++17 on. Below that the macro has no
// effect and the scalar validator is used; it accepts and rejects exactly the
// same input, so only throughput differs. macro_scope.hpp is included above to
// have JSON_HAS_CPP_17 available for this test.
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
#include <simdutf.h>
#endif
// This file contains the byte-level string-scanning helpers used by the lexer's
// contiguous fast path. They operate purely on raw bytes (no dependency on the
// lexer's template parameters) so they are free functions, keeping the lexer
// itself focused on the state machine; see lexer::scan_string_bulk().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
// classify a single byte as needing individual string handling: the closing
// quote, an escape, a control character, or a non-ASCII (UTF-8)
// lead/continuation byte. Ordinary bytes (0x20..0x7F except '"' and '\\') are
// copied verbatim, which the bulk scanner does 8 bytes at a time.
inline bool is_string_special(unsigned char c) noexcept
{
return c == '\"' || c == '\\' || c < 0x20u || c >= 0x80u;
}
// SWAR helper: return a word whose high bit is set in every byte of @a v that
// is_string_special(); zero if the 8 bytes are all ordinary.
inline std::uint64_t swar_string_special(std::uint64_t v) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
const std::uint64_t has_quote = (q - ones) & ~q & high;
const std::uint64_t has_backslash = (b - ones) & ~b & high;
const std::uint64_t has_control = (v - 0x2020202020202020ull) & ~v & high; // < 0x20
const std::uint64_t has_non_ascii = v & high; // >= 0x80
return has_quote | has_backslash | has_control | has_non_ascii;
}
// return the index of the first is_string_special() byte in [data, data+n), or
// n if every byte is ordinary; scans 8 bytes at a time
inline std::size_t find_string_special(const unsigned char* data, std::size_t n) noexcept
{
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t word = 0;
std::memcpy(&word, data + i, sizeof(word));
if (swar_string_special(word) != 0)
{
// a special byte is in this word; locate it (endian-agnostic)
for (std::size_t j = 0; j < 8; ++j)
{
if (is_string_special(data[i + j]))
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
if (is_string_special(data[i]))
{
return i;
}
}
return n;
}
// classify a byte as one the serializer must NOT copy verbatim when
// ensure_ascii is requested: the closing quote, an escape, a control character
// (< 0x20), DEL (0x7F), or any non-ASCII byte (>= 0x80). Everything else -
// printable ASCII except '"' and '\\' - is emitted unchanged. Note this differs
// from is_string_special() only in that 0x7F is also a stop (it is escaped as
// \u007f under ensure_ascii).
inline bool is_ascii_copyable(unsigned char c) noexcept
{
return c >= 0x20u && c < 0x7Fu && c != '"' && c != '\\';
}
// return the index of the first byte in [data, data+n) that is NOT
// is_ascii_copyable(), or n if every byte can be copied verbatim; scans 8 bytes
// at a time. Used by the serializer's ensure_ascii fast path.
inline std::size_t find_ascii_copyable_run(const unsigned char* data, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t v = 0;
std::memcpy(&v, data + i, sizeof(v));
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
const std::uint64_t d = v ^ 0x7F7F7F7F7F7F7F7Full; // DEL (0x7F)
const std::uint64_t stop = ((q - ones) & ~q & high) // == '"'
| ((b - ones) & ~b & high) // == '\\'
| ((d - ones) & ~d & high) // == 0x7F
| ((v - 0x2020202020202020ull) & ~v & high) // < 0x20
| (v & high); // >= 0x80
if (stop != 0)
{
break;
}
}
for (; i < n; ++i)
{
if (!is_ascii_copyable(data[i]))
{
return i;
}
}
return n;
}
// Validate one UTF-8 sequence at the front of [data, data+avail). Returns its
// length (2..4) only when the bytes form a *well-formed* sequence using exactly
// the same ranges as scan_string()'s per-byte switch, so the bulk path accepts
// precisely what the byte path accepts. Returns 0 for anything that is invalid,
// incomplete, or that the byte path must diagnose (the caller then defers to
// that path, keeping error messages unchanged). Lead bytes < 0x80 are handled
// by the caller and never passed here.
inline std::size_t validate_one_utf8(const unsigned char* data, std::size_t avail) noexcept
{
const unsigned char c0 = data[0];
if (c0 >= 0xC2 && c0 <= 0xDF) // U+0080..U+07FF
{
if (avail >= 2 && data[1] >= 0x80 && data[1] <= 0xBF)
{
return 2;
}
}
else if (c0 == 0xE0) // U+0800..U+0FFF
{
if (avail >= 3 && data[1] >= 0xA0 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if ((c0 >= 0xE1 && c0 <= 0xEC) || c0 == 0xEE || c0 == 0xEF) // U+1000..U+CFFF, U+E000..U+FFFF
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xED) // U+D000..U+D7FF (excludes surrogates)
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0x9F && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xF0) // U+10000..U+3FFFF
{
if (avail >= 4 && data[1] >= 0x90 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 >= 0xF1 && c0 <= 0xF3) // U+40000..U+FFFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 == 0xF4) // U+100000..U+10FFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0x8F && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
return 0; // invalid, incomplete, or must be diagnosed by the byte path
}
// Scalar (C++11) computation of the bulk run length: the number of leading
// bytes in [data, data+n) that are ordinary ASCII or complete well-formed UTF-8
// sequences, stopping before the first byte that needs individual handling (the
// closing quote, an escape, a control character, or an ill-formed/truncated
// sequence). ASCII is skipped 8 bytes at a time.
inline std::size_t scalar_string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
std::size_t pos = 0;
while (pos < n)
{
pos += find_string_special(data + pos, n - pos);
if (pos >= n || data[pos] < 0x80u)
{
break; // end of buffer, or a quote/escape/control byte
}
const std::size_t seq = validate_one_utf8(data + pos, n - pos);
if (seq == 0)
{
break; // ill-formed or truncated: let the byte path diagnose it
}
pos += seq;
}
return pos;
}
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
// Index of the first quote/escape/control byte in [data, data+n) (non-ASCII
// bytes are *not* stops here - the whole run is handed to simdutf), or n.
inline std::size_t find_string_delimiter(const unsigned char* data, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t v = 0;
std::memcpy(&v, data + i, sizeof(v));
const std::uint64_t q = v ^ 0x2222222222222222ull;
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull;
const std::uint64_t hit = ((q - ones) & ~q & high)
| ((b - ones) & ~b & high)
| ((v - 0x2020202020202020ull) & ~v & high);
if (hit != 0)
{
for (std::size_t j = 0; j < 8; ++j)
{
const unsigned char c = data[i + j];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
const unsigned char c = data[i];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i;
}
}
return n;
}
#endif
// Backend-dispatched bulk run length. With JSON_USE_SIMDUTF the run up to the
// next delimiter is validated in one shot by simdutf; on the rare failure the
// scalar helper recomputes the exact valid prefix so the byte path still
// produces the precise diagnostic. Without it, the pure scalar path is used.
inline std::size_t string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
const std::size_t run = find_string_delimiter(data, n);
if (run != 0 && simdutf::validate_utf8(reinterpret_cast<const char*>(data), run))
{
return run;
}
#endif
return scalar_string_bulk_run(data, n);
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -88,8 +88,13 @@ class iter_impl // NOLINT(cppcoreguidelines-special-member-functions,hicpp-speci
iter_impl() = default;
~iter_impl() = default;
iter_impl(iter_impl&&) noexcept = default;
iter_impl& operator=(iter_impl&&) noexcept = default;
// the exception specification is left to be computed rather than declared:
// an array or object type whose iterator is not nothrow move constructible
// (std::deque's is not before libstdc++ 11) would make a declared noexcept
// differ from the implicit one, which deletes the function -- and is an
// error outright with older compilers
iter_impl(iter_impl&&) = default; // NOLINT(hicpp-noexcept-move,performance-noexcept-move-constructor,cppcoreguidelines-noexcept-move-operations)
iter_impl& operator=(iter_impl&&) = default; // NOLINT(hicpp-noexcept-move,performance-noexcept-move-constructor,cppcoreguidelines-noexcept-move-operations)
/*!
@brief constructor for a given JSON instance
@@ -18,6 +18,7 @@
#endif
#include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/string_utils.hpp>
#include <nlohmann/detail/value_t.hpp>
@@ -206,10 +207,10 @@ NLOHMANN_JSON_NAMESPACE_END
namespace std
{
// Fix: https://github.com/nlohmann/json/issues/1401
#if defined(__clang__)
// Fix: https://github.com/nlohmann/json/issues/1401
#pragma clang diagnostic push
#pragma clang diagnostic ignored "-Wmismatched-tags"
JSON_HEDLEY_DIAGNOSTIC_PUSH
JSON_HEDLEY_PRAGMA(clang diagnostic ignored "-Wmismatched-tags")
#endif
template<typename IteratorType>
class tuple_size<::nlohmann::detail::iteration_proxy_value<IteratorType>> // NOLINT(cert-dcl58-cpp)
@@ -224,7 +225,7 @@ class tuple_element<N, ::nlohmann::detail::iteration_proxy_value<IteratorType >>
::nlohmann::detail::iteration_proxy_value<IteratorType >> ()));
};
#if defined(__clang__)
#pragma clang diagnostic pop
JSON_HEDLEY_DIAGNOSTIC_POP
#endif
} // namespace std
+61 -8
View File
@@ -17,6 +17,7 @@
#endif // JSON_NO_IO
#include <limits> // max
#include <numeric> // accumulate
#include <set> // set
#include <string> // string
#include <utility> // move
#include <vector> // vector
@@ -71,7 +72,7 @@ class json_pointer
string_t{},
[](const string_t& a, const string_t& b)
{
return detail::concat(a, '/', detail::escape(b));
return detail::concat<string_t>(a, '/', detail::escape(b));
});
}
@@ -265,7 +266,7 @@ class json_pointer
JSON_THROW(detail::parse_error::create(109, 0, detail::concat("array index '", s, "' is not a number"), nullptr));
}
const char* p = s.c_str();
const char* p = s.data();
char* p_end = nullptr; // NOLINT(misc-const-correctness)
errno = 0; // strtoull doesn't reset errno
const unsigned long long res = std::strtoull(p, &p_end, 10); // NOLINT(runtime/int)
@@ -300,19 +301,35 @@ class json_pointer
}
private:
/*!
@brief the reference token sequences that denote arrays
@ref unflatten collects the pointer prefixes that have a reference token 0
among their children; @ref get_and_create creates arrays exactly below
those prefixes and objects everywhere else. Deciding this up front keeps
the result independent of the order in which the flattened object is
iterated, which is unspecified for some object types.
*/
using array_parents_t = std::set<std::vector<string_t>>;
/*!
@brief create and return a reference to the pointed to value
@complexity Linear in the number of reference tokens.
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if array index is not a number
@throw type_error.313 if value cannot be unflattened
*/
template<typename BasicJsonType>
BasicJsonType& get_and_create(BasicJsonType& j) const
BasicJsonType& get_and_create(BasicJsonType& j, const array_parents_t& array_parents) const
{
auto* result = &j;
// the reference tokens that have been consumed so far; used to look up
// whether the value to be created below is an array or an object
std::vector<string_t> prefix;
// in case no reference tokens exist, return a reference to the JSON value
// j which will be overwritten by a primitive value
for (const auto& reference_token : reference_tokens)
@@ -321,10 +338,11 @@ class json_pointer
{
case detail::value_t::null:
{
if (reference_token == "0")
if (array_parents.find(prefix) != array_parents.end())
{
// start a new array if the reference token is 0
result = &result->operator[](0);
// some reference token below this position is 0, so the
// value is an array
result = &result->operator[](array_index<BasicJsonType>(reference_token));
}
else
{
@@ -364,6 +382,8 @@ class json_pointer
default:
JSON_THROW(detail::type_error::create(313, "invalid value to unflatten", &j));
}
prefix.push_back(reference_token);
}
return *result;
@@ -748,6 +768,20 @@ class json_pointer
}
}
// the reference token consists only of digits at this point (cf. checks
// above); however, its numeric value might not be representable, in which
// case array_index() would throw out_of_range.404/410 -- contains() must
// not throw (see #5395), so such a reference token is treated as "not found"
errno = 0; // strtoull() does not reset errno on success
char* p_end = nullptr; // NOLINT(misc-const-correctness)
const unsigned long long magnitude = std::strtoull(reference_token.c_str(), &p_end, 10); // NOLINT(runtime/int)
if (JSON_HEDLEY_UNLIKELY(errno == ERANGE // the value exceeds ULLONG_MAX
|| magnitude >= static_cast<unsigned long long>((std::numeric_limits<typename BasicJsonType::size_type>::max)()))) // NOLINT(runtime/int)
{
// the array index cannot be represented as size_type
return false;
}
const auto idx = array_index<BasicJsonType>(reference_token);
if (idx >= ptr->size())
{
@@ -823,7 +857,8 @@ class json_pointer
{
// use the text between the beginning of the reference token
// (start) and the last slash (slash).
auto reference_token = reference_string.substr(start, slash - start);
const auto count = (slash == string_t::npos ? reference_string.size() : slash) - start;
auto reference_token = string_t(reference_string.data() + start, count);
// check reference tokens are properly escaped
for (std::size_t pos = reference_token.find_first_of('~');
@@ -939,6 +974,24 @@ class json_pointer
BasicJsonType result;
// collect the pointer prefixes that have a reference token 0 among
// their children; the values below them are arrays, all others are
// objects (see array_parents_t)
array_parents_t array_parents;
for (const auto& element : *value.m_data.m_value.object)
{
json_pointer ptr(element.first);
std::vector<string_t> prefix;
for (auto& reference_token : ptr.reference_tokens)
{
if (reference_token == "0")
{
array_parents.insert(prefix);
}
prefix.push_back(std::move(reference_token));
}
}
// iterate the JSON object values
for (const auto& element : *value.m_data.m_value.object)
{
@@ -951,7 +1004,7 @@ class json_pointer
// that if the JSON pointer is "" (i.e., points to the whole value),
// function get_and_create returns a reference to the result itself.
// An assignment will then create a primitive value.
json_pointer(element.first).get_and_create(result) = element.second;
json_pointer(element.first).get_and_create(result, array_parents) = element.second;
}
return result;
+4
View File
@@ -807,3 +807,7 @@ void templated_json_throw(ExceptionType exception)
#ifndef JSON_BRACE_INIT_COPY_SEMANTICS
#define JSON_BRACE_INIT_COPY_SEMANTICS 0
#endif
#ifndef JSON_STRICT_NUL_HANDLING
#define JSON_STRICT_NUL_HANDLING 0
#endif
@@ -27,6 +27,7 @@
#undef JSON_DISABLE_ENUM_SERIALIZATION
#undef JSON_USE_GLOBAL_UDLS
#undef JSON_BRACE_INIT_COPY_SEMANTICS
#undef JSON_STRICT_NUL_HANDLING
#ifndef JSON_TEST_KEEP_MACROS
#undef JSON_CATCH
+23 -6
View File
@@ -172,17 +172,18 @@ struct has_to_json < BasicJsonType, T, enable_if_t < !is_basic_json<T>::value >>
template<typename T>
using detect_key_compare = typename T::key_compare;
template<typename T>
struct has_key_compare : std::integral_constant<bool, is_detected<detect_key_compare, T>::value> {};
// obtains the actual object key comparator
// obtains the actual object key comparator: object_t::key_compare if the
// object type defines it, and default_object_comparator_t otherwise
//
// note detected_or_t is used rather than std::conditional, because the latter
// names both of its type arguments eagerly; object_t::key_compare would then
// be a hard error for an object type that does not define it
template<typename BasicJsonType>
struct actual_object_comparator
{
using object_t = typename BasicJsonType::object_t;
using object_comparator_t = typename BasicJsonType::default_object_comparator_t;
using type = typename std::conditional < has_key_compare<object_t>::value,
typename object_t::key_compare, object_comparator_t>::type;
using type = detected_or_t<object_comparator_t, detect_key_compare, object_t>;
};
template<typename BasicJsonType>
@@ -778,6 +779,22 @@ using has_erase_with_key_type = typename std::conditional <
std::true_type,
std::false_type >::type;
template<typename ObjectType, typename IteratorType>
using detect_erase_with_iterator = decltype(std::declval<ObjectType&>().erase(std::declval<IteratorType>()));
// type trait to check if erase(iterator) returns void instead of the following
// iterator, as the object types that do not compute a successor the caller may
// not need do
template<typename ObjectType, typename IteratorType>
using erase_returns_void = is_detected_exact<void, detect_erase_with_iterator, ObjectType, IteratorType>;
template<typename T>
using detect_capacity = decltype(std::declval<const T&>().capacity());
// type trait to check if a type has a capacity() member function
template<typename T>
struct has_capacity : std::integral_constant<bool, is_detected<detect_capacity, T>::value> {};
// a naive helper to check if a type is an ordered_map (exploits the fact that
// ordered_map inherits capacity() from std::vector)
template <typename T>
File diff suppressed because it is too large Load Diff
@@ -13,6 +13,7 @@
#include <iterator> // back_inserter
#include <memory> // shared_ptr, make_shared
#include <string> // basic_string
#include <utility> // move
#include <vector> // vector
#ifndef JSON_NO_IO
@@ -44,22 +45,32 @@ template<typename CharType> struct output_adapter_protocol
template<typename CharType>
using output_adapter_t = std::shared_ptr<output_adapter_protocol<CharType>>;
/// output adapter for byte vectors
/// @brief non-virtual output sink writing into a std::vector
///
/// This sink is not part of the virtual output_adapter_protocol hierarchy: it is
/// passed to binary_writer by value as a template parameter, so
/// write_character()/write_characters() are ordinary (inlinable) calls with no
/// vtable lookup and no shared_ptr. It is used for the common
/// `to_cbor`/`to_msgpack`/... into a std::vector. output_vector_adapter below
/// wraps this same sink to provide the virtual interface.
template<typename CharType, typename AllocatorType = std::allocator<CharType>>
class output_vector_adapter : public output_adapter_protocol<CharType>
class output_vector_sink
{
public:
explicit output_vector_adapter(std::vector<CharType, AllocatorType>& vec) noexcept
explicit output_vector_sink(std::vector<CharType, AllocatorType>& vec) noexcept
: v(vec)
{}
void write_character(CharType c) override
void write_character(CharType c)
{
v.push_back(c);
}
JSON_HEDLEY_NON_NULL(2)
void write_characters(const CharType* s, std::size_t length) override
// no JSON_HEDLEY_NON_NULL here: binary_writer legitimately passes a null
// pointer with length 0 for empty strings/binary values. Appending an empty
// range is a no-op; the type-erased path tolerates this via the (unattributed)
// virtual base, and the concrete sink must do the same.
void write_characters(const CharType* s, std::size_t length)
{
v.insert(v.end(), s, s + length);
}
@@ -68,6 +79,34 @@ class output_vector_adapter : public output_adapter_protocol<CharType>
std::vector<CharType, AllocatorType>& v;
};
/// output adapter for byte vectors
///
/// The appending itself lives in output_vector_sink; this class only adds the
/// virtual output_adapter_protocol interface on top of it, so both the
/// type-erased and the templated path share one implementation.
template<typename CharType, typename AllocatorType = std::allocator<CharType>>
class output_vector_adapter : public output_adapter_protocol<CharType>
{
public:
explicit output_vector_adapter(std::vector<CharType, AllocatorType>& vec) noexcept
: sink(vec)
{}
void write_character(CharType c) override
{
sink.write_character(c);
}
JSON_HEDLEY_NON_NULL(2)
void write_characters(const CharType* s, std::size_t length) override
{
sink.write_characters(s, length);
}
private:
output_vector_sink<CharType, AllocatorType> sink;
};
#ifndef JSON_NO_IO
/// output adapter for output streams
template<typename CharType>
@@ -118,6 +157,39 @@ class output_string_adapter : public output_adapter_protocol<CharType>
StringType& str;
};
/// @brief output sink forwarding to a type-erased output adapter
///
/// Wraps the polymorphic output_adapter_t so the same binary_writer template can
/// also target arbitrary adapters (output streams, strings, user-provided
/// adapters) via the `output_adapter`-based overloads. Each write still goes
/// through one virtual call, exactly as before; only the concrete sinks above
/// avoid it.
template<typename CharType>
class output_adapter_sink
{
public:
explicit output_adapter_sink(output_adapter_t<CharType> adapter)
: oa(std::move(adapter))
{
JSON_ASSERT(oa);
}
void write_character(CharType c)
{
oa->write_character(c);
}
// no JSON_HEDLEY_NON_NULL: forwards (null, 0) for empty payloads, exactly as
// the type-erased path already did before this sink existed
void write_characters(const CharType* s, std::size_t length)
{
oa->write_characters(s, length);
}
private:
output_adapter_t<CharType> oa;
};
template<typename CharType, typename StringType = std::basic_string<CharType>>
class output_adapter
{
File diff suppressed because it is too large Load Diff
+68 -31
View File
@@ -8,50 +8,56 @@
#pragma once
#include <cstddef> // size_t
#include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief replace all occurrences of a substring by another string
@param[in,out] s the string to manipulate; changed so that all
occurrences of @a f are replaced with @a t
@param[in] f the substring to replace with @a t
@param[in] t the string to replace @a f
@pre The search string @a f must not be empty. **This precondition is
enforced with an assertion.**
@since version 2.0.0
*/
template<typename StringType>
inline void replace_substring(StringType& s, const StringType& f,
const StringType& t)
{
JSON_ASSERT(!f.empty());
for (auto pos = s.find(f); // find the first occurrence of f
pos != StringType::npos; // make sure f was found
s.replace(pos, f.size(), t), // replace with t, and
pos = s.find(f, pos + t.size())) // find the next occurrence of f
{}
}
/*!
* @brief string escaping as described in RFC 6901 (Sect. 4)
* @param[in] s string to escape
* @return escaped string
*
* Note the order of escaping "~" to "~0" and "/" to "~1" is important.
*
* The string is rebuilt in a single pass, appending whole runs between the
* characters that need escaping. Scanning with find_first_of() keeps the
* common case -- nothing to escape -- as fast as a single search, while
* repeated replace() calls would move the tail of the string once per
* escaped character.
*/
template<typename StringType>
inline StringType escape(StringType s)
inline StringType escape(const StringType& s)
{
replace_substring(s, StringType{"~"}, StringType{"~0"});
replace_substring(s, StringType{"/"}, StringType{"~1"});
return s;
auto next_special = [&s](std::size_t from)
{
const auto tilde = s.find_first_of('~', from);
const auto slash = s.find_first_of('/', from);
return tilde < slash ? tilde : slash; // npos is the largest value
};
auto pos = next_special(0);
if (pos == StringType::npos)
{
return s;
}
StringType result;
result.reserve(s.size() + 2);
std::size_t run = 0;
while (pos != StringType::npos)
{
result.append(s.data() + run, pos - run);
result.append(s[pos] == '~' ? "~0" : "~1", 2);
run = pos + 1;
pos = next_special(run);
}
result.append(s.data() + run, s.size() - run);
return result;
}
/*!
@@ -60,12 +66,43 @@ inline StringType escape(StringType s)
* @return unescaped string
*
* Note the order of escaping "~1" to "/" and "~0" to "~" is important.
*
* Rebuilt in a single pass, see @ref escape. A "~" that is followed by
* neither "0" nor "1" is passed through unchanged; @ref json_pointer rejects
* such input before it gets here.
*/
template<typename StringType>
inline void unescape(StringType& s)
{
replace_substring(s, StringType{"~1"}, StringType{"/"});
replace_substring(s, StringType{"~0"}, StringType{"~"});
auto pos = s.find_first_of('~', 0);
if (pos == StringType::npos)
{
return;
}
StringType result;
result.reserve(s.size());
std::size_t run = 0;
while (pos != StringType::npos)
{
result.append(s.data() + run, pos - run);
const auto next = pos + 1;
if (next < s.size() && (s[next] == '0' || s[next] == '1'))
{
result.append(s[next] == '0' ? "~" : "/", 1);
run = pos + 2;
}
else
{
result.append("~", 1);
run = pos + 1;
}
pos = s.find_first_of('~', run);
}
result.append(s.data() + run, s.size() - run);
s = result;
}
} // namespace detail
+410 -111
View File
File diff suppressed because it is too large Load Diff
+4 -1
View File
@@ -17,7 +17,7 @@
#undef JSON_HEDLEY_CLANG_HAS_ATTRIBUTE
#undef JSON_HEDLEY_CLANG_HAS_BUILTIN
#undef JSON_HEDLEY_CLANG_HAS_CPP_ATTRIBUTE
#undef JSON_HEDLEY_CLANG_HAS_DECLSPEC_DECLSPEC_ATTRIBUTE
#undef JSON_HEDLEY_CLANG_HAS_DECLSPEC_ATTRIBUTE
#undef JSON_HEDLEY_CLANG_HAS_EXTENSION
#undef JSON_HEDLEY_CLANG_HAS_FEATURE
#undef JSON_HEDLEY_CLANG_HAS_WARNING
@@ -108,7 +108,10 @@
#undef JSON_HEDLEY_PELLES_VERSION_CHECK
#undef JSON_HEDLEY_PGI_VERSION
#undef JSON_HEDLEY_PGI_VERSION_CHECK
#undef JSON_HEDLEY_PRAGMA
#undef JSON_HEDLEY_PREDICT
#undef JSON_HEDLEY_PREDICT_FALSE
#undef JSON_HEDLEY_PREDICT_TRUE
#undef JSON_HEDLEY_PRINTF_FORMAT
#undef JSON_HEDLEY_PRIVATE
#undef JSON_HEDLEY_PUBLIC
File diff suppressed because it is too large Load Diff
+139
View File
@@ -2,6 +2,9 @@ cmake_minimum_required(VERSION 3.13...4.0)
option(JSON_Valgrind "Execute test suite with Valgrind." OFF)
option(JSON_FastTests "Skip expensive/slow tests." OFF)
option(JSON_TestSimdutf "Build the unit tests against the simdutf UTF-8 validation backend." OFF)
set(JSON_SIMDUTF_VERSION 9.1.0 CACHE STRING "The simdutf version used by JSON_TestSimdutf.")
set(JSON_32bitTest AUTO CACHE STRING "Enable the 32bit unit test (ON/OFF/AUTO/ONLY).")
set(JSON_TestStandards "" CACHE STRING "The list of standards to test explicitly.")
@@ -125,6 +128,51 @@ json_test_set_test_options(test-unicode4 TEST_PROPERTIES TIMEOUT 3000)
# add unit tests
#############################################################################
# Generate the leak checks for every JSON_HEDLEY_* macro defined in
# hedley.hpp; tests/src/unit-no-macro-leak.cpp #include-s the result after
# nlohmann/json.hpp (see issue #5408). Using the shared
# cmake/scripts/gen_hedley_undef_check.cmake script (also used by `make
# update_hedley_undef`) instead of a hand-maintained list of macro names
# means this test can never go stale after a future `make update_hedley`.
set(hedley_hpp "${PROJECT_SOURCE_DIR}/include/nlohmann/thirdparty/hedley/hedley.hpp")
set(hedley_undef_check_script "${PROJECT_SOURCE_DIR}/cmake/scripts/gen_hedley_undef_check.cmake")
set(hedley_undef_checks "${PROJECT_BINARY_DIR}/include/hedley_undef_checks.inc")
# Reconfigure whenever the vendored header or the generator script changes,
# so a `cmake --build` after `make update_hedley` does not silently keep a
# stale generated file around.
set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS
"${hedley_hpp}"
"${hedley_undef_check_script}")
# Generate once at configure time, so the very first build (before any
# custom-command build step has run) already has an up-to-date file.
execute_process(
COMMAND ${CMAKE_COMMAND}
"-DHEDLEY_HPP=${hedley_hpp}"
"-DOUTPUT=${hedley_undef_checks}"
-DMODE=checks
-P "${hedley_undef_check_script}"
RESULT_VARIABLE hedley_undef_check_result
)
if(NOT hedley_undef_check_result EQUAL 0)
message(FATAL_ERROR "Failed to generate ${hedley_undef_checks}")
endif()
# Also (re)generate as a build step, so an incremental build after editing
# hedley.hpp without a full reconfigure still picks up the change.
add_custom_command(
OUTPUT "${hedley_undef_checks}"
COMMAND ${CMAKE_COMMAND}
"-DHEDLEY_HPP=${hedley_hpp}"
"-DOUTPUT=${hedley_undef_checks}"
-DMODE=checks
-P "${hedley_undef_check_script}"
DEPENDS "${hedley_hpp}" "${hedley_undef_check_script}"
COMMENT "Generating Hedley undef leak checks"
VERBATIM)
add_custom_target(generate_hedley_undef_checks DEPENDS "${hedley_undef_checks}")
if("${JSON_TestStandards}" STREQUAL "")
set(test_cxx_standards 11 14 17 20 23)
unset(test_force)
@@ -149,6 +197,71 @@ if(test_force)
endif()
message(STATUS "${msg}")
#############################################################################
# optionally validate UTF-8 with simdutf (JSON_USE_SIMDUTF)
#############################################################################
# The simdutf backend is opt-in and not vendored, so it is fetched here rather
# than being a checked-in dependency. Everything below hangs off test_main,
# whose usage requirements every test target inherits; the library target and
# the installed CMake package are deliberately left untouched.
if (JSON_TestSimdutf)
# simdutf requires C++17, both to compile itself and to be reachable from
# the library, which keeps its scalar validator below that. Find a tested
# standard that satisfies it.
set(simdutf_standard "")
foreach(cxx_standard ${test_cxx_standards})
if(NOT cxx_standard LESS 17 AND compiler_supports_cpp_${cxx_standard})
set(simdutf_standard ${cxx_standard})
break()
endif()
endforeach()
if("${simdutf_standard}" STREQUAL "")
# Building simdutf would fail outright without a C++17 compiler, and
# even with one it would go unused if no C++17-or-later standard is
# tested. Say so and fall back to the scalar validator rather than
# failing the build.
if(NOT compiler_supports_cpp_17)
set(simdutf_reason "the compiler does not support C++17")
else()
set(simdutf_reason "no tested standard is C++17 or later (testing ${msg_standards})")
endif()
message(WARNING
"JSON_TestSimdutf is enabled, but ${simdutf_reason}. simdutf requires C++17, so it "
"is not fetched and JSON_USE_SIMDUTF is not defined: the tests run against the "
"built-in scalar UTF-8 validator instead. Set JSON_TestStandards to include 17 or "
"later, or build with a compiler that supports C++17.")
else()
if (CMAKE_VERSION VERSION_LESS 3.18)
message(FATAL_ERROR "JSON_TestSimdutf requires CMake 3.18 or later (simdutf's minimum).")
endif()
include(FetchContent)
# simdutf builds its tests and tools by default, and its tests pull
# further dependencies of their own; only the library is needed here
set(SIMDUTF_TESTS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_TOOLS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_BENCHMARKS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_ICONV OFF CACHE BOOL "" FORCE)
FetchContent_Declare(simdutf
URL https://github.com/simdutf/simdutf/archive/refs/tags/v${JSON_SIMDUTF_VERSION}.tar.gz
DOWNLOAD_EXTRACT_TIMESTAMP TRUE
)
FetchContent_MakeAvailable(simdutf)
target_compile_definitions(test_main PUBLIC JSON_USE_SIMDUTF)
target_link_libraries(test_main PUBLIC simdutf::simdutf)
# simdutf.h requires C++17; below that the library keeps its scalar
# validator, so any C++11/14 test targets exercise the fallback and the
# C++17-and-later ones exercise simdutf. Both must agree.
message(STATUS "UTF-8 validation delegated to simdutf ${JSON_SIMDUTF_VERSION} for C++17 and later (JSON_USE_SIMDUTF)")
endif()
endif()
# *DO* use json_test_set_test_options() above this line
json_test_should_build_32bit_test(json_32bit_test json_32bit_test_only "${JSON_32bitTest}")
@@ -163,6 +276,14 @@ foreach(file ${files})
json_test_add_test_for(${file} MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force})
endforeach()
# tests/src/unit-no-macro-leak.cpp #include-s the generated leak-check file,
# so its test targets must be built after generate_hedley_undef_checks.
foreach(cxx_standard ${test_cxx_standards})
if(TARGET test-no-macro-leak_cpp${cxx_standard})
add_dependencies(test-no-macro-leak_cpp${cxx_standard} generate_hedley_undef_checks)
endif()
endforeach()
if(json_32bit_test_only)
# Skip all other tests in this file
return()
@@ -177,6 +298,24 @@ json_test_add_test_for(src/unit-comparison.cpp
MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force}
)
# test the parser again with JSON_DIAGNOSTIC_POSITIONS enabled
json_test_set_test_options(test-class_parser_diagnostic_positions
COMPILE_DEFINITIONS JSON_DIAGNOSTIC_POSITIONS=1
)
json_test_add_test_for(src/unit-class_parser.cpp
NAME test-class_parser_diagnostic_positions
MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force}
)
# test diagnostic positions again without regular diagnostics (JSON pointer paths)
json_test_set_test_options(test-diagnostic-positions_only
COMPILE_DEFINITIONS JSON_DIAGNOSTICS=0
)
json_test_add_test_for(src/unit-diagnostic-positions.cpp
NAME test-diagnostic-positions_only
MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force}
)
# *DO NOT* use json_test_set_test_options() below this line
#############################################################################
+21 -15
View File
@@ -1,4 +1,4 @@
cmake_minimum_required(VERSION 3.11...3.14)
cmake_minimum_required(VERSION 3.14)
project(JSON_Benchmarks LANGUAGES CXX)
# set compiler flags
@@ -6,29 +6,35 @@ if((CMAKE_CXX_COMPILER_ID MATCHES GNU) OR (CMAKE_CXX_COMPILER_ID MATCHES Clang))
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -flto -DNDEBUG -O3")
endif()
# configure Google Benchmarks
# configure Google Benchmark; a fixed release, so that results stay comparable
set(JSON_GOOGLE_BENCHMARK_VERSION 1.9.5)
include(FetchContent)
FetchContent_Declare(
benchmark
GIT_REPOSITORY https://github.com/google/benchmark.git
GIT_TAG origin/main
GIT_SHALLOW TRUE
)
FetchContent_GetProperties(benchmark)
if(NOT benchmark_POPULATED)
FetchContent_Populate(benchmark)
set(BENCHMARK_ENABLE_TESTING OFF CACHE INTERNAL "" FORCE)
add_subdirectory(${benchmark_SOURCE_DIR} ${benchmark_BINARY_DIR})
endif()
# only the library is needed; -Werror would break the pinned release as soon as
# a newer compiler adds a warning
set(BENCHMARK_ENABLE_TESTING OFF CACHE BOOL "" FORCE)
set(BENCHMARK_ENABLE_INSTALL OFF CACHE BOOL "" FORCE)
set(BENCHMARK_ENABLE_WERROR OFF CACHE BOOL "" FORCE)
FetchContent_Declare(benchmark
URL https://github.com/google/benchmark/archive/refs/tags/v${JSON_GOOGLE_BENCHMARK_VERSION}.tar.gz
URL_HASH SHA256=9631341c82bac4a288bef951f8b26b41f69021794184ece969f8473977eaa340
DOWNLOAD_EXTRACT_TIMESTAMP TRUE
)
FetchContent_MakeAvailable(benchmark)
# download test data
set(CMAKE_MODULE_PATH ${CMAKE_CURRENT_SOURCE_DIR}/../../cmake ${CMAKE_MODULE_PATH})
include(download_test_data)
# the header to benchmark; point this at a directory holding another version's
# nlohmann/json.hpp to compare versions (see README.md)
set(JSON_BENCHMARK_INCLUDE_DIR "${CMAKE_CURRENT_SOURCE_DIR}/../../single_include" CACHE PATH
"directory containing the nlohmann/json.hpp to benchmark")
# benchmark binary
add_executable(json_benchmarks src/benchmarks.cpp)
target_compile_features(json_benchmarks PRIVATE cxx_std_11)
target_link_libraries(json_benchmarks benchmark ${CMAKE_THREAD_LIBS_INIT})
add_dependencies(json_benchmarks download_test_data)
target_include_directories(json_benchmarks PRIVATE ${CMAKE_SOURCE_DIR}/../../single_include ${CMAKE_BINARY_DIR}/include)
target_include_directories(json_benchmarks PRIVATE ${JSON_BENCHMARK_INCLUDE_DIR} ${CMAKE_BINARY_DIR}/include)
+130
View File
@@ -0,0 +1,130 @@
# Benchmarks
Micro-benchmarks for parsing, serialization and the binary formats, written with
[Google Benchmark](https://github.com/google/benchmark). They are not run by CI; see
[When to run them](#when-to-run-them).
## What is measured
| benchmark | what it does |
|---|---|
| `ParseFile`, `ParseString` | parse JSON from a file stream or a string |
| `ParseIndented` | parse the large files re-indented by 4 spaces, for the lexer's whitespace handling |
| `Dump` | serialize, compact (`-`) and indented (`4`) |
| `ToCbor`, `BinaryToCbor` | write CBOR; `BinaryToCbor` writes binary values of growing size |
| `FromMsgpack` | read MessagePack; unchanged over the years, so its numbers stay comparable across releases |
| `FromBinaryBuffer`, `FromBinaryFile` | read CBOR, MessagePack, UBJSON, BJData and BSON from a buffer or a `FILE*` |
| `FromBinaryShape` | read deeply nested, container-heavy and scalar-heavy documents in every binary format |
| `FromCborChunkedString` | read CBOR strings split into indefinite-length chunks |
The input files are those of [nativejson-benchmark](https://github.com/miloyip/nativejson-benchmark) (`canada`,
`citm_catalog`, `twitter`), a large `jeopardy` file, and number-heavy files (`floats`, `signed_ints`, ...).
`bytes_per_second` counts the bytes read or written: the JSON text when parsing, the output when serializing.
## Requirements
- CMake 3.14 or later, a C++11 compiler, and Ninja for the `make` target.
- Network access on the first configure: CMake downloads Google Benchmark and the
[test data](https://github.com/nlohmann/json_test_data) into the build directory. To reuse a download of the test
data, pass `-DJSON_TestDataDirectory=<build directory>/test_files`.
- Google Benchmark is pinned to a release (1.9.5), so that results from different days stay comparable. To update it,
change `JSON_GOOGLE_BENCHMARK_VERSION` and the archive's `URL_HASH` in `CMakeLists.txt` together.
- The benchmarks include `single_include/nlohmann/json.hpp`, so run `make amalgamate` after changing anything in
`include/`.
GCC and Clang builds use `-O3 -flto -DNDEBUG`.
## Running them
From the repository root, this builds everything from scratch in `cmake-build-benchmarks` and runs all benchmarks:
```sh
make run_benchmarks
```
To build once and run selectively:
```sh
cmake -S tests/benchmarks -B build-benchmarks -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build-benchmarks
build-benchmarks/json_benchmarks --benchmark_filter='ParseString|Dump'
```
Useful options of `json_benchmarks`:
| option | effect |
|---|---|
| `--benchmark_list_tests` | list the benchmarks instead of running them |
| `--benchmark_filter=<regex>` | run only the benchmarks whose names match |
| `--benchmark_repetitions=<n>` | run every benchmark `n` times and add mean, median, standard deviation and coefficient of variation |
| `--benchmark_enable_random_interleaving=true` | run the repetitions in random order, which spreads out drifts such as thermal throttling |
| `--benchmark_min_time=<seconds>s` | run each benchmark at least this long (e.g. `2s`) |
| `--benchmark_out=<file> --benchmark_out_format=json` | also write the results to a file, e.g. for `compare.py` |
## Reading the output
Each line shows the wall-clock `Time` and the `CPU` time per iteration, the number of `Iterations` Google Benchmark
chose, and the throughput in `bytes_per_second`. With repetitions, the lines ending in `_median` are the ones to
compare. A `_cv` (coefficient of variation) above a few percent means the machine was too noisy for small
differences to mean anything.
## Comparing two versions
To see what a change or a release did, build the same benchmarks twice: once against the header of the version to
compare with, and once against the current one. `JSON_BENCHMARK_INCLUDE_DIR` names the directory holding the
`nlohmann/json.hpp` to benchmark. For example, to compare the current checkout with 3.12.0:
```sh
# the header of the version to compare with
mkdir -p build-baseline-header/nlohmann
git show v3.12.0:single_include/nlohmann/json.hpp > build-baseline-header/nlohmann/json.hpp
# the same benchmarks, built against either header
cmake -S tests/benchmarks -B build-baseline -G Ninja -DCMAKE_BUILD_TYPE=Release \
-DJSON_BENCHMARK_INCLUDE_DIR="$PWD/build-baseline-header"
cmake -S tests/benchmarks -B build-current -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build-baseline
cmake --build build-current
# run both, back to back
build-baseline/json_benchmarks --benchmark_repetitions=10 --benchmark_enable_random_interleaving=true \
--benchmark_out=build-baseline/results.json --benchmark_out_format=json
build-current/json_benchmarks --benchmark_repetitions=10 --benchmark_enable_random_interleaving=true \
--benchmark_out=build-current/results.json --benchmark_out_format=json
```
Google Benchmark ships a tool to compare the two result files. It needs NumPy and SciPy:
```sh
python3 -m venv build-venv
build-venv/bin/pip install numpy scipy
build-venv/bin/python build-current/_deps/benchmark-src/tools/compare.py -a benchmarks build-baseline/results.json build-current/results.json
```
The tool's own `tools/requirements.txt` pins NumPy and SciPy versions that need Python 3.11 or later; with an older
Python, unpinned versions work as well. In its output:
- the `Time` and `CPU` columns are relative changes: `-0.35` means 35% faster, `+0.10` means 10% slower;
- `_pvalue` lines report a Mann-Whitney U test of whether the two versions differ. It needs at least 9
repetitions, and a p-value below 0.05 means the difference is unlikely to be noise;
- `OVERALL_GEOMEAN` summarizes all benchmarks;
- `-a` shows only the aggregates, not every repetition.
The header you compare with must support everything the benchmarks use. The current benchmarks build against 3.12.0.
Only benchmarks present in both result files are compared, so for older releases, either filter the benchmarks or
build that release's own `tests/benchmarks` against its own header.
## Getting stable numbers
- Build and run both versions on the same machine, one right after the other.
- Keep the machine otherwise idle: no builds, no browser, and a laptop plugged in.
- On Linux, set the CPU frequency governor to `performance`, e.g. `sudo cpupower frequency-set --governor performance`.
Google Benchmark prints a warning when frequency scaling is enabled. Pinning the process to a core
(`taskset -c 2 ...`) helps as well.
- Use 10 or more repetitions with random interleaving, compare medians, and treat changes within the `_cv` as noise.
## When to run them
They are a manual step, not part of CI: shared CI runners vary more between runs than most of the effects measured.
Run the comparison above before a release, comparing the previous release tag with `develop`, and for pull requests
that claim to change performance.
+359 -1
View File
@@ -81,6 +81,44 @@ BENCHMARK_CAPTURE(ParseString, signed_ints, TEST_DATA_DIRECTORY "/regressi
BENCHMARK_CAPTURE(ParseString, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json");
BENCHMARK_CAPTURE(ParseString, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json");
//////////////////////////////////////////////////////////////////////////////
// parse pretty-printed JSON from string
//
// Every file in the corpus above is minified or only lightly spaced, so none of
// them exercise the lexer's whitespace handling. Real-world JSON is frequently
// indented - configuration files, pretty-printed API responses, anything kept
// under version control - where insignificant whitespace can outweigh the data.
// Re-serializing a document with an indentation and parsing that keeps the
// content identical to the ParseString row above, so the pair isolates the cost
// of the whitespace alone.
//////////////////////////////////////////////////////////////////////////////
static void ParseIndented(benchmark::State& state, const char* filename, int indent)
{
std::ifstream f(filename);
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
const std::string indented = json::parse(str).dump(indent);
while (state.KeepRunning())
{
state.PauseTiming();
auto* j = new json();
state.ResumeTiming();
*j = json::parse(indented);
state.PauseTiming();
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * indented.size());
}
BENCHMARK_CAPTURE(ParseIndented, jeopardy / 4, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", 4);
BENCHMARK_CAPTURE(ParseIndented, canada / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", 4);
BENCHMARK_CAPTURE(ParseIndented, citm_catalog / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", 4);
BENCHMARK_CAPTURE(ParseIndented, twitter / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", 4);
//////////////////////////////////////////////////////////////////////////////
// serialize JSON
//////////////////////////////////////////////////////////////////////////////
@@ -93,7 +131,8 @@ static void Dump(benchmark::State& state, const char* filename, int indent)
while (state.KeepRunning())
{
j.dump(indent);
std::string output = j.dump(indent);
benchmark::DoNotOptimize(output);
}
state.SetBytesProcessed(state.iterations() * j.dump(indent).size());
@@ -214,4 +253,323 @@ static void BinaryToCbor(benchmark::State& state)
}
BENCHMARK(BinaryToCbor)->RangeMultiplier(2)->Range(8, 8 << 12);
//////////////////////////////////////////////////////////////////////////////
// parse binary formats
//////////////////////////////////////////////////////////////////////////////
// Only MessagePack had a read benchmark (FromMsgpack above, left untouched so
// its numbers stay comparable across releases). The benchmarks below cover the
// other formats, and read from a contiguous buffer as well as from a FILE*:
// most callers pass a container, and the two adapters compile to different
// code. The test data repository ships JSON only, so the input for each is
// derived at setup time by serializing a parsed test file.
/// binary format to benchmark; the _optimized variants add UBJSON/BJData size
/// and type annotations, which the readers handle in a separate code path
enum class binary_format
{
cbor,
msgpack,
ubjson,
ubjson_optimized,
bjdata,
bjdata_optimized,
bson
};
static std::vector<std::uint8_t> to_binary(const json& j, const binary_format format)
{
switch (format)
{
case binary_format::cbor:
return json::to_cbor(j);
case binary_format::msgpack:
return json::to_msgpack(j);
case binary_format::ubjson:
return json::to_ubjson(j);
case binary_format::ubjson_optimized:
return json::to_ubjson(j, true, true);
case binary_format::bjdata:
return json::to_bjdata(j);
case binary_format::bjdata_optimized:
return json::to_bjdata(j, true, true);
case binary_format::bson:
default:
return json::to_bson(j);
}
}
static json from_binary(const std::vector<std::uint8_t>& bytes, const binary_format format)
{
switch (format)
{
case binary_format::cbor:
return json::from_cbor(bytes);
case binary_format::msgpack:
return json::from_msgpack(bytes);
case binary_format::ubjson:
case binary_format::ubjson_optimized:
return json::from_ubjson(bytes);
case binary_format::bjdata:
case binary_format::bjdata_optimized:
return json::from_bjdata(bytes);
case binary_format::bson:
default:
return json::from_bson(bytes);
}
}
static json from_binary(std::FILE* file, const binary_format format)
{
switch (format)
{
case binary_format::cbor:
return json::from_cbor(file);
case binary_format::msgpack:
return json::from_msgpack(file);
case binary_format::ubjson:
case binary_format::ubjson_optimized:
return json::from_ubjson(file);
case binary_format::bjdata:
case binary_format::bjdata_optimized:
return json::from_bjdata(file);
case binary_format::bson:
default:
return json::from_bson(file);
}
}
/*!
@brief serialize a parsed test file to @a format
Returns an empty vector and marks the benchmark as skipped if the file cannot
be represented in the format, rather than letting the exception escape: BSON
requires an object at the top level, and several test files are arrays.
*/
static std::vector<std::uint8_t> binary_input(benchmark::State& state, const char* filename, const binary_format format)
{
std::ifstream f(filename);
std::string const str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
const json j = json::parse(str);
if (format == binary_format::bson && !j.is_object())
{
state.SkipWithError("BSON requires an object at the top level");
return {};
}
return to_binary(j, format);
}
static void FromBinaryBuffer(benchmark::State& state, const char* filename, const binary_format format)
{
const std::vector<std::uint8_t> bytes = binary_input(state, filename, format);
if (bytes.empty())
{
return;
}
for (auto _ : state)
{
// the value is destroyed outside the timed section, because destroying
// a large DOM is not what this benchmark measures
state.PauseTiming();
auto* j = new json();
state.ResumeTiming();
*j = from_binary(bytes, format);
state.PauseTiming();
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / floats, TEST_DATA_DIRECTORY "/regression/floats.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson_optimized / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson_optimized);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson_optimized / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson_optimized);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata_optimized / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bjdata_optimized);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata_optimized / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata_optimized);
// BSON requires an object at the top level, so the array-rooted test files
// (jeopardy and the regression files) cannot be captured here
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bson);
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::bson);
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bson);
static void FromBinaryFile(benchmark::State& state, const char* filename, const binary_format format)
{
const std::vector<std::uint8_t> bytes = binary_input(state, filename, format);
if (bytes.empty())
{
return;
}
const char* tmp = "benchmark_input.bin";
std::ofstream o(tmp, std::ios::binary);
o.write(reinterpret_cast<const char*>(bytes.data()), static_cast<std::streamsize>(bytes.size()));
o.flush();
o.close();
for (auto _ : state)
{
state.PauseTiming();
auto* j = new json();
auto* file = std::fopen(tmp, "rb");
state.ResumeTiming();
*j = from_binary(file, format);
state.PauseTiming();
std::fclose(file);
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromBinaryFile, cbor / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryFile, cbor / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryFile, ubjson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryFile, ubjson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryFile, bjdata / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryFile, bson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bson);
//////////////////////////////////////////////////////////////////////////////
// parse binary formats: value shapes
//////////////////////////////////////////////////////////////////////////////
// The test files above are wide and shallow, but the readers' cost is per
// container, so these cover the shapes that stress the container handling
// itself. Every shape is wrapped in an object so that BSON, which requires an
// object at the top level, measures the same value as the other formats.
/// deeply nested arrays: one container per level, no other work
static json make_nested()
{
json nested = json::array();
json* p = &nested;
for (std::size_t i = 1; i < 1000; ++i)
{
p->push_back(json::array());
p = &p->operator[](0);
}
json j = json::object();
j["data"] = std::move(nested);
return j;
}
/// many sibling containers: maximum container churn, minimum nesting
static json make_containers()
{
json data = json::array();
for (std::size_t i = 0; i < 100000; ++i)
{
data.push_back(json::array({1, 2}));
}
json j = json::object();
j["data"] = std::move(data);
return j;
}
/// one flat array of numbers: the scalar decoding path, which must not move
static json make_scalars()
{
json data = json::array();
for (std::size_t i = 0; i < 1000000; ++i)
{
data.push_back(i);
}
json j = json::object();
j["data"] = std::move(data);
return j;
}
static void FromBinaryShape(benchmark::State& state, json (*build)(), const binary_format format)
{
const std::vector<std::uint8_t> bytes = to_binary(build(), format);
for (auto _ : state)
{
state.PauseTiming();
auto* j = new json();
state.ResumeTiming();
*j = from_binary(bytes, format);
state.PauseTiming();
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromBinaryShape, nested / cbor, make_nested, binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryShape, nested / msgpack, make_nested, binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryShape, nested / ubjson, make_nested, binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryShape, nested / bjdata, make_nested, binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryShape, nested / bson, make_nested, binary_format::bson);
BENCHMARK_CAPTURE(FromBinaryShape, containers / cbor, make_containers, binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryShape, containers / msgpack, make_containers, binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryShape, containers / ubjson, make_containers, binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryShape, containers / ubjson_optimized, make_containers, binary_format::ubjson_optimized);
BENCHMARK_CAPTURE(FromBinaryShape, containers / bjdata, make_containers, binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryShape, containers / bson, make_containers, binary_format::bson);
// BSON names every array element, so a large array measures key generation
// rather than scalar decoding and is left out here
BENCHMARK_CAPTURE(FromBinaryShape, scalars / cbor, make_scalars, binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryShape, scalars / msgpack, make_scalars, binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryShape, scalars / ubjson, make_scalars, binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryShape, scalars / bjdata, make_scalars, binary_format::bjdata);
/*!
@brief parse an indefinite-length CBOR string
The writer never emits this form, so the input is assembled by hand: 0x7F
opens the string, each chunk is a one-character string, and 0xFF closes it.
*/
static void FromCborChunkedString(benchmark::State& state, const std::size_t chunks)
{
std::vector<std::uint8_t> bytes;
bytes.reserve(2 * chunks + 2);
bytes.push_back(0x7F);
for (std::size_t i = 0; i < chunks; ++i)
{
bytes.push_back(0x61); // string of length 1
bytes.push_back(0x61); // 'a'
}
bytes.push_back(0xFF);
for (auto _ : state)
{
json j = json::from_cbor(bytes);
benchmark::DoNotOptimize(j);
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromCborChunkedString, 10000 chunks, 10000);
BENCHMARK_MAIN();
+34 -4
View File
@@ -21,6 +21,27 @@ array data, it performs the following steps:
- j4 = from_bjdata(vec3)
- assert(j1 == j4)
Re-serializing j2/j3/j4 with the same use_size/use_type settings is checked
for value-stability rather than byte-exact stability: from_bjdata(to_bjdata(j2))
must equal j2 (and likewise for j3, j4). Byte-exact stability does not hold in
general, because a BJData value can lose type fidelity across a round trip
(e.g. a binary_t value serialized without the optimized "$U#" array header is
parsed back as a plain array of numbers, see #5398 and the discussion on
PR #5494) - the numeric value is preserved, but the writer's smallest-type
selection for the now-plain numbers may legitimately pick a different, but
equally valid, single-byte type marker than the dedicated binary-data writer
would have. Both encodings are valid BJData and both decode to the same
value, so this is not treated as a round-trip failure here.
"Value-stable" is checked by comparing dump()s rather than with operator==
directly: a BJData/UBJSON payload can decode to a non-finite double (NaN or
+-Infinity), and IEEE 754 NaN is never equal to itself, so operator== would
report two structurally-identical trees as different whenever a NaN is
involved -- not a round-trip bug, just NaN's ordinary (non-)reflexivity.
dump() serializes any non-finite double the same deterministic way (as JSON
`null`, since JSON itself cannot represent NaN/Infinity), so comparing
dumps is stable under exactly the same values that break operator==.
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -31,6 +52,13 @@ drivers.
using json = nlohmann::json;
// value-stable comparison for the round-trip checks below; see the note
// above on why this compares dump()s rather than the json values directly
static bool is_value_stable(const json& lhs, const json& rhs)
{
return lhs.dump() == rhs.dump();
}
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
@@ -56,10 +84,12 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
json const j3 = json::from_bjdata(vec3);
json const j4 = json::from_bjdata(vec4);
// serializations must match
assert(json::to_bjdata(j2, false, false) == vec2);
assert(json::to_bjdata(j3, true, false) == vec3);
assert(json::to_bjdata(j4, true, true) == vec4);
// re-serializing must be value-stable (see the notes above on
// why byte-exact stability is not guaranteed in general, and
// why this compares dump()s rather than the values directly)
assert(is_value_stable(json::from_bjdata(json::to_bjdata(j2, false, false)), j2));
assert(is_value_stable(json::from_bjdata(json::to_bjdata(j3, true, false)), j3));
assert(is_value_stable(json::from_bjdata(json::to_bjdata(j4, true, true)), j4));
}
catch (const json::parse_error&)
{
+61
View File
@@ -0,0 +1,61 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Standalone compile-and-run check for the JSON_SKIP_LIBRARY_VERSION_CHECK
// configuration macro, which (per #5423) was never exercised anywhere in the
// test matrix.
//
// include/nlohmann/detail/abi_macros.hpp normally emits a #warning if
// NLOHMANN_JSON_VERSION_MAJOR/MINOR/PATCH are already defined (as they would
// be by an earlier inclusion of a different version of the library) with
// values that mismatch the version about to be defined -- unless
// JSON_SKIP_LIBRARY_VERSION_CHECK is defined, in which case the check (and
// that #warning) is skipped.
//
// This file deliberately is not named tests/src/unit-*.cpp: it is compiled
// directly (with a modest, non-strict warning set) by the dedicated
// ci_test_skiplibraryversioncheck target in cmake/ci.cmake, rather than being
// folded into the library's own -Weverything/-Werror unit test matrix. That
// is because the scenario simulated here -- mixing two different, already
// differently-versioned inclusions of the library in one translation unit --
// unavoidably also triggers the *compiler's own* "macro redefined" warning,
// independent of (and unaffected by) JSON_SKIP_LIBRARY_VERSION_CHECK, which
// only ever silences the library's own #warning. Building this file under
// -Weverything -Werror would therefore fail for a reason unrelated to the
// macro under test.
#define NLOHMANN_JSON_VERSION_MAJOR 0
#define NLOHMANN_JSON_VERSION_MINOR 0
#define NLOHMANN_JSON_VERSION_PATCH 0
#define JSON_SKIP_LIBRARY_VERSION_CHECK 1
#include <nlohmann/json.hpp>
int main()
{
// reaching this point at all already proves that the mismatched,
// pre-defined version macros above did not stop compilation -- which is
// exactly what JSON_SKIP_LIBRARY_VERSION_CHECK is for. The library must
// also still be fully usable.
const nlohmann::json j = {{"a", 1}, {"b", {1, 2, 3}}};
if (j.dump() != "{\"a\":1,\"b\":[1,2,3]}")
{
return 1;
}
// include/nlohmann/detail/abi_macros.hpp unconditionally (re)defines the
// version macros to the library's real, current version right after the
// (here, skipped) mismatch check, regardless of the deliberately wrong
// stand-in values defined above.
if (NLOHMANN_JSON_VERSION_MAJOR == 0 && NLOHMANN_JSON_VERSION_MINOR == 0 && NLOHMANN_JSON_VERSION_PATCH == 0)
{
return 1;
}
return 0;
}
+9
View File
@@ -15,6 +15,15 @@
namespace utils
{
// Some tests intentionally discard the [[nodiscard]]/JSON_HEDLEY_WARN_UNUSED_RESULT
// return value of a call they only make to exercise its side effects (e.g. checking
// that it does not throw). A plain (void) cast on the call expression does not
// suppress GCC's warning for functions using the GNU __attribute__((warn_unused_result))
// form (as opposed to the C++17 [[nodiscard]] attribute) -- passing the value into an
// ordinary function call does.
template<typename T>
inline void ignore_return_value(T&& /*unused*/) noexcept {}
inline std::vector<std::uint8_t> read_binary_file(const std::string& filename)
{
std::ifstream file(filename, std::ios::binary);
+40 -32
View File
@@ -11,8 +11,10 @@
#include <nlohmann/json.hpp>
#include <cstdint>
#include <string>
#include <utility>
#include <vector>
/* forward declarations */
class alt_string;
@@ -22,6 +24,10 @@ void int_to_string(alt_string& target, std::size_t value); // NOLINT(misc-use-in
/*
* This is virtually a string class.
* It covers std::string under the hood.
*
* It deliberately does not provide c_str(), back(), find(str, pos), replace(),
* or substr(): the library must not rely on them. Do not add members here
* without checking that the library actually needs them.
*/
class alt_string
{
@@ -106,11 +112,6 @@ class alt_string
return str_impl < op.str_impl;
}
const char* c_str() const
{
return str_impl.c_str();
}
char& operator[](std::size_t index)
{
return str_impl[index];
@@ -121,16 +122,6 @@ class alt_string
return str_impl[index];
}
char& back()
{
return str_impl.back();
}
const char& back() const
{
return str_impl.back();
}
void clear()
{
str_impl.clear();
@@ -146,28 +137,11 @@ class alt_string
return str_impl.empty();
}
std::size_t find(const alt_string& str, std::size_t pos = 0) const
{
return str_impl.find(str.str_impl, pos);
}
std::size_t find_first_of(char c, std::size_t pos = 0) const
{
return str_impl.find_first_of(c, pos);
}
alt_string substr(std::size_t pos = 0, std::size_t count = npos) const
{
const std::string s = str_impl.substr(pos, count);
return {s.data(), s.size()};
}
alt_string& replace(std::size_t pos, std::size_t count, const alt_string& str)
{
str_impl.replace(pos, count, str.str_impl);
return *this;
}
void reserve( std::size_t new_cap = 0 )
{
str_impl.reserve(new_cap);
@@ -202,6 +176,31 @@ bool operator<(const char* op1, const alt_string& op2) noexcept
TEST_CASE("alternative string type")
{
SECTION("binary formats")
{
alt_json doc;
doc["pi"] = 3.141;
doc["happy"] = true;
doc["list"] = {1, 2, 3};
CHECK(alt_json::from_cbor(alt_json::to_cbor(doc)) == doc);
CHECK(alt_json::from_msgpack(alt_json::to_msgpack(doc)) == doc);
// BSON is not covered: it additionally needs string_t::find(value_type),
// which alt_string does not provide
CHECK(alt_json::from_ubjson(alt_json::to_ubjson(doc)) == doc);
// a UBJSON high-precision number is parsed into a std::string that the
// reader has to hand to the SAX interface as an alt_string
const std::vector<uint8_t> high_precision =
{
'H', 'i', 0x16, '3', '.', '1', '4', '1', '5', '9', '2', '6', '5', '3',
'5', '8', '9', '7', '9', '3', '2', '3', '8', '4', '6'
};
const auto number = alt_json::from_ubjson(high_precision);
CHECK(number.is_number_float());
CHECK(number.get<double>() == doctest::Approx(3.14159265358979323846));
}
SECTION("dump")
{
{
@@ -332,6 +331,15 @@ TEST_CASE("alternative string type")
CHECK(j.at(alt_json::json_pointer("/foo/0")) == j["foo"][0]);
CHECK(j.at(alt_json::json_pointer("/foo/1")) == j["foo"][1]);
// RFC 6901 escaping works without string_t::find(str, pos), replace(),
// and substr()
auto j2 = alt_json::parse(R"({"a/b": 1, "m~n": 2, "~/~~//": 3})");
CHECK(j2.at(alt_json::json_pointer("/a~1b")) == 1);
CHECK(j2.at(alt_json::json_pointer("/m~0n")) == 2);
CHECK(j2.at(alt_json::json_pointer("/~0~1~0~0~1~1")) == 3);
CHECK(alt_json::json_pointer("/~0~1~0~0~1~1").to_string() == alt_string("/~0~1~0~0~1~1"));
CHECK(j2.flatten().unflatten() == j2);
}
SECTION("patch")

Some files were not shown because too many files have changed in this diff Show More