Conflicts:
- include/nlohmann/detail/output/binary_writer.hpp: kept write_msgpack_unsigned() for both msgpack integer cases; it already reads the active union member, which is what develop's #5694 fixes in the old ladders
- single_include/nlohmann/json.hpp: regenerated with make amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
operator>> parsed directly into its basic_json& target, so a parse
error left the target holding whatever was parsed before the error
instead of its previous value. With JSON_DIAGNOSTICS=1, that partial
value also violated the class invariant, because the parent pointers
of an array or object's elements are only set when the container is
closed, which a failed parse never reaches; copying such a value then
aborted in assert_invariant().
Fix it the way basic_json::parse() already handles this: parse into a
temporary and move it into the target only once parsing succeeds, so
the target is left unchanged if an exception is thrown.
Fixes#5652.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.
Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.
Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).
Fixes#5666.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add JSON_NO_UDLS to leave out the user-defined string literals
The bodies of operator""_json and operator""_json_pointer call the
parser, so every translation unit including the library instantiates it,
even if it never parses anything. Defining JSON_NO_UDLS leaves the
literals out entirely, which saves 15-35% compile time for such
translation units (#5294). Nothing changes if the macro is not defined.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mention JSON_NO_UDLS in the list of exported module symbols
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the user-defined string literals to <nlohmann/json_literals.hpp>
Following the review in #5294, the literals now live in their own header
instead of being removed entirely: <nlohmann/json.hpp> includes it at the
end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined,
so a project can opt out globally and include the header only where the
literals are used.
The header only uses public and standard macros, because the library's
internal macros are undefined at the end of json.hpp and the amalgamation
inlines macro_scope.hpp only once. For the same reason, the library no
longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition
is still visible to the header. The single-header copy is identical to
the multi-header one, as it only includes <nlohmann/json.hpp>. The module
always exports the literals.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI: include cycle, GCC 4.8 literal operator spacing, and global UDLs off in the JSON_NO_AUTOMATIC_UDLS test
- Suppress clang-tidy misc-header-include-cycle on the intentional mutual
include of json.hpp and json_literals.hpp.
- Use operator"" _json with a space for GCC 4.8 in the test's detection
aliases, as the header does.
- Only test the global literal operators when JSON_USE_GLOBAL_UDLS is on.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Declare the literal operators through a local macro
The GCC 4.8 spacing condition was repeated for both operator definitions
and the global using-declarations. NLOHMANN_JSON_LITERAL_OPERATOR(suffix)
now selects operator""##suffix or operator"" suffix in one place and is
undefined at the end of json_literals.hpp.
Suggested by gregmarr in review.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When copying a value nested deeper than 128 levels, and an allocation
fails while an inner array or object is being copied, the partially
built copy ended up with an element typed array/object but holding a
null pointer. That element was already a fully constructed member of
its parent's container, so destroying the parent during stack
unwinding dereferenced the null pointer (release builds) or failed
assert_invariant() (debug builds), instead of letting std::bad_alloc
reach the caller.
copy_iteratively() set a pending worklist element's type right after
popping it, before the next loop iteration created its container in
copy_array_level()/copy_object_level(). Move that type assignment into
those two functions, right after the container is successfully
created, and drop the premature one in copy_iteratively(), so a
half-built element stays a null value - as copy_shallow()'s comment
already promised - until it can safely hold one.
Add a regression test to tests/src/unit-allocator.cpp that copies a
value nested 130 levels deep (both arrays and objects, with a
std::map- and an ordered_map-backed object_t) and fails every
allocation of the copy in turn: each attempt must throw std::bad_alloc
without crashing, and the source must stay unchanged.
Fixes#5640.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
For an object type whose comparator treats unequal keys as equivalent
(for example a std::map with a case-insensitive comparator),
compare_iteratively() looked up a mismatched left key in the right
object with find(), which uses the object's own comparator, and
accepted whatever entry it found without checking that the keys are
actually equal. A case-insensitive comparator then found "KEY" for
"key", so two objects nested past the recursion bound (or at every
depth with JSON_NO_THREAD_LOCAL) could compare equal even though the
object type's own operator== - and basic_json itself, below the bound
- consider them different.
Accept the found entry only if its key equals (not just compares
equivalent to) the looked-up key.
Fixes#5655.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json(first, last) treated value_t::binary like the structured
types (array, object) in the range check, so it always copied the
whole binary value regardless of the iterators, even for an empty
range such as (b.end(), b.end()). The other primitive types (number,
boolean, string) already reject such a range with
invalid_iterator.204, and erase(first, last) already does the same
for binary values, so this made the constructor inconsistent with
both. Move case value_t::binary into the group of checked primitive
types.
Also update the two matching passages in basic_json.md that describe
overload 7, and add a version-history note.
Fixes#5670.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When a parser callback rejects an object's or array's start event,
json_sax_dom_callback_parser kept calling it for everything inside
that container anyway: nested keys, values, and the start/end events
of containers below it. This contradicts parser_callback_t's own
documentation, which promises that discarding a container at its
start event also hides its content from the callback.
The same code path also kept a full copy of every key inside such a
discarded container in key_stack until the whole parse finished,
because the early return for values that are not stored skipped the
matching pop. Filtering out a large subtree is the main reason to use
a callback, so this made peak memory during the parse scale with the
size of the very subtree the callback was trying to skip.
Fix start_object(), start_array(), and key() so that a container
whose own start event was discarded, or that is nested inside one, is
never handed to the callback, and no longer pushes onto the key
stacks. A container whose start event was accepted but whose key was
rejected still gets its content reported, as documented ("the
callback is still called for the associated value, but its return
value has no further effect"); only its own bookkeeping is skipped
since it will not be stored.
Fixes#5643.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
j.contains(0), j.find(0), and j.count(0) used to compile: the literal 0 is
a null pointer constant, so it converts to a null const char*, and the
overloads taking const typename object_t::key_type& accepted it by
constructing a std::string from that null pointer, which is undefined
behavior (a crash with both libc++ and libstdc++). value(0, default_value)
had the same problem in C++11, where the object comparator is not
transparent.
Add deleted overloads for integral arguments to contains(), find()
(const and non-const), count(), and value() so that these calls are
compile errors in every supported language mode instead of crashing.
Calls with string, string_view, json_pointer, and size-typed element
access (at(), operator[](), erase()) are unaffected.
Fixes#5657.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
to_bson() rejected a binary value's subtype above 255 (out_of_range.415)
in write_bson_binary(), which only has the binary_t, not the basic_json
value that holds it, so the exception was created with no JSON_DIAGNOSTICS
context even though the equivalent to_msgpack() check names the value's
path. The check also ran after the document size, all preceding elements,
and this element's header and length had already reached the output
adapter, so a caller-provided std::vector or std::string ended up holding
a truncated document.
calc_bson_sizes() already walks every value before anything is written,
to size embedded documents and arrays and to reject invalid keys
(out_of_range.409) up front. The subtype check now runs there instead,
in calc_bson_binary_size(), which is given the basic_json value so the
exception can use it as context. The now-redundant check in
write_bson_binary() is removed, since calc_bson_sizes() always throws
first if any binary value in the document has an oversized subtype.
Fixes#5675.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16,
UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]"))
failed with parse_error.101 at the terminating NUL of the literal,
and accept() returned false. The array overload of input_adapter()
only dropped the compiler-added trailing '\0' for arrays of char,
so for wchar_t, char16_t, char32_t, and char8_t arrays that
terminator was passed to the parser as data, which the macro then
rejected.
Broaden the trimming to every character type that a string literal
can use (char, wchar_t, char16_t, char32_t, and, since C++20,
char8_t). Arrays of any other element type (unsigned char,
std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a
trailing zero byte there is still read as genuine data.
Fixes#5658.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Parsing from a std::istream crashed in two unusual but valid stream
states, both in input_stream_adapter:
- With eofbit in the stream's exceptions() mask, get_character() sets
eofbit via is->clear(), which throws std::ios_base::failure. While
that exception unwinds, ~input_stream_adapter() called clear() again
to reset eofbit, which is still set and still in the exception mask,
so it throws a second time out of the (implicitly noexcept)
destructor and std::terminate() is called. The destructor now only
calls clear() if a bit other than eofbit remains set, so the first
exception can propagate normally.
- For an std::istream without a stream buffer (rdbuf() == nullptr,
e.g. std::istream(nullptr)), the constructor stored the null
pointer without checking it, and get_character() dereferenced it.
input_adapter(std::istream&) now throws parse_error.101 for such a
stream, the same as it already does for a null FILE* or char*.
Added regression tests to unit-deserialization.cpp and, for the
JSON_PRECISE_STREAM_POSITION variant of get_character(), to
unit-precise-stream-position.cpp; both crashed before this fix.
Documented the two exceptions in parse.md and operator_gtgt.md.
Fixes#5646.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
from_json built its out_of_range.410 message with "..." + j.dump(). If the
unmatched value is (or contains) a string with invalid UTF-8, that dump()
itself throws type_error.316, so the caller got type_error.316 instead of
the documented out_of_range.410; such strings can reach get<Enum>()
unvalidated, e.g. from from_cbor()/from_msgpack(). With a custom string_t,
j.dump() returns that type, and "const char*" + string_t does not compile
unless the type happens to provide operator+, so the macro failed to
compile for such types.
Build the message with detail::concat(), which appends any type exposing
data()/size() and always yields a std::string, and dump with
error_handler_t::replace so building the message itself cannot throw.
Added regression tests: an invalid-UTF-8 case in the existing strict-enum
test in unit-conversions.cpp, and a strict-enum use with alt_string (the
custom string_t from unit-alt-string.cpp) to cover the compile failure.
Fixes#5667.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json::swap() (and the friend swap() and the pre-C++20 std::swap
overload that forward to it) only exchanged m_data.m_type/m_data.m_value,
leaving each value's json_base_class_t subobject in place. This is
inconsistent with the copy and move constructors and copy assignment,
which all carry the base class along with the value, so after
a.swap(b) any metadata stored in a CustomBaseClass ended up attached to
the wrong value. Algorithms that mix swap() with moves, such as
std::sort, scrambled the metadata across the whole container.
Fix the member swap() to also exchange the json_base_class_t subobject
and extend the noexcept specifications of swap() and the friend swap()
accordingly.
Fixes#5653.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix to_msgpack() reading the inactive number union member
basic_json stores number_integer and number_unsigned in a union, and
number_unsigned_t only has to be at least as wide as number_integer_t
(with the default types, both are 64-bit and have the same
representation). When number_integer_t is narrower, write_msgpack()
read the wrong union member in two places:
- The number_unsigned case wrote number_integer's bits instead of
number_unsigned's, silently writing the wrong value whenever it
did not fit in number_integer_t.
- The number_integer case (non-negative branch) picked the encoded
width by comparing number_unsigned's bits, which is undefined
behavior, though the value written was still number_integer's, so
at worst a too-wide encoding was chosen.
Read the active member in both cases, like the other binary writers
(CBOR, UBJSON, BJData, BSON, BON8) already do.
Fixes#5644.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Cast number_integer to number_unsigned_t only once in to_msgpack()
Addresses review comment by @gregmarr.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Copy values before inserting an initializer list into an array
insert(pos, {...}) inserted wrong values when the initializer list
contained const references to elements of the array being inserted
into. json_ref stores only a pointer for a const lvalue, so the
initializer_list_t range passed straight to the array's range insert
aliased the array's own storage; std::vector::insert(pos, first, last)
may move or shift elements before copying from that range, so the
source elements were already stale by the time they were read
(different wrong results on libc++ and libstdc++).
Copy the referenced values into a temporary array_t first, then move
that temporary into place, so the source range never aliases the
array being modified.
Fixes#5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the reserve_array helper in the initializer_list insert fix
The previous commit called array_t::reserve() directly on the
temporary buffer used to copy an ilist's values before inserting.
std::deque, a documented ArrayType (tests/src/unit-custom-array-type.cpp),
has no reserve(), so insert(pos, initializer_list) no longer compiled
for it. Use the existing detail::reserve_array() SFINAE helper (already
used by the SAX DOM parser) instead, which leaves array types without
reserve() untouched.
Added a regression check that deque_json::insert(pos, {...}) compiles
and handles the aliasing case from #5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Take diff()'s fast path unless the object type reorders members
For every object type except an insertion-ordered one like ordered_map,
diff() no longer produced a member-by-member patch when target had a key
that sorts before a key the two objects share: it fell through to the
slow path, which removes every member of source and re-adds every member
of target, instead of just adding the new key.
#5465 added an order check to require the fast path to also reproduce
target's member order, needed because ordered_map's patch()-driven "add"
appends a new member at the end. The check compared the common keys'
order between source and target and also required that every added key
come after every common key in target's order ("new_keys_form_suffix").
The comment above it argued this check is always true for std::map, and
that reasoning is correct for the order of the common keys themselves,
but not for new_keys_form_suffix: a std::map iterates in sorted key
order, so a new key that sorts before an existing common key is
enumerated between common keys, making new_keys_form_suffix false even
though std::map's own key order does not need reordering at all - it
places every member itself, regardless of insertion history, so a
member-by-member diff already reproduces target's iteration order.
Only require the order check for an object type that keeps insertion
order, using the same detail::is_ordered_map trait the library already
uses to recognize such an object type in set_parent(). Every other
object type - std::map in key order, a hash map in an order its
operator== ignores - always takes the fast path.
Added a regression test to unit-json_patch.cpp: the issue's example now
yields a single "add" op for json, while ordered_json still takes the
slow path to reproduce target's member order.
Fixes#5639.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Only track the target key order in diff() for insertion-ordered objects
common_keys_target_order and new_keys_form_suffix are only read when object_t keeps its members in insertion order; skip building them otherwise. Addresses review comment by @gregmarr.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
For idx == SIZE_MAX, the non-const array operator[] computed the new size as
idx + 1, which wraps to 0. resize(0) then emptied the array, and the
subsequent operator[](idx) on the now-empty vector wrote one element before
its buffer. Every other too-large index (e.g. SIZE_MAX - 1) already went
through resize(), which throws std::length_error and leaves the array
unchanged; SIZE_MAX was the one value for which the overflow bypassed that
safety net.
Add a guard that throws std::length_error before computing idx + 1 when idx
is the largest representable size_type value, so the array is left
unchanged, matching the exception vector::resize() already throws for
smaller (but still too large) indices.
Fixes#5647.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Let to_json(std::optional<T>) propagate exceptions from T's to_json
The overload was marked noexcept even though its body assigns *opt to
the JSON value, which calls T's to_json (or allocates for std::string,
std::vector, or json). Any exception from there -- a user-defined
to_json reporting an error, or std::bad_alloc -- called std::terminate()
instead of propagating. The noexcept also made basic_json's converting
constructor noexcept(true) for std::optional<T>, so json j = opt; could
not report the error either.
Fixes#5642.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI: make to_json(std::optional<T>) conditionally noexcept
GCC's -Wnoexcept (an error in ci_test_gcc) fired at to_json_fn's
noexcept(noexcept(to_json(j, val))): after dropping the unconditional
noexcept, to_json(std::optional<int>) had no exception specification
although GCC could prove its body cannot throw. It also made
json(std::optional<int>) lose its noexcept.
Declare the overload noexcept exactly when assigning the contained
value to the JSON value is (std::is_nothrow_assignable<BasicJsonType&,
const T&>), which is what the body does. std::optional<int> is noexcept
again; a T whose to_json may throw still propagates the exception.
Static assertions in the test check both cases.
The test's throwing to_json triggered -Wmissing-prototypes and
-Wmissing-noreturn (clang) and -Wmissing-declarations and
-Wsuggest-attribute=noreturn (GCC). Move the type and its to_json into
an anonymous namespace, mark the function [[noreturn]], and compile
them only without JSON_NOEXCEPTION, like the test that uses them.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
value(KeyType&&, default) is constrained on is_comparable_with_object_key,
which passes KeyType as a reference. is_comparable's dispatch on
is_json_pointer_of<A, B> only matches a json_pointer as a plain type or a
plain reference, so a const-qualified reference (as produced when KeyType
is deduced from a json_pointer argument) fell through to
is_comparable_no_json_pointer, which instantiates the deprecated
json_pointer/string comparison operators. This made ordered_json's
transparent comparator (and any transparent comparator on a custom string
type) warn under -Wdeprecated-declarations when calling
value(json_pointer, default), even though no such comparison is ever
performed. at() was already fixed for this in #5289, which does not use
is_comparable_with_object_key.
Strip references and cv-qualifiers with uncvref_t before the
is_json_pointer_of dispatch, so any reference-to-json_pointer is
recognized regardless of qualifiers.
Fixes#5664.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When converting a JSON array to std::map or std::unordered_map with a
non-string key, each element must itself be a [key, value] array. If an
element is not an array, from_json() threw type_error 302 with the outer
array's value (&j) as the exception context, so with JSON_DIAGNOSTICS
enabled the message pointed at the whole array instead of the offending
element (e.g. "(/outer/m)" instead of "(/outer/m/2)"), even though the
message text already described the element's type.
Both from_json() overloads now pass the element (&p) as the context, so
the reported JSON Pointer matches the type named in the message, the
same way std::vector<std::vector<T>> and similar conversions already do.
Fixes#5668.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clear() to also reset the subtype of a binary value
clear() on a binary value cleared the bytes but left the subtype
untouched, so the result was not equal to a default-constructed
binary value even though the documentation says clear() has the same
effect as *this = basic_json(type()). The fix calls
byte_container_with_subtype::clear_subtype() alongside the existing
clear() call.
Extended the "filled binary" clear() test in unit-modifiers.cpp with
a case that uses a subtype, since the existing cases only covered
binary values without one.
Fixes#5669.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the table alignment in clear.md
Addresses review comment by @gregmarr.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_arithmetic_value() rejects boolean_t, so the default from_json()
for enums failed to compile for an enum whose underlying type is
bool (e.g. enum class Flag : bool { off, on }), even though the
matching to_json() serializes such enums as an unsigned number.
Read the underlying value through number_unsigned_t in that case,
matching what to_json() writes, then cast back to the underlying
type before constructing the enum.
Fixes#5671.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Copying and comparing stopped their descent at
basic_json::nesting_depth_limit(), while serializing, hashing and
merging used detail::recursion_depth_limit(). Both were 128, but
nothing kept them equal. The thread-local count now tests against
detail::recursion_depth_limit() as well, and a static_assert keeps
the limit small enough for the byte that holds the count.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Do not throw in contains() for an empty array reference token
json_pointer::contains() rejected malformed array indices, but an empty
reference token (e.g. "/a/" where "a" is an array, or "/" on an array)
passed every check and reached array_index(), which throws
out_of_range.404. contains() must not throw (cf. #5395), so it now
returns false for an empty token. at() still throws out_of_range.404.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI: bind j_nested_const by reference in the json_pointer test
clang-tidy (ci_clang_tidy) flagged the new test with
performance-unnecessary-copy-initialization: the local copy
j_nested_const of j_nested is never modified. Bind it as a const
reference instead; it still exercises the const overloads of at() and
contains().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move instead of deep-copy ordered_json values when an object grows
ordered_map keeps its elements in a std::vector<std::pair<const Key, T>>.
With a std::string key, that pair is not nothrow move constructible (the
const key has to be copied), so std::vector copies every element when it
reallocates. For ordered_json, this deep-copies every member value an
object already holds, including whole nested subtrees, on each growth
step.
Grow the storage in ordered_map instead, copying the keys and moving the
values. This happens in two phases, so the strong exception guarantee is
kept without try/catch. The first phase may throw, but only touches a
temporary buffer: it copies the keys, value-initializes the values, and
constructs the new element. The second phase moves the values (noexcept)
and swaps the buffers. Because the new element is constructed before any
value is moved, arguments that refer to elements of the container stay
valid, as with std::vector. Types that cannot take this path keep the
std::vector behavior.
Parsing into ordered_json (ParseStringOrdered, Apple M1 Max, clang -O3):
twitter 3.20 -> 1.70 ms, citm_catalog 7.73 -> 3.67 ms, jeopardy 219 ->
177 ms, canada unchanged. The number of allocations for twitter and
citm_catalog drops by two thirds.
Also add ParseStringOrdered rows to the benchmarks, and document the
growth behavior and the exception safety of ordered_map.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI: skip std::pair noexcept assumptions on EDG-based compilers
ci_icpc and ci_nvhpc failed to compile unit-ordered_map.cpp: the static
assertion that std::pair<const std::string, ordered_json> is not nothrow
move-constructible fails there. The EDG front end (Intel icpc 2021.10,
NVIDIA nvc++ 25.5) considers the defaulted move constructor of
std::pair<const Key, T> noexcept even if copying Key can throw. With these
compilers, std::vector already moves such elements itself when it grows,
and ordered_map correctly leaves growing to it.
The same misjudgement makes std::vector call std::terminate when a key copy
throws during growth, so the exception-safety test with throwing_key would
abort on these compilers as well.
Skip the static assertion and the exception-safety section when __EDG__ is
defined. Verified with icpc 2021.10 (-std=gnu++11) and nvc++ 25.5 (C++11 and
C++17) on Compiler Explorer: unit-ordered_map and unit-disabled_exceptions
build and pass.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Convert floats with Eisel-Lemire when std::from_chars is unavailable
Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit
coordinates of canada.json) went to strtod unless std::from_chars was
available. It is not used in C++11/14, and not with libc++, which does not
define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's
compute_float) now converts them with integer arithmetic, correctly rounded
for any token with at most 19 significant digits. Longer tokens are
truncated; the result is used if w and w + 1 round alike, else strtod
decides as before. Overflow still yields infinity (out_of_range.406).
The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit
test recomputes every entry with big-integer arithmetic. Further tests:
known values generated with Python (whose float() is correctly rounded),
200,000 round trips through to_chars, and the 128-bit multiplication and
leading-zero count against big-integer references (both with and without a
128-bit type). Checked against strtod on 6.5 million tokens, among them
60,000 exact halfway cases: no difference.
json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang);
other files unchanged. Compile time of a TU including json.hpp: +0.7%.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI: unused parse result and find() == npos in the Eisel-Lemire tests
GCC (-Werror=unused-result) rejected CHECK_THROWS_WITH_AS(json::parse(...))
because parse() is [[nodiscard]]; assign the result to a dummy json as the
other tests do. clang-tidy flagged longer.find('.') == npos with
abseil-string-find-str-contains; store the position in a variable first.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
lexer::convert_number() converted float tokens with std::from_chars (when
available), Clinger's fast path, and the locale-aware strtod fallback, all
as lexer members. They are now free functions in number_parse.hpp:
- convert_float_fast(): std::from_chars, then Clinger's fast path, skipped
when the mantissa has too many significant digits
- convert_float_locale_aware(): strtof/strtod/strtold with the decimal point
of the current locale, retried when the locale changed (#5198)
so that other code converting JSON number tokens gets the same values. No
change in behavior; the lexer no longer includes <clocale> and <cstdlib>.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A header that builds on json.hpp (such as the planned json_view.hpp) must
not inline json.hpp: its single-header version would contain a second copy
of the library, and that copy would change with every library change.
The optional config key "external" lists include paths that are kept as
#include directives. Only the first directive per path is kept; repeated
ones are commented out, as the tool already does for inlined headers.
config_json_view.json uses it for json_view.hpp; the existing configs do
not set it, and json.hpp and json_fwd.hpp regenerate byte-identically.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently
sized signed integer and then repeated the same "reject negative with
error 113" check; only 'L' additionally checked value_in_range_of for
the out_of_range.408 case. Any change to that error path had to be
made four times.
Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read,
the negative check and the range check once, and route all four
markers through it. The range check is a no-op for 'i'/'I'/'l' (their
values always fit std::size_t) and only live for 'L' on a 32-bit
std::size_t target, matching today's behavior exactly.
In the ndarray dimension-product loop, the preceding loop already
returns early on any zero dimension and result starts at 1, so `i > 0`
in the pre-multiplication overflow check was always true, and
`result == 0` in the post-multiplication check could not be reached
either: two positive factors whose product does not overflow (as the
pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&`
and narrow the post-check to `result == npos`, the one case the
pre-check cannot rule out (an exact, non-overflowing match with the
sentinel reserved for unknown-size containers), with a comment
explaining why.
Verified byte-for-byte identical behavior before/after with a
standalone probe covering negative counts for every marker, a matching
positive count, and ndarray inputs, plus the full unit-ubjson and
unit-bjdata suites (same assertion counts as before this change).
Overlaps #5601 (rewrites the four parse_error calls and the overflow
checks touched here) and #5607/#5707 (touch neighboring lines in the
same functions).
#5711 item 6
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
binary_reader's constructor stores the format in the input_format
member, and sax_parse(format, sax_, strict, tag_handler) took the same
value again purely to dispatch on it. Every in-tree caller passed the
same value both times (all 16 from_cbor/from_msgpack/from_ubjson/
from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three
public basic_json::sax_parse() overloads), so nothing was broken
today, but a caller of the detail class directly (only reachable via
JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could
pass a mismatched pair - say bjdata to the constructor and ubjson to
sax_parse - and dispatch on one format while applying the other
format's rules; the default-constructed input_format_t::json reader
would additionally hit JSON_ASSERT(false) in exception_message() on
its first error.
Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the
existing overload with the stored input_format, and switch every
caller to it: the 16 from_*() sites (keeping their
`// cppcheck-suppress[accessMoved]` comments) and the three
basic_json::sax_parse() overloads, all of which already had the format
available from their own `format` parameter. The four-argument overload
is kept for anyone still calling it, now with
JSON_ASSERT(format == input_format) so a mismatch fails immediately
in a debug build (assert-enabled binaries, including the fuzzers and
test suite) instead of misbehaving; verified with a probe that
constructs a reader for one format and calls the explicit overload
with another, which aborts on that assertion as expected.
Removing or asserting against the constructor's input_format_t::json
default, which would affect direct detail users, is left as a separate
decision per #5711 item 5.
Overlaps #5601, which is expected to add an AllowRecovery template
parameter to sax_parse() and touch these same call sites in json.hpp.
#5711 item 5
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Every container is opened through enter_container(), whose docs
promise that a check placed there runs before every start event. The
close side had no equivalent: the same
"container_stack.pop_back(); dispatch to end_object() or end_array()"
sequence was written out separately in BSON, CBOR, MessagePack,
UBJSON/BJData and BON8, each copying the pattern of keeping an
is_object flag around the pop_back() that would otherwise invalidate
a reference to it. A check needed on close would have had to be added
in five places, and a sixth copy could go unnoticed.
Add leave_container() next to enter_container(), doing the same
pop-then-dispatch, and replace the five sites with it. Each site keeps
its own surrounding logic (BSON's check_bson_document_size() call
before popping, MessagePack's is_object copy used again below,
UBJSON/BJData's remaining-container handling after popping, BON8's
top used again below); only the repeated pop/dispatch line pair is
now shared.
Verified all six binary-format unit suites and unit-regression2's
deep-nesting tests (dependent count/reuse count and the bjdata ndarray
depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan.
Overlaps #5601, which is expected to add a sixth close site in its own
skip loop; that site can route through leave_container() too once it
lands.
#5711 item 4
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte
big-endian unsigned integer" ladder four times over:
- twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore
branch and once, nearly identically, in the ::store branch (~90
lines to read one integer);
- twice more for container lengths, once for array heads 0x98-0x9B and
once for map heads 0xB8-0xBB, where the 1/2-byte forms called
enter_array()/enter_object() directly and the 4/8-byte forms
additionally went through get_cbor_container_size().
Add get_cbor_argument(std::uint64_t&), reading the width selected by
current & 0x1F via the same get_number() calls as before (so EOF is
reported exactly as before), and route all four sites through it:
- 0xD8-0xDB now read the argument once per branch instead of switching
on `current` a second time; behavior split cleanly from embedded tags
0xC0-0xD7 (tag value in the head, no argument to read), which is now
its own case block that no longer has to fall into the ::store
switch's "default" case to reach the same tag_pending = true; return
true; outcome.
- 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always
going through get_cbor_container_size() (harmless for 1/2-byte
lengths, which already always fit).
Verified byte-for-byte identical behavior before/after with a
standalone probe covering embedded and multi-byte tags under all three
tag_handler_t settings, a tag over a byte string (subtype path),
truncated tag/length arguments of every width, and array/map lengths
of every width, including the out_of_range.408 "excessive size" case:
same exceptions, same messages, same chars_read, same successful
results.
Left the string/byte-string length ladders in get_cbor_string()/
get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is
expected to touch them separately.
Overlaps #5601 (adds a branch right above the embedded-tag case) and
#5607 (touches the integer cases 0x18-0x1B, which share this ladder's
shape in separate hunks).
#5711 item 2
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
binary_reader held bjd_optimized_type_markers and bjd_types_map as
non-static const members (12 string_t objects for the type-name table),
built and destroyed on every from_cbor/from_msgpack/from_bson/
from_ubjson/from_bon8/from_bjdata call even though only from_bjdata
ever reads them. They also needed the #define/decltype/#undef
workaround from #3637 and two NOLINTNEXTLINE suppressions, and
binary_writer already carries the same two lists in another form
(is_bjdata_excluded_type_marker() and a local std::map in
write_bjdata_ndarray(), the latter removed by the item-4 commit), so
the excluded-marker lists could drift apart.
Replace bjd_optimized_type_markers with static constexpr
is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain
as binary_writer's is_bjdata_excluded_type_marker(). Replace
bjd_types_map with a non-constexpr static bjd_type_name(char_int_type)
switch returning nullptr for an unknown marker (a C++11 constexpr
function cannot contain a switch). Delete both
JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the
NOLINTNEXTLINE suppressions, detail::make_array() (no longer used
anywhere), and the now-unused <algorithm> and <array> includes.
Update the two call sites (the ND-array excluded-type check and the
_ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's
"LUT arrays are sorted" section, which only checked the two tables'
internal ordering, with a check of all 12 type names and all 8
excluded markers against both new functions.
#5711 item 1
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
calc_bson_sizes() and write_bson_document() are a hand-synchronized
pair of passes over the same object/array tree, introduced by #5553:
the size pass appends to nested_sizes in visiting order, and the write
pass consumes the table by position with nested_sizes[next_size++].
Nothing checked that the write pass consumed the whole table. If a
future change touched only one of the two passes - for example to skip
or reject an entry - every later size prefix in the document would be
silently wrong.
Add JSON_ASSERT(next_size == nested_sizes.size()) where
write_bson_document() returns, so such a future drift between the two
passes is caught immediately (JSON_ASSERT expands to nothing in
release builds using assert(), and the fuzzers/tests already build
with it enabled). The two passes agree today, so this changes nothing
observable; it only guards against the risk described in #5710 item 5.
Extracting a shared stepper for the two passes (the second half of the
proposed change) is left for a follow-up: it only saves ~30 lines and
the issue asks for it only if the result reads clearly, which needs
more room to get right than a mechanical cleanup pass allows.
#5710 item 5
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on
every call just to translate the _ArrayType_ name to a dtype marker
(the only reason binary_writer.hpp included <map>), then mapped dtype
to C++ type twice more: once as a switch for the range-check pass and
once as a separate if/else chain for the write pass, with nothing
checking that the two agreed. The caller also ran three at() lookups,
and the callee called value.at(key) about ten more times for the same
three members.
Replace the map with bjdata_ndarray_type_marker(), a plain string
comparison chain (a C++11 constexpr function cannot contain a switch,
so this mirrors binary_reader's own static table style). Replace the
switch/if-chain pair with one write_bjdata_ndarray_elements() that
switches on dtype once and calls a per-type helper -
write_bjdata_ndarray_element<T>() for the eight integer dtypes and
write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag
selecting the range check or the actual write, so the two passes can
no longer disagree on the type. _ArrayType_, _ArraySize_ and
_ArrayData_ are now looked up once into references, and the four
header marker bytes ('[', '$', '#') are written through to_char_type()
like the rest of the UBJSON/BJData writer.
The 'd' (single-precision) rule is left exactly as before, since #5707
is expected to change it separately.
Verified byte-for-byte identical output before/after for every dtype
(including the Draft 2/Draft 3 'byte' fallback and the use_count/
use_type combinations) via a standalone probe, plus round-tripping
through from_bjdata().
Overlaps #5707, which is expected to touch the 'd' dtype case, and
#5518, which is expected to move the write_bjdata_ndarray() call site.
#5710 item 4
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Four formats picked between a float32 and float64 marker through four
different helper styles: dummy-argument overloads for CBOR and
MessagePack, an std::is_same template for BON8, and a runtime if-chain
on input_format_t for write_compact_float(). With number_float_t set
to long double, to_cbor, to_msgpack and to_ubjson failed inside the
library with "call to 'get_cbor_float_prefix' is ambiguous", while
to_bson kept working because write_bson_double() takes a plain double.
Change write_compact_float() to take the two marker bytes directly
(each of its three callers already knows them at compile time) instead
of an input_format_t it only forwarded, and delete the now-unused
get_cbor_float_prefix(), get_msgpack_float_prefix(),
get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn
the two get_ubjson_float_prefix() overloads into one template. Both
write_compact_float() and get_ubjson_float_prefix() now report an
unsupported number_float_t with a static_assert naming the requirement,
rather than an ambiguous-overload error; the assert lives in the
function body, not the class scope, so to_bson with long double is
unaffected.
Verified with a probe basic_json<..., long double>: to_bson still
compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail
to compile with the new static_assert message.
This changes the text of an existing compile error for users with an
unsupported number_float_t (documented as a public-API-visible change
in #5710).
#5710 item 1
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The number_integer (non-negative branch) and number_unsigned cases in
write_msgpack() each held their own copy of the fixint/uint8/16/32/64
ladder, kept in lockstep only by a comment ("we used the code from the
value_t::number_unsigned case here"). Both copies mixed union members:
the signed copy compared number_unsigned but wrote number_integer, and
vice versa.
Extract write_msgpack_unsigned(std::uint64_t), mirroring how
write_cbor_head() already avoids the same duplication for CBOR, and
call it from both cases. Each case now reads only its own active
union member. Output bytes are unchanged for the default 64-bit
number types.
#5710 item 3
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
binary_reader had two ~45-line copies of the IEEE 754 half-precision
decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is
normalised, the two blocks were identical except for the byte order
used to assemble the 16-bit half (CBOR is big endian, BJData is little
endian). Any future change to half-float decoding had to be made and
kept in sync in both places.
Add one get_half_float(format, little_endian) helper that does the two
get()/unexpect_eof() reads, assembles the half in the requested byte
order, decodes it per RFC 8949 Appendix D, and calls sax->number_float.
Both cases now just call it with their byte order; the BJData case
keeps its bjdata-only guard.
Behavior, the public API and the ABI are unchanged. Verified with a
scratch probe comparing the old and new decoders bit-for-bit (NaN by
isnan()) over all 65536 wire byte pairs, in both formats, and by
running unit-cbor and unit-bjdata (offline, against the stubbed
test_data.hpp).
Part of #5711
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-recursive rewrite of the binary readers (#5505, #5506, #5507)
left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char
parameters dead: parse_cbor_internal() has one caller and it always
passes true, and parse_ubjson_internal() has one caller and it always
uses the true default. Both parameters, and the @param docs describing
the "reuse the last character" mode they used to select, no longer
correspond to anything.
Drop both parameters, initialise fetch/prefix unconditionally, and
update the two call sites in sax_parse(). parse_cbor_value()'s and
get_ubjson_string()'s own get_char parameters are unrelated and are
left alone; both still have a false caller.
Also delete a stray `@return whether a valid MessagePack value was
passed to the SAX parser` doxygen block that sits directly above
parse_msgpack_value()'s real doc comment, a leftover of the same
rewrite.
Behavior, the public API and the ABI are unchanged; these are private
members of detail::binary_reader. Verified by compiling with
-Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and
unit-msgpack (offline, against the stubbed test_data.hpp).
Part of #5711
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
write_number_with_ubjson_prefix() (unsigned and signed overloads) and
ubjson_prefix() (number_integer and number_unsigned cases) each picked
the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their
own independent if/else ladder, and the values beyond 64 bits were
handled by a second, tag-dispatched pair of ladders. An optimized
container announces the marker of its first element via ubjson_prefix()
and then writes every element through write_number_with_ubjson_prefix(),
so the two had to be kept in lockstep by hand across four call sites.
Replace all of that with one ubjson_integer_prefix() built on
value_in_range_of<T>, and one write_ubjson_integer_payload() that
writes the value (or, for 'H', the decimal digits) for a given marker.
write_number_with_ubjson_prefix() and ubjson_prefix() keep their
signatures and now just call these two helpers.
Behavior, the public API and the ABI are unchanged. Verified with a
new regression test covering scalars and $-optimized arrays/objects at
every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for
to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing
to_ubjson/to_bjdata output before and after over the json_test_data
corpus (bit-identical).
Part of #5710
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The doc block of write_number() ended up above the byte_swap() helpers
added in #5286, about 80 lines from the function. It was also a plain
comment that Doxygen skips, said "write a number to output input", and
left BON8 out of the big-endian formats. Move it back onto
write_number() as a /*! block and fix the text.
write_bson() documented "@pre j.type() == value_t::object", but it
throws type_error.317 for every other type, and to_bson() relies on
that. Document the exception instead.
Explain why the CBOR binary subtype is always written with a 0xD8..0xDB
head and never in the one-byte tag form: binary_reader with
cbor_tag_handler_t::store only keeps those heads as a subtype, so
switching to write_cbor_head() would break round trips for subtypes
0..23.
Also fix the grammar of the to_char_type comment. Comments only; no
change in behavior, API or ABI.
Part of #5710
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:
- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
calls whose result was discarded. Check the result, like the other
resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
callbacks, which cannot throw but were not declared noexcept.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Look up the locale decimal point at conversion time, not lexer construction
The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.
token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.
Fixes#5198
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop the strtod retry loop when the decimal point is unchanged
convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.
Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix -Weffc++ errors in the #5198 locale test
GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the key type when rejecting non-string CBOR/MessagePack map keys
CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:
syntax error while parsing MessagePack object key: only string keys
are supported, but found nil; last byte: 0xC0
The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.
Refs #2766, #3381
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point the MessagePack key note to the spec's profile section
The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The "Size above uint32" tests for arrays and objects fake a container
size of 2^32 and expect to_msgpack() to throw out_of_range.412. But
to_msgpack(j) first reserves binary_reserve_hint(j) bytes, which is
size + 1 for arrays and 2 * size + 1 for objects, i.e. 4 or 8 GiB.
Linux and macOS overcommit, so the reservation succeeds; on Windows it
throws std::bad_alloc before the size check is reached (seen with
msvc-vs2026 Debug x64 on the object test).
Write into a caller-owned vector instead, so nothing is reserved.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5559 added a 224-character line to cbor_tag_handler_t.md; the
documentation style check allows at most 160.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The line added in #5559 exceeded the 160-character limit enforced by
docs/mkdocs/scripts/check_structure.py, breaking the documentation build.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests added by #5515 fail two ways on develop:
- clang-tidy reports the size() overrides of huge_string and huge_binary
(readability-convert-member-functions-to-static) and the non-const
test value (misc-const-correctness); mark them like the #5584 types
- clang with libstdc++ 10 cannot compile the file for C++17: the
std::filesystem::path conversion considered for huge_string, a class
derived from std::string, is ambiguous. Guard it with
JSON_TEST_BEYOND_UINT32_STRING, which #5584 introduced for the same
reason, and define that macro before both test blocks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add BON8 support
Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.
The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.
The round-trip tests need the .bon8 files of json_test_data 3.2.0.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Address review comments
- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Select the BON8 float prefix by type
get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Rename a test variable that Flawfinder mistakes for read()
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures
- compare the float in write_bon8_float with number_float_t constants,
so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BON8 strings in bulk from contiguous input
- copy the valid UTF-8 of a string in one step when the input is
contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Link the BON8 functions from the other binary format pages
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the bulk scan flag after the input, not BON8
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Read BSON keys in bulk from contiguous input
BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the BON8 CI failures of the bulk-read tests
- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
when exceptions are disabled: they catch the parse errors of invalid
input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the explicit basic_json instantiation into its own test file
Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).
Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Convert the bytes of the BON8 test strings explicitly
The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The tests fake a container size of UINT32_MAX + 1, which does not fit
into a 32-bit std::size_t: MSVC rejects the truncation (C4305/C4309
with /WX), and clang-cl wraps the size to 0 so nothing throws. Guard
them with SIZE_MAX > UINT32_MAX like the tests from #5584.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
When using cbor_tag_handler_t::store, tags 0xD8-0xDB previously assumed
that the tagged item was a byte string, unconditionally attempting to
parse binary data and failing on valid CBOR documents containing tags
applied to integers, strings, arrays, or objects (such as self-describe
tag 55799).
Check whether the tagged data item is a byte string (0x40-0x5B or 0x5F).
If it is a byte string, store the subtype on the binary value as before.
Otherwise, iteratively process the tagged value in the driver loop using
item_read so that chained tags do not consume native stack space.
Part of #5316.
Signed-off-by: ReturnKartikey <kartikeynegi2000.work@gmail.com>
* Cut test suite runtime in binary roundtrips and integer sweeps
The Linux CI jobs pass --no-skip, so skip() does not help there.
Parse each corpus file once in the binary roundtrip loops instead of
four times. Sample the 16-bit integer ranges with stride 7 (still hits
every low byte) and always keep the endpoints.
Also drop the 5M-node parse test to 500k, which still covers the
non-recursive destructor, and move jeopardy.json into its own skipped
test so the cheaper binary-format size checks actually run.
See #5418.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Drop useless int32_t casts in the sampled integer loops
ci_test_gcc compiles with -Werror=useless-cast. On that compiler
int32_t is int, so static_cast<int32_t> of the loop bound is an
error. The bounds are already int, and the sampled values do not
change.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Revert unit-binary_formats.cpp to develop and fix comment
Revert tests/src/unit-binary_formats.cpp to its develop state.
The test-case split made valgrind jobs slower instead of faster,
because the cheaper corpus files (canada/twitter/citm/sample)
now ran under valgrind where they never did before.
Fix the next_integer_sample comment: the function has no 'first'
parameter, so describe what the function actually does.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
---------
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
* Throw instead of writing MessagePack lengths beyond UINT32_MAX
MessagePack stores the length of a string, binary value, array, or
object in at most 32 bits. For a larger value, to_msgpack wrote no length
at all, so the output could not be read back. It now throws
out_of_range.412, which BSON already uses for its 32-bit length fields.
The check lives in one function, so each length is written by an
if/else chain that ends in a plain else, without a condition that can
never be false. It is tested with string and binary types that report a
size beyond UINT32_MAX without allocating it, like the BSON tests do.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the CI failures of the MessagePack length check
- mark to_msgpack_length's value as used when exceptions are disabled
(-Wunused-parameter, misc-unused-parameters)
- put "Exception safety" before "Exceptions" in to_msgpack.md, as the
documentation style check requires
- create the test's string value from its type: constructing it from a
beyond_uint32_string_t considers the std::filesystem::path conversion,
which libstdc++ 10 reports as ambiguous for a class derived from
std::string (clang 13)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the MessagePack string length test for clang with libstdc++ 10
C++17 builds consider the std::filesystem::path conversion for the
string type, and with clang and libstdc++ 10 that conversion is
ambiguous for a class derived from std::string. Creating the value from
its type did not avoid it, since any basic_json with that string type
instantiates the check. The binary and ext cases are still tested there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep the MessagePack string test type and its alias in one block
astyle indented the alias oddly when it had an #ifdef of its own after
the binary alias; declare it right after the string type, in the same
block.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare unordered objects by key below the nesting bound
Values nested deeper than the nesting bound are compared without the
call stack, walking both objects entry by entry. Two equal objects of a
type that enumerates its entries in no fixed order - std::unordered_map,
say - can be walked in different orders, so they compared unequal, and
a deep copy compared unequal to its original. std::unordered_map's own
operator== does not depend on the order, which is what applies above the
bound.
Where the keys differ, equality now finds the entry by its key instead.
An ordering, and ordered_map, whose operator== compares its entries in
sequence, still decide by the key.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test unordered object equality without std::unordered_map
basic_json<std::unordered_map> instantiates std::pair<const string,
basic_json> while basic_json is still incomplete. The standard does not
require std::unordered_map to support that, and libstdc++ 6 to 9 as well
as the EDG front ends of icpc and nvc++ reject it, which broke the build
of unit-comparison on those CI jobs.
The test now uses an object type derived from std::map (which, as the
default object type, works everywhere) whose comparator orders keys
ascending or descending as chosen at construction, and whose operator==
does not depend on the order of the entries - the property of
std::unordered_map the test is about.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare the test object type's entries with std::all_of
clang-tidy (readability-use-anyofallof) asked for std::all_of instead of
the loop in unordered_object_t's operator==. The entry type is spelled
out, as C++11 needs typename for base_type::value_type and C++20
reports it as redundant.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add tests for uncovered code paths
Cover code the test suite did not reach, found from the Coveralls report
of develop and a local coverage run of HEAD:
- dump() of every kind of value below the bound of the recursive descent
(pretty-printed objects, binary values, discarded values, scalars), and
flushes of the escape and write buffers mid-string and mid-binary
- the iterative comparison: objects with different keys, containers that
are a prefix of each other, and elements that cannot be ordered, each
both at the top level and below the nesting bound
- SAX handlers that stop at any event, including the end of a nested
container, in the BSON, CBOR, MessagePack, UBJSON and BJData readers
- from_bson/cbor/msgpack/ubjson/bjdata returning a discarded value
through the iterator and pointer overloads
- JSON Patch, diff, merge_patch and update(..., true) on ordered_json
- smaller gaps: get_allocator(), to_ubjson/to_bjdata into a string,
value() with an unresolvable JSON pointer, integer/float comparison
below the integer range and with negative fractions, conversion to a
custom binary type, std::formatter::parse on a spec without '}',
unescape() of a lone '~', and the callback parser's start_array()
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Cover more paths that were thought unreachable
- parse_float_fast() declining malformed or inexact input, called
directly since the lexer only passes well-formed numbers to it
- a UTF-16 high surrogate followed by a unit above the low surrogates
- self-assignment of a const_iterator
- a truncated CBOR string read through non-contiguous iterators
- serializing a long double under the de_DE locale, which undoes the
locale's decimal point and thousands separator
- values read from a binary format carrying no diagnostic positions,
with and without a parser callback
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the CI failures of the new coverage tests
- declare the self-assignment reference const (misc-const-correctness)
- expect the (/path) prefix that JSON_DIAGNOSTICS adds to the messages
of the failing ordered_json patch operations
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Expect the byte range JSON_DIAGNOSTIC_POSITIONS adds to the patch errors
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Build the expected dump of the nested-object test with +=
clang-tidy (performance-inefficient-string-concatenation) reported the
chain of operator+ calls that assembled the expected indented output.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Compare the BJData and UBJSON test outputs byte by byte
Building a std::string from the byte vector converts each byte
implicitly, which -fsanitize=integer reports for bytes of 0x80 and
above (ci_test_clang_sanitizer).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Coverage reported conditions in the binary writer that can never be
false, and marked the code behind them with LCOV_EXCL. Remove them
instead of excluding them:
- CBOR writes the length of a string, binary value, array, or object
exactly like an unsigned integer, only with another major type. One
function, write_cbor_head(), now writes both, so the integer tests
cover every width and the four excluded 64-bit length branches are
gone.
- A last `else if` whose condition holds for every remaining value
(an unsigned value at most UINT64_MAX, a signed one in the range of
int64_t) is now a plain `else`.
- Whether a signed integer fits into an int64 for UBJSON and BJData is
decided by its type at compile time. Only an integer type wider than
64 bits gets a range check and the high-precision fallback.
- The private get_impl(boolean_t*) was never called.
The UBJSON type prefix 'H' of an optimized container of unsigned
integers beyond the range of int64 was reachable although excluded; it
is tested now.
The output is unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Write BSON in linear time, without recursing per nesting level
to_bson() had two problems with nested values:
- It recursed once per nesting level, so a value nested deeply enough -
100,000 levels on an 8 MiB stack - exhausted the call stack and
terminated the process, although parse() accepts such values without
complaint.
- BSON prefixes every document and array with its length. The writer
computed that length by walking the entire value below it, again for
every nested document it wrote, which made serializing O(size x depth).
A 200-level document took 30 ms instead of 1.
Both passes are now iterative, and each length is computed exactly once:
- calc_bson_sizes() computes the length of every document and array in
one pass, each from the lengths of its entries, into a table ordered
the way they are written.
- write_bson_document() then writes the document, taking each length from
the table.
Everything observable is unchanged, as a differential test against
develop confirms byte for byte:
- The same bytes are written.
- A key containing U+0000 still throws out_of_range.409 for the same
first key, with the same diagnostics path, before anything is written.
- A document too large for BSON still throws out_of_range.412 before
anything is written.
- A binary subtype above 255 still throws out_of_range.415 after the
same partial output.
Only the enclosing objects and arrays are kept on a stack, so a flat
document allocates nothing for it. Measured against develop (clang -O3,
median of 201 runs): flat objects unchanged, flat arrays 37% faster (the
array length was computed twice), a nested 3,000-object document 2x
faster, a 200-level document 33x faster.
to_bson.md documented the quadratic complexity since #5334; it is linear
again.
Fixes#5392 for BSON, and #5308.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Do not require a default-constructible string_t in the BSON writer
GCC 4.9 and MSVC rejected the test's huge_string_t, which has no default
constructor; develop never default-constructed string_t here either.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Let the BSON index-name helper only fill its output parameter
It returned a reference to the string it filled, so callers held a second
name for index_name. Addresses review feedback.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make JSON_STRICT_NUL_HANDLING part of the ABI tag
JSON_STRICT_NUL_HANDLING (#5534) changes the bodies of inline functions:
the lexer's handling of '\0' and input_adapter() for char arrays. So
translation units compiled with and without it define the same functions
differently, an ODR violation - the case the ABI tag exists for, as with
JSON_BRACE_INIT_COPY_SEMANTICS (_bics). It now appends _snul to the inline
namespace. The macro is new in 3.13.0, so no existing namespace changes.
Its default moves to abi_macros.hpp, and it is only #undef'd without
JSON_TEST_KEEP_MACROS, as for the other ABI macros. The ABI config tests,
the namespace docs, the macro's docs and the Natvis file cover the new tag.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Allocate the deep copy's key scratch space with the provided allocator
The iterative deep copy builds each object's keys in a temporary vector of
key/value pairs before handing them to the object's range constructor. That
vector holds basic_json values, so like the values themselves it now uses
AllocatorType instead of std::allocator.
Also document that AllocatorType covers the JSON values, while most
temporary storage still uses std::allocator.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Count allocate_at_least in the scratch-counting test allocator
From C++23 on, libc++'s containers allocate through allocate_at_least when
the allocator has one. The test allocator inherited it from std::allocator,
so the scratch allocations were not counted and the test failed on Xcode.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 fix BSON conformance issue
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* 🐛 reject ill-formed UTF-8 in CBOR/MessagePack/BSON text strings at decode time (#5531)
from_cbor()/from_msgpack()/from_bson() copied the raw bytes of a decoded
text string into the resulting json value without any UTF-8 validation,
even though RFC 8949 §3.1 (CBOR) and the MessagePack/BSON specifications
all require text strings to be valid UTF-8. Malformed input only failed
later, if the value was dump()'d, with a type_error.316 - so the
allow_exceptions=false pattern used specifically to get a discarded
sentinel instead of an exception did not discard this category of
malformed input, unlike every other kind of malformed binary input this
library rejects at decode time (see #5529).
Fix this at the single choke point shared by BSON/CBOR/MessagePack/UBJSON
string reads, binary_reader::get_string(): validate the bytes with the
UTF-8 DFA right after they are read, and report failures the same way as
every other binary_reader error (parse_error.113), so allow_exceptions
and strict discarding behave consistently. get_binary()/binary blob reads
are untouched and still accept arbitrary bytes, since only text strings
are required to be UTF-8.
There were two independent implementations of a UTF-8 validator: the
lexer's streaming scanner, and the serializer's Hoehrmann DFA used by
dump_escaped_impl(). Rather than write a third, the serializer's decode()
function, its utf8d table and the UTF8_ACCEPT/UTF8_REJECT constants are
extracted into detail/string_utils.hpp (a low-level header already
included before both detail/input/ and detail/output/), alongside a new
is_valid_utf8() helper built on the same decode() step. serializer.hpp's
dump_escaped_impl() now calls the shared decode(), so there is exactly
one UTF-8 validator in the codebase; dump()'s exact type_error.316
messages and byte-index reporting are unchanged (see the added
regression-guard test in unit-serialization.cpp).
Claude-Session: https://claude.ai/code/session_01N4RQ1Ahan5YAGbnAQGjZTY
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* ⚡ validate only newly read bytes of binary-format strings
get_string() validated the whole result after each call, but get_bytes()
appends to it and CBOR indefinite-length strings collect all chunks in
the same result, so every chunk re-validated everything read before it.
An input of many small chunks took quadratic time (80000 one-byte chunks,
160 KB of input, took about 7 seconds). Only the newly read bytes are
validated now, which also matches RFC 8949's requirement that every
chunk is valid UTF-8 on its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Support zero-member types in NLOHMANN_DEFINE_TYPE_* macros (#4041)
NLOHMANN_DEFINE_TYPE_INTRUSIVE(Type) and its 11 sibling macros produced
broken code for types with no members to serialize. Invoking a variadic
macro so __VA_ARGS__ is empty is only standard-conforming since C++20,
so a plain __VA_OPT__ fix (as tried in #5142) breaks every pre-C++20
build under -pedantic. Instead, make all 12 macros purely variadic and
dispatch on argument count using a sentinel-padded extension of the
existing NLOHMANN_JSON_GET_MACRO idiom, giving full C++11-C++26 support
with no feature-test gate.
Verified against real GCC 16 and Clang at -std=c++11/14/17/20 with
-pedantic -Werror -Wvariadic-macros: zero regressions in the existing
unit-udt_macro.cpp suite plus 12 new zero-member test cases.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI failures in zero-member NLOHMANN_DEFINE_TYPE_* macros
Three issues surfaced on PR #5272's real CI that weren't caught by
local testing against a narrower flag set:
- GCC -Werror=noexcept: the four truly-empty from_json bodies (plain
INTRUSIVE/NON_INTRUSIVE, with and without _WITH_DEFAULT) provably
never throw but weren't declared noexcept; mark them noexcept
explicitly. to_json and the derived-type from_json overloads are
left alone since they genuinely can throw (object assignment /
delegating to the base class's from_json).
- clang-tidy bugprone-macro-parentheses: false positive on the same
8 zero-member bodies (Type/BaseType used purely as declarator
types); suppressed with NOLINTNEXTLINE comments in the same style
already used elsewhere in this file (see NLOHMANN_JSON_SERIALIZE_ENUM).
- MSVC's traditional preprocessor doesn't fully expand
NLOHMANN_JSON_CAT(prefix, NLOHMANN_JSON_TYPE_TAG(...))(...) in one
pass, which broke a pre-existing one-member usage in
unit-regression2.cpp with syntax errors. Wrap all 12 public
dispatcher macros in an extra outer NLOHMANN_JSON_EXPAND(...),
matching the pattern NLOHMANN_JSON_PASTE already uses for the same
MSVC quirk.
Re-verified against real GCC 16 and Clang at -std=c++11/14/17/20 with
-pedantic -Werror -Wvariadic-macros -Wnoexcept, including the exact
files that failed in CI (unit-udt_macro.cpp, unit-regression2.cpp),
against both the modular headers and the re-amalgamated single header.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy misc-const-correctness in unit-udt_macro.cpp
The four zero-member ONLY_SERIALIZE test objects are only ever read
(via to_json), never mutated, so mark them const per clang-tidy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix derived-type macro dispatch capping members at 62 instead of 63
NLOHMANN_JSON_GET_MACRO resolves 64 positional arguments, with NAME at
position 65. NLOHMANN_JSON_TYPE_TAG dispatches on Type plus the member
list, so it resolves correctly up to the 63 members NLOHMANN_JSON_PASTE
supports. NLOHMANN_JSON_DERIVED_TYPE_TAG dispatched on the two-token
Type,BaseType prefix plus the member list, running out one slot early:
at 63 members, position 65 landed on the last member name instead of a
sentinel and NLOHMANN_JSON_CAT built an undefined identifier such as
NLOHMANN_JSON_DEFINE_DERIVED_TYPE_INTRUSIVE_m63, with the compiler
reporting "unknown type name 'm1'" once per member and nothing pointing
at an argument-count limit.
That silently reduced all six NLOHMANN_DEFINE_DERIVED_TYPE_* macros from
63 members to 62, contradicting the "up to 63 members" contract in
docs/mkdocs/docs/api/macros/nlohmann_define_derived_type.md.
Drop the leading Type and defer to NLOHMANN_JSON_TYPE_TAG so the tag is
computed from BaseType plus the member list, which fits the available
slots. The zero-own-member derived bodies are therefore selected by tag
1 rather than 2, and the sentinel table for the derived tag is no longer
needed.
Add a regression test at the documented maximum for both the plain and
the derived macros; it fails to compile against the previous dispatch.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the zero-member macro bodies by intent, not argument count
The dispatch tag was the literal token 1 or N, pasted onto a macro prefix
to select the zero-member or member-carrying body. For the derived-type
macros that reads wrong: their tag is computed after dropping the leading
Type, so the zero-member body was named _1 while taking two parameters
(Type, BaseType).
Emit EMPTY and MEMBERS instead. The mechanism is unchanged -- the tag is
still a token pasted onto the prefix by NLOHMANN_JSON_CAT -- but the body
names now say what they are rather than encoding an argument count that
only lines up for half of the macros.
Collapse the four duplicated zero-member bodies while here: with no
members there is nothing to default, so each _WITH_DEFAULT_EMPTY body was
a byte-for-byte copy of its plain counterpart. They are now one-line
aliases, leaving a single definition of what an empty object serializes
to per intrusive/non-intrusive and base/derived combination.
No functional change: for both zero-member and member-carrying types the
preprocessed to_json/from_json output is token-for-token identical, and
the arity limits are unchanged (63 members, base and derived).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document zero-member support in the macro API reference
docs/mkdocs/docs/features/arbitrary_types.md already gained a note, but
the three api/macros pages are where the parameter contract is actually
specified and they still described member as a non-empty list.
State that the list may be empty on each page, and add a note showing
what the zero-member case generates: an empty JSON object for the plain
macros, and base-type-only serialization for the derived ones. Both notes
record that the WITH_NAMES variants do not support this.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep user macros named EMPTY or MEMBERS out of the member-count dispatch
The dispatch produced the bare token EMPTY or MEMBERS and pasted it onto
the macro prefix afterwards. In between, the token was rescanned, so a
user macro with either name replaced it: with `#define MEMBERS x` in
scope, even NLOHMANN_DEFINE_TYPE_INTRUSIVE(A, member) -- which compiled
before -- expanded to garbage, and `#define EMPTY` broke the zero-member
form.
Paste the suffix onto the prefix directly in the GET_MACRO slot table
instead. Operands of ## are not macro-expanded, so the selected body name
is formed before any user macro can interfere. NLOHMANN_JSON_TYPE_TAG and
NLOHMANN_JSON_DERIVED_TYPE_TAG become NLOHMANN_JSON_TYPE_BODY and
NLOHMANN_JSON_DERIVED_TYPE_BODY, taking the prefix as their first
argument; NLOHMANN_JSON_CAT is no longer needed. The body macro names are
unchanged, and so is the generated code.
Add a regression test that defines EMPTY and MEMBERS around plain and
derived types, with and without members.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test for EMPTY and MEMBERS so -Wunused-macros accepts them
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document the benchmarks, and make them build and compare versions again
The benchmark project hasn't configured since #4793: download_test_data.cmake
compiles cmake/detect_libcpp_version.cpp relative to CMAKE_SOURCE_DIR, which
is tests/benchmarks when that is the top-level project, so try_run fails and
so does `make run_benchmarks`. The path is now relative to the module itself,
which is the same file for the main build.
The Dump benchmark discarded dump()'s result, which is [[nodiscard]] by now;
it warned, and left the optimizer free to shorten the loop. The result is
now kept with benchmark::DoNotOptimize.
A new cache variable, JSON_BENCHMARK_INCLUDE_DIR, names the directory holding
the nlohmann/json.hpp to benchmark (single_include by default, as before),
so the same benchmarks can be built against two versions and compared.
tests/benchmarks/README.md documents what is measured, how to build and run
the benchmarks, how to read the output, and how to compare two versions with
Google Benchmark's compare.py; it recommends doing so by hand before a
release rather than in CI.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Point ci_benchmarks at tests/benchmarks
The target has configured ${PROJECT_SOURCE_DIR}/benchmarks since it was
added in #2561, but the benchmarks live in tests/benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Pin Google Benchmark to release 1.9.5
The benchmarks fetched Google Benchmark's main branch, so two builds on
different days could measure with different library code, and CMake 3.30
and later warn that the single-argument FetchContent_Populate() is
deprecated. Fetch the 1.9.5 release archive, verified by its SHA-256,
with FetchContent_MakeAvailable() instead. That needs CMake 3.14; Google
Benchmark itself already needed 3.13.
Its -Werror is switched off, so a newer compiler's new warnings cannot
break the pinned release, and its install rules are no longer added.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Ubuntu, Windows, macOS, and CodeQL already cancel an older run of the same
workflow on the same ref. Check amalgamation, CIFuzz, Dependency Review,
Flawfinder, Semgrep, Scorecard, and the labeler did not, so every push
to a pull request left their earlier runs going. Give them the same
concurrency group. The labeler runs on pull_request_target, where
github.ref is the base branch, so it groups by pull request number.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Answer the OpenSSF Best Practices criteria that asked for policies the
project follows but had not written down:
- SECURITY.md: a first response within 14 days, publishing an advisory
with credit once a fix is released, and that only the latest release
receives security fixes.
- Governance: who has access to the project's resources, how write or
admin access is granted, and how CI secrets are stored and rotated.
- Quality assurance: how dependencies of the build, test, and
documentation tooling are pinned, scanned, and kept free of known
vulnerabilities.
Also update the assurance case, since comparison no longer recurses per
nesting level (#5390), and point the best practices badge and links to
bestpractices.dev under the program's current name.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5344 added two lines to unit-deserialization.cpp that clang-tidy
reports: modernize-return-braced-init-list for the remaining() helper
and readability-isolate-declaration for "json j1, j2, j3;".
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Improve error message for const fields
* Reject const arguments to get_to() with a clear message
Reword the static_assert, add it to the C array overload of get_to() as well,
and document that v must not be const.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Niels Lohmann <mail@nlohmann.me>
* Compare values without recursing, and without comparing them twice
Comparing two values compared their containers, which compare their elements,
which brought the comparison back once per nesting level. Two values nested
deeply enough exhausted the call stack and terminated the process with a
segmentation fault - the same bug as #5387, in the last operation that still
had it.
Worse, an ordered comparison took exponentially long in the nesting depth
before C++20. std::vector's operator< is a lexicographical comparison, which
asks whether an element is less than its counterpart and then whether the
counterpart is less than it - two full comparisons of everything below that
element, at every level. Comparing two equal values nested 30 levels deep,
which is nothing unusual, took 3.8 seconds; 40 levels would have taken an
hour, and nothing about the value has to be pathological to get there. C++20
is unaffected: std::lexicographical_compare_three_way asks once.
Compare a value that is nested too deeply to descend into on an explicit
stack instead, in a single pass that yields less, equal, greater or unordered
at once. Equality and the three-way comparison descend as they always did for
the first 128 levels, which nothing measurable costs them; an ordered
comparison no longer descends at all, which is what takes the exponent out of
it. Objects and arrays that are not nested deeply are otherwise compared
exactly as before.
The results are unchanged for every pair of values: 68121 comparisons of a
corpus that covers NaN, discarded values, mixed number types, binary values,
empty containers and both object types are identical to develop, in C++11,
C++17 and C++20, with and without thread_local storage and legacy discarded
comparison. Reproducing that meant reproducing two subtleties: a lexicographic
comparison steps over a pair it cannot order, where a three-way comparison
stops at it, and an object compares its keys with < where its entries are
ordered but with == where they are only checked for equality - not with the
object's own comparator, which for nlohmann::ordered_map tells equality.
Equality needs no ordering, so it no longer asks for any: a key or string type
that can only be compared for equality still works.
Measured (medians of 7 interleaved runs, clang -O3, C++11): comparing two
equal values nested 30 levels deep 3778 ms -> 0.002 ms; ordering flat objects
-33.6%; ordering flat arrays of numbers +27.3%, the one shape that pays for
the single pass; equality unchanged throughout.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Describe comparison in the no-thread-local docs and CI target
Comparing two values now bounds its descent with a thread_local counter
just as copying does, so the JSON_NO_THREAD_LOCAL page, the macro
overview and the ci_test_no_thread_local target cover both rather than
copying alone.
Also record what switching the macro on costs a comparison: on the
benchmark documents, comparing two equal values takes 10% to 90% longer.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Take the descent flag as an argument rather than testing it
MSVC reports the test of a constant as C4127 ("conditional expression is
constant"), which the Windows builds treat as an error: may_descend is
false for operator<, so the operand short-circuits the whole condition.
Passing it to compare_descent_exhausted() puts the test where the value
is an ordinary parameter, and leaves the call sites with no condition of
their own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Note the comparison fallback in the no-thread-local documentation
The macro page describes what the library defines JSON_NO_THREAD_LOCAL for
by itself in terms of copying alone; comparing falls back the same way.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Parenthesise the reserve() computation in the comparison test
clang-tidy reports the mixed * and + as readability-math-missing-
parentheses, as it does for the identical line in the copy test.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the shared descent bookkeeping rather than a second set
Comparing kept a thread_local count, a limit and a guard of its own beside
the ones copying already had, all three the same thing under a different
name. They are gone; the shared count, limit and guard do the work.
The guard grows a second constructor here, because the comparison
operators are written as a macro and a macro cannot use the preprocessor:
it cannot look the count up behind an #ifdef the way copy_structured does,
so the guard looks it up for it. nesting_depth_exhausted() arrives for the
same reason - whether an operator descends at all is a constant at every
call site, and testing it there is what MSVC reports as C4127.
Also say in compare_leaves what happens to a pair that is an array on one
side and an object on the other, since the answer is not obvious from the
code: an operator only descends into two values of the same type, so such
a pair is told apart by its types alone - unequal, and ordered the way the
types are - exactly as it is above the bound.
And record what the explicit stack costs: the comparison operators are
noexcept and the container comparison this replaces allocated nothing, so
running out of memory here ends the process instead of throwing. It takes
a value nested past the bound and an exhausted heap to reach, and the same
comparison used to exhaust the call stack, but it is a new way to fail.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* docs: qualify the operator>> stream positioning guarantee
operator>>'s notes state that it leaves the stream positioned right
after the parsed value, so that concatenated JSON values can be read
back to back. That does not hold when the value is a number: a number
is only terminated by the character that follows it, and the lexer's
unget() is simulated (it rewinds only the lexer's own bookkeeping),
so that character stays consumed from the stream.
Document the actual behaviour: the guarantee holds for all value types
except numbers, which must be followed by whitespace. Also qualify the
cross-reference on the JSON Lines page, which repeated the unqualified
claim.
Documentation only; the behaviour itself is tracked in #5340.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* fix: restore the character that terminates a number (#5340)
operator>> is documented to leave the stream positioned right after the
parsed value, so that concatenated JSON values can be read back to back.
That did not hold for numbers: a number is only terminated by the
character following it, and lexer::scan_number() reads that character
and calls unget() -- which is simulated and rewinds only the lexer's own
bookkeeping. input_stream_adapter consumes via sbumpc() with no matching
sungetc(), so the terminating character stayed consumed and the next
extraction started one byte too late ('1true' left the stream at 'rue').
Propagating unget() to the adapter directly does not work: next_unget
makes the following get() replay the cached character, so the terminator
would be delivered twice. Instead, restore the still-pending character
once at the end of a non-strict parse, where the input is handed back to
the caller:
- input_stream_adapter gains unget_character() (sungetc()) and advertises
it via supports_unget, detected the same way as supports_seek.
- lexer::restore_pending_unget() turns a pending simulated unget of a
real (non-EOF) character into a real one and clears next_unget so the
character is not also replayed. It is a no-op for adapters that cannot
unget, and reports failure when sungetc() fails, in which case the
input is left as it was before.
- parser calls it on the three non-strict paths, i.e. for operator>> and
sax_parse(strict = false).
Strict parse()/accept() are unaffected: they require the input to end
after the value, so the character is consumed by the end-of-input check
anyway. Parse error messages and reported positions are unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* tests: fix CI failures in the #5340 test helpers
Four CI failures, all in the new test code:
- GCC (-Werror=useless-cast): drop the `json(...)` wrapper around
`json::parse(...)`, which already returns a `json`.
- GCC (-Werror=unused-result): assign the discarded `json::parse()`
result to a dummy, the idiom used elsewhere in the test suite, and
catch `json::parse_error&` for consistency.
- clang-tidy (google-default-arguments): remove the default argument
from the `pbackfail()` override; `sungetc()` supplies the base
declaration's default.
- MSVC (bad allocation): `no_putback_streambuf::underflow()` set a
one-character get area without advancing `m_pos`, so an implementation
whose `istream::get` peeks before it bumps re-read the same character
forever. Keep no get area at all: `underflow()` peeks, `uflow()`
consumes, and `sungetc()` still always lands in `pbackfail()`, which
is what the test needs.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* fix: leave the character that terminates a number in the input
Read the character following a number without consuming it, instead of
consuming it and putting it back. input_stream_adapter now peeks with
sgetc() and only steps over the character when the next one is requested
or when the adapter is destroyed, so releasing it cannot fail - no
putback position is required from the streambuf.
Suggested by gregmarr in #5344.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* docs: match the version history wording to the peek-based fix
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* docs: drop the whitespace-separator caveat from the parsing pages
The caveat added in #5343 describes the behavior this branch fixes: a
number no longer consumes the character that terminates it, so
concatenated values need no separator.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* refactor: split the strict and non-strict paths in parser
Folding the release_lookahead() call into the existing strict check left
the "in strict mode" comment on an else-if branch, and made the strict
condition in sax_parse() redundant with the branch it followed.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Put the stream position fix behind JSON_PRECISE_STREAM_POSITION
Leaving the character that terminates a number in the stream is observable:
reading "1,2,3" with repeated operator>> works today only because the comma
after each number is swallowed, and std::getline after a number skips the
line break. Both break with the fix, so make it opt-in for 3.x, as suggested
by @gregmarr in the review.
- JSON_PRECISE_STREAM_POSITION (default 0) selects the peek-based
input_stream_adapter. Without it, the adapter is the consuming one from
develop and has no supports_lookahead, so lexer::release_lookahead() and
the parser's calls to it compile to nothing.
- The macro changes input_stream_adapter's layout and member functions, so
it gets the ABI tag _psp, after _bics. The ABI config tests, the natvis
generator, and nlohmann_json.natvis (regenerated) know the tag.
- The tests for the fix move to unit-precise-stream-position.cpp, which
defines the macro itself and runs in every build, and gain the two cases
above. unit-deserialization.cpp pins the default behavior instead.
- The docs describe the default behavior again and point to the new macro
page; version history says "added in 3.13.0, planned default in 4.0.0".
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The windows-11-arm runner now ships MSVC 19.51, which reports doctest's
forward declaration of std::tuple as C5285 ("cannot declare a
specialization for 'std::tuple'"). With /WX this breaks the msvc-arm64
job on develop and on every open pull request. Disable the warning for
the test targets, like the other MSVC warnings already disabled there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This should have no effect for header only libraries as mentioned.
It was previously removed in e509007d but then accidentally added
again in 26cfec34.
Signed-off-by: Jeremy Nimmer <jeremy.nimmer@tri.global>
* Keep JSON_DIAGNOSTICS parent pointers of ordered_json members after erase() and update()
ordered_json stores its members in a vector, and two operations moved
members without restoring their parent pointers afterwards:
- ordered_map::erase() re-constructs every member after the erased one in
place. The basic_json move constructor leaves m_parent at nullptr, and
none of the object branches of basic_json::erase() (by key, iterator, or
iterator range) called set_parents(). This also affected merge_patch()
with a null member and patch() with a remove operation.
- update() only set the parent pointer of the inserted member. Adding a key
can reallocate the vector, which copies all other members and leaves
their m_parent at nullptr. The set_parents() call added for #4813 only
repaired this for the nested object of a merge, not for the target.
The next assert_invariant() on such an object (for instance, when copying
it) aborted, and diagnostic messages lost the path prefix above the moved
member. std::map-based json was not affected, because its nodes do not
move.
Erasing from an ordered_map object now calls set_parents(), and update()
uses set_parent(), which already refreshes all members for vector-based
objects. This makes the #4813 workaround redundant.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Account for JSON_DIAGNOSTIC_POSITIONS in the ordered_json parent-pointer test
The merge_patch() case parses its input, so with JSON_DIAGNOSTIC_POSITIONS
the exception message also carries the byte range of the parsed value.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Silence clang-tidy for the intentional copy in the ordered_json parent-pointer test
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep parent pointers when update() merges past its descent bound
The iterative path of update() only set the parent pointer of the member
it inserted, like the recursive one did before. It now uses set_parent()
too, so ordered_json members that move when a nested object grows keep
their parents, and the set_parents() calls that patched this up after
each nested merge are gone.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests
The strongest correctness checks for the UBJSON and BJData writers lived
only in the OSS-Fuzz drivers: anything from_ubjson()/from_bjdata()
returns must serialize with every option combination, parse back, and
re-serialize stably. Those checks only run at OSS-Fuzz, so regressions
surfaced days later as external reports - the same BJData assert pair
was reported five times over three years, and #5494's harness change
was followed by OSS-Fuzz 563659413 within a day.
Add "UBJSON round-trip invariants" and "BJData round-trip invariants"
test cases that run the drivers' checks on a fixed, deterministic corpus
(tests/src/round_trip_corpus.hpp): integer and float boundaries,
non-finite numbers, strings, binary values, optimized containers, deep
nesting, the JData annotated-array matrix, and seeded random containers.
They also check two properties the drivers do not: the first round trip
preserves the value, and re-serializing reproduces the exact bytes. For
BJData both exclude values containing a binary value, which is read back
as an array of integers unless it was written as a Draft 3 optimized
binary array; this carve-out is now documented in bjdata.md. Run against
the headers before #5542, the BJData test fails, including on the shape
from OSS-Fuzz 563659413.
Also document how OSS-Fuzz reports are handled (reference them as
"OSS-Fuzz: <id>", turn the reproducer into a unit test, keep drivers and
unit tests in sync) in tests/fuzzing.md, and link it from the PR
template and the quality assurance page.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add the OSS-Fuzz reproducers for 474400817 and 474480402 as unit tests
Following the convention added to tests/fuzzing.md, the reproducers of
the two BJData fuzzer asserts tracked since January are now unit tests:
- 474400817 (assert(false)): an empty object _ArraySize_ was written as
the ND-array header length, which from_bjdata() could not read back.
Fixed by #5455.
- 474480402 (to_bjdata(j2, false, false) == vec2): a one-byte Draft 3
binary array is written in Draft 2 mode as a uint8 array and then
re-serialized with the int8 marker. This is the documented exception to
byte stability, not a library bug; OSS-Fuzz closed it after #5494
relaxed the harness to value stability. The test pins the exact bytes
so the exception stays deliberate.
The 563659413 reproducer is already a unit test (#5542). A comment also
ties the existing UBJSON excessive-count test to the timeout OSS-Fuzz
reported for that shape (testcase 6347769435193344).
OSS-Fuzz: 474400817
OSS-Fuzz: 474480402
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix GCC -Weffc++ and -Wuseless-cast warnings in the round-trip corpus
Initialize the atoms in the member initialization list, and drop the cast of
the generator's result, which already is std::size_t on 64-bit Linux.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Describe the threat model, the trust boundaries, the secure-design
argument, and how common weaknesses are countered, with links to the
quality assurance page as evidence.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Describe what the project will and will not do over the next year,
and point to issue #3453 for the open question of a 4.0 release.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Complete the architecture documentation page
Replace the placeholder bullets and TODOs with a description of the
component pipeline (with a diagram), the source layout, the template
parameters, the value storage (now in struct data), the input and
output adapters, and the SAX interface.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Link sources and basic_json, document full input adapter interface
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Align the default column of the template parameter table
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The windows-11-arm runner image moved to windows-11-vs2026-arm64, which no
longer ships Visual Studio 2022, so the msvc-arm64 job failed at configure
time. Use the "Visual Studio 18 2026" generator like the msvc2026 job.
clang-tidy's modernize-raw-string-literal check flagged two string literals
in the nesting tests added by #5546 and #5547.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Merge deeply nested objects without recursing per nesting level
merge_patch() and update(j, true) merged a nested object by calling
themselves on it, once per nesting level. A value nested deeply enough -
50,000 levels of objects on an 8 MiB stack - exhausted the call stack
and terminated the process, although parse() accepts such values without
complaint.
Bound the descent the same way dump() does. The recursion now carries
the nesting level, and once merge_depth_limit() (128) levels have been
entered, update_members_iteratively() and merge_patch_iteratively()
finish the merge on an explicit stack. They still merge a nested object
completely before the next member, and in the same order, so the results,
including the parents JSON_DIAGNOSTICS reports paths from, are unchanged.
Values nested less deeply than the bound run the same code as before, so
the common case does not pay for the stack: merging only on it cost
10-14% in a first version.
The public signatures are unchanged. The recursive worker behind
merge_patch() has its own name rather than being a private overload, so
that &basic_json::merge_patch stays unambiguous.
Tests check every depth up to 300 against recursive reference
implementations of both operations, check the diagnostic paths past the
bound, and merge objects nested 100,000 levels deep.
Fixes#5545 for update(j, true), and #5393 for merge_patch().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the shared recursion limit in update() and merge_patch()
merge_depth_limit() is gone in favor of detail::recursion_depth_limit().
The two identical function-local frame structs become one member struct,
merge_frame, with a constructor, so both loops emplace_back() their
frames. merge_patch_iteratively() copies the frame it works on out of the
stack and changes it only through stack.back().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Build the update()/merge_patch() diagnostics test values instead of parsing them
Parsed values carry byte positions under JSON_DIAGNOSTIC_POSITIONS, which
the expected messages do not include.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Hash deeply nested values without recursing per nesting level
std::hash<basic_json> hashed an array or object by hashing each element,
which called detail::hash again once per nesting level. A value nested
deeply enough - 50,000 levels of objects on an 8 MiB stack - exhausted
the call stack and terminated the process. parse() accepts such values
without complaint, since the parser is iterative, and a parsed value is
hashed wherever it is used as a key in an unordered container.
Bound the descent the same way dump() does: detail::hash takes the
nesting level, and once hash_depth_limit() (128) levels have been entered,
hash_iteratively() hashes what is left on an explicit stack. It combines
the seeds in exactly the same order, so hash values are unchanged. A value
nested less deeply than the bound is hashed by the same code as before,
without allocating, and is as fast as before.
Tests check that every depth up to twice the bound hashes exactly like
the recursive definition of the hash, and that values nested 100,000
levels deep hash without crashing.
Fixes#5545 for std::hash.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Declare hash_frame's constructor noexcept
GCC's -Wnoexcept (an error in CI) flags the emplace_back() into the
hash stack under C++26: the constructor cannot throw, since cbegin() is
noexcept, but it did not say so. dump_frame's constructor is noexcept
for the same reason.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Share one recursion depth limit, and copy the hash frame out of the stack
dump() and hash() each defined their own limit on how many nesting levels
they recurse into, and the operations still to come would have added more,
free to diverge over time. They now all use detail::recursion_depth_limit(),
in a header of its own; serializer::dump_depth_limit() and
hash_depth_limit() are gone.
hash_iteratively() now copies the frame it works on out of the stack and
changes the frame only through stack.back(), so nothing can refer into
the stack after entering an element has grown it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Parenthesize multiplications in the hash test for clang-tidy
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-24 17:12:02 +02:00
186 changed files with 21073 additions and 3456 deletions
- [ ] The changes are described in detail, both the what and why.
- [ ] If applicable, an [existing issue](https://github.com/nlohmann/json/issues) is referenced.
- [ ] If applicable, a fixed [OSS-Fuzz](https://issues.oss-fuzz.com) issue is referenced as `OSS-Fuzz: <id>` (see [fuzz testing](https://github.com/nlohmann/json/blob/develop/tests/fuzzing.md#handling-oss-fuzz-reports)).
- [ ] The [Code coverage](https://coveralls.io/github/nlohmann/json) remained at 100%. A test case for every new line of code.
- [ ] If applicable, the [documentation](https://json.nlohmann.me) is updated.
- [ ] The source code is amalgamated by running `make amalgamate`.
@diff BUILD.bazel BUILD.bazel~ ||(echo"===================================================================\n BUILD.bazel is out of date! Please run 'make BUILD.bazel'.\n==================================================================="; mv BUILD.bazel~ BUILD.bazel ;false)
@@ -213,7 +231,7 @@ json.tar.xz:
# We use `-X` to make the resulting ZIP file reproducible, see
[](https://isitmaintained.com/project/nlohmann/json "Average time to resolve an issue")
[](https://bestpractices.coreinfrastructure.org/projects/289)
[](https://www.bestpractices.dev/projects/289)
- [Binary formats (BSON, CBOR, MessagePack, UBJSON, and BJData)](#binary-formats-bson-cbor-messagepack-ubjson-and-bjdata)
- [Binary formats (BSON, CBOR, MessagePack, UBJSON, BJData, and BON8)](#binary-formats-bson-cbor-messagepack-ubjson-bjdata-and-bon8)
- [Customers](#customers)
- [Ecosystem](#ecosystem)
- [Supported compilers](#supported-compilers)
@@ -63,7 +63,7 @@ There are myriads of [JSON](https://json.org) libraries out there, and each may
- **Trivial integration**. Our whole code consists of a single header file [`json.hpp`](https://github.com/nlohmann/json/blob/develop/single_include/nlohmann/json.hpp). That's it. No library, no subproject, no dependencies, no complex build system. The class is written in vanilla C++11. All in all, everything should require no adjustment of your compiler flags or project settings. The library is also included in all popular [package managers](https://json.nlohmann.me/integration/package_managers/).
- **Serious testing**. Our code is heavily [unit-tested](https://github.com/nlohmann/json/tree/develop/tests/src) and covers [100%](https://coveralls.io/r/nlohmann/json) of the code, including all exceptional behavior. Furthermore, we checked with [Valgrind](https://valgrind.org) and the [Clang Sanitizers](https://clang.llvm.org/docs/index.html) that there are no memory leaks. [Google OSS-Fuzz](https://github.com/google/oss-fuzz/tree/master/projects/json) additionally runs fuzz tests against all parsers 24/7, effectively executing billions of tests so far. To maintain high quality, the project is following the [Core Infrastructure Initiative (CII) best practices](https://bestpractices.coreinfrastructure.org/projects/289). See the [quality assurance](https://json.nlohmann.me/community/quality_assurance) overview documentation.
- **Serious testing**. Our code is heavily [unit-tested](https://github.com/nlohmann/json/tree/develop/tests/src) and covers [100%](https://coveralls.io/r/nlohmann/json) of the code, including all exceptional behavior. Furthermore, we checked with [Valgrind](https://valgrind.org) and the [Clang Sanitizers](https://clang.llvm.org/docs/index.html) that there are no memory leaks. [Google OSS-Fuzz](https://github.com/google/oss-fuzz/tree/master/projects/json) additionally runs fuzz tests against all parsers 24/7, effectively executing billions of tests so far. To maintain high quality, the project is following the [OpenSSF Best Practices](https://www.bestpractices.dev/projects/289). See the [quality assurance](https://json.nlohmann.me/community/quality_assurance) overview documentation.
Other aspects were not so important to us:
@@ -128,7 +128,7 @@ There is also a [**docset**](https://github.com/Kapeli/Dash-User-Contributions/t
- When using `get<ENUM_TYPE>()`, undefined JSON values will default to the first pair specified in your map. Select this default pair carefully. If you desire an exception in this circumstance use `NLOHMANN_JSON_SERIALIZE_ENUM_STRICT()` which behaves identically except for throwing an exception on unrecognized values.
- If an enum or JSON value is specified more than once in your map, the first matching occurrence from the top of the map will be returned when converting to or from JSON.
### Binary formats (BSON, CBOR, MessagePack, UBJSON, and BJData)
### Binary formats (BSON, CBOR, MessagePack, UBJSON, BJData, and BON8)
Though JSON is a ubiquitous data format, it is not a very compact format suitable for data exchange, for instance over a network. Hence, the library supports [BSON](https://bsonspec.org) (Binary JSON), [CBOR](https://cbor.io) (Concise Binary Object Representation), [MessagePack](https://msgpack.org), [UBJSON](https://ubjson.org) (Universal Binary JSON Specification) and [BJData](https://neurojson.org/bjdata) (Binary JData) to efficiently encode JSON values to byte vectors and to decode such vectors.
Though JSON is a ubiquitous data format, it is not a very compact format suitable for data exchange, for instance over a network. Hence, the library supports [BSON](https://bsonspec.org) (Binary JSON), [CBOR](https://cbor.io) (Concise Binary Object Representation), [MessagePack](https://msgpack.org), [UBJSON](https://ubjson.org) (Universal Binary JSON Specification), [BJData](https://neurojson.org/bjdata) (Binary JData), and [BON8](https://github.com/hikoworks/hikogui/blob/main/docs/BON8.md) (Binary Object Notation 8) to efficiently encode JSON values to byte vectors and to decode such vectors.
The library also supports binary types from BSON, CBOR (byte strings), and MessagePack (bin, ext, fixext). They are stored by default as `std::vector<std::uint8_t>` to be processed outside the library.
@@ -1387,6 +1395,7 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I
- The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
- The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0).
| Out-of-bounds read/write ([CWE-125](https://cwe.mitre.org/data/definitions/125.html), [CWE-787](https://cwe.mitre.org/data/definitions/787.html)) | bounds checks on all reads from the input; AddressSanitizer and Valgrind on the test suite; OSS-Fuzz |
| Use after free, double free ([CWE-416](https://cwe.mitre.org/data/definitions/416.html), [CWE-415](https://cwe.mitre.org/data/definitions/415.html)) | ownership of all memory by values; AddressSanitizer and Valgrind; Clang Static Analyzer |
| Memory leaks ([CWE-401](https://cwe.mitre.org/data/definitions/401.html)) | Valgrind (Memcheck) on the test suite |
| Uncontrolled recursion ([CWE-674](https://cwe.mitre.org/data/definitions/674.html)) | iterative parser, binary readers, and destructor; bounded recursion in value operations; tests with deeply nested inputs |
| Uncontrolled resource consumption ([CWE-400](https://cwe.mitre.org/data/definitions/400.html)) | allocations based on announced sizes are capped; OSS-Fuzz with memory limits |
| Undefined behavior in general ([CWE-758](https://cwe.mitre.org/data/definitions/758.html)) | UndefinedBehaviorSanitizer; runtime assertions; Clang-Tidy, Cppcheck, Clang Static Analyzer, Infer |
In addition, every line of the library is covered by the unit tests, and all parsers are fuzz-tested around the clock
by [OSS-Fuzz](https://github.com/google/oss-fuzz/tree/master/projects/json).
@@ -174,11 +174,34 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Object keys"
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps with keys other than UTF-8 strings are rejected.
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps
with keys other than text strings (major type 3) are rejected with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set
to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. This is a deliberate restriction of the library's JSON value model, not an oversight: formats built on CBOR
maps with integer keys, such as COSE ([RFC 9052](https://www.rfc-editor.org/rfc/rfc9052.html)) or CWT
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "UTF-8 validation of text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
text.
!!! warning "Tagged items"
Tagged items (0xC0..0xDB) will throw a parse error by default. They can be ignored by passing `cbor_tag_handler_t::ignore` to function `from_cbor`, in which case the tag is skipped and the enclosed data item is parsed on its own. They can be stored by passing `cbor_tag_handler_t::store` to function `from_cbor`. Note that no tag is ever interpreted: for instance, a text string tagged with tag 0 (date/time) stays a string.
Tagged items (0xC0..0xDB) will throw a parse error by default. They can be ignored by passing `cbor_tag_handler_t::ignore` to function `from_cbor`, in which case the tag is skipped and the enclosed data item is parsed on its own. Passing `cbor_tag_handler_t::store` to function `from_cbor` stores tagged byte strings (for bytes 0xd8..0xdb) as binary values with the tag as subtype; other tagged values are read as if the tag were ignored. If several tags precede a byte string, only the innermost one is stored. Note that no tag is ever interpreted: for instance, a text string tagged with tag 0 (date/time) stays a string.
Serializing such a value throws [`out_of_range.412`](../../home/exceptions.md#jsonexceptionout_of_range412).
!!! info "NaN/infinity handling"
`NaN`, `Infinity`, and `-Infinity` are serialized as a MessagePack float 32 (type 0xCA, 5 bytes total),
@@ -136,6 +138,29 @@ The library maps MessagePack types to JSON value types as follows:
Any MessagePack output created by `to_msgpack` can be successfully parsed by `from_msgpack`.
!!! warning "Object keys"
MessagePack allows map keys of any type, whereas JSON only allows strings as keys in object values. Like the
JSON-compatible [profile](https://github.com/msgpack/msgpack/blob/master/spec.md#profile) sketched in the
MessagePack specification, this library restricts map keys to `str` values. Maps with keys of any other type are
rejected with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
This library validates the bytes of every such string (object keys included) at decode time and rejects
ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
| [`diff`](../../api/basic_json/diff.md), [`items`](../../api/basic_json/items.md), [`std::hash`](../../api/basic_json/std_hash.md) | conversion of a `#!cpp std::size_t` to `StringType`: either assignability from the result of `#!cpp std::to_string`, or an ADL overload `#!cpp void int_to_string(StringType&, std::size_t)` |
| [`operator/(std::size_t)`](../../api/json_pointer/operator_slash.md) | the same conversion of a `#!cpp std::size_t` to `StringType` as `diff`, `items`, and `std::hash` above |
| [`std::hash<basic_json>`](../../api/basic_json/std_hash.md) | additionally a specialization of `#!cpp std::hash<StringType>` |
| [`to_bson`](../../api/basic_json/to_bson.md) | `find(value_type)` and `npos` |
| [`parse`](../../api/basic_json/parse.md) from a `string_t` | the input adapters must accept it; otherwise pass a character range |
@@ -547,7 +548,7 @@ Grisu2 algorithm, which produces the shortest representation that round-trips. O
### Required for the binary formats
`NumberFloatType` must be `#!cpp float` or `#!cpp double`. The writers for
[CBOR, MessagePack, UBJSON, BJData, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
[CBOR, MessagePack, UBJSON, BJData, BON8, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
binary32 or binary64 field and have no encoding for `#!cpp long double`.
### Compatible types
@@ -562,7 +563,11 @@ binary32 or binary64 field and have no encoding for `#!cpp long double`.
## `AllocatorType`
`AllocatorType` is instantiated with **one** argument, for each of `object_t`, `array_t`, `string_t`, `binary_t`,
[`byte_container_with_subtype`](../api/byte_container_with_subtype/index.md), and
[`ordered_map`](../api/ordered_map.md).
Everything else lives in [`detail/`](https://github.com/nlohmann/json/tree/develop/include/nlohmann/detail) and namespace `nlohmann::detail`, which is not part of the public API. Paths
below are relative to `include/nlohmann`.
| Component | Location |
|-----------|----------|
| Value type enumeration | [`detail/value_t.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/value_t.hpp) |
| SAX interface and DOM builders | [`detail/input/json_sax.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/json_sax.hpp) |
| Binary format readers | [`detail/input/binary_reader.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/binary_reader.hpp) |
@@ -6,7 +6,7 @@ There are myriads of [JSON](https://json.org) libraries out there, and each may
- **Trivial integration**. Our whole code consists of a single header file [`json.hpp`](https://github.com/nlohmann/json/blob/develop/single_include/nlohmann/json.hpp). That's it. No library, no subproject, no dependencies, no complex build system. The class is written in vanilla C++11. All in all, everything should require no adjustment of your compiler flags or project settings.
- **Serious testing**. Our class is heavily [unit-tested](https://github.com/nlohmann/json/tree/develop/tests/src) and covers [100%](https://coveralls.io/r/nlohmann/json) of the code, including all exceptional behavior. Furthermore, we checked with [Valgrind](http://valgrind.org) and the [Clang Sanitizers](https://clang.llvm.org/docs/index.html) that there are no memory leaks. [Google OSS-Fuzz](https://github.com/google/oss-fuzz/tree/master/projects/json) additionally runs fuzz tests against all parsers 24/7, effectively executing billions of tests so far. To maintain high quality, the project is following the [Core Infrastructure Initiative (CII) best practices](https://bestpractices.coreinfrastructure.org/projects/289).
- **Serious testing**. Our class is heavily [unit-tested](https://github.com/nlohmann/json/tree/develop/tests/src) and covers [100%](https://coveralls.io/r/nlohmann/json) of the code, including all exceptional behavior. Furthermore, we checked with [Valgrind](http://valgrind.org) and the [Clang Sanitizers](https://clang.llvm.org/docs/index.html) that there are no memory leaks. [Google OSS-Fuzz](https://github.com/google/oss-fuzz/tree/master/projects/json) additionally runs fuzz tests against all parsers 24/7, effectively executing billions of tests so far. To maintain high quality, the project is following the [OpenSSF Best Practices](https://www.bestpractices.dev/projects/289).
@@ -340,15 +340,23 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
### json.exception.parse_error.113
A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
string was read where one was required (for instance as a map key), or the string's length specification is invalid.
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
[CBOR](../features/binary_formats/cbor.md) and [MessagePack](../features/binary_formats/messagepack.md).
!!! failure "Example messages"
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82
@@ -356,6 +364,9 @@ string was read where one was required (for instance as a map key), or the strin
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative
```
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte
```
### json.exception.parse_error.114
@@ -928,19 +939,25 @@ A JSON Patch `add` operation cannot be applied because the target location's par
### json.exception.out_of_range.412
BSON stores the length of documents, arrays, strings, and binary values in a signed 32-bit integer. This exception is thrown when a value is too large to be described by such a length field.
BSON stores the length of documents, arrays, strings, and binary values in a signed 32-bit integer, and MessagePack
stores the length of strings, binary values, arrays, and objects in at most an unsigned 32-bit integer. This exception
is thrown when a value is too large to be described by such a length field.
!!! failure "Example message"
!!! failure "Example messages"
```
BSON length 2147483661 exceeds maximum of 2147483647
```
```
MessagePack length 4294967296 exceeds maximum of 4294967295
```
!!! note
This exception was added in version 3.13.0. Before that, the length was silently truncated, and
This exception was added in version 3.13.0. Before that, the BSON length was silently truncated, and
[`to_bson`](../api/basic_json/to_bson.md) produced documents with negative length prefixes that
The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
JSON_THROW(parse_error::create(101,0,"attempting to parse an empty input; check that your input string or stream contains the expected JSON",nullptr));
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.