Compare commits

..
Author SHA1 Message Date
Niels Lohmann 0da083744a Address two Clang-Tidy findings the custom container tests exposed
Both come from instantiating basic_json with containers other than the
default ones, and neither shows up with the Clang-Tidy version available
outside CI:

- insert(const_iterator, basic_json&&) forwards its by-value iterator to
  the const-reference overload. performance-unnecessary-value-param asks
  for the copy to be a move; it only fires for an iterator that is not
  trivially copyable, as std::deque's is not. The NOLINT on the function
  does not cover it, because the finding is reported where the parameter
  is used rather than where it is declared. Move it, which is what the
  check asks for and is a (very small) improvement in its own right.

- cppcoreguidelines-use-enum-class rejects the unnamed enum that shadowed
  the inherited key_compare member type. An enum class would not do, since
  it declares a type of that name and the probe would find it again; a
  member function declaration hides the name just as well.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 19:49:20 +00:00
Niels Lohmann 5a2b8a274d Do not require the container iterators to be nothrow move constructible
iter_impl declared its defaulted move operations noexcept. The exception
specification a defaulted function gets implicitly follows from its members,
here internal_iterator, which holds the object and array iterators. libstdc++
gives std::deque's iterator a user-provided copy constructor without noexcept
before version 11, so the implicit specification is noexcept(false) and does
not match the declared one. That deletes the function -- and with g++ 4.8,
which predates CWG 1778, it is an error outright:

    error: function 'iter_impl<basic_json<std::map, std::deque> >::iter_impl(
    iter_impl&&)' defaulted on its first declaration with an
    exception-specification that differs from the implicit declaration

So std::deque, which this branch documents as a usable array type, could not
be used with an older standard library. Leaving the specification to be
computed cannot mismatch; iteration_proxy_value already spells out the same
condition next door.

The default configuration is unaffected: json::iterator, json::const_iterator
and ordered_json::iterator stay nothrow move constructible and move
assignable, which the test now checks so it cannot regress unnoticed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 19:09:51 +00:00
Niels Lohmann 8ce64b9c16 Move the custom BinaryType tests into their own translation unit
The two sections added to unit-regression2.cpp brought a third full
basic_json instantiation into a translation unit that was already large.
With Clang on MinGW that pushed the object over the reach of a 32-bit
relocation and test-regression2_cpp20.exe failed to link:

    relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'

unit-regression2.cpp is restored to exactly what it was before, and the
coverage moves to unit-custom-binary-type.cpp, next to the object and array
type tests it belongs with. The signed value type is now also covered in
C++11, where std::byte is not available.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 19:02:04 +00:00
Niels Lohmann 3b28316ee4 Do not parse the value in the array-at() test
JSON_DIAGNOSTIC_POSITIONS adds the byte range of the value to the exception
message, which a parsed value has and an in-memory one does not, so the two
message checks failed in that configuration. Build the array in memory
instead of parsing it; the test is about at(size_type) not needing
array_t::at(), and the byte range is beside the point.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 18:48:32 +00:00
Niels Lohmann 22c8a9554f Keep diagnostic key paths null-terminated
Building the token from data() and size() kept an embedded null byte in the
key, and since what() hands out a C string, that truncated the whole message
rather than just the key: to_bson() on a key containing U+0000 reported
"[json.exception.out_of_range.409] (/en" instead of the full explanation.
This broke test-bson under JSON_DIAGNOSTICS.

Constructing from data() alone stops at the first null byte, which is what
c_str() did before, so the message is unchanged -- without requiring
string_t to provide c_str().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 18:36:22 +00:00
Niels Lohmann 43afb5bebc Do not instantiate a hash map with an incomplete basic_json in the tests
object_t is probed for key_compare inside the definition of basic_json, so
it is instantiated while basic_json is still incomplete. Whether a hash map
survives that depends on the standard library: libstdc++ 9 needs the size of
the mapped type to instantiate std::unordered_map's node type and rejects
the adapter, which broke the GCC 9 builds.

The test now derives its no-key_compare object type from std::map -- which
does cope -- and shadows the inherited key_compare member type with an
entity that is not a type, so the library's probe finds none, exactly as for
a hash map. The unflatten() order-independence checks in unit-json_pointer
already cover the behaviour that the unordered object type was there for.
The limitation is documented for std::unordered_map.

Also address two Clang-Tidy findings the earlier commits introduced:
erase_from_object() declares its iterator with auto, and at(size_type) checks
the type first and then falls through to the return instead of throwing from
an else branch.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 18:17:15 +00:00
Niels Lohmann 110cd31e8f Use character literals for the signed BinaryType test
MSVC rejects char(0xFF) with C4310 (cast truncates constant value),
which the Windows workflow treats as an error. The character literals
carry the same byte values without a narrowing cast.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:53:02 +00:00
Niels Lohmann 0e4ad2e8da docs: record the reduced string_t and array_t requirements
Drop c_str(), back(), find(str, pos), replace(), and substr() from the
StringType requirements and at(size_type) from the ArrayType ones, and
note the string assignment the JSON pointer code performs. Streaming a
json_pointer no longer needs assignability from a std::string.

Add the non-null-terminated data() to the list of violations that are not
diagnosed at compile time -- it was described in the StringType section
but missing from the summary at the top -- and correct the QString row,
which no longer fails for the c_str() it lacks.

JSON_CATCH_USER no longer wraps a catch of std::out_of_range: the last one
went away with array_t::at(). Describe what the library actually catches.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:58 +00:00
Niels Lohmann ccb290facf Reduce the string_t and array_t members the library requires
Several members were required only because of how the library happened to
be written, not because the functionality needs them. Dropping them widens
the set of usable string and array types, and one of them was also a
performance problem.

string_t:

- c_str() is gone. Every call site already knew the length and passed it
  along, so data() is enough. The one place that did not, the diagnostics
  path in exceptions.hpp, now builds the token from data() and size(),
  which also stops it from truncating keys that contain a null byte.
- back() is gone; the serializer indexes the last character instead.
- find(str, pos), replace(), and substr() are gone. escape() and
  unescape() rebuilt the string with one replace() per escaped character,
  which moves the tail every time: escaping a string of n characters that
  all need escaping cost O(n^2). Both now scan with find_first_of() -- a
  member the pointer parser already required -- and append whole runs, so
  the common case is one search and one copy. Escaping 64000 tildes drops
  from 717 ms to 20 ms; a string with nothing to escape gets faster too
  (8.4 ms to 5.8 ms), because the scan is still a single memchr per pass.
  json_pointer::split() takes its reference tokens with the
  (const char*, size_type) constructor rather than substr().
- json_pointer::to_string() accumulates with concat<string_t> instead of
  letting concat default to std::string and converting afterwards, so
  streaming a json_pointer no longer requires string_t to be assignable
  from a std::string.

array_t:

- at(size_type) is gone. basic_json::at(size_type) checked the index by
  calling array_t::at() and translating std::out_of_range, which also
  required the array type to throw that exact exception. It now compares
  against size() and uses operator[]. The thrown exception, its message,
  and the behaviour under JSON_NOEXCEPTION are unchanged.

The BSON writer wrote the terminating null byte out of the string's own
buffer (size() + 1). It now writes the byte itself, so string_t::data()
need not be null-terminated for to_bson().

The tests pin the reduced API: alt_string loses the five dropped members
and gains coverage of the escaping paths, and a std::vector whose at() is
hidden is used as an ArrayType.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:58 +00:00
Niels Lohmann 26b10a7b18 docs: correct the template parameter requirements after independent verification
Every claim on the page was re-checked by compiling and running it, including
the rows that say a type cannot be used, which were checked to fail for the
documented reason and not merely to fail. Twenty-four claims were wrong.

The most consequential: the incomplete-type constraint applies to ObjectType
only. object_t is instantiated inside the class definition, because it is
probed for key_compare; array_t is only named there and is not instantiated
until basic_json is complete. So eastl::vector, QList and QVector are not
excluded by incomplete types at all -- they simply have no max_size() -- and
absl::InlinedVector is excluded for a subtler reason of its own.

Further corrections: ObjectType does not need erase(key), which has a fallback,
but does need at(key) for UBJSON output; only == and < are used, or == and <=>
under C++20, not all six; the documented adapter does not fit ankerl or
robin_hood. ArrayType needs no initializer-list insert, and value_type, the
(count, value) constructor and swappability are per-function, not always.
BinaryType needs a range insert for CBOR indefinite-length byte strings and
does not need push_back. StringType needs append(const StringType&)
unconditionally, and does not need operator!= or operator== against const
char*; empty(), resize(n) and reserve(n) are per-subsystem; int_to_string is
needed by diff, items and std::hash rather than by JSON Pointer or flatten.
BooleanType must be implicitly convertible from bool, and JSONSerializer's
second parameter need not carry a default.

std::pmr::string was wrong in the other direction this time: a moved-in string
does keep its memory resource, and later growth allocates from it. Only copies
land on the default resource.

Five requirement violations are not caught at compile time rather than the two
the page claimed; they are now listed together up front. Split every
compatibility table into what works and what does not, as the reasons in the
second half are the useful part.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:58 +00:00
Niels Lohmann 690c3be01d docs: remove a duplicated StringType compatibility section
The StringType section carried two 'Compatible types' tables and two copies of
the reference-implementation tip. The second table was a stale copy from before
the binary format string fixes and still listed std::pmr::string and
std::basic_string with a custom allocator as unusable, contradicting the
corrected table a few lines above it, and it dragged along the old explanation
that blamed int_to_string.

Drop the stale copy and put the surviving table before the notes, so the
'see below' in the std::pmr::string row points forwards.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:57 +00:00
Niels Lohmann ca47dd539d docs: qualify the std::pmr::string support claim
Listing std::pmr::string as fully supported was an overclaim: it was only ever
checked with the default memory resource, which is not what PMR is for.

basic_json cannot be given an allocator or a memory resource, so a pmr string
inside a value always allocates from std::pmr::get_default_resource(), and
assigning an arena-backed string into a value silently drops its resource,
because polymorphic_allocator does not propagate on copy construction. Passing
polymorphic_allocator as AllocatorType does not compile either. Only the
process-global set_default_resource() redirects these allocations.

Say so, and separate the row from std::basic_string with a custom stateless
allocator, which is unaffected.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:57 +00:00
Niels Lohmann 681fb07eb2 docs: cover fifo_map, gtl, folly::sorted_vector_map, and Qt
nlohmann::fifo_map works through the adapter that has always been documented
for it, and preserves the insertion order. Restore its mention in the object
order page, which was dropped together with the tsl::ordered_map one: unlike
ordered_map it keeps a lookup index, so it is the insertion-ordered option
without the quadratic cost.

gtl::flat_hash_map and folly::sorted_vector_map work as well, the latter
through an alias that drops the allocator, whose value type it disagrees on.
gtl::btree_map does not, for the same reason as the other btree containers.

None of the Qt containers can be used, each for its own reason: QMap has no
value_type, QHash iterators yield the mapped value rather than a pair, QList
has no max_size(), QByteArray spells empty() as isEmpty(), and QString is
UTF-16.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:57 +00:00
Niels Lohmann a02741fd28 docs: record Folly and the remaining vector replacements
Folly works, with the caveat that its headers need C++20: folly::fbstring as
StringType, folly::fbvector and folly::small_vector as ArrayType,
folly::fbvector<std::uint8_t> as BinaryType, and folly::F14NodeMap as
ObjectType through the usual argument-order adapter. folly::F14FastMap is the
exception and requires a complete value type.

For ArrayType, boost::container::devector, boost::container::static_vector
(within its fixed capacity), and std::pmr::vector work as well.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:57 +00:00
Niels Lohmann b8482ed7f4 docs: record compatibility for the common header-only hash maps
ankerl::unordered_dense (map and segmented_map), phmap (flat_hash_map and
node_hash_map), and robin_hood::unordered_flat_map all work as ObjectType
through the same adapter as Abseil's and Boost's hash maps, which only has to
restore the template argument order.

phmap::btree_map and robin_hood::unordered_node_map do not: like the other
btree containers they require a complete value type.

Note that none of these hash maps defines key_compare, so every one of them
depends on object_comparator_t falling back to default_object_comparator_t.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:56 +00:00
Niels Lohmann d386e0aa52 Do not require string_t to be convertible from std::string
Three places built a std::string and handed it to something expecting a
string_t: the UBJSON high-precision number reader, which every binary reader
instantiates, and the BSON writer's array element size calculation and write.
That silently required string_t to be implicitly convertible from std::string,
which std::string itself and types with a string_view conversion satisfy, but
many string types do not.

Construct the string_t explicitly from the data and size, which the
requirements already cover. This makes boost::container::string, eastl::string,
std::pmr::string, and std::basic_string with a custom allocator work as
StringType, none of which could previously be used with any binary format.

Add binary format coverage to the alt_string test, which had none, including a
UBJSON high-precision number -- the case that goes through the reader path.
BSON stays uncovered there: it additionally needs string_t::find(value_type),
which alt_string does not provide.

Also record which containers from Boost, Abseil, and EASTL work for each
template parameter, and correct two claims: std::pmr::string is usable after
this change, and tsl::ordered_map is not usable at all, because its iterators
expose the mapped value as const while basic_json modifies it in place.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:56 +00:00
Niels LohmannandClaude Opus 5 5d93f35463 Relax the ArrayType and ObjectType requirements
Two requirements forced users of otherwise suitable containers to write a
wrapper, and neither was load-bearing.

array_t::capacity() was read in push_back(), emplace_back(), operator+=(), and
operator[](size_type), but set_parent() only looks at the value under
JSON_DIAGNOSTICS; without diagnostics it was computed and discarded. Read it
through array_capacity(), which reports unknown_size() when diagnostics are off
or when the array type has no capacity() at all, and treat an unknown capacity
as "the elements may have moved" so the parent pointers are refreshed
conservatively. std::deque now works as ArrayType, in both builds, and
capacity() is no longer named at all in a default build. Since the capacity is
now only meaningful for array insertions, it moves out of set_parent() into
set_parent_after_array_insert().

basic_json::erase(iterator) assigned the object's erase() return value, which
requires the container to return the following iterator. Abseil's hash maps
return void to avoid computing a successor the caller may not need. Detect that
and compute the successor before erasing; containers that return an iterator,
including the vector-backed ordered_map where a precomputed successor would be
wrong, keep the existing path.

Together these leave an Abseil hash map needing only an alias that restores the
template argument order, and no adapter at all for std::deque.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:56 +00:00
Niels LohmannandClaude Opus 5 96806af2dc docs: note which Abseil containers can be used as template arguments
Checked against Abseil release 20250127.0 with the same workload as the other
entries on the page (DOM access, dump, parse, CBOR/MessagePack/UBJSON
round-trip, flatten, hash), with and without JSON_DIAGNOSTICS.

absl::flat_hash_map and absl::node_hash_map work as ObjectType through an
adapter that restores the template argument order and makes erase(iterator)
return the following iterator, which Abseil's returns as void. The page now
carries that adapter, and notes that absl::flat_hash_map does not keep
references to the mapped values valid across insertions while
absl::node_hash_map does. Both have a capacity() member, so JSON_DIAGNOSTICS
already refreshes the parent pointers conservatively for them.

absl::btree_map and absl::InlinedVector cannot be used at all: object_t and
array_t are formed while basic_json is still incomplete, and both inspect
their value type at class scope. std::map and std::vector are required by the
standard to tolerate this, third-party containers generally are not, so the
page states the constraint on its own rather than only per container.

absl::InlinedVector does work as BinaryType, where it is instantiated with a
complete type. absl::FixedArray and absl::Cord are not usable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:56 +00:00
Niels LohmannandClaude Opus 5 7a37a27a67 Fix unflatten and binary dumping for non-default configurations
unflatten() decided between array and object by looking at the first reference
token it happened to see for a node: it started an array only when that token
was 0. With a sorted object type the token 0 always arrives first, so the
result was correct by accident; with an object type whose iteration order is
unspecified, {"/c/2":3,"/c/1":2,"/c/0":1} unflattened to an object with the
keys "0", "1", and "2" instead of an array.

Collect the pointer prefixes that have a reference token 0 among their children
before building the result, and let get_and_create() consult that set. The
outcome is now independent of the iteration order and matches, for every input,
what a sorted object type produced before: a value is restored as an array if
and only if one of its keys is 0. Iterating the flattened object in a different
order would have been simpler, but it would have changed the key order of the
result for insertion-ordered object types.

The serializer, std::hash, and the UBJSON writer converted the elements of a
binary value to an integer implicitly, which does not compile for a BinaryType
whose value type is std::byte, and which made dump() write the bytes of a
signed value type as negative numbers. Convert to std::uint8_t explicitly in
all three places, so every byte type dumps as 0..255. The default
std::vector<std::uint8_t> configuration is unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:55 +00:00
Niels LohmannandClaude Opus 5 06feaa8d04 docs: list the types that are known to work for each template parameter
Follow up on the template parameter requirements page: state, for every
template parameter, which concrete types work and where they stop working.
Each entry was verified by compiling and running a common workload (DOM
access, dump, parse, CBOR/MessagePack round-trip, flatten, hash) against that
instantiation.

Findings worth calling out:

- ObjectType no longer needs a key_compare member type, so the std::unordered_map
  adapter only has to restore the template argument order. A hash-ordered
  ObjectType works everywhere except unflatten(), which reconstructs an array
  only when it meets the reference token 0 before the other indices.
- ArrayType: std::deque works when wrapped to add capacity(); std::list does not.
- StringType: std::pmr::string and std::basic_string with a custom allocator
  compile for the DOM, dump, and parse, but not for the binary readers, flatten,
  or diff, because the library assigns std::string values to string_t and
  int_to_string cannot be overloaded for a type in namespace std.
- NumberFloatType: long double works for dump and parse but not for the binary
  formats, which have no encoding for it.
- BinaryType: std::vector<std::byte> supports assignment, get, and the binary
  formats, but neither dump nor std::hash<basic_json>.

Also record the object_comparator_t fix in its version history.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:55 +00:00
Niels LohmannandClaude Opus 5 b1c9a68b9b Fix object_comparator_t for object types without key_compare
detail::actual_object_comparator selected between object_t::key_compare and
default_object_comparator_t with std::conditional. Both type arguments of
std::conditional are named eagerly, so object_t::key_compare had to exist
regardless of the condition, and the has_key_compare guard added in 3.11.0
never took effect: any ObjectType without a key_compare member type failed to
compile while instantiating basic_json itself.

Use detected_or_t instead, which resolves through a SFINAE partial
specialization and only names object_t::key_compare when it exists. The
selected type is unchanged for every object type that compiled before, so
object_comparator_t -- a public member type -- keeps its meaning and ABI.

has_key_compare had no other users and is removed.

Add a regression test using an adapter around std::unordered_map, which has no
key_compare; it fails to compile without this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:55 +00:00
Niels LohmannandClaude Opus 5 599bb1b68c docs: document the implicit requirements on basic_json's template parameters
The requirements that basic_json places on its eleven template parameters
were only implied by how the library uses the resulting object_t, array_t,
string_t, etc. Consumers had to discover them by trial and error.

Add "Template Parameter Requirements" collecting them, split into what is
always required and what is only required when a particular part of the API
is instantiated. Notable findings that were previously undocumented:

- ObjectType must provide a key_compare member type (actual_object_comparator
  names object_t::key_compare in both arms of a std::conditional), and its
  third template parameter is used as a comparator, so std::unordered_map
  cannot be used without a wrapper.
- ArrayType must provide capacity() -- push_back(), emplace_back(),
  operator+=(), and operator[](size_type) call it unconditionally -- and
  needs random-access iterators, so std::deque and std::list do not work.
- StringType needs contiguous, null-terminated data(), a one-byte value_type,
  and either assignability from std::to_string or an ADL int_to_string().
- NumberFloatType must be float, double, or long double for parsing and
  serialization; the integer types must satisfy std::is_integral.
- AllocatorType must be stateless, support incomplete types, and use plain
  pointers.
- BooleanType and the number types are union members and must be trivial.

Link the new page from the basic_json overview, the types feature page, and
the individual type alias pages, and correct the container examples given for
ObjectType (std::unordered_map) and ArrayType (std::list), which do not work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018hxZxz8svM54c6ATEvXp5E
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-28 17:38:55 +00:00
elix3randGitHub 35705d79d8 Fix update(merge_objects=true) throwing on primitive-to-object merge (#5414)
When merge_objects is true, recurse only if the existing value is an
object. Otherwise overwrite, matching the documented "all other values
are overwritten as usual" behavior.

Fixes #5402

Signed-off-by: elix3r <157088510+22elix3r@users.noreply.github.com>
2026-08-28 13:38:28 +01:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
892be68ca4 Bump the codeql-action group across 1 directory with 4 updates (#5388)
Bumps the codeql-action group with 4 updates in the / directory: [github/codeql-action/init](https://github.com/github/codeql-action), [github/codeql-action/autobuild](https://github.com/github/codeql-action), [github/codeql-action/analyze](https://github.com/github/codeql-action) and [github/codeql-action/upload-sarif](https://github.com/github/codeql-action).


Updates `github/codeql-action/init` from 4.37.6 to 4.37.7
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/5595ccaf912efad79be6eef63a5619ff05969be3...ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd)

Updates `github/codeql-action/autobuild` from 4.37.6 to 4.37.7
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/5595ccaf912efad79be6eef63a5619ff05969be3...ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd)

Updates `github/codeql-action/analyze` from 4.37.6 to 4.37.7
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/5595ccaf912efad79be6eef63a5619ff05969be3...ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd)

Updates `github/codeql-action/upload-sarif` from 4.37.6 to 4.37.7
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/5595ccaf912efad79be6eef63a5619ff05969be3...ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd)

---
updated-dependencies:
- dependency-name: github/codeql-action/analyze
  dependency-version: 4.37.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/autobuild
  dependency-version: 4.37.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/init
  dependency-version: 4.37.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.37.7
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: codeql-action
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-28 13:30:33 +01:00
whnandGitHub 1ac268d409 docs: correct to_bson complexity (#5334)
Signed-off-by: whn <142425816+Whning0513@users.noreply.github.com>
2026-08-28 13:27:18 +01:00
3fa93dac65 docs: document std::pair/std::tuple serializing as an object for string-keyed pairs (#5442)
A std::pair or std::tuple whose every element is itself a two-element array
with a string first element (e.g. std::pair<std::string, int>) serializes to
a JSON object instead of a JSON array, because to_json builds the value with a
brace initializer and the initializer-list object-detection rule fires. The
resulting object cannot be read back into the original type and collapses
duplicate keys. Document this quirk in the conversions guide, together with the
unaffected cases and the idiom to force an array.


Claude-Session: https://claude.ai/code/session_016cwQq8WQRFzQcGQtbJTtJg

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-08-28 13:26:01 +01:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
1876493f87 Bump step-security/harden-runner from 2.20.1 to 2.21.0 (#5394)
Bumps [step-security/harden-runner](https://github.com/step-security/harden-runner) from 2.20.1 to 2.21.0.
- [Release notes](https://github.com/step-security/harden-runner/releases)
- [Commits](https://github.com/step-security/harden-runner/compare/b09bb98e06d4d774595224525879c09bc6e98c40...05e31511f85b41b11d1cf0ef85d0992719546e2c)

---
updated-dependencies:
- dependency-name: step-security/harden-runner
  dependency-version: 2.21.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-27 07:39:56 +01:00
whnandGitHub 01853ed6bc docs: document lenient BSON input handling (#5333)
Signed-off-by: whn <142425816+Whning0513@users.noreply.github.com>
2026-08-25 07:18:50 +01:00
Krishnanand GandGitHub 2f025f401e Throw other_error.502 when UBJSON use_type is set without use_size (#5380)
* Throw other_error.502 when UBJSON use_type is set without use_size

Fixes #5321

Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>

* Scope UBJSON use_type check to container branches and expand tests

Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>

* Re-amalgamate single_include/json.hpp

The previous commit updated the split headers but the amalgamated
file didn't go back through astyle before I committed it, so CI's
amalgamation check caught formatting drift in json_fwd.hpp and a
few noexcept clauses in basic_json, plus one doc example. None of
it touches the UBJSON logic. Applied the patch CI generated to
bring single_include back in sync.

Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>

---------

Signed-off-by: Krishnanand G <118352827+Krishnanand-G@users.noreply.github.com>
2026-08-25 07:16:10 +01:00
Niels LohmannandGitHub 734fd305a1 Format-check the documentation examples in CI (#5386)
* Reformat parser_callback_t example with astyle

The file uses "json & /*parsed*/" in three lambda parameter lists, which
astyle rewrites to "json& /*parsed*/" per --align-reference=type. The
drift went unnoticed because CI never format-checked the documentation
examples; "make pretty" does cover them.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Format-check the documentation examples in CI

The examples live in docs/mkdocs/docs/examples, but both format checks
still referenced the long-gone docs/examples path:

- check_amalgamation.yml passed it to find, which printed an error for
  the missing path and carried on, so astyle only ever saw include and
  tests. The step still exited 0.
- ci.cmake globbed it into INDENT_FILES, and a GLOB_RECURSE over a
  missing directory silently yields nothing, so the ci_test_amalgamation
  target skipped the examples too.

Either way the 231 example files have never been format-checked. Point
both at the real path, and guard the workflow with an explicit directory
check so a future rename fails the job instead of quietly shrinking the
file list again.

Also drop the dead docs/examples/** path filter from
publish_documentation.yml; docs/mkdocs/** already covers the examples.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-20 12:32:50 +02:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
36187cacfb ⬆️ Bump wheel from 0.47.0 to 0.48.0 in /docs/mkdocs (#5385)
Bumps [wheel](https://github.com/pypa/wheel) from 0.47.0 to 0.48.0.
- [Release notes](https://github.com/pypa/wheel/releases)
- [Changelog](https://github.com/pypa/wheel/blob/main/docs/news.rst)
- [Commits](https://github.com/pypa/wheel/compare/0.47.0...0.48.0)

---
updated-dependencies:
- dependency-name: wheel
  dependency-version: 0.48.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-20 12:32:34 +02:00
Sahil_KamateandGitHub b5378e8deb Fix CBOR tag handlers not recognizing tags 0-5 and 21-23 (#5331)
* Fix CBOR tag handlers not recognizing tags 0-5 and 21-23

The tagged-item switch in binary_reader::parse_cbor_internal() only handled
head bytes 0xC6-0xD4 and 0xD8-0xDB. Bytes 0xC0-0xC5 (tags 0-5: date/time,
epoch, bignum, decimal, bigfloat) and 0xD5-0xD7 (tags 21-23: base64url,
base64, base16 conversion hints) fell through to the default case and were
reported as invalid bytes, even under cbor_tag_handler_t::ignore and ::store,
despite being valid CBOR major-type-6 tags per RFC 8949.

Add the missing case labels so the full 0xC0-0xDB range is handled
uniformly. Extend the "Tagged values" test in unit-cbor.cpp to cover
0xC0-0xD7, and update the CBOR docs to state the corrected tag range.

Fixes #5315

Signed-off-by: sahilkamate03 <45514385+sahilkamate03@users.noreply.github.com>

* Fix stale CBOR tag docs and add store-mode binary-payload test

The "Incomplete mapping" warning still listed tags 0-5 (date/time,
bignum, decimal fraction, bigfloat) and 21-23 (expected conversions)
as unsupported, even though they now parse correctly under
cbor_tag_handler_t::ignore/store, same as 0xC6..0xD4/0xD8..0xDB.
Remove those five bullets and cross-reference the "Tagged items"
warning below, matching the equivalent docs fix landed independently
in PR #5367.

Also add a cbor_tag_handler_t::store test that wraps a binary
payload (not just a string) for every byte in 0xC0..0xD7, confirming
these tags are unwrapped the same way as 0xC6..0xD4 rather than
mistaken for the 0xD8..0xDB binary-subtype marker syntax, per review
feedback on #5331.

Signed-off-by: sahilkamate03 <45514385+sahilkamate03@users.noreply.github.com>

---------

Signed-off-by: sahilkamate03 <45514385+sahilkamate03@users.noreply.github.com>
2026-08-19 20:19:47 +02:00
70 changed files with 2182 additions and 3060 deletions
+13 -3
View File
@@ -11,7 +11,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -34,7 +34,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -67,8 +67,18 @@ jobs:
${{ github.workspace }}/venv/bin/astyle --project=tools/astyle/.astylerc --suffix=none --quiet \
$INCLUDE_DIR/json.hpp $INCLUDE_DIR/json_fwd.hpp
# fail loudly if a directory is renamed or removed: find would only warn
# about the missing path and silently drop its files from the check
SOURCE_DIRS="docs/mkdocs/docs/examples include tests"
for DIR in $SOURCE_DIRS; do
if [ ! -d "$DIR" ]; then
echo "::error::source directory '$DIR' does not exist"
exit 1
fi
done
${{ github.workspace }}/venv/bin/astyle --project=tools/astyle/.astylerc --suffix=none --quiet \
$(find docs/examples include tests -type f \( -name '*.hpp' -o -name '*.cpp' -o -name '*.cu' \) -not -path 'tests/thirdparty/*' -not -path 'tests/abi/include/nlohmann/*' | sort)
$(find $SOURCE_DIRS -type f \( -name '*.hpp' -o -name '*.cpp' -o -name '*.cu' \) -not -path 'tests/thirdparty/*' -not -path 'tests/abi/include/nlohmann/*' | sort)
- name: Build patch and check for differences
id: diff
+1 -1
View File
@@ -9,7 +9,7 @@ jobs:
runs-on: ubuntu-22.04
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
+4 -4
View File
@@ -27,7 +27,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -38,14 +38,14 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
uses: github/codeql-action/init@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7
with:
languages: c-cpp
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
uses: github/codeql-action/autobuild@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
uses: github/codeql-action/analyze@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7
@@ -19,7 +19,7 @@ jobs:
pull-requests: write
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
+1 -1
View File
@@ -17,7 +17,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
+2 -2
View File
@@ -27,7 +27,7 @@ jobs:
security-events: write
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -43,6 +43,6 @@ jobs:
output: 'flawfinder_results.sarif'
- name: Upload analysis results to GitHub Security tab
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7
with:
sarif_file: ${{github.workspace}}/flawfinder_results.sarif
+1 -1
View File
@@ -17,7 +17,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
+1 -2
View File
@@ -7,7 +7,6 @@ on:
- develop
paths:
- docs/mkdocs/**
- docs/examples/**
workflow_dispatch:
# we don't want to have concurrent jobs, and we don't want to cancel running jobs to avoid broken publications
@@ -27,7 +26,7 @@ jobs:
runs-on: ubuntu-22.04
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
+2 -2
View File
@@ -36,7 +36,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -76,6 +76,6 @@ jobs:
# Upload the results to GitHub's code scanning dashboard.
- name: "Upload to code-scanning"
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7
with:
sarif_file: results.sarif
+2 -2
View File
@@ -32,7 +32,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -61,7 +61,7 @@ jobs:
# Upload SARIF file generated in previous step
- name: Upload SARIF file
uses: github/codeql-action/upload-sarif@5595ccaf912efad79be6eef63a5619ff05969be3 # v4.37.6
uses: github/codeql-action/upload-sarif@ff2f1c621b7f889edc0d3c761ac2e6a3f8cdb0dd # v4.37.7
with:
sarif_file: semgrep.sarif
if: always()
+1 -1
View File
@@ -16,7 +16,7 @@ jobs:
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
+6 -6
View File
@@ -35,7 +35,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -60,7 +60,7 @@ jobs:
target: [ci_test_amalgamation, ci_test_single_header, ci_cppcheck, ci_cpplint, ci_reproducible_tests, ci_non_git_tests, ci_offline_testdata, ci_reuse_compliance, ci_test_valgrind]
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -100,7 +100,7 @@ jobs:
container: ubuntu:focal
strategy:
matrix:
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_simdutf]
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls]
steps:
- name: Install build-essential
run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev
@@ -118,7 +118,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -369,7 +369,7 @@ jobs:
runs-on: ubuntu-latest
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
@@ -392,7 +392,7 @@ jobs:
target: [ci_test_examples, ci_test_build_documentation]
steps:
- name: Harden Runner
uses: step-security/harden-runner@b09bb98e06d4d774595224525879c09bc6e98c40 # v2.20.1
uses: step-security/harden-runner@05e31511f85b41b11d1cf0ef85d0992719546e2c # v2.21.0
with:
egress-policy: audit
+1 -19
View File
@@ -212,24 +212,6 @@ add_custom_target(ci_test_legacycomparison
COMMENT "Compile and test with legacy discarded value comparison enabled"
)
###############################################################################
# Validate UTF-8 with simdutf.
###############################################################################
add_custom_target(ci_test_simdutf
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON -DJSON_TestSimdutf=ON
# simdutf needs C++17, so the library falls back to its scalar validator
# below that: build the suite at C++11 to cover the fallback with the macro
# defined, and at C++17 to run every test against simdutf itself
"-DJSON_TestStandards=11\;17"
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_simdutf
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_simdutf
COMMAND cd ${PROJECT_BINARY_DIR}/build_simdutf && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure
COMMENT "Compile and test with simdutf UTF-8 validation enabled"
)
###############################################################################
# Enable brace-init copy semantics.
###############################################################################
@@ -312,7 +294,7 @@ file(GLOB_RECURSE INDENT_FILES
${PROJECT_SOURCE_DIR}/tests/src/*.cpp
${PROJECT_SOURCE_DIR}/tests/src/*.hpp
${PROJECT_SOURCE_DIR}/tests/benchmarks/src/benchmarks.cpp
${PROJECT_SOURCE_DIR}/docs/examples/*.cpp
${PROJECT_SOURCE_DIR}/docs/mkdocs/docs/examples/*.cpp
)
set(include_dir ${PROJECT_SOURCE_DIR}/single_include/nlohmann)
+6 -1
View File
@@ -14,7 +14,11 @@ To store objects in C++, a type is defined by the template parameters explained
## Template parameters
`ArrayType`
: container type to store arrays (e.g., `std::vector` or `std::list`)
: container type to store arrays. It must be a vector-like container: the library uses `operator[]`, `at()`, and
`resize()`, and requires random-access iterators. `#!cpp std::vector` and `#!cpp std::deque` qualify;
`#!cpp std::list` does not. See
[Template Parameter Requirements](../../features/types/template_parameters.md#arraytype) for the full list of
requirements.
`AllocatorType`
: the allocator to use for objects (e.g., `std::allocator`)
@@ -66,3 +70,4 @@ Arrays are stored as pointers in a `basic_json` type. That is, for any access to
## Version history
- Added in version 1.0.0.
- Made `capacity()` optional, so that array types such as `#!cpp std::deque` can be used, in version 3.13.0.
+11 -1
View File
@@ -42,7 +42,9 @@ represent a byte array in modern C++.
`value_type` must additionally be exactly one byte wide (e.g., `std::uint8_t`/`char`/`std::byte`): the binary
serializers (CBOR, MessagePack, BSON, UBJSON) read and write the container's raw bytes via
`reinterpret_cast`, which is only correct for byte-sized elements -- a container like
`#!cpp std::vector<std::intptr_t>` will not work as `BinaryType`.
`#!cpp std::vector<std::intptr_t>` will not work as `BinaryType`. The elements must be stored contiguously, and
the binary readers additionally require `resize()` and `operator[]`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#binarytype) for the full list.
## Notes
@@ -50,6 +52,11 @@ represent a byte array in modern C++.
The default values for `BinaryType` is `#!cpp std::vector<std::uint8_t>`.
#### Supported byte types
`#!cpp std::vector<std::uint8_t>`, `#!cpp std::vector<char>`, and `#!cpp std::vector<std::byte>` are supported.
Regardless of which of them is configured, [`dump`](dump.md) writes the bytes as the numbers 0..255.
#### Custom BinaryType behavior
When a custom `BinaryType` is configured (other than the default `#!cpp std::vector<std::uint8_t>`), you can assign
@@ -126,3 +133,6 @@ type `#!cpp binary_t*` must be dereferenced.
## Version history
- Added in version 3.8.0. Changed the type of subtype to `std::uint64_t` in version 3.10.0.
- Fixed [`dump`](dump.md), [`std::hash`](std_hash.md), and [`to_ubjson`](to_ubjson.md) for byte types that are not
integers (e.g., `#!cpp std::byte`) in version 3.13.0. `dump` now writes the bytes of a signed byte type (e.g.,
`#!cpp char`) as 0..255 rather than as negative numbers.
@@ -11,6 +11,14 @@ literals `#!json true` and `#!json false`.
To store boolean values in C++, a type is defined by the template parameter `BooleanType` which chooses the type to use.
## Template parameters
`BooleanType`
: the type to store booleans. As it is stored directly inside a `basic_json` value (in a union), it must be a
trivially default-constructible, trivially copyable, and trivially destructible type that is convertible to and
from `#!cpp bool`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#booleantype).
## Notes
#### Default type
+4
View File
@@ -35,6 +35,10 @@ class basic_json;
| `BinaryType` | type for binary arrays | [`binary_t`](binary_t.md) |
| `CustomBaseClass` | extension point for user code | [`json_base_class_t`](json_base_class_t.md) |
The library imposes a number of requirements on these types that are not expressed as C++ concepts, such as the
container operations `object_t` and `array_t` must provide, or the fact that `StringType` must be `char`-based. They
are collected in [Template Parameter Requirements](../../features/types/template_parameters.md).
## Specializations
- [**json**](../json.md) - default specialization
@@ -21,8 +21,11 @@ The default value for `CustomBaseClass` is `void`. In this case, an
#### Limitations
The type `CustomBaseClass` has to be a default-constructible class.
The type `CustomBaseClass` has to be a default-constructible, non-`final` class.
`basic_json` only supports copy/move construction/assignment if `CustomBaseClass` does so as well.
A `CustomBaseClass` with non-static data members forfeits `basic_json`'s
[standard layout](https://en.cppreference.com/w/cpp/named_req/StandardLayoutType) guarantee. See
[Template Parameter Requirements](../../features/types/template_parameters.md#custombaseclass).
## Examples
@@ -19,6 +19,12 @@ using json_serializer = JSONSerializer<T, SFINAE>;
The default values for `json_serializer` is [`adl_serializer`](../adl_serializer/index.md).
#### Requirements
A custom serializer must provide `#!cpp static void to_json(basic_json&, T)` for every type it serializes, and either
`#!cpp static void from_json(const basic_json&, T&)` or `#!cpp static T from_json(const basic_json&)` for every type it
deserializes. See [Template Parameter Requirements](../../features/types/template_parameters.md#jsonserializer).
## Examples
??? example
@@ -20,6 +20,16 @@ used.
To store floating-point numbers in C++, a type is defined by the template parameter `NumberFloatType` which chooses the
type to use.
## Template parameters
`NumberFloatType`
: the type to store floating-point numbers. Parsing and serialization are implemented in terms of
`#!cpp std::strtof`/`#!cpp std::strtod`/`#!cpp std::strtold` and `#!cpp std::snprintf`, so the type must be
`#!cpp float`, `#!cpp double`, or `#!cpp long double`. The
[binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`,
because they have no encoding for `#!cpp long double`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype).
## Notes
#### Default type
@@ -20,6 +20,13 @@ used.
To store integer numbers in C++, a type is defined by the template parameter `NumberIntegerType` which chooses the type
to use.
## Template parameters
`NumberIntegerType`
: the type to store signed integers. It must be a **signed integral** type (`#!cpp std::is_integral`) with a
`#!cpp std::numeric_limits` specialization, and it is stored directly inside a `basic_json` value. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberintegertype-and-numberunsignedtype).
## Notes
#### Default type
@@ -20,6 +20,14 @@ used.
To store unsigned integer numbers in C++, a type is defined by the template parameter `NumberUnsignedType` which chooses
the type to use.
## Template parameters
`NumberUnsignedType`
: the type to store unsigned integers. It must be an **unsigned integral** type (`#!cpp std::is_integral`) with a
`#!cpp std::numeric_limits` specialization, and it must be able to represent the absolute value of every
[`number_integer_t`](number_integer_t.md) value. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberintegertype-and-numberunsignedtype).
## Notes
#### Default type
@@ -30,3 +30,5 @@ and [`default_object_comparator_t`](default_object_comparator_t.md) otherwise.
- Added in version 3.0.0.
- Changed to be conditionally defined as `#!cpp typename object_t::key_compare` or `default_object_comparator_t` in
version 3.11.0.
- Fixed the fallback to `default_object_comparator_t`, which previously failed to compile for object types without a
`key_compare` member type, in version 3.13.0.
+6 -1
View File
@@ -18,7 +18,11 @@ To store objects in C++, a type is defined by the template parameters described
## Template parameters
`ObjectType`
: the container to store objects (e.g., `std::map` or `std::unordered_map`)
: the container to store objects. Its template parameters must have the same order and meaning as those of
`std::map`; in particular, the third parameter is a comparator. `#!cpp std::unordered_map`, whose third parameter
is a hash function, therefore needs an adapter -- see
[Template Parameter Requirements](../../features/types/template_parameters.md#objecttype) for the full list of
requirements, an adapter example, and the containers that are known to work.
`StringType`
: the type of the keys or names (e.g., `std::string`). The comparison function `std::less<StringType>` is used to
@@ -122,3 +126,4 @@ the object is silently converted as an array of key-value pairs, which is incorr
## Version history
- Added in version 1.0.0.
- Allowed object types whose `erase(iterator)` returns `#!cpp void` in version 3.13.0.
@@ -23,6 +23,11 @@ JSON class into byte-sized characters during deserialization.
`StringType`. To work with wide-character data, convert it to/from UTF-8 at the boundary instead -- see the
FAQ's [wide string handling](../../home/faq.md#wide-string-handling) section for a conversion recipe.
Beyond the character type, the library expects a substantial part of the `#!cpp std::string` interface (contiguous
null-terminated `data()`, `substr()`, `find()`, `append()`, ...). See
[Template Parameter Requirements](../../features/types/template_parameters.md#stringtype) for the full list and
for the string types that are known to work.
## Notes
#### Default type
@@ -78,3 +83,5 @@ and an example.
## Version history
- Added in version 1.0.0.
- Removed the requirement that `string_t` be implicitly convertible from `#!cpp std::string`, which the BSON writer and
the UBJSON reader relied on, in version 3.13.0.
@@ -52,6 +52,11 @@ optional, `#!cpp bjdata_version_t::draft2` by default.
Strong guarantee: if an exception is thrown, there are no changes in the JSON value.
## Exceptions
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
## Complexity
Linear in the size of the JSON value `j`.
+3 -1
View File
@@ -46,7 +46,9 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
## Complexity
Linear in the size of the JSON value `j`.
Proportional to the size of the JSON value `j` multiplied by its maximum nesting
depth, `O(n × d)`. BSON length prefixes are computed recursively before nested
values are written.
## Examples
@@ -45,6 +45,11 @@ The exact mapping and its limitations are described on a [dedicated page](../../
Strong guarantee: if an exception is thrown, there are no changes in the JSON value.
## Exceptions
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
## Complexity
Linear in the size of the JSON value `j`.
+9 -1
View File
@@ -37,7 +37,14 @@ Linear in the size of the JSON value.
## Notes
Empty objects and arrays are flattened by [`flatten()`](flatten.md) to `#!json null` values and cannot unflattened to
their original type. Apart from this example, for a JSON value `j`, the following is always true:
their original type.
A flattened array and a flattened object whose keys are array indices are indistinguishable, because both are
described by the same JSON pointers. A value is therefore restored as an array if and only if one of its keys is the
reference token `0`, and as an object otherwise: `#!json {"2": 1}` is restored unchanged, whereas `#!json {"0": 1}` is
restored as `#!json [1]`. This decision does not depend on the order in which the flattened object is iterated.
Apart from these two cases, for a JSON value `j`, the following is always true:
`#!cpp j == j.flatten().unflatten()`.
## Examples
@@ -63,3 +70,4 @@ their original type. Apart from this example, for a JSON value `j`, the followin
## Version history
- Added in version 2.0.0.
- Made the array/object decision independent of the object's iteration order in version 3.13.0.
-1
View File
@@ -24,7 +24,6 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
- [**JSON_USE_SIMDUTF**](json_use_simdutf.md) - use the simdutf library to accelerate UTF-8 validation
## Library version
@@ -12,9 +12,11 @@
Controls how exceptions are handled by the library.
1. This macro overrides [`#!cpp catch`](https://en.cppreference.com/w/cpp/language/try_catch) calls inside the library.
The argument is the type of the exception to catch. As of version 3.8.0, the library only catches `std::out_of_range`
exceptions internally to rethrow them as [`json::out_of_range`](../../home/exceptions.md#out-of-range) exceptions.
The macro is always followed by a scope.
The argument is the type of the exception to catch. The library uses it in a single place: to swallow any exception
escaping the parent-pointer check that [`JSON_DIAGNOSTICS`](json_diagnostics.md) adds to the class invariant. The
places where the library catches its own [`json::out_of_range`](../../home/exceptions.md#out-of-range) exceptions
use `JSON_INTERNAL_CATCH` instead, which `JSON_CATCH_USER` also overrides unless `JSON_INTERNAL_CATCH_USER` is
defined. The macro is always followed by a scope.
2. This macro overrides `#!cpp throw` calls inside the library. The argument is the exception to be thrown. Note that
`JSON_THROW_USER` should leave the current scope (e.g., by throwing or aborting), as continuing after it may yield
undefined behavior.
@@ -1,71 +0,0 @@
# JSON_USE_SIMDUTF
```cpp
#define JSON_USE_SIMDUTF
```
When defined, the parser validates the UTF-8 content of JSON strings that come from a **contiguous byte input**
(`std::string`, `std::vector<char>`/`<std::uint8_t>`, string literals, `const char*` ranges, …) using the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. On text with many
non-ASCII characters (e.g. CJK or emoji) this can validate several times faster.
This is an **opt-in external dependency**. The library itself remains header-only and its behavior is unchanged: the
same input is accepted or rejected either way, and every parse error is reported at the same position with the same
message (simdutf is only used to fast-path *valid* runs; anything it flags falls back to the scalar path so the exact
diagnostic is preserved). Streaming inputs (files, `std::istream`, wide strings, user-defined adapters) always use the
scalar path.
When `JSON_USE_SIMDUTF` is defined you must make the `simdutf.h` header available on the include path and link the
simdutf library. When it is not defined, no simdutf header is included and there is no dependency.
!!! note "Requires C++17"
simdutf requires C++17 and its header rejects older standards with an `#!cpp #error`. The backend is therefore only
compiled in from C++17 on. In C++11 and C++14 the macro has no effect and the scalar validator is used, which
accepts and rejects exactly the same input -- only throughput differs. Setting the macro project-wide is therefore
safe even when some translation units are built with an older standard.
!!! warning "Define consistently"
The macro selects between two definitions of the same inline validation function. It must therefore be defined
identically for **every** translation unit that includes the library; mixing translation units that define it with
ones that do not is an ODR violation. Prefer setting it as a compile definition on the target rather than with
`#!cpp #define` in individual source files.
## Default definition
By default, `#!cpp JSON_USE_SIMDUTF` is not defined and the portable C++11 scalar validator is used.
```cpp
#undef JSON_USE_SIMDUTF
```
## Examples
??? example
The code below enables the simdutf backend for UTF-8 validation.
```cpp
#define JSON_USE_SIMDUTF 1
#include <nlohmann/json.hpp>
...
```
The project must also link against simdutf, e.g. with CMake:
```cmake
target_compile_definitions(your_target PRIVATE JSON_USE_SIMDUTF)
target_link_libraries(your_target PRIVATE simdutf::simdutf)
```
!!! hint "Testing this configuration"
The unit tests can be built against the simdutf backend with the CMake option `JSON_TestSimdutf` (`OFF` by
default), which fetches simdutf and defines `JSON_USE_SIMDUTF` for every test target. The `ci_test_simdutf` target
runs the whole test suite in that configuration.
## Version history
- Added in version 3.13.0.
@@ -98,6 +98,17 @@ The library maps BSON record types to JSON value types as follows:
This library deserializes BSON type `0x11` (Timestamp) as a `number_unsigned` value. The 64-bit value is preserved,
but the Timestamp type information is not.
!!! warning "Lenient BSON input handling"
The BSON reader is lenient in a few areas where the BSON specification is more restrictive:
- array element keys are not checked against the required decimal sequence (`0`, `1`, `2`, ...),
- any non-zero byte is accepted as `true` for the boolean type, and
- the payload for binary subtype `0x02` is returned as-is, including its inner length prefix.
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
`from_bson()`.
??? example
```cpp
@@ -160,14 +160,11 @@ The library maps CBOR types to JSON value types as follows:
The mapping is **incomplete** in the sense that not all CBOR types can be converted to a JSON value. The following CBOR types are not supported and will yield parse errors:
- date/time (0xC0..0xC1)
- bignum (0xC2..0xC3)
- decimal fraction (0xC4)
- bigfloat (0xC5)
- expected conversions (0xD5..0xD7)
- simple values (0xE0..0xF3, 0xF8)
- undefined (0xF7)
Tagged items (0xC0..0xDB) are not interpreted either; see the note on tagged items below.
!!! warning "Negative integer overflow"
CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the
@@ -181,7 +178,7 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Tagged items"
Tagged items will throw a parse error by default. They can be ignored by passing `cbor_tag_handler_t::ignore` to function `from_cbor`. They can be stored by passing `cbor_tag_handler_t::store` to function `from_cbor`.
Tagged items (0xC0..0xDB) will throw a parse error by default. They can be ignored by passing `cbor_tag_handler_t::ignore` to function `from_cbor`, in which case the tag is skipped and the enclosed data item is parsed on its own. They can be stored by passing `cbor_tag_handler_t::store` to function `from_cbor`. Note that no tag is ever interpreted: for instance, a text string tagged with tag 0 (date/time) stays a string.
??? example
+24
View File
@@ -54,6 +54,30 @@ json j = {1.0, "hello", 42};
auto t = j.get<std::tuple<double, std::string, int>>(); // {1.0, "hello", 42}
```
!!! warning "Serializing a `std::pair`/`std::tuple` whose every element is a string-keyed pair"
When *every* element of a `#!cpp std::pair` or `#!cpp std::tuple` is itself a two-element array whose first
element is a string (for example `#!cpp std::pair<std::string, int>`), serializing it produces a JSON **object**
instead of the expected array:
```cpp
using kv = std::pair<std::string, int>;
json j = std::pair<kv, kv>{{"a", 1}, {"b", 2}}; // {"a":1,"b":2}, not [["a",1],["b",2]]
```
This is a consequence of the [brace-initializer object-detection rule](creating_values.md): the same rule that
lets `#!cpp json{{"a", 1}, {"b", 2}}` create an object also fires here. The resulting object cannot be read back
into the original type (`#!cpp get<std::pair<kv, kv>>()` throws [`type_error.302`](../home/exceptions.md#jsonexceptiontype_error302)),
and duplicate keys collapse into one, losing elements. This only affects `#!cpp std::pair`/`#!cpp std::tuple`
themselves; a `#!cpp std::vector<std::pair<std::string, int>>`, or a pair/tuple with at least one element that is
not a string-keyed pair, serializes to an array as expected. To force an array, build one explicitly from the
elements with [`array`](../api/basic_json/array.md):
```cpp
std::pair<kv, kv> p{{"a", 1}, {"b", 2}};
json a = json::array({p.first, p.second}); // [["a",1],["b",2]]
```
!!! info "Extracting references into a tuple"
A tuple type may also hold references (e.g. `#!cpp std::tuple<double&, std::string&>`) to avoid copying: `get`
-8
View File
@@ -137,14 +137,6 @@ behavior is deprecated and switched off (`0`) by default.
See [full documentation of `JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`](../api/macros/json_use_legacy_discarded_value_comparison.md).
## `JSON_USE_SIMDUTF`
When defined, UTF-8 validation of JSON strings read from contiguous byte input is delegated to the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. This is an opt-in
external dependency and is not defined by default.
See [full documentation of `JSON_USE_SIMDUTF`](../api/macros/json_use_simdutf.md).
## `NLOHMANN_DEFINE_TYPE_*(...)`, `NLOHMANN_DEFINE_DERIVED_TYPE_*(...)`
The library defines 12 macros to simplify the serialization/deserialization of types. See the page on
+5 -1
View File
@@ -51,7 +51,11 @@ If you do want to preserve the **insertion order**, you can use the type [`nlohm
--8<-- "examples/ordered_json.output"
```
Alternatively, you can use a more sophisticated ordered map like [`tsl::ordered_map`](https://github.com/Tessil/ordered-map) ([integration](https://github.com/nlohmann/json/issues/546#issuecomment-304447518)) or [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) ([integration](https://github.com/nlohmann/json/issues/485#issuecomment-333652309)).
Alternatively, [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) also preserves the insertion order and, unlike [`ordered_map`](../api/ordered_map.md), keeps a lookup index, so it does not have the quadratic cost described below. It is used through a small adapter ([integration](https://github.com/nlohmann/json/issues/485#issuecomment-333652309)).
If the order does not matter and you only want faster lookup, `boost::unordered_flat_map`, `absl::flat_hash_map`, `absl::node_hash_map`, and several other hash maps work through an adapter that restores the template argument order `basic_json` expects; see [Template Parameter Requirements](types/template_parameters.md#objecttype). Note these are *unordered*, not insertion-ordered.
[`tsl::ordered_map`](https://github.com/Tessil/ordered-map) cannot be used: its iterators expose the mapped value as `const`, while `basic_json` needs to modify it in place.
The [`ordered_map`](../api/ordered_map.md) behind `nlohmann::ordered_json` is deliberately minimal and has no lookup
index, so every key access is a linear scan and building an object of `n` keys costs O(n²). This is unnoticeable at
+6 -1
View File
@@ -79,7 +79,8 @@ template<
class NumberFloatType = double,
template<typename U> class AllocatorType = std::allocator,
template<typename T, typename SFINAE = void> class JSONSerializer = adl_serializer,
class BinaryType = std::vector<std::uint8_t>
class BinaryType = std::vector<std::uint8_t>,
class CustomBaseClass = void
>
class basic_json;
```
@@ -106,6 +107,10 @@ using number_float_t = NumberFloatType;
using binary_t = nlohmann::byte_container_with_subtype<BinaryType>;
```
Not every type can be passed for these template arguments: the library uses the resulting types in ways that imply a
number of requirements, for instance that `StringType` is `char`-based or that `ArrayType` is vector-like. These
requirements are collected in [Template Parameter Requirements](template_parameters.md).
## Objects
@@ -0,0 +1,663 @@
# Template Parameter Requirements
Class [`basic_json`](../../api/basic_json/index.md) is configurable through eleven template parameters. The library
never formally states what a type passed for one of these parameters has to provide -- the requirements are implied by
the way the library uses the resulting [`object_t`](../../api/basic_json/object_t.md),
[`array_t`](../../api/basic_json/array_t.md), [`string_t`](../../api/basic_json/string_t.md), etc. This page collects
these requirements so they do not have to be discovered by trial and error. Each section lists the concrete types
that are known to work for that parameter and the ones that do not, checked against Boost 1.83, Abseil 20250127.0,
Folly, EASTL 3.21, `ankerl::unordered_dense`, `phmap`, `gtl`, `robin_hood`, `tsl::ordered_map`, and Qt 6.
## How to read this page
Requirements are split into two groups:
- **Always required** -- needed to instantiate `basic_json` at all, or needed by functions that virtually every program
uses (construction, element access, [`dump`](../../api/basic_json/dump.md)).
- **Required for ...** -- only needed when a particular part of the API is instantiated. Member function templates are
only instantiated when they are used, so a type may be perfectly usable even though it does not satisfy these
requirements, as long as the corresponding functions are never called.
!!! warning "Requirements are not checked"
Apart from a `#!cpp static_assert` on the array iterator category, the requirements below are not diagnosed with
dedicated error messages. Violating most of them results in a compiler error somewhere inside the library. Six
violations are not caught at compile time at all:
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers,
because the lexer hands the buffer to `#!cpp std::strtoull`/`#!cpp std::strtoll`/`#!cpp std::strtod`.
- A [`BinaryType`](#binarytype) whose `value_type` is wider than one byte compiles and silently produces wrong
results, because the readers and writers reinterpret its storage as raw bytes.
- A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation,
and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance.
- A [`NumberUnsignedType`](#numberintegertype-and-numberunsignedtype) too narrow to hold the absolute value of
every `NumberIntegerType` value silently corrupts: with `#!cpp std::int64_t`/`#!cpp std::uint32_t`,
`#!cpp basic_json(INT64_MIN).dump()` yields `#!json -0`.
- The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion
in a normal build, and only fail silently under `#!cpp NDEBUG`.
## Overview
| Template parameter | Default | Notable substitutes |
|-------------------------------------------------------------------|-----------------------------------|------------------------------------------------------------------------|
| [`ObjectType`](#objecttype) | `std::map` | [`nlohmann::ordered_map`](../../api/ordered_map.md), Abseil hash maps |
| [`ArrayType`](#arraytype) | `std::vector` | `#!cpp std::deque` |
| [`StringType`](#stringtype) | `std::string` | `std::string`-like types over `char` |
| [`BooleanType`](#booleantype) | `bool` | none worth using |
| [`NumberIntegerType`](#numberintegertype-and-numberunsignedtype) | `std::int64_t` | any signed integer type |
| [`NumberUnsignedType`](#numberintegertype-and-numberunsignedtype) | `std::uint64_t` | any unsigned integer type |
| [`NumberFloatType`](#numberfloattype) | `double` | `float` (`long double`: no binary formats) |
| [`AllocatorType`](#allocatortype) | `std::allocator` | stateless allocators |
| [`JSONSerializer`](#jsonserializer) | `adl_serializer` | serializers with the same interface |
| [`BinaryType`](#binarytype) | `#!cpp std::vector<std::uint8_t>` | `#!cpp std::vector<char>` |
| [`CustomBaseClass`](#custombaseclass) | `void` | any default-constructible class |
!!! warning "Third-party containers and incomplete types"
`object_t` is instantiated inside the definition of `basic_json` -- it is probed for a `key_compare` member to
form [`object_comparator_t`](../../api/basic_json/object_comparator_t.md) -- i.e. while `basic_json` is still an
incomplete type. `#!cpp std::map` is required by the standard to support incomplete mapped types; most
third-party maps are not, and inspecting the mapped type at class scope (for instance with
`#!cpp std::is_trivially_move_assignable`) makes them unusable as `ObjectType`, no matter how their template
arguments are adapted. This rules out `absl::btree_map`, `phmap::btree_map`, `gtl::btree_map`,
`robin_hood::unordered_node_map`, `folly::F14FastMap`, and `eastl::hash_map`.
`array_t` is only *named* in the class definition and is not instantiated until `basic_json` is complete, so an
`ArrayType` that inspects its value type at class scope is generally fine -- `boost::container::small_vector` and
`static_vector` both reject incomplete value types yet work here. `absl::InlinedVector` is the exception: the
`#!cpp std::is_trivially_move_assignable<basic_json>` it evaluates while instantiating itself re-enters the
library's own trait machinery mid-instantiation.
!!! note "Folly requires C++20"
Folly's headers use `#!cpp consteval` and `#!cpp std::type_identity`, so any `basic_json` specialization that
names a Folly type has to be compiled as C++20 or later, whatever the rest of the library supports.
## `ObjectType`
`ObjectType` is instantiated as
```cpp
using object_t = ObjectType<StringType, // key_type
basic_json, // mapped_type
default_object_comparator_t, // key_compare
AllocatorType<std::pair<const StringType,
basic_json>>>; // allocator_type
```
i.e., the template arguments follow the order and meaning of `std::map`.
### Always required
- The template must be usable with **four** type arguments in the order shown above. The third argument is a
**comparator**; containers that expect something else in this position (e.g., a hash function) need an alias template
or wrapper -- see [Notes](#notes).
- An optional member type `key_compare`. If it is present it becomes
[`object_comparator_t`](../../api/basic_json/object_comparator_t.md); otherwise
[`default_object_comparator_t`](../../api/basic_json/default_object_comparator_t.md) is used.
- Member types `key_type`, `mapped_type`, `value_type`, and `iterator`.
- `value_type` must behave like `#!cpp std::pair<const key_type, mapped_type>`; the library accesses `.first` and
`.second` on it.
- `iterator` must be default-constructible and satisfy
[LegacyBidirectionalIterator](https://en.cppreference.com/w/cpp/named_req/BidirectionalIterator). The type returned
by `cbegin()`/`cend()` must satisfy the same requirements.
- Constructors: default, copy, move, and from an iterator range `(first, last)`.
- Member functions `begin()`, `end()`, `cbegin()`, `cend()`, `empty()`, `size()`, `max_size()`, `clear()`,
`find(key)`, `count(key)`, `emplace(key, value)`, `insert(value_type)`, `insert(first, last)`, `operator[](key)`,
`erase(iterator)`, and `erase(first, last)`. `erase(iterator)` may return the following iterator or `#!cpp void`;
in the latter case the library computes the successor itself, before erasing.
- `erase(key)` is **optional**: if the container does not provide one, the library falls back to `find(key)` followed
by `erase(iterator)`.
- `at(key)` is required only by [`to_ubjson`](../../api/basic_json/to_ubjson.md) and
[`to_bjdata`](../../api/basic_json/to_bjdata.md), but every container tried here provides it.
- `emplace` and `insert(value_type)` must return `#!cpp std::pair<iterator, bool>` and must have **unique-key**
semantics; multimaps cannot be used.
- The type must be swappable (via `std::swap` or an ADL `swap`).
- The comparison operators `==` and `<`; `!=`, `<=`, `>`, and `>=` are derived from them. Where the library uses
three-way comparison (C++20), `==` and `<=>` are required **instead** -- the six two-way operators do not satisfy
it. They implement [`basic_json`'s comparison operators](../../api/basic_json/operator_eq.md).
### Required for heterogeneous key lookup
The overloads of [`at`](../../api/basic_json/at.md), [`operator[]`](../../api/basic_json/operator%5B%5D.md),
[`find`](../../api/basic_json/find.md), [`contains`](../../api/basic_json/contains.md),
[`count`](../../api/basic_json/count.md), [`erase`](../../api/basic_json/erase.md), and
[`value`](../../api/basic_json/value.md) that accept a key type other than `object_t::key_type` require
- a **transparent** comparator, i.e. [`object_comparator_t`](../../api/basic_json/object_comparator_t.md) has a member
type `is_transparent` (this is why the default comparator is `#!cpp std::less<>` since C++14), and
- corresponding heterogeneous `find`, `count`, `erase`, and `operator[]` overloads on the container.
### Notes
#### `std::unordered_map` needs an adapter
`#!cpp std::unordered_map` cannot be passed directly: its third template parameter is a hash function, but
`basic_json` passes a comparator in that position. An alias template or wrapper that restores the expected argument
order makes it usable:
```cpp
template<class Key, class T, class IgnoredCompare, class Allocator>
struct unordered_map_object
: std::unordered_map<Key, T, std::hash<Key>, std::equal_to<Key>, Allocator>
{
using base_t = std::unordered_map<Key, T, std::hash<Key>, std::equal_to<Key>, Allocator>;
using base_t::base_t;
};
using unordered_json = nlohmann::basic_json<unordered_map_object>;
```
Whether `#!cpp std::unordered_map` can be instantiated at all depends on the standard library: `object_t` is formed
while `basic_json` is still incomplete (see the warning above), and libstdc++ 9 needs the size of the mapped type to
instantiate the hash map's node type, so the adapter does not compile there. Newer libstdc++ versions, and the hash
maps listed below, do not have that problem.
The adapter above works verbatim for Abseil's, Boost's, `phmap`'s and `gtl`'s hash maps, which all place the hash
function third and take a `#!cpp std::pair<const Key, T>` allocator fifth. Two need a different adapter:
- `ankerl::unordered_dense` expects an allocator over `#!cpp std::pair<Key, T>` (non-const key), so the allocator has
to be rebound to that or dropped.
- `robin_hood`'s fifth parameter is the non-type `MaxLoadFactor100`, so its adapter must drop the allocator entirely.
None of these hash maps defines `key_compare`, so all of them additionally rely on `object_comparator_t` falling back
to [`default_object_comparator_t`](../../api/basic_json/default_object_comparator_t.md); see
[`object_comparator_t`](../../api/basic_json/object_comparator_t.md).
#### Abseil hash maps
`absl::flat_hash_map` and `absl::node_hash_map` tolerate an incomplete value type, but they take a hash function as
their third template argument. The same adapter as for `#!cpp std::unordered_map` makes them usable:
```cpp
template<class Key, class T, class IgnoredCompare, class Allocator>
struct flat_hash_object
: absl::flat_hash_map<Key, T, absl::Hash<Key>, std::equal_to<Key>, Allocator>
{
using base_t = absl::flat_hash_map<Key, T, absl::Hash<Key>, std::equal_to<Key>, Allocator>;
using base_t::base_t;
};
using flat_hash_json = nlohmann::basic_json<flat_hash_object>;
```
`absl::node_hash_map` keeps references to the mapped values valid across insertions; `absl::flat_hash_map` does not,
which makes it behave like [`ordered_json`](../../api/ordered_json.md) with respect to
[iterator invalidation](../../api/basic_json/index.md#iterator-invalidation). Both expose a `capacity()` member
function, so [`JSON_DIAGNOSTICS`](../../api/macros/json_diagnostics.md) treats them conservatively and keeps the
parent pointers correct either way.
#### Iteration order
The library never relies on the container's iteration order for correctness; it does determine the order in which
object keys are serialized by [`dump`](../../api/basic_json/dump.md) and visited by
[`items`](../../api/basic_json/items.md). See [Object Order](../object_order.md).
#### `capacity()` marks a container as insertion-ordered
With [`JSON_DIAGNOSTICS`](../../api/macros/json_diagnostics.md) enabled, the library detects insertion-ordered maps by
probing for a `capacity()` member function (`nlohmann::ordered_map` inherits it from `std::vector`) and refreshes all
parent pointers after every insertion. An `ObjectType` that happens to have a `capacity()` member is therefore treated
conservatively -- this is correct, but slower.
#### Key order and duplicate keys
The library does not sort or de-duplicate keys itself; the behavior described in
[`object_t`](../../api/basic_json/object_t.md) is entirely the behavior of the chosen container.
### Compatible containers
| Container | Notes |
|---------------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------|
| `#!cpp std::map` (default) | |
| [`nlohmann::ordered_map`](../../api/ordered_map.md) | used by [`ordered_json`](../../api/ordered_json.md); keeps insertion order |
| [`nlohmann::fifo_map`](https://github.com/nlohmann/fifo_map) | keeps insertion order; adapter puts `fifo_map_compare` in the comparator slot |
| `boost::container::map`, `boost::container::flat_map` | no adapter needed |
| `#!cpp std::unordered_map` | through the adapter above; not with libstdc++ 9, see the note |
| `boost::unordered_map`, `boost::unordered_flat_map`, `boost::unordered_node_map` | through the adapter above |
| `absl::flat_hash_map`, `absl::node_hash_map` | through the adapter above; `flat_hash_map` moves mapped values on rehash |
| `phmap::flat_hash_map`, `phmap::node_hash_map`, `gtl::flat_hash_map` | through the adapter above |
| `ankerl::unordered_dense::map` and `segmented_map` | adapter must rebind or drop the allocator |
| `robin_hood::unordered_flat_map` | adapter must drop the allocator |
| `folly::F14NodeMap` | through the adapter above; requires C++20, see the note above |
| `folly::sorted_vector_map` | alias must drop the allocator, whose value type it disagrees on |
### Containers that cannot be used
| Container | Reason |
|---------------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------|
| `absl::btree_map`, `phmap::btree_map`, `gtl::btree_map`, `robin_hood::unordered_node_map`, `folly::F14FastMap`, `eastl::hash_map` | require a complete mapped type |
| `eastl::map` | EASTL iterators do not work with `#!cpp std::iterator_traits` |
| `tsl::ordered_map` | its iterators expose the mapped value as `#!cpp const` |
| `QMap` | no `value_type` member type |
| `QHash` | its `value_type` is the mapped type rather than a key/value pair, and its iterators dereference to the mapped value |
| `#!cpp std::multimap`, `#!cpp std::unordered_multimap` | `emplace` does not return `#!cpp std::pair<iterator, bool>` |
## `ArrayType`
`ArrayType` is instantiated as
```cpp
using array_t = ArrayType<basic_json, AllocatorType<basic_json>>;
```
### Always required
- The template must be usable with **two** type arguments (value type and allocator).
- Member types `value_type` and `iterator`.
- Constructors: default, copy, and move; and from an iterator range `(first, last)`.
- Member functions `begin()`, `end()`, `cbegin()`, `cend()`, `empty()`, `size()`, `max_size()`, `clear()`,
`operator[](size_type)`, `back()`, `push_back()`, `emplace_back()`, `pop_back()`, `resize()`,
`insert()` (single element, count, and range), `erase(pos)`, and `erase(first, last)`.
`basic_json::insert(pos, initializer_list)` goes through the range overload, so no initializer-list `insert` is
needed. `at(size_type)` is **not** required: [`basic_json::at(size_type)`](../../api/basic_json/at.md) checks the
index itself and then uses `operator[]`.
- `iterator` must be default-constructible, and it as well as the type returned by `cbegin()`/`cend()` must satisfy
[LegacyRandomAccessIterator](https://en.cppreference.com/w/cpp/named_req/RandomAccessIterator).
A `#!cpp static_assert` only checks for
[LegacyBidirectionalIterator](https://en.cppreference.com/w/cpp/named_req/BidirectionalIterator), but
[`dump`](../../api/basic_json/dump.md) (`cend() - 1`),
[`erase(idx)`](../../api/basic_json/erase.md) (`begin() + idx`), and the random-access operations of
[`basic_json::iterator`](../../api/basic_json/begin.md) require random access.
- The comparison operators, as for [`ObjectType`](#objecttype): `==` and `<`, or `==` and `<=>` under C++20.
### Required for individual functions
- A member type `value_type`, for [`to_bson`](../../api/basic_json/to_bson.md) of an array.
- A constructor from `(count, value)`, for
[`basic_json(size_type, const basic_json&)`](../../api/basic_json/basic_json.md).
- Swappability, via `#!cpp std::swap` or an ADL `swap`, for [`swap(array_t&)`](../../api/basic_json/swap.md).
!!! note "`capacity()` is optional"
With [`JSON_DIAGNOSTICS`](../../api/macros/json_diagnostics.md) enabled, the library reads `array_t::capacity()`
to find out whether adding an element reallocated the array and moved its elements, which would invalidate the
parent pointers. An array type without a `capacity()` member function is handled conservatively: the parent
pointers of all elements are refreshed after every insertion, which makes adding *n* elements cost O(*n*²). Only
diagnostics builds pay this; without them `capacity()` is never called.
### Compatible containers
| Container | Notes |
|-------------------------------------------------------------------------------------------|----------------------------------------------------------|
| `#!cpp std::vector` (default) | |
| `#!cpp std::deque` | references survive appends, but not insertions elsewhere; see the `capacity()` note above |
| `#!cpp std::pmr::vector` | through an alias, as the allocator comes from `AllocatorType` instead |
| `boost::container::vector`, `deque`, `devector` | |
| `boost::container::stable_vector` | the only one tried that keeps references valid across *every* insertion |
| `boost::container::small_vector`, `folly::small_vector` | through an alias that fixes the inline capacity |
| `boost::container::static_vector` | through the same kind of alias, for arrays that stay within the fixed capacity |
| `folly::fbvector` | requires C++20, see the note above |
### Containers that cannot be used
| Container | Reason |
|----------------------------------|----------------------------------------------------------------------------------------------|
| `#!cpp std::list` | no `operator[]`, and no random-access iterators |
| `eastl::vector`, `QList`, `QVector` | no `max_size()`; they handle the incomplete value type fine |
| `absl::InlinedVector` | requires a complete value type, see the note above |
| `absl::FixedArray` | the size is fixed at construction, so `resize`, `push_back`, `insert` and `erase` are missing |
## `StringType`
`StringType` is used **both** for JSON string values and for the keys of JSON objects
(`string_t` and `object_t::key_type`).
### Always required
- A member type `value_type` that is one byte wide and `char`-compatible. The library stores and processes UTF-8
encoded `char` data and hands `data()` to `#!cpp std::strtoull`/`#!cpp std::strtoll`.
`#!cpp std::wstring`, `#!cpp std::u16string`, and `#!cpp std::u32string` are **not** valid choices; see the FAQ on
[wide string handling](../../home/faq.md#wide-string-handling).
- Constructors: default, copy, move, from `#!cpp const char*` (which must not be `#!cpp explicit`), from
`#!cpp (const char*, size_type)`, and from `#!cpp (size_type, char)`; and copy or move assignment.
- Member functions `size()`, `clear()`, `resize(n, c)`, `data()`, `push_back(char)`, and `operator[]`
(const and non-const, returning references). `c_str()` and `back()` are **not** required.
- `data()` must return a pointer to a contiguous, **null-terminated** buffer -- the parser hands it to
`#!cpp std::strtoull`. A type whose `data()` is not null-terminated does not fail to compile; it silently
misparses numbers.
- `append(const char*, size_type)`, used by [`dump`](../../api/basic_json/dump.md), and `append(const StringType&)`,
used by the CBOR reader for indefinite-length strings. The library's internal string concatenation additionally has
to append a `#!cpp char` and a `#!cpp const char*`; for each it selects between `append(arg)`, `#!cpp operator+=`,
`append(first, last)`, and `append(data, size)`.
- The comparison operator `==` against another `StringType`, and `<` for use as a key of the chosen
[`ObjectType`](#objecttype) (with the default comparator, `#!cpp std::less<>` must be able to compare two
`StringType` values, and a `StringType` with the key types used for lookup). `!=` is never applied to a
`StringType`, and `==` against `#!cpp const char*` is resolved by the implicit `#!cpp const char*` constructor.
### Required for the binary formats
- `resize(n)`, used by the readers to make room for a block of bytes.
- Non-const `operator[]`, into which the readers `#!cpp std::memcpy` those bytes.
### Required for JSON Pointer, `flatten`, and `diff`
- A static member `npos` and the member function `find_first_of(char, size_type)` -- together with `data()`,
`reserve(n)`, and `append(const char*, size_type)` they implement the escaping and unescaping of reference tokens
described in RFC 6901. Neither `find(const StringType&, size_type)`, nor `substr(pos, count)`, nor
`replace(pos, count, const StringType&)` is required.
- `empty()`.
- `begin()` and `end()` -- used by
[`operator[](const json_pointer&)`](../../api/basic_json/operator%5B%5D.md) to decide whether a reference token
denotes an array index.
### Required for other functionality
| Functionality | Additional requirement |
|--------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------|
| [`diff`](../../api/basic_json/diff.md), [`items`](../../api/basic_json/items.md), [`std::hash`](../../api/basic_json/std_hash.md) | conversion of a `#!cpp std::size_t` to `StringType`: either assignability from the result of `#!cpp std::to_string`, or an ADL overload `#!cpp void int_to_string(StringType&, std::size_t)` |
| [`std::hash<basic_json>`](../../api/basic_json/std_hash.md) | additionally a specialization of `#!cpp std::hash<StringType>` |
| [`to_bson`](../../api/basic_json/to_bson.md) | `find(value_type)` and `npos` |
| [`parse`](../../api/basic_json/parse.md) from a `string_t` | the input adapters must accept it; otherwise pass a character range |
| `#!cpp operator<<(std::ostream&, const json_pointer&)` | streamability to `#!cpp std::ostream` |
| exception messages | `data()` and `size()`, or `begin()` and `end()` |
### Compatible types
| Type | Notes |
|-------------------------------------------------------------------|---------------------------------------------------------------------------|
| `#!cpp std::string` (default) | |
| `#!cpp std::basic_string` with a custom **stateless** allocator | |
| `#!cpp std::pmr::string` | see the warning below before relying on the memory resource |
| `boost::container::string` | needs a user-supplied `#!cpp std::hash` specialization (Boost provides `boost::hash` instead) |
| `folly::fbstring` | requires C++20, see the note above |
| `eastl::string` | needs a user-supplied `#!cpp std::hash` and an ADL `int_to_string` (it is not assignable from a `#!cpp std::string`); [`parse`](../../api/basic_json/parse.md) does not accept it directly -- pass a character range or a `#!cpp std::string` |
| a custom string class in a user-defined namespace | if the requirements above are met |
### Types that cannot be used
| Type | Reason |
|-----------------------------------------------------------------------|-----------------------------------------------------------------------|
| `#!cpp std::wstring`, `#!cpp std::u16string`, `#!cpp std::u32string` | the character type is not one byte wide |
| `#!cpp std::u8string` | one byte wide, but `#!cpp char8_t` is not `#!cpp char`-compatible |
| `absl::Cord` | no `value_type`, and the storage is not contiguous |
| `QString` | no `append(const char*, size_type)`; its `QChar` is also two bytes wide, though that is never diagnosed |
!!! warning "A `std::pmr::string` mostly does not use the memory resource you choose"
`basic_json` cannot be given an allocator or a memory resource. `AllocatorType` is default-constructed at every
allocation and has to be stateless (see [`AllocatorType`](#allocatortype)), and string values the library creates
are constructed with their own default allocator. So:
- Every string the library itself produces -- from [`parse`](../../api/basic_json/parse.md), from
[`dump`](../../api/basic_json/dump.md), or by default construction -- allocates from
`#!cpp std::pmr::get_default_resource()`.
- **Copying** an arena-backed string into a value silently drops its memory resource: the copy lands on the
default resource, because `#!cpp std::pmr::polymorphic_allocator` does not propagate on copy construction.
Nothing warns about this.
- **Moving** one in does keep it, and later growth still allocates from that arena -- but it does not survive a
copy of the enclosing `basic_json`.
- Passing `#!cpp std::pmr::polymorphic_allocator` as `AllocatorType` does not work around any of this; it does
not compile.
Apart from moving a string in, the only way to redirect these allocations is the process-global
`#!cpp std::pmr::set_default_resource()`.
!!! tip "Reference implementation"
The unit test `tests/src/unit-alt-string.cpp` contains `alt_string`, a minimal string type that satisfies the
requirements needed for the tested subset of the API. It is a good starting point for a custom `StringType`.
## `BooleanType`
`boolean_t` is stored **directly** inside `basic_json`, as a member of an anonymous union.
### Always required
- A literal type that is trivially default-constructible, trivially copyable, and trivially destructible; otherwise the
union's special member functions are deleted.
- **Implicitly** convertible from `#!cpp bool` -- an `#!cpp explicit` constructor is not enough, because the
`to_json` overload for a custom `BooleanType` is constrained on `#!cpp std::is_convertible` -- and contextually
convertible to `#!cpp bool` (here an `#!cpp explicit operator bool` is fine).
- Comparison operators `==`, `!=`, `<`, `<=`, `>`, `>=` (or `<=>`).
- Convertible from and to `#!cpp bool` through the serializer, because
[`get<bool>()`](../../api/basic_json/get.md) is used internally.
There is little reason to use anything other than `#!cpp bool` here.
### Compatible types
`#!cpp bool` is the only usable choice. Another trivially copyable type that is implicitly convertible to and from
`#!cpp bool` -- `#!cpp std::uint8_t`, say -- does compile, and JSON booleans still round-trip, but the type then
serves as both `boolean_t` and an ordinary integer: `basic_json` can no longer be constructed or assigned from a
`#!cpp std::uint8_t` at all (the boolean and unsigned-integer `to_json` overloads become ambiguous), and
[`get<std::uint8_t>()`](../../api/basic_json/get.md) on a number throws
[`type_error.302`](../../home/exceptions.md#jsonexceptiontype_error302) instead of returning the value.
## `NumberIntegerType` and `NumberUnsignedType`
Both types are stored **directly** inside `basic_json`'s union.
### Always required
- `#!cpp std::is_integral` must be satisfied: `NumberIntegerType` must be a **signed** integer type,
`NumberUnsignedType` an **unsigned** integer type. Class types are not supported -- among others, the constructors
taking integer values are constrained on `#!cpp std::is_integral`.
- Trivially default-constructible, trivially copyable, and trivially destructible (union member).
- `#!cpp std::numeric_limits` must be specialized for both types.
- `NumberUnsignedType` must be able to represent the absolute value of every `NumberIntegerType` value; serialization
of negative numbers converts the value to `NumberUnsignedType`.
- Both types must fit into the internal 64-character number buffer used by
[`dump`](../../api/basic_json/dump.md), which is the case for all standard integer types.
- [`std::hash<basic_json>`](../../api/basic_json/std_hash.md) additionally requires `#!cpp std::hash` specializations.
### Notes
The number types influence what the parser accepts: an integer literal that does not round-trip through the chosen type
is stored as [`number_float_t`](../../api/basic_json/number_float_t.md) instead. Choosing types narrower than 64 bits
therefore silently changes parse results rather than raising an error. See
[Number Handling](number_handling.md) for details.
### Compatible types
| Type pair | Support |
|----------------------------------------------------------------------------------------------|---------------------------------------------------------------------|
| `#!cpp std::int64_t` / `#!cpp std::uint64_t` (default) | full |
| `#!cpp std::int32_t` / `#!cpp std::uint32_t`, `#!cpp long long` / `#!cpp unsigned long long` | full; narrower types change which literals the parser can represent |
| any other pair of standard signed/unsigned integer types | full |
| class types, enumerations | not usable; `#!cpp std::is_integral` must hold |
| `#!cpp bool`, or a type already used for another member of the union | not usable; `#!cpp std::is_integral<bool>` is in fact `#!cpp true`, but the `get_impl_ptr` overloads for `boolean_t`, `number_integer_t`, `number_unsigned_t` and `number_float_t` would collide |
## `NumberFloatType`
`number_float_t` is stored **directly** inside `basic_json`'s union.
### Always required
- Trivially default-constructible, trivially copyable, and trivially destructible (union member).
- `#!cpp std::numeric_limits` must be specialized; `max_digits10` is used to size the conversion.
- `#!cpp std::isfinite` must be applicable to the type.
### Required for parsing and serialization
`NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`:
- The [parser](../parsing/index.md) converts number literals with `#!cpp std::strtof`, `#!cpp std::strtod`, or
`#!cpp std::strtold`; the library provides overloads for exactly these three types.
- [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
(`#!cpp float` is promoted to `#!cpp double`).
If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary32 or binary64 number, `dump` uses the
Grisu2 algorithm, which produces the shortest representation that round-trips. Otherwise the `snprintf` fallback with
`max_digits10` digits is used.
### Required for the binary formats
`NumberFloatType` must be `#!cpp float` or `#!cpp double`. The writers for
[CBOR, MessagePack, UBJSON, BJData, and BSON](../binary_formats/index.md) map a floating-point value onto an IEEE 754
binary32 or binary64 field and have no encoding for `#!cpp long double`.
### Compatible types
| Type | Support |
|-----------------------------|-----------------------------------------------------------------------------------------------------|
| `#!cpp double` (default) | full; short round-trip output through Grisu2 |
| `#!cpp float` | full; short round-trip output through Grisu2 |
| `#!cpp long double` | `dump` and `parse` only; the binary format writers do not compile, as they only handle IEEE 754 binary32 and binary64 |
| any other type | not usable |
## `AllocatorType`
`AllocatorType` is instantiated with **one** argument, for each of `object_t`, `array_t`, `string_t`, `binary_t`,
`basic_json`, and `#!cpp std::pair<const StringType, basic_json>`.
### Always required
- The template must be usable with exactly one type argument. The library instantiates `AllocatorType<T>` directly and
never uses `#!cpp std::allocator_traits<...>::rebind_alloc`.
- It must satisfy the [Allocator](https://en.cppreference.com/w/cpp/named_req/Allocator) named requirement so that
`#!cpp std::allocator_traits` can be used with it.
- It must be **default-constructible and stateless**. Objects are allocated with a default-constructed allocator and
deallocated with a *different* default-constructed allocator, and
[`get_allocator()`](../../api/basic_json/get_allocator.md) returns a default-constructed instance. Allocators
carrying state are not supported, so there is no way to tell a `basic_json` where to allocate from; see the note
under [`StringType`](#stringtype) for what that means in practice. A stateful allocator is **not diagnosed**: it
compiles and silently ignores the state.
- It must support **incomplete types**: `AllocatorType<basic_json>` is instantiated inside the definition of
`basic_json` itself.
- `#!cpp std::allocator_traits<AllocatorType<basic_json>>::pointer` becomes
[`basic_json::pointer`](../../api/basic_json/index.md#container-types), and iterators are constructed from raw
`#!cpp basic_json*` values. The `pointer` type must therefore be a plain pointer; fancy pointers are not supported.
### Compatible types
| Type | Support |
|-----------------------------------------------------------------|----------------------------------------------------|
| `#!cpp std::allocator` (default) | full |
| a custom stateless allocator template | full |
| stateful allocators, e.g. `#!cpp std::pmr::polymorphic_allocator`| not usable; see the requirements above |
## `JSONSerializer`
`JSONSerializer` is instantiated as `JSONSerializer<T, void>` and defaults to
[`adl_serializer`](../../api/adl_serializer/index.md).
### Always required
- The template must accept **two** type arguments. It does not have to give the second one a default -- `basic_json`
declares the parameter as `#!cpp template<typename T, typename SFINAE = void> class JSONSerializer`, so uses such as
`#!cpp JSONSerializer<T>` inside the library supply `#!cpp void` themselves. The second parameter exists so that
partial specializations can be constrained by SFINAE.
- For every type `T` that is converted **to** a JSON value, a static member function
`#!cpp static void to_json(basic_json&, T)` must exist.
- For every type `T` that is converted **from** a JSON value, either
`#!cpp static void from_json(const basic_json&, T&)` or `#!cpp static T from_json(const basic_json&)` must exist.
The latter form is required for types that are not default-constructible; see
[Arbitrary Types Conversions](../arbitrary_types.md).
- To support the [converting constructor](../../api/basic_json/basic_json.md) between different `basic_json`
specializations, `to_json` must be available for `boolean_t`, `number_integer_t`, `number_unsigned_t`,
`number_float_t`, `string_t`, `object_t`, `array_t`, and `binary_t` of the *source* specialization.
### Compatible types
| Type | Support |
|-------------------------------------------------------------------|-------------------------------------------------------------------------|
| [`nlohmann::adl_serializer`](../../api/adl_serializer/index.md) (default) | full |
| a class template deriving from `adl_serializer` | full; the usual way to change behavior while keeping the defaults |
| an unrelated template with the same interface | full, but it has to handle every type the library converts |
## `BinaryType`
`BinaryType` is not a JSON type; it is used for the byte strings of the
[binary formats](../binary_formats/index.md). It is wrapped as
```cpp
using binary_t = nlohmann::byte_container_with_subtype<BinaryType>;
```
### Always required
- A non-`final` class type -- [`byte_container_with_subtype`](../../api/byte_container_with_subtype/index.md) derives
from it publicly.
- A member type `value_type` that is **exactly one byte** wide (e.g., `#!cpp std::uint8_t`, `#!cpp char`, or
`#!cpp std::byte`). Readers and writers reinterpret the container's storage as raw bytes. A wider `value_type` is
**not diagnosed**: it compiles and silently produces wrong results.
- Contiguous storage: the binary readers `#!cpp std::memcpy` into `#!cpp &binary[n]`, the writers `reinterpret_cast`
`data()`.
- Default-constructible, copy-constructible, and move-constructible.
- Member functions `size()`, `empty()`, `data()`, `resize()`, `operator[]`, `back()`, `begin()`, `end()`, `cbegin()`,
and `cend()` with random-access iterators, and `insert(pos, first, last)`, which the CBOR reader uses to join the
chunks of an indefinite-length byte string. `push_back()` is **not** required.
- Comparison operators: `==` is used by
[`byte_container_with_subtype`](../../api/byte_container_with_subtype/index.md), the relational operators by
[`basic_json`'s comparison operators](../../api/basic_json/operator_le.md).
### Required for individual functions
- `clear()`, for [`basic_json::clear()`](../../api/basic_json/clear.md).
`max_size()`, `at()`, `reserve()`, `erase()`, `pop_back()`, and `emplace_back()` are **not** used at all.
See [`binary_t`](../../api/basic_json/binary_t.md) for how a non-default `BinaryType` changes the meaning of assigning
such a container to a `basic_json` value.
### Compatible containers
| Container | Notes |
|------------------------------------------------------------------------------------------------------------------------------|--------------------------------------------------------------------|
| `#!cpp std::vector<std::uint8_t>` (default) | |
| `#!cpp std::vector<char>`, `#!cpp std::vector<std::byte>` | `dump()` writes the bytes as 0..255 whichever is used |
| `boost::container::vector<std::uint8_t>`, `boost::container::small_vector<std::uint8_t, N>` | |
| `absl::InlinedVector<std::uint8_t, N>` | usable here, unlike as an `ArrayType`, because the value type is complete |
| `eastl::vector<std::uint8_t>` | usable here, unlike as an `ArrayType`, because `max_size()` is not needed |
| `folly::fbvector<std::uint8_t>` | requires C++20, see the note above |
### Containers that cannot be used
| Container | Reason |
|--------------------------------------------------------|-----------------------------------------------------------------------------------------|
| `QByteArray` | no `empty()` (it spells that `isEmpty()`); its `insert` takes an index rather than an iterator; and it converts to `string_t`, which makes `to_json` ambiguous between a string and a binary value |
| `#!cpp std::string` | `binary_t::container_type` and `string_t` would be the same type, so the two [`swap`](../../api/basic_json/swap.md) overloads collide and `basic_json` cannot be instantiated at all |
| `#!cpp std::deque<std::uint8_t>` | storage is not contiguous, so there is no `data()` |
| containers whose `value_type` is wider than one byte | see above -- accepted by the compiler, wrong at runtime |
## `CustomBaseClass`
`CustomBaseClass` is an extension point: unless it is `#!cpp void` (the default, which selects the empty
`nlohmann::json_default_base`), `basic_json` publicly derives from it.
### Always required
- A non-`final`, default-constructible class type.
- `basic_json` is copy-/move-constructible and copy-/move-assignable only if `CustomBaseClass` is.
### Notes
`basic_json` is documented to be a
[StandardLayoutType](https://en.cppreference.com/w/cpp/named_req/StandardLayoutType). Because `basic_json` has
non-static data members of its own, a `CustomBaseClass` with non-static data members forfeits this guarantee.
Note the namespace of `CustomBaseClass` becomes an associated namespace of `basic_json` for the purpose of
argument-dependent lookup.
See [`json_base_class_t`](../../api/basic_json/json_base_class_t.md) for an example.
### Compatible types
| Type | Support |
|----------------------------------------------------------|--------------------------------------------------------------|
| `#!cpp void` (default) | an empty base class is used; no effect on `basic_json` |
| any default-constructible, non-`final` class | full; see [`json_base_class_t`](../../api/basic_json/json_base_class_t.md) |
## Cross-specialization conversions
Converting a value from one `basic_json` specialization into another (see the
[converting constructor](../../api/basic_json/basic_json.md)) imposes two additional requirements that are not
diagnosed at compile time. With assertions enabled they abort on the `#!cpp JSON_ASSERT` at the end of the converting
constructor; under `#!cpp NDEBUG` they fail **silently** at runtime:
- The target `string_t` must be directly constructible from the source `string_t`. Otherwise the string is converted to
an array of character codes.
- The target `object_t::key_type` must be directly constructible from the source object's key type. Otherwise the
object is converted to an array of key/value pairs.
See [issue #3425](https://github.com/nlohmann/json/issues/3425), [`string_t`](../../api/basic_json/string_t.md), and
[`object_t`](../../api/basic_json/object_t.md).
## See also
- [Types](index.md) -- overview of how JSON values are stored
- [Number Handling](number_handling.md) -- how the number types affect parsing and serialization
- [Object Order](../object_order.md) -- using an insertion-ordered `ObjectType`
- [`basic_json`](../../api/basic_json/index.md) -- API documentation of the class template
+16
View File
@@ -965,3 +965,19 @@ A JSON Patch operation 'test' failed. The unsuccessful operation is also printed
```
[json.exception.other_error.501] unsuccessful: {"op":"test","path":"/baz","value":"bar"}
```
### json.exception.other_error.502
[`to_ubjson`](../api/basic_json/to_ubjson.md) and [`to_bjdata`](../api/basic_json/to_bjdata.md) were called with
`use_type = true` but `use_size = false`. UBJSON requires a size marker (`#`) after a type marker (`$`).
!!! failure "Example message"
```
[json.exception.other_error.502] use_type requires use_size = true
```
!!! note
This exception was added in version 3.13.0. Before that, debug builds aborted on an assertion and release builds
wrote a `$` marker without `#`, which [`from_ubjson`](../api/basic_json/from_ubjson.md) then rejected.
+1 -1
View File
@@ -98,6 +98,7 @@ nav:
- Types:
- features/types/index.md
- features/types/number_handling.md
- features/types/template_parameters.md
- Integration:
- integration/index.md
- integration/migration_guide.md
@@ -296,7 +297,6 @@ nav:
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
- 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md
- 'JSON_USE_SIMDUTF': api/macros/json_use_simdutf.md
- 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md
- 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md
- 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md
+1 -1
View File
@@ -1,4 +1,4 @@
wheel==0.47.0
wheel==0.48.0
mkdocs==1.6.1 # documentation framework
mkdocs-git-revision-date-localized-plugin==1.5.3 # plugin "git-revision-date-localized"
+4 -1
View File
@@ -101,7 +101,10 @@ class exception : public std::exception
{
if (&element.second == current)
{
tokens.emplace_back(element.first.c_str());
// data() is null-terminated, so a key containing
// a null byte is cut short here rather than
// truncating the whole message at what()
tokens.emplace_back(element.first.data());
break;
}
}
+3 -1
View File
@@ -114,7 +114,9 @@ std::size_t hash(const BasicJsonType& j)
seed = combine(seed, static_cast<std::size_t>(j.get_binary().subtype()));
for (const auto byte : j.get_binary())
{
seed = combine(seed, std::hash<std::uint8_t> {}(byte));
// the cast is needed for binary types whose value type is not
// an integer (e.g., std::byte)
seed = combine(seed, std::hash<std::uint8_t> {}(static_cast<std::uint8_t>(byte)));
}
return seed;
}
@@ -773,7 +773,13 @@ class binary_reader
case 0xBF: // map (indefinite length)
return get_cbor_object(detail::unknown_size(), tag_handler);
case 0xC6: // tagged item
case 0xC0: // tagged item
case 0xC1:
case 0xC2:
case 0xC3:
case 0xC4:
case 0xC5:
case 0xC6:
case 0xC7:
case 0xC8:
case 0xC9:
@@ -788,6 +794,9 @@ class binary_reader
case 0xD2:
case 0xD3:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8: // tagged item (1 byte follows)
case 0xD9: // tagged item (2 bytes follow)
case 0xDA: // tagged item (4 bytes follow)
@@ -2890,7 +2899,10 @@ class binary_reader
number_string,
out_of_range::create(406, concat("number overflow parsing '", number_string, '\''), nullptr));
}
return sax->number_float(parsed_float, std::move(number_string));
// number_string is a std::string, while the SAX interface takes a
// string_t; convert explicitly, as the two are only implicitly
// convertible for some string types
return sax->number_float(parsed_float, string_t(number_string.data(), number_string.size()));
}
case token_type::uninitialized:
case token_type::literal_true:
+21 -131
View File
@@ -155,31 +155,11 @@ class input_stream_adapter
// General-purpose iterator-based adapter. It might not be as fast as
// theoretically possible for some containers, but it is extremely versatile.
// SentinelType defaults to IteratorType for backward compatibility, but may be
// a different type, e.g. a C++20 sentinel such as std::default_sentinel_t when
// IteratorType is a std::counted_iterator.
// SentinelType defaults to IteratorType for backward compatibility, but may
// be a different type (e.g., a C++20 sentinel or counted_iterator).
template<typename IteratorType, typename SentinelType = IteratorType>
class iterator_input_adapter
{
// Whether the number of elements between two positions can be computed in
// O(1): either the iterator and the sentinel have the same type (plain
// std::distance) or, in C++20, the sentinel is a sized sentinel for the
// iterator (std::ranges::distance), e.g. std::default_sentinel_t paired
// with std::counted_iterator.
//
// JSON_HAS_RANGES gates the C++20 branch: on standard libraries with an
// incomplete <ranges> (libstdc++ < 11, see #4440) evaluating
// std::contiguous_iterator on a std::counted_iterator is a hard error
// instead of yielding false, and these traits are instantiated for every
// adapter. Such toolchains fall back to the pointer-only test and simply
// use the byte-at-a-time scanner.
static constexpr bool sentinel_is_sized =
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>;
#else
std::is_same<IteratorType, SentinelType>::value;
#endif
public:
using char_type = typename std::iterator_traits<IteratorType>::value_type;
@@ -191,7 +171,7 @@ class iterator_input_adapter
// in wide_string_input_adapter, which does not expose this).
static constexpr bool supports_seek =
std::is_same<typename std::iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value
&& sentinel_is_sized
&& std::is_same<IteratorType, SentinelType>::value
&& sizeof(char_type) == 1;
iterator_input_adapter(IteratorType first, SentinelType last)
@@ -239,60 +219,30 @@ class iterator_input_adapter
private:
// whether IteratorType refers to a contiguous range and therefore supports
// a std::memcpy fast path (pointers always do; in C++20 we can also detect
// library iterators such as those of std::vector and std::string). The
// available element count must also be computable in O(1), hence
// sentinel_is_sized.
static constexpr bool iterator_is_contiguous = sentinel_is_sized &&
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
// library iterators such as those of std::vector and std::string).
// Computing the available element count needs either same-type iterators
// (plain std::distance) or, in C++20, a sized sentinel (std::ranges::distance),
// e.g. std::counted_iterator paired with std::default_sentinel_t.
static constexpr bool iterator_is_contiguous =
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>)
&& (std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
#else
std::is_pointer<IteratorType>::value;
std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
#endif
// number of unread elements in [current, end)
std::size_t remaining_count() const
{
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
return static_cast<std::size_t>(std::ranges::distance(current, end));
#else
return static_cast<std::size_t>(std::distance(current, end));
#endif
}
public:
// Whether the remaining input is a single contiguous block of 1-byte
// elements that the lexer can inspect directly (used for the SWAR string
// fast path).
static constexpr bool supports_bulk_scan =
iterator_is_contiguous && sizeof(char_type) == 1;
// Pointer to the next unread element; only valid when bulk_remaining() > 0.
const char_type* bulk_data() const
{
return &*current;
}
// Number of unread elements available as one contiguous block.
std::size_t bulk_remaining() const
{
return remaining_count();
}
// Consume @a n elements previously inspected via bulk_data().
void bulk_skip(std::size_t n)
{
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(n));
}
private:
// contiguous fast path: bulk copy the remaining range with std::memcpy
template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
{
const std::size_t wanted = count * sizeof(T);
const std::size_t available = remaining_count() * sizeof(char_type);
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
const std::size_t available = static_cast<std::size_t>(std::ranges::distance(current, end)) * sizeof(char_type);
#else
const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type);
#endif
const std::size_t copied = (std::min)(wanted, available);
if (JSON_HEDLEY_LIKELY(copied != 0))
{
@@ -620,46 +570,6 @@ typename iterator_input_adapter_factory<IteratorType, SentinelType>::adapter_typ
return factory_type::create(first, last);
}
// The element type a container's data() points at, cv-qualifiers removed.
// Ill-formed - and therefore SFINAE-friendly - for types without data().
template<typename ContainerType>
using container_data_t = typename std::remove_cv<typename std::remove_pointer <
decltype(std::declval<const ContainerType&>().data()) >::type >::type;
// The container's own element type, cv-qualifiers removed. It is looked up on
// the bare type so it is also found when ContainerType is deduced as a
// reference by the forwarding-reference overload below.
template<typename ContainerType>
using container_value_t = typename std::remove_cv <
typename std::remove_cv<typename std::remove_reference<ContainerType>::type>::type::value_type >::type;
// Detect a container that stores its elements contiguously as single bytes
// (std::string, std::vector<char/unsigned char>, std::array<char, N>,
// std::string_view, ...). Such inputs are wrapped in a pointer-based adapter so
// they benefit from the contiguous fast paths (bulk string scanning, memcpy for
// binary formats) in every C++ standard - not only in C++20, where the standard
// library iterators model std::contiguous_iterator and are detected directly.
//
// data() and size() on their own would be duck typing: they say nothing about
// size() counting the units data() points at, and reading [data(), data() +
// size()) as bytes would be wrong for a type where it does not. Requiring the
// container's own value_type to be that same single-byte element ties the two
// together; every contiguous standard container satisfies it. Anything else
// keeps the iterator-based adapter, which is always correct - only slower.
template<typename ContainerType, typename = void>
struct is_contiguous_byte_container : std::false_type {};
template<typename ContainerType>
struct is_contiguous_byte_container < ContainerType, void_t <
container_data_t<ContainerType>,
container_value_t<ContainerType>,
decltype(std::declval<const ContainerType&>().size()) >>
: std::integral_constant < bool,
std::is_pointer<decltype(std::declval<const ContainerType&>().data())>::value&&
std::is_integral<container_data_t<ContainerType>>::value&&
sizeof(container_data_t<ContainerType>) == 1 &&
std::is_same<container_data_t<ContainerType>, container_value_t<ContainerType>>::value > {};
// Convenience shorthand from container to iterator
// Enables ADL on begin(container) and end(container)
// Encloses the using declarations in namespace for not to leak them to outside scope
@@ -687,32 +597,12 @@ struct container_input_adapter_factory< ContainerType,
} // namespace container_input_adapter_factory_impl
// General container path (iterator-based). Contiguous single-byte containers
// are excluded here and routed through the pointer-based overload below.
template < typename ContainerType,
enable_if_t < !is_contiguous_byte_container<ContainerType>::value, int > = 0 >
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType && container)
template<typename ContainerType>
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType&& container)
{
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
}
// Contiguous single-byte containers (std::string, std::vector<char>, ...) are
// wrapped in a pointer-based adapter so the contiguous fast paths apply in every
// standard. The pointer keeps the container's own element type (const char* for
// std::string, const std::uint8_t* for std::vector<std::uint8_t>, ...), so the
// resulting char_type - and therefore the parsing behavior - is byte-for-byte
// identical to the iterator-based path; only the raw pointer additionally
// enables the bulk fast paths. The container outlives the adapter for the whole
// parse (temporaries live until the end of the full expression), exactly as the
// iterators it replaces did.
template < typename ContainerType,
enable_if_t < is_contiguous_byte_container<ContainerType>::value, int > = 0 >
auto input_adapter(const ContainerType& container)
-> decltype(input_adapter(container.data(), container.data() + container.size()))
{
return input_adapter(container.data(), container.data() + container.size());
}
// specialization for std::string
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
+27 -292
View File
@@ -19,9 +19,7 @@
#include <vector> // vector
#include <nlohmann/detail/input/input_adapters.hpp>
#include <nlohmann/detail/input/number_parse.hpp>
#include <nlohmann/detail/input/position_t.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
@@ -127,25 +125,6 @@ constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
return false;
}
// Detect whether an input adapter exposes a contiguous byte block that the
// lexer can scan directly (see iterator_input_adapter::supports_bulk_scan).
// Adapters without the flag - file, stream, wide-string, user-defined - fall
// back to the character-at-a-time string scanner.
template<typename InputAdapterType>
using detect_supports_bulk_scan = decltype(InputAdapterType::supports_bulk_scan);
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::true_type /*detected*/)
{
return InputAdapterType::supports_bulk_scan;
}
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::false_type /*detected*/)
{
return false;
}
/*!
@brief lexical analysis
@@ -167,14 +146,6 @@ class lexer : public lexer_base<BasicJsonType>
static constexpr bool lazy_token_string =
input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {});
/// whether string scanning may bulk-consume runs of ordinary characters
/// directly from a contiguous input buffer (SWAR fast path). This requires
/// the token to be reconstructible lazily (lazy_token_string), so bypassing
/// the per-character capture in get() cannot lose error diagnostics.
static constexpr bool bulk_scan =
lazy_token_string
&& input_adapter_supports_bulk_scan<InputAdapterType>(is_detected<detect_supports_bulk_scan, InputAdapterType> {});
public:
using token_type = typename lexer_base<BasicJsonType>::token_type;
@@ -294,40 +265,6 @@ class lexer : public lexer_base<BasicJsonType>
return true;
}
/// contiguous input: bulk-append the run of ordinary characters and complete
/// well-formed UTF-8 sequences starting at the current read position, leaving
/// the first byte that needs individual handling (the closing quote, an
/// escape, a control character, or an ill-formed UTF-8 byte) for get()
void scan_string_bulk(std::true_type /*bulk*/)
{
// a pending unget must be consumed through the normal path first
if (next_unget)
{
return;
}
const std::size_t remaining = ia.bulk_remaining();
if (remaining == 0)
{
return;
}
const auto* const data = reinterpret_cast<const unsigned char*>(ia.bulk_data());
const std::size_t pos = string_bulk_run(data, remaining);
if (pos == 0)
{
return;
}
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), pos);
ia.bulk_skip(pos);
// the run contains no newline (all bytes < 0x20 are treated as special),
// so only the flat character counters advance
position.chars_read_total += pos;
position.chars_read_current_line += pos;
}
/// streaming input: no bulk fast path
void scan_string_bulk(std::false_type /*bulk*/) const noexcept {}
/*!
@brief scan a string literal
@@ -353,10 +290,6 @@ class lexer : public lexer_base<BasicJsonType>
while (true)
{
// bulk-consume ordinary characters from contiguous input, then
// handle the next special byte through the switch below
scan_string_bulk(std::integral_constant<bool, bulk_scan> {});
// get the next character
switch (get())
{
@@ -1346,78 +1279,45 @@ scan_number_done:
// we are done scanning a number)
unget();
return convert_number(number_type);
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
errno = 0;
/*!
@brief convert an already-validated integer token to its value
The digit sequence in [first, last) has been validated by the caller, so a
dedicated parser can avoid the locale/errno overhead of std::strtoull.
@return the token type on success; token_type::uninitialized if @a
number_type is not an integer type or the value does not fit, in
which case the caller falls back to the floating-point conversion
(matching the previous std::strtoull/std::strtoll behavior)
*/
token_type convert_integer(token_type number_type, const char* first, const char* last)
{
// try to parse integers first and fall back to floats
if (number_type == token_type::value_unsigned)
{
if (parse_integer_unsigned(first, last, value_unsigned))
const auto x = std::strtoull(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
{
return token_type::value_unsigned;
value_unsigned = static_cast<number_unsigned_t>(x);
if (value_unsigned == x)
{
return token_type::value_unsigned;
}
}
}
else if (number_type == token_type::value_integer)
{
if (parse_integer_signed(first, last, value_integer))
const auto x = std::strtoll(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
{
return token_type::value_integer;
}
}
return token_type::uninitialized;
}
/*!
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
*/
token_type convert_number(token_type number_type)
{
const char* const num_begin = token_buffer.data();
const char* const num_end = num_begin + token_buffer.size();
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, num_begin, num_end);
if (integer_result != token_type::uninitialized)
{
return integer_result;
value_integer = static_cast<number_integer_t>(x);
if (value_integer == x)
{
return token_type::value_integer;
}
}
}
// this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
}
if (parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
{
return token_type::value_float;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
// integer conversion above failed
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
@@ -1426,156 +1326,6 @@ scan_number_done:
return token_type::value_float;
}
/*!
@brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
then produces the exact diagnostic. @a current is the first digit or the
leading minus (already read); the remaining bytes are taken from the adapter.
*/
token_type scan_number_bulk_contiguous()
{
// a pending unget offsets the buffer position from current; fall back
if (next_unget)
{
return token_type::uninitialized;
}
const std::size_t rem = ia.bulk_remaining();
if (rem == 0)
{
// the first digit is the last input byte; let scan_number() finish
return token_type::uninitialized;
}
// the byte before the next unread one is current (contiguous input)
const char* const data = reinterpret_cast<const char*>(ia.bulk_data()) - 1;
const std::size_t avail = rem + 1;
// validate + classify the number extent (mirrors scan_number()'s grammar)
std::size_t i = 0;
std::size_t dot_index = std::string::npos;
token_type number_type = token_type::value_unsigned;
if (data[0] == '-')
{
number_type = token_type::value_integer;
i = 1;
if (i >= avail)
{
return token_type::uninitialized;
}
}
if (data[i] == '0')
{
++i;
}
else if (data[i] >= '1' && data[i] <= '9')
{
++i;
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
else
{
return token_type::uninitialized;
}
if (i < avail && data[i] == '.')
{
number_type = token_type::value_float;
dot_index = i;
++i;
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
if (i < avail && (data[i] == 'e' || data[i] == 'E'))
{
number_type = token_type::value_float;
++i;
if (i < avail && (data[i] == '+' || data[i] == '-'))
{
++i;
}
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
const std::size_t len = i;
// reset() records where this token starts (for diagnostics), so it has
// to run before the input position advances below
reset();
// An integer token needs no token_buffer: the SAX callbacks for
// number_integer/number_unsigned take only the value, and the overflow
// diagnostic rebuilds the text from the input. Convert straight from the
// input buffer and leave token_buffer empty. (JSON_DIAGNOSTIC_POSITIONS
// derives a number's start position from get_string().size(), so there
// the token still has to be materialized.)
#if !JSON_DIAGNOSTIC_POSITIONS
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, data, data + len);
if (JSON_HEDLEY_LIKELY(integer_result != token_type::uninitialized))
{
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return integer_result;
}
// The value does not fit an integer, so this token converts as a
// float. Recording that here keeps convert_number() below from
// repeating the integer attempt that just failed.
number_type = token_type::value_float;
}
#endif
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return convert_number(number_type);
}
/// contiguous input: try the number fast path, else the byte-path scanner
token_type scan_number_dispatch(std::true_type /*bulk*/)
{
const token_type t = scan_number_bulk_contiguous();
return (t != token_type::uninitialized) ? t : scan_number();
}
/// streaming input: always use the byte-path scanner
token_type scan_number_dispatch(std::false_type /*bulk*/)
{
return scan_number();
}
/*!
@param[in] literal_text the literal text to expect
@param[in] length the length of the passed literal text
@@ -1663,9 +1413,6 @@ scan_number_done:
if (current == '\n')
{
++position.lines_read;
// remember the column the newline was read at: chars_read_current_line
// is about to be cleared, and a matching unget() cannot reconstruct it
chars_read_before_newline = position.chars_read_current_line;
position.chars_read_current_line = 0;
}
@@ -1699,20 +1446,12 @@ scan_number_done:
--position.chars_read_total;
// in case we "unget" a newline, we have to also decrement the lines_read
// and restore the column that get() cleared when it saw the newline;
// chars_read_current_line == 0 can only mean the last get() read one
if (position.chars_read_current_line == 0)
{
if (position.lines_read > 0)
{
--position.lines_read;
}
// chars_read_before_newline counts the newline itself, which is the
// character being ungotten, hence the -1
position.chars_read_current_line = (chars_read_before_newline > 0)
? chars_read_before_newline - 1
: 0;
}
else
{
@@ -1955,7 +1694,7 @@ scan_number_done:
case '7':
case '8':
case '9':
return scan_number_dispatch(std::integral_constant<bool, bulk_scan> {});
return scan_number();
// end of input (the null byte is needed when parsing from
// string literals)
@@ -1986,10 +1725,6 @@ scan_number_done:
/// the start position of the current token
position_t position {};
/// the value chars_read_current_line had when the last newline was read, so
/// that unget() can restore the column instead of leaving it at 0
std::size_t chars_read_before_newline = 0;
/// raw input token string for error messages; only populated for streaming
/// adapters (seekable adapters reconstruct it lazily via token_string_start)
std::vector<char_type> token_string {};
@@ -1,302 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint64_t
#include <limits> // numeric_limits
#include <nlohmann/detail/macro_scope.hpp>
// std::from_chars lives in <charconv>, but being in C++17 mode does not
// guarantee the header exists: GCC 7 sets __cplusplus to C++17 yet ships no
// <charconv> (added in GCC 8; floating-point support in GCC 11). Guard the
// include with __has_include so such toolchains fall back to the scalar path.
#if defined(JSON_HAS_CPP_17) && defined(__has_include)
#if __has_include(<charconv>)
#include <charconv> // from_chars (only used when __cpp_lib_to_chars is defined)
#include <system_error> // errc
#endif
#endif
// This file contains the value-conversion helpers used by the lexer to turn an
// already-validated number token into a value, without the locale/errno
// overhead of std::strtoull/std::strtod. They are free functions so the lexer
// stays focused on scanning; see lexer::convert_number().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief fast integer parser for an already-validated unsigned integer
The number scanner has already checked that [first, last) is a valid JSON
integer, so this only needs to accumulate the digits and detect overflow. This
avoids the locale/errno machinery of std::strtoull, which dominates
integer-heavy inputs.
@param[in] first pointer to the first character (a digit)
@param[in] last pointer past the last character
@param[out] value the parsed value on success
@return true if the value fit into @a NumberUnsignedType; false on overflow, in
which case the caller falls back to floating-point parsing (matching the
previous std::strtoull behavior)
*/
template<typename NumberUnsignedType>
bool parse_integer_unsigned(const char* first, const char* last, NumberUnsignedType& value) noexcept
{
// accumulate in the widest unsigned type used by the previous strtoull
// path so the overflow behavior is unchanged for custom number types
std::uint64_t x = 0;
constexpr std::uint64_t cutoff = (std::numeric_limits<std::uint64_t>::max)() / 10u;
constexpr std::uint64_t cutlim = (std::numeric_limits<std::uint64_t>::max)() % 10u;
for (const char* p = first; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(x > cutoff || (x == cutoff && digit > cutlim)))
{
return false;
}
x = (x * 10u) + digit;
}
value = static_cast<NumberUnsignedType>(x);
// reject values that do not round-trip into a narrower NumberUnsignedType
return static_cast<std::uint64_t>(value) == x;
}
/*!
@brief fast integer parser for an already-validated negative integer
@param[in] first pointer to the leading '-'
@param[in] last pointer past the last character
@param[out] value the parsed (negative) value on success
@return true on success; false on overflow (caller falls back to float)
*/
template<typename NumberIntegerType>
bool parse_integer_signed(const char* first, const char* last, NumberIntegerType& value) noexcept
{
// the state machine only reaches the signed path via a leading '-'
JSON_ASSERT(first != last && *first == '-');
std::uint64_t magnitude = 0;
// |INT64_MIN| == INT64_MAX + 1; this is the largest admissible magnitude
constexpr std::uint64_t limit = static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()) + 1u;
for (const char* p = first + 1; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(magnitude > (limit - digit) / 10u))
{
return false;
}
magnitude = (magnitude * 10u) + digit;
}
const std::int64_t x = (magnitude == limit)
? (std::numeric_limits<std::int64_t>::min)()
: -static_cast<std::int64_t>(magnitude);
value = static_cast<NumberIntegerType>(x);
// reject values that do not round-trip into a narrower NumberIntegerType
return static_cast<std::int64_t>(value) == x;
}
/*!
@brief exact fast path for parsing a `double` (Clinger's algorithm)
For the common case - at most 19 significant digits, a decimal exponent in
[-22, 22], and a significand below 2^53 - the value equals significand *
10^exp computed in IEEE-754 double arithmetic, which is exact under
round-to-nearest because both operands are exactly representable. This is the
same fast path used by fast_float/simdjson; the general cases are left to
std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
// true double precision. On platforms that keep intermediates in extended
// precision (e.g. the x87 FPU on 32-bit x86, where FLT_EVAL_METHOD == 2) the
// single significand * 10^scale step is double-rounded and can be 1 ULP off,
// so decline and let the caller fall back to the correctly-rounded
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
static const std::array<double, 23> powers_of_ten =
{
{
1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11,
1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22
}
};
const char* p = first;
bool negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
negative = (*p == '-');
++p;
}
std::uint64_t significand = 0;
int num_digits = 0;
int fractional_digits = 0;
bool seen_dot = false;
bool any_digit = false;
for (; p != last; ++p)
{
const char c = *p;
if (c >= '0' && c <= '9')
{
any_digit = true;
if (JSON_HEDLEY_UNLIKELY(num_digits >= 19))
{
return false; // significand may not fit into uint64_t
}
significand = (significand * 10u) + static_cast<std::uint64_t>(c - '0');
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (static_cast<DecimalPointType>(c) == decimal_point)
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
return false;
}
seen_dot = true;
}
else if (c == 'e' || c == 'E')
{
++p;
break;
}
else
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_digit))
{
return false;
}
int exponent = 0;
if (p != last) // an exponent part remains
{
bool exp_negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
exp_negative = (*p == '-');
++p;
}
bool any_exp_digit = false;
for (; p != last; ++p)
{
if (JSON_HEDLEY_UNLIKELY(*p < '0' || *p > '9'))
{
return false;
}
exponent = (exponent * 10) + (*p - '0');
any_exp_digit = true;
if (JSON_HEDLEY_UNLIKELY(exponent > 9999))
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_exp_digit))
{
return false;
}
if (exp_negative)
{
exponent = -exponent;
}
}
const int scale = exponent - fractional_digits;
if (JSON_HEDLEY_UNLIKELY(significand >= (static_cast<std::uint64_t>(1) << 53)))
{
return false; // significand not exactly representable as double
}
auto result = static_cast<double>(significand);
if (scale >= 0)
{
if (JSON_HEDLEY_UNLIKELY(scale > 22))
{
return false;
}
result *= powers_of_ten[static_cast<std::size_t>(scale)];
}
else
{
if (JSON_HEDLEY_UNLIKELY(-scale > 22))
{
return false;
}
result /= powers_of_ten[static_cast<std::size_t>(-scale)];
}
out = negative ? -result : result;
return true;
#endif
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
{
return false;
}
/*!
@brief parse a float with std::from_chars (Eisel-Lemire) when available
std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
@return true if the value was parsed exactly and fully; false to fall back
*/
template<typename FloatType>
bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept
{
// JSON_HAS_CPP_17 must gate the use as well as the <charconv> include above:
// some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even
// in C++14 mode, where <charconv> is not included.
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
const auto result = std::from_chars(first, last, out);
return result.ec == std::errc() && result.ptr == last;
#else
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(out);
return false;
#endif
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -1,241 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef> // size_t
#include <cstdint> // uint64_t
#include <cstring> // memcpy
#include <nlohmann/detail/macro_scope.hpp>
// Optional SIMD backend for bulk UTF-8 validation. This is an opt-in external
// dependency: nlohmann/json itself stays header-only and the C++11 scalar
// validator below is always available; defining JSON_USE_SIMDUTF additionally
// requires the simdutf headers on the include path and linking the simdutf
// library. See string_bulk_run().
//
// simdutf.h itself requires C++17 - it rejects older standards with an #error -
// so the backend is only compiled in from C++17 on. Below that the macro has no
// effect and the scalar validator is used; it accepts and rejects exactly the
// same input, so only throughput differs. macro_scope.hpp is included above to
// have JSON_HAS_CPP_17 available for this test.
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
#include <simdutf.h>
#endif
// This file contains the byte-level string-scanning helpers used by the lexer's
// contiguous fast path. They operate purely on raw bytes (no dependency on the
// lexer's template parameters) so they are free functions, keeping the lexer
// itself focused on the state machine; see lexer::scan_string_bulk().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
// classify a single byte as needing individual string handling: the closing
// quote, an escape, a control character, or a non-ASCII (UTF-8)
// lead/continuation byte. Ordinary bytes (0x20..0x7F except '"' and '\\') are
// copied verbatim, which the bulk scanner does 8 bytes at a time.
inline bool is_string_special(unsigned char c) noexcept
{
return c == '\"' || c == '\\' || c < 0x20u || c >= 0x80u;
}
// SWAR helper: return a word whose high bit is set in every byte of @a v that
// is_string_special(); zero if the 8 bytes are all ordinary.
inline std::uint64_t swar_string_special(std::uint64_t v) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
const std::uint64_t has_quote = (q - ones) & ~q & high;
const std::uint64_t has_backslash = (b - ones) & ~b & high;
const std::uint64_t has_control = (v - 0x2020202020202020ull) & ~v & high; // < 0x20
const std::uint64_t has_non_ascii = v & high; // >= 0x80
return has_quote | has_backslash | has_control | has_non_ascii;
}
// return the index of the first is_string_special() byte in [data, data+n), or
// n if every byte is ordinary; scans 8 bytes at a time
inline std::size_t find_string_special(const unsigned char* data, std::size_t n) noexcept
{
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t word = 0;
std::memcpy(&word, data + i, sizeof(word));
if (swar_string_special(word) != 0)
{
// a special byte is in this word; locate it (endian-agnostic)
for (std::size_t j = 0; j < 8; ++j)
{
if (is_string_special(data[i + j]))
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
if (is_string_special(data[i]))
{
return i;
}
}
return n;
}
// Validate one UTF-8 sequence at the front of [data, data+avail). Returns its
// length (2..4) only when the bytes form a *well-formed* sequence using exactly
// the same ranges as scan_string()'s per-byte switch, so the bulk path accepts
// precisely what the byte path accepts. Returns 0 for anything that is invalid,
// incomplete, or that the byte path must diagnose (the caller then defers to
// that path, keeping error messages unchanged). Lead bytes < 0x80 are handled
// by the caller and never passed here.
inline std::size_t validate_one_utf8(const unsigned char* data, std::size_t avail) noexcept
{
const unsigned char c0 = data[0];
if (c0 >= 0xC2 && c0 <= 0xDF) // U+0080..U+07FF
{
if (avail >= 2 && data[1] >= 0x80 && data[1] <= 0xBF)
{
return 2;
}
}
else if (c0 == 0xE0) // U+0800..U+0FFF
{
if (avail >= 3 && data[1] >= 0xA0 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if ((c0 >= 0xE1 && c0 <= 0xEC) || c0 == 0xEE || c0 == 0xEF) // U+1000..U+CFFF, U+E000..U+FFFF
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xED) // U+D000..U+D7FF (excludes surrogates)
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0x9F && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xF0) // U+10000..U+3FFFF
{
if (avail >= 4 && data[1] >= 0x90 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 >= 0xF1 && c0 <= 0xF3) // U+40000..U+FFFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 == 0xF4) // U+100000..U+10FFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0x8F && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
return 0; // invalid, incomplete, or must be diagnosed by the byte path
}
// Scalar (C++11) computation of the bulk run length: the number of leading
// bytes in [data, data+n) that are ordinary ASCII or complete well-formed UTF-8
// sequences, stopping before the first byte that needs individual handling (the
// closing quote, an escape, a control character, or an ill-formed/truncated
// sequence). ASCII is skipped 8 bytes at a time.
inline std::size_t scalar_string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
std::size_t pos = 0;
while (pos < n)
{
pos += find_string_special(data + pos, n - pos);
if (pos >= n || data[pos] < 0x80u)
{
break; // end of buffer, or a quote/escape/control byte
}
const std::size_t seq = validate_one_utf8(data + pos, n - pos);
if (seq == 0)
{
break; // ill-formed or truncated: let the byte path diagnose it
}
pos += seq;
}
return pos;
}
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
// Index of the first quote/escape/control byte in [data, data+n) (non-ASCII
// bytes are *not* stops here - the whole run is handed to simdutf), or n.
inline std::size_t find_string_delimiter(const unsigned char* data, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t v = 0;
std::memcpy(&v, data + i, sizeof(v));
const std::uint64_t q = v ^ 0x2222222222222222ull;
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull;
const std::uint64_t hit = ((q - ones) & ~q & high)
| ((b - ones) & ~b & high)
| ((v - 0x2020202020202020ull) & ~v & high);
if (hit != 0)
{
for (std::size_t j = 0; j < 8; ++j)
{
const unsigned char c = data[i + j];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
const unsigned char c = data[i];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i;
}
}
return n;
}
#endif
// Backend-dispatched bulk run length. With JSON_USE_SIMDUTF the run up to the
// next delimiter is validated in one shot by simdutf; on the rare failure the
// scalar helper recomputes the exact valid prefix so the byte path still
// produces the precise diagnostic. Without it, the pure scalar path is used.
inline std::size_t string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
const std::size_t run = find_string_delimiter(data, n);
if (run != 0 && simdutf::validate_utf8(reinterpret_cast<const char*>(data), run))
{
return run;
}
#endif
return scalar_string_bulk_run(data, n);
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -88,8 +88,13 @@ class iter_impl // NOLINT(cppcoreguidelines-special-member-functions,hicpp-speci
iter_impl() = default;
~iter_impl() = default;
iter_impl(iter_impl&&) noexcept = default;
iter_impl& operator=(iter_impl&&) noexcept = default;
// the exception specification is left to be computed rather than declared:
// an array or object type whose iterator is not nothrow move constructible
// (std::deque's is not before libstdc++ 11) would make a declared noexcept
// differ from the implicit one, which deletes the function -- and is an
// error outright with older compilers
iter_impl(iter_impl&&) = default; // NOLINT(hicpp-noexcept-move,performance-noexcept-move-constructor,cppcoreguidelines-noexcept-move-operations)
iter_impl& operator=(iter_impl&&) = default; // NOLINT(hicpp-noexcept-move,performance-noexcept-move-constructor,cppcoreguidelines-noexcept-move-operations)
/*!
@brief constructor for a given JSON instance
+47 -8
View File
@@ -17,6 +17,7 @@
#endif // JSON_NO_IO
#include <limits> // max
#include <numeric> // accumulate
#include <set> // set
#include <string> // string
#include <utility> // move
#include <vector> // vector
@@ -71,7 +72,7 @@ class json_pointer
string_t{},
[](const string_t& a, const string_t& b)
{
return detail::concat(a, '/', detail::escape(b));
return detail::concat<string_t>(a, '/', detail::escape(b));
});
}
@@ -265,7 +266,7 @@ class json_pointer
JSON_THROW(detail::parse_error::create(109, 0, detail::concat("array index '", s, "' is not a number"), nullptr));
}
const char* p = s.c_str();
const char* p = s.data();
char* p_end = nullptr; // NOLINT(misc-const-correctness)
errno = 0; // strtoull doesn't reset errno
const unsigned long long res = std::strtoull(p, &p_end, 10); // NOLINT(runtime/int)
@@ -300,19 +301,35 @@ class json_pointer
}
private:
/*!
@brief the reference token sequences that denote arrays
@ref unflatten collects the pointer prefixes that have a reference token 0
among their children; @ref get_and_create creates arrays exactly below
those prefixes and objects everywhere else. Deciding this up front keeps
the result independent of the order in which the flattened object is
iterated, which is unspecified for some object types.
*/
using array_parents_t = std::set<std::vector<string_t>>;
/*!
@brief create and return a reference to the pointed to value
@complexity Linear in the number of reference tokens.
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if array index is not a number
@throw type_error.313 if value cannot be unflattened
*/
template<typename BasicJsonType>
BasicJsonType& get_and_create(BasicJsonType& j) const
BasicJsonType& get_and_create(BasicJsonType& j, const array_parents_t& array_parents) const
{
auto* result = &j;
// the reference tokens that have been consumed so far; used to look up
// whether the value to be created below is an array or an object
std::vector<string_t> prefix;
// in case no reference tokens exist, return a reference to the JSON value
// j which will be overwritten by a primitive value
for (const auto& reference_token : reference_tokens)
@@ -321,10 +338,11 @@ class json_pointer
{
case detail::value_t::null:
{
if (reference_token == "0")
if (array_parents.find(prefix) != array_parents.end())
{
// start a new array if the reference token is 0
result = &result->operator[](0);
// some reference token below this position is 0, so the
// value is an array
result = &result->operator[](array_index<BasicJsonType>(reference_token));
}
else
{
@@ -364,6 +382,8 @@ class json_pointer
default:
JSON_THROW(detail::type_error::create(313, "invalid value to unflatten", &j));
}
prefix.push_back(reference_token);
}
return *result;
@@ -823,7 +843,8 @@ class json_pointer
{
// use the text between the beginning of the reference token
// (start) and the last slash (slash).
auto reference_token = reference_string.substr(start, slash - start);
const auto count = (slash == string_t::npos ? reference_string.size() : slash) - start;
auto reference_token = string_t(reference_string.data() + start, count);
// check reference tokens are properly escaped
for (std::size_t pos = reference_token.find_first_of('~');
@@ -939,6 +960,24 @@ class json_pointer
BasicJsonType result;
// collect the pointer prefixes that have a reference token 0 among
// their children; the values below them are arrays, all others are
// objects (see array_parents_t)
array_parents_t array_parents;
for (const auto& element : *value.m_data.m_value.object)
{
json_pointer ptr(element.first);
std::vector<string_t> prefix;
for (auto& reference_token : ptr.reference_tokens)
{
if (reference_token == "0")
{
array_parents.insert(prefix);
}
prefix.push_back(std::move(reference_token));
}
}
// iterate the JSON object values
for (const auto& element : *value.m_data.m_value.object)
{
@@ -951,7 +990,7 @@ class json_pointer
// that if the JSON pointer is "" (i.e., points to the whole value),
// function get_and_create returns a reference to the result itself.
// An assignment will then create a primitive value.
json_pointer(element.first).get_and_create(result) = element.second;
json_pointer(element.first).get_and_create(result, array_parents) = element.second;
}
return result;
+14 -6
View File
@@ -172,17 +172,18 @@ struct has_to_json < BasicJsonType, T, enable_if_t < !is_basic_json<T>::value >>
template<typename T>
using detect_key_compare = typename T::key_compare;
template<typename T>
struct has_key_compare : std::integral_constant<bool, is_detected<detect_key_compare, T>::value> {};
// obtains the actual object key comparator
// obtains the actual object key comparator: object_t::key_compare if the
// object type defines it, and default_object_comparator_t otherwise
//
// note detected_or_t is used rather than std::conditional, because the latter
// names both of its type arguments eagerly; object_t::key_compare would then
// be a hard error for an object type that does not define it
template<typename BasicJsonType>
struct actual_object_comparator
{
using object_t = typename BasicJsonType::object_t;
using object_comparator_t = typename BasicJsonType::default_object_comparator_t;
using type = typename std::conditional < has_key_compare<object_t>::value,
typename object_t::key_compare, object_comparator_t>::type;
using type = detected_or_t<object_comparator_t, detect_key_compare, object_t>;
};
template<typename BasicJsonType>
@@ -778,6 +779,13 @@ using has_erase_with_key_type = typename std::conditional <
std::true_type,
std::false_type >::type;
template<typename T>
using detect_capacity = decltype(std::declval<const T&>().capacity());
// type trait to check if a type has a capacity() member function
template<typename T>
struct has_capacity : std::integral_constant<bool, is_detected<detect_capacity, T>::value> {};
// a naive helper to check if a type is an ordered_map (exploits the fact that
// ordered_map inherits capacity() from std::vector)
template <typename T>
@@ -261,7 +261,7 @@ class binary_writer
// step 2: write the string
oa->write_characters(
reinterpret_cast<const CharType*>(j.m_data.m_value.string->c_str()),
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
break;
}
@@ -581,7 +581,7 @@ class binary_writer
// step 2: write the string
oa->write_characters(
reinterpret_cast<const CharType*>(j.m_data.m_value.string->c_str()),
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
break;
}
@@ -798,7 +798,7 @@ class binary_writer
}
write_number_with_ubjson_prefix(j.m_data.m_value.string->size(), true, use_bjdata);
oa->write_characters(
reinterpret_cast<const CharType*>(j.m_data.m_value.string->c_str()),
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
break;
}
@@ -813,7 +813,10 @@ class binary_writer
bool prefix_required = true;
if (use_type && !j.m_data.m_value.array->empty())
{
JSON_ASSERT(use_count);
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin() + 1, j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
@@ -859,7 +862,10 @@ class binary_writer
if (use_type && (bjdata_draft3 || !j.m_data.m_value.binary->empty()))
{
JSON_ASSERT(use_count);
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
oa->write_character(to_char_type('$'));
oa->write_character(bjdata_draft3 ? 'B' : 'U');
}
@@ -881,7 +887,9 @@ class binary_writer
for (size_t i = 0; i < j.m_data.m_value.binary->size(); ++i)
{
oa->write_character(to_char_type(bjdata_draft3 ? 'B' : 'U'));
oa->write_character(to_char_type(j.m_data.m_value.binary->data()[i]));
// the cast is needed for binary types whose value type
// is not an integer (e.g., std::byte)
oa->write_character(to_char_type(static_cast<std::uint8_t>(j.m_data.m_value.binary->data()[i])));
}
}
@@ -911,7 +919,10 @@ class binary_writer
bool prefix_required = true;
if (use_type && !j.m_data.m_value.object->empty())
{
JSON_ASSERT(use_count);
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin(), j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
@@ -939,7 +950,7 @@ class binary_writer
{
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
oa->write_characters(
reinterpret_cast<const CharType*>(el.first.c_str()),
reinterpret_cast<const CharType*>(el.first.data()),
el.first.size());
write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
}
@@ -1002,8 +1013,11 @@ class binary_writer
{
oa->write_character(to_char_type(element_type));
oa->write_characters(
reinterpret_cast<const CharType*>(name.c_str()),
name.size() + 1u);
reinterpret_cast<const CharType*>(name.data()),
name.size());
// the terminating null byte is written explicitly rather than taken
// from the buffer, so that string_t::data() need not be null-terminated
oa->write_character(to_char_type(0x00));
}
/*!
@@ -1044,8 +1058,9 @@ class binary_writer
write_number<std::int32_t>(to_bson_length(value.size() + 1ul), true);
oa->write_characters(
reinterpret_cast<const CharType*>(value.c_str()),
value.size() + 1);
reinterpret_cast<const CharType*>(value.data()),
value.size());
oa->write_character(to_char_type(0x00));
}
/*!
@@ -1136,7 +1151,8 @@ class binary_writer
const std::size_t embedded_document_size = std::accumulate(std::begin(value), std::end(value), static_cast<std::size_t>(0), [&array_index](std::size_t result, const typename BasicJsonType::array_t::value_type & el)
{
return result + calc_bson_element_size(std::to_string(array_index++), el);
const auto key = std::to_string(array_index++);
return result + calc_bson_element_size(string_t(key.data(), key.size()), el);
});
return sizeof(std::int32_t) + embedded_document_size + 1ul;
@@ -1163,7 +1179,11 @@ class binary_writer
for (const auto& el : value)
{
write_bson_element(std::to_string(array_index++), el);
// the index is built as a std::string, while write_bson_element takes
// a string_t; convert explicitly, as the two are only implicitly
// convertible for some string types
const auto key = std::to_string(array_index++);
write_bson_element(string_t(key.data(), key.size()), el);
}
oa->write_character(to_char_type(0x00));
+28 -16
View File
@@ -135,7 +135,7 @@ class serializer
auto i = val.m_data.m_value.object->cbegin();
for (std::size_t cnt = 0; cnt < val.m_data.m_value.object->size() - 1; ++cnt, ++i)
{
o->write_characters(indent_string.c_str(), new_indent);
o->write_characters(indent_string.data(), new_indent);
o->write_character('\"');
dump_escaped(i->first, ensure_ascii);
o->write_characters("\": ", 3);
@@ -146,14 +146,14 @@ class serializer
// last element
JSON_ASSERT(i != val.m_data.m_value.object->cend());
JSON_ASSERT(std::next(i) == val.m_data.m_value.object->cend());
o->write_characters(indent_string.c_str(), new_indent);
o->write_characters(indent_string.data(), new_indent);
o->write_character('\"');
dump_escaped(i->first, ensure_ascii);
o->write_characters("\": ", 3);
dump(i->second, true, ensure_ascii, indent_step, new_indent);
o->write_character('\n');
o->write_characters(indent_string.c_str(), current_indent);
o->write_characters(indent_string.data(), current_indent);
o->write_character('}');
}
else
@@ -208,18 +208,18 @@ class serializer
for (auto i = val.m_data.m_value.array->cbegin();
i != val.m_data.m_value.array->cend() - 1; ++i)
{
o->write_characters(indent_string.c_str(), new_indent);
o->write_characters(indent_string.data(), new_indent);
dump(*i, true, ensure_ascii, indent_step, new_indent);
o->write_characters(",\n", 2);
}
// last element
JSON_ASSERT(!val.m_data.m_value.array->empty());
o->write_characters(indent_string.c_str(), new_indent);
o->write_characters(indent_string.data(), new_indent);
dump(val.m_data.m_value.array->back(), true, ensure_ascii, indent_step, new_indent);
o->write_character('\n');
o->write_characters(indent_string.c_str(), current_indent);
o->write_characters(indent_string.data(), current_indent);
o->write_character(']');
}
else
@@ -265,7 +265,7 @@ class serializer
indent_string.resize(indent_string.size() * 2, ' ');
}
o->write_characters(indent_string.c_str(), new_indent);
o->write_characters(indent_string.data(), new_indent);
o->write_characters("\"bytes\": [", 10);
@@ -274,14 +274,14 @@ class serializer
for (auto i = val.m_data.m_value.binary->cbegin();
i != val.m_data.m_value.binary->cend() - 1; ++i)
{
dump_integer(*i);
dump_integer(to_byte_value(*i));
o->write_characters(", ", 2);
}
dump_integer(val.m_data.m_value.binary->back());
dump_integer(to_byte_value(val.m_data.m_value.binary->back()));
}
o->write_characters("],\n", 3);
o->write_characters(indent_string.c_str(), new_indent);
o->write_characters(indent_string.data(), new_indent);
o->write_characters("\"subtype\": ", 11);
if (val.m_data.m_value.binary->has_subtype())
@@ -293,7 +293,7 @@ class serializer
o->write_characters("null", 4);
}
o->write_character('\n');
o->write_characters(indent_string.c_str(), current_indent);
o->write_characters(indent_string.data(), current_indent);
o->write_character('}');
}
else
@@ -305,10 +305,10 @@ class serializer
for (auto i = val.m_data.m_value.binary->cbegin();
i != val.m_data.m_value.binary->cend() - 1; ++i)
{
dump_integer(*i);
dump_integer(to_byte_value(*i));
o->write_character(',');
}
dump_integer(val.m_data.m_value.binary->back());
dump_integer(to_byte_value(val.m_data.m_value.binary->back()));
}
o->write_characters("],\"subtype\":", 12);
@@ -596,7 +596,7 @@ class serializer
{
case error_handler_t::strict:
{
JSON_THROW(type_error::create(316, concat("incomplete UTF-8 string; last byte: 0x", hex_bytes(static_cast<std::uint8_t>(s.back() | 0))), nullptr));
JSON_THROW(type_error::create(316, concat("incomplete UTF-8 string; last byte: 0x", hex_bytes(static_cast<std::uint8_t>(s[s.size() - 1] | 0))), nullptr));
}
case error_handler_t::ignore:
@@ -703,6 +703,19 @@ class serializer
pos += 6;
}
/*!
@brief convert a single element of a binary value to its byte value
The elements of a binary value are dumped as the numbers 0..255, regardless
of the value type of the configured BinaryType: that type may be signed
(`char`), unsigned (`std::uint8_t`), or not an integer at all
(`std::byte`), none of which @ref dump_integer can handle uniformly.
*/
static std::uint8_t to_byte_value(binary_char_t x) noexcept
{
return static_cast<std::uint8_t>(x);
}
// templates to avoid warnings about useless casts
template <typename NumberType, enable_if_t<std::is_signed<NumberType>::value, int> = 0>
bool is_negative_number(NumberType x)
@@ -728,8 +741,7 @@ class serializer
template < typename NumberType, detail::enable_if_t <
std::is_integral<NumberType>::value ||
std::is_same<NumberType, number_unsigned_t>::value ||
std::is_same<NumberType, number_integer_t>::value ||
std::is_same<NumberType, binary_char_t>::value,
std::is_same<NumberType, number_integer_t>::value,
int > = 0 >
void dump_integer(NumberType x)
{
+68 -31
View File
@@ -8,50 +8,56 @@
#pragma once
#include <cstddef> // size_t
#include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief replace all occurrences of a substring by another string
@param[in,out] s the string to manipulate; changed so that all
occurrences of @a f are replaced with @a t
@param[in] f the substring to replace with @a t
@param[in] t the string to replace @a f
@pre The search string @a f must not be empty. **This precondition is
enforced with an assertion.**
@since version 2.0.0
*/
template<typename StringType>
inline void replace_substring(StringType& s, const StringType& f,
const StringType& t)
{
JSON_ASSERT(!f.empty());
for (auto pos = s.find(f); // find the first occurrence of f
pos != StringType::npos; // make sure f was found
s.replace(pos, f.size(), t), // replace with t, and
pos = s.find(f, pos + t.size())) // find the next occurrence of f
{}
}
/*!
* @brief string escaping as described in RFC 6901 (Sect. 4)
* @param[in] s string to escape
* @return escaped string
*
* Note the order of escaping "~" to "~0" and "/" to "~1" is important.
*
* The string is rebuilt in a single pass, appending whole runs between the
* characters that need escaping. Scanning with find_first_of() keeps the
* common case -- nothing to escape -- as fast as a single search, while
* repeated replace() calls would move the tail of the string once per
* escaped character.
*/
template<typename StringType>
inline StringType escape(StringType s)
inline StringType escape(const StringType& s)
{
replace_substring(s, StringType{"~"}, StringType{"~0"});
replace_substring(s, StringType{"/"}, StringType{"~1"});
return s;
auto next_special = [&s](std::size_t from)
{
const auto tilde = s.find_first_of('~', from);
const auto slash = s.find_first_of('/', from);
return tilde < slash ? tilde : slash; // npos is the largest value
};
auto pos = next_special(0);
if (pos == StringType::npos)
{
return s;
}
StringType result;
result.reserve(s.size() + 2);
std::size_t run = 0;
while (pos != StringType::npos)
{
result.append(s.data() + run, pos - run);
result.append(s[pos] == '~' ? "~0" : "~1", 2);
run = pos + 1;
pos = next_special(run);
}
result.append(s.data() + run, s.size() - run);
return result;
}
/*!
@@ -60,12 +66,43 @@ inline StringType escape(StringType s)
* @return unescaped string
*
* Note the order of escaping "~1" to "/" and "~0" to "~" is important.
*
* Rebuilt in a single pass, see @ref escape. A "~" that is followed by
* neither "0" nor "1" is passed through unchanged; @ref json_pointer rejects
* such input before it gets here.
*/
template<typename StringType>
inline void unescape(StringType& s)
{
replace_substring(s, StringType{"~1"}, StringType{"/"});
replace_substring(s, StringType{"~0"}, StringType{"~"});
auto pos = s.find_first_of('~', 0);
if (pos == StringType::npos)
{
return;
}
StringType result;
result.reserve(s.size());
std::size_t run = 0;
while (pos != StringType::npos)
{
result.append(s.data() + run, pos - run);
const auto next = pos + 1;
if (next < s.size() && (s[next] == '0' || s[next] == '1'))
{
result.append(s[next] == '0' ? "~" : "/", 1);
run = pos + 2;
}
else
{
result.append("~", 1);
run = pos + 1;
}
pos = s.find_first_of('~', run);
}
result.append(s.data() + run, s.size() - run);
s = result;
}
} // namespace detail
+96 -48
View File
@@ -783,21 +783,76 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
return it;
}
reference set_parent(reference j, std::size_t old_capacity = detail::unknown_size())
/// @brief erase an element from the object and return the following one
/// Not every map returns an iterator from erase(iterator): some containers
/// (e.g., Abseil's hash maps) return void to avoid computing a successor
/// the caller may not need. Compute it before erasing for those.
template < typename It, detail::enable_if_t <
!std::is_void<decltype(std::declval<object_t&>().erase(std::declval<It>()))>::value, int > = 0 >
typename object_t::iterator erase_from_object(It pos)
{
return m_data.m_value.object->erase(pos);
}
template < typename It, detail::enable_if_t <
std::is_void<decltype(std::declval<object_t&>().erase(std::declval<It>()))>::value, int > = 0 >
typename object_t::iterator erase_from_object(It pos)
{
auto next = std::next(pos);
m_data.m_value.object->erase(pos);
return next;
}
/// @brief the capacity of the stored array, or unknown_size()
/// Only JSON_DIAGNOSTICS uses the value, to detect a reallocation that
/// would invalidate the parent pointers. Array types that do not have a
/// capacity() member function report unknown_size(), which is treated as
/// "the elements may have moved".
#if JSON_DIAGNOSTICS
template < typename A = array_t, detail::enable_if_t < detail::has_capacity<A>::value, int > = 0 >
std::size_t array_capacity() const noexcept
{
return m_data.m_value.array->capacity();
}
template < typename A = array_t, detail::enable_if_t < !detail::has_capacity<A>::value, int > = 0 >
std::size_t array_capacity() const noexcept
{
return detail::unknown_size();
}
#else
static constexpr std::size_t array_capacity() noexcept
{
return detail::unknown_size();
}
#endif
/// @brief set the parent of a value that has just been added to an array
/// @param j the added value
/// @param old_capacity the value @ref array_capacity() returned before the
/// insertion
reference set_parent_after_array_insert(reference j, std::size_t old_capacity)
{
#if JSON_DIAGNOSTICS
if (old_capacity != detail::unknown_size())
// see https://github.com/nlohmann/json/issues/2838
JSON_ASSERT(type() == value_t::array);
if (JSON_HEDLEY_UNLIKELY(old_capacity == detail::unknown_size()
|| array_capacity() != old_capacity))
{
// see https://github.com/nlohmann/json/issues/2838
JSON_ASSERT(type() == value_t::array);
if (JSON_HEDLEY_UNLIKELY(m_data.m_value.array->capacity() != old_capacity))
{
// capacity has changed: update all parents
set_parents();
return j;
}
// the capacity has changed, or the array type does not let us tell:
// the elements may have moved, so update all parents
set_parents();
return j;
}
#else
static_cast<void>(old_capacity);
#endif
return set_parent(j);
}
reference set_parent(reference j)
{
#if JSON_DIAGNOSTICS
// ordered_json uses a vector internally, so pointers could have
// been invalidated; see https://github.com/nlohmann/json/issues/2962
#ifdef JSON_HEDLEY_MSVC_VERSION
@@ -816,7 +871,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
j.m_parent = this;
#else
static_cast<void>(j);
static_cast<void>(old_capacity);
#endif
return j;
}
@@ -2009,22 +2063,17 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
reference at(size_type idx)
{
// at only works for arrays
if (JSON_HEDLEY_LIKELY(is_array()))
{
JSON_TRY
{
return set_parent(m_data.m_value.array->at(idx));
}
JSON_CATCH (std::out_of_range&)
{
// create a better exception explanation
JSON_THROW(out_of_range::create(401, detail::concat("array index ", std::to_string(idx), " is out of range"), this));
} // cppcheck-suppress[missingReturn]
}
else
if (JSON_HEDLEY_UNLIKELY(!is_array()))
{
JSON_THROW(type_error::create(304, detail::concat("cannot use at() with ", type_name()), this));
}
if (JSON_HEDLEY_UNLIKELY(idx >= m_data.m_value.array->size()))
{
JSON_THROW(out_of_range::create(401, detail::concat("array index ", std::to_string(idx), " is out of range"), this));
}
return set_parent((*m_data.m_value.array)[idx]);
}
/// @brief access specified array element with bounds checking
@@ -2032,22 +2081,17 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
const_reference at(size_type idx) const
{
// at only works for arrays
if (JSON_HEDLEY_LIKELY(is_array()))
{
JSON_TRY
{
return m_data.m_value.array->at(idx);
}
JSON_CATCH (std::out_of_range&)
{
// create a better exception explanation
JSON_THROW(out_of_range::create(401, detail::concat("array index ", std::to_string(idx), " is out of range"), this));
} // cppcheck-suppress[missingReturn]
}
else
if (JSON_HEDLEY_UNLIKELY(!is_array()))
{
JSON_THROW(type_error::create(304, detail::concat("cannot use at() with ", type_name()), this));
}
if (JSON_HEDLEY_UNLIKELY(idx >= m_data.m_value.array->size()))
{
JSON_THROW(out_of_range::create(401, detail::concat("array index ", std::to_string(idx), " is out of range"), this));
}
return (*m_data.m_value.array)[idx];
}
/// @brief access specified object element with bounds checking
@@ -2147,12 +2191,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
#if JSON_DIAGNOSTICS
// remember array size & capacity before resizing
const auto old_size = m_data.m_value.array->size();
const auto old_capacity = m_data.m_value.array->capacity();
const auto old_capacity = array_capacity();
#endif
m_data.m_value.array->resize(idx + 1);
#if JSON_DIAGNOSTICS
if (JSON_HEDLEY_UNLIKELY(m_data.m_value.array->capacity() != old_capacity))
if (JSON_HEDLEY_UNLIKELY(old_capacity == detail::unknown_size()
|| array_capacity() != old_capacity))
{
// capacity has changed: update all parents
set_parents();
@@ -2543,7 +2588,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
case value_t::object:
{
result.m_it.object_iterator = m_data.m_value.object->erase(pos.m_it.object_iterator);
result.m_it.object_iterator = erase_from_object(pos.m_it.object_iterator);
break;
}
@@ -3173,9 +3218,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
}
// add the element to the array (move semantics)
const auto old_capacity = m_data.m_value.array->capacity();
const auto old_capacity = array_capacity();
m_data.m_value.array->push_back(std::move(val));
set_parent(m_data.m_value.array->back(), old_capacity);
set_parent_after_array_insert(m_data.m_value.array->back(), old_capacity);
// if val is moved from, basic_json move constructor marks it null, so we do not call the destructor
}
@@ -3206,9 +3251,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
}
// add the element to the array
const auto old_capacity = m_data.m_value.array->capacity();
const auto old_capacity = array_capacity();
m_data.m_value.array->push_back(val);
set_parent(m_data.m_value.array->back(), old_capacity);
set_parent_after_array_insert(m_data.m_value.array->back(), old_capacity);
}
/// @brief add an object to an array
@@ -3294,9 +3339,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
}
// add the element to the array (perfect forwarding)
const auto old_capacity = m_data.m_value.array->capacity();
const auto old_capacity = array_capacity();
m_data.m_value.array->emplace_back(std::forward<Args>(args)...);
return set_parent(m_data.m_value.array->back(), old_capacity);
return set_parent_after_array_insert(m_data.m_value.array->back(), old_capacity);
}
/// @brief add an object to an object if key does not exist
@@ -3375,7 +3420,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/insert/
iterator insert(const_iterator pos, basic_json&& val) // NOLINT(performance-unnecessary-value-param)
{
return insert(pos, val);
return insert(std::move(pos), val);
}
/// @brief inserts copies of element into array
@@ -3516,7 +3561,10 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
if (merge_objects && it.value().is_object())
{
auto it2 = m_data.m_value.object->find(it.key());
if (it2 != m_data.m_value.object->end())
// Only recurse when the existing value is itself an object.
// Otherwise overwrite, matching the documented "all other values
// are overwritten as usual" behavior (see #5402).
if (it2 != m_data.m_value.object->end() && it2->second.is_object())
{
it2->second.update(it.value(), true);
#if JSON_DIAGNOSTICS
File diff suppressed because it is too large Load Diff
-68
View File
@@ -2,9 +2,6 @@ cmake_minimum_required(VERSION 3.13...4.0)
option(JSON_Valgrind "Execute test suite with Valgrind." OFF)
option(JSON_FastTests "Skip expensive/slow tests." OFF)
option(JSON_TestSimdutf "Build the unit tests against the simdutf UTF-8 validation backend." OFF)
set(JSON_SIMDUTF_VERSION 9.1.0 CACHE STRING "The simdutf version used by JSON_TestSimdutf.")
set(JSON_32bitTest AUTO CACHE STRING "Enable the 32bit unit test (ON/OFF/AUTO/ONLY).")
set(JSON_TestStandards "" CACHE STRING "The list of standards to test explicitly.")
@@ -152,71 +149,6 @@ if(test_force)
endif()
message(STATUS "${msg}")
#############################################################################
# optionally validate UTF-8 with simdutf (JSON_USE_SIMDUTF)
#############################################################################
# The simdutf backend is opt-in and not vendored, so it is fetched here rather
# than being a checked-in dependency. Everything below hangs off test_main,
# whose usage requirements every test target inherits; the library target and
# the installed CMake package are deliberately left untouched.
if (JSON_TestSimdutf)
# simdutf requires C++17, both to compile itself and to be reachable from
# the library, which keeps its scalar validator below that. Find a tested
# standard that satisfies it.
set(simdutf_standard "")
foreach(cxx_standard ${test_cxx_standards})
if(NOT cxx_standard LESS 17 AND compiler_supports_cpp_${cxx_standard})
set(simdutf_standard ${cxx_standard})
break()
endif()
endforeach()
if("${simdutf_standard}" STREQUAL "")
# Building simdutf would fail outright without a C++17 compiler, and
# even with one it would go unused if no C++17-or-later standard is
# tested. Say so and fall back to the scalar validator rather than
# failing the build.
if(NOT compiler_supports_cpp_17)
set(simdutf_reason "the compiler does not support C++17")
else()
set(simdutf_reason "no tested standard is C++17 or later (testing ${msg_standards})")
endif()
message(WARNING
"JSON_TestSimdutf is enabled, but ${simdutf_reason}. simdutf requires C++17, so it "
"is not fetched and JSON_USE_SIMDUTF is not defined: the tests run against the "
"built-in scalar UTF-8 validator instead. Set JSON_TestStandards to include 17 or "
"later, or build with a compiler that supports C++17.")
else()
if (CMAKE_VERSION VERSION_LESS 3.18)
message(FATAL_ERROR "JSON_TestSimdutf requires CMake 3.18 or later (simdutf's minimum).")
endif()
include(FetchContent)
# simdutf builds its tests and tools by default, and its tests pull
# further dependencies of their own; only the library is needed here
set(SIMDUTF_TESTS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_TOOLS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_BENCHMARKS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_ICONV OFF CACHE BOOL "" FORCE)
FetchContent_Declare(simdutf
URL https://github.com/simdutf/simdutf/archive/refs/tags/v${JSON_SIMDUTF_VERSION}.tar.gz
DOWNLOAD_EXTRACT_TIMESTAMP TRUE
)
FetchContent_MakeAvailable(simdutf)
target_compile_definitions(test_main PUBLIC JSON_USE_SIMDUTF)
target_link_libraries(test_main PUBLIC simdutf::simdutf)
# simdutf.h requires C++17; below that the library keeps its scalar
# validator, so any C++11/14 test targets exercise the fallback and the
# C++17-and-later ones exercise simdutf. Both must agree.
message(STATUS "UTF-8 validation delegated to simdutf ${JSON_SIMDUTF_VERSION} for C++17 and later (JSON_USE_SIMDUTF)")
endif()
endif()
# *DO* use json_test_set_test_options() above this line
json_test_should_build_32bit_test(json_32bit_test json_32bit_test_only "${JSON_32bitTest}")
+40 -32
View File
@@ -11,8 +11,10 @@
#include <nlohmann/json.hpp>
#include <cstdint>
#include <string>
#include <utility>
#include <vector>
/* forward declarations */
class alt_string;
@@ -22,6 +24,10 @@ void int_to_string(alt_string& target, std::size_t value); // NOLINT(misc-use-in
/*
* This is virtually a string class.
* It covers std::string under the hood.
*
* It deliberately does not provide c_str(), back(), find(str, pos), replace(),
* or substr(): the library must not rely on them. Do not add members here
* without checking that the library actually needs them.
*/
class alt_string
{
@@ -106,11 +112,6 @@ class alt_string
return str_impl < op.str_impl;
}
const char* c_str() const
{
return str_impl.c_str();
}
char& operator[](std::size_t index)
{
return str_impl[index];
@@ -121,16 +122,6 @@ class alt_string
return str_impl[index];
}
char& back()
{
return str_impl.back();
}
const char& back() const
{
return str_impl.back();
}
void clear()
{
str_impl.clear();
@@ -146,28 +137,11 @@ class alt_string
return str_impl.empty();
}
std::size_t find(const alt_string& str, std::size_t pos = 0) const
{
return str_impl.find(str.str_impl, pos);
}
std::size_t find_first_of(char c, std::size_t pos = 0) const
{
return str_impl.find_first_of(c, pos);
}
alt_string substr(std::size_t pos = 0, std::size_t count = npos) const
{
const std::string s = str_impl.substr(pos, count);
return {s.data(), s.size()};
}
alt_string& replace(std::size_t pos, std::size_t count, const alt_string& str)
{
str_impl.replace(pos, count, str.str_impl);
return *this;
}
void reserve( std::size_t new_cap = 0 )
{
str_impl.reserve(new_cap);
@@ -202,6 +176,31 @@ bool operator<(const char* op1, const alt_string& op2) noexcept
TEST_CASE("alternative string type")
{
SECTION("binary formats")
{
alt_json doc;
doc["pi"] = 3.141;
doc["happy"] = true;
doc["list"] = {1, 2, 3};
CHECK(alt_json::from_cbor(alt_json::to_cbor(doc)) == doc);
CHECK(alt_json::from_msgpack(alt_json::to_msgpack(doc)) == doc);
// BSON is not covered: it additionally needs string_t::find(value_type),
// which alt_string does not provide
CHECK(alt_json::from_ubjson(alt_json::to_ubjson(doc)) == doc);
// a UBJSON high-precision number is parsed into a std::string that the
// reader has to hand to the SAX interface as an alt_string
const std::vector<uint8_t> high_precision =
{
'H', 'i', 0x16, '3', '.', '1', '4', '1', '5', '9', '2', '6', '5', '3',
'5', '8', '9', '7', '9', '3', '2', '3', '8', '4', '6'
};
const auto number = alt_json::from_ubjson(high_precision);
CHECK(number.is_number_float());
CHECK(number.get<double>() == doctest::Approx(3.14159265358979323846));
}
SECTION("dump")
{
{
@@ -332,6 +331,15 @@ TEST_CASE("alternative string type")
CHECK(j.at(alt_json::json_pointer("/foo/0")) == j["foo"][0]);
CHECK(j.at(alt_json::json_pointer("/foo/1")) == j["foo"][1]);
// RFC 6901 escaping works without string_t::find(str, pos), replace(),
// and substr()
auto j2 = alt_json::parse(R"({"a/b": 1, "m~n": 2, "~/~~//": 3})");
CHECK(j2.at(alt_json::json_pointer("/a~1b")) == 1);
CHECK(j2.at(alt_json::json_pointer("/m~0n")) == 2);
CHECK(j2.at(alt_json::json_pointer("/~0~1~0~0~1~1")) == 3);
CHECK(alt_json::json_pointer("/~0~1~0~0~1~1").to_string() == alt_string("/~0~1~0~0~1~1"));
CHECK(j2.flatten().unflatten() == j2);
}
SECTION("patch")
+42
View File
@@ -3843,6 +3843,48 @@ TEST_CASE("all BJData first bytes")
}
#endif
TEST_CASE("BJData use_type requires use_size")
{
SECTION("non-empty object throws other_error.502")
{
const json j = {{"a", 1}, {"b", 2}};
CHECK_THROWS_WITH_AS(json::to_bjdata(j, false, true),
"[json.exception.other_error.502] use_type requires use_size = true",
json::other_error&);
}
SECTION("non-empty array throws other_error.502")
{
const json j = {1, 2, 3};
CHECK_THROWS_WITH_AS(json::to_bjdata(j, false, true),
"[json.exception.other_error.502] use_type requires use_size = true",
json::other_error&);
}
SECTION("scalars do not throw with use_type=true, use_count=false")
{
CHECK_NOTHROW(json::to_bjdata(42, false, true));
CHECK_NOTHROW(json::to_bjdata(3.14, false, true));
CHECK_NOTHROW(json::to_bjdata("hello", false, true));
CHECK_NOTHROW(json::to_bjdata(true, false, true));
CHECK_NOTHROW(json::to_bjdata(nullptr, false, true));
}
SECTION("empty containers do not throw with use_type=true, use_count=false")
{
CHECK_NOTHROW(json::to_bjdata(json::array(), false, true));
CHECK_NOTHROW(json::to_bjdata(json::object(), false, true));
}
SECTION("valid combinations on non-empty containers")
{
const json j = {{"a", 1}, {"b", 2}};
CHECK_NOTHROW(json::to_bjdata(j, false, false));
CHECK_NOTHROW(json::to_bjdata(j, true, false));
CHECK_NOTHROW(json::to_bjdata(j, true, true));
}
}
TEST_CASE("BJData roundtrips" * doctest::skip())
{
SECTION("input from self-generated BJData files")
+13 -2
View File
@@ -2565,11 +2565,16 @@ TEST_CASE("Tagged values")
const json j = "s";
auto v = json::to_cbor(j);
SECTION("0xC6..0xD4")
const json j_bin_payload = json::binary(std::vector<std::uint8_t> {0x01, 0x02, 0x03});
auto v_bin_payload = json::to_cbor(j_bin_payload);
SECTION("0xC0..0xD7")
{
for (const auto b : std::vector<std::uint8_t>
{
0xC6, 0xC7, 0xC8, 0xC9, 0xCA, 0xCB, 0xCC, 0xCD, 0xCE, 0xCF, 0xD0, 0xD1, 0xD2, 0xD3, 0xD4
0xC0, 0xC1, 0xC2, 0xC3, 0xC4, 0xC5,
0xC6, 0xC7, 0xC8, 0xC9, 0xCA, 0xCB, 0xCC, 0xCD, 0xCE, 0xCF, 0xD0, 0xD1, 0xD2, 0xD3, 0xD4,
0xD5, 0xD6, 0xD7
})
{
CAPTURE(b);
@@ -2589,6 +2594,12 @@ TEST_CASE("Tagged values")
auto j_tagged_stored = json::from_cbor(v_tagged, true, true, json::cbor_tag_handler_t::store);
CHECK(j_tagged_stored == j);
auto v_binary_tagged = v_bin_payload;
v_binary_tagged.insert(v_binary_tagged.begin(), b);
auto j_binary_tagged_stored = json::from_cbor(v_binary_tagged, true, true, json::cbor_tag_handler_t::store);
CHECK(j_binary_tagged_stored == j_bin_payload);
CHECK(!j_binary_tagged_stored.get_binary().has_subtype());
}
}
-382
View File
@@ -12,10 +12,6 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <sstream> // stringstream
#include <string> // string
#include <vector> // vector
namespace
{
// shortcut to scan a string literal
@@ -228,381 +224,3 @@ TEST_CASE("lexer class")
CHECK((scan_string("/**//**//**/", true) == json::lexer::token_type::end_of_input));
}
}
TEST_CASE("lexer number fast path")
{
// The contiguous fast path (used for pointer/string input) must agree with
// the streaming byte path (used for std::istream) on token type, numeric
// value, and round-trip text for every well-formed number, and reject the
// same malformed numbers with the same message.
SECTION("contiguous vs streaming parity")
{
const std::vector<std::string> numbers =
{
"0", "-0", "1", "-1", "42", "-42", "10", "100", "1234567890",
"0.0", "-0.0", "3.14", "-3.14", "0.5", "-0.001", "123.456789",
"1e0", "1E0", "1e10", "1e-10", "1e+10", "1.5e3", "-2.5E-4",
"9223372036854775807", // INT64_MAX -> unsigned
"9223372036854775808", // INT64_MAX + 1 -> unsigned
"18446744073709551615", // UINT64_MAX -> unsigned
"18446744073709551616", // UINT64_MAX + 1 -> float
"-9223372036854775808", // INT64_MIN -> integer
"-9223372036854775809", // INT64_MIN - 1 -> float
"123456789012345678901234567890", // huge -> float
"0.30000000000000004", "2.2250738585072014e-308", "1e308",
// high-precision / wide-exponent values that exercise the
// std::from_chars (Eisel-Lemire) path beyond the Clinger subset
"1.7976931348623157e308", "1.2345678901234567e-250",
"9007199254740993", "5e-324", "1e-320"
};
for (const auto& n : numbers)
{
const std::string doc = "[" + n + "]";
// contiguous fast path
const json a = json::parse(doc);
// streaming byte path
std::stringstream ss(doc);
const json b = json::parse(ss);
CAPTURE(n);
CHECK(a == b);
CHECK(a.dump() == b.dump());
CHECK(a[0].type() == b[0].type());
}
}
SECTION("token type classification")
{
CHECK((scan_string("0") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("-1") == json::lexer::token_type::value_integer));
CHECK((scan_string("1.5") == json::lexer::token_type::value_float));
CHECK((scan_string("1e5") == json::lexer::token_type::value_float));
CHECK((scan_string("18446744073709551615") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("18446744073709551616") == json::lexer::token_type::value_float));
CHECK((scan_string("-9223372036854775808") == json::lexer::token_type::value_integer));
CHECK((scan_string("-9223372036854775809") == json::lexer::token_type::value_float));
}
SECTION("malformed numbers are rejected identically")
{
for (const char* bad :
{"-", "1.", "1e", "1e+", "1.2e", "01", "-01", "1..2", "1.2.3"
})
{
CAPTURE(bad);
// the contiguous fast path must decline and let the byte path report
const std::string doc = std::string("[") + bad + "]";
CHECK_FALSE(json::accept(doc));
std::stringstream ss(doc);
CHECK_FALSE(json::accept(ss));
}
}
#if !defined(JSON_NOEXCEPTION)
// these sections parse invalid input, which aborts when exceptions are off
SECTION("exhaustive grammar parity with the streaming path")
{
// The JSON number grammar is encoded twice: once as the scan_number()
// state machine and once as the contiguous fast path. Enumerate every
// short string over the number alphabet and require the two encodings to
// agree exactly - on acceptance, on the reported error, and on the parsed
// value - so they cannot drift apart.
const std::string alphabet = "01.eE+-";
// full outcome of parsing @a doc, so a mismatch in type, value, or error
// message is caught, not just a mismatch in acceptance
const auto outcome = [](const std::string & doc, bool streaming) -> std::string
{
try
{
if (streaming)
{
std::stringstream ss(doc);
const json j = json::parse(ss);
return std::string(j[0].type_name()) + '|' + j.dump();
}
const json j = json::parse(doc);
return std::string(j[0].type_name()) + '|' + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
std::vector<std::string> mismatches;
std::vector<std::string> tokens{""};
for (std::size_t length = 1; length <= 4; ++length)
{
std::vector<std::string> next;
next.reserve(tokens.size() * alphabet.size());
for (const auto& prefix : tokens)
{
for (const char c : alphabet)
{
next.push_back(prefix + c);
}
}
tokens = next;
for (const auto& token : tokens)
{
const std::string doc = "[" + token + "]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
// 7 + 49 + 343 + 2401 tokens
CHECK(tokens.size() == 2401);
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
SECTION("error positions match the streaming path")
{
// Rejecting identically is not enough: the fast path must also report the
// error at the same position as the byte path. A number directly followed
// by a newline is the interesting case, because the byte path reaches the
// newline (which resets the column) and then ungets it.
// returns the parse_error message, or "" if the document parsed
const auto contiguous_error = [](const std::string & doc) -> std::string
{
try
{
const json j = json::parse(doc);
static_cast<void>(j);
}
catch (const json::parse_error& e)
{
return {e.what()};
}
return {};
};
const auto streaming_error = [](const std::string & doc) -> std::string
{
try
{
std::stringstream ss(doc);
const json j = json::parse(ss);
static_cast<void>(j);
}
catch (const json::parse_error& e)
{
return {e.what()};
}
return {};
};
for (const char* bad :
{"[01\n]", "[00\n]", "[-01\n]", "{1\n}", "[1\n2]", "[1.2.3\n]",
"[1 \n2]", "[\n1\n2]", "1\n2", "[01\r\n]", "[1e\n]", "[-\n]"
})
{
CAPTURE(bad);
const std::string doc = bad;
const std::string contiguous_what = contiguous_error(doc);
CHECK_FALSE(contiguous_what.empty());
CHECK(contiguous_what == streaming_error(doc));
}
// A number terminated by a newline must report the same position as the
// same number terminated by anything else: scan_number() reads the
// terminator and ungets it, so the reported column is the one reached
// after the number's last character - not the 0 that an unget() across
// the newline used to leave behind.
CHECK(contiguous_error("[01\n]") == contiguous_error("[01 ]"));
CHECK(contiguous_error("[01\n]") ==
"[json.exception.parse_error.101] parse error at line 1, column 3: "
"syntax error while parsing array - unexpected number literal; expected ']'");
// the same for a multi-character token, where the column of the last
// character (the '3' of "-2.5e3") differs from the column it starts at
CHECK(contiguous_error("null -2.5e3\nfalse") == contiguous_error("null -2.5e3 false"));
CHECK(contiguous_error("null -2.5e3\nfalse") ==
"[json.exception.parse_error.101] parse error at line 1, column 11: "
"syntax error while parsing value - unexpected number literal; expected end of input");
}
#endif
}
TEST_CASE("lexer string fast path")
{
// Build a byte string from explicit values: a hex escape in a string
// literal swallows every following hex digit, which makes sequences like
// "\xC3\xA9b" mean something other than they look like.
const auto bytes = [](std::initializer_list<int> values)
{
std::string result;
for (const int value : values)
{
result.push_back(static_cast<char>(value));
}
return result;
};
#if !defined(JSON_NOEXCEPTION)
// the full outcome of parsing @a doc: the parsed value, or the exact error
// message, so a mismatch in either is caught. Only usable with exceptions
// on: parsing invalid input aborts when they are off.
const auto outcome = [](const std::string & doc, bool streaming) -> std::string
{
try
{
if (streaming)
{
std::stringstream ss(doc);
const json j = json::parse(ss);
return j.dump();
}
const json j = json::parse(doc);
return j.dump();
}
// not just parse_error: if a bulk scanner ever let ill-formed UTF-8
// through, dump() would throw type_error.316, and that has to surface
// as a reported mismatch rather than as an uncaught exception
catch (const json::exception& e)
{
return {e.what()};
}
};
#endif
// once at the start of the string, once past the first 8-byte SWAR word, so
// the bulk scanner sees each case with and without a run behind it
const std::vector<std::size_t> offsets{0, 9};
#if !defined(JSON_NOEXCEPTION)
SECTION("exhaustive contiguous vs streaming parity")
{
// ordinary ASCII, both specials, a control byte, characters that make
// the preceding backslash a valid escape, a UTF-8 lead byte of each
// length, a continuation byte, and a byte that is never valid
const std::vector<std::string> alphabet =
{
"a", "\"", "\\", "n", "u", "0", bytes({0x01}),
bytes({0xC3}), bytes({0xA9}), bytes({0xE4}), bytes({0xF0}),
bytes({0x80}), bytes({0xFF})
};
std::vector<std::string> mismatches;
std::vector<std::string> tokens{""};
for (std::size_t length = 1; length <= 3; ++length)
{
std::vector<std::string> next;
next.reserve(tokens.size() * alphabet.size());
for (const auto& prefix : tokens)
{
for (const auto& symbol : alphabet)
{
next.push_back(prefix + symbol);
}
}
tokens = next;
for (const auto& token : tokens)
{
for (const std::size_t offset : offsets)
{
const std::string doc = "[\"" + std::string(offset, 'a') + token + "\"]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
}
// 13 + 169 + 2197 tokens, each at two offsets
CHECK(tokens.size() == 2197);
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
SECTION("special bytes at every offset of the SWAR stride")
{
// The bulk scanner consumes 8 bytes at a time and then a tail; place
// every kind of byte that ends a run at each offset across two words,
// so multibyte sequences also straddle the word boundary.
const std::vector<std::string> specials =
{
"\"", "\\", bytes({0x01}), bytes({0x1F}), bytes({0x7F}),
bytes({0xC3, 0xA9}), bytes({0xE4, 0xB8, 0xAD}), bytes({0xF0, 0x9F, 0x98, 0x80}),
bytes({0xFF}), bytes({0xC3}), bytes({0xE4, 0xB8})
};
std::vector<std::string> mismatches;
for (std::size_t offset = 0; offset <= 17; ++offset)
{
for (const auto& special : specials)
{
const std::string doc = "[\"" + std::string(offset, 'a') + special + "\"]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
#endif
// json::accept() never throws, so the ranges stay covered without exceptions
SECTION("UTF-8 ranges are accepted and rejected as documented")
{
// The bulk validator must accept exactly what the byte-at-a-time
// scanner accepts, so pin the boundaries of every range it recognizes.
// aggregate, only ever brace-initialized below; default member
// initializers would stop it being an aggregate in C++11
struct utf8_case // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init)
{
std::string sequence;
bool valid;
const char* description;
};
const std::vector<utf8_case> cases =
{
{bytes({0xC2, 0x80}), true, "U+0080, shortest two-byte"},
{bytes({0xDF, 0xBF}), true, "U+07FF, longest two-byte"},
{bytes({0xC1, 0xBF}), false, "overlong two-byte"},
{bytes({0xC2, 0x7F}), false, "two-byte with bad continuation"},
{bytes({0xE0, 0xA0, 0x80}), true, "U+0800, shortest three-byte"},
{bytes({0xE0, 0x9F, 0xBF}), false, "overlong three-byte"},
{bytes({0xED, 0x9F, 0xBF}), true, "U+D7FF, just below the surrogates"},
{bytes({0xED, 0xA0, 0x80}), false, "surrogate U+D800"},
{bytes({0xED, 0xBF, 0xBF}), false, "surrogate U+DFFF"},
{bytes({0xEE, 0x80, 0x80}), true, "U+E000, just above the surrogates"},
{bytes({0xEF, 0xBF, 0xBF}), true, "U+FFFF"},
{bytes({0xF0, 0x90, 0x80, 0x80}), true, "U+10000, shortest four-byte"},
{bytes({0xF0, 0x8F, 0xBF, 0xBF}), false, "overlong four-byte"},
{bytes({0xF4, 0x8F, 0xBF, 0xBF}), true, "U+10FFFF, highest code point"},
{bytes({0xF4, 0x90, 0x80, 0x80}), false, "above U+10FFFF"},
{bytes({0xF5, 0x80, 0x80, 0x80}), false, "lead byte out of range"},
{bytes({0x80}), false, "bare continuation byte"},
{bytes({0xFF}), false, "byte that never appears in UTF-8"},
{bytes({0xC3}), false, "truncated two-byte"},
{bytes({0xE4, 0xB8}), false, "truncated three-byte"},
{bytes({0xF0, 0x9F, 0x98}), false, "truncated four-byte"}
};
for (const auto& test_case : cases)
{
CAPTURE(test_case.description);
for (const std::size_t offset : offsets)
{
CAPTURE(offset);
const std::string doc = "[\"" + std::string(offset, 'a') + test_case.sequence + "\"]";
CHECK(json::accept(doc) == test_case.valid);
#if !defined(JSON_NOEXCEPTION)
CHECK(outcome(doc, false) == outcome(doc, true));
#endif
}
}
}
}
+136
View File
@@ -0,0 +1,136 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
#include <deque>
#include <map>
#include <memory>
#include <string>
#include <type_traits>
#include <vector>
namespace
{
// std::deque has no capacity() member function, which the library only needs
// to detect a reallocation for JSON_DIAGNOSTICS
using deque_json = nlohmann::basic_json<std::map, std::deque>;
// a std::vector whose at() is hidden: the library performs its own bounds
// check and must not fall back to the container's checked accessor
template<class T, class Allocator = std::allocator<T>>
class vector_without_at : public std::vector<T, Allocator>
{
public:
using std::vector<T, Allocator>::vector;
void at() = delete;
};
using no_at_json = nlohmann::basic_json<std::map, vector_without_at>;
} // namespace
TEST_CASE("array type without capacity()")
{
SECTION("the iterators of the default configuration stay nothrow movable")
{
// basic_json's iterators take their exception specification from the
// container iterators; std::deque's is not nothrow move constructible
// with older standard libraries, which must not cost the default
// configuration its noexcept
CHECK(std::is_nothrow_move_constructible<nlohmann::json::iterator>::value);
CHECK(std::is_nothrow_move_assignable<nlohmann::json::iterator>::value);
CHECK(std::is_nothrow_move_constructible<nlohmann::json::const_iterator>::value);
CHECK(std::is_move_constructible<deque_json::iterator>::value);
CHECK(std::is_move_assignable<deque_json::iterator>::value);
}
SECTION("adding elements")
{
deque_json j = deque_json::array();
j.push_back(1);
j.push_back("two");
j.emplace_back(3);
j += 4;
CHECK(j.size() == 4);
CHECK(j == deque_json({1, "two", 3, 4}));
CHECK(j.back() == 4);
CHECK(j.front() == 1);
}
SECTION("accessing and modifying elements")
{
auto j = deque_json::parse(R"([1,2,3])");
CHECK(j[1] == 2);
CHECK(j.at(2) == 3);
// growing through operator[] fills up with null values
j[5] = 6;
CHECK(j.size() == 6);
CHECK(j[4].is_null());
CHECK(j[5] == 6);
j.erase(0);
CHECK(j == deque_json({2, 3, nullptr, nullptr, 6}));
auto it = j.erase(j.begin());
CHECK(*it == 3);
j.insert(j.begin(), 1);
CHECK(j.front() == 1);
}
SECTION("serialization and deserialization")
{
const auto j = deque_json::parse(R"({"a":[1,[2,3]],"b":[]})");
CHECK(j.dump() == R"({"a":[1,[2,3]],"b":[]})");
CHECK(deque_json::parse(j.dump()) == j);
CHECK(deque_json::from_cbor(deque_json::to_cbor(j)) == j);
// empty containers are flattened to null and cannot be restored
const auto nested = deque_json::parse(R"({"a":[1,[2,3]]})");
CHECK(nested.flatten().unflatten() == nested);
}
SECTION("references stay valid while the array grows")
{
deque_json j = deque_json::array();
j.push_back(1);
auto& first = j[0];
for (int i = 0; i < 100; ++i)
{
j.push_back(i);
}
CHECK(&first == &j[0]);
CHECK(first == 1);
}
}
TEST_CASE("array type without at()")
{
// built in memory rather than parsed, so that the exception message does
// not gain a byte range with JSON_DIAGNOSTIC_POSITIONS
no_at_json j = {1, 2, 3};
const auto& jc = j;
CHECK(j.at(0) == 1);
CHECK(j.at(2) == 3);
CHECK(jc.at(2) == 3);
CHECK_THROWS_WITH_AS(j.at(3), "[json.exception.out_of_range.401] array index 3 is out of range", no_at_json::out_of_range);
CHECK_THROWS_WITH_AS(jc.at(3), "[json.exception.out_of_range.401] array index 3 is out of range", no_at_json::out_of_range);
CHECK(j.at(no_at_json::json_pointer("/1")) == 2);
CHECK_THROWS_AS(j.at(no_at_json::json_pointer("/3")), no_at_json::out_of_range);
}
+79
View File
@@ -0,0 +1,79 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
#include <cstdint>
#include <functional>
#include <map>
#include <memory>
#include <string>
#include <vector>
#ifdef JSON_HAS_CPP_17
#include <cstddef>
#endif
namespace
{
// a BinaryType whose value type is signed: the elements must still be
// processed as the numbers 0..255
using char_binary_json = nlohmann::basic_json <
std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t,
double, std::allocator, nlohmann::adl_serializer, std::vector<char>, void >;
#ifdef JSON_HAS_CPP_17
// a BinaryType whose value type is not an integer type at all
using byte_binary_json = nlohmann::basic_json <
std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t,
double, std::allocator, nlohmann::adl_serializer, std::vector<std::byte>, void >;
#endif
} // namespace
TEST_CASE("binary type whose value type is not std::uint8_t")
{
SECTION("a signed value type does not dump negative numbers")
{
const std::vector<char> chars{'\0', '\x01', '\xFF'};
CHECK(char_binary_json::binary(chars).dump() == R"({"bytes":[0,1,255],"subtype":null})");
CHECK(char_binary_json::binary(chars, 42).dump() == R"({"bytes":[0,1,255],"subtype":42})");
CHECK(char_binary_json::binary({}).dump() == R"({"bytes":[],"subtype":null})");
}
SECTION("the default binary type is unchanged")
{
CHECK(nlohmann::json::binary({0, 1, 255}, 42).dump() == R"({"bytes":[0,1,255],"subtype":42})");
}
#ifdef JSON_HAS_CPP_17
SECTION("dumping a value type that is not an integer")
{
const std::vector<std::byte> bytes{std::byte{0}, std::byte{1}, std::byte{0xFF}};
CHECK(byte_binary_json::binary(bytes).dump() == R"({"bytes":[0,1,255],"subtype":null})");
CHECK(byte_binary_json::binary(bytes, 42).dump() == R"({"bytes":[0,1,255],"subtype":42})");
CHECK(byte_binary_json::binary({}).dump() == R"({"bytes":[],"subtype":null})");
}
SECTION("hashing and the binary formats")
{
const std::vector<std::byte> bytes{std::byte{0}, std::byte{1}, std::byte{0xFF}};
const auto j = byte_binary_json::binary(bytes);
CHECK(std::hash<byte_binary_json> {}(j) == std::hash<byte_binary_json> {}(j));
CHECK(byte_binary_json::from_cbor(byte_binary_json::to_cbor(j)) == j);
CHECK(byte_binary_json::from_msgpack(byte_binary_json::to_msgpack(j)) == j);
// UBJSON has no binary type, so binary values are written as an array
CHECK(byte_binary_json::from_ubjson(byte_binary_json::to_ubjson(j)) == byte_binary_json({0, 1, 255}));
}
#endif
}
+187
View File
@@ -0,0 +1,187 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
#include <cstdint>
#include <map>
#include <string>
#include <type_traits>
#include <utility>
#include <vector>
namespace
{
// An ObjectType that does *not* define a key_compare member type, which is
// what every hash map looks like to the library.
//
// A hash map is deliberately not used here: object_t is probed for
// key_compare inside the definition of basic_json, that is, while basic_json
// is still an incomplete type, and whether a hash map can be instantiated
// with an incomplete mapped type depends on the standard library (libstdc++ 9
// needs the size of the mapped type for its node type and rejects it). So the
// object type is built from std::map, and the inherited key_compare member
// type is shadowed by an entity that is not a type -- the library's probe
// then finds no type, exactly as for a hash map.
template<class Key, class T, class Compare, class Allocator>
struct no_key_compare_map : std::map<Key, T, Compare, Allocator>
{
using base_t = std::map<Key, T, Compare, Allocator>;
using base_t::base_t;
// shadows base_t::key_compare, which is a type; never defined or called
void key_compare();
};
using no_key_compare_json = nlohmann::basic_json<no_key_compare_map>;
// An ObjectType whose erase(iterator) returns void rather than the following
// iterator, as for instance Abseil's hash maps do
template<class Key, class T, class Compare, class Allocator>
struct void_erase_map : std::map<Key, T, Compare, Allocator>
{
using base_t = std::map<Key, T, Compare, Allocator>;
using base_t::base_t;
using iterator = typename base_t::iterator;
using base_t::erase;
void erase(iterator pos)
{
base_t::erase(pos);
}
};
using void_erase_json = nlohmann::basic_json<void_erase_map>;
} // namespace
TEST_CASE("object type whose erase() returns void")
{
SECTION("erasing every element through the returned iterator")
{
void_erase_json j;
for (int i = 0; i < 8; ++i)
{
j["k" + std::to_string(i)] = i;
}
std::size_t erased = 0;
for (auto it = j.begin(); it != j.end(); ++erased)
{
it = j.erase(it);
}
CHECK(erased == 8);
CHECK(j.empty());
}
SECTION("erasing in the middle returns the following element")
{
void_erase_json j;
for (int i = 0; i < 4; ++i)
{
j["k" + std::to_string(i)] = i;
}
auto it = j.begin();
++it;
const auto after = j.erase(it);
CHECK(j.size() == 3);
CHECK(after.key() == "k2");
CHECK(after.value() == 2);
CHECK(!j.contains("k1"));
}
SECTION("the other erase overloads are unaffected")
{
void_erase_json j;
j["a"] = 1;
j["b"] = 2;
j["c"] = 3;
CHECK(j.erase("a") == 1);
CHECK(j.erase("nope") == 0);
j.erase(j.begin(), j.end());
CHECK(j.empty());
}
}
TEST_CASE("object type without key_compare")
{
SECTION("object_comparator_t falls back to default_object_comparator_t")
{
CHECK(std::is_same < no_key_compare_json::object_comparator_t,
no_key_compare_json::default_object_comparator_t >::value);
}
SECTION("object types defining key_compare are unaffected")
{
CHECK(std::is_same<nlohmann::json::object_comparator_t,
nlohmann::json::object_t::key_compare>::value);
CHECK(std::is_same<nlohmann::ordered_json::object_comparator_t,
nlohmann::ordered_json::object_t::key_compare>::value);
}
SECTION("creating and accessing values")
{
no_key_compare_json j;
j["one"] = 1;
j["two"] = "zwei";
j["three"]["nested"] = true;
CHECK(j.size() == 3);
CHECK(j.at("one") == 1);
CHECK(j["two"] == "zwei");
CHECK(j["three"]["nested"] == true);
CHECK(j.contains("one"));
CHECK(!j.contains("four"));
CHECK(j.find("one") != j.end());
CHECK(j.count("one") == 1);
CHECK(j.erase("one") == 1);
CHECK(j.size() == 2);
}
SECTION("serialization and deserialization")
{
const auto j = no_key_compare_json::parse(R"({"a":[1,2,3],"b":{"c":null}})");
CHECK(j["a"].size() == 3);
CHECK(j["a"][2] == 3);
CHECK(j["b"]["c"].is_null());
CHECK(no_key_compare_json::parse(j.dump()) == j);
}
SECTION("binary formats")
{
const auto j = no_key_compare_json::parse(R"({"a":[1,2,3],"b":"x"})");
CHECK(no_key_compare_json::from_cbor(no_key_compare_json::to_cbor(j)) == j);
CHECK(no_key_compare_json::from_msgpack(no_key_compare_json::to_msgpack(j)) == j);
}
SECTION("flatten and unflatten")
{
// "o" has a key that looks like an array index, so unflatten() must
// not turn it into an array
const auto j = no_key_compare_json::parse(
R"({"c":[1,2,3],"d":{"e":"s"},"n":[[0,1],[2]],"o":{"2":"x"}})");
CHECK(j.flatten().unflatten() == j);
}
SECTION("conversion to and from nlohmann::json")
{
const auto j = no_key_compare_json::parse(R"({"a":1,"b":[true,null]})");
const nlohmann::json converted(j);
CHECK(converted.is_object());
CHECK(converted["a"] == 1);
CHECK(converted["b"][0] == true);
CHECK(converted["b"][1].is_null());
CHECK(no_key_compare_json(converted) == j);
}
}
+10
View File
@@ -465,6 +465,16 @@ TEST_CASE("JSON pointers")
// explicit roundtrip check
CHECK(j.flatten().unflatten() == j);
// an object is only unflattened to an array if one of its keys is the
// reference token 0; this must not depend on which key is seen first
CHECK(json({{"/2", "x"}}).unflatten() == json({{"2", "x"}}));
CHECK(json({{"/10", "y"}, {"/2", "z"}}).unflatten() == json({{"10", "y"}, {"2", "z"}}));
CHECK(json({{"/0", 1}, {"/1", 2}}).unflatten() == json({1, 2}));
CHECK(json({{"/1", 2}, {"/0", 1}}).unflatten() == json({1, 2}));
CHECK(json({{"/0", 1}, {"/2", 3}}).unflatten() == json({1, nullptr, 3}));
CHECK(json({{"/a/1", 2}, {"/a/0", 1}}).unflatten() == json({{"a", {1, 2}}}));
CHECK(json({{"/a/1", 2}, {"/a/x", 1}}).unflatten() == json({{"a", {{"1", 2}, {"x", 1}}}}));
// roundtrip for primitive values
json j_null;
CHECK(j_null.flatten().unflatten() == j_null);
+24
View File
@@ -801,6 +801,30 @@ TEST_CASE("modifiers")
j1.update(j2, true);
CHECK(j1 == json({{"string", "t"}, {"numbers", 1}}));
}
SECTION("overwrite primitive with object")
{
json j1 = {{"k", 1}};
json const j2 = {{"k", {{"x", 2}}}};
j1.update(j2, true);
CHECK(j1 == json({{"k", {{"x", 2}}}}));
}
SECTION("overwrite array with object")
{
json j1 = {{"k", {1, 2}}};
json const j2 = {{"k", {{"x", 2}}}};
j1.update(j2, true);
CHECK(j1 == json({{"k", {{"x", 2}}}}));
}
SECTION("overwrite nested primitive with object")
{
json j1 = {{"k", {{"inner", 1}}}};
json const j2 = {{"k", {{"inner", {{"x", 2}}}}}};
j1.update(j2, true);
CHECK(j1 == json({{"k", {{"inner", {{"x", 2}}}}}}));
}
}
}
}
+11
View File
@@ -1555,4 +1555,15 @@ TEST_CASE("issue #5338 - truncated CBOR tagged binary subtype is rejected")
}
}
TEST_CASE("issue #5402 - update(merge_objects=true) overwrites a primitive with an object")
{
json t = {{"k", 1}};
t.update(json{{"k", {{"x", 2}}}}, true);
CHECK(t == json({{"k", {{"x", 2}}}}));
json mixed = {{"keep", {{"a", 1}}}, {"replace", 1}};
mixed.update(json{{"keep", {{"b", 2}}}, {"replace", {{"x", 2}}}}, true);
CHECK(mixed == json({{"keep", {{"a", 1}, {"b", 2}}}, {"replace", {{"x", 2}}}}));
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
+42
View File
@@ -2503,6 +2503,48 @@ TEST_CASE("all UBJSON first bytes")
}
#endif
TEST_CASE("UBJSON use_type requires use_size")
{
SECTION("non-empty array throws other_error.502")
{
const json j = {1, 2, 3};
CHECK_THROWS_WITH_AS(json::to_ubjson(j, false, true),
"[json.exception.other_error.502] use_type requires use_size = true",
json::other_error&);
}
SECTION("non-empty object throws other_error.502")
{
const json j = {{"a", 1}, {"b", 2}};
CHECK_THROWS_WITH_AS(json::to_ubjson(j, false, true),
"[json.exception.other_error.502] use_type requires use_size = true",
json::other_error&);
}
SECTION("scalars do not throw with use_type=true, use_count=false")
{
CHECK_NOTHROW(json::to_ubjson(42, false, true));
CHECK_NOTHROW(json::to_ubjson(3.14, false, true));
CHECK_NOTHROW(json::to_ubjson("hello", false, true));
CHECK_NOTHROW(json::to_ubjson(true, false, true));
CHECK_NOTHROW(json::to_ubjson(nullptr, false, true));
}
SECTION("empty containers do not throw with use_type=true, use_count=false")
{
CHECK_NOTHROW(json::to_ubjson(json::array(), false, true));
CHECK_NOTHROW(json::to_ubjson(json::object(), false, true));
}
SECTION("valid combinations on non-empty containers")
{
const json j = {1, 2, 3};
CHECK_NOTHROW(json::to_ubjson(j, false, false));
CHECK_NOTHROW(json::to_ubjson(j, true, false));
CHECK_NOTHROW(json::to_ubjson(j, true, true));
}
}
TEST_CASE("UBJSON roundtrips" * doctest::skip())
{
SECTION("input from self-generated UBJSON files")
-238
View File
@@ -18,12 +18,7 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint8_t
#include <list>
#include <string> // string
#include <vector> // vector
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
#include <iterator>
@@ -217,65 +212,6 @@ TEST_CASE("Parse with heterogeneous iterator and sentinel types")
CHECK(j2.at(0) == 1);
}
// A type whose data() hands out raw bytes but whose size() counts something
// else - here fixed-size records. Reading [data(), data() + size()) as bytes
// would silently truncate the input, so data() and size() alone must not be
// taken as evidence of contiguous byte storage.
struct record_buffer
{
using value_type = std::array<char, 4>;
std::string bytes;
const char* data() const noexcept
{
return bytes.data();
}
std::size_t size() const noexcept
{
return bytes.size() / sizeof(value_type);
}
const char* begin() const noexcept
{
return bytes.data();
}
const char* end() const noexcept
{
return bytes.data() + bytes.size();
}
};
TEST_CASE("Contiguous byte containers take the pointer adapter")
{
// Containers with contiguous single-byte storage are routed through the
// pointer-based adapter so the bulk fast paths apply in every standard, not
// only in C++20 where the library iterators model std::contiguous_iterator.
CHECK(nlohmann::detail::is_contiguous_byte_container<std::string>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<std::vector<char>>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<std::vector<std::uint8_t>>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<std::array<char, 4>>::value);
// input_adapter() takes its container by forwarding reference, so the trait
// is also asked about reference types
CHECK(nlohmann::detail::is_contiguous_byte_container<std::string&>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<const std::string&>::value);
// everything else keeps the iterator-based adapter
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<std::list<char>>::value);
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<std::vector<int>>::value);
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<const char*>::value);
// including a type that has data() and size() but whose size() does not
// count the units data() points at: its value_type says so
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<record_buffer>::value);
// and such a container still parses through its iterators, in full - taking
// it for a byte container would stop after data() + size() bytes
const record_buffer buffer{"[1,2,3,4,5]"};
CHECK(buffer.size() * sizeof(record_buffer::value_type) < buffer.bytes.size());
CHECK(json::parse(buffer) == json({1, 2, 3, 4, 5}));
}
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// JSON_HAS_CPP_20 (do not remove; see note at top of file)
TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t")
@@ -292,180 +228,6 @@ TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t")
const std::counted_iterator<iterator_type> first2(json_str.begin(), len);
CHECK(json::accept(first2, std::default_sentinel));
}
TEST_CASE("std::counted_iterator reaches the contiguous fast paths")
{
// A sized sentinel makes the remaining element count computable in O(1), so
// std::counted_iterator over a contiguous iterator must reach the same bulk
// string/number scanners as a plain pointer - not just the byte-at-a-time
// fallback (see #5268 for the equivalent memcpy fast path).
#if JSON_HAS_RANGES
// JSON_HAS_RANGES is 0 on standard libraries with an incomplete <ranges>
// (libstdc++ < 11, libc++ < 16), where the adapter deliberately falls back
// to the byte-at-a-time scanner; everything below still has to work there.
using adapter_type = nlohmann::detail::iterator_input_adapter<std::counted_iterator<const char*>, std::default_sentinel_t>;
CHECK(adapter_type::supports_bulk_scan);
CHECK(adapter_type::supports_seek);
#endif
// exercise every fast path: long ASCII run, multibyte UTF-8, escapes, and
// integer/floating-point numbers
const std::string json_str =
R"({"ascii":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",)"
"\"utf8\":\"\xe4\xb8\xad\xe6\x96\x87\xf0\x9f\x98\x80\xc3\xa9\","
R"("escaped":"aéb\n\\","ints":[0,-1,18446744073709551615,-9223372036854775808],)"
R"("floats":[1.5,-2.25e3,0.30000000000000004]})";
const auto len = static_cast<std::iter_difference_t<const char*>>(json_str.size());
const std::counted_iterator<const char*> first(json_str.data(), len);
const json j = json::parse(first, std::default_sentinel);
// parsing through the pointer adapter must give exactly the same result
CHECK(j == json::parse(json_str));
#if !defined(JSON_NOEXCEPTION)
// Diagnostics that quote the offending token are reconstructed from the
// already-consumed input (supports_seek), a path a sized sentinel only
// reaches now; check a few that include the "last read" text. Parsing
// invalid input aborts when exceptions are off, hence the guard.
// Raw strings and explicit bytes: an escaped literal and two literals
// written next to each other both read as mistakes to static analysis.
const auto byte = [](int value)
{
return std::string(1, static_cast<char>(value));
};
const std::vector<std::string> diagnostic_docs =
{
"1\nx",
"truX",
"[tru]",
R"("abc)",
R"(["\ud834"])",
R"(["a)" + byte(0x01) + R"(b"])",
R"([")" + byte(0xC3) + byte(0x28) + R"("])",
"[1e]",
R"(["aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaX)"
};
for (const auto& text : diagnostic_docs)
{
CAPTURE(text);
const std::counted_iterator<const char*> it(text.data(), static_cast<std::iter_difference_t<const char*>>(text.size()));
std::string counted_message;
std::string string_message;
try
{
const json counted_result = json::parse(it, std::default_sentinel);
static_cast<void>(counted_result);
}
catch (const json::parse_error& e)
{
counted_message = e.what();
}
try
{
const json string_result = json::parse(text);
static_cast<void>(string_result);
}
catch (const json::parse_error& e)
{
string_message = e.what();
}
CHECK_FALSE(counted_message.empty());
CHECK(counted_message == string_message);
}
// and errors must still be reported identically
const std::string bad = "[01\n]";
const std::counted_iterator<const char*> bad_first(bad.data(), static_cast<std::iter_difference_t<const char*>>(bad.size()));
std::string counted_what;
std::string string_what;
try
{
const json counted_result = json::parse(bad_first, std::default_sentinel);
static_cast<void>(counted_result);
}
catch (const json::parse_error& e)
{
counted_what = e.what();
}
try
{
const json string_result = json::parse(bad);
static_cast<void>(string_result);
}
catch (const json::parse_error& e)
{
string_what = e.what();
}
CHECK_FALSE(counted_what.empty());
CHECK(counted_what == string_what);
#endif
}
#if !defined(JSON_NOEXCEPTION)
// several cases below are truncated on purpose, and parsing invalid input
// aborts when exceptions are off
TEST_CASE("std::counted_iterator bulk scanning stops at the counted end")
{
// The count, not the size of the underlying buffer, is the end of the
// input: the bulk scanners must never look at the bytes behind it, even
// though they are readable. Each case is compared against parsing the
// equivalent prefix as a std::string.
const auto via_counted = [](const std::string & buf, std::size_t n) -> std::string
{
const std::counted_iterator<const char*> first(buf.data(), static_cast<std::iter_difference_t<const char*>>(n));
try
{
const json j = json::parse(first, std::default_sentinel);
return "OK|" + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
const auto via_prefix = [](const std::string & buf, std::size_t n) -> std::string
{
try
{
const json j = json::parse(buf.substr(0, n));
return "OK|" + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
struct testcase // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init)
{
const char* buffer;
std::size_t count;
};
const std::vector<testcase> cases =
{
{"[\"abc\"]____TRAILING____", 7}, // exact fit, tail hidden
{"[\"abcdefghijklmnop\"]____", 8}, // cut inside a string
{"[\"abc\"]____", 6}, // cut just before the closing quote
{"[12345]xxxxx", 4}, // cut inside a number
{"[123]999999", 5}, // number ends exactly at the count
{"[\"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\"]", 12}, // closing quote only behind the count
{"[\"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\"]", 19}, // cut inside an 8-byte SWAR stride
{"[\"\xe4\xb8\xad\xe6\x96\x87\"]", 5}, // cut inside a UTF-8 sequence
{"[\"\xe4\xb8\xad\xe6\x96\x87\"]____", 10}, // complete UTF-8, tail hidden
{"[1.25e3]TRAILINGDIGITS999", 7}, // number token reaches the count
};
for (const auto& tc : cases)
{
CAPTURE(tc.buffer);
CAPTURE(tc.count);
const std::string buffer = tc.buffer;
CHECK(via_counted(buffer, tc.count) == via_prefix(buffer, tc.count));
}
}
#endif
#endif
} // namespace