Compare commits

..
Author SHA1 Message Date
Niels LohmannandClaude Opus 5 e7cca81d9a Mark the parse callback in the new test noexcept
GCC's -Wnoexcept (part of the ci_test_gcc warning set) rejects the
lambda passed to parse() because it cannot throw but is not declared
noexcept, matching the existing callback in unit-regression2.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 20:42:45 +02:00
Niels LohmannandClaude Opus 5 ee211df64a Only reserve array capacity if the array type supports it
PR #5476 added an unconditional `reserve()` call to the `start_array()`
implementations of both `json_sax_dom_parser` and
`json_sax_dom_callback_parser`. `ArrayType` is a template parameter of
`basic_json`, however, and is not required to have a `reserve()` member
function: instantiating either parser for, e.g., `basic_json<std::map,
std::deque>` fails to compile, which breaks every use of `parse()` and
the binary readers for such a type.

Move the capped reservation into a `reserve_array()` helper that selects
a no-op overload through `priority_tag` when `ArrayType` has no
`reserve()`, mirroring `from_json_object_reserve()`. `std::vector` array
types keep reserving as before, and the 16384-element cap is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 18:02:01 +02:00
Niels Lohmann c0b2878a44 Swap diagnostic positions in basic_json::swap() (#5493)
basic_json::swap() (and the friend swap() that forwards to it) only
exchanged m_data.m_type/m_data.m_value, leaving start_position/end_position
untouched under JSON_DIAGNOSTIC_POSITIONS. This is inconsistent with
copy-assignment's operator=(basic_json), which swaps positions as part of
its copy-and-swap implementation, so after swap(a, b) each value ended up
with the other value's content but its own original position.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-13 17:43:59 +02:00
Niels Lohmann a50c2537eb Reserve capped array capacity in json_sax_dom_parser::start_array() for definite-length binary arrays (#5476)
* Reserve capped array capacity for definite-length binary arrays

CBOR, MessagePack, and the optimized [$type#count UBJSON/BJData form all
pass an exact element count to sax->start_array(len), but
json_sax_dom_parser::start_array() (and the callback variant) only used
len for an overflow check against max_size() and never reserved the
underlying vector, so each element triggered a reallocation cascade via
emplace_back().

Reserve upfront, but cap the reservation at 16384 elements: max_size()
for a std::vector is far larger than any realistic input, so an
unbounded reserve(len) would let a crafted/truncated header (e.g. CBOR
0x9A + a huge uint32 count with no data) trigger a multi-gigabyte
allocation attempt instead of the normal graceful parse_error. With the
cap, a hostile length still fails fast with the existing parse_error,
while realistic arrays get a single up-front allocation.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the huge-claimed-length DoS regression tests portable across size_t widths

On a platform where size_t is narrower than 64 bits (e.g. 32-bit mingw/msvc
x86), the previously-hardcoded huge test lengths either collide with that
platform's unknown_size() sentinel (CBOR/MessagePack, both using exactly
SIZE_MAX) or exceed the platform's smaller vector<json>::max_size()
(UBJSON/BJData's 0x7FFFFFFF), so the header is now rejected outright
(out_of_range.408) instead of being accepted and only found short of data
(parse_error.110). Both are safe, bounded rejections of the hostile input;
the property under test -- no attempt to allocate space for billions of
elements -- holds either way. Accept both outcomes instead of pinning the
64-bit-only exact result.

Also fixed an unrelated clang-tidy finding (google-readability-casting) on
the functional-style std::size_t(...) casts in the neighboring "arrays of
various sizes" section.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix remaining CI failures in the huge-claimed-length DoS regression tests

- Apply the same google-readability-casting fix (std::size_t{N} instead
  of std::size_t(N)) to the "arrays of various sizes" section in
  unit-msgpack.cpp, unit-ubjson.cpp, and unit-bjdata.cpp; only
  unit-cbor.cpp had been fixed previously, since clang-tidy's build
  didn't get far enough to report the other three in the same pass.

- json_sax_dom_parser::start_array()'s max_size() check calls JSON_THROW
  directly rather than going through sax->parse_error(), so unlike the
  scanner's own "not enough data" parse_error it is not gated by
  allow_exceptions=false. On a platform where a header's claimed count
  exceeds max_size() (e.g. 32-bit, for UBJSON/BJData's 0x7FFFFFFF test
  value), from_ubjson/from_bjdata(input, true, false) can therefore still
  throw instead of returning a discarded value. Make that assertion
  tolerant of either outcome, same as the main exception-catching check
  above it.

- Guard all four "a huge claimed length..." SECTIONs with
  #if !defined(JSON_NOEXCEPTION), matching this test suite's existing
  convention for exception-dependent tests: under JSON_NOEXCEPTION,
  JSON_THROW never produces a catchable C++ exception at all (it aborts
  the process), so a section that relies on try/catch to distinguish
  between two acceptable outcomes cannot be expressed under that build
  configuration regardless of platform.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use (std::min)(len, reserve_cap) instead of a ternary in start_array()

Addresses review feedback from @gregmarr on PR #5476.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-12 21:09:33 +02:00
Niels Lohmann aa391dc0a5 Fix two develop CI regressions: dump() nodiscard warning and binary-reader const-correctness (#5520)
* Discard dump()'s [[nodiscard]] return value in an exception-only check

CHECK_THROWS_WITH_AS(j.dump(), ...) called dump() only to trigger and
catch the exception, but never used the return value. dump() is
warn_unused_result, so GCC's pedantic build (-Werror --all-warnings)
rejected it as -Werror=unused-result, breaking ci_test_gcc. Wrapped in
utils::ignore_return_value(), matching every other such call in this
file.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Mark container_frame top as const in CBOR/UBJSON readers

clang-tidy's misc-const-correctness flagged these on PR #5520's CI: the
BSON sibling copy was already const, but these two were left mutable
even though only container_stack.back().remaining is ever written.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 17:55:46 +02:00
Niels Lohmann d2514a46f7 Add benchmarks for the binary readers (#5510)
The benchmark suite covered parsing JSON text, dumping, and serializing to
CBOR, but only one binary read: FromMsgpack. Nothing measured from_cbor,
from_ubjson, from_bjdata or from_bson, so a change to binary_reader.hpp had
no baseline to be compared against.

Add read benchmarks for every format, in the two shapes that matter: from a
contiguous buffer, which is what most callers pass, and from a FILE*, which
is what FromMsgpack already measures and which compiles to different code.
FromMsgpack itself is left untouched so its numbers stay comparable across
releases. The input is derived at setup time by serializing a parsed test
file, because the test data repository ships JSON only.

The test files are wide and shallow, but the readers' cost is per container,
so add three value shapes they do not cover -- deeply nested containers, many
sibling containers, and one flat array of numbers -- plus an indefinite-length
CBOR string, a form the writer never emits and which therefore has to be
assembled by hand. UBJSON and BJData are also captured in their size- and
type-annotated form, which the readers handle in a separate code path.

BSON requires an object at the top level, so it cannot reuse the array-rooted
test files; it is captured on the object-rooted ones, and the shapes are
wrapped in an object so every format measures the same value. The setup marks
the benchmark as skipped rather than letting the exception escape if that
requirement is ever violated.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-11 13:15:09 +02:00
27 changed files with 1031 additions and 1180 deletions
+1 -1
View File
@@ -100,7 +100,7 @@ jobs:
container: ubuntu:focal
strategy:
matrix:
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_simdutf, ci_test_no_thread_local]
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_simdutf]
steps:
- name: Install build-essential
run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev
-4
View File
@@ -158,10 +158,6 @@ jobs:
# to fit: IMAGE_REL_AMD64_SECREL against `.debug_line'" because the
# MinGW linker cannot relocate the debug sections this test produces.
# The tests are only built and run here, so the debug info is not used.
# Do not add -O1 here to shrink the objects further: it does make them
# link, but the binaries clang 11.0.1 and clang 18.1.8 then produce crash
# before doctest prints its first line - 39 of 102 tests on clang 18.
# Keep the objects small by splitting the test files instead.
- name: Run CMake
run: cmake -S . -B build ^
-DCMAKE_CXX_COMPILER="C:/Program Files/LLVM/bin/clang++.exe" ^
-19
View File
@@ -260,25 +260,6 @@ add_custom_target(ci_test_noglobaludls
COMMENT "Compile and test with global UDLs disabled"
)
###############################################################################
# Disable thread-local storage.
###############################################################################
# Without thread-local storage, the copy constructor cannot bound its descent
# and copies every object and array without the call stack. That path is
# otherwise only reached by values nested deeper than the bound, so this target
# is what runs the whole test suite through it.
add_custom_target(ci_test_no_thread_local
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON
-DCMAKE_CXX_FLAGS=-DJSON_NO_THREAD_LOCAL
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_no_thread_local
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_no_thread_local
COMMAND cd ${PROJECT_BINARY_DIR}/build_no_thread_local && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure
COMMENT "Compile and test without thread-local storage"
)
###############################################################################
# Coverage.
###############################################################################
+6 -2
View File
@@ -34,10 +34,14 @@ void swap(typename binary_t::container_type& other);
```
1. Exchanges the contents of the JSON value with those of `other`. Does not invoke any move, copy, or swap operations on
individual elements. All iterators and references remain valid. The past-the-end iterator is invalidated.
individual elements. All iterators and references remain valid. The past-the-end iterator is invalidated. If macro
[`JSON_DIAGNOSTIC_POSITIONS`](../macros/json_diagnostic_positions.md) is defined to `#!cpp 1`, the
[`start_pos()`](start_pos.md)/[`end_pos()`](end_pos.md) diagnostic positions are exchanged along with the value.
2. Exchanges the contents of the JSON value from `left` with those of `right`. Does not invoke any move, copy, or swap
operations on individual elements. All iterators and references remain valid. The past-the-end iterator is
invalidated. Implemented as a friend function callable via ADL.
invalidated. Implemented as a friend function callable via ADL. If macro
[`JSON_DIAGNOSTIC_POSITIONS`](../macros/json_diagnostic_positions.md) is defined to `#!cpp 1`, the
[`start_pos()`](start_pos.md)/[`end_pos()`](end_pos.md) diagnostic positions are exchanged along with the value.
3. Exchanges the contents of a JSON array with those of `other`. Does not invoke any move, copy, or swap operations on
individual elements. All iterators and references remain valid. The past-the-end iterator is invalidated.
4. Exchanges the contents of a JSON object with those of `other`. Does not invoke any move, copy, or swap operations on
-1
View File
@@ -22,7 +22,6 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_HAS_STD_FORMAT**](json_has_std_format.md) - control `std::format`/`std::formatter` support
- [**JSON_HAS_THREE_WAY_COMPARISON**](json_has_three_way_comparison.md) - control 3-way comparison support
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
- [**JSON_NO_THREAD_LOCAL**](json_no_thread_local.md) - switch off the use of `thread_local` storage
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
- [**JSON_USE_SIMDUTF**](json_use_simdutf.md) - use the simdutf library to accelerate UTF-8 validation
@@ -1,47 +0,0 @@
# JSON_NO_THREAD_LOCAL
```cpp
#define JSON_NO_THREAD_LOCAL
```
When defined, the library does not use `#!cpp thread_local` storage. This is relevant for the few environments whose
toolchain does not support it.
The copy constructor copies the first levels of a value by copying the containers, which copy their elements, and
completes whatever is nested deeper than that without the call stack, so that copying a value cannot exhaust the stack
however deeply it is nested. It counts the levels it has descended into in a `#!cpp thread_local` variable, as a counter
shared between threads would be raced.
Without that counter, no descent can be bounded safely, so objects and arrays are copied without the call stack right
away. Copying keeps working exactly as it does otherwise - the same values come out, and deeply nested values are copied
just as safely - but copying is slower, because the containers no longer copy themselves. Copying the benchmark
documents takes 9% (`canada.json`) to 34% (`twitter.json`) longer; values built mostly from objects are affected the
most.
## Default definition
By default, `#!cpp JSON_NO_THREAD_LOCAL` is not defined.
```cpp
#undef JSON_NO_THREAD_LOCAL
```
The library defines it by itself for Clang targeting MinGW, which does not survive the `#!cpp thread_local` storage:
copying a value segfaults there, with both old and current Clang versions, while GCC targeting MinGW is unaffected.
## Examples
??? example
The code below forces the library not to use `#!cpp thread_local` storage.
```cpp
#define JSON_NO_THREAD_LOCAL 1
#include <nlohmann/json.hpp>
...
```
## Version history
- Added in version 3.12.1.
-7
View File
@@ -91,13 +91,6 @@ security reasons (e.g., Intel Software Guard Extensions (SGX)).
See [full documentation of `JSON_NO_IO`](../api/macros/json_no_io.md).
## `JSON_NO_THREAD_LOCAL`
When defined, the library does not use `#!cpp thread_local` storage. Copying a value then always avoids the call stack
rather than descending into a bounded number of levels first, which is slower but yields the same values.
See [full documentation of `JSON_NO_THREAD_LOCAL`](../api/macros/json_no_thread_local.md).
## `JSON_SKIP_LIBRARY_VERSION_CHECK`
When defined, the library will not create a compiler warning when a different version of the library was already
-1
View File
@@ -291,7 +291,6 @@ nav:
- 'JSON_HAS_THREE_WAY_COMPARISON': api/macros/json_has_three_way_comparison.md
- 'JSON_NOEXCEPTION': api/macros/json_noexception.md
- 'JSON_NO_IO': api/macros/json_no_io.md
- 'JSON_NO_THREAD_LOCAL': api/macros/json_no_thread_local.md
- 'JSON_SKIP_LIBRARY_VERSION_CHECK': api/macros/json_skip_library_version_check.md
- 'JSON_SKIP_UNSUPPORTED_COMPILER_CHECK': api/macros/json_skip_unsupported_compiler_check.md
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
@@ -1434,7 +1434,7 @@ class binary_reader
// a copy, not a reference: it must stay valid across the
// pop_back() below, which destroys the container_stack element
// it would otherwise alias
container_frame top = container_stack.back();
const container_frame top = container_stack.back();
bool at_end = false;
if (top.remaining != npos)
@@ -2211,7 +2211,7 @@ class binary_reader
// would otherwise alias.
for (;;)
{
container_frame top = container_stack.back();
const container_frame top = container_stack.back();
if (top.remaining != npos)
{
@@ -8,6 +8,7 @@
#pragma once
#include <algorithm> // min
#include <cstddef>
#include <string> // string
#include <type_traits> // enable_if_t
@@ -17,6 +18,7 @@
#include <nlohmann/detail/exceptions.hpp>
#include <nlohmann/detail/input/lexer.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/cpp_future.hpp>
#include <nlohmann/detail/string_concat.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
@@ -150,6 +152,29 @@ constexpr std::size_t unknown_size()
return (std::numeric_limits<std::size_t>::max)();
}
/*!
@brief reserve capacity for @a len elements in array @a arr
Reserving upfront avoids repeated reallocations while the elements are added,
but the reservation is capped so a bogus/hostile length (which is not bounded
by max_size(), unlike e.g. std::vector) cannot trigger an oversized allocation
for a small or truncated input.
The overload below is selected for array types without reserve() (e.g.,
std::deque), which are then left untouched.
*/
template<typename ArrayType>
auto reserve_array(ArrayType& arr, std::size_t len, priority_tag<1> /*unused*/)
-> decltype(arr.reserve(len), void())
{
constexpr std::size_t reserve_cap = 16384;
arr.reserve((std::min)(len, reserve_cap));
}
template<typename ArrayType>
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
{}
/*!
@brief SAX implementation to create a JSON value from SAX events
@@ -305,6 +330,11 @@ class json_sax_dom_parser
JSON_THROW(out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
}
if (len != detail::unknown_size())
{
reserve_array(*ref_stack.back()->m_data.m_value.array, len, priority_tag<1> {});
}
return true;
}
@@ -683,6 +713,11 @@ class json_sax_dom_callback_parser
{
JSON_THROW(out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
}
if (len != detail::unknown_size())
{
reserve_array(*ref_stack.back()->m_data.m_value.array, len, priority_tag<1> {});
}
}
return true;
-9
View File
@@ -186,15 +186,6 @@
#define JSON_NO_UNIQUE_ADDRESS
#endif
// Clang targeting MinGW does not survive the thread_local storage the copy
// constructor uses to bound its descent: every test that copies a value
// segfaults with clang 11.0.1 and clang 18.1.8, while the same tests pass with
// GCC targeting MinGW and with every other toolchain the library is tested on.
// Copying works the same way without the counter, only more slowly.
#if !defined(JSON_NO_THREAD_LOCAL) && defined(__clang__) && defined(__MINGW32__)
#define JSON_NO_THREAD_LOCAL 1
#endif
// disable documentation warnings on clang
#if defined(__clang__)
#pragma clang diagnostic push
+60 -355
View File
@@ -28,14 +28,14 @@
#pragma GCC diagnostic ignored "-Wignored-attributes"
#endif
#include <algorithm> // all_of, find, for_each, none_of
#include <algorithm> // all_of, find, for_each
#include <cstddef> // nullptr_t, ptrdiff_t, size_t
#include <functional> // hash, less
#include <initializer_list> // initializer_list
#ifndef JSON_NO_IO
#include <iosfwd> // istream, ostream
#endif // JSON_NO_IO
#include <iterator> // make_move_iterator, random_access_iterator_tag
#include <iterator> // random_access_iterator_tag
#include <memory> // unique_ptr
#include <string> // string, stoi, to_string
#include <utility> // declval, forward, move, pair, swap
@@ -822,351 +822,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
return j;
}
#ifndef JSON_NO_THREAD_LOCAL
/// the number of levels an operation descends into before it finishes the
/// value below it without the call stack
static constexpr std::uint8_t nesting_depth_limit()
{
return 128;
}
/*!
@brief how many levels the operation going on in this thread has descended into
Copying a value and comparing two values share this count. The library never
nests one inside the other - copying a value does not compare one, and
comparing two values does not copy them - and where user code nests them
anyway, sharing the count only ends a descent sooner than it had to, which
costs a little speed and is never wrong.
A byte is enough: the count never exceeds the limit by more than the single
level that notices the limit has been reached.
*/
static std::uint8_t& nesting_depth() noexcept
{
static thread_local std::uint8_t depth = 0; // NOLINT(misc-use-internal-linkage)
return depth;
}
#endif
/*!
@brief counts one level of a bounded descent for as long as it runs, and
reports whether the descent was still within the limit when it began
Looks the count up and tests it against the limit itself, rather than
leaving that to the caller: either way it is reached exactly once, so
there is nothing to be gained by making the caller do it.
Does nothing and is never @ref okay without thread-local storage, where no
descent can be bounded at all: a caller that only descends while this says
it may always ends up finishing without the call stack, exactly as if every
value were nested past the limit.
*/
class nesting_depth_guard
{
public:
nesting_depth_guard() noexcept
#ifdef JSON_NO_THREAD_LOCAL
: m_okay(false)
#else
: m_okay(nesting_depth() < nesting_depth_limit())
#endif
{
#ifndef JSON_NO_THREAD_LOCAL
++nesting_depth();
#endif
}
~nesting_depth_guard()
{
#ifndef JSON_NO_THREAD_LOCAL
--nesting_depth();
#endif
}
nesting_depth_guard(const nesting_depth_guard&) = delete;
nesting_depth_guard& operator=(const nesting_depth_guard&) = delete;
nesting_depth_guard(nesting_depth_guard&&) = delete;
nesting_depth_guard& operator=(nesting_depth_guard&&) = delete;
bool okay() const noexcept
{
return m_okay;
}
private:
bool m_okay;
};
/// an entry of the iterative deep copy's worklist: a structured value and
/// the value that is to become its copy
using copy_worklist_t = std::vector<std::pair<const basic_json*, basic_json*>>;
/// scratch space to build the key skeleton of an object copy in one go
using copy_scratch_t = std::vector<std::pair<typename object_t::key_type, basic_json>>;
/// @brief copy everything of @a src into @a dst but its type and value
static void copy_metadata(const basic_json& src, basic_json& dst)
{
// a custom base class is only required to be copy-constructible and
// move-assignable, so the copy has to go through a temporary
static_cast<json_base_class_t&>(dst) = json_base_class_t(static_cast<const json_base_class_t&>(src));
#if JSON_DIAGNOSTIC_POSITIONS
dst.start_position = src.start_position;
dst.end_position = src.end_position;
#else
static_cast<void>(src);
static_cast<void>(dst);
#endif
}
/*!
@brief copy the value of @a src into @a dst, which must not be structured
Objects and arrays are left alone: creating those is the one thing the copy
constructor and @ref copy_shallow do differently from one another, and it is
the reason copying a value can descend at all.
*/
/// @note inlined on purpose: both callers have already told an object or an
/// array apart from the rest, and letting the compiler fold that test
/// into this switch is worth a few percent when copying a value made
/// mostly of numbers
JSON_HEDLEY_ALWAYS_INLINE
static void copy_leaf_value(const basic_json& src, basic_json& dst)
{
switch (src.m_data.m_type)
{
case value_t::string:
{
dst.m_data.m_value = *src.m_data.m_value.string;
break;
}
case value_t::binary:
{
dst.m_data.m_value = *src.m_data.m_value.binary;
break;
}
case value_t::boolean:
{
dst.m_data.m_value = src.m_data.m_value.boolean;
break;
}
case value_t::number_integer:
{
dst.m_data.m_value = src.m_data.m_value.number_integer;
break;
}
case value_t::number_unsigned:
{
dst.m_data.m_value = src.m_data.m_value.number_unsigned;
break;
}
case value_t::number_float:
{
dst.m_data.m_value = src.m_data.m_value.number_float;
break;
}
case value_t::object:
case value_t::array:
case value_t::null:
case value_t::discarded:
default:
break;
}
}
/*!
@brief copy everything of @a src into the null value @a dst but the children
Objects and arrays are not copied here; they are appended to @a worklist to
be created later by @ref copy_iteratively. Until that happens, @a dst remains
a null value, so that a partially built copy can be destroyed at any point
without ever violating the class invariants.
*/
static void copy_shallow(const basic_json& src, basic_json& dst, copy_worklist_t& worklist)
{
copy_metadata(src, dst);
if (src.m_data.m_type == value_t::object || src.m_data.m_type == value_t::array)
{
// defer: dst stays a null value until its container exists
worklist.emplace_back(&src, &dst);
return;
}
copy_leaf_value(src, dst);
// only now that the value exists may the type be set: had the creation
// of the value thrown, dst would have been left as a valid null value
dst.m_data.m_type = src.m_data.m_type;
}
/// @brief create the copy of the array @a src in @a dst
/// @note structured elements are appended to @a worklist instead
static void copy_array_level(const basic_json& src, basic_json& dst, copy_worklist_t& worklist)
{
const array_t& src_array = *src.m_data.m_value.array;
// create all elements up front: growing the array afterwards could
// invalidate the pointers that are handed to the worklist
dst.m_data.m_value.array = create<array_t>(src_array.size(), basic_json());
auto dst_it = dst.m_data.m_value.array->begin();
for (auto src_it = src_array.cbegin(); src_it != src_array.cend(); ++src_it, ++dst_it)
{
copy_shallow(*src_it, *dst_it, worklist);
}
}
/// @brief create the copy of the object @a src in @a dst
/// @note structured values are appended to @a worklist instead
static void copy_object_level(const basic_json& src, basic_json& dst,
copy_worklist_t& worklist, copy_scratch_t& scratch)
{
const object_t& src_object = *src.m_data.m_value.object;
// build the complete key skeleton and hand it to the object's range
// constructor: adding the keys one by one would be quadratic for object
// types that are backed by a vector, such as nlohmann::ordered_map
scratch.clear();
scratch.reserve(src_object.size());
for (const auto& element : src_object)
{
scratch.emplace_back(element.first, basic_json());
}
dst.m_data.m_value.object = create<object_t>(std::make_move_iterator(scratch.begin()),
std::make_move_iterator(scratch.end()));
scratch.clear();
// pair every value of the copy with its counterpart in the original;
// both are enumerated in the same order for every object type with a
// deterministic order, so the lookup is only needed for exotic ones
auto src_it = src_object.cbegin();
for (auto& element : *dst.m_data.m_value.object)
{
if (JSON_HEDLEY_LIKELY(src_it != src_object.cend() && src_it->first == element.first))
{
copy_shallow(src_it->second, element.second, worklist);
++src_it;
}
else
{
const auto found = src_object.find(element.first);
JSON_ASSERT(found != src_object.cend());
copy_shallow(found->second, element.second, worklist);
}
}
}
/*!
@brief deep-copy the object or array @a src into this value without recursing
The values whose copy has not been created yet are kept on an explicit
worklist rather than on the call stack. This is only reached for values
nested deeper than @ref nesting_depth_limit levels, which is why it copies
every container by hand instead of letting the container do it: the fast
ways of doing so would descend into the elements and defeat the purpose.
*/
void copy_iteratively(const basic_json& src)
{
copy_worklist_t worklist;
copy_scratch_t scratch;
const basic_json* src_value = &src;
basic_json* dst_value = this;
for (;;)
{
if (src_value->m_data.m_type == value_t::array)
{
copy_array_level(*src_value, *dst_value, worklist);
}
else
{
copy_object_level(*src_value, *dst_value, worklist, scratch);
}
// the container is complete and will not be modified again
dst_value->set_parents();
if (worklist.empty())
{
break;
}
const auto& next = worklist.back();
src_value = next.first;
dst_value = next.second;
worklist.pop_back();
// the value stops being a null value exactly here
dst_value->m_data.m_type = src_value->m_data.m_type;
}
}
/*!
@brief copy one level of the object or array @a src into this value
The container copies its own elements, which is the fastest way to fill it.
Every element that is structured itself comes back to @ref copy_structured.
*/
void copy_level(const basic_json& src)
{
if (m_data.m_type == value_t::object)
{
m_data.m_value = *src.m_data.m_value.object;
}
else
{
m_data.m_value = *src.m_data.m_value.array;
}
set_parents();
}
/*!
@brief deep-copy the object or array @a src into this value
Copying a container copies its elements, so a value nested deeply enough
used to exhaust the call stack. The descent is bounded here: the first
@ref nesting_depth_limit levels are copied by the containers themselves, just
as they always were, and anything below that is copied without the call
stack by @ref copy_iteratively. Copying a value can therefore no longer
exhaust the stack, however deeply it is nested, just like destroying one
cannot since #1436.
Nothing has to be scanned or built by hand to reach that: a value that is
not nested deeper than the limit - all but a vanishing minority - is copied
exactly as it was before, and this whole detour costs it one counter.
@sa https://github.com/nlohmann/json/issues/5387
*/
void copy_structured(const basic_json& src)
{
const nesting_depth_guard guard;
if (JSON_HEDLEY_LIKELY(guard.okay()))
{
copy_level(src);
return;
}
// Finish this value without descending any further. It is completed
// before this returns, so a copy made by a custom base class - or by
// anything else that runs while a copy is going on - is unaffected by
// the copy it is nested in.
copy_iteratively(src);
}
public:
//////////////////////////
// JSON parser callback //
@@ -1546,15 +1201,60 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// check of passed value is valid
other.assert_invariant();
if (m_data.m_type == value_t::object || m_data.m_type == value_t::array)
switch (m_data.m_type)
{
// copying the container directly would call this constructor again
// for every element, once per nesting level
copy_structured(other);
}
else
{
copy_leaf_value(other, *this);
case value_t::object:
{
m_data.m_value = *other.m_data.m_value.object;
break;
}
case value_t::array:
{
m_data.m_value = *other.m_data.m_value.array;
break;
}
case value_t::string:
{
m_data.m_value = *other.m_data.m_value.string;
break;
}
case value_t::boolean:
{
m_data.m_value = other.m_data.m_value.boolean;
break;
}
case value_t::number_integer:
{
m_data.m_value = other.m_data.m_value.number_integer;
break;
}
case value_t::number_unsigned:
{
m_data.m_value = other.m_data.m_value.number_unsigned;
break;
}
case value_t::number_float:
{
m_data.m_value = other.m_data.m_value.number_float;
break;
}
case value_t::binary:
{
m_data.m_value = *other.m_data.m_value.binary;
break;
}
case value_t::null:
case value_t::discarded:
default:
break;
}
set_parents();
@@ -3876,6 +3576,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
std::swap(m_data.m_type, other.m_data.m_type);
std::swap(m_data.m_value, other.m_data.m_value);
#if JSON_DIAGNOSTIC_POSITIONS
std::swap(start_position, other.start_position);
std::swap(end_position, other.end_position);
#endif
set_parents();
other.set_parents();
assert_invariant();
+98 -366
View File
@@ -28,14 +28,14 @@
#pragma GCC diagnostic ignored "-Wignored-attributes"
#endif
#include <algorithm> // all_of, find, for_each, none_of
#include <algorithm> // all_of, find, for_each
#include <cstddef> // nullptr_t, ptrdiff_t, size_t
#include <functional> // hash, less
#include <initializer_list> // initializer_list
#ifndef JSON_NO_IO
#include <iosfwd> // istream, ostream
#endif // JSON_NO_IO
#include <iterator> // make_move_iterator, random_access_iterator_tag
#include <iterator> // random_access_iterator_tag
#include <memory> // unique_ptr
#include <string> // string, stoi, to_string
#include <utility> // declval, forward, move, pair, swap
@@ -2564,15 +2564,6 @@ JSON_HEDLEY_DIAGNOSTIC_POP
#define JSON_NO_UNIQUE_ADDRESS
#endif
// Clang targeting MinGW does not survive the thread_local storage the copy
// constructor uses to bound its descent: every test that copies a value
// segfaults with clang 11.0.1 and clang 18.1.8, while the same tests pass with
// GCC targeting MinGW and with every other toolchain the library is tested on.
// Copying works the same way without the counter, only more slowly.
#if !defined(JSON_NO_THREAD_LOCAL) && defined(__clang__) && defined(__MINGW32__)
#define JSON_NO_THREAD_LOCAL 1
#endif
// disable documentation warnings on clang
#if defined(__clang__)
#pragma clang diagnostic push
@@ -7901,6 +7892,7 @@ NLOHMANN_JSON_NAMESPACE_END
#include <algorithm> // min
#include <cstddef>
#include <string> // string
#include <type_traits> // enable_if_t
@@ -10738,6 +10730,8 @@ NLOHMANN_JSON_NAMESPACE_END
// #include <nlohmann/detail/macro_scope.hpp>
// #include <nlohmann/detail/meta/cpp_future.hpp>
// #include <nlohmann/detail/string_concat.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
@@ -10872,6 +10866,29 @@ constexpr std::size_t unknown_size()
return (std::numeric_limits<std::size_t>::max)();
}
/*!
@brief reserve capacity for @a len elements in array @a arr
Reserving upfront avoids repeated reallocations while the elements are added,
but the reservation is capped so a bogus/hostile length (which is not bounded
by max_size(), unlike e.g. std::vector) cannot trigger an oversized allocation
for a small or truncated input.
The overload below is selected for array types without reserve() (e.g.,
std::deque), which are then left untouched.
*/
template<typename ArrayType>
auto reserve_array(ArrayType& arr, std::size_t len, priority_tag<1> /*unused*/)
-> decltype(arr.reserve(len), void())
{
constexpr std::size_t reserve_cap = 16384;
arr.reserve((std::min)(len, reserve_cap));
}
template<typename ArrayType>
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
{}
/*!
@brief SAX implementation to create a JSON value from SAX events
@@ -11027,6 +11044,11 @@ class json_sax_dom_parser
JSON_THROW(out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
}
if (len != detail::unknown_size())
{
reserve_array(*ref_stack.back()->m_data.m_value.array, len, priority_tag<1> {});
}
return true;
}
@@ -11405,6 +11427,11 @@ class json_sax_dom_callback_parser
{
JSON_THROW(out_of_range::create(408, concat("excessive array size: ", std::to_string(len)), ref_stack.back()));
}
if (len != detail::unknown_size())
{
reserve_array(*ref_stack.back()->m_data.m_value.array, len, priority_tag<1> {});
}
}
return true;
@@ -13392,7 +13419,7 @@ class binary_reader
// a copy, not a reference: it must stay valid across the
// pop_back() below, which destroys the container_stack element
// it would otherwise alias
container_frame top = container_stack.back();
const container_frame top = container_stack.back();
bool at_end = false;
if (top.remaining != npos)
@@ -14169,7 +14196,7 @@ class binary_reader
// would otherwise alias.
for (;;)
{
container_frame top = container_stack.back();
const container_frame top = container_stack.back();
if (top.remaining != npos)
{
@@ -24564,351 +24591,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
return j;
}
#ifndef JSON_NO_THREAD_LOCAL
/// the number of levels an operation descends into before it finishes the
/// value below it without the call stack
static constexpr std::uint8_t nesting_depth_limit()
{
return 128;
}
/*!
@brief how many levels the operation going on in this thread has descended into
Copying a value and comparing two values share this count. The library never
nests one inside the other - copying a value does not compare one, and
comparing two values does not copy them - and where user code nests them
anyway, sharing the count only ends a descent sooner than it had to, which
costs a little speed and is never wrong.
A byte is enough: the count never exceeds the limit by more than the single
level that notices the limit has been reached.
*/
static std::uint8_t& nesting_depth() noexcept
{
static thread_local std::uint8_t depth = 0; // NOLINT(misc-use-internal-linkage)
return depth;
}
#endif
/*!
@brief counts one level of a bounded descent for as long as it runs, and
reports whether the descent was still within the limit when it began
Looks the count up and tests it against the limit itself, rather than
leaving that to the caller: either way it is reached exactly once, so
there is nothing to be gained by making the caller do it.
Does nothing and is never @ref okay without thread-local storage, where no
descent can be bounded at all: a caller that only descends while this says
it may always ends up finishing without the call stack, exactly as if every
value were nested past the limit.
*/
class nesting_depth_guard
{
public:
nesting_depth_guard() noexcept
#ifdef JSON_NO_THREAD_LOCAL
: m_okay(false)
#else
: m_okay(nesting_depth() < nesting_depth_limit())
#endif
{
#ifndef JSON_NO_THREAD_LOCAL
++nesting_depth();
#endif
}
~nesting_depth_guard()
{
#ifndef JSON_NO_THREAD_LOCAL
--nesting_depth();
#endif
}
nesting_depth_guard(const nesting_depth_guard&) = delete;
nesting_depth_guard& operator=(const nesting_depth_guard&) = delete;
nesting_depth_guard(nesting_depth_guard&&) = delete;
nesting_depth_guard& operator=(nesting_depth_guard&&) = delete;
bool okay() const noexcept
{
return m_okay;
}
private:
bool m_okay;
};
/// an entry of the iterative deep copy's worklist: a structured value and
/// the value that is to become its copy
using copy_worklist_t = std::vector<std::pair<const basic_json*, basic_json*>>;
/// scratch space to build the key skeleton of an object copy in one go
using copy_scratch_t = std::vector<std::pair<typename object_t::key_type, basic_json>>;
/// @brief copy everything of @a src into @a dst but its type and value
static void copy_metadata(const basic_json& src, basic_json& dst)
{
// a custom base class is only required to be copy-constructible and
// move-assignable, so the copy has to go through a temporary
static_cast<json_base_class_t&>(dst) = json_base_class_t(static_cast<const json_base_class_t&>(src));
#if JSON_DIAGNOSTIC_POSITIONS
dst.start_position = src.start_position;
dst.end_position = src.end_position;
#else
static_cast<void>(src);
static_cast<void>(dst);
#endif
}
/*!
@brief copy the value of @a src into @a dst, which must not be structured
Objects and arrays are left alone: creating those is the one thing the copy
constructor and @ref copy_shallow do differently from one another, and it is
the reason copying a value can descend at all.
*/
/// @note inlined on purpose: both callers have already told an object or an
/// array apart from the rest, and letting the compiler fold that test
/// into this switch is worth a few percent when copying a value made
/// mostly of numbers
JSON_HEDLEY_ALWAYS_INLINE
static void copy_leaf_value(const basic_json& src, basic_json& dst)
{
switch (src.m_data.m_type)
{
case value_t::string:
{
dst.m_data.m_value = *src.m_data.m_value.string;
break;
}
case value_t::binary:
{
dst.m_data.m_value = *src.m_data.m_value.binary;
break;
}
case value_t::boolean:
{
dst.m_data.m_value = src.m_data.m_value.boolean;
break;
}
case value_t::number_integer:
{
dst.m_data.m_value = src.m_data.m_value.number_integer;
break;
}
case value_t::number_unsigned:
{
dst.m_data.m_value = src.m_data.m_value.number_unsigned;
break;
}
case value_t::number_float:
{
dst.m_data.m_value = src.m_data.m_value.number_float;
break;
}
case value_t::object:
case value_t::array:
case value_t::null:
case value_t::discarded:
default:
break;
}
}
/*!
@brief copy everything of @a src into the null value @a dst but the children
Objects and arrays are not copied here; they are appended to @a worklist to
be created later by @ref copy_iteratively. Until that happens, @a dst remains
a null value, so that a partially built copy can be destroyed at any point
without ever violating the class invariants.
*/
static void copy_shallow(const basic_json& src, basic_json& dst, copy_worklist_t& worklist)
{
copy_metadata(src, dst);
if (src.m_data.m_type == value_t::object || src.m_data.m_type == value_t::array)
{
// defer: dst stays a null value until its container exists
worklist.emplace_back(&src, &dst);
return;
}
copy_leaf_value(src, dst);
// only now that the value exists may the type be set: had the creation
// of the value thrown, dst would have been left as a valid null value
dst.m_data.m_type = src.m_data.m_type;
}
/// @brief create the copy of the array @a src in @a dst
/// @note structured elements are appended to @a worklist instead
static void copy_array_level(const basic_json& src, basic_json& dst, copy_worklist_t& worklist)
{
const array_t& src_array = *src.m_data.m_value.array;
// create all elements up front: growing the array afterwards could
// invalidate the pointers that are handed to the worklist
dst.m_data.m_value.array = create<array_t>(src_array.size(), basic_json());
auto dst_it = dst.m_data.m_value.array->begin();
for (auto src_it = src_array.cbegin(); src_it != src_array.cend(); ++src_it, ++dst_it)
{
copy_shallow(*src_it, *dst_it, worklist);
}
}
/// @brief create the copy of the object @a src in @a dst
/// @note structured values are appended to @a worklist instead
static void copy_object_level(const basic_json& src, basic_json& dst,
copy_worklist_t& worklist, copy_scratch_t& scratch)
{
const object_t& src_object = *src.m_data.m_value.object;
// build the complete key skeleton and hand it to the object's range
// constructor: adding the keys one by one would be quadratic for object
// types that are backed by a vector, such as nlohmann::ordered_map
scratch.clear();
scratch.reserve(src_object.size());
for (const auto& element : src_object)
{
scratch.emplace_back(element.first, basic_json());
}
dst.m_data.m_value.object = create<object_t>(std::make_move_iterator(scratch.begin()),
std::make_move_iterator(scratch.end()));
scratch.clear();
// pair every value of the copy with its counterpart in the original;
// both are enumerated in the same order for every object type with a
// deterministic order, so the lookup is only needed for exotic ones
auto src_it = src_object.cbegin();
for (auto& element : *dst.m_data.m_value.object)
{
if (JSON_HEDLEY_LIKELY(src_it != src_object.cend() && src_it->first == element.first))
{
copy_shallow(src_it->second, element.second, worklist);
++src_it;
}
else
{
const auto found = src_object.find(element.first);
JSON_ASSERT(found != src_object.cend());
copy_shallow(found->second, element.second, worklist);
}
}
}
/*!
@brief deep-copy the object or array @a src into this value without recursing
The values whose copy has not been created yet are kept on an explicit
worklist rather than on the call stack. This is only reached for values
nested deeper than @ref nesting_depth_limit levels, which is why it copies
every container by hand instead of letting the container do it: the fast
ways of doing so would descend into the elements and defeat the purpose.
*/
void copy_iteratively(const basic_json& src)
{
copy_worklist_t worklist;
copy_scratch_t scratch;
const basic_json* src_value = &src;
basic_json* dst_value = this;
for (;;)
{
if (src_value->m_data.m_type == value_t::array)
{
copy_array_level(*src_value, *dst_value, worklist);
}
else
{
copy_object_level(*src_value, *dst_value, worklist, scratch);
}
// the container is complete and will not be modified again
dst_value->set_parents();
if (worklist.empty())
{
break;
}
const auto& next = worklist.back();
src_value = next.first;
dst_value = next.second;
worklist.pop_back();
// the value stops being a null value exactly here
dst_value->m_data.m_type = src_value->m_data.m_type;
}
}
/*!
@brief copy one level of the object or array @a src into this value
The container copies its own elements, which is the fastest way to fill it.
Every element that is structured itself comes back to @ref copy_structured.
*/
void copy_level(const basic_json& src)
{
if (m_data.m_type == value_t::object)
{
m_data.m_value = *src.m_data.m_value.object;
}
else
{
m_data.m_value = *src.m_data.m_value.array;
}
set_parents();
}
/*!
@brief deep-copy the object or array @a src into this value
Copying a container copies its elements, so a value nested deeply enough
used to exhaust the call stack. The descent is bounded here: the first
@ref nesting_depth_limit levels are copied by the containers themselves, just
as they always were, and anything below that is copied without the call
stack by @ref copy_iteratively. Copying a value can therefore no longer
exhaust the stack, however deeply it is nested, just like destroying one
cannot since #1436.
Nothing has to be scanned or built by hand to reach that: a value that is
not nested deeper than the limit - all but a vanishing minority - is copied
exactly as it was before, and this whole detour costs it one counter.
@sa https://github.com/nlohmann/json/issues/5387
*/
void copy_structured(const basic_json& src)
{
const nesting_depth_guard guard;
if (JSON_HEDLEY_LIKELY(guard.okay()))
{
copy_level(src);
return;
}
// Finish this value without descending any further. It is completed
// before this returns, so a copy made by a custom base class - or by
// anything else that runs while a copy is going on - is unaffected by
// the copy it is nested in.
copy_iteratively(src);
}
public:
//////////////////////////
// JSON parser callback //
@@ -25288,15 +24970,60 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// check of passed value is valid
other.assert_invariant();
if (m_data.m_type == value_t::object || m_data.m_type == value_t::array)
switch (m_data.m_type)
{
// copying the container directly would call this constructor again
// for every element, once per nesting level
copy_structured(other);
}
else
{
copy_leaf_value(other, *this);
case value_t::object:
{
m_data.m_value = *other.m_data.m_value.object;
break;
}
case value_t::array:
{
m_data.m_value = *other.m_data.m_value.array;
break;
}
case value_t::string:
{
m_data.m_value = *other.m_data.m_value.string;
break;
}
case value_t::boolean:
{
m_data.m_value = other.m_data.m_value.boolean;
break;
}
case value_t::number_integer:
{
m_data.m_value = other.m_data.m_value.number_integer;
break;
}
case value_t::number_unsigned:
{
m_data.m_value = other.m_data.m_value.number_unsigned;
break;
}
case value_t::number_float:
{
m_data.m_value = other.m_data.m_value.number_float;
break;
}
case value_t::binary:
{
m_data.m_value = *other.m_data.m_value.binary;
break;
}
case value_t::null:
case value_t::discarded:
default:
break;
}
set_parents();
@@ -27618,6 +27345,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
std::swap(m_data.m_type, other.m_data.m_type);
std::swap(m_data.m_value, other.m_data.m_value);
#if JSON_DIAGNOSTIC_POSITIONS
std::swap(start_position, other.start_position);
std::swap(end_position, other.end_position);
#endif
set_parents();
other.set_parents();
assert_invariant();
+1 -6
View File
@@ -75,12 +75,7 @@ target_compile_options(test_main PUBLIC
# is annotated JSON_HEDLEY_NO_RETURN (it always throws), which
# makes MSVC flag the code following its call in binary_reader.hpp
# as unreachable for that instantiation, in both Debug and Release
# Disable warning C4503: decorated name length exceeded, name was truncated; the deep
# copy support added for #5387 pushes the mangled name of
# std::allocator_traits<...>::construct for the custom-base-class
# test's map type past VS2015's limit. The name is only used for
# debug info, so truncation does not affect the build.
$<$<CXX_COMPILER_ID:MSVC>:/W4;/wd4566;/wd4996;/wd4702;/wd4503>
$<$<CXX_COMPILER_ID:MSVC>:/W4;/wd4566;/wd4996;/wd4702>
# https://github.com/nlohmann/json/issues/1114
$<$<CXX_COMPILER_ID:MSVC>:/bigobj> $<$<BOOL:${MINGW}>:-Wa,-mbig-obj>
+319
View File
@@ -252,4 +252,323 @@ static void BinaryToCbor(benchmark::State& state)
}
BENCHMARK(BinaryToCbor)->RangeMultiplier(2)->Range(8, 8 << 12);
//////////////////////////////////////////////////////////////////////////////
// parse binary formats
//////////////////////////////////////////////////////////////////////////////
// Only MessagePack had a read benchmark (FromMsgpack above, left untouched so
// its numbers stay comparable across releases). The benchmarks below cover the
// other formats, and read from a contiguous buffer as well as from a FILE*:
// most callers pass a container, and the two adapters compile to different
// code. The test data repository ships JSON only, so the input for each is
// derived at setup time by serializing a parsed test file.
/// binary format to benchmark; the _optimized variants add UBJSON/BJData size
/// and type annotations, which the readers handle in a separate code path
enum class binary_format
{
cbor,
msgpack,
ubjson,
ubjson_optimized,
bjdata,
bjdata_optimized,
bson
};
static std::vector<std::uint8_t> to_binary(const json& j, const binary_format format)
{
switch (format)
{
case binary_format::cbor:
return json::to_cbor(j);
case binary_format::msgpack:
return json::to_msgpack(j);
case binary_format::ubjson:
return json::to_ubjson(j);
case binary_format::ubjson_optimized:
return json::to_ubjson(j, true, true);
case binary_format::bjdata:
return json::to_bjdata(j);
case binary_format::bjdata_optimized:
return json::to_bjdata(j, true, true);
case binary_format::bson:
default:
return json::to_bson(j);
}
}
static json from_binary(const std::vector<std::uint8_t>& bytes, const binary_format format)
{
switch (format)
{
case binary_format::cbor:
return json::from_cbor(bytes);
case binary_format::msgpack:
return json::from_msgpack(bytes);
case binary_format::ubjson:
case binary_format::ubjson_optimized:
return json::from_ubjson(bytes);
case binary_format::bjdata:
case binary_format::bjdata_optimized:
return json::from_bjdata(bytes);
case binary_format::bson:
default:
return json::from_bson(bytes);
}
}
static json from_binary(std::FILE* file, const binary_format format)
{
switch (format)
{
case binary_format::cbor:
return json::from_cbor(file);
case binary_format::msgpack:
return json::from_msgpack(file);
case binary_format::ubjson:
case binary_format::ubjson_optimized:
return json::from_ubjson(file);
case binary_format::bjdata:
case binary_format::bjdata_optimized:
return json::from_bjdata(file);
case binary_format::bson:
default:
return json::from_bson(file);
}
}
/*!
@brief serialize a parsed test file to @a format
Returns an empty vector and marks the benchmark as skipped if the file cannot
be represented in the format, rather than letting the exception escape: BSON
requires an object at the top level, and several test files are arrays.
*/
static std::vector<std::uint8_t> binary_input(benchmark::State& state, const char* filename, const binary_format format)
{
std::ifstream f(filename);
std::string const str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
const json j = json::parse(str);
if (format == binary_format::bson && !j.is_object())
{
state.SkipWithError("BSON requires an object at the top level");
return {};
}
return to_binary(j, format);
}
static void FromBinaryBuffer(benchmark::State& state, const char* filename, const binary_format format)
{
const std::vector<std::uint8_t> bytes = binary_input(state, filename, format);
if (bytes.empty())
{
return;
}
for (auto _ : state)
{
// the value is destroyed outside the timed section, because destroying
// a large DOM is not what this benchmark measures
state.PauseTiming();
auto* j = new json();
state.ResumeTiming();
*j = from_binary(bytes, format);
state.PauseTiming();
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / floats, TEST_DATA_DIRECTORY "/regression/floats.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, cbor / signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, msgpack / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson_optimized / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson_optimized);
BENCHMARK_CAPTURE(FromBinaryBuffer, ubjson_optimized / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson_optimized);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata_optimized / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bjdata_optimized);
BENCHMARK_CAPTURE(FromBinaryBuffer, bjdata_optimized / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata_optimized);
// BSON requires an object at the top level, so the array-rooted test files
// (jeopardy and the regression files) cannot be captured here
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::bson);
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", binary_format::bson);
BENCHMARK_CAPTURE(FromBinaryBuffer, bson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bson);
static void FromBinaryFile(benchmark::State& state, const char* filename, const binary_format format)
{
const std::vector<std::uint8_t> bytes = binary_input(state, filename, format);
if (bytes.empty())
{
return;
}
const char* tmp = "benchmark_input.bin";
std::ofstream o(tmp, std::ios::binary);
o.write(reinterpret_cast<const char*>(bytes.data()), static_cast<std::streamsize>(bytes.size()));
o.flush();
o.close();
for (auto _ : state)
{
state.PauseTiming();
auto* j = new json();
auto* file = std::fopen(tmp, "rb");
state.ResumeTiming();
*j = from_binary(file, format);
state.PauseTiming();
std::fclose(file);
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromBinaryFile, cbor / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryFile, cbor / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryFile, ubjson / canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryFile, ubjson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryFile, bjdata / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryFile, bson / twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", binary_format::bson);
//////////////////////////////////////////////////////////////////////////////
// parse binary formats: value shapes
//////////////////////////////////////////////////////////////////////////////
// The test files above are wide and shallow, but the readers' cost is per
// container, so these cover the shapes that stress the container handling
// itself. Every shape is wrapped in an object so that BSON, which requires an
// object at the top level, measures the same value as the other formats.
/// deeply nested arrays: one container per level, no other work
static json make_nested()
{
json nested = json::array();
json* p = &nested;
for (std::size_t i = 1; i < 1000; ++i)
{
p->push_back(json::array());
p = &p->operator[](0);
}
json j = json::object();
j["data"] = std::move(nested);
return j;
}
/// many sibling containers: maximum container churn, minimum nesting
static json make_containers()
{
json data = json::array();
for (std::size_t i = 0; i < 100000; ++i)
{
data.push_back(json::array({1, 2}));
}
json j = json::object();
j["data"] = std::move(data);
return j;
}
/// one flat array of numbers: the scalar decoding path, which must not move
static json make_scalars()
{
json data = json::array();
for (std::size_t i = 0; i < 1000000; ++i)
{
data.push_back(i);
}
json j = json::object();
j["data"] = std::move(data);
return j;
}
static void FromBinaryShape(benchmark::State& state, json (*build)(), const binary_format format)
{
const std::vector<std::uint8_t> bytes = to_binary(build(), format);
for (auto _ : state)
{
state.PauseTiming();
auto* j = new json();
state.ResumeTiming();
*j = from_binary(bytes, format);
state.PauseTiming();
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromBinaryShape, nested / cbor, make_nested, binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryShape, nested / msgpack, make_nested, binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryShape, nested / ubjson, make_nested, binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryShape, nested / bjdata, make_nested, binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryShape, nested / bson, make_nested, binary_format::bson);
BENCHMARK_CAPTURE(FromBinaryShape, containers / cbor, make_containers, binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryShape, containers / msgpack, make_containers, binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryShape, containers / ubjson, make_containers, binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryShape, containers / ubjson_optimized, make_containers, binary_format::ubjson_optimized);
BENCHMARK_CAPTURE(FromBinaryShape, containers / bjdata, make_containers, binary_format::bjdata);
BENCHMARK_CAPTURE(FromBinaryShape, containers / bson, make_containers, binary_format::bson);
// BSON names every array element, so a large array measures key generation
// rather than scalar decoding and is left out here
BENCHMARK_CAPTURE(FromBinaryShape, scalars / cbor, make_scalars, binary_format::cbor);
BENCHMARK_CAPTURE(FromBinaryShape, scalars / msgpack, make_scalars, binary_format::msgpack);
BENCHMARK_CAPTURE(FromBinaryShape, scalars / ubjson, make_scalars, binary_format::ubjson);
BENCHMARK_CAPTURE(FromBinaryShape, scalars / bjdata, make_scalars, binary_format::bjdata);
/*!
@brief parse an indefinite-length CBOR string
The writer never emits this form, so the input is assembled by hand: 0x7F
opens the string, each chunk is a one-character string, and 0xFF closes it.
*/
static void FromCborChunkedString(benchmark::State& state, const std::size_t chunks)
{
std::vector<std::uint8_t> bytes;
bytes.reserve(2 * chunks + 2);
bytes.push_back(0x7F);
for (std::size_t i = 0; i < chunks; ++i)
{
bytes.push_back(0x61); // string of length 1
bytes.push_back(0x61); // 'a'
}
bytes.push_back(0xFF);
for (auto _ : state)
{
json j = json::from_cbor(bytes);
benchmark::DoNotOptimize(j);
}
state.SetBytesProcessed(state.iterations() * bytes.size());
}
BENCHMARK_CAPTURE(FromCborChunkedString, 10000 chunks, 10000);
BENCHMARK_MAIN();
-51
View File
@@ -216,57 +216,6 @@ TEST_CASE("controlled bad_alloc")
CHECK_THROWS_AS(my_json(s), std::bad_alloc&);
next_construct_fails = false;
}
SECTION("basic_json(const basic_json&) of a deeply nested value (#5387)")
{
// Copying a value nested deeper than the descent bound builds the
// copy from the top down: every value whose own copy has not been
// made yet stays a null value until it is. Failing an allocation
// part-way through is what proves such a half-built copy can still
// be destroyed.
//
// Which path the failure lands in depends on the build: the first
// allocation of a copy belongs to the outermost level, so here it
// is the descending one. Built with JSON_NO_THREAD_LOCAL - as the
// ci_test_no_thread_local target builds the whole suite - no
// descent is made at all and the very same failure lands in the
// iterative path instead, part-way through its worklist.
const auto check_deep_copy = [](bool objects)
{
CAPTURE(objects);
next_construct_fails = false;
// deeper than the 128 levels the copy constructor descends into
const std::size_t depth = 300;
my_json j = 1;
for (std::size_t i = 0; i < depth; ++i)
{
if (objects)
{
my_json wrapper = my_json::object();
wrapper["a"] = std::move(j);
j = std::move(wrapper);
}
else
{
j = my_json::array({std::move(j)});
}
}
// NOLINTNEXTLINE(performance-unnecessary-copy-initialization): the copy is what is tested
CHECK_NOTHROW(my_json(j));
next_construct_fails = true;
// NOLINTNEXTLINE(performance-unnecessary-copy-initialization): the copy is what is tested
CHECK_THROWS_AS(my_json(j), std::bad_alloc&);
next_construct_fails = false;
};
check_deep_copy(false);
check_deep_copy(true);
}
}
}
+105
View File
@@ -3551,6 +3551,111 @@ TEST_CASE("BJData")
}
}
TEST_CASE("issue #5405 - array reserve for definite-length BJData arrays")
{
#if !defined(JSON_NOEXCEPTION)
// this SECTION relies on catching a thrown exception to distinguish
// which of two acceptable, bounded rejections a hostile header took;
// under JSON_NOEXCEPTION, JSON_THROW never produces a catchable C++
// exception (it aborts instead), so this cannot be tested that way here
SECTION("a huge claimed length with no element data must not over-allocate")
{
// optimized form [$type#count: type 'i' (int8), count as a four-byte
// little-endian 'l' (int32) of 0x7FFFFFFF (2147483647), but no
// element data at all. max_size() for a std::vector is far larger
// than this count, so it does not reject the header outright; the
// (capped) reservation must not attempt to allocate space for
// billions of elements before the missing data is detected.
json _;
const std::vector<uint8_t> input = {'[', '$', 'i', '#', 'l', 0xFF, 0xFF, 0xFF, 0x7F};
// On a platform where std::vector<json>::max_size() is smaller than
// the claimed count (e.g. 32-bit, where max_size() is bounded by a
// 32-bit SIZE_MAX divided by sizeof(json)), the SAX consumer's own
// check rejects the header outright (out_of_range.408, with the
// claimed count in the message) instead of accepting it and only
// finding it short of data once the (capped) reservation looks for
// element bytes that were never provided (parse_error.110). Either
// is an acceptable, bounded rejection of the hostile header -- the
// property under test is that no path attempts to allocate space
// for billions of elements.
bool threw = false;
try
{
_ = json::from_bjdata(input);
}
catch (const json::parse_error& e)
{
threw = true;
CHECK(e.id == 110);
CHECK(std::string(e.what()) == "[json.exception.parse_error.110] parse error at byte 10: syntax error while parsing BJData number: unexpected end of input");
}
catch (const json::out_of_range& e)
{
threw = true;
CHECK(e.id == 408);
CHECK(std::string(e.what()).find("excessive array size") != std::string::npos);
}
CHECK(threw);
// json_sax_dom_parser::start_array()'s max_size() check (unlike the
// scanner's own parse_error path) throws unconditionally via
// JSON_THROW rather than going through sax->parse_error(), so it is
// not gated by allow_exceptions=false on a platform where this
// header hits that check (e.g. 32-bit, see above) -- allow either
// a discarded result or the same out_of_range it throws with
// exceptions enabled.
try
{
CHECK(json::from_bjdata(input, true, false).is_discarded());
}
catch (const json::out_of_range& e)
{
CHECK(e.id == 408);
}
}
#endif
SECTION("arrays of various sizes decode to the same value as before the reserve optimization")
{
for (const auto size :
{
std::size_t{0}, std::size_t{1}, std::size_t{5}, // small
std::size_t{16384}, // exactly at the reserve cap
std::size_t{20000} // above the reserve cap
})
{
CAPTURE(size)
json j = json::array();
for (std::size_t i = 0; i < size; ++i)
{
j.push_back(static_cast<int>(i % 1000));
}
// exercise both the plain and the optimized [$type#count encoding
const auto packed_plain = json::to_bjdata(j);
CHECK(json::from_bjdata(packed_plain) == j);
const auto packed_optimized = json::to_bjdata(j, true, true);
CHECK(json::from_bjdata(packed_optimized) == j);
}
}
SECTION("a user-defined SAX consumer is unaffected by the internal DOM reserve optimization")
{
// the reserve() call is local to json_sax_dom_parser / json_sax_dom_callback_parser;
// a custom SAX consumer that does not touch a DOM array sees identical events
json j = json::array();
for (int i = 0; i < 100; ++i)
{
j.push_back(i);
}
const auto packed = json::to_bjdata(j, true, true);
SaxCountdown scp(1000000); // large enough to never trigger an abort
CHECK(json::sax_parse(packed, &scp, json::input_format_t::bjdata));
}
}
TEST_CASE("Universal Binary JSON Specification Examples 1")
{
SECTION("Null Value")
+86
View File
@@ -2174,6 +2174,92 @@ TEST_CASE("CBOR indefinite-length strings do not recurse per chunk")
}
}
TEST_CASE("issue #5405 - array reserve for definite-length CBOR arrays")
{
#if !defined(JSON_NOEXCEPTION)
// this SECTION relies on catching a thrown exception to distinguish
// which of two acceptable, bounded rejections a hostile header took;
// under JSON_NOEXCEPTION, JSON_THROW never produces a catchable C++
// exception (it aborts instead), so this cannot be tested that way here
SECTION("a huge claimed length with no element data must not over-allocate")
{
// 0x9A: array with a four-byte length; claims 0xFFFFFFFF (4294967295)
// elements but provides none. max_size() for a std::vector is far
// larger than this count, so it does not reject the header outright;
// the (capped) reservation must not attempt to allocate space for
// billions of elements before the missing data is detected.
json _;
const std::vector<uint8_t> input = {0x9A, 0xFF, 0xFF, 0xFF, 0xFF};
// On a platform where std::size_t is narrower than 64 bits (e.g.
// 32-bit), the claimed count 0xFFFFFFFF coincides with that
// platform's detail::unknown_size() sentinel (SIZE_MAX), so the
// format-level size check rejects it outright (out_of_range.408,
// "excessive ... size") before the SAX consumer's own max_size()
// check would even run; on a 64-bit platform it passes both of
// those checks and is only found short of data once the (capped)
// reservation looks for element bytes that were never provided
// (parse_error.110). Either is an acceptable, bounded rejection of
// the hostile header -- the property under test is that no path
// attempts to allocate space for billions of elements.
bool threw = false;
try
{
_ = json::from_cbor(input);
}
catch (const json::parse_error& e)
{
threw = true;
CHECK(e.id == 110);
CHECK(std::string(e.what()) == "[json.exception.parse_error.110] parse error at byte 6: syntax error while parsing CBOR value: unexpected end of input");
}
catch (const json::out_of_range& e)
{
threw = true;
CHECK(e.id == 408);
CHECK(std::string(e.what()).find("excessive") != std::string::npos);
}
CHECK(threw);
CHECK(json::from_cbor(input, true, false).is_discarded());
}
#endif
SECTION("arrays of various sizes decode to the same value as before the reserve optimization")
{
for (const auto size :
{
std::size_t{0}, std::size_t{1}, std::size_t{5}, // small
std::size_t{16384}, // exactly at the reserve cap
std::size_t{20000} // above the reserve cap
})
{
CAPTURE(size)
json j = json::array();
for (std::size_t i = 0; i < size; ++i)
{
j.push_back(static_cast<int>(i % 1000));
}
const auto packed = json::to_cbor(j);
CHECK(json::from_cbor(packed) == j);
}
}
SECTION("a user-defined SAX consumer is unaffected by the internal DOM reserve optimization")
{
// the reserve() call is local to json_sax_dom_parser / json_sax_dom_callback_parser;
// a custom SAX consumer that does not touch a DOM array sees identical events
json j = json::array();
for (int i = 0; i < 100; ++i)
{
j.push_back(i);
}
const auto packed = json::to_cbor(j);
SaxCountdown scp(1000000); // large enough to never trigger an abort
CHECK(json::sax_parse(packed, &scp, json::input_format_t::cbor));
}
}
TEST_CASE("CBOR roundtrips" * doctest::skip())
{
SECTION("input from flynn")
+80
View File
@@ -2259,6 +2259,86 @@ TEST_CASE("parser class")
#endif
}
#if JSON_DIAGNOSTIC_POSITIONS
TEST_CASE("diagnostic positions: value lifetime")
{
SECTION("copy constructor copies positions, recursively")
{
const std::string s = R"({"a":1,"b":[1,2,3]})";
const json a = json::parse(s);
const json b = a; // NOLINT(performance-unnecessary-copy-initialization)
CHECK(b.start_pos() == a.start_pos());
CHECK(b.end_pos() == a.end_pos());
CHECK(b["b"].start_pos() == a["b"].start_pos());
CHECK(b["b"].end_pos() == a["b"].end_pos());
}
SECTION("move constructor resets the moved-from value to npos")
{
const std::string s = R"({"a":1,"b":[1,2,3]})";
json a = json::parse(s);
const auto a_start = a.start_pos();
const auto a_end = a.end_pos();
const json b(std::move(a));
CHECK(b.start_pos() == a_start);
CHECK(b.end_pos() == a_end);
CHECK(a.start_pos() == std::string::npos); // NOLINT(bugprone-use-after-move,clang-analyzer-cplusplus.Move)
CHECK(a.end_pos() == std::string::npos); // NOLINT(bugprone-use-after-move,clang-analyzer-cplusplus.Move)
}
SECTION("swap() exchanges positions along with the values")
{
// basic_json::swap() (and the friend swap() that forwards to it) used
// to swap only m_data.m_type/m_data.m_value, leaving
// start_position/end_position untouched -- unlike copy-assignment's
// operator=(basic_json), which swaps positions as part of its
// copy-and-swap implementation. After swap(a, b), each value ended up
// with the *other* value's content but its *own* original position.
// This is now fixed so that swap() is consistent with copy-assignment.
json a = json::parse(R"({"a":1})");
json b = json::parse(R"([1,2,3,4,5])");
const auto a_start = a.start_pos();
const auto a_end = a.end_pos();
const auto b_start = b.start_pos();
const auto b_end = b.end_pos();
// lengths (and thus end positions) differ, which is enough to tell
// after the swap whether positions actually moved with the values
CHECK(a_end != b_end);
using std::swap;
swap(a, b);
CHECK(a == json::parse(R"([1,2,3,4,5])"));
CHECK(b == json::parse(R"({"a":1})"));
CHECK(a.start_pos() == b_start);
CHECK(a.end_pos() == b_end);
CHECK(b.start_pos() == a_start);
CHECK(b.end_pos() == a_end);
// member swap() behaves the same as the free function
json c = json::parse(R"({"a":1})");
json d = json::parse(R"([1,2,3,4,5])");
const auto c_start = c.start_pos();
const auto c_end = c.end_pos();
const auto d_start = d.start_pos();
const auto d_end = d.end_pos();
c.swap(d);
CHECK(c.start_pos() == d_start);
CHECK(c.end_pos() == d_end);
CHECK(d.start_pos() == c_start);
CHECK(d.end_pos() == c_end);
}
}
#endif
// this test relies on parse errors being thrown, so it is skipped when
// exceptions are disabled (json::parse aborts instead of throwing there)
#if !defined(JSON_NOEXCEPTION)
-66
View File
@@ -75,72 +75,6 @@ TEST_CASE("Better diagnostics with positions")
CHECK(j.end_pos() == root.size());
}
SECTION("copying keeps the positions of nested values (#5387)")
{
// Values nested deeper than the copy constructor's descent bound are
// copied without the call stack, on a path that has to carry the
// positions over itself; shallower ones copy their containers, which
// bring the positions along. Both sides of the bound are checked here.
const auto check_copy = [](std::size_t depth, bool objects)
{
CAPTURE(depth)
CAPTURE(objects)
const std::string opening = objects ? R"({"a":)" : "[";
const std::string closing = objects ? "}" : "]";
std::string text;
for (std::size_t i = 0; i < depth; ++i)
{
text += opening;
}
text += "12";
for (std::size_t i = 0; i < depth; ++i)
{
text += closing;
}
const json original = json::parse(text);
const json copy(original); // NOLINT(performance-unnecessary-copy-initialization)
const json* o = &original;
const json* c = &copy;
for (std::size_t level = 0; level <= depth; ++level)
{
CAPTURE(level)
REQUIRE(c->start_pos() == o->start_pos());
REQUIRE(c->end_pos() == o->end_pos());
if (level < depth)
{
o = objects ? &o->at("a") : &o->at(0);
c = objects ? &c->at("a") : &c->at(0);
}
}
};
const auto check_arrays = [&check_copy](std::size_t depth)
{
check_copy(depth, false);
};
const auto check_objects = [&check_copy](std::size_t depth)
{
check_copy(depth, true);
};
check_arrays(1);
check_arrays(127);
check_arrays(128);
check_arrays(129);
check_arrays(300);
check_objects(1);
check_objects(127);
check_objects(128);
check_objects(129);
check_objects(300);
}
SECTION("JSON patch add to primitive parent (#4292)")
{
// the JSON Patch "add" target /foo/bar/baz has a string parent
-57
View File
@@ -273,62 +273,5 @@ TEST_CASE("Regression tests for extended diagnostics")
CHECK(j1["numbers"]["two"] == 2);
CHECK(j1["string"] == "t");
}
SECTION("Regression test for issue #5387 - copying keeps the parents of nested values")
{
// A value nested deeper than the copy constructor's descent bound is
// copied without the call stack. Every container that path creates has
// to have the parents of its children set, or the JSON Pointer in the
// diagnostic is cut short.
const std::size_t depth = 300;
SECTION("objects")
{
json j = "not a number";
std::string pointer;
for (std::size_t i = 0; i < depth; ++i)
{
j = json{{"a", j}};
pointer += "/a";
}
json const copy(j); // NOLINT(performance-unnecessary-copy-initialization)
const json* inner = &copy;
for (std::size_t i = 0; i < depth; ++i)
{
inner = &inner->at("a");
}
std::string const expected = "[json.exception.type_error.302] (" + pointer + ") type must be number, but is string";
int i = 0;
CHECK_THROWS_WITH_AS(i = inner->get<int>(), expected.c_str(), json::type_error);
CHECK(i == 0);
}
SECTION("arrays")
{
json j = "not a number";
std::string pointer;
for (std::size_t i = 0; i < depth; ++i)
{
j = json::array({j});
pointer += "/0";
}
json const copy(j); // NOLINT(performance-unnecessary-copy-initialization)
const json* inner = &copy;
for (std::size_t i = 0; i < depth; ++i)
{
inner = &inner->at(0);
}
std::string const expected = "[json.exception.type_error.302] (" + pointer + ") type must be number, but is string";
int i = 0;
CHECK_THROWS_WITH_AS(i = inner->get<int>(), expected.c_str(), json::type_error);
CHECK(i == 0);
}
}
}
-151
View File
@@ -12,7 +12,6 @@
using nlohmann::json;
#include <algorithm>
#include <string>
TEST_CASE("tests on very large JSONs")
{
@@ -28,153 +27,3 @@ TEST_CASE("tests on very large JSONs")
}
}
namespace
{
// Descend a chain of single-element containers and return the value at its end,
// reporting the number of levels traversed in @a depth.
//
// The values in the test case below are nested far deeper than the call stack
// can follow, so they must not be inspected with operator== or dump(): both are
// still recursive and would overflow the stack themselves.
const json* innermost_value(const json& j, std::size_t& depth)
{
const json* current = &j;
depth = 0;
while ((current->is_array() || current->is_object()) && !current->empty())
{
current = current->is_array()
? &current->front()
: &current->begin().value();
++depth;
}
return current;
}
} // namespace
TEST_CASE("tests on deeply nested JSONs")
{
// deep enough to exhaust the call stack, but small enough to stay cheap:
// parsing is iterative, so building the values below costs little
const std::size_t depth = 100000;
SECTION("issue #5387 - stack overflow in the copy constructor")
{
SECTION("array")
{
const json j = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
const json copy(j); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t copy_depth = 0;
CHECK(*innermost_value(copy, copy_depth) == 0);
CHECK(copy_depth == depth);
}
SECTION("object")
{
std::string s;
s.reserve((6 * depth) + 1);
for (std::size_t i = 0; i < depth; ++i)
{
s += "{\"a\":";
}
s += '1';
s.append(depth, '}');
const json j = json::parse(s);
const json copy(j); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t copy_depth = 0;
CHECK(*innermost_value(copy, copy_depth) == 1);
CHECK(copy_depth == depth);
}
SECTION("copy assignment")
{
// operator=(basic_json) takes its argument by value, so the deep
// copy happens in the copy constructor
const json j = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
json target;
target = j;
std::size_t target_depth = 0;
CHECK(*innermost_value(target, target_depth) == 0);
CHECK(target_depth == depth);
}
SECTION("depths around the bound of the recursive descent")
{
// The copy constructor descends into a bounded number of levels and
// completes whatever is below that without the call stack. Cover
// every depth around that bound, so that the two ways of copying
// are known to meet cleanly - wherever the bound is set.
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d);
const json array = json::parse(std::string(d, '[') + '0' + std::string(d, ']'));
const json array_copy(array); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t array_depth = 0;
CHECK(*innermost_value(array_copy, array_depth) == 0);
CHECK(array_depth == d);
std::string object_text;
for (std::size_t i = 0; i < d; ++i)
{
object_text += "{\"a\":";
}
object_text += '1';
object_text.append(d, '}');
const json object = json::parse(object_text);
const json object_copy(object); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t object_depth = 0;
CHECK(*innermost_value(object_copy, object_depth) == 1);
CHECK(object_depth == d);
}
}
SECTION("a value that is deep in one place only")
{
json j = json::object();
j["shallow"] = 1;
j["deep"] = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
j["also_shallow"] = json::array({1, 2, 3});
const json copy(j);
CHECK(copy["shallow"] == 1);
CHECK(copy["also_shallow"] == json::array({1, 2, 3}));
std::size_t deep_depth = 0;
CHECK(*innermost_value(copy["deep"], deep_depth) == 0);
CHECK(deep_depth == depth);
}
SECTION("the copy is independent of the original")
{
const json j = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
json copy(j);
// reach the innermost value without recursing and replace it
json* current = &copy;
while (current->is_array() && !current->empty())
{
current = &current->front();
}
*current = 42;
std::size_t unused = 0;
CHECK(*innermost_value(copy, unused) == 42);
CHECK(*innermost_value(j, unused) == 0);
}
}
}
+85
View File
@@ -1597,6 +1597,91 @@ TEST_CASE("MessagePack")
}
}
TEST_CASE("issue #5405 - array reserve for definite-length MessagePack arrays")
{
#if !defined(JSON_NOEXCEPTION)
// this SECTION relies on catching a thrown exception to distinguish
// which of two acceptable, bounded rejections a hostile header took;
// under JSON_NOEXCEPTION, JSON_THROW never produces a catchable C++
// exception (it aborts instead), so this cannot be tested that way here
SECTION("a huge claimed length with no element data must not over-allocate")
{
// 0xdd: array 32 (four-byte length); claims 0xFFFFFFFF (4294967295)
// elements but provides none. max_size() for a std::vector is far
// larger than this count, so it does not reject the header outright;
// the (capped) reservation must not attempt to allocate space for
// billions of elements before the missing data is detected.
json _;
const std::vector<uint8_t> input = {0xdd, 0xFF, 0xFF, 0xFF, 0xFF};
// On a platform where std::size_t is narrower than 64 bits (e.g.
// 32-bit), the claimed count 0xFFFFFFFF coincides with that
// platform's SIZE_MAX, which some size-narrowing checks treat the
// same as detail::unknown_size(); it may then be rejected before
// the SAX consumer's own max_size() check (out_of_range.408) rather
// than being accepted and only found short of data once the
// (capped) reservation looks for element bytes that were never
// provided (parse_error.110). Either is an acceptable, bounded
// rejection of the hostile header -- the property under test is
// that no path attempts to allocate space for billions of elements.
bool threw = false;
try
{
_ = json::from_msgpack(input);
}
catch (const json::parse_error& e)
{
threw = true;
CHECK(e.id == 110);
CHECK(std::string(e.what()) == "[json.exception.parse_error.110] parse error at byte 6: syntax error while parsing MessagePack value: unexpected end of input");
}
catch (const json::out_of_range& e)
{
threw = true;
CHECK(e.id == 408);
CHECK(std::string(e.what()).find("excessive") != std::string::npos);
}
CHECK(threw);
CHECK(json::from_msgpack(input, true, false).is_discarded());
}
#endif
SECTION("arrays of various sizes decode to the same value as before the reserve optimization")
{
for (const auto size :
{
std::size_t{0}, std::size_t{1}, std::size_t{5}, // small
std::size_t{16384}, // exactly at the reserve cap
std::size_t{20000} // above the reserve cap
})
{
CAPTURE(size)
json j = json::array();
for (std::size_t i = 0; i < size; ++i)
{
j.push_back(static_cast<int>(i % 1000));
}
const auto packed = json::to_msgpack(j);
CHECK(json::from_msgpack(packed) == j);
}
}
SECTION("a user-defined SAX consumer is unaffected by the internal DOM reserve optimization")
{
// the reserve() call is local to json_sax_dom_parser / json_sax_dom_callback_parser;
// a custom SAX consumer that does not touch a DOM array sees identical events
json j = json::array();
for (int i = 0; i < 100; ++i)
{
j.push_back(i);
}
const auto packed = json::to_msgpack(j);
SaxCountdown scp(1000000); // large enough to never trigger an abort
CHECK(json::sax_parse(packed, &scp, json::input_format_t::msgpack));
}
}
// use this testcase outside [hide] to run it with Valgrind
TEST_CASE("MessagePack nesting does not consume the call stack")
{
-34
View File
@@ -81,37 +81,3 @@ TEST_CASE("regression test for issue #3732 - iteration_proxy_value<iter_impl<ord
};
static_cast<void>(fn);
}
TEST_CASE("copying an ordered_json with nested values")
{
// ordered_map is backed by a vector, so copying an object that has
// structured values takes a different route than copying a std::map-backed
// one; see https://github.com/nlohmann/json/issues/5387
ordered_json oj;
oj["z"] = 1;
oj["a"]["y"] = 2;
oj["a"]["b"]["x"] = 3;
oj["m"] = {1, 2, {{"w", 4}}};
const ordered_json copy(oj);
SECTION("the copy is equal to the original")
{
CHECK(copy == oj);
CHECK(copy.dump() == oj.dump());
}
SECTION("the key order is preserved at every level")
{
CHECK(copy.dump() == R"({"z":1,"a":{"y":2,"b":{"x":3}},"m":[1,2,{"w":4}]})");
}
SECTION("the copy is independent of the original")
{
ordered_json mutated(oj);
mutated["a"]["b"]["x"] = 99;
CHECK(oj["a"]["b"]["x"] == 3);
CHECK(mutated["a"]["b"]["x"] == 99);
}
}
+46
View File
@@ -27,6 +27,7 @@ using ordered_json = nlohmann::ordered_json;
#endif
#include <cstdio>
#include <deque>
#include <list>
#include <type_traits>
#include <utility>
@@ -896,4 +897,49 @@ TEST_CASE("issue #5402 - update(merge_objects=true) overwrites a primitive with
}
TEST_CASE("regression test #5476 - array type without reserve()")
{
// the capacity reserved for definite-length arrays must not require the
// array type to have a reserve() member function
using deque_json = nlohmann::basic_json<std::map, std::deque>;
SECTION("std::deque")
{
const auto j = deque_json::parse(R"({"a":[1,[2,3]],"b":[]})");
CHECK(j.dump() == R"({"a":[1,[2,3]],"b":[]})");
// the binary formats pass a definite length to start_array()
CHECK(deque_json::from_cbor(deque_json::to_cbor(j)) == j);
CHECK(deque_json::from_msgpack(deque_json::to_msgpack(j)) == j);
// parse() instantiates the callback parser as well, which reserves too
const auto with_callback = deque_json::parse(R"([1,2,3])", [](int /*depth*/, deque_json::parse_event_t /*event*/, deque_json& /*parsed*/) noexcept
{
return true;
});
CHECK(with_callback == deque_json({1, 2, 3}));
}
SECTION("std::vector still reserves")
{
json array = json::array();
for (int i = 0; i < 100; ++i)
{
array.push_back(i);
}
const auto j = json::from_cbor(json::to_cbor(array));
CHECK(j == array);
CHECK(j.get_ref<const json::array_t&>().capacity() >= 100);
}
SECTION("the reservation stays capped")
{
// CBOR array announcing 2^32-1 elements, but truncated right after the
// header: the input must be rejected without reserving that capacity
const std::vector<std::uint8_t> truncated = {0x9A, 0xFF, 0xFF, 0xFF, 0xFF};
CHECK(json::from_cbor(truncated, true, false).is_discarded());
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
+1 -1
View File
@@ -469,7 +469,7 @@ TEST_CASE("serialization of strings (bulk fast path)")
SECTION("invalid UTF-8 handling is unaffected by the fast path")
{
const json j = std::string("valid\xff" "more");
CHECK_THROWS_WITH_AS(j.dump(), "[json.exception.type_error.316] invalid UTF-8 byte at index 5: 0xFF", json::type_error&);
CHECK_THROWS_WITH_AS(utils::ignore_return_value(j.dump()), "[json.exception.type_error.316] invalid UTF-8 byte at index 5: 0xFF", json::type_error&);
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"valid\xef\xbf\xbd" "more\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"valid\\ufffdmore\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"validmore\"");
+106
View File
@@ -2315,6 +2315,112 @@ TEST_CASE("UBJSON optimized arrays of a valueless type are bounded")
}
}
TEST_CASE("issue #5405 - array reserve for definite-length UBJSON arrays")
{
#if !defined(JSON_NOEXCEPTION)
// this SECTION relies on catching a thrown exception to distinguish
// which of two acceptable, bounded rejections a hostile header took;
// under JSON_NOEXCEPTION, JSON_THROW never produces a catchable C++
// exception (it aborts instead), so this cannot be tested that way here
SECTION("a huge claimed length with no element data must not over-allocate")
{
// optimized form [$type#count: type 'i' (int8), count as a four-byte
// 'l' (int32) of 0x7FFFFFFF (2147483647), but no element data at all.
// max_size() for a std::vector is far larger than this count, so it
// does not reject the header outright; the (capped) reservation must
// not attempt to allocate space for billions of elements before the
// missing data is detected.
json _;
const std::vector<uint8_t> input = {'[', '$', 'i', '#', 'l', 0x7F, 0xFF, 0xFF, 0xFF};
// On a platform where std::vector<json>::max_size() is smaller than
// the claimed count (e.g. 32-bit, where max_size() is bounded by a
// 32-bit SIZE_MAX divided by sizeof(json)), the SAX consumer's own
// check rejects the header outright (out_of_range.408, with the
// claimed count in the message) instead of accepting it and only
// finding it short of data once the (capped) reservation looks for
// element bytes that were never provided (parse_error.110). Either
// is an acceptable, bounded rejection of the hostile header -- the
// property under test is that no path attempts to allocate space
// for billions of elements.
bool threw = false;
try
{
_ = json::from_ubjson(input);
}
catch (const json::parse_error& e)
{
threw = true;
CHECK(e.id == 110);
CHECK(std::string(e.what()) == "[json.exception.parse_error.110] parse error at byte 10: syntax error while parsing UBJSON number: unexpected end of input");
}
catch (const json::out_of_range& e)
{
threw = true;
CHECK(e.id == 408);
CHECK(std::string(e.what()).find("excessive array size") != std::string::npos);
}
CHECK(threw);
// json_sax_dom_parser::start_array()'s max_size() check (unlike the
// scanner's own parse_error path) throws unconditionally via
// JSON_THROW rather than going through sax->parse_error(), so it is
// not gated by allow_exceptions=false on a platform where this
// header hits that check (e.g. 32-bit, see above) -- allow either
// a discarded result or the same out_of_range it throws with
// exceptions enabled.
try
{
CHECK(json::from_ubjson(input, true, false).is_discarded());
}
catch (const json::out_of_range& e)
{
CHECK(e.id == 408);
}
}
#endif
SECTION("arrays of various sizes decode to the same value as before the reserve optimization")
{
for (const auto size :
{
std::size_t{0}, std::size_t{1}, std::size_t{5}, // small
std::size_t{16384}, // exactly at the reserve cap
std::size_t{20000} // above the reserve cap
})
{
CAPTURE(size)
json j = json::array();
for (std::size_t i = 0; i < size; ++i)
{
j.push_back(static_cast<int>(i % 1000));
}
// exercise both the plain and the optimized [$type#count encoding
const auto packed_plain = json::to_ubjson(j);
CHECK(json::from_ubjson(packed_plain) == j);
const auto packed_optimized = json::to_ubjson(j, true, true);
CHECK(json::from_ubjson(packed_optimized) == j);
}
}
SECTION("a user-defined SAX consumer is unaffected by the internal DOM reserve optimization")
{
// the reserve() call is local to json_sax_dom_parser / json_sax_dom_callback_parser;
// a custom SAX consumer that does not touch a DOM array sees identical events
json j = json::array();
for (int i = 0; i < 100; ++i)
{
j.push_back(i);
}
const auto packed = json::to_ubjson(j, true, true);
SaxCountdown scp(1000000); // large enough to never trigger an abort
CHECK(json::sax_parse(packed, &scp, json::input_format_t::ubjson));
}
}
TEST_CASE("Universal Binary JSON Specification Examples 1")
{
SECTION("Null Value")