Compare commits

..
Author SHA1 Message Date
Niels Lohmann 7bfb3aafd1 Name the test's locals so Flawfinder stops matching them
The code scanning job reports CWE-362 - "check when opening files" - for
a test that opens no files: Flawfinder matched a local variable called
open. Rename it and its partner.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 08:50:03 +02:00
Niels Lohmann b905e81cc1 Check that an abandoned copy can still be destroyed
Copying a value without the call stack builds the copy from the top down,
and every value whose own copy has not been made yet stays a null value
until it is. That is what lets a copy be abandoned half-built: the
destructor finds nothing but complete values and null ones.

Nothing tested it. Failing an allocation part-way through a copy of a
deeply nested value does, with the allocator the file already has for
exactly this kind of test.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 08:50:03 +02:00
Niels Lohmann ad3b07f7e5 Keep the descent bookkeeping in one place
Copying carried a depth count, a depth limit and a guard of its own, and
the comparison in the follow-up added a second set beside them. Neither
operation needs its own: they are never nested inside one another by the
library - copying a value does not compare one, and comparing two values
does not copy them - and where user code nests them anyway, sharing the
count only ends a descent sooner than it had to.

So there is now one nesting_depth(), one nesting_depth_limit(), one
nesting_depth_exhausted() and one nesting_depth_guard, which the follow-up
uses instead of adding its own. Inverting the test in copy_structured
leaves the too-deep case and the no-thread-local case as the same code.

The guard takes the count where the caller has already looked it up to
test it, and looks it up itself where the caller cannot - the comparison
operators are written as a macro, and a macro cannot use the preprocessor.
Whether an operator descends at all is passed to nesting_depth_exhausted()
rather than tested at the call site, where the constant makes MSVC report
C4127.

The switch that copies the value of anything that is not an object or an
array was written twice - once in the copy constructor, once in
copy_shallow - so that adding a value_t meant editing both, and missing
one would have been silent. It is copy_leaf_value now, and inlined: both
callers have already sorted the containers out, and folding that test into
the switch is what keeps a value made mostly of numbers copying as fast as
it did.

Copying canada.json, citm_catalog.json and twitter.json is within 0.6% of
what it was before, measured as a paired ratio over 18 interleaved rounds
against a run-to-run spread of 0.3%.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 08:49:54 +02:00
Niels Lohmann e66a55066e Include <span> where the split moved its only use
The #2546 test case guards itself with __has_include(<span>), but the
include itself sat in unit-regression2.cpp's preamble and stayed behind,
so the section compiled without a declaration wherever the guard passed -
which nvhpc reported and libc++ builds do not, as they skip the section
altogether.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 03:19:53 +02:00
Niels Lohmann 02013c7f0d Move the #4804 alias to the file that uses it
The split left the json_4804 alias behind in unit-regression2.cpp while
the test case that uses it went to unit-regression3.cpp, which does not
build for C++17 and C++20 as a result.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 03:00:18 +02:00
Niels Lohmann cdf24bccde Split the regression tests far enough to leave room
The first split left unit-regression2.cpp 0.7% below the size develop
links at, which the comparison change in the follow-up immediately used
up: the MinGW linker fails on test-regression2_cpp20 again, naming
copy_shallow and to_partial_ordering among the relocations it cannot fit.

Move the sections from "issue #2067" on, and the helper types they use,
so that the file stops being the one that decides whether the tests can
be linked at all. At -O0 and C++20, unit-regression2.cpp is now 2,964,944
bytes against develop's 4,708,248, and 3,070,568 bytes with the follow-up
applied - roughly a third smaller either way, rather than a fraction of a
percent larger.

The 135 assertions are the same ones as before, now spread over three
test cases in two files.

Also silence the clang-tidy findings the deep-nesting tests draw: the
copies they make are what is being tested, and the reserve() computation
gets its parentheses.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 02:53:10 +02:00
Niels Lohmann 9d75c87de5 Check both shapes without a C-style array
clang-tidy rejects the array the two shapes were iterated over
(cppcoreguidelines-avoid-c-arrays). The array only existed because astyle
reformats a range-for over a braced initializer list into something
unreadable; naming the two cases avoids both.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 01:41:05 +02:00
Niels Lohmann 4fd940f91b Balance the warning suppression the split separated
unit-regression2.cpp opens a DOCTEST_CLANG_SUPPRESS_WARNING_PUSH block at
the top and closed it at the very bottom, which the split moved into
unit-regression3.cpp: one file was left with a push and no pop, the other
with a pop and no push, which clang reports as an error.

Give each file the pair it needs.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 01:01:57 +02:00
Niels Lohmann 540f6d11cc Do not use thread_local storage with Clang targeting MinGW
Every test that copies a value segfaults there - 42 of 105 on clang
11.0.1, 39 of 102 on clang 18.1.8 - while the same tests pass with GCC
targeting MinGW, with Clang targeting MSVC, and with every other
toolchain the library is tested on. The counter that bounds the copy
constructor's descent is the library's first use of thread_local, so
that job had never exercised it before.

JSON_NO_THREAD_LOCAL already covers toolchains without thread_local
storage, and copying yields the same values with it, only more slowly.
Define it for this one automatically.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 00:48:42 +02:00
Niels Lohmann 6305cc9f72 Split the regression tests so that they keep linking
Linking test-regression2 fails with "relocation truncated to fit:
IMAGE_REL_AMD64_REL32 against `.rdata'" once its object grows past what
the MinGW linker copes with, and the copy constructor's helpers push it
over: the object grows by 6.3%, from 4,654,128 to 4,944,920 bytes at -O0,
and develop links at the smaller of the two.

Building the tests optimized shrinks the object enough to link, but the
binaries clang 11.0.1 and clang 18.1.8 then produce crash before doctest
prints its first line - 39 of 102 tests on clang 18 - so the objects have
to become smaller rather than denser.

Moving the test cases that follow "regression tests 2" into a file of
their own brings that object to 4,687,888 bytes, which is 0.7% above the
size that links today rather than 6.3%. Both files still build for C++11,
C++17 and C++20, and run the same 9 test cases and 135 assertions as
before, now spread over two binaries.

New regression tests belong in unit-regression3.cpp from here on, which
is what CONTRIBUTING.md now says.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 00:36:05 +02:00
Niels Lohmann fa9b76283a Test the copy constructor's iterative path in CI
The copy constructor descends into 128 levels before it finishes a value
without the call stack, so the iterative path is otherwise only reached
by the few tests that nest deeper than that.

JSON_NO_THREAD_LOCAL switches the descent off, which sends every value
down that path. Running the whole test suite that way covers it with
every object type, string type, allocator, and base class the suite
already exercises. The new ci_test_no_thread_local target does that; the
macro had no build coverage at all before.

Copying a nested value also has to carry over what the element-wise copy
constructor would have copied: the parents that JSON_DIAGNOSTICS relies
on, and the positions that JSON_DIAGNOSTIC_POSITIONS reports. Both are
now checked on either side of the descent bound, for objects and arrays.
Neither was tested before, and dropping either one makes the new tests
fail.

Also quantify what JSON_NO_THREAD_LOCAL costs a copy instead of calling
it "measurably slower".

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-21 00:35:45 +02:00
Niels Lohmann e486005583 Bound the descent of the copy constructor
basic_json's copy constructor copied objects and arrays by handing the
container to its own copy constructor, which copy-constructs every element
and so reaches this constructor again, once per nesting level. A value
nested deeply enough exhausted the call stack and terminated the process
with a segmentation fault - no exception, nothing the caller could catch.
Parsing such a value works, as the parser is iterative, and so does
destroying one, as #1436 made destruction iterative.

Bound how far the copy descends rather than take the call stack away from
it. The first levels are copied exactly as they were - the containers copy
their own elements, which is by far the fastest way to fill them - and only
once the copy has descended 128 levels is the value below it finished
without the call stack, through an explicit worklist. Copying can therefore
no longer exhaust the stack, however deeply a value is nested, while a value
nested less deeply than the bound - all but a vanishing minority - is copied
by the very same code as before and pays only for one counter.

That counter lives in thread_local storage, as one shared between threads
would be raced. JSON_NO_THREAD_LOCAL switches it off for toolchains without
thread_local; copying then goes through the worklist right away, which
yields the same values but is measurably slower.

The deferred values are completed before the copy they belong to returns, so
a value copied while another copy is going on - by a custom base class, say -
is unaffected by the copy it is nested in.

operator= takes its argument by value, so copy assignment is fixed as well.

Copying is as fast as it was, within measurement noise (medians of 9
interleaved runs, clang -O3): -1.3% for an array of strings, +0.0% for a
flat object, +0.1% for a flat array of numbers, +0.3% for nested arrays,
+0.6% for nested objects and +1.2% for a twitter-like document. Copying a
three-key object costs about ten nanoseconds more, the counter. Deferring
every level instead, rather than only those below the bound, measured
between 3% and 9% slower depending on the shape of the value.

This fixes #5387 for the copy constructor. dump() is still recursive.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-20 19:22:16 +02:00
Niels LohmannandGitHub 734fd305a1 Format-check the documentation examples in CI (#5386)
* Reformat parser_callback_t example with astyle

The file uses "json & /*parsed*/" in three lambda parameter lists, which
astyle rewrites to "json& /*parsed*/" per --align-reference=type. The
drift went unnoticed because CI never format-checked the documentation
examples; "make pretty" does cover them.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Format-check the documentation examples in CI

The examples live in docs/mkdocs/docs/examples, but both format checks
still referenced the long-gone docs/examples path:

- check_amalgamation.yml passed it to find, which printed an error for
  the missing path and carried on, so astyle only ever saw include and
  tests. The step still exited 0.
- ci.cmake globbed it into INDENT_FILES, and a GLOB_RECURSE over a
  missing directory silently yields nothing, so the ci_test_amalgamation
  target skipped the examples too.

Either way the 231 example files have never been format-checked. Point
both at the real path, and guard the workflow with an explicit directory
check so a future rename fails the job instead of quietly shrinking the
file list again.

Also drop the dead docs/examples/** path filter from
publish_documentation.yml; docs/mkdocs/** already covers the examples.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-20 12:32:50 +02:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
36187cacfb ⬆️ Bump wheel from 0.47.0 to 0.48.0 in /docs/mkdocs (#5385)
Bumps [wheel](https://github.com/pypa/wheel) from 0.47.0 to 0.48.0.
- [Release notes](https://github.com/pypa/wheel/releases)
- [Changelog](https://github.com/pypa/wheel/blob/main/docs/news.rst)
- [Commits](https://github.com/pypa/wheel/compare/0.47.0...0.48.0)

---
updated-dependencies:
- dependency-name: wheel
  dependency-version: 0.48.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-20 12:32:34 +02:00
Sahil_KamateandGitHub b5378e8deb Fix CBOR tag handlers not recognizing tags 0-5 and 21-23 (#5331)
* Fix CBOR tag handlers not recognizing tags 0-5 and 21-23

The tagged-item switch in binary_reader::parse_cbor_internal() only handled
head bytes 0xC6-0xD4 and 0xD8-0xDB. Bytes 0xC0-0xC5 (tags 0-5: date/time,
epoch, bignum, decimal, bigfloat) and 0xD5-0xD7 (tags 21-23: base64url,
base64, base16 conversion hints) fell through to the default case and were
reported as invalid bytes, even under cbor_tag_handler_t::ignore and ::store,
despite being valid CBOR major-type-6 tags per RFC 8949.

Add the missing case labels so the full 0xC0-0xDB range is handled
uniformly. Extend the "Tagged values" test in unit-cbor.cpp to cover
0xC0-0xD7, and update the CBOR docs to state the corrected tag range.

Fixes #5315

Signed-off-by: sahilkamate03 <45514385+sahilkamate03@users.noreply.github.com>

* Fix stale CBOR tag docs and add store-mode binary-payload test

The "Incomplete mapping" warning still listed tags 0-5 (date/time,
bignum, decimal fraction, bigfloat) and 21-23 (expected conversions)
as unsupported, even though they now parse correctly under
cbor_tag_handler_t::ignore/store, same as 0xC6..0xD4/0xD8..0xDB.
Remove those five bullets and cross-reference the "Tagged items"
warning below, matching the equivalent docs fix landed independently
in PR #5367.

Also add a cbor_tag_handler_t::store test that wraps a binary
payload (not just a string) for every byte in 0xC0..0xD7, confirming
these tags are unwrapped the same way as 0xC6..0xD4 rather than
mistaken for the 0xD8..0xDB binary-subtype marker syntax, per review
feedback on #5331.

Signed-off-by: sahilkamate03 <45514385+sahilkamate03@users.noreply.github.com>

---------

Signed-off-by: sahilkamate03 <45514385+sahilkamate03@users.noreply.github.com>
2026-08-19 20:19:47 +02:00
31 changed files with 2553 additions and 3730 deletions
+3 -1
View File
@@ -108,7 +108,9 @@ The tests are located in [`tests/src/unit-*.cpp`](https://github.com/nlohmann/js
are structured along the features of the library or the nature of the tests. Usually, it should be clear from the
context which existing file needs to be extended, and only very few cases require creating new test files.
When fixing a bug, edit `unit-regression2.cpp` and add a section referencing the fixed issue.
When fixing a bug, edit `unit-regression3.cpp` and add a test case referencing the fixed issue. Its predecessors
`unit-regression1.cpp` and `unit-regression2.cpp` stay as they are: the MinGW linker fails on the object a file this
size produces, which is why the tests are spread over several files in the first place.
#### Exceptions
+11 -1
View File
@@ -67,8 +67,18 @@ jobs:
${{ github.workspace }}/venv/bin/astyle --project=tools/astyle/.astylerc --suffix=none --quiet \
$INCLUDE_DIR/json.hpp $INCLUDE_DIR/json_fwd.hpp
# fail loudly if a directory is renamed or removed: find would only warn
# about the missing path and silently drop its files from the check
SOURCE_DIRS="docs/mkdocs/docs/examples include tests"
for DIR in $SOURCE_DIRS; do
if [ ! -d "$DIR" ]; then
echo "::error::source directory '$DIR' does not exist"
exit 1
fi
done
${{ github.workspace }}/venv/bin/astyle --project=tools/astyle/.astylerc --suffix=none --quiet \
$(find docs/examples include tests -type f \( -name '*.hpp' -o -name '*.cpp' -o -name '*.cu' \) -not -path 'tests/thirdparty/*' -not -path 'tests/abi/include/nlohmann/*' | sort)
$(find $SOURCE_DIRS -type f \( -name '*.hpp' -o -name '*.cpp' -o -name '*.cu' \) -not -path 'tests/thirdparty/*' -not -path 'tests/abi/include/nlohmann/*' | sort)
- name: Build patch and check for differences
id: diff
@@ -7,7 +7,6 @@ on:
- develop
paths:
- docs/mkdocs/**
- docs/examples/**
workflow_dispatch:
# we don't want to have concurrent jobs, and we don't want to cancel running jobs to avoid broken publications
+1 -1
View File
@@ -100,7 +100,7 @@ jobs:
container: ubuntu:focal
strategy:
matrix:
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls]
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_no_thread_local]
steps:
- name: Install build-essential
run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev
+4
View File
@@ -158,6 +158,10 @@ jobs:
# to fit: IMAGE_REL_AMD64_SECREL against `.debug_line'" because the
# MinGW linker cannot relocate the debug sections this test produces.
# The tests are only built and run here, so the debug info is not used.
# Do not add -O1 here to shrink the objects further: it does make them
# link, but the binaries clang 11.0.1 and clang 18.1.8 then produce crash
# before doctest prints its first line - 39 of 102 tests on clang 18.
# Keep the objects small by splitting the test files instead.
- name: Run CMake
run: cmake -S . -B build ^
-DCMAKE_CXX_COMPILER="C:/Program Files/LLVM/bin/clang++.exe" ^
+20 -1
View File
@@ -242,6 +242,25 @@ add_custom_target(ci_test_noglobaludls
COMMENT "Compile and test with global UDLs disabled"
)
###############################################################################
# Disable thread-local storage.
###############################################################################
# Without thread-local storage, the copy constructor cannot bound its descent
# and copies every object and array without the call stack. That path is
# otherwise only reached by values nested deeper than the bound, so this target
# is what runs the whole test suite through it.
add_custom_target(ci_test_no_thread_local
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON
-DCMAKE_CXX_FLAGS=-DJSON_NO_THREAD_LOCAL
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_no_thread_local
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_no_thread_local
COMMAND cd ${PROJECT_BINARY_DIR}/build_no_thread_local && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure
COMMENT "Compile and test without thread-local storage"
)
###############################################################################
# Coverage.
###############################################################################
@@ -294,7 +313,7 @@ file(GLOB_RECURSE INDENT_FILES
${PROJECT_SOURCE_DIR}/tests/src/*.cpp
${PROJECT_SOURCE_DIR}/tests/src/*.hpp
${PROJECT_SOURCE_DIR}/tests/benchmarks/src/benchmarks.cpp
${PROJECT_SOURCE_DIR}/docs/examples/*.cpp
${PROJECT_SOURCE_DIR}/docs/mkdocs/docs/examples/*.cpp
)
set(include_dir ${PROJECT_SOURCE_DIR}/single_include/nlohmann)
+1 -1
View File
@@ -22,9 +22,9 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_HAS_STD_FORMAT**](json_has_std_format.md) - control `std::format`/`std::formatter` support
- [**JSON_HAS_THREE_WAY_COMPARISON**](json_has_three_way_comparison.md) - control 3-way comparison support
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
- [**JSON_NO_THREAD_LOCAL**](json_no_thread_local.md) - switch off the use of `thread_local` storage
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
- [**JSON_USE_SIMDUTF**](json_use_simdutf.md) - use the simdutf library to accelerate UTF-8 validation
## Library version
@@ -0,0 +1,47 @@
# JSON_NO_THREAD_LOCAL
```cpp
#define JSON_NO_THREAD_LOCAL
```
When defined, the library does not use `#!cpp thread_local` storage. This is relevant for the few environments whose
toolchain does not support it.
The copy constructor copies the first levels of a value by copying the containers, which copy their elements, and
completes whatever is nested deeper than that without the call stack, so that copying a value cannot exhaust the stack
however deeply it is nested. It counts the levels it has descended into in a `#!cpp thread_local` variable, as a counter
shared between threads would be raced.
Without that counter, no descent can be bounded safely, so objects and arrays are copied without the call stack right
away. Copying keeps working exactly as it does otherwise - the same values come out, and deeply nested values are copied
just as safely - but copying is slower, because the containers no longer copy themselves. Copying the benchmark
documents takes 9% (`canada.json`) to 34% (`twitter.json`) longer; values built mostly from objects are affected the
most.
## Default definition
By default, `#!cpp JSON_NO_THREAD_LOCAL` is not defined.
```cpp
#undef JSON_NO_THREAD_LOCAL
```
The library defines it by itself for Clang targeting MinGW, which does not survive the `#!cpp thread_local` storage:
copying a value segfaults there, with both old and current Clang versions, while GCC targeting MinGW is unaffected.
## Examples
??? example
The code below forces the library not to use `#!cpp thread_local` storage.
```cpp
#define JSON_NO_THREAD_LOCAL 1
#include <nlohmann/json.hpp>
...
```
## Version history
- Added in version 3.12.1.
@@ -1,59 +0,0 @@
# JSON_USE_SIMDUTF
```cpp
#define JSON_USE_SIMDUTF
```
When defined, the parser validates the UTF-8 content of JSON strings that come from a **contiguous byte input**
(`std::string`, `std::vector<char>`/`<std::uint8_t>`, string literals, `const char*` ranges, …) using the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. On text with many
non-ASCII characters (e.g. CJK or emoji) this can validate several times faster.
This is an **opt-in external dependency**. The library itself remains header-only and its behavior is unchanged: the
same input is accepted or rejected either way, and every parse error is reported at the same position with the same
message (simdutf is only used to fast-path *valid* runs; anything it flags falls back to the scalar path so the exact
diagnostic is preserved). Streaming inputs (files, `std::istream`, wide strings, user-defined adapters) always use the
scalar path.
When `JSON_USE_SIMDUTF` is defined you must make the `simdutf.h` header available on the include path and link the
simdutf library. When it is not defined, no simdutf header is included and there is no dependency.
!!! warning "Define consistently"
The macro selects between two definitions of the same inline validation function. It must therefore be defined
identically for **every** translation unit that includes the library; mixing translation units that define it with
ones that do not is an ODR violation. Prefer setting it as a compile definition on the target rather than with
`#!cpp #define` in individual source files.
## Default definition
By default, `#!cpp JSON_USE_SIMDUTF` is not defined and the portable C++11 scalar validator is used.
```cpp
#undef JSON_USE_SIMDUTF
```
## Examples
??? example
The code below enables the simdutf backend for UTF-8 validation.
```cpp
#define JSON_USE_SIMDUTF 1
#include <simdutf.h>
#include <nlohmann/json.hpp>
...
```
The project must also link against simdutf, e.g. with CMake:
```cmake
target_compile_definitions(your_target PRIVATE JSON_USE_SIMDUTF)
target_link_libraries(your_target PRIVATE simdutf::simdutf)
```
## Version history
- Added in version 3.12.1.
@@ -160,14 +160,11 @@ The library maps CBOR types to JSON value types as follows:
The mapping is **incomplete** in the sense that not all CBOR types can be converted to a JSON value. The following CBOR types are not supported and will yield parse errors:
- date/time (0xC0..0xC1)
- bignum (0xC2..0xC3)
- decimal fraction (0xC4)
- bigfloat (0xC5)
- expected conversions (0xD5..0xD7)
- simple values (0xE0..0xF3, 0xF8)
- undefined (0xF7)
Tagged items (0xC0..0xDB) are not interpreted either; see the note on tagged items below.
!!! warning "Negative integer overflow"
CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the
@@ -181,7 +178,7 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Tagged items"
Tagged items will throw a parse error by default. They can be ignored by passing `cbor_tag_handler_t::ignore` to function `from_cbor`. They can be stored by passing `cbor_tag_handler_t::store` to function `from_cbor`.
Tagged items (0xC0..0xDB) will throw a parse error by default. They can be ignored by passing `cbor_tag_handler_t::ignore` to function `from_cbor`, in which case the tag is skipped and the enclosed data item is parsed on its own. They can be stored by passing `cbor_tag_handler_t::store` to function `from_cbor`. Note that no tag is ever interpreted: for instance, a text string tagged with tag 0 (date/time) stays a string.
??? example
+7 -8
View File
@@ -91,6 +91,13 @@ security reasons (e.g., Intel Software Guard Extensions (SGX)).
See [full documentation of `JSON_NO_IO`](../api/macros/json_no_io.md).
## `JSON_NO_THREAD_LOCAL`
When defined, the library does not use `#!cpp thread_local` storage. Copying a value then always avoids the call stack
rather than descending into a bounded number of levels first, which is slower but yields the same values.
See [full documentation of `JSON_NO_THREAD_LOCAL`](../api/macros/json_no_thread_local.md).
## `JSON_SKIP_LIBRARY_VERSION_CHECK`
When defined, the library will not create a compiler warning when a different version of the library was already
@@ -137,14 +144,6 @@ behavior is deprecated and switched off (`0`) by default.
See [full documentation of `JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`](../api/macros/json_use_legacy_discarded_value_comparison.md).
## `JSON_USE_SIMDUTF`
When defined, UTF-8 validation of JSON strings read from contiguous byte input is delegated to the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. This is an opt-in
external dependency and is not defined by default.
See [full documentation of `JSON_USE_SIMDUTF`](../api/macros/json_use_simdutf.md).
## `NLOHMANN_DEFINE_TYPE_*(...)`, `NLOHMANN_DEFINE_DERIVED_TYPE_*(...)`
The library defines 12 macros to simplify the serialization/deserialization of types. See the page on
+1 -1
View File
@@ -291,12 +291,12 @@ nav:
- 'JSON_HAS_THREE_WAY_COMPARISON': api/macros/json_has_three_way_comparison.md
- 'JSON_NOEXCEPTION': api/macros/json_noexception.md
- 'JSON_NO_IO': api/macros/json_no_io.md
- 'JSON_NO_THREAD_LOCAL': api/macros/json_no_thread_local.md
- 'JSON_SKIP_LIBRARY_VERSION_CHECK': api/macros/json_skip_library_version_check.md
- 'JSON_SKIP_UNSUPPORTED_COMPILER_CHECK': api/macros/json_skip_unsupported_compiler_check.md
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
- 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md
- 'JSON_USE_SIMDUTF': api/macros/json_use_simdutf.md
- 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md
- 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md
- 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md
+1 -1
View File
@@ -1,4 +1,4 @@
wheel==0.47.0
wheel==0.48.0
mkdocs==1.6.1 # documentation framework
mkdocs-git-revision-date-localized-plugin==1.5.3 # plugin "git-revision-date-localized"
@@ -773,7 +773,13 @@ class binary_reader
case 0xBF: // map (indefinite length)
return get_cbor_object(detail::unknown_size(), tag_handler);
case 0xC6: // tagged item
case 0xC0: // tagged item
case 0xC1:
case 0xC2:
case 0xC3:
case 0xC4:
case 0xC5:
case 0xC6:
case 0xC7:
case 0xC8:
case 0xC9:
@@ -788,6 +794,9 @@ class binary_reader
case 0xD2:
case 0xD3:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8: // tagged item (1 byte follows)
case 0xD9: // tagged item (2 bytes follow)
case 0xDA: // tagged item (4 bytes follow)
+21 -109
View File
@@ -155,31 +155,11 @@ class input_stream_adapter
// General-purpose iterator-based adapter. It might not be as fast as
// theoretically possible for some containers, but it is extremely versatile.
// SentinelType defaults to IteratorType for backward compatibility, but may be
// a different type, e.g. a C++20 sentinel such as std::default_sentinel_t when
// IteratorType is a std::counted_iterator.
// SentinelType defaults to IteratorType for backward compatibility, but may
// be a different type (e.g., a C++20 sentinel or counted_iterator).
template<typename IteratorType, typename SentinelType = IteratorType>
class iterator_input_adapter
{
// Whether the number of elements between two positions can be computed in
// O(1): either the iterator and the sentinel have the same type (plain
// std::distance) or, in C++20, the sentinel is a sized sentinel for the
// iterator (std::ranges::distance), e.g. std::default_sentinel_t paired
// with std::counted_iterator.
//
// JSON_HAS_RANGES gates the C++20 branch: on standard libraries with an
// incomplete <ranges> (libstdc++ < 11, see #4440) evaluating
// std::contiguous_iterator on a std::counted_iterator is a hard error
// instead of yielding false, and these traits are instantiated for every
// adapter. Such toolchains fall back to the pointer-only test and simply
// use the byte-at-a-time scanner.
static constexpr bool sentinel_is_sized =
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>;
#else
std::is_same<IteratorType, SentinelType>::value;
#endif
public:
using char_type = typename std::iterator_traits<IteratorType>::value_type;
@@ -191,7 +171,7 @@ class iterator_input_adapter
// in wide_string_input_adapter, which does not expose this).
static constexpr bool supports_seek =
std::is_same<typename std::iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value
&& sentinel_is_sized
&& std::is_same<IteratorType, SentinelType>::value
&& sizeof(char_type) == 1;
iterator_input_adapter(IteratorType first, SentinelType last)
@@ -239,60 +219,30 @@ class iterator_input_adapter
private:
// whether IteratorType refers to a contiguous range and therefore supports
// a std::memcpy fast path (pointers always do; in C++20 we can also detect
// library iterators such as those of std::vector and std::string). The
// available element count must also be computable in O(1), hence
// sentinel_is_sized.
static constexpr bool iterator_is_contiguous = sentinel_is_sized &&
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
// library iterators such as those of std::vector and std::string).
// Computing the available element count needs either same-type iterators
// (plain std::distance) or, in C++20, a sized sentinel (std::ranges::distance),
// e.g. std::counted_iterator paired with std::default_sentinel_t.
static constexpr bool iterator_is_contiguous =
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>)
&& (std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
#else
std::is_pointer<IteratorType>::value;
std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
#endif
// number of unread elements in [current, end)
std::size_t remaining_count() const
{
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
return static_cast<std::size_t>(std::ranges::distance(current, end));
#else
return static_cast<std::size_t>(std::distance(current, end));
#endif
}
public:
// Whether the remaining input is a single contiguous block of 1-byte
// elements that the lexer can inspect directly (used for the SWAR string
// fast path).
static constexpr bool supports_bulk_scan =
iterator_is_contiguous && sizeof(char_type) == 1;
// Pointer to the next unread element; only valid when bulk_remaining() > 0.
const char_type* bulk_data() const
{
return &*current;
}
// Number of unread elements available as one contiguous block.
std::size_t bulk_remaining() const
{
return remaining_count();
}
// Consume @a n elements previously inspected via bulk_data().
void bulk_skip(std::size_t n)
{
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(n));
}
private:
// contiguous fast path: bulk copy the remaining range with std::memcpy
template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
{
const std::size_t wanted = count * sizeof(T);
const std::size_t available = remaining_count() * sizeof(char_type);
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
const std::size_t available = static_cast<std::size_t>(std::ranges::distance(current, end)) * sizeof(char_type);
#else
const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type);
#endif
const std::size_t copied = (std::min)(wanted, available);
if (JSON_HEDLEY_LIKELY(copied != 0))
{
@@ -620,24 +570,6 @@ typename iterator_input_adapter_factory<IteratorType, SentinelType>::adapter_typ
return factory_type::create(first, last);
}
// Detect a container that stores its elements contiguously as single bytes
// (std::string, std::vector<char/unsigned char>, std::array<char, N>,
// std::string_view, ...). Such inputs are wrapped in a pointer-based adapter so
// they benefit from the contiguous fast paths (bulk string scanning, memcpy for
// binary formats) in every C++ standard - not only in C++20, where the standard
// library iterators model std::contiguous_iterator and are detected directly.
template<typename ContainerType, typename = void>
struct is_contiguous_byte_container : std::false_type {};
template<typename ContainerType>
struct is_contiguous_byte_container < ContainerType, void_t <
decltype(std::declval<const ContainerType&>().data()),
decltype(std::declval<const ContainerType&>().size()) >>
: std::integral_constant < bool,
std::is_pointer<decltype(std::declval<const ContainerType&>().data())>::value&&
std::is_integral<typename std::remove_pointer<decltype(std::declval<const ContainerType&>().data())>::type>::value&&
sizeof(typename std::remove_pointer<decltype(std::declval<const ContainerType&>().data())>::type) == 1 > {};
// Convenience shorthand from container to iterator
// Enables ADL on begin(container) and end(container)
// Encloses the using declarations in namespace for not to leak them to outside scope
@@ -665,32 +597,12 @@ struct container_input_adapter_factory< ContainerType,
} // namespace container_input_adapter_factory_impl
// General container path (iterator-based). Contiguous single-byte containers
// are excluded here and routed through the pointer-based overload below.
template < typename ContainerType,
enable_if_t < !is_contiguous_byte_container<ContainerType>::value, int > = 0 >
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType && container)
template<typename ContainerType>
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType&& container)
{
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
}
// Contiguous single-byte containers (std::string, std::vector<char>, ...) are
// wrapped in a pointer-based adapter so the contiguous fast paths apply in every
// standard. The pointer keeps the container's own element type (const char* for
// std::string, const std::uint8_t* for std::vector<std::uint8_t>, ...), so the
// resulting char_type - and therefore the parsing behavior - is byte-for-byte
// identical to the iterator-based path; only the raw pointer additionally
// enables the bulk fast paths. The container outlives the adapter for the whole
// parse (temporaries live until the end of the full expression), exactly as the
// iterators it replaces did.
template < typename ContainerType,
enable_if_t < is_contiguous_byte_container<ContainerType>::value, int > = 0 >
auto input_adapter(const ContainerType& container)
-> decltype(input_adapter(container.data(), container.data() + container.size()))
{
return input_adapter(container.data(), container.data() + container.size());
}
// specialization for std::string
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
+27 -289
View File
@@ -19,9 +19,7 @@
#include <vector> // vector
#include <nlohmann/detail/input/input_adapters.hpp>
#include <nlohmann/detail/input/number_parse.hpp>
#include <nlohmann/detail/input/position_t.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
@@ -127,25 +125,6 @@ constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
return false;
}
// Detect whether an input adapter exposes a contiguous byte block that the
// lexer can scan directly (see iterator_input_adapter::supports_bulk_scan).
// Adapters without the flag - file, stream, wide-string, user-defined - fall
// back to the character-at-a-time string scanner.
template<typename InputAdapterType>
using detect_supports_bulk_scan = decltype(InputAdapterType::supports_bulk_scan);
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::true_type /*detected*/)
{
return InputAdapterType::supports_bulk_scan;
}
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::false_type /*detected*/)
{
return false;
}
/*!
@brief lexical analysis
@@ -167,14 +146,6 @@ class lexer : public lexer_base<BasicJsonType>
static constexpr bool lazy_token_string =
input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {});
/// whether string scanning may bulk-consume runs of ordinary characters
/// directly from a contiguous input buffer (SWAR fast path). This requires
/// the token to be reconstructible lazily (lazy_token_string), so bypassing
/// the per-character capture in get() cannot lose error diagnostics.
static constexpr bool bulk_scan =
lazy_token_string
&& input_adapter_supports_bulk_scan<InputAdapterType>(is_detected<detect_supports_bulk_scan, InputAdapterType> {});
public:
using token_type = typename lexer_base<BasicJsonType>::token_type;
@@ -294,40 +265,6 @@ class lexer : public lexer_base<BasicJsonType>
return true;
}
/// contiguous input: bulk-append the run of ordinary characters and complete
/// well-formed UTF-8 sequences starting at the current read position, leaving
/// the first byte that needs individual handling (the closing quote, an
/// escape, a control character, or an ill-formed UTF-8 byte) for get()
void scan_string_bulk(std::true_type /*bulk*/)
{
// a pending unget must be consumed through the normal path first
if (next_unget)
{
return;
}
const std::size_t remaining = ia.bulk_remaining();
if (remaining == 0)
{
return;
}
const auto* const data = reinterpret_cast<const unsigned char*>(ia.bulk_data());
const std::size_t pos = string_bulk_run(data, remaining);
if (pos == 0)
{
return;
}
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), pos);
ia.bulk_skip(pos);
// the run contains no newline (all bytes < 0x20 are treated as special),
// so only the flat character counters advance
position.chars_read_total += pos;
position.chars_read_current_line += pos;
}
/// streaming input: no bulk fast path
void scan_string_bulk(std::false_type /*bulk*/) const noexcept {}
/*!
@brief scan a string literal
@@ -353,10 +290,6 @@ class lexer : public lexer_base<BasicJsonType>
while (true)
{
// bulk-consume ordinary characters from contiguous input, then
// handle the next special byte through the switch below
scan_string_bulk(std::integral_constant<bool, bulk_scan> {});
// get the next character
switch (get())
{
@@ -1346,78 +1279,45 @@ scan_number_done:
// we are done scanning a number)
unget();
return convert_number(number_type);
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
errno = 0;
/*!
@brief convert an already-validated integer token to its value
The digit sequence in [first, last) has been validated by the caller, so a
dedicated parser can avoid the locale/errno overhead of std::strtoull.
@return the token type on success; token_type::uninitialized if @a
number_type is not an integer type or the value does not fit, in
which case the caller falls back to the floating-point conversion
(matching the previous std::strtoull/std::strtoll behavior)
*/
token_type convert_integer(token_type number_type, const char* first, const char* last)
{
// try to parse integers first and fall back to floats
if (number_type == token_type::value_unsigned)
{
if (parse_integer_unsigned(first, last, value_unsigned))
const auto x = std::strtoull(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
{
return token_type::value_unsigned;
value_unsigned = static_cast<number_unsigned_t>(x);
if (value_unsigned == x)
{
return token_type::value_unsigned;
}
}
}
else if (number_type == token_type::value_integer)
{
if (parse_integer_signed(first, last, value_integer))
const auto x = std::strtoll(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
{
return token_type::value_integer;
}
}
return token_type::uninitialized;
}
/*!
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
*/
token_type convert_number(token_type number_type)
{
const char* const num_begin = token_buffer.data();
const char* const num_end = num_begin + token_buffer.size();
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, num_begin, num_end);
if (integer_result != token_type::uninitialized)
{
return integer_result;
value_integer = static_cast<number_integer_t>(x);
if (value_integer == x)
{
return token_type::value_integer;
}
}
}
// this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
}
if (parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
{
return token_type::value_float;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
// integer conversion above failed
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
@@ -1426,153 +1326,6 @@ scan_number_done:
return token_type::value_float;
}
/*!
@brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
then produces the exact diagnostic. @a current is the first digit or the
leading minus (already read); the remaining bytes are taken from the adapter.
*/
token_type scan_number_bulk_contiguous()
{
// a pending unget offsets the buffer position from current; fall back
if (next_unget)
{
return token_type::uninitialized;
}
const std::size_t rem = ia.bulk_remaining();
if (rem == 0)
{
// the first digit is the last input byte; let scan_number() finish
return token_type::uninitialized;
}
// the byte before the next unread one is current (contiguous input)
const char* const data = reinterpret_cast<const char*>(ia.bulk_data()) - 1;
const std::size_t avail = rem + 1;
// validate + classify the number extent (mirrors scan_number()'s grammar)
std::size_t i = 0;
std::size_t dot_index = std::string::npos;
token_type number_type = token_type::value_unsigned;
if (data[0] == '-')
{
number_type = token_type::value_integer;
i = 1;
if (i >= avail)
{
return token_type::uninitialized;
}
}
if (data[i] == '0')
{
++i;
}
else if (data[i] >= '1' && data[i] <= '9')
{
++i;
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
else
{
return token_type::uninitialized;
}
if (i < avail && data[i] == '.')
{
number_type = token_type::value_float;
dot_index = i;
++i;
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
if (i < avail && (data[i] == 'e' || data[i] == 'E'))
{
number_type = token_type::value_float;
++i;
if (i < avail && (data[i] == '+' || data[i] == '-'))
{
++i;
}
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
const std::size_t len = i;
// reset() records where this token starts (for diagnostics), so it has
// to run before the input position advances below
reset();
// An integer token needs no token_buffer: the SAX callbacks for
// number_integer/number_unsigned take only the value, and the overflow
// diagnostic rebuilds the text from the input. Convert straight from the
// input buffer and leave token_buffer empty. (JSON_DIAGNOSTIC_POSITIONS
// derives a number's start position from get_string().size(), so there
// the token still has to be materialized.)
#if !JSON_DIAGNOSTIC_POSITIONS
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, data, data + len);
if (JSON_HEDLEY_LIKELY(integer_result != token_type::uninitialized))
{
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return integer_result;
}
// the value overflowed: fall through and let the float tail handle it
}
#endif
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return convert_number(number_type);
}
/// contiguous input: try the number fast path, else the byte-path scanner
token_type scan_number_dispatch(std::true_type /*bulk*/)
{
const token_type t = scan_number_bulk_contiguous();
return (t != token_type::uninitialized) ? t : scan_number();
}
/// streaming input: always use the byte-path scanner
token_type scan_number_dispatch(std::false_type /*bulk*/)
{
return scan_number();
}
/*!
@param[in] literal_text the literal text to expect
@param[in] length the length of the passed literal text
@@ -1660,9 +1413,6 @@ scan_number_done:
if (current == '\n')
{
++position.lines_read;
// remember the column the newline was read at: chars_read_current_line
// is about to be cleared, and a matching unget() cannot reconstruct it
chars_read_before_newline = position.chars_read_current_line;
position.chars_read_current_line = 0;
}
@@ -1696,20 +1446,12 @@ scan_number_done:
--position.chars_read_total;
// in case we "unget" a newline, we have to also decrement the lines_read
// and restore the column that get() cleared when it saw the newline;
// chars_read_current_line == 0 can only mean the last get() read one
if (position.chars_read_current_line == 0)
{
if (position.lines_read > 0)
{
--position.lines_read;
}
// chars_read_before_newline counts the newline itself, which is the
// character being ungotten, hence the -1
position.chars_read_current_line = (chars_read_before_newline > 0)
? chars_read_before_newline - 1
: 0;
}
else
{
@@ -1952,7 +1694,7 @@ scan_number_done:
case '7':
case '8':
case '9':
return scan_number_dispatch(std::integral_constant<bool, bulk_scan> {});
return scan_number();
// end of input (the null byte is needed when parsing from
// string literals)
@@ -1983,10 +1725,6 @@ scan_number_done:
/// the start position of the current token
position_t position {};
/// the value chars_read_current_line had when the last newline was read, so
/// that unget() can restore the column instead of leaving it at 0
std::size_t chars_read_before_newline = 0;
/// raw input token string for error messages; only populated for streaming
/// adapters (seekable adapters reconstruct it lazily via token_string_start)
std::vector<char_type> token_string {};
@@ -1,302 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint64_t
#include <limits> // numeric_limits
#include <nlohmann/detail/macro_scope.hpp>
// std::from_chars lives in <charconv>, but being in C++17 mode does not
// guarantee the header exists: GCC 7 sets __cplusplus to C++17 yet ships no
// <charconv> (added in GCC 8; floating-point support in GCC 11). Guard the
// include with __has_include so such toolchains fall back to the scalar path.
#if defined(JSON_HAS_CPP_17) && defined(__has_include)
#if __has_include(<charconv>)
#include <charconv> // from_chars (only used when __cpp_lib_to_chars is defined)
#include <system_error> // errc
#endif
#endif
// This file contains the value-conversion helpers used by the lexer to turn an
// already-validated number token into a value, without the locale/errno
// overhead of std::strtoull/std::strtod. They are free functions so the lexer
// stays focused on scanning; see lexer::convert_number().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief fast integer parser for an already-validated unsigned integer
The number scanner has already checked that [first, last) is a valid JSON
integer, so this only needs to accumulate the digits and detect overflow. This
avoids the locale/errno machinery of std::strtoull, which dominates
integer-heavy inputs.
@param[in] first pointer to the first character (a digit)
@param[in] last pointer past the last character
@param[out] value the parsed value on success
@return true if the value fit into @a NumberUnsignedType; false on overflow, in
which case the caller falls back to floating-point parsing (matching the
previous std::strtoull behavior)
*/
template<typename NumberUnsignedType>
bool parse_integer_unsigned(const char* first, const char* last, NumberUnsignedType& value) noexcept
{
// accumulate in the widest unsigned type used by the previous strtoull
// path so the overflow behavior is unchanged for custom number types
std::uint64_t x = 0;
constexpr std::uint64_t cutoff = (std::numeric_limits<std::uint64_t>::max)() / 10u;
constexpr std::uint64_t cutlim = (std::numeric_limits<std::uint64_t>::max)() % 10u;
for (const char* p = first; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(x > cutoff || (x == cutoff && digit > cutlim)))
{
return false;
}
x = (x * 10u) + digit;
}
value = static_cast<NumberUnsignedType>(x);
// reject values that do not round-trip into a narrower NumberUnsignedType
return static_cast<std::uint64_t>(value) == x;
}
/*!
@brief fast integer parser for an already-validated negative integer
@param[in] first pointer to the leading '-'
@param[in] last pointer past the last character
@param[out] value the parsed (negative) value on success
@return true on success; false on overflow (caller falls back to float)
*/
template<typename NumberIntegerType>
bool parse_integer_signed(const char* first, const char* last, NumberIntegerType& value) noexcept
{
// the state machine only reaches the signed path via a leading '-'
JSON_ASSERT(first != last && *first == '-');
std::uint64_t magnitude = 0;
// |INT64_MIN| == INT64_MAX + 1; this is the largest admissible magnitude
constexpr std::uint64_t limit = static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()) + 1u;
for (const char* p = first + 1; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(magnitude > (limit - digit) / 10u))
{
return false;
}
magnitude = (magnitude * 10u) + digit;
}
const std::int64_t x = (magnitude == limit)
? (std::numeric_limits<std::int64_t>::min)()
: -static_cast<std::int64_t>(magnitude);
value = static_cast<NumberIntegerType>(x);
// reject values that do not round-trip into a narrower NumberIntegerType
return static_cast<std::int64_t>(value) == x;
}
/*!
@brief exact fast path for parsing a `double` (Clinger's algorithm)
For the common case - at most 19 significant digits, a decimal exponent in
[-22, 22], and a significand below 2^53 - the value equals significand *
10^exp computed in IEEE-754 double arithmetic, which is exact under
round-to-nearest because both operands are exactly representable. This is the
same fast path used by fast_float/simdjson; the general cases are left to
std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
// true double precision. On platforms that keep intermediates in extended
// precision (e.g. the x87 FPU on 32-bit x86, where FLT_EVAL_METHOD == 2) the
// single significand * 10^scale step is double-rounded and can be 1 ULP off,
// so decline and let the caller fall back to the correctly-rounded
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
static const std::array<double, 23> powers_of_ten =
{
{
1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11,
1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22
}
};
const char* p = first;
bool negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
negative = (*p == '-');
++p;
}
std::uint64_t significand = 0;
int num_digits = 0;
int fractional_digits = 0;
bool seen_dot = false;
bool any_digit = false;
for (; p != last; ++p)
{
const char c = *p;
if (c >= '0' && c <= '9')
{
any_digit = true;
if (JSON_HEDLEY_UNLIKELY(num_digits >= 19))
{
return false; // significand may not fit into uint64_t
}
significand = (significand * 10u) + static_cast<std::uint64_t>(c - '0');
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (static_cast<DecimalPointType>(c) == decimal_point)
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
return false;
}
seen_dot = true;
}
else if (c == 'e' || c == 'E')
{
++p;
break;
}
else
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_digit))
{
return false;
}
int exponent = 0;
if (p != last) // an exponent part remains
{
bool exp_negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
exp_negative = (*p == '-');
++p;
}
bool any_exp_digit = false;
for (; p != last; ++p)
{
if (JSON_HEDLEY_UNLIKELY(*p < '0' || *p > '9'))
{
return false;
}
exponent = (exponent * 10) + (*p - '0');
any_exp_digit = true;
if (JSON_HEDLEY_UNLIKELY(exponent > 9999))
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_exp_digit))
{
return false;
}
if (exp_negative)
{
exponent = -exponent;
}
}
const int scale = exponent - fractional_digits;
if (JSON_HEDLEY_UNLIKELY(significand >= (static_cast<std::uint64_t>(1) << 53)))
{
return false; // significand not exactly representable as double
}
auto result = static_cast<double>(significand);
if (scale >= 0)
{
if (JSON_HEDLEY_UNLIKELY(scale > 22))
{
return false;
}
result *= powers_of_ten[static_cast<std::size_t>(scale)];
}
else
{
if (JSON_HEDLEY_UNLIKELY(-scale > 22))
{
return false;
}
result /= powers_of_ten[static_cast<std::size_t>(-scale)];
}
out = negative ? -result : result;
return true;
#endif
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
{
return false;
}
/*!
@brief parse a float with std::from_chars (Eisel-Lemire) when available
std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
@return true if the value was parsed exactly and fully; false to fall back
*/
template<typename FloatType>
bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept
{
// JSON_HAS_CPP_17 must gate the use as well as the <charconv> include above:
// some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even
// in C++14 mode, where <charconv> is not included.
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
const auto result = std::from_chars(first, last, out);
return result.ec == std::errc() && result.ptr == last;
#else
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(out);
return false;
#endif
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -1,237 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef> // size_t
#include <cstdint> // uint64_t
#include <cstring> // memcpy
#if defined(JSON_USE_SIMDUTF)
// Optional SIMD backend for bulk UTF-8 validation. This is an opt-in
// external dependency: nlohmann/json itself stays header-only and the C++11
// scalar validator below is always available; defining JSON_USE_SIMDUTF
// additionally requires the simdutf headers on the include path and linking
// the simdutf library. See string_bulk_run().
#include <simdutf.h>
#endif
#include <nlohmann/detail/macro_scope.hpp>
// This file contains the byte-level string-scanning helpers used by the lexer's
// contiguous fast path. They operate purely on raw bytes (no dependency on the
// lexer's template parameters) so they are free functions, keeping the lexer
// itself focused on the state machine; see lexer::scan_string_bulk().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
// classify a single byte as needing individual string handling: the closing
// quote, an escape, a control character, or a non-ASCII (UTF-8)
// lead/continuation byte. Ordinary bytes (0x20..0x7F except '"' and '\\') are
// copied verbatim, which the bulk scanner does 8 bytes at a time.
inline bool is_string_special(unsigned char c) noexcept
{
return c == '\"' || c == '\\' || c < 0x20u || c >= 0x80u;
}
// SWAR helper: return a word whose high bit is set in every byte of @a v that
// is_string_special(); zero if the 8 bytes are all ordinary.
inline std::uint64_t swar_string_special(std::uint64_t v) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
const std::uint64_t has_quote = (q - ones) & ~q & high;
const std::uint64_t has_backslash = (b - ones) & ~b & high;
const std::uint64_t has_control = (v - 0x2020202020202020ull) & ~v & high; // < 0x20
const std::uint64_t has_non_ascii = v & high; // >= 0x80
return has_quote | has_backslash | has_control | has_non_ascii;
}
// return the index of the first is_string_special() byte in [data, data+n), or
// n if every byte is ordinary; scans 8 bytes at a time
inline std::size_t find_string_special(const unsigned char* data, std::size_t n) noexcept
{
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t word = 0;
std::memcpy(&word, data + i, sizeof(word));
if (swar_string_special(word) != 0)
{
// a special byte is in this word; locate it (endian-agnostic)
for (std::size_t j = 0; j < 8; ++j)
{
if (is_string_special(data[i + j]))
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
if (is_string_special(data[i]))
{
return i;
}
}
return n;
}
// Validate one UTF-8 sequence at the front of [data, data+avail). Returns its
// length (2..4) only when the bytes form a *well-formed* sequence using exactly
// the same ranges as scan_string()'s per-byte switch, so the bulk path accepts
// precisely what the byte path accepts. Returns 0 for anything that is invalid,
// incomplete, or that the byte path must diagnose (the caller then defers to
// that path, keeping error messages unchanged). Lead bytes < 0x80 are handled
// by the caller and never passed here.
inline std::size_t validate_one_utf8(const unsigned char* data, std::size_t avail) noexcept
{
const unsigned char c0 = data[0];
if (c0 >= 0xC2 && c0 <= 0xDF) // U+0080..U+07FF
{
if (avail >= 2 && data[1] >= 0x80 && data[1] <= 0xBF)
{
return 2;
}
}
else if (c0 == 0xE0) // U+0800..U+0FFF
{
if (avail >= 3 && data[1] >= 0xA0 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if ((c0 >= 0xE1 && c0 <= 0xEC) || c0 == 0xEE || c0 == 0xEF) // U+1000..U+CFFF, U+E000..U+FFFF
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xED) // U+D000..U+D7FF (excludes surrogates)
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0x9F && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xF0) // U+10000..U+3FFFF
{
if (avail >= 4 && data[1] >= 0x90 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 >= 0xF1 && c0 <= 0xF3) // U+40000..U+FFFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 == 0xF4) // U+100000..U+10FFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0x8F && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
return 0; // invalid, incomplete, or must be diagnosed by the byte path
}
// Scalar (C++11) computation of the bulk run length: the number of leading
// bytes in [data, data+n) that are ordinary ASCII or complete well-formed UTF-8
// sequences, stopping before the first byte that needs individual handling (the
// closing quote, an escape, a control character, or an ill-formed/truncated
// sequence). ASCII is skipped 8 bytes at a time.
inline std::size_t scalar_string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
std::size_t pos = 0;
while (pos < n)
{
pos += find_string_special(data + pos, n - pos);
if (pos >= n || data[pos] < 0x80u)
{
break; // end of buffer, or a quote/escape/control byte
}
const std::size_t seq = validate_one_utf8(data + pos, n - pos);
if (seq == 0)
{
break; // ill-formed or truncated: let the byte path diagnose it
}
pos += seq;
}
return pos;
}
#if defined(JSON_USE_SIMDUTF)
// Index of the first quote/escape/control byte in [data, data+n) (non-ASCII
// bytes are *not* stops here - the whole run is handed to simdutf), or n.
inline std::size_t find_string_delimiter(const unsigned char* data, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t v = 0;
std::memcpy(&v, data + i, sizeof(v));
const std::uint64_t q = v ^ 0x2222222222222222ull;
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull;
const std::uint64_t hit = ((q - ones) & ~q & high)
| ((b - ones) & ~b & high)
| ((v - 0x2020202020202020ull) & ~v & high);
if (hit != 0)
{
for (std::size_t j = 0; j < 8; ++j)
{
const unsigned char c = data[i + j];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
const unsigned char c = data[i];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i;
}
}
return n;
}
#endif
// Backend-dispatched bulk run length. With JSON_USE_SIMDUTF the run up to the
// next delimiter is validated in one shot by simdutf; on the rare failure the
// scalar helper recomputes the exact valid prefix so the byte path still
// produces the precise diagnostic. Without it, the pure scalar path is used.
inline std::size_t string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
#if defined(JSON_USE_SIMDUTF)
const std::size_t run = find_string_delimiter(data, n);
if (run != 0 && simdutf::validate_utf8(reinterpret_cast<const char*>(data), run))
{
return run;
}
return scalar_string_bulk_run(data, n);
#else
return scalar_string_bulk_run(data, n);
#endif
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+9
View File
@@ -186,6 +186,15 @@
#define JSON_NO_UNIQUE_ADDRESS
#endif
// Clang targeting MinGW does not survive the thread_local storage the copy
// constructor uses to bound its descent: every test that copies a value
// segfaults with clang 11.0.1 and clang 18.1.8, while the same tests pass with
// GCC targeting MinGW and with every other toolchain the library is tested on.
// Copying works the same way without the counter, only more slowly.
#if !defined(JSON_NO_THREAD_LOCAL) && defined(__clang__) && defined(__MINGW32__)
#define JSON_NO_THREAD_LOCAL 1
#endif
// disable documentation warnings on clang
#if defined(__clang__)
#pragma clang diagnostic push
+383 -55
View File
@@ -28,14 +28,14 @@
#pragma GCC diagnostic ignored "-Wignored-attributes"
#endif
#include <algorithm> // all_of, find, for_each
#include <algorithm> // all_of, find, for_each, none_of
#include <cstddef> // nullptr_t, ptrdiff_t, size_t
#include <functional> // hash, less
#include <initializer_list> // initializer_list
#ifndef JSON_NO_IO
#include <iosfwd> // istream, ostream
#endif // JSON_NO_IO
#include <iterator> // random_access_iterator_tag
#include <iterator> // make_move_iterator, random_access_iterator_tag
#include <memory> // unique_ptr
#include <string> // string, stoi, to_string
#include <utility> // declval, forward, move, pair, swap
@@ -821,6 +821,379 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
return j;
}
/// the number of levels an operation descends into before it finishes the
/// value below it without the call stack
static constexpr std::size_t nesting_depth_limit()
{
return 128;
}
#ifndef JSON_NO_THREAD_LOCAL
/*!
@brief how many levels the operation going on in this thread has descended into
Copying a value and comparing two values share this count. The library never
nests one inside the other - copying a value does not compare one, and
comparing two values does not copy them - and where user code nests them
anyway, sharing the count only ends a descent sooner than it had to, which
costs a little speed and is never wrong.
A byte is enough: the count never exceeds the limit by more than the single
level that notices the limit has been reached.
*/
static std::size_t& nesting_depth() noexcept
{
static thread_local std::size_t depth = 0; // NOLINT(misc-use-internal-linkage)
return depth;
}
#endif
/*!
@brief whether a descent must stop here and finish without the call stack
@a may_descend says whether the operation descends at all; it is a constant
at every call site, and is passed rather than tested by the caller so that
the test does not become a constant condition there, which MSVC reports as
C4127.
*/
static bool nesting_depth_exhausted(bool may_descend = true) noexcept
{
#ifdef JSON_NO_THREAD_LOCAL
// without a count of its own per thread, a descent cannot be bounded
// without racing another one, so none is made
static_cast<void>(may_descend);
return true;
#else
return !may_descend || nesting_depth() >= nesting_depth_limit();
#endif
}
/*!
@brief counts one level of a bounded descent for as long as it runs
The constructor taking the count is for callers that have looked it up
already to test it: reaching thread-local storage is not free, and the path
that is taken almost every time should reach it once rather than twice. The
other is for callers that cannot look it up - the comparison operators are
written as a macro, and a macro cannot use the preprocessor.
*/
class nesting_depth_guard
{
public:
explicit nesting_depth_guard(std::size_t& depth) noexcept
: m_depth(&depth)
{
++*m_depth;
}
nesting_depth_guard() noexcept
: m_depth(countable())
{
if (m_depth != nullptr)
{
++*m_depth;
}
}
~nesting_depth_guard()
{
if (m_depth != nullptr)
{
--*m_depth;
}
}
nesting_depth_guard(const nesting_depth_guard&) = delete;
nesting_depth_guard& operator=(const nesting_depth_guard&) = delete;
nesting_depth_guard(nesting_depth_guard&&) = delete;
nesting_depth_guard& operator=(nesting_depth_guard&&) = delete;
private:
/// @brief the count to keep, or nullptr where there is none to keep
static std::size_t* countable() noexcept
{
#ifdef JSON_NO_THREAD_LOCAL
return nullptr;
#else
return &nesting_depth();
#endif
}
std::size_t* m_depth;
};
/// an entry of the iterative deep copy's worklist: a structured value and
/// the value that is to become its copy
using copy_worklist_t = std::vector<std::pair<const basic_json*, basic_json*>>;
/// scratch space to build the key skeleton of an object copy in one go
using copy_scratch_t = std::vector<std::pair<typename object_t::key_type, basic_json>>;
/// @brief copy everything of @a src into @a dst but its type and value
static void copy_metadata(const basic_json& src, basic_json& dst)
{
// a custom base class is only required to be copy-constructible and
// move-assignable, so the copy has to go through a temporary
static_cast<json_base_class_t&>(dst) = json_base_class_t(static_cast<const json_base_class_t&>(src));
#if JSON_DIAGNOSTIC_POSITIONS
dst.start_position = src.start_position;
dst.end_position = src.end_position;
#else
static_cast<void>(src);
static_cast<void>(dst);
#endif
}
/*!
@brief copy the value of @a src into @a dst, which must not be structured
Objects and arrays are left alone: creating those is the one thing the copy
constructor and @ref copy_shallow do differently from one another, and it is
the reason copying a value can descend at all.
*/
/// @note inlined on purpose: both callers have already told an object or an
/// array apart from the rest, and letting the compiler fold that test
/// into this switch is worth a few percent when copying a value made
/// mostly of numbers
JSON_HEDLEY_ALWAYS_INLINE
static void copy_leaf_value(const basic_json& src, basic_json& dst)
{
switch (src.m_data.m_type)
{
case value_t::string:
{
dst.m_data.m_value = *src.m_data.m_value.string;
break;
}
case value_t::binary:
{
dst.m_data.m_value = *src.m_data.m_value.binary;
break;
}
case value_t::boolean:
{
dst.m_data.m_value = src.m_data.m_value.boolean;
break;
}
case value_t::number_integer:
{
dst.m_data.m_value = src.m_data.m_value.number_integer;
break;
}
case value_t::number_unsigned:
{
dst.m_data.m_value = src.m_data.m_value.number_unsigned;
break;
}
case value_t::number_float:
{
dst.m_data.m_value = src.m_data.m_value.number_float;
break;
}
case value_t::object:
case value_t::array:
case value_t::null:
case value_t::discarded:
default:
break;
}
}
/*!
@brief copy everything of @a src into the null value @a dst but the children
Objects and arrays are not copied here; they are appended to @a worklist to
be created later by @ref copy_iteratively. Until that happens, @a dst remains
a null value, so that a partially built copy can be destroyed at any point
without ever violating the class invariants.
*/
static void copy_shallow(const basic_json& src, basic_json& dst, copy_worklist_t& worklist)
{
copy_metadata(src, dst);
if (src.m_data.m_type == value_t::object || src.m_data.m_type == value_t::array)
{
// defer: dst stays a null value until its container exists
worklist.emplace_back(&src, &dst);
return;
}
copy_leaf_value(src, dst);
// only now that the value exists may the type be set: had the creation
// of the value thrown, dst would have been left as a valid null value
dst.m_data.m_type = src.m_data.m_type;
}
/// @brief create the copy of the array @a src in @a dst
/// @note structured elements are appended to @a worklist instead
static void copy_array_level(const basic_json& src, basic_json& dst, copy_worklist_t& worklist)
{
const array_t& src_array = *src.m_data.m_value.array;
// create all elements up front: growing the array afterwards could
// invalidate the pointers that are handed to the worklist
dst.m_data.m_value.array = create<array_t>(src_array.size(), basic_json());
auto dst_it = dst.m_data.m_value.array->begin();
for (auto src_it = src_array.cbegin(); src_it != src_array.cend(); ++src_it, ++dst_it)
{
copy_shallow(*src_it, *dst_it, worklist);
}
}
/// @brief create the copy of the object @a src in @a dst
/// @note structured values are appended to @a worklist instead
static void copy_object_level(const basic_json& src, basic_json& dst,
copy_worklist_t& worklist, copy_scratch_t& scratch)
{
const object_t& src_object = *src.m_data.m_value.object;
// build the complete key skeleton and hand it to the object's range
// constructor: adding the keys one by one would be quadratic for object
// types that are backed by a vector, such as nlohmann::ordered_map
scratch.clear();
scratch.reserve(src_object.size());
for (const auto& element : src_object)
{
scratch.emplace_back(element.first, basic_json());
}
dst.m_data.m_value.object = create<object_t>(std::make_move_iterator(scratch.begin()),
std::make_move_iterator(scratch.end()));
scratch.clear();
// pair every value of the copy with its counterpart in the original;
// both are enumerated in the same order for every object type with a
// deterministic order, so the lookup is only needed for exotic ones
auto src_it = src_object.cbegin();
for (auto& element : *dst.m_data.m_value.object)
{
if (JSON_HEDLEY_LIKELY(src_it != src_object.cend() && src_it->first == element.first))
{
copy_shallow(src_it->second, element.second, worklist);
++src_it;
}
else
{
const auto found = src_object.find(element.first);
JSON_ASSERT(found != src_object.cend());
copy_shallow(found->second, element.second, worklist);
}
}
}
/*!
@brief deep-copy the object or array @a src into this value without recursing
The values whose copy has not been created yet are kept on an explicit
worklist rather than on the call stack. This is only reached for values
nested deeper than @ref nesting_depth_limit levels, which is why it copies
every container by hand instead of letting the container do it: the fast
ways of doing so would descend into the elements and defeat the purpose.
*/
void copy_iteratively(const basic_json& src)
{
copy_worklist_t worklist;
copy_scratch_t scratch;
const basic_json* src_value = &src;
basic_json* dst_value = this;
for (;;)
{
if (src_value->m_data.m_type == value_t::array)
{
copy_array_level(*src_value, *dst_value, worklist);
}
else
{
copy_object_level(*src_value, *dst_value, worklist, scratch);
}
// the container is complete and will not be modified again
dst_value->set_parents();
if (worklist.empty())
{
break;
}
const auto& next = worklist.back();
src_value = next.first;
dst_value = next.second;
worklist.pop_back();
// the value stops being a null value exactly here
dst_value->m_data.m_type = src_value->m_data.m_type;
}
}
/*!
@brief copy one level of the object or array @a src into this value
The container copies its own elements, which is the fastest way to fill it.
Every element that is structured itself comes back to @ref copy_structured.
*/
void copy_level(const basic_json& src)
{
if (m_data.m_type == value_t::object)
{
m_data.m_value = *src.m_data.m_value.object;
}
else
{
m_data.m_value = *src.m_data.m_value.array;
}
set_parents();
}
/*!
@brief deep-copy the object or array @a src into this value
Copying a container copies its elements, so a value nested deeply enough
used to exhaust the call stack. The descent is bounded here: the first
@ref nesting_depth_limit levels are copied by the containers themselves, just
as they always were, and anything below that is copied without the call
stack by @ref copy_iteratively. Copying a value can therefore no longer
exhaust the stack, however deeply it is nested, just like destroying one
cannot since #1436.
Nothing has to be scanned or built by hand to reach that: a value that is
not nested deeper than the limit - all but a vanishing minority - is copied
exactly as it was before, and this whole detour costs it one counter.
@sa https://github.com/nlohmann/json/issues/5387
*/
void copy_structured(const basic_json& src)
{
#ifndef JSON_NO_THREAD_LOCAL
std::size_t& depth = nesting_depth();
if (JSON_HEDLEY_LIKELY(depth < nesting_depth_limit()))
{
const nesting_depth_guard guard(depth);
copy_level(src);
return;
}
#endif
// Finish this value without descending any further. It is completed
// before this returns, so a copy made by a custom base class - or by
// anything else that runs while a copy is going on - is unaffected by
// the copy it is nested in.
copy_iteratively(src);
}
public:
//////////////////////////
// JSON parser callback //
@@ -1200,60 +1573,15 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// check of passed value is valid
other.assert_invariant();
switch (m_data.m_type)
if (m_data.m_type == value_t::object || m_data.m_type == value_t::array)
{
case value_t::object:
{
m_data.m_value = *other.m_data.m_value.object;
break;
}
case value_t::array:
{
m_data.m_value = *other.m_data.m_value.array;
break;
}
case value_t::string:
{
m_data.m_value = *other.m_data.m_value.string;
break;
}
case value_t::boolean:
{
m_data.m_value = other.m_data.m_value.boolean;
break;
}
case value_t::number_integer:
{
m_data.m_value = other.m_data.m_value.number_integer;
break;
}
case value_t::number_unsigned:
{
m_data.m_value = other.m_data.m_value.number_unsigned;
break;
}
case value_t::number_float:
{
m_data.m_value = other.m_data.m_value.number_float;
break;
}
case value_t::binary:
{
m_data.m_value = *other.m_data.m_value.binary;
break;
}
case value_t::null:
case value_t::discarded:
default:
break;
// copying the container directly would call this constructor again
// for every element, once per nesting level
copy_structured(other);
}
else
{
copy_leaf_value(other, *this);
}
set_parents();
File diff suppressed because it is too large Load Diff
+51
View File
@@ -216,6 +216,57 @@ TEST_CASE("controlled bad_alloc")
CHECK_THROWS_AS(my_json(s), std::bad_alloc&);
next_construct_fails = false;
}
SECTION("basic_json(const basic_json&) of a deeply nested value (#5387)")
{
// Copying a value nested deeper than the descent bound builds the
// copy from the top down: every value whose own copy has not been
// made yet stays a null value until it is. Failing an allocation
// part-way through is what proves such a half-built copy can still
// be destroyed.
//
// Which path the failure lands in depends on the build: the first
// allocation of a copy belongs to the outermost level, so here it
// is the descending one. Built with JSON_NO_THREAD_LOCAL - as the
// ci_test_no_thread_local target builds the whole suite - no
// descent is made at all and the very same failure lands in the
// iterative path instead, part-way through its worklist.
const auto check_deep_copy = [](bool objects)
{
CAPTURE(objects);
next_construct_fails = false;
// deeper than the 128 levels the copy constructor descends into
const std::size_t depth = 300;
my_json j = 1;
for (std::size_t i = 0; i < depth; ++i)
{
if (objects)
{
my_json wrapper = my_json::object();
wrapper["a"] = std::move(j);
j = std::move(wrapper);
}
else
{
j = my_json::array({std::move(j)});
}
}
// NOLINTNEXTLINE(performance-unnecessary-copy-initialization): the copy is what is tested
CHECK_NOTHROW(my_json(j));
next_construct_fails = true;
// NOLINTNEXTLINE(performance-unnecessary-copy-initialization): the copy is what is tested
CHECK_THROWS_AS(my_json(j), std::bad_alloc&);
next_construct_fails = false;
};
check_deep_copy(false);
check_deep_copy(true);
}
}
}
+13 -2
View File
@@ -2565,11 +2565,16 @@ TEST_CASE("Tagged values")
const json j = "s";
auto v = json::to_cbor(j);
SECTION("0xC6..0xD4")
const json j_bin_payload = json::binary(std::vector<std::uint8_t> {0x01, 0x02, 0x03});
auto v_bin_payload = json::to_cbor(j_bin_payload);
SECTION("0xC0..0xD7")
{
for (const auto b : std::vector<std::uint8_t>
{
0xC6, 0xC7, 0xC8, 0xC9, 0xCA, 0xCB, 0xCC, 0xCD, 0xCE, 0xCF, 0xD0, 0xD1, 0xD2, 0xD3, 0xD4
0xC0, 0xC1, 0xC2, 0xC3, 0xC4, 0xC5,
0xC6, 0xC7, 0xC8, 0xC9, 0xCA, 0xCB, 0xCC, 0xCD, 0xCE, 0xCF, 0xD0, 0xD1, 0xD2, 0xD3, 0xD4,
0xD5, 0xD6, 0xD7
})
{
CAPTURE(b);
@@ -2589,6 +2594,12 @@ TEST_CASE("Tagged values")
auto j_tagged_stored = json::from_cbor(v_tagged, true, true, json::cbor_tag_handler_t::store);
CHECK(j_tagged_stored == j);
auto v_binary_tagged = v_bin_payload;
v_binary_tagged.insert(v_binary_tagged.begin(), b);
auto j_binary_tagged_stored = json::from_cbor(v_binary_tagged, true, true, json::cbor_tag_handler_t::store);
CHECK(j_binary_tagged_stored == j_bin_payload);
CHECK(!j_binary_tagged_stored.get_binary().has_subtype());
}
}
-371
View File
@@ -12,10 +12,6 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <sstream> // stringstream
#include <string> // string
#include <vector> // vector
namespace
{
// shortcut to scan a string literal
@@ -228,370 +224,3 @@ TEST_CASE("lexer class")
CHECK((scan_string("/**//**//**/", true) == json::lexer::token_type::end_of_input));
}
}
TEST_CASE("lexer number fast path")
{
// The contiguous fast path (used for pointer/string input) must agree with
// the streaming byte path (used for std::istream) on token type, numeric
// value, and round-trip text for every well-formed number, and reject the
// same malformed numbers with the same message.
SECTION("contiguous vs streaming parity")
{
const std::vector<std::string> numbers =
{
"0", "-0", "1", "-1", "42", "-42", "10", "100", "1234567890",
"0.0", "-0.0", "3.14", "-3.14", "0.5", "-0.001", "123.456789",
"1e0", "1E0", "1e10", "1e-10", "1e+10", "1.5e3", "-2.5E-4",
"9223372036854775807", // INT64_MAX -> unsigned
"9223372036854775808", // INT64_MAX + 1 -> unsigned
"18446744073709551615", // UINT64_MAX -> unsigned
"18446744073709551616", // UINT64_MAX + 1 -> float
"-9223372036854775808", // INT64_MIN -> integer
"-9223372036854775809", // INT64_MIN - 1 -> float
"123456789012345678901234567890", // huge -> float
"0.30000000000000004", "2.2250738585072014e-308", "1e308",
// high-precision / wide-exponent values that exercise the
// std::from_chars (Eisel-Lemire) path beyond the Clinger subset
"1.7976931348623157e308", "1.2345678901234567e-250",
"9007199254740993", "5e-324", "1e-320"
};
for (const auto& n : numbers)
{
const std::string doc = "[" + n + "]";
// contiguous fast path
const json a = json::parse(doc);
// streaming byte path
std::stringstream ss(doc);
const json b = json::parse(ss);
CAPTURE(n);
CHECK(a == b);
CHECK(a.dump() == b.dump());
CHECK(a[0].type() == b[0].type());
}
}
SECTION("token type classification")
{
CHECK((scan_string("0") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("-1") == json::lexer::token_type::value_integer));
CHECK((scan_string("1.5") == json::lexer::token_type::value_float));
CHECK((scan_string("1e5") == json::lexer::token_type::value_float));
CHECK((scan_string("18446744073709551615") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("18446744073709551616") == json::lexer::token_type::value_float));
CHECK((scan_string("-9223372036854775808") == json::lexer::token_type::value_integer));
CHECK((scan_string("-9223372036854775809") == json::lexer::token_type::value_float));
}
SECTION("malformed numbers are rejected identically")
{
for (const char* bad :
{"-", "1.", "1e", "1e+", "1.2e", "01", "-01", "1..2", "1.2.3"
})
{
CAPTURE(bad);
// the contiguous fast path must decline and let the byte path report
const std::string doc = std::string("[") + bad + "]";
CHECK_FALSE(json::accept(doc));
std::stringstream ss(doc);
CHECK_FALSE(json::accept(ss));
}
}
#if !defined(JSON_NOEXCEPTION)
// these sections parse invalid input, which aborts when exceptions are off
SECTION("exhaustive grammar parity with the streaming path")
{
// The JSON number grammar is encoded twice: once as the scan_number()
// state machine and once as the contiguous fast path. Enumerate every
// short string over the number alphabet and require the two encodings to
// agree exactly - on acceptance, on the reported error, and on the parsed
// value - so they cannot drift apart.
const std::string alphabet = "01.eE+-";
// full outcome of parsing @a doc, so a mismatch in type, value, or error
// message is caught, not just a mismatch in acceptance
const auto outcome = [](const std::string & doc, bool streaming) -> std::string
{
try
{
if (streaming)
{
std::stringstream ss(doc);
const json j = json::parse(ss);
return std::string(j[0].type_name()) + '|' + j.dump();
}
const json j = json::parse(doc);
return std::string(j[0].type_name()) + '|' + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
std::vector<std::string> mismatches;
std::vector<std::string> tokens{""};
for (std::size_t length = 1; length <= 4; ++length)
{
std::vector<std::string> next;
next.reserve(tokens.size() * alphabet.size());
for (const auto& prefix : tokens)
{
for (const char c : alphabet)
{
next.push_back(prefix + c);
}
}
tokens = next;
for (const auto& token : tokens)
{
const std::string doc = "[" + token + "]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
// 7 + 49 + 343 + 2401 tokens
CHECK(tokens.size() == 2401);
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
SECTION("error positions match the streaming path")
{
// Rejecting identically is not enough: the fast path must also report the
// error at the same position as the byte path. A number directly followed
// by a newline is the interesting case, because the byte path reaches the
// newline (which resets the column) and then ungets it.
// returns the parse_error message, or "" if the document parsed
const auto contiguous_error = [](const std::string & doc) -> std::string
{
try
{
const json j = json::parse(doc);
static_cast<void>(j);
}
catch (const json::parse_error& e)
{
return {e.what()};
}
return {};
};
const auto streaming_error = [](const std::string & doc) -> std::string
{
try
{
std::stringstream ss(doc);
const json j = json::parse(ss);
static_cast<void>(j);
}
catch (const json::parse_error& e)
{
return {e.what()};
}
return {};
};
for (const char* bad :
{"[01\n]", "[00\n]", "[-01\n]", "{1\n}", "[1\n2]", "[1.2.3\n]",
"[1 \n2]", "[\n1\n2]", "1\n2", "[01\r\n]", "[1e\n]", "[-\n]"
})
{
CAPTURE(bad);
const std::string doc = bad;
const std::string contiguous_what = contiguous_error(doc);
CHECK_FALSE(contiguous_what.empty());
CHECK(contiguous_what == streaming_error(doc));
}
// the column must be the one the offending token actually starts at,
// not the 0 that an unget() across the newline used to leave behind
CHECK(contiguous_error("[01\n]") ==
"[json.exception.parse_error.101] parse error at line 1, column 3: "
"syntax error while parsing array - unexpected number literal; expected ']'");
}
#endif
}
TEST_CASE("lexer string fast path")
{
// Build a byte string from explicit values: a hex escape in a string
// literal swallows every following hex digit, which makes sequences like
// "\xC3\xA9b" mean something other than they look like.
const auto bytes = [](std::initializer_list<int> values)
{
std::string result;
for (const int value : values)
{
result.push_back(static_cast<char>(value));
}
return result;
};
#if !defined(JSON_NOEXCEPTION)
// the full outcome of parsing @a doc: the parsed value, or the exact error
// message, so a mismatch in either is caught. Only usable with exceptions
// on: parsing invalid input aborts when they are off.
const auto outcome = [](const std::string & doc, bool streaming) -> std::string
{
try
{
if (streaming)
{
std::stringstream ss(doc);
const json j = json::parse(ss);
return j.dump();
}
const json j = json::parse(doc);
return j.dump();
}
// not just parse_error: if a bulk scanner ever let ill-formed UTF-8
// through, dump() would throw type_error.316, and that has to surface
// as a reported mismatch rather than as an uncaught exception
catch (const json::exception& e)
{
return {e.what()};
}
};
#endif
// once at the start of the string, once past the first 8-byte SWAR word, so
// the bulk scanner sees each case with and without a run behind it
const std::vector<std::size_t> offsets{0, 9};
#if !defined(JSON_NOEXCEPTION)
SECTION("exhaustive contiguous vs streaming parity")
{
// ordinary ASCII, both specials, a control byte, characters that make
// the preceding backslash a valid escape, a UTF-8 lead byte of each
// length, a continuation byte, and a byte that is never valid
const std::vector<std::string> alphabet =
{
"a", "\"", "\\", "n", "u", "0", bytes({0x01}),
bytes({0xC3}), bytes({0xA9}), bytes({0xE4}), bytes({0xF0}),
bytes({0x80}), bytes({0xFF})
};
std::vector<std::string> mismatches;
std::vector<std::string> tokens{""};
for (std::size_t length = 1; length <= 3; ++length)
{
std::vector<std::string> next;
next.reserve(tokens.size() * alphabet.size());
for (const auto& prefix : tokens)
{
for (const auto& symbol : alphabet)
{
next.push_back(prefix + symbol);
}
}
tokens = next;
for (const auto& token : tokens)
{
for (const std::size_t offset : offsets)
{
const std::string doc = "[\"" + std::string(offset, 'a') + token + "\"]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
}
// 13 + 169 + 2197 tokens, each at two offsets
CHECK(tokens.size() == 2197);
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
SECTION("special bytes at every offset of the SWAR stride")
{
// The bulk scanner consumes 8 bytes at a time and then a tail; place
// every kind of byte that ends a run at each offset across two words,
// so multibyte sequences also straddle the word boundary.
const std::vector<std::string> specials =
{
"\"", "\\", bytes({0x01}), bytes({0x1F}), bytes({0x7F}),
bytes({0xC3, 0xA9}), bytes({0xE4, 0xB8, 0xAD}), bytes({0xF0, 0x9F, 0x98, 0x80}),
bytes({0xFF}), bytes({0xC3}), bytes({0xE4, 0xB8})
};
std::vector<std::string> mismatches;
for (std::size_t offset = 0; offset <= 17; ++offset)
{
for (const auto& special : specials)
{
const std::string doc = "[\"" + std::string(offset, 'a') + special + "\"]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
#endif
// json::accept() never throws, so the ranges stay covered without exceptions
SECTION("UTF-8 ranges are accepted and rejected as documented")
{
// The bulk validator must accept exactly what the byte-at-a-time
// scanner accepts, so pin the boundaries of every range it recognizes.
// aggregate, only ever brace-initialized below; default member
// initializers would stop it being an aggregate in C++11
struct utf8_case // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init)
{
std::string sequence;
bool valid;
const char* description;
};
const std::vector<utf8_case> cases =
{
{bytes({0xC2, 0x80}), true, "U+0080, shortest two-byte"},
{bytes({0xDF, 0xBF}), true, "U+07FF, longest two-byte"},
{bytes({0xC1, 0xBF}), false, "overlong two-byte"},
{bytes({0xC2, 0x7F}), false, "two-byte with bad continuation"},
{bytes({0xE0, 0xA0, 0x80}), true, "U+0800, shortest three-byte"},
{bytes({0xE0, 0x9F, 0xBF}), false, "overlong three-byte"},
{bytes({0xED, 0x9F, 0xBF}), true, "U+D7FF, just below the surrogates"},
{bytes({0xED, 0xA0, 0x80}), false, "surrogate U+D800"},
{bytes({0xED, 0xBF, 0xBF}), false, "surrogate U+DFFF"},
{bytes({0xEE, 0x80, 0x80}), true, "U+E000, just above the surrogates"},
{bytes({0xEF, 0xBF, 0xBF}), true, "U+FFFF"},
{bytes({0xF0, 0x90, 0x80, 0x80}), true, "U+10000, shortest four-byte"},
{bytes({0xF0, 0x8F, 0xBF, 0xBF}), false, "overlong four-byte"},
{bytes({0xF4, 0x8F, 0xBF, 0xBF}), true, "U+10FFFF, highest code point"},
{bytes({0xF4, 0x90, 0x80, 0x80}), false, "above U+10FFFF"},
{bytes({0xF5, 0x80, 0x80, 0x80}), false, "lead byte out of range"},
{bytes({0x80}), false, "bare continuation byte"},
{bytes({0xFF}), false, "byte that never appears in UTF-8"},
{bytes({0xC3}), false, "truncated two-byte"},
{bytes({0xE4, 0xB8}), false, "truncated three-byte"},
{bytes({0xF0, 0x9F, 0x98}), false, "truncated four-byte"}
};
for (const auto& test_case : cases)
{
CAPTURE(test_case.description);
for (const std::size_t offset : offsets)
{
CAPTURE(offset);
const std::string doc = "[\"" + std::string(offset, 'a') + test_case.sequence + "\"]";
CHECK(json::accept(doc) == test_case.valid);
#if !defined(JSON_NOEXCEPTION)
CHECK(outcome(doc, false) == outcome(doc, true));
#endif
}
}
}
}
+66
View File
@@ -68,6 +68,72 @@ TEST_CASE("Better diagnostics with positions")
CHECK(j.end_pos() == root.size());
}
SECTION("copying keeps the positions of nested values (#5387)")
{
// Values nested deeper than the copy constructor's descent bound are
// copied without the call stack, on a path that has to carry the
// positions over itself; shallower ones copy their containers, which
// bring the positions along. Both sides of the bound are checked here.
const auto check_copy = [](std::size_t depth, bool objects)
{
CAPTURE(depth)
CAPTURE(objects)
const std::string opening = objects ? R"({"a":)" : "[";
const std::string closing = objects ? "}" : "]";
std::string text;
for (std::size_t i = 0; i < depth; ++i)
{
text += opening;
}
text += "12";
for (std::size_t i = 0; i < depth; ++i)
{
text += closing;
}
const json original = json::parse(text);
const json copy(original); // NOLINT(performance-unnecessary-copy-initialization)
const json* o = &original;
const json* c = &copy;
for (std::size_t level = 0; level <= depth; ++level)
{
CAPTURE(level)
REQUIRE(c->start_pos() == o->start_pos());
REQUIRE(c->end_pos() == o->end_pos());
if (level < depth)
{
o = objects ? &o->at("a") : &o->at(0);
c = objects ? &c->at("a") : &c->at(0);
}
}
};
const auto check_arrays = [&check_copy](std::size_t depth)
{
check_copy(depth, false);
};
const auto check_objects = [&check_copy](std::size_t depth)
{
check_copy(depth, true);
};
check_arrays(1);
check_arrays(127);
check_arrays(128);
check_arrays(129);
check_arrays(300);
check_objects(1);
check_objects(127);
check_objects(128);
check_objects(129);
check_objects(300);
}
SECTION("JSON patch add to primitive parent (#4292)")
{
// the JSON Patch "add" target /foo/bar/baz has a string parent
+57
View File
@@ -273,5 +273,62 @@ TEST_CASE("Regression tests for extended diagnostics")
CHECK(j1["numbers"]["two"] == 2);
CHECK(j1["string"] == "t");
}
SECTION("Regression test for issue #5387 - copying keeps the parents of nested values")
{
// A value nested deeper than the copy constructor's descent bound is
// copied without the call stack. Every container that path creates has
// to have the parents of its children set, or the JSON Pointer in the
// diagnostic is cut short.
const std::size_t depth = 300;
SECTION("objects")
{
json j = "not a number";
std::string pointer;
for (std::size_t i = 0; i < depth; ++i)
{
j = json{{"a", j}};
pointer += "/a";
}
json const copy(j); // NOLINT(performance-unnecessary-copy-initialization)
const json* inner = &copy;
for (std::size_t i = 0; i < depth; ++i)
{
inner = &inner->at("a");
}
std::string const expected = "[json.exception.type_error.302] (" + pointer + ") type must be number, but is string";
int i = 0;
CHECK_THROWS_WITH_AS(i = inner->get<int>(), expected.c_str(), json::type_error);
CHECK(i == 0);
}
SECTION("arrays")
{
json j = "not a number";
std::string pointer;
for (std::size_t i = 0; i < depth; ++i)
{
j = json::array({j});
pointer += "/0";
}
json const copy(j); // NOLINT(performance-unnecessary-copy-initialization)
const json* inner = &copy;
for (std::size_t i = 0; i < depth; ++i)
{
inner = &inner->at(0);
}
std::string const expected = "[json.exception.type_error.302] (" + pointer + ") type must be number, but is string";
int i = 0;
CHECK_THROWS_WITH_AS(i = inner->get<int>(), expected.c_str(), json::type_error);
CHECK(i == 0);
}
}
}
+151
View File
@@ -12,6 +12,7 @@
using nlohmann::json;
#include <algorithm>
#include <string>
TEST_CASE("tests on very large JSONs")
{
@@ -27,3 +28,153 @@ TEST_CASE("tests on very large JSONs")
}
}
namespace
{
// Descend a chain of single-element containers and return the value at its end,
// reporting the number of levels traversed in @a depth.
//
// The values in the test case below are nested far deeper than the call stack
// can follow, so they must not be inspected with operator== or dump(): both are
// still recursive and would overflow the stack themselves.
const json* innermost_value(const json& j, std::size_t& depth)
{
const json* current = &j;
depth = 0;
while ((current->is_array() || current->is_object()) && !current->empty())
{
current = current->is_array()
? &current->front()
: &current->begin().value();
++depth;
}
return current;
}
} // namespace
TEST_CASE("tests on deeply nested JSONs")
{
// deep enough to exhaust the call stack, but small enough to stay cheap:
// parsing is iterative, so building the values below costs little
const std::size_t depth = 100000;
SECTION("issue #5387 - stack overflow in the copy constructor")
{
SECTION("array")
{
const json j = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
const json copy(j); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t copy_depth = 0;
CHECK(*innermost_value(copy, copy_depth) == 0);
CHECK(copy_depth == depth);
}
SECTION("object")
{
std::string s;
s.reserve((6 * depth) + 1);
for (std::size_t i = 0; i < depth; ++i)
{
s += "{\"a\":";
}
s += '1';
s.append(depth, '}');
const json j = json::parse(s);
const json copy(j); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t copy_depth = 0;
CHECK(*innermost_value(copy, copy_depth) == 1);
CHECK(copy_depth == depth);
}
SECTION("copy assignment")
{
// operator=(basic_json) takes its argument by value, so the deep
// copy happens in the copy constructor
const json j = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
json target;
target = j;
std::size_t target_depth = 0;
CHECK(*innermost_value(target, target_depth) == 0);
CHECK(target_depth == depth);
}
SECTION("depths around the bound of the recursive descent")
{
// The copy constructor descends into a bounded number of levels and
// completes whatever is below that without the call stack. Cover
// every depth around that bound, so that the two ways of copying
// are known to meet cleanly - wherever the bound is set.
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d);
const json array = json::parse(std::string(d, '[') + '0' + std::string(d, ']'));
const json array_copy(array); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t array_depth = 0;
CHECK(*innermost_value(array_copy, array_depth) == 0);
CHECK(array_depth == d);
std::string object_text;
for (std::size_t i = 0; i < d; ++i)
{
object_text += "{\"a\":";
}
object_text += '1';
object_text.append(d, '}');
const json object = json::parse(object_text);
const json object_copy(object); // NOLINT(performance-unnecessary-copy-initialization): the copy is what is tested
std::size_t object_depth = 0;
CHECK(*innermost_value(object_copy, object_depth) == 1);
CHECK(object_depth == d);
}
}
SECTION("a value that is deep in one place only")
{
json j = json::object();
j["shallow"] = 1;
j["deep"] = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
j["also_shallow"] = json::array({1, 2, 3});
const json copy(j);
CHECK(copy["shallow"] == 1);
CHECK(copy["also_shallow"] == json::array({1, 2, 3}));
std::size_t deep_depth = 0;
CHECK(*innermost_value(copy["deep"], deep_depth) == 0);
CHECK(deep_depth == depth);
}
SECTION("the copy is independent of the original")
{
const json j = json::parse(std::string(depth, '[') + '0' + std::string(depth, ']'));
json copy(j);
// reach the innermost value without recursing and replace it
json* current = &copy;
while (current->is_array() && !current->empty())
{
current = &current->front();
}
*current = 42;
std::size_t unused = 0;
CHECK(*innermost_value(copy, unused) == 42);
CHECK(*innermost_value(j, unused) == 0);
}
}
}
+34
View File
@@ -81,3 +81,37 @@ TEST_CASE("regression test for issue #3732 - iteration_proxy_value<iter_impl<ord
};
static_cast<void>(fn);
}
TEST_CASE("copying an ordered_json with nested values")
{
// ordered_map is backed by a vector, so copying an object that has
// structured values takes a different route than copying a std::map-backed
// one; see https://github.com/nlohmann/json/issues/5387
ordered_json oj;
oj["z"] = 1;
oj["a"]["y"] = 2;
oj["a"]["b"]["x"] = 3;
oj["m"] = {1, 2, {{"w", 4}}};
const ordered_json copy(oj);
SECTION("the copy is equal to the original")
{
CHECK(copy == oj);
CHECK(copy.dump() == oj.dump());
}
SECTION("the key order is preserved at every level")
{
CHECK(copy.dump() == R"({"z":1,"a":{"y":2,"b":{"x":3}},"m":[1,2,{"w":4}]})");
}
SECTION("the copy is independent of the original")
{
ordered_json mutated(oj);
mutated["a"]["b"]["x"] = 99;
CHECK(oj["a"]["b"]["x"] == 3);
CHECK(mutated["a"]["b"]["x"] == 99);
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
-176
View File
@@ -19,8 +19,6 @@
using nlohmann::json;
#include <list>
#include <string> // string
#include <vector> // vector
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
#include <iterator>
@@ -230,180 +228,6 @@ TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t")
const std::counted_iterator<iterator_type> first2(json_str.begin(), len);
CHECK(json::accept(first2, std::default_sentinel));
}
TEST_CASE("std::counted_iterator reaches the contiguous fast paths")
{
// A sized sentinel makes the remaining element count computable in O(1), so
// std::counted_iterator over a contiguous iterator must reach the same bulk
// string/number scanners as a plain pointer - not just the byte-at-a-time
// fallback (see #5268 for the equivalent memcpy fast path).
#if JSON_HAS_RANGES
// JSON_HAS_RANGES is 0 on standard libraries with an incomplete <ranges>
// (libstdc++ < 11, libc++ < 16), where the adapter deliberately falls back
// to the byte-at-a-time scanner; everything below still has to work there.
using adapter_type = nlohmann::detail::iterator_input_adapter<std::counted_iterator<const char*>, std::default_sentinel_t>;
CHECK(adapter_type::supports_bulk_scan);
CHECK(adapter_type::supports_seek);
#endif
// exercise every fast path: long ASCII run, multibyte UTF-8, escapes, and
// integer/floating-point numbers
const std::string json_str =
R"({"ascii":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",)"
"\"utf8\":\"\xe4\xb8\xad\xe6\x96\x87\xf0\x9f\x98\x80\xc3\xa9\","
R"("escaped":"aéb\n\\","ints":[0,-1,18446744073709551615,-9223372036854775808],)"
R"("floats":[1.5,-2.25e3,0.30000000000000004]})";
const auto len = static_cast<std::iter_difference_t<const char*>>(json_str.size());
const std::counted_iterator<const char*> first(json_str.data(), len);
const json j = json::parse(first, std::default_sentinel);
// parsing through the pointer adapter must give exactly the same result
CHECK(j == json::parse(json_str));
#if !defined(JSON_NOEXCEPTION)
// Diagnostics that quote the offending token are reconstructed from the
// already-consumed input (supports_seek), a path a sized sentinel only
// reaches now; check a few that include the "last read" text. Parsing
// invalid input aborts when exceptions are off, hence the guard.
// Raw strings and explicit bytes: an escaped literal and two literals
// written next to each other both read as mistakes to static analysis.
const auto byte = [](int value)
{
return std::string(1, static_cast<char>(value));
};
const std::vector<std::string> diagnostic_docs =
{
"1\nx",
"truX",
"[tru]",
R"("abc)",
R"(["\ud834"])",
R"(["a)" + byte(0x01) + R"(b"])",
R"([")" + byte(0xC3) + byte(0x28) + R"("])",
"[1e]",
R"(["aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaX)"
};
for (const auto& text : diagnostic_docs)
{
CAPTURE(text);
const std::counted_iterator<const char*> it(text.data(), static_cast<std::iter_difference_t<const char*>>(text.size()));
std::string counted_message;
std::string string_message;
try
{
const json counted_result = json::parse(it, std::default_sentinel);
static_cast<void>(counted_result);
}
catch (const json::parse_error& e)
{
counted_message = e.what();
}
try
{
const json string_result = json::parse(text);
static_cast<void>(string_result);
}
catch (const json::parse_error& e)
{
string_message = e.what();
}
CHECK_FALSE(counted_message.empty());
CHECK(counted_message == string_message);
}
// and errors must still be reported identically
const std::string bad = "[01\n]";
const std::counted_iterator<const char*> bad_first(bad.data(), static_cast<std::iter_difference_t<const char*>>(bad.size()));
std::string counted_what;
std::string string_what;
try
{
const json counted_result = json::parse(bad_first, std::default_sentinel);
static_cast<void>(counted_result);
}
catch (const json::parse_error& e)
{
counted_what = e.what();
}
try
{
const json string_result = json::parse(bad);
static_cast<void>(string_result);
}
catch (const json::parse_error& e)
{
string_what = e.what();
}
CHECK_FALSE(counted_what.empty());
CHECK(counted_what == string_what);
#endif
}
#if !defined(JSON_NOEXCEPTION)
// several cases below are truncated on purpose, and parsing invalid input
// aborts when exceptions are off
TEST_CASE("std::counted_iterator bulk scanning stops at the counted end")
{
// The count, not the size of the underlying buffer, is the end of the
// input: the bulk scanners must never look at the bytes behind it, even
// though they are readable. Each case is compared against parsing the
// equivalent prefix as a std::string.
const auto via_counted = [](const std::string & buf, std::size_t n) -> std::string
{
const std::counted_iterator<const char*> first(buf.data(), static_cast<std::iter_difference_t<const char*>>(n));
try
{
const json j = json::parse(first, std::default_sentinel);
return "OK|" + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
const auto via_prefix = [](const std::string & buf, std::size_t n) -> std::string
{
try
{
const json j = json::parse(buf.substr(0, n));
return "OK|" + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
struct testcase // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init)
{
const char* buffer;
std::size_t count;
};
const std::vector<testcase> cases =
{
{"[\"abc\"]____TRAILING____", 7}, // exact fit, tail hidden
{"[\"abcdefghijklmnop\"]____", 8}, // cut inside a string
{"[\"abc\"]____", 6}, // cut just before the closing quote
{"[12345]xxxxx", 4}, // cut inside a number
{"[123]999999", 5}, // number ends exactly at the count
{"[\"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\"]", 12}, // closing quote only behind the count
{"[\"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\"]", 19}, // cut inside an 8-byte SWAR stride
{"[\"\xe4\xb8\xad\xe6\x96\x87\"]", 5}, // cut inside a UTF-8 sequence
{"[\"\xe4\xb8\xad\xe6\x96\x87\"]____", 10}, // complete UTF-8, tail hidden
{"[1.25e3]TRAILINGDIGITS999", 7}, // number token reaches the count
};
for (const auto& tc : cases)
{
CAPTURE(tc.buffer);
CAPTURE(tc.count);
const std::string buffer = tc.buffer;
CHECK(via_counted(buffer, tc.count) == via_prefix(buffer, tc.count));
}
}
#endif
#endif
} // namespace