Compare commits

..
Author SHA1 Message Date
Niels Lohmann d91dff77e6 Split unit-regression2.cpp so the MinGW linker can relocate it
Linking test-regression2 with clang and MinGW fails with

    relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'

once the translation unit grows past a certain size: the code can no longer
reach the read-only data it references within the range of a 32-bit
relocation. The file is one of the largest in the test suite and had been
sitting just under that limit, so an unrelated change elsewhere in the
library is enough to tip it over. It is already the second such file --
unit-regression1.cpp was split for size before -- and windows.yml already
carries a workaround for the same limit hitting the debug sections of this
same target, where -g0 was enough because that relocation was against
`.debug_line'. This one is against `.rdata', which no compiler flag avoids.

Move the second half of the regression tests, and the helper types only they
use, into unit-regression3.cpp. The sections are independent -- every
statement in "regression tests 2" was already inside a SECTION -- so they
move unchanged, and the counts confirm nothing was lost: 168 assertions
before the split, 50 plus 118 after.

The result is that both files are comfortably smaller than the one that used
to link, measured with clang at -O1 for C++20:

                        read-only data        text     object
    before                      58,233   1,287,764  3,158,120
    unit-regression2.cpp        48,161   1,012,988  2,522,296
    unit-regression3.cpp        41,710     772,704  1,878,880

No CMake change is needed: tests/CMakeLists.txt globs src/unit-*.cpp, so the
new file is picked up and built for every standard like its siblings.

CONTRIBUTING.md pointed contributors at unit-regression2.cpp for new bug
tests; it now points at the smaller file and says why the two exist, so the
split does not quietly undo itself.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 10:21:21 +02:00
Niels Lohmann 9edfb53906 Note the 1,048,576 valueless-array limit as (1 << 20) in the docs
Addresses review feedback from @gregmarr on PR #5504: spell out the
binary/hex form next to the decimal count so it reads as the round
power-of-two it is, matching how include/nlohmann/detail/input/binary_reader.hpp
defines max_valueless_container_size. Applied in both docs/exceptions.md
and ubjson.md, as requested.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 10:21:20 +02:00
Niels Lohmann e27d1192e0 Bound UBJSON optimized arrays of a valueless type
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized UBJSON array of one of those has no
payload: reading an element consumes no input at all. Its declared count is
therefore the only thing that decides how much is allocated, and nothing
bounded it. "[$Z#l" and a four-byte count is nine bytes of input describing
two billion values; #2793 reports 35 GB and 150 seconds from ten bytes, and
OSS-Fuzz has an out-of-memory and a timeout report for the same shape.

Every other type costs at least one byte per element, so the end of the input
bounds it. 'N' (no-op) is already skipped rather than stored. Objects are not
affected either: each element is preceded by its key, which costs bytes. And
BJData already refuses these markers as an optimized type, so this is a plain
UBJSON matter.

Reject a count above 1,048,576 elements for those three types with
out_of_range.408, the code this reader already uses for a declared size it
will not honour. The check runs before the SAX start event, so no container
is opened and then abandoned.

Rejecting on the read side alone would break the guarantee that anything
to_ubjson() writes can be read back, and would trip the round-trip assertion
in fuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized
encoding, one byte per element, for arrays of these types above the same
limit. Its decision depends only on the array's size, which is identical for
a value and for anything parsed back from it, so the round trip is stable.

No existing test changes: the largest such count in the test suite is 65,793.
The excessive-size test that already used this shape still passes, now
rejected a little earlier than by the max_size() check it used to reach.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 10:21:20 +02:00
Niels Lohmann 47643785a6 Reject a nested BJData ndarray dimension vector where it is read
get_ubjson_size_type() takes an inside_ndarray parameter saying whether it is
being called for an ndarray's dimension vector, where another ndarray is not
allowed. It then seeded the flag it passes down to get_ubjson_size_value()
with `false` rather than with that parameter, and only consulted
inside_ndarray afterwards, on the '$' branch.

So on the '#' branch nothing stopped the descent: every "#[" pair of an input
like "[" followed by "#[#[#[..." opened another dimension vector, several
native stack frames deeper each time, and the recursion was only reported on
the way back out. 100,000 pairs crash the process. This is #5104 again, in a
path that has nothing to do with containers.

Seed the flag with inside_ndarray, which is what get_ubjson_size_value()
documents it wants: "for input, `true` means already inside an ndarray vector
or ndarray dimension is not allowed". The nested '[' is then refused where it
is read, so the length of the chain no longer matters.

Both post-checks gain `&& !inside_ndarray`, because an ndarray was found
*here* only if the flag flipped -- get_ubjson_size_value() only ever returns
`true` when its initial value was `false`, as its documentation says. With
that, the "ndarray can not be recursive" branch is unreachable: a recursive
ndarray is now caught one level earlier, and reported as "ndarray dimensional
vector is not allowed" like every other nested dimension vector.

Three existing expectations move accordingly (vR2, vR4, vR6). All three now
fail earlier, and all three now report the same error that vR1, vR5 and vH
already reported for the same shape, which is the more consistent outcome.
Everything else is unchanged: valid 1D and 2D ndarrays, optimized containers
and plain arrays produce identical results, and unit-ubjson is untouched.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 10:21:20 +02:00
Niels Lohmann d61ef62fc7 Stop CBOR indefinite-length strings from recursing per chunk
get_cbor_string() and get_cbor_binary() handled the indefinite-length forms
(0x7F and 0x5F) by calling themselves once per chunk. Each chunk therefore
cost a native stack frame, and since a chunk may itself be an indefinite-
length string, an input of repeated 0x7F bytes reached one frame per input
byte: 200,000 of them crash the process with SIGSEGV before a single byte is
rejected. This is the same defect as #5104, in a path the container-level
work does not touch.

Count the open levels instead of recursing through them. That is enough here
because every chunk is appended to the same result -- get_bytes() writes at
result.size() -- so there is no per-level state to keep. The temporary chunk
string and its copy into the result go away with the recursion.

The definite-length cases move to get_cbor_string_chunk() and
get_cbor_binary_chunk() unchanged, including their error messages, which
still name 0x7F and 0x5F because those are handled one level up.

Behaviour is unchanged. Comparing against develop over the interesting byte
sequences -- empty, single-chunk, nested, over-closed and truncated forms,
both strings and byte arrays, and an indefinite-length map key -- produces
identical values, error codes, messages and byte offsets. The 200,000-level
input now reports parse_error.110 at byte 200001 instead of crashing.

Note that nesting these is not valid CBOR: RFC 8949, Section 3.2.3 forbids
it. This does not change that either way -- it has always been accepted, and
rejecting it is a separate decision (#5317, #5325). Should it be rejected
later, that is now one condition on the level counter rather than a change to
the control flow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 10:21:19 +02:00
Niels Lohmann 79f990dd04 Return the parsed value by move from from_cbor() and friends
The binary entry points end with

    return res ? result : basic_json(value_t::discarded);

The condition operator's second operand is an lvalue, so this is not a case
where the return value can be elided or implicitly moved from: every
successful from_cbor(), from_msgpack(), from_ubjson(), from_bjdata() and
from_bson() call deep-copies the value it just parsed, and then destroys the
original.

The copy is not cheap, and it is not incidental: basic_json's copy
constructor walks the whole value. Parsing a 2 MB CBOR document with 60,000
objects, median of 25 runs, clang 17 -O3:

    from_cbor      26.99 ms  ->  14.65 ms
    from_msgpack   26.82 ms  ->  14.82 ms

Moving instead of copying is the entire change; the parsed value is not used
again after the return expression is evaluated.

There is a second reason to prefer the move. The copy constructor recurses
once per nesting level, so the copy is also a stack-overflow path on the
return side, on a value the reader has already accepted. That is currently
masked because the readers themselves recurse and overflow first (#5104), but
it has to be fixed for making them iterative to have any effect.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-09 10:21:19 +02:00
28 changed files with 1905 additions and 5875 deletions
+3 -1
View File
@@ -108,7 +108,9 @@ The tests are located in [`tests/src/unit-*.cpp`](https://github.com/nlohmann/js
are structured along the features of the library or the nature of the tests. Usually, it should be clear from the are structured along the features of the library or the nature of the tests. Usually, it should be clear from the
context which existing file needs to be extended, and only very few cases require creating new test files. context which existing file needs to be extended, and only very few cases require creating new test files.
When fixing a bug, edit `unit-regression2.cpp` and add a section referencing the fixed issue. When fixing a bug, edit `unit-regression3.cpp` and add a section referencing the fixed issue.
`unit-regression2.cpp` holds the older tests; the two files exist because a single one grew large enough for the
MinGW linker to fail relocating it, so please keep adding to the smaller file rather than growing the larger one.
#### Exceptions #### Exceptions
+1 -1
View File
@@ -100,7 +100,7 @@ jobs:
container: ubuntu:focal container: ubuntu:focal
strategy: strategy:
matrix: matrix:
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_simdutf] target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls]
steps: steps:
- name: Install build-essential - name: Install build-essential
run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev
-18
View File
@@ -212,24 +212,6 @@ add_custom_target(ci_test_legacycomparison
COMMENT "Compile and test with legacy discarded value comparison enabled" COMMENT "Compile and test with legacy discarded value comparison enabled"
) )
###############################################################################
# Validate UTF-8 with simdutf.
###############################################################################
add_custom_target(ci_test_simdutf
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON -DJSON_TestSimdutf=ON
# simdutf needs C++17, so the library falls back to its scalar validator
# below that: build the suite at C++11 to cover the fallback with the macro
# defined, and at C++17 to run every test against simdutf itself
"-DJSON_TestStandards=11\;17"
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_simdutf
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_simdutf
COMMAND cd ${PROJECT_BINARY_DIR}/build_simdutf && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure
COMMENT "Compile and test with simdutf UTF-8 validation enabled"
)
############################################################################### ###############################################################################
# Enable brace-init copy semantics. # Enable brace-init copy semantics.
############################################################################### ###############################################################################
-1
View File
@@ -24,7 +24,6 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers - [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers - [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace - [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
- [**JSON_USE_SIMDUTF**](json_use_simdutf.md) - use the simdutf library to accelerate UTF-8 validation
## Library version ## Library version
@@ -1,71 +0,0 @@
# JSON_USE_SIMDUTF
```cpp
#define JSON_USE_SIMDUTF
```
When defined, the parser validates the UTF-8 content of JSON strings that come from a **contiguous byte input**
(`std::string`, `std::vector<char>`/`<std::uint8_t>`, string literals, `const char*` ranges, …) using the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. On text with many
non-ASCII characters (e.g. CJK or emoji) this can validate several times faster.
This is an **opt-in external dependency**. The library itself remains header-only and its behavior is unchanged: the
same input is accepted or rejected either way, and every parse error is reported at the same position with the same
message (simdutf is only used to fast-path *valid* runs; anything it flags falls back to the scalar path so the exact
diagnostic is preserved). Streaming inputs (files, `std::istream`, wide strings, user-defined adapters) always use the
scalar path.
When `JSON_USE_SIMDUTF` is defined you must make the `simdutf.h` header available on the include path and link the
simdutf library. When it is not defined, no simdutf header is included and there is no dependency.
!!! note "Requires C++17"
simdutf requires C++17 and its header rejects older standards with an `#!cpp #error`. The backend is therefore only
compiled in from C++17 on. In C++11 and C++14 the macro has no effect and the scalar validator is used, which
accepts and rejects exactly the same input -- only throughput differs. Setting the macro project-wide is therefore
safe even when some translation units are built with an older standard.
!!! warning "Define consistently"
The macro selects between two definitions of the same inline validation function. It must therefore be defined
identically for **every** translation unit that includes the library; mixing translation units that define it with
ones that do not is an ODR violation. Prefer setting it as a compile definition on the target rather than with
`#!cpp #define` in individual source files.
## Default definition
By default, `#!cpp JSON_USE_SIMDUTF` is not defined and the portable C++11 scalar validator is used.
```cpp
#undef JSON_USE_SIMDUTF
```
## Examples
??? example
The code below enables the simdutf backend for UTF-8 validation.
```cpp
#define JSON_USE_SIMDUTF 1
#include <nlohmann/json.hpp>
...
```
The project must also link against simdutf, e.g. with CMake:
```cmake
target_compile_definitions(your_target PRIVATE JSON_USE_SIMDUTF)
target_link_libraries(your_target PRIVATE simdutf::simdutf)
```
!!! hint "Testing this configuration"
The unit tests can be built against the simdutf backend with the CMake option `JSON_TestSimdutf` (`OFF` by
default), which fetches simdutf and defines `JSON_USE_SIMDUTF` for every test target. The `ci_test_simdutf` target
runs the whole test suite in that configuration.
## Version history
- Added in version 3.13.0.
@@ -69,6 +69,13 @@ The library uses the following mapping from JSON values types to UBJSON types ac
Note that `use_size = true` alone may result in larger representations - the benefit of this parameter is that the Note that `use_size = true` alone may result in larger representations - the benefit of this parameter is that the
receiving side is immediately informed on the number of elements of the container. receiving side is immediately informed on the number of elements of the container.
An array whose type marker is `Z` (null), `T` (true) or `F` (false) stores no payload at all, because the marker
already is the value. Its declared count is therefore the only thing that decides how much memory the receiving side
allocates, and a handful of bytes can describe billions of elements. `from_ubjson` rejects such an array with
[`out_of_range.408`](../../home/exceptions.md#jsonexceptionout_of_range408) when the count exceeds 1,048,576
(`1 << 20`), and `to_ubjson` writes longer arrays of these types without the annotation, so any value it produces
can be read back.
!!! info "Binary values" !!! info "Binary values"
If the JSON data contains the binary type, the value stored is a list of integers, as suggested by the UBJSON If the JSON data contains the binary type, the value stored is a list of integers, as suggested by the UBJSON
-8
View File
@@ -137,14 +137,6 @@ behavior is deprecated and switched off (`0`) by default.
See [full documentation of `JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`](../api/macros/json_use_legacy_discarded_value_comparison.md). See [full documentation of `JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`](../api/macros/json_use_legacy_discarded_value_comparison.md).
## `JSON_USE_SIMDUTF`
When defined, UTF-8 validation of JSON strings read from contiguous byte input is delegated to the
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. This is an opt-in
external dependency and is not defined by default.
See [full documentation of `JSON_USE_SIMDUTF`](../api/macros/json_use_simdutf.md).
## `NLOHMANN_DEFINE_TYPE_*(...)`, `NLOHMANN_DEFINE_DERIVED_TYPE_*(...)` ## `NLOHMANN_DEFINE_TYPE_*(...)`, `NLOHMANN_DEFINE_DERIVED_TYPE_*(...)`
The library defines 12 macros to simplify the serialization/deserialization of types. See the page on The library defines 12 macros to simplify the serialization/deserialization of types. See the page on
+9
View File
@@ -868,6 +868,12 @@ The size of an array or object in a [binary format](../features/binary_formats/i
the size following `#` for [UBJSON](../features/binary_formats/ubjson.md)/[BJData](../features/binary_formats/bjdata.md), the size following `#` for [UBJSON](../features/binary_formats/ubjson.md)/[BJData](../features/binary_formats/bjdata.md),
or the encoded length for [CBOR](../features/binary_formats/cbor.md). or the encoded length for [CBOR](../features/binary_formats/cbor.md).
The exception is also thrown for a [UBJSON](../features/binary_formats/ubjson.md) array of a type that is encoded by its
marker alone (`Z`, `T` or `F`) whose declared count exceeds 1,048,576 (`1 << 20`). Such an array has no payload, so its
count alone decides how much memory is allocated, and a handful of bytes would otherwise describe billions of values.
[`to_ubjson`](../api/basic_json/to_ubjson.md) writes longer arrays of these types without the size and type annotation,
so any value it produces can still be read back.
!!! failure "Example messages" !!! failure "Example messages"
``` ```
@@ -879,6 +885,9 @@ or the encoded length for [CBOR](../features/binary_formats/cbor.md).
``` ```
[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size [json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size
``` ```
```
[json.exception.out_of_range.408] syntax error while parsing UBJSON size: excessive array size
```
### json.exception.out_of_range.409 ### json.exception.out_of_range.409
-1
View File
@@ -296,7 +296,6 @@ nav:
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md - 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md - 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
- 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md - 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md
- 'JSON_USE_SIMDUTF': api/macros/json_use_simdutf.md
- 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md - 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md
- 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md - 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md
- 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md - 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md
+179 -59
View File
@@ -58,6 +58,26 @@ inline bool little_endianness(int num = 1) noexcept
return *reinterpret_cast<char*>(&num) == 1; return *reinterpret_cast<char*>(&num) == 1;
} }
/*!
@brief largest element count accepted for a UBJSON container of a valueless type
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized container of one of those types has no
payload at all and its declared count is the only thing that decides how much
is allocated: `[$Z#L` followed by a large count turns some ten bytes of input
into that many values (see #2793, which reports 35 GB and 150 seconds). Every
other type costs at least one byte per element and is bounded by the end of
the input.
This is a sanity bound rather than a security boundary, and it is far above
any container met in practice. @ref binary_writer falls back to the
unoptimized encoding for longer containers, so that a value serialized by
this library can always be read back.
@sa https://github.com/nlohmann/json/issues/2793
*/
JSON_INLINE_VARIABLE constexpr std::size_t max_valueless_container_size = 1 << 20;
/////////////////// ///////////////////
// binary reader // // binary reader //
/////////////////// ///////////////////
@@ -996,23 +1016,21 @@ class binary_reader
} }
/*! /*!
@brief reads a CBOR string @brief reads a definite-length CBOR string
This function first reads starting bytes to determine the expected Reads everything @ref get_cbor_string accepts except the indefinite-length
string length and then copies this number of bytes into a string. form, which that function handles itself. The bytes are appended to @a
Additionally, CBOR's strings with indefinite lengths are supported. result, so consecutive chunks of an indefinite-length string can be read
into the same string.
@param[out] result created string @param[out] result string the bytes are appended to
@return whether string creation completed @return whether string creation completed
*/
bool get_cbor_string(string_t& result)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "string")))
{
return false;
}
@pre @a current is not EOF
*/
bool get_cbor_string_chunk(string_t& result)
{
switch (current) switch (current)
{ {
// UTF-8 string (0x00..0x17 bytes follow) // UTF-8 string (0x00..0x17 bytes follow)
@@ -1068,20 +1086,6 @@ class binary_reader
return get_number(input_format_t::cbor, len) && get_string(input_format_t::cbor, len, result); return get_number(input_format_t::cbor, len) && get_string(input_format_t::cbor, len, result);
} }
case 0x7F: // UTF-8 string (indefinite length)
{
while (get() != 0xFF)
{
string_t chunk;
if (!get_cbor_string(chunk))
{
return false;
}
result.append(chunk);
}
return true;
}
default: default:
{ {
auto last_token = get_token_string(); auto last_token = get_token_string();
@@ -1092,23 +1096,82 @@ class binary_reader
} }
/*! /*!
@brief reads a CBOR byte array @brief reads a CBOR string
This function first reads starting bytes to determine the expected This function first reads starting bytes to determine the expected
byte array length and then copies this number of bytes into the byte array. string length and then copies this number of bytes into a string.
Additionally, CBOR's byte arrays with indefinite lengths are supported. Additionally, CBOR's strings with indefinite lengths are supported.
@param[out] result created byte array @param[out] result created string
@return whether string creation completed
*/
bool get_cbor_string(string_t& result)
{
// number of indefinite-length strings that have been opened and not
// closed yet. RFC 8949, Section 3.2.3 does not permit nesting them,
// but this reader has always accepted it, so the open levels are
// counted instead of recursed through, which overflowed the stack for
// an input of repeated 0x7F bytes (see #5104). Every chunk is appended
// to the same result, so no per-level state is needed.
std::size_t open = 0;
while (true)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "string")))
{
return false;
}
if (current == 0x7F) // UTF-8 string (indefinite length)
{
++open;
get();
continue;
}
// a break marker closes the innermost indefinite-length string;
// outside of one it is not a string and falls through to the error
if (open != 0 && current == 0xFF)
{
if (--open == 0)
{
return true;
}
get();
continue;
}
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string_chunk(result)))
{
return false;
}
if (open == 0)
{
return true;
}
get();
}
}
/*!
@brief reads a definite-length CBOR byte array
Reads everything @ref get_cbor_binary accepts except the indefinite-length
form, which that function handles itself. The bytes are appended to @a
result, so consecutive chunks of an indefinite-length byte array can be
read into the same byte array.
@param[out] result byte array the bytes are appended to
@return whether byte array creation completed @return whether byte array creation completed
*/
bool get_cbor_binary(binary_t& result)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "binary")))
{
return false;
}
@pre @a current is not EOF
*/
bool get_cbor_binary_chunk(binary_t& result)
{
switch (current) switch (current)
{ {
// Binary data (0x00..0x17 bytes follow) // Binary data (0x00..0x17 bytes follow)
@@ -1168,20 +1231,6 @@ class binary_reader
get_binary(input_format_t::cbor, len, result); get_binary(input_format_t::cbor, len, result);
} }
case 0x5F: // Binary data (indefinite length)
{
while (get() != 0xFF)
{
binary_t chunk;
if (!get_cbor_binary(chunk))
{
return false;
}
result.insert(result.end(), chunk.begin(), chunk.end());
}
return true;
}
default: default:
{ {
auto last_token = get_token_string(); auto last_token = get_token_string();
@@ -1191,6 +1240,63 @@ class binary_reader
} }
} }
/*!
@brief reads a CBOR byte array
This function first reads starting bytes to determine the expected
byte array length and then copies this number of bytes into the byte array.
Additionally, CBOR's byte arrays with indefinite lengths are supported.
@param[out] result created byte array
@return whether byte array creation completed
*/
bool get_cbor_binary(binary_t& result)
{
// the open indefinite-length byte arrays are counted rather than
// recursed through, for the reason given in @ref get_cbor_string
std::size_t open = 0;
while (true)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "binary")))
{
return false;
}
if (current == 0x5F) // Binary data (indefinite length)
{
++open;
get();
continue;
}
// a break marker closes the innermost indefinite-length byte
// array; outside of one it falls through to the error below
if (open != 0 && current == 0xFF)
{
if (--open == 0)
{
return true;
}
get();
continue;
}
if (JSON_HEDLEY_UNLIKELY(!get_cbor_binary_chunk(result)))
{
return false;
}
if (open == 0)
{
return true;
}
get();
}
}
/*! /*!
@brief narrow a definite CBOR array/map length to std::size_t @brief narrow a definite CBOR array/map length to std::size_t
@@ -2391,7 +2497,12 @@ class binary_reader
{ {
result.first = npos; // size result.first = npos; // size
result.second = 0; // type result.second = 0; // type
bool is_ndarray = false; // seed the flag with the caller's context: inside an ndarray dimension
// vector another ndarray is not allowed, and get_ubjson_size_value()
// rejects it up front instead of reading it and reporting afterwards.
// Seeding it with `false` made every '#' of a "[#[#[..." chain descend
// another level, which overflowed the stack (see #5104).
bool is_ndarray = inside_ndarray;
get_ignore_noop(); get_ignore_noop();
@@ -2424,13 +2535,11 @@ class binary_reader
} }
const bool is_error = get_ubjson_size_value(result.first, is_ndarray); const bool is_error = get_ubjson_size_value(result.first, is_ndarray);
if (input_format == input_format_t::bjdata && is_ndarray) // an ndarray was read here only if the flag flipped; when it was
// seeded true, get_ubjson_size_value() already rejected the nested
// dimension vector
if (input_format == input_format_t::bjdata && is_ndarray && !inside_ndarray)
{ {
if (inside_ndarray)
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format, "ndarray can not be recursive", "size"), nullptr));
}
result.second |= (1 << 8); // use bit 8 to indicate ndarray, all UBJSON and BJData markers should be ASCII letters result.second |= (1 << 8); // use bit 8 to indicate ndarray, all UBJSON and BJData markers should be ASCII letters
} }
return is_error; return is_error;
@@ -2439,7 +2548,7 @@ class binary_reader
if (current == '#') if (current == '#')
{ {
const bool is_error = get_ubjson_size_value(result.first, is_ndarray); const bool is_error = get_ubjson_size_value(result.first, is_ndarray);
if (input_format == input_format_t::bjdata && is_ndarray) if (input_format == input_format_t::bjdata && is_ndarray && !inside_ndarray)
{ {
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read, return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format, "ndarray requires both type and size", "size"), nullptr)); exception_message(input_format, "ndarray requires both type and size", "size"), nullptr));
@@ -2710,6 +2819,17 @@ class binary_reader
if (size_and_type.first != npos) if (size_and_type.first != npos)
{ {
// reading an element of a valueless type consumes no input, so the
// declared count alone decides how much is allocated; the check is
// made before the start event so that no container is opened that
// is then abandoned. See @ref max_valueless_container_size.
if (JSON_HEDLEY_UNLIKELY((size_and_type.second == 'Z' || size_and_type.second == 'T' || size_and_type.second == 'F')
&& size_and_type.first > max_valueless_container_size))
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408,
exception_message(input_format, "excessive array size", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(size_and_type.first))) if (JSON_HEDLEY_UNLIKELY(!sax->start_array(size_and_type.first)))
{ {
return false; return false;
+21 -131
View File
@@ -155,31 +155,11 @@ class input_stream_adapter
// General-purpose iterator-based adapter. It might not be as fast as // General-purpose iterator-based adapter. It might not be as fast as
// theoretically possible for some containers, but it is extremely versatile. // theoretically possible for some containers, but it is extremely versatile.
// SentinelType defaults to IteratorType for backward compatibility, but may be // SentinelType defaults to IteratorType for backward compatibility, but may
// a different type, e.g. a C++20 sentinel such as std::default_sentinel_t when // be a different type (e.g., a C++20 sentinel or counted_iterator).
// IteratorType is a std::counted_iterator.
template<typename IteratorType, typename SentinelType = IteratorType> template<typename IteratorType, typename SentinelType = IteratorType>
class iterator_input_adapter class iterator_input_adapter
{ {
// Whether the number of elements between two positions can be computed in
// O(1): either the iterator and the sentinel have the same type (plain
// std::distance) or, in C++20, the sentinel is a sized sentinel for the
// iterator (std::ranges::distance), e.g. std::default_sentinel_t paired
// with std::counted_iterator.
//
// JSON_HAS_RANGES gates the C++20 branch: on standard libraries with an
// incomplete <ranges> (libstdc++ < 11, see #4440) evaluating
// std::contiguous_iterator on a std::counted_iterator is a hard error
// instead of yielding false, and these traits are instantiated for every
// adapter. Such toolchains fall back to the pointer-only test and simply
// use the byte-at-a-time scanner.
static constexpr bool sentinel_is_sized =
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>;
#else
std::is_same<IteratorType, SentinelType>::value;
#endif
public: public:
using char_type = typename std::iterator_traits<IteratorType>::value_type; using char_type = typename std::iterator_traits<IteratorType>::value_type;
@@ -191,7 +171,7 @@ class iterator_input_adapter
// in wide_string_input_adapter, which does not expose this). // in wide_string_input_adapter, which does not expose this).
static constexpr bool supports_seek = static constexpr bool supports_seek =
std::is_same<typename std::iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value std::is_same<typename std::iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value
&& sentinel_is_sized && std::is_same<IteratorType, SentinelType>::value
&& sizeof(char_type) == 1; && sizeof(char_type) == 1;
iterator_input_adapter(IteratorType first, SentinelType last) iterator_input_adapter(IteratorType first, SentinelType last)
@@ -239,60 +219,30 @@ class iterator_input_adapter
private: private:
// whether IteratorType refers to a contiguous range and therefore supports // whether IteratorType refers to a contiguous range and therefore supports
// a std::memcpy fast path (pointers always do; in C++20 we can also detect // a std::memcpy fast path (pointers always do; in C++20 we can also detect
// library iterators such as those of std::vector and std::string). The // library iterators such as those of std::vector and std::string).
// available element count must also be computable in O(1), hence // Computing the available element count needs either same-type iterators
// sentinel_is_sized. // (plain std::distance) or, in C++20, a sized sentinel (std::ranges::distance),
static constexpr bool iterator_is_contiguous = sentinel_is_sized && // e.g. std::counted_iterator paired with std::default_sentinel_t.
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20) static constexpr bool iterator_is_contiguous =
(std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value); #if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>)
&& (std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
#else #else
std::is_pointer<IteratorType>::value; std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
#endif #endif
// number of unread elements in [current, end)
std::size_t remaining_count() const
{
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
return static_cast<std::size_t>(std::ranges::distance(current, end));
#else
return static_cast<std::size_t>(std::distance(current, end));
#endif
}
public:
// Whether the remaining input is a single contiguous block of 1-byte
// elements that the lexer can inspect directly (used for the SWAR string
// fast path).
static constexpr bool supports_bulk_scan =
iterator_is_contiguous && sizeof(char_type) == 1;
// Pointer to the next unread element; only valid when bulk_remaining() > 0.
const char_type* bulk_data() const
{
return &*current;
}
// Number of unread elements available as one contiguous block.
std::size_t bulk_remaining() const
{
return remaining_count();
}
// Consume @a n elements previously inspected via bulk_data().
void bulk_skip(std::size_t n)
{
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(n));
}
private:
// contiguous fast path: bulk copy the remaining range with std::memcpy // contiguous fast path: bulk copy the remaining range with std::memcpy
template<class T> template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/) std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
{ {
const std::size_t wanted = count * sizeof(T); const std::size_t wanted = count * sizeof(T);
const std::size_t available = remaining_count() * sizeof(char_type); #if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
const std::size_t available = static_cast<std::size_t>(std::ranges::distance(current, end)) * sizeof(char_type);
#else
const std::size_t available = static_cast<std::size_t>(std::distance(current, end)) * sizeof(char_type);
#endif
const std::size_t copied = (std::min)(wanted, available); const std::size_t copied = (std::min)(wanted, available);
if (JSON_HEDLEY_LIKELY(copied != 0)) if (JSON_HEDLEY_LIKELY(copied != 0))
{ {
@@ -620,46 +570,6 @@ typename iterator_input_adapter_factory<IteratorType, SentinelType>::adapter_typ
return factory_type::create(first, last); return factory_type::create(first, last);
} }
// The element type a container's data() points at, cv-qualifiers removed.
// Ill-formed - and therefore SFINAE-friendly - for types without data().
template<typename ContainerType>
using container_data_t = typename std::remove_cv<typename std::remove_pointer <
decltype(std::declval<const ContainerType&>().data()) >::type >::type;
// The container's own element type, cv-qualifiers removed. It is looked up on
// the bare type so it is also found when ContainerType is deduced as a
// reference by the forwarding-reference overload below.
template<typename ContainerType>
using container_value_t = typename std::remove_cv <
typename std::remove_cv<typename std::remove_reference<ContainerType>::type>::type::value_type >::type;
// Detect a container that stores its elements contiguously as single bytes
// (std::string, std::vector<char/unsigned char>, std::array<char, N>,
// std::string_view, ...). Such inputs are wrapped in a pointer-based adapter so
// they benefit from the contiguous fast paths (bulk string scanning, memcpy for
// binary formats) in every C++ standard - not only in C++20, where the standard
// library iterators model std::contiguous_iterator and are detected directly.
//
// data() and size() on their own would be duck typing: they say nothing about
// size() counting the units data() points at, and reading [data(), data() +
// size()) as bytes would be wrong for a type where it does not. Requiring the
// container's own value_type to be that same single-byte element ties the two
// together; every contiguous standard container satisfies it. Anything else
// keeps the iterator-based adapter, which is always correct - only slower.
template<typename ContainerType, typename = void>
struct is_contiguous_byte_container : std::false_type {};
template<typename ContainerType>
struct is_contiguous_byte_container < ContainerType, void_t <
container_data_t<ContainerType>,
container_value_t<ContainerType>,
decltype(std::declval<const ContainerType&>().size()) >>
: std::integral_constant < bool,
std::is_pointer<decltype(std::declval<const ContainerType&>().data())>::value&&
std::is_integral<container_data_t<ContainerType>>::value&&
sizeof(container_data_t<ContainerType>) == 1 &&
std::is_same<container_data_t<ContainerType>, container_value_t<ContainerType>>::value > {};
// Convenience shorthand from container to iterator // Convenience shorthand from container to iterator
// Enables ADL on begin(container) and end(container) // Enables ADL on begin(container) and end(container)
// Encloses the using declarations in namespace for not to leak them to outside scope // Encloses the using declarations in namespace for not to leak them to outside scope
@@ -687,32 +597,12 @@ struct container_input_adapter_factory< ContainerType,
} // namespace container_input_adapter_factory_impl } // namespace container_input_adapter_factory_impl
// General container path (iterator-based). Contiguous single-byte containers template<typename ContainerType>
// are excluded here and routed through the pointer-based overload below. typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType&& container)
template < typename ContainerType,
enable_if_t < !is_contiguous_byte_container<ContainerType>::value, int > = 0 >
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType && container)
{ {
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container)); return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
} }
// Contiguous single-byte containers (std::string, std::vector<char>, ...) are
// wrapped in a pointer-based adapter so the contiguous fast paths apply in every
// standard. The pointer keeps the container's own element type (const char* for
// std::string, const std::uint8_t* for std::vector<std::uint8_t>, ...), so the
// resulting char_type - and therefore the parsing behavior - is byte-for-byte
// identical to the iterator-based path; only the raw pointer additionally
// enables the bulk fast paths. The container outlives the adapter for the whole
// parse (temporaries live until the end of the full expression), exactly as the
// iterators it replaces did.
template < typename ContainerType,
enable_if_t < is_contiguous_byte_container<ContainerType>::value, int > = 0 >
auto input_adapter(const ContainerType& container)
-> decltype(input_adapter(container.data(), container.data() + container.size()))
{
return input_adapter(container.data(), container.data() + container.size());
}
// specialization for std::string // specialization for std::string
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>())); using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
+33 -378
View File
@@ -19,9 +19,7 @@
#include <vector> // vector #include <vector> // vector
#include <nlohmann/detail/input/input_adapters.hpp> #include <nlohmann/detail/input/input_adapters.hpp>
#include <nlohmann/detail/input/number_parse.hpp>
#include <nlohmann/detail/input/position_t.hpp> #include <nlohmann/detail/input/position_t.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp> #include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp> #include <nlohmann/detail/meta/type_traits.hpp>
@@ -127,25 +125,6 @@ constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
return false; return false;
} }
// Detect whether an input adapter exposes a contiguous byte block that the
// lexer can scan directly (see iterator_input_adapter::supports_bulk_scan).
// Adapters without the flag - file, stream, wide-string, user-defined - fall
// back to the character-at-a-time string scanner.
template<typename InputAdapterType>
using detect_supports_bulk_scan = decltype(InputAdapterType::supports_bulk_scan);
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::true_type /*detected*/)
{
return InputAdapterType::supports_bulk_scan;
}
template<typename InputAdapterType>
constexpr bool input_adapter_supports_bulk_scan(std::false_type /*detected*/)
{
return false;
}
/*! /*!
@brief lexical analysis @brief lexical analysis
@@ -167,14 +146,6 @@ class lexer : public lexer_base<BasicJsonType>
static constexpr bool lazy_token_string = static constexpr bool lazy_token_string =
input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {}); input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {});
/// whether string scanning may bulk-consume runs of ordinary characters
/// directly from a contiguous input buffer (SWAR fast path). This requires
/// the token to be reconstructible lazily (lazy_token_string), so bypassing
/// the per-character capture in get() cannot lose error diagnostics.
static constexpr bool bulk_scan =
lazy_token_string
&& input_adapter_supports_bulk_scan<InputAdapterType>(is_detected<detect_supports_bulk_scan, InputAdapterType> {});
public: public:
using token_type = typename lexer_base<BasicJsonType>::token_type; using token_type = typename lexer_base<BasicJsonType>::token_type;
@@ -295,40 +266,6 @@ class lexer : public lexer_base<BasicJsonType>
return true; return true;
} }
/// contiguous input: bulk-append the run of ordinary characters and complete
/// well-formed UTF-8 sequences starting at the current read position, leaving
/// the first byte that needs individual handling (the closing quote, an
/// escape, a control character, or an ill-formed UTF-8 byte) for get()
void scan_string_bulk(std::true_type /*bulk*/)
{
// a pending unget must be consumed through the normal path first
if (next_unget)
{
return;
}
const std::size_t remaining = ia.bulk_remaining();
if (remaining == 0)
{
return;
}
const auto* const data = reinterpret_cast<const unsigned char*>(ia.bulk_data());
const std::size_t pos = string_bulk_run(data, remaining);
if (pos == 0)
{
return;
}
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), pos);
ia.bulk_skip(pos);
// the run contains no newline (all bytes < 0x20 are treated as special),
// so only the flat character counters advance
position.chars_read_total += pos;
position.chars_read_current_line += pos;
}
/// streaming input: no bulk fast path
void scan_string_bulk(std::false_type /*bulk*/) const noexcept {}
/*! /*!
@brief scan a string literal @brief scan a string literal
@@ -354,10 +291,6 @@ class lexer : public lexer_base<BasicJsonType>
while (true) while (true)
{ {
// bulk-consume ordinary characters from contiguous input, then
// handle the next special byte through the switch below
scan_string_bulk(std::integral_constant<bool, bulk_scan> {});
// get the next character // get the next character
switch (get()) switch (get())
{ {
@@ -1076,12 +1009,6 @@ class lexer : public lexer_base<BasicJsonType>
// changed if minus sign, decimal point, or exponent is read // changed if minus sign, decimal point, or exponent is read
token_type number_type = token_type::value_unsigned; token_type number_type = token_type::value_unsigned;
// offset just past the last mantissa byte in token_buffer (i.e. the
// index of 'e'/'E', or the whole token when there is no exponent).
// convert_number() uses it to count significant digits; npos means
// "not seen an exponent yet" and is resolved at scan_number_done
std::size_t mantissa_end = std::string::npos;
// state (init): we just found out we need to scan a number // state (init): we just found out we need to scan a number
switch (current) switch (current)
{ {
@@ -1267,9 +1194,6 @@ scan_number_decimal2:
scan_number_exponent: scan_number_exponent:
// we just parsed an exponent // we just parsed an exponent
number_type = token_type::value_float; number_type = token_type::value_float;
// this label is reached only right after the 'e'/'E' was appended (from
// the zero, any1, and decimal2 states), so the mantissa ends before it
mantissa_end = token_buffer.size() - 1;
switch (get()) switch (get())
{ {
case '+': case '+':
@@ -1356,116 +1280,6 @@ scan_number_done:
// we are done scanning a number) // we are done scanning a number)
unget(); unget();
// no exponent was scanned: the mantissa spans the whole token
if (mantissa_end == std::string::npos)
{
mantissa_end = token_buffer.size();
}
return convert_number(number_type, mantissa_end);
}
/*!
@brief convert an already-validated integer token to its value
The digit sequence in [first, last) has been validated by the caller, so a
dedicated parser can avoid the locale/errno overhead of std::strtoull.
@return the token type on success; token_type::uninitialized if @a
number_type is not an integer type or the value does not fit, in
which case the caller falls back to the floating-point conversion
(matching the previous std::strtoull/std::strtoll behavior)
*/
token_type convert_integer(token_type number_type, const char* first, const char* last)
{
if (number_type == token_type::value_unsigned)
{
if (parse_integer_unsigned(first, last, value_unsigned))
{
return token_type::value_unsigned;
}
}
else if (number_type == token_type::value_integer)
{
if (parse_integer_signed(first, last, value_integer))
{
return token_type::value_integer;
}
}
return token_type::uninitialized;
}
/*!
@brief check whether Clinger's fast path can still succeed for this token
parse_float_fast() needs a significand below 2^53. A mantissa with 17 or
more significant digits is at least 10^16 and therefore always exceeds it,
so calling the fast path would walk the token one extra time only to
decline before strtod has to run anyway.
Significant digits are the mantissa's digits from the first nonzero one on;
the sign, the decimal point, leading zeros, and the exponent do not count.
The answer is derived from indices - the digits are not scanned again - so
this stays off the hot path of the number scanners.
@param[in] mantissa_end offset just past the last mantissa byte in
token_buffer
@return false if parse_float_fast() is guaranteed to decline
*/
bool mantissa_fits_clinger(std::size_t mantissa_end) const
{
// 10^16 already exceeds 2^53, so 17 digits can never fit
constexpr std::size_t limit = 17;
const std::size_t neg = (!token_buffer.empty() && token_buffer[0] == '-') ? 1u : 0u;
const std::size_t has_dot = (decimal_point_position != std::string::npos) ? 1u : 0u;
// the JSON grammar restricts the integer part to "0" or [1-9][0-9]*, so
// a leading zero can only be a lone "0", which is not significant
const std::size_t lead_zero = (token_buffer[neg] == '0') ? 1u : 0u;
JSON_ASSERT(mantissa_end >= neg + has_dot + lead_zero);
std::size_t digits = mantissa_end - neg - has_dot - lead_zero;
if (JSON_HEDLEY_LIKELY(digits < limit))
{
return true;
}
// Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. Note
// token_buffer holds the locale's decimal point, so the fraction is
// located through decimal_point_position rather than by searching '.'.
if (lead_zero != 0)
{
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
for (std::size_t i = decimal_point_position + 1;
digits >= limit && i < mantissa_end && token_buffer[i] == '0'; ++i)
{
--digits;
}
}
return digits < limit;
}
/*!
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
@param[in] mantissa_end offset just past the last mantissa byte in
token_buffer (the index of 'e'/'E', or
token_buffer.size() when there is no exponent);
used to skip Clinger's fast path when it cannot
possibly succeed - see mantissa_fits_clinger()
*/
token_type convert_number(token_type number_type, std::size_t mantissa_end)
{
// If the caller does not need the converted value (only whether the // If the caller does not need the converted value (only whether the
// input is syntactically valid; see json_sax_acceptor/accept()), an // input is syntactically valid; see json_sax_acceptor/accept()), an
// unsigned/integer token can be reported without calling // unsigned/integer token can be reported without calling
@@ -1518,37 +1332,45 @@ scan_number_done:
} }
} }
const char* const num_begin = token_buffer.data(); char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
const char* const num_end = num_begin + token_buffer.size(); errno = 0;
if (number_type != token_type::value_float) // try to parse integers first and fall back to floats
if (number_type == token_type::value_unsigned)
{ {
const token_type integer_result = convert_integer(number_type, num_begin, num_end); const auto x = std::strtoull(token_buffer.data(), &endptr, 10);
if (integer_result != token_type::uninitialized)
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
{ {
return integer_result; value_unsigned = static_cast<number_unsigned_t>(x);
if (value_unsigned == x)
{
return token_type::value_unsigned;
}
}
}
else if (number_type == token_type::value_integer)
{
const auto x = std::strtoll(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (errno != ERANGE)
{
value_integer = static_cast<number_integer_t>(x);
if (value_integer == x)
{
return token_type::value_integer;
}
} }
} }
// this code is reached if we parse a floating-point number or if an // this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. Prefer std::from_chars // integer conversion above failed
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
}
// Skipping a fast path that cannot succeed is lossless and saves a full
// extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
{
return token_type::value_float;
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr); strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before // we checked the number format before
@@ -1557,158 +1379,6 @@ scan_number_done:
return token_type::value_float; return token_type::value_float;
} }
/*!
@brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
then produces the exact diagnostic. @a current is the first digit or the
leading minus (already read); the remaining bytes are taken from the adapter.
*/
token_type scan_number_bulk_contiguous()
{
// a pending unget offsets the buffer position from current; fall back
if (next_unget)
{
return token_type::uninitialized;
}
const std::size_t rem = ia.bulk_remaining();
if (rem == 0)
{
// the first digit is the last input byte; let scan_number() finish
return token_type::uninitialized;
}
// the byte before the next unread one is current (contiguous input)
const char* const data = reinterpret_cast<const char*>(ia.bulk_data()) - 1;
const std::size_t avail = rem + 1;
// validate + classify the number extent (mirrors scan_number()'s grammar)
std::size_t i = 0;
std::size_t dot_index = std::string::npos;
token_type number_type = token_type::value_unsigned;
if (data[0] == '-')
{
number_type = token_type::value_integer;
i = 1;
if (i >= avail)
{
return token_type::uninitialized;
}
}
if (data[i] == '0')
{
++i;
}
else if (data[i] >= '1' && data[i] <= '9')
{
++i;
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
else
{
return token_type::uninitialized;
}
if (i < avail && data[i] == '.')
{
number_type = token_type::value_float;
dot_index = i;
++i;
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
// the mantissa ends here, whether or not an exponent part follows
const std::size_t mantissa_end = i;
if (i < avail && (data[i] == 'e' || data[i] == 'E'))
{
number_type = token_type::value_float;
++i;
if (i < avail && (data[i] == '+' || data[i] == '-'))
{
++i;
}
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
{
return token_type::uninitialized;
}
while (i < avail && data[i] >= '0' && data[i] <= '9')
{
++i;
}
}
const std::size_t len = i;
// reset() records where this token starts (for diagnostics), so it has
// to run before the input position advances below
reset();
// An integer token needs no token_buffer: the SAX callbacks for
// number_integer/number_unsigned take only the value, and the overflow
// diagnostic rebuilds the text from the input. Convert straight from the
// input buffer and leave token_buffer empty. (JSON_DIAGNOSTIC_POSITIONS
// derives a number's start position from get_string().size(), so there
// the token still has to be materialized.)
#if !JSON_DIAGNOSTIC_POSITIONS
if (number_type != token_type::value_float)
{
const token_type integer_result = convert_integer(number_type, data, data + len);
if (JSON_HEDLEY_LIKELY(integer_result != token_type::uninitialized))
{
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return integer_result;
}
// The value does not fit an integer, so this token converts as a
// float. Recording that here keeps convert_number() below from
// repeating the integer attempt that just failed.
number_type = token_type::value_float;
}
#endif
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
position.chars_read_current_line += (len - 1);
return convert_number(number_type, mantissa_end);
}
/// contiguous input: try the number fast path, else the byte-path scanner
token_type scan_number_dispatch(std::true_type /*bulk*/)
{
const token_type t = scan_number_bulk_contiguous();
return (t != token_type::uninitialized) ? t : scan_number();
}
/// streaming input: always use the byte-path scanner
token_type scan_number_dispatch(std::false_type /*bulk*/)
{
return scan_number();
}
/*! /*!
@param[in] literal_text the literal text to expect @param[in] literal_text the literal text to expect
@param[in] length the length of the passed literal text @param[in] length the length of the passed literal text
@@ -1812,9 +1482,6 @@ scan_number_done:
if (current == '\n') if (current == '\n')
{ {
++position.lines_read; ++position.lines_read;
// remember the column the newline was read at: chars_read_current_line
// is about to be cleared, and a matching unget() cannot reconstruct it
chars_read_before_newline = position.chars_read_current_line;
position.chars_read_current_line = 0; position.chars_read_current_line = 0;
} }
@@ -1871,20 +1538,12 @@ scan_number_done:
--position.chars_read_total; --position.chars_read_total;
// in case we "unget" a newline, we have to also decrement the lines_read // in case we "unget" a newline, we have to also decrement the lines_read
// and restore the column that get() cleared when it saw the newline;
// chars_read_current_line == 0 can only mean the last get() read one
if (position.chars_read_current_line == 0) if (position.chars_read_current_line == 0)
{ {
if (position.lines_read > 0) if (position.lines_read > 0)
{ {
--position.lines_read; --position.lines_read;
} }
// chars_read_before_newline counts the newline itself, which is the
// character being ungotten, hence the -1
position.chars_read_current_line = (chars_read_before_newline > 0)
? chars_read_before_newline - 1
: 0;
} }
else else
{ {
@@ -2151,7 +1810,7 @@ scan_number_done:
case '7': case '7':
case '8': case '8':
case '9': case '9':
return scan_number_dispatch(std::integral_constant<bool, bulk_scan> {}); return scan_number();
// end of input (the null byte is needed when parsing from // end of input (the null byte is needed when parsing from
// string literals) // string literals)
@@ -2182,10 +1841,6 @@ scan_number_done:
/// the start position of the current token /// the start position of the current token
position_t position {}; position_t position {};
/// the value chars_read_current_line had when the last newline was read, so
/// that unget() can restore the column instead of leaving it at 0
std::size_t chars_read_before_newline = 0;
/// raw input token string for error messages; only populated for streaming /// raw input token string for error messages; only populated for streaming
/// adapters (seekable adapters reconstruct it lazily via token_string_start) /// adapters (seekable adapters reconstruct it lazily via token_string_start)
std::vector<char_type> token_string {}; std::vector<char_type> token_string {};
@@ -1,302 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint64_t
#include <limits> // numeric_limits
#include <nlohmann/detail/macro_scope.hpp>
// std::from_chars lives in <charconv>, but being in C++17 mode does not
// guarantee the header exists: GCC 7 sets __cplusplus to C++17 yet ships no
// <charconv> (added in GCC 8; floating-point support in GCC 11). Guard the
// include with __has_include so such toolchains fall back to the scalar path.
#if defined(JSON_HAS_CPP_17) && defined(__has_include)
#if __has_include(<charconv>)
#include <charconv> // from_chars (only used when __cpp_lib_to_chars is defined)
#include <system_error> // errc
#endif
#endif
// This file contains the value-conversion helpers used by the lexer to turn an
// already-validated number token into a value, without the locale/errno
// overhead of std::strtoull/std::strtod. They are free functions so the lexer
// stays focused on scanning; see lexer::convert_number().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief fast integer parser for an already-validated unsigned integer
The number scanner has already checked that [first, last) is a valid JSON
integer, so this only needs to accumulate the digits and detect overflow. This
avoids the locale/errno machinery of std::strtoull, which dominates
integer-heavy inputs.
@param[in] first pointer to the first character (a digit)
@param[in] last pointer past the last character
@param[out] value the parsed value on success
@return true if the value fit into @a NumberUnsignedType; false on overflow, in
which case the caller falls back to floating-point parsing (matching the
previous std::strtoull behavior)
*/
template<typename NumberUnsignedType>
bool parse_integer_unsigned(const char* first, const char* last, NumberUnsignedType& value) noexcept
{
// accumulate in the widest unsigned type used by the previous strtoull
// path so the overflow behavior is unchanged for custom number types
std::uint64_t x = 0;
constexpr std::uint64_t cutoff = (std::numeric_limits<std::uint64_t>::max)() / 10u;
constexpr std::uint64_t cutlim = (std::numeric_limits<std::uint64_t>::max)() % 10u;
for (const char* p = first; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(x > cutoff || (x == cutoff && digit > cutlim)))
{
return false;
}
x = (x * 10u) + digit;
}
value = static_cast<NumberUnsignedType>(x);
// reject values that do not round-trip into a narrower NumberUnsignedType
return static_cast<std::uint64_t>(value) == x;
}
/*!
@brief fast integer parser for an already-validated negative integer
@param[in] first pointer to the leading '-'
@param[in] last pointer past the last character
@param[out] value the parsed (negative) value on success
@return true on success; false on overflow (caller falls back to float)
*/
template<typename NumberIntegerType>
bool parse_integer_signed(const char* first, const char* last, NumberIntegerType& value) noexcept
{
// the state machine only reaches the signed path via a leading '-'
JSON_ASSERT(first != last && *first == '-');
std::uint64_t magnitude = 0;
// |INT64_MIN| == INT64_MAX + 1; this is the largest admissible magnitude
constexpr std::uint64_t limit = static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()) + 1u;
for (const char* p = first + 1; p != last; ++p)
{
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
if (JSON_HEDLEY_UNLIKELY(magnitude > (limit - digit) / 10u))
{
return false;
}
magnitude = (magnitude * 10u) + digit;
}
const std::int64_t x = (magnitude == limit)
? (std::numeric_limits<std::int64_t>::min)()
: -static_cast<std::int64_t>(magnitude);
value = static_cast<NumberIntegerType>(x);
// reject values that do not round-trip into a narrower NumberIntegerType
return static_cast<std::int64_t>(value) == x;
}
/*!
@brief exact fast path for parsing a `double` (Clinger's algorithm)
For the common case - at most 19 significant digits, a decimal exponent in
[-22, 22], and a significand below 2^53 - the value equals significand *
10^exp computed in IEEE-754 double arithmetic, which is exact under
round-to-nearest because both operands are exactly representable. This is the
same fast path used by fast_float/simdjson; the general cases are left to
std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
// true double precision. On platforms that keep intermediates in extended
// precision (e.g. the x87 FPU on 32-bit x86, where FLT_EVAL_METHOD == 2) the
// single significand * 10^scale step is double-rounded and can be 1 ULP off,
// so decline and let the caller fall back to the correctly-rounded
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
static const std::array<double, 23> powers_of_ten =
{
{
1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11,
1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22
}
};
const char* p = first;
bool negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
negative = (*p == '-');
++p;
}
std::uint64_t significand = 0;
int num_digits = 0;
int fractional_digits = 0;
bool seen_dot = false;
bool any_digit = false;
for (; p != last; ++p)
{
const char c = *p;
if (c >= '0' && c <= '9')
{
any_digit = true;
if (JSON_HEDLEY_UNLIKELY(num_digits >= 19))
{
return false; // significand may not fit into uint64_t
}
significand = (significand * 10u) + static_cast<std::uint64_t>(c - '0');
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (static_cast<DecimalPointType>(c) == decimal_point)
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
return false;
}
seen_dot = true;
}
else if (c == 'e' || c == 'E')
{
++p;
break;
}
else
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_digit))
{
return false;
}
int exponent = 0;
if (p != last) // an exponent part remains
{
bool exp_negative = false;
if (p != last && (*p == '-' || *p == '+'))
{
exp_negative = (*p == '-');
++p;
}
bool any_exp_digit = false;
for (; p != last; ++p)
{
if (JSON_HEDLEY_UNLIKELY(*p < '0' || *p > '9'))
{
return false;
}
exponent = (exponent * 10) + (*p - '0');
any_exp_digit = true;
if (JSON_HEDLEY_UNLIKELY(exponent > 9999))
{
return false;
}
}
if (JSON_HEDLEY_UNLIKELY(!any_exp_digit))
{
return false;
}
if (exp_negative)
{
exponent = -exponent;
}
}
const int scale = exponent - fractional_digits;
if (JSON_HEDLEY_UNLIKELY(significand >= (static_cast<std::uint64_t>(1) << 53)))
{
return false; // significand not exactly representable as double
}
auto result = static_cast<double>(significand);
if (scale >= 0)
{
if (JSON_HEDLEY_UNLIKELY(scale > 22))
{
return false;
}
result *= powers_of_ten[static_cast<std::size_t>(scale)];
}
else
{
if (JSON_HEDLEY_UNLIKELY(-scale > 22))
{
return false;
}
result /= powers_of_ten[static_cast<std::size_t>(-scale)];
}
out = negative ? -result : result;
return true;
#endif
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
{
return false;
}
/*!
@brief parse a float with std::from_chars (Eisel-Lemire) when available
std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
@return true if the value was parsed exactly and fully; false to fall back
*/
template<typename FloatType>
bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept
{
// JSON_HAS_CPP_17 must gate the use as well as the <charconv> include above:
// some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even
// in C++14 mode, where <charconv> is not included.
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
const auto result = std::from_chars(first, last, out);
return result.ec == std::errc() && result.ptr == last;
#else
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(out);
return false;
#endif
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -1,293 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef> // size_t
#include <cstdint> // uint64_t
#include <cstring> // memcpy
#include <nlohmann/detail/macro_scope.hpp>
// Optional SIMD backend for bulk UTF-8 validation. This is an opt-in external
// dependency: nlohmann/json itself stays header-only and the C++11 scalar
// validator below is always available; defining JSON_USE_SIMDUTF additionally
// requires the simdutf headers on the include path and linking the simdutf
// library. See string_bulk_run().
//
// simdutf.h itself requires C++17 - it rejects older standards with an #error -
// so the backend is only compiled in from C++17 on. Below that the macro has no
// effect and the scalar validator is used; it accepts and rejects exactly the
// same input, so only throughput differs. macro_scope.hpp is included above to
// have JSON_HAS_CPP_17 available for this test.
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
#include <simdutf.h>
#endif
// This file contains the byte-level string-scanning helpers used by the lexer's
// contiguous fast path. They operate purely on raw bytes (no dependency on the
// lexer's template parameters) so they are free functions, keeping the lexer
// itself focused on the state machine; see lexer::scan_string_bulk().
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
// classify a single byte as needing individual string handling: the closing
// quote, an escape, a control character, or a non-ASCII (UTF-8)
// lead/continuation byte. Ordinary bytes (0x20..0x7F except '"' and '\\') are
// copied verbatim, which the bulk scanner does 8 bytes at a time.
inline bool is_string_special(unsigned char c) noexcept
{
return c == '\"' || c == '\\' || c < 0x20u || c >= 0x80u;
}
// SWAR helper: return a word whose high bit is set in every byte of @a v that
// is_string_special(); zero if the 8 bytes are all ordinary.
inline std::uint64_t swar_string_special(std::uint64_t v) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
const std::uint64_t has_quote = (q - ones) & ~q & high;
const std::uint64_t has_backslash = (b - ones) & ~b & high;
const std::uint64_t has_control = (v - 0x2020202020202020ull) & ~v & high; // < 0x20
const std::uint64_t has_non_ascii = v & high; // >= 0x80
return has_quote | has_backslash | has_control | has_non_ascii;
}
// return the index of the first is_string_special() byte in [data, data+n), or
// n if every byte is ordinary; scans 8 bytes at a time
inline std::size_t find_string_special(const unsigned char* data, std::size_t n) noexcept
{
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t word = 0;
std::memcpy(&word, data + i, sizeof(word));
if (swar_string_special(word) != 0)
{
// a special byte is in this word; locate it (endian-agnostic)
for (std::size_t j = 0; j < 8; ++j)
{
if (is_string_special(data[i + j]))
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
if (is_string_special(data[i]))
{
return i;
}
}
return n;
}
// classify a byte as one the serializer must NOT copy verbatim when
// ensure_ascii is requested: the closing quote, an escape, a control character
// (< 0x20), DEL (0x7F), or any non-ASCII byte (>= 0x80). Everything else -
// printable ASCII except '"' and '\\' - is emitted unchanged. Note this differs
// from is_string_special() only in that 0x7F is also a stop (it is escaped as
// \u007f under ensure_ascii).
inline bool is_ascii_copyable(unsigned char c) noexcept
{
return c >= 0x20u && c < 0x7Fu && c != '\"' && c != '\\';
}
// return the index of the first byte in [data, data+n) that is NOT
// is_ascii_copyable(), or n if every byte can be copied verbatim; scans 8 bytes
// at a time. Used by the serializer's ensure_ascii fast path.
inline std::size_t find_ascii_copyable_run(const unsigned char* data, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t v = 0;
std::memcpy(&v, data + i, sizeof(v));
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
const std::uint64_t d = v ^ 0x7F7F7F7F7F7F7F7Full; // DEL (0x7F)
const std::uint64_t stop = ((q - ones) & ~q & high) // == '"'
| ((b - ones) & ~b & high) // == '\\'
| ((d - ones) & ~d & high) // == 0x7F
| ((v - 0x2020202020202020ull) & ~v & high) // < 0x20
| (v & high); // >= 0x80
if (stop != 0)
{
for (std::size_t j = 0; j < 8; ++j)
{
if (!is_ascii_copyable(data[i + j]))
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
if (!is_ascii_copyable(data[i]))
{
return i;
}
}
return n;
}
// Validate one UTF-8 sequence at the front of [data, data+avail). Returns its
// length (2..4) only when the bytes form a *well-formed* sequence using exactly
// the same ranges as scan_string()'s per-byte switch, so the bulk path accepts
// precisely what the byte path accepts. Returns 0 for anything that is invalid,
// incomplete, or that the byte path must diagnose (the caller then defers to
// that path, keeping error messages unchanged). Lead bytes < 0x80 are handled
// by the caller and never passed here.
inline std::size_t validate_one_utf8(const unsigned char* data, std::size_t avail) noexcept
{
const unsigned char c0 = data[0];
if (c0 >= 0xC2 && c0 <= 0xDF) // U+0080..U+07FF
{
if (avail >= 2 && data[1] >= 0x80 && data[1] <= 0xBF)
{
return 2;
}
}
else if (c0 == 0xE0) // U+0800..U+0FFF
{
if (avail >= 3 && data[1] >= 0xA0 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if ((c0 >= 0xE1 && c0 <= 0xEC) || c0 == 0xEE || c0 == 0xEF) // U+1000..U+CFFF, U+E000..U+FFFF
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xED) // U+D000..U+D7FF (excludes surrogates)
{
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0x9F && data[2] >= 0x80 && data[2] <= 0xBF)
{
return 3;
}
}
else if (c0 == 0xF0) // U+10000..U+3FFFF
{
if (avail >= 4 && data[1] >= 0x90 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 >= 0xF1 && c0 <= 0xF3) // U+40000..U+FFFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
else if (c0 == 0xF4) // U+100000..U+10FFFF
{
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0x8F && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
{
return 4;
}
}
return 0; // invalid, incomplete, or must be diagnosed by the byte path
}
// Scalar (C++11) computation of the bulk run length: the number of leading
// bytes in [data, data+n) that are ordinary ASCII or complete well-formed UTF-8
// sequences, stopping before the first byte that needs individual handling (the
// closing quote, an escape, a control character, or an ill-formed/truncated
// sequence). ASCII is skipped 8 bytes at a time.
inline std::size_t scalar_string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
std::size_t pos = 0;
while (pos < n)
{
pos += find_string_special(data + pos, n - pos);
if (pos >= n || data[pos] < 0x80u)
{
break; // end of buffer, or a quote/escape/control byte
}
const std::size_t seq = validate_one_utf8(data + pos, n - pos);
if (seq == 0)
{
break; // ill-formed or truncated: let the byte path diagnose it
}
pos += seq;
}
return pos;
}
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
// Index of the first quote/escape/control byte in [data, data+n) (non-ASCII
// bytes are *not* stops here - the whole run is handed to simdutf), or n.
inline std::size_t find_string_delimiter(const unsigned char* data, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
std::uint64_t v = 0;
std::memcpy(&v, data + i, sizeof(v));
const std::uint64_t q = v ^ 0x2222222222222222ull;
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull;
const std::uint64_t hit = ((q - ones) & ~q & high)
| ((b - ones) & ~b & high)
| ((v - 0x2020202020202020ull) & ~v & high);
if (hit != 0)
{
for (std::size_t j = 0; j < 8; ++j)
{
const unsigned char c = data[i + j];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i + j;
}
}
}
}
for (; i < n; ++i)
{
const unsigned char c = data[i];
if (c == '\"' || c == '\\' || c < 0x20u)
{
return i;
}
}
return n;
}
#endif
// Backend-dispatched bulk run length. With JSON_USE_SIMDUTF the run up to the
// next delimiter is validated in one shot by simdutf; on the rare failure the
// scalar helper recomputes the exact valid prefix so the byte path still
// produces the precise diagnostic. Without it, the pure scalar path is used.
inline std::size_t string_bulk_run(const unsigned char* data, std::size_t n) noexcept
{
#if defined(JSON_USE_SIMDUTF) && defined(JSON_HAS_CPP_17)
const std::size_t run = find_string_delimiter(data, n);
if (run != 0 && simdutf::validate_utf8(reinterpret_cast<const char*>(data), run))
{
return run;
}
#endif
return scalar_string_bulk_run(data, n);
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -826,7 +826,17 @@ class binary_writer
std::vector<CharType> bjdx = {'[', '{', 'S', 'H', 'T', 'F', 'N', 'Z'}; // excluded markers in bjdata optimized type std::vector<CharType> bjdx = {'[', '{', 'S', 'H', 'T', 'F', 'N', 'Z'}; // excluded markers in bjdata optimized type
if (same_prefix && !(use_bjdata && std::find(bjdx.begin(), bjdx.end(), first_prefix) != bjdx.end())) // an optimized array of a valueless type carries no payload, so a
// reader has nothing but the declared count to bound the allocation
// by and refuses an excessive one. Write the unoptimized form for
// those, at one byte per element, so the result can be read back.
// Objects are not affected: every element is preceded by its key.
const bool valueless_type = (first_prefix == 'Z' || first_prefix == 'T' || first_prefix == 'F');
const bool excessive_valueless = valueless_type
&& j.m_data.m_value.array->size() > detail::max_valueless_container_size;
if (same_prefix && !excessive_valueless
&& !(use_bjdata && std::find(bjdx.begin(), bjdx.end(), first_prefix) != bjdx.end()))
{ {
prefix_required = false; prefix_required = false;
oa->write_character(to_char_type('$')); oa->write_character(to_char_type('$'));
File diff suppressed because it is too large Load Diff
+73 -33
View File
@@ -1343,12 +1343,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
const error_handler_t error_handler = error_handler_t::strict) const const error_handler_t error_handler = error_handler_t::strict) const
{ {
string_t result; string_t result;
detail::output_string_adapter<char, string_t> string_adapter(result); serializer s(detail::output_adapter<char, string_t>(result), indent_char, error_handler);
serializer s(string_adapter, indent_char, error_handler);
if (indent >= 0) if (indent >= 0)
{ {
s.dump(*this, true, ensure_ascii, static_cast<std::size_t>(indent)); s.dump(*this, true, ensure_ascii, static_cast<unsigned int>(indent));
} }
else else
{ {
@@ -4081,8 +4080,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
o.width(0); o.width(0);
// do the actual serialization // do the actual serialization
detail::output_stream_adapter<char> stream_adapter(o); serializer s(detail::output_adapter<char>(o), o.fill());
serializer s(stream_adapter, o.fill());
s.dump(j, pretty_print, false, static_cast<unsigned int>(indentation)); s.dump(j, pretty_print, false, static_cast<unsigned int>(indentation));
return o; return o;
} }
@@ -4503,8 +4501,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in CBOR format (iterator pair, or iterator+sentinel pair for C++20 ranges support) /// @brief create a JSON value from an input in CBOR format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4520,8 +4521,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
template<typename T> template<typename T>
@@ -4546,8 +4550,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in MessagePack format /// @brief create a JSON value from an input in MessagePack format
@@ -4561,8 +4568,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in MessagePack format (iterator pair, or iterator+sentinel pair for C++20 ranges support) /// @brief create a JSON value from an input in MessagePack format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4577,8 +4587,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
template<typename T> template<typename T>
@@ -4601,8 +4614,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in UBJSON format /// @brief create a JSON value from an input in UBJSON format
@@ -4616,8 +4632,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in UBJSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support) /// @brief create a JSON value from an input in UBJSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4632,8 +4651,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
template<typename T> template<typename T>
@@ -4656,8 +4678,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in BJData format /// @brief create a JSON value from an input in BJData format
@@ -4671,8 +4696,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in BJData format (iterator pair, or iterator+sentinel pair for C++20 ranges support) /// @brief create a JSON value from an input in BJData format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4687,8 +4715,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in BSON format /// @brief create a JSON value from an input in BSON format
@@ -4702,8 +4733,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @brief create a JSON value from an input in BSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support) /// @brief create a JSON value from an input in BSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4718,8 +4752,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
template<typename T> template<typename T>
@@ -4742,8 +4779,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict)) // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded); {
result = value_t::discarded;
}
return result;
} }
/// @} /// @}
File diff suppressed because it is too large Load Diff
-68
View File
@@ -2,9 +2,6 @@ cmake_minimum_required(VERSION 3.13...4.0)
option(JSON_Valgrind "Execute test suite with Valgrind." OFF) option(JSON_Valgrind "Execute test suite with Valgrind." OFF)
option(JSON_FastTests "Skip expensive/slow tests." OFF) option(JSON_FastTests "Skip expensive/slow tests." OFF)
option(JSON_TestSimdutf "Build the unit tests against the simdutf UTF-8 validation backend." OFF)
set(JSON_SIMDUTF_VERSION 9.1.0 CACHE STRING "The simdutf version used by JSON_TestSimdutf.")
set(JSON_32bitTest AUTO CACHE STRING "Enable the 32bit unit test (ON/OFF/AUTO/ONLY).") set(JSON_32bitTest AUTO CACHE STRING "Enable the 32bit unit test (ON/OFF/AUTO/ONLY).")
set(JSON_TestStandards "" CACHE STRING "The list of standards to test explicitly.") set(JSON_TestStandards "" CACHE STRING "The list of standards to test explicitly.")
@@ -197,71 +194,6 @@ if(test_force)
endif() endif()
message(STATUS "${msg}") message(STATUS "${msg}")
#############################################################################
# optionally validate UTF-8 with simdutf (JSON_USE_SIMDUTF)
#############################################################################
# The simdutf backend is opt-in and not vendored, so it is fetched here rather
# than being a checked-in dependency. Everything below hangs off test_main,
# whose usage requirements every test target inherits; the library target and
# the installed CMake package are deliberately left untouched.
if (JSON_TestSimdutf)
# simdutf requires C++17, both to compile itself and to be reachable from
# the library, which keeps its scalar validator below that. Find a tested
# standard that satisfies it.
set(simdutf_standard "")
foreach(cxx_standard ${test_cxx_standards})
if(NOT cxx_standard LESS 17 AND compiler_supports_cpp_${cxx_standard})
set(simdutf_standard ${cxx_standard})
break()
endif()
endforeach()
if("${simdutf_standard}" STREQUAL "")
# Building simdutf would fail outright without a C++17 compiler, and
# even with one it would go unused if no C++17-or-later standard is
# tested. Say so and fall back to the scalar validator rather than
# failing the build.
if(NOT compiler_supports_cpp_17)
set(simdutf_reason "the compiler does not support C++17")
else()
set(simdutf_reason "no tested standard is C++17 or later (testing ${msg_standards})")
endif()
message(WARNING
"JSON_TestSimdutf is enabled, but ${simdutf_reason}. simdutf requires C++17, so it "
"is not fetched and JSON_USE_SIMDUTF is not defined: the tests run against the "
"built-in scalar UTF-8 validator instead. Set JSON_TestStandards to include 17 or "
"later, or build with a compiler that supports C++17.")
else()
if (CMAKE_VERSION VERSION_LESS 3.18)
message(FATAL_ERROR "JSON_TestSimdutf requires CMake 3.18 or later (simdutf's minimum).")
endif()
include(FetchContent)
# simdutf builds its tests and tools by default, and its tests pull
# further dependencies of their own; only the library is needed here
set(SIMDUTF_TESTS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_TOOLS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_BENCHMARKS OFF CACHE BOOL "" FORCE)
set(SIMDUTF_ICONV OFF CACHE BOOL "" FORCE)
FetchContent_Declare(simdutf
URL https://github.com/simdutf/simdutf/archive/refs/tags/v${JSON_SIMDUTF_VERSION}.tar.gz
DOWNLOAD_EXTRACT_TIMESTAMP TRUE
)
FetchContent_MakeAvailable(simdutf)
target_compile_definitions(test_main PUBLIC JSON_USE_SIMDUTF)
target_link_libraries(test_main PUBLIC simdutf::simdutf)
# simdutf.h requires C++17; below that the library keeps its scalar
# validator, so any C++11/14 test targets exercise the fallback and the
# C++17-and-later ones exercise simdutf. Both must agree.
message(STATUS "UTF-8 validation delegated to simdutf ${JSON_SIMDUTF_VERSION} for C++17 and later (JSON_USE_SIMDUTF)")
endif()
endif()
# *DO* use json_test_set_test_options() above this line # *DO* use json_test_set_test_options() above this line
json_test_should_build_32bit_test(json_32bit_test json_32bit_test_only "${JSON_32bitTest}") json_test_should_build_32bit_test(json_32bit_test json_32bit_test_only "${JSON_32bitTest}")
+18 -3
View File
@@ -3288,8 +3288,10 @@ TEST_CASE("BJData")
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR1), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR1), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR1, true, false).is_discarded()); CHECK(json::from_bjdata(vR1, true, false).is_discarded());
// a dimension vector that opens another one is rejected where the
// nested '[' is read, rather than after it has been descended into
std::vector<uint8_t> const vR2 = {'[', '$', 'i', '#', '[', '#', '[', 'i', 1, ']', ']', 1}; std::vector<uint8_t> const vR2 = {'[', '$', 'i', '#', '[', '#', '[', 'i', 1, ']', ']', 1};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR2), "[json.exception.parse_error.113] parse error at byte 11: syntax error while parsing BJData size: expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x5D", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR2), "[json.exception.parse_error.113] parse error at byte 7: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR2, true, false).is_discarded()); CHECK(json::from_bjdata(vR2, true, false).is_discarded());
std::vector<uint8_t> const vR3 = {'[', '#', '[', 'i', '2', 'i', 2, ']'}; std::vector<uint8_t> const vR3 = {'[', '#', '[', 'i', '2', 'i', 2, ']'};
@@ -3297,7 +3299,7 @@ TEST_CASE("BJData")
CHECK(json::from_bjdata(vR3, true, false).is_discarded()); CHECK(json::from_bjdata(vR3, true, false).is_discarded());
std::vector<uint8_t> const vR4 = {'[', '$', 'i', '#', '[', '$', 'i', '#', '[', 'i', 1, ']', 1}; std::vector<uint8_t> const vR4 = {'[', '$', 'i', '#', '[', '$', 'i', '#', '[', 'i', 1, ']', 1};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR4), "[json.exception.parse_error.110] parse error at byte 14: syntax error while parsing BJData number: unexpected end of input", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR4), "[json.exception.parse_error.113] parse error at byte 9: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR4, true, false).is_discarded()); CHECK(json::from_bjdata(vR4, true, false).is_discarded());
std::vector<uint8_t> const vR5 = {'[', '$', 'i', '#', '[', '[', '[', ']', ']', ']'}; std::vector<uint8_t> const vR5 = {'[', '$', 'i', '#', '[', '[', '[', ']', ']', ']'};
@@ -3305,12 +3307,25 @@ TEST_CASE("BJData")
CHECK(json::from_bjdata(vR5, true, false).is_discarded()); CHECK(json::from_bjdata(vR5, true, false).is_discarded());
std::vector<uint8_t> const vR6 = {'[', '$', 'i', '#', '[', '$', 'i', '#', '[', 'i', '2', 'i', 2, ']'}; std::vector<uint8_t> const vR6 = {'[', '$', 'i', '#', '[', '$', 'i', '#', '[', 'i', '2', 'i', 2, ']'};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR6), "[json.exception.parse_error.112] parse error at byte 14: syntax error while parsing BJData size: ndarray can not be recursive", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR6), "[json.exception.parse_error.113] parse error at byte 9: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR6, true, false).is_discarded()); CHECK(json::from_bjdata(vR6, true, false).is_discarded());
std::vector<uint8_t> const vH = {'[', 'H', '[', '#', '[', '$', 'i', '#', '[', 'i', '2', 'i', 2, ']'}; std::vector<uint8_t> const vH = {'[', 'H', '[', '#', '[', '$', 'i', '#', '[', 'i', '2', 'i', 2, ']'};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vH), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vH), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vH, true, false).is_discarded()); CHECK(json::from_bjdata(vH, true, false).is_discarded());
// Every "#[" of this chain used to open another dimension vector
// and cost several stack frames before anything was rejected, so a
// long enough chain crashed the process (see #5104). The nested
// vector is refused where it is read, so the length is irrelevant.
std::vector<uint8_t> vRdeep = {'['};
for (std::size_t i = 0; i < 100000; ++i)
{
vRdeep.push_back('#');
vRdeep.push_back('[');
}
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vRdeep), "[json.exception.parse_error.113] parse error at byte 5: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vRdeep, true, false).is_discarded());
} }
SECTION("objects") SECTION("objects")
+52
View File
@@ -2035,6 +2035,58 @@ TEST_CASE("CBOR definite length equal to the indefinite-length sentinel")
} }
} }
TEST_CASE("CBOR indefinite-length strings do not recurse per chunk")
{
// Reading an indefinite-length string or byte array used to call itself
// once per chunk, so a payload of repeated 0x7F (or 0x5F) bytes exhausted
// the call stack before any of the input was rejected. The open levels are
// counted now, and the levels below prove the reader still reads the same
// values and reports the same errors at the same byte offsets.
json _;
SECTION("many open levels are reported, not crashed on")
{
const std::vector<uint8_t> input(200000, 0x7F);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), "[json.exception.parse_error.110] parse error at byte 200001: syntax error while parsing CBOR string: unexpected end of input", json::parse_error&);
CHECK(json::from_cbor(input, true, false).is_discarded());
}
SECTION("many open levels are reported, not crashed on (binary)")
{
const std::vector<uint8_t> input(200000, 0x5F);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), "[json.exception.parse_error.110] parse error at byte 200001: syntax error while parsing CBOR binary: unexpected end of input", json::parse_error&);
CHECK(json::from_cbor(input, true, false).is_discarded());
}
SECTION("chunks are still concatenated")
{
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0xFF})) == json(""));
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0x61, 0x61, 0xFF})) == json("a"));
// nested indefinite-length strings are concatenated across levels
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0x7F, 0x61, 0x61, 0xFF, 0x61, 0x62, 0xFF})) == json("ab"));
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0x7F, 0x7F, 0x61, 0x7A, 0xFF, 0xFF, 0xFF})) == json("z"));
CHECK(json::from_cbor(std::vector<uint8_t>({0xA1, 0x7F, 0x61, 0x61, 0xFF, 0x01})) == json({{"a", 1}}));
}
SECTION("chunks are still concatenated (binary)")
{
CHECK(json::from_cbor(std::vector<uint8_t>({0x5F, 0x41, 0x61, 0xFF})) == json::binary({0x61}));
CHECK(json::from_cbor(std::vector<uint8_t>({0x5F, 0x5F, 0x41, 0x61, 0xFF, 0x41, 0x62, 0xFF})) == json::binary({0x61, 0x62}));
}
SECTION("a chunk that is not a string is still rejected")
{
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7F, 0x7F, 0x00})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x00", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x5F, 0x5F, 0x00})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR binary: expected length specification (0x40-0x5B) or indefinite binary array type (0x5F); last byte: 0x00", json::parse_error&);
}
SECTION("a break marker outside an indefinite-length string is not a string")
{
// 0xFF only closes a string that was opened; on its own it is not one
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&);
}
}
TEST_CASE("CBOR roundtrips" * doctest::skip()) TEST_CASE("CBOR roundtrips" * doctest::skip())
{ {
SECTION("input from flynn") SECTION("input from flynn")
-433
View File
@@ -12,11 +12,6 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <cstdlib> // strtod
#include <sstream> // stringstream
#include <string> // string
#include <vector> // vector
namespace namespace
{ {
// shortcut to scan a string literal // shortcut to scan a string literal
@@ -229,431 +224,3 @@ TEST_CASE("lexer class")
CHECK((scan_string("/**//**//**/", true) == json::lexer::token_type::end_of_input)); CHECK((scan_string("/**//**//**/", true) == json::lexer::token_type::end_of_input));
} }
} }
TEST_CASE("lexer number fast path")
{
// The contiguous fast path (used for pointer/string input) must agree with
// the streaming byte path (used for std::istream) on token type, numeric
// value, and round-trip text for every well-formed number, and reject the
// same malformed numbers with the same message.
SECTION("contiguous vs streaming parity")
{
const std::vector<std::string> numbers =
{
"0", "-0", "1", "-1", "42", "-42", "10", "100", "1234567890",
"0.0", "-0.0", "3.14", "-3.14", "0.5", "-0.001", "123.456789",
"1e0", "1E0", "1e10", "1e-10", "1e+10", "1.5e3", "-2.5E-4",
"9223372036854775807", // INT64_MAX -> unsigned
"9223372036854775808", // INT64_MAX + 1 -> unsigned
"18446744073709551615", // UINT64_MAX -> unsigned
"18446744073709551616", // UINT64_MAX + 1 -> float
"-9223372036854775808", // INT64_MIN -> integer
"-9223372036854775809", // INT64_MIN - 1 -> float
"123456789012345678901234567890", // huge -> float
"0.30000000000000004", "2.2250738585072014e-308", "1e308",
// high-precision / wide-exponent values that exercise the
// std::from_chars (Eisel-Lemire) path beyond the Clinger subset
"1.7976931348623157e308", "1.2345678901234567e-250",
"9007199254740993", "5e-324", "1e-320"
};
for (const auto& n : numbers)
{
const std::string doc = "[" + n + "]";
// contiguous fast path
const json a = json::parse(doc);
// streaming byte path
std::stringstream ss(doc);
const json b = json::parse(ss);
CAPTURE(n);
CHECK(a == b);
CHECK(a.dump() == b.dump());
CHECK(a[0].type() == b[0].type());
}
}
SECTION("significant-digit gate for the Clinger fast path")
{
// Clinger's fast path needs a significand below 2^53, so it cannot
// succeed once the mantissa has 17 or more significant digits (the
// significand would be at least 10^16). The lexer skips the attempt
// there. That is only allowed to save work: every value must still come
// out bit-exactly, and both scanners must agree. In particular the gate
// must not fire for tokens whose leading zeros merely look like extra
// digits - "0.1234567890123456" has 16 significant digits, not 17.
const std::vector<std::string> numbers =
{
"1234567890123456", // 16 significant digits
"12345678901234567", // 17 -> attempt skipped
"123456789012345678", // 18 -> attempt skipped
"0.1234567890123456", // 16: the leading "0" is not significant
"0.12345678901234567", // 17
"0.00000000000000001", // 1, in a long token
"0.000000000000000012345678901234", // 14, in a long token
"-0.0000000000000000000001", // 1, negative
"1.0000000000000000", // 17: trailing zeros are significant here
"10000000000000000", // 17
"9007199254740992", // 2^53
"9007199254740993", // 2^53 + 1
"-65.613616999999977", // canada.json shape
"1.2345678901234567e-250", // 17 with an exponent
"1.234567890123456e-250", // 16 with an exponent
"1e10", "0.0", "-0.0", "0e0", "0.000123"
};
for (const auto& n : numbers)
{
CAPTURE(n);
const std::string doc = "[" + n + "]";
const json a = json::parse(doc); // contiguous fast path
std::stringstream ss(doc);
const json b = json::parse(ss); // streaming byte path
CHECK(a[0].type() == b[0].type());
CHECK(a == b);
if (a[0].is_number_float())
{
const double expected = std::strtod(n.c_str(), nullptr);
CHECK(a[0].get<double>() == expected);
CHECK(b[0].get<double>() == expected);
}
}
}
SECTION("token type classification")
{
CHECK((scan_string("0") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("-1") == json::lexer::token_type::value_integer));
CHECK((scan_string("1.5") == json::lexer::token_type::value_float));
CHECK((scan_string("1e5") == json::lexer::token_type::value_float));
CHECK((scan_string("18446744073709551615") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("18446744073709551616") == json::lexer::token_type::value_float));
CHECK((scan_string("-9223372036854775808") == json::lexer::token_type::value_integer));
CHECK((scan_string("-9223372036854775809") == json::lexer::token_type::value_float));
}
SECTION("malformed numbers are rejected identically")
{
for (const char* bad :
{"-", "1.", "1e", "1e+", "1.2e", "01", "-01", "1..2", "1.2.3"
})
{
CAPTURE(bad);
// the contiguous fast path must decline and let the byte path report
const std::string doc = std::string("[") + bad + "]";
CHECK_FALSE(json::accept(doc));
std::stringstream ss(doc);
CHECK_FALSE(json::accept(ss));
}
}
#if !defined(JSON_NOEXCEPTION)
// these sections parse invalid input, which aborts when exceptions are off
SECTION("exhaustive grammar parity with the streaming path")
{
// The JSON number grammar is encoded twice: once as the scan_number()
// state machine and once as the contiguous fast path. Enumerate every
// short string over the number alphabet and require the two encodings to
// agree exactly - on acceptance, on the reported error, and on the parsed
// value - so they cannot drift apart.
const std::string alphabet = "01.eE+-";
// full outcome of parsing @a doc, so a mismatch in type, value, or error
// message is caught, not just a mismatch in acceptance
const auto outcome = [](const std::string & doc, bool streaming) -> std::string
{
try
{
if (streaming)
{
std::stringstream ss(doc);
const json j = json::parse(ss);
return std::string(j[0].type_name()) + '|' + j.dump();
}
const json j = json::parse(doc);
return std::string(j[0].type_name()) + '|' + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
std::vector<std::string> mismatches;
std::vector<std::string> tokens{""};
for (std::size_t length = 1; length <= 4; ++length)
{
std::vector<std::string> next;
next.reserve(tokens.size() * alphabet.size());
for (const auto& prefix : tokens)
{
for (const char c : alphabet)
{
next.push_back(prefix + c);
}
}
tokens = next;
for (const auto& token : tokens)
{
const std::string doc = "[" + token + "]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
// 7 + 49 + 343 + 2401 tokens
CHECK(tokens.size() == 2401);
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
SECTION("error positions match the streaming path")
{
// Rejecting identically is not enough: the fast path must also report the
// error at the same position as the byte path. A number directly followed
// by a newline is the interesting case, because the byte path reaches the
// newline (which resets the column) and then ungets it.
// returns the parse_error message, or "" if the document parsed
const auto contiguous_error = [](const std::string & doc) -> std::string
{
try
{
const json j = json::parse(doc);
static_cast<void>(j);
}
catch (const json::parse_error& e)
{
return {e.what()};
}
return {};
};
const auto streaming_error = [](const std::string & doc) -> std::string
{
try
{
std::stringstream ss(doc);
const json j = json::parse(ss);
static_cast<void>(j);
}
catch (const json::parse_error& e)
{
return {e.what()};
}
return {};
};
for (const char* bad :
{"[01\n]", "[00\n]", "[-01\n]", "{1\n}", "[1\n2]", "[1.2.3\n]",
"[1 \n2]", "[\n1\n2]", "1\n2", "[01\r\n]", "[1e\n]", "[-\n]"
})
{
CAPTURE(bad);
const std::string doc = bad;
const std::string contiguous_what = contiguous_error(doc);
CHECK_FALSE(contiguous_what.empty());
CHECK(contiguous_what == streaming_error(doc));
}
// A number terminated by a newline must report the same position as the
// same number terminated by anything else: scan_number() reads the
// terminator and ungets it, so the reported column is the one reached
// after the number's last character - not the 0 that an unget() across
// the newline used to leave behind.
CHECK(contiguous_error("[01\n]") == contiguous_error("[01 ]"));
CHECK(contiguous_error("[01\n]") ==
"[json.exception.parse_error.101] parse error at line 1, column 3: "
"syntax error while parsing array - unexpected number literal; expected ']'");
// the same for a multi-character token, where the column of the last
// character (the '3' of "-2.5e3") differs from the column it starts at
CHECK(contiguous_error("null -2.5e3\nfalse") == contiguous_error("null -2.5e3 false"));
CHECK(contiguous_error("null -2.5e3\nfalse") ==
"[json.exception.parse_error.101] parse error at line 1, column 11: "
"syntax error while parsing value - unexpected number literal; expected end of input");
}
#endif
}
TEST_CASE("lexer string fast path")
{
// Build a byte string from explicit values: a hex escape in a string
// literal swallows every following hex digit, which makes sequences like
// "\xC3\xA9b" mean something other than they look like.
const auto bytes = [](std::initializer_list<int> values)
{
std::string result;
for (const int value : values)
{
result.push_back(static_cast<char>(value));
}
return result;
};
#if !defined(JSON_NOEXCEPTION)
// the full outcome of parsing @a doc: the parsed value, or the exact error
// message, so a mismatch in either is caught. Only usable with exceptions
// on: parsing invalid input aborts when they are off.
const auto outcome = [](const std::string & doc, bool streaming) -> std::string
{
try
{
if (streaming)
{
std::stringstream ss(doc);
const json j = json::parse(ss);
return j.dump();
}
const json j = json::parse(doc);
return j.dump();
}
// not just parse_error: if a bulk scanner ever let ill-formed UTF-8
// through, dump() would throw type_error.316, and that has to surface
// as a reported mismatch rather than as an uncaught exception
catch (const json::exception& e)
{
return {e.what()};
}
};
#endif
// once at the start of the string, once past the first 8-byte SWAR word, so
// the bulk scanner sees each case with and without a run behind it
const std::vector<std::size_t> offsets{0, 9};
#if !defined(JSON_NOEXCEPTION)
SECTION("exhaustive contiguous vs streaming parity")
{
// ordinary ASCII, both specials, a control byte, characters that make
// the preceding backslash a valid escape, a UTF-8 lead byte of each
// length, a continuation byte, and a byte that is never valid
const std::vector<std::string> alphabet =
{
"a", "\"", "\\", "n", "u", "0", bytes({0x01}),
bytes({0xC3}), bytes({0xA9}), bytes({0xE4}), bytes({0xF0}),
bytes({0x80}), bytes({0xFF})
};
std::vector<std::string> mismatches;
std::vector<std::string> tokens{""};
for (std::size_t length = 1; length <= 3; ++length)
{
std::vector<std::string> next;
next.reserve(tokens.size() * alphabet.size());
for (const auto& prefix : tokens)
{
for (const auto& symbol : alphabet)
{
next.push_back(prefix + symbol);
}
}
tokens = next;
for (const auto& token : tokens)
{
for (const std::size_t offset : offsets)
{
const std::string doc = "[\"" + std::string(offset, 'a') + token + "\"]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
}
// 13 + 169 + 2197 tokens, each at two offsets
CHECK(tokens.size() == 2197);
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
SECTION("special bytes at every offset of the SWAR stride")
{
// The bulk scanner consumes 8 bytes at a time and then a tail; place
// every kind of byte that ends a run at each offset across two words,
// so multibyte sequences also straddle the word boundary.
const std::vector<std::string> specials =
{
"\"", "\\", bytes({0x01}), bytes({0x1F}), bytes({0x7F}),
bytes({0xC3, 0xA9}), bytes({0xE4, 0xB8, 0xAD}), bytes({0xF0, 0x9F, 0x98, 0x80}),
bytes({0xFF}), bytes({0xC3}), bytes({0xE4, 0xB8})
};
std::vector<std::string> mismatches;
for (std::size_t offset = 0; offset <= 17; ++offset)
{
for (const auto& special : specials)
{
const std::string doc = "[\"" + std::string(offset, 'a') + special + "\"]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
}
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
#endif
// json::accept() never throws, so the ranges stay covered without exceptions
SECTION("UTF-8 ranges are accepted and rejected as documented")
{
// The bulk validator must accept exactly what the byte-at-a-time
// scanner accepts, so pin the boundaries of every range it recognizes.
// aggregate, only ever brace-initialized below; default member
// initializers would stop it being an aggregate in C++11
struct utf8_case // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init)
{
std::string sequence;
bool valid;
const char* description;
};
const std::vector<utf8_case> cases =
{
{bytes({0xC2, 0x80}), true, "U+0080, shortest two-byte"},
{bytes({0xDF, 0xBF}), true, "U+07FF, longest two-byte"},
{bytes({0xC1, 0xBF}), false, "overlong two-byte"},
{bytes({0xC2, 0x7F}), false, "two-byte with bad continuation"},
{bytes({0xE0, 0xA0, 0x80}), true, "U+0800, shortest three-byte"},
{bytes({0xE0, 0x9F, 0xBF}), false, "overlong three-byte"},
{bytes({0xED, 0x9F, 0xBF}), true, "U+D7FF, just below the surrogates"},
{bytes({0xED, 0xA0, 0x80}), false, "surrogate U+D800"},
{bytes({0xED, 0xBF, 0xBF}), false, "surrogate U+DFFF"},
{bytes({0xEE, 0x80, 0x80}), true, "U+E000, just above the surrogates"},
{bytes({0xEF, 0xBF, 0xBF}), true, "U+FFFF"},
{bytes({0xF0, 0x90, 0x80, 0x80}), true, "U+10000, shortest four-byte"},
{bytes({0xF0, 0x8F, 0xBF, 0xBF}), false, "overlong four-byte"},
{bytes({0xF4, 0x8F, 0xBF, 0xBF}), true, "U+10FFFF, highest code point"},
{bytes({0xF4, 0x90, 0x80, 0x80}), false, "above U+10FFFF"},
{bytes({0xF5, 0x80, 0x80, 0x80}), false, "lead byte out of range"},
{bytes({0x80}), false, "bare continuation byte"},
{bytes({0xFF}), false, "byte that never appears in UTF-8"},
{bytes({0xC3}), false, "truncated two-byte"},
{bytes({0xE4, 0xB8}), false, "truncated three-byte"},
{bytes({0xF0, 0x9F, 0x98}), false, "truncated four-byte"}
};
for (const auto& test_case : cases)
{
CAPTURE(test_case.description);
for (const std::size_t offset : offsets)
{
CAPTURE(offset);
const std::string doc = "[\"" + std::string(offset, 'a') + test_case.sequence + "\"]";
CHECK(json::accept(doc) == test_case.valid);
#if !defined(JSON_NOEXCEPTION)
CHECK(outcome(doc, false) == outcome(doc, true));
#endif
}
}
}
}
+1 -3
View File
@@ -98,10 +98,8 @@ void check_escaped(const char* original, const char* escaped = "", bool ensure_a
void check_escaped(const char* original, const char* escaped, const bool ensure_ascii) void check_escaped(const char* original, const char* escaped, const bool ensure_ascii)
{ {
std::stringstream ss; std::stringstream ss;
nlohmann::detail::output_stream_adapter<char> adapter(ss); json::serializer s(nlohmann::detail::output_adapter<char>(ss), ' ');
json::serializer s(adapter, ' ');
s.dump_escaped(original, ensure_ascii); s.dump_escaped(original, ensure_ascii);
s.flush(); // dump_escaped writes into the serializer's internal buffer
CHECK(ss.str() == escaped); CHECK(ss.str() == escaped);
} }
} // namespace } // namespace
-808
View File
@@ -241,209 +241,6 @@ class my_allocator : public std::allocator<T>
}; };
}; };
/////////////////////////////////////////////////////////////////////
// for #3077
/////////////////////////////////////////////////////////////////////
class FooAlloc
{};
class Foo
{
public:
explicit Foo(const FooAlloc& /* unused */ = FooAlloc()) {}
bool value = false;
};
class FooBar
{
public:
Foo foo{}; // NOLINT(readability-redundant-member-init)
};
inline void from_json(const nlohmann::json& j, FooBar& fb) // NOLINT(misc-use-internal-linkage)
{
j.at("value").get_to(fb.foo.value);
}
/////////////////////////////////////////////////////////////////////
// for #3171
/////////////////////////////////////////////////////////////////////
struct for_3171_base // NOLINT(cppcoreguidelines-special-member-functions)
{
for_3171_base(const std::string& /*unused*/ = {}) {}
virtual ~for_3171_base();
for_3171_base(const for_3171_base& other) // NOLINT(hicpp-use-equals-default,modernize-use-equals-default)
: str(other.str)
{}
for_3171_base& operator=(const for_3171_base& other)
{
if (this != &other)
{
str = other.str;
}
return *this;
}
for_3171_base(for_3171_base&& other) noexcept
: str(std::move(other.str))
{}
for_3171_base& operator=(for_3171_base&& other) noexcept
{
if (this != &other)
{
str = std::move(other.str);
}
return *this;
}
virtual void _from_json(const json& j)
{
j.at("str").get_to(str);
}
std::string str{}; // NOLINT(readability-redundant-member-init)
};
for_3171_base::~for_3171_base() = default;
struct for_3171_derived : public for_3171_base
{
for_3171_derived() = default;
~for_3171_derived() override;
explicit for_3171_derived(const std::string& /*unused*/) { }
for_3171_derived(const for_3171_derived& other) // NOLINT(hicpp-use-equals-default,modernize-use-equals-default)
: for_3171_base(other)
{}
for_3171_derived& operator=(const for_3171_derived& other)
{
if (this != &other)
{
for_3171_base::operator=(other); // Call base class assignment operator
}
return *this;
}
for_3171_derived(for_3171_derived&& other) noexcept
: for_3171_base(std::move(other))
{}
for_3171_derived& operator=(for_3171_derived&& other) noexcept
{
if (this != &other)
{
for_3171_base::operator=(std::move(other)); // Call base class move assignment operator
}
return *this;
}
};
for_3171_derived::~for_3171_derived() = default;
inline void from_json(const json& j, for_3171_base& tb) // NOLINT(misc-use-internal-linkage)
{
tb._from_json(j);
}
/////////////////////////////////////////////////////////////////////
// for #3312
/////////////////////////////////////////////////////////////////////
#ifdef JSON_HAS_CPP_20
struct for_3312
{
std::string name;
};
inline void from_json(const json& j, for_3312& obj) // NOLINT(misc-use-internal-linkage)
{
j.at("name").get_to(obj.name);
}
#endif
/////////////////////////////////////////////////////////////////////
// for #3204
/////////////////////////////////////////////////////////////////////
struct for_3204_foo
{
for_3204_foo() = default;
explicit for_3204_foo(std::string /*unused*/) {} // NOLINT(performance-unnecessary-value-param)
};
struct for_3204_bar
{
enum constructed_from_t // NOLINT(cppcoreguidelines-use-enum-class)
{
constructed_from_none = 0,
constructed_from_foo = 1,
constructed_from_json = 2
};
explicit for_3204_bar(std::function<void(for_3204_foo)> /*unused*/) noexcept // NOLINT(performance-unnecessary-value-param)
: constructed_from(constructed_from_foo) {}
explicit for_3204_bar(std::function<void(json)> /*unused*/) noexcept // NOLINT(performance-unnecessary-value-param)
: constructed_from(constructed_from_json) {}
constructed_from_t constructed_from = constructed_from_none;
};
/////////////////////////////////////////////////////////////////////
// for #3333
/////////////////////////////////////////////////////////////////////
struct for_3333 final
{
for_3333(int x_ = 0, int y_ = 0) : x(x_), y(y_) {}
template <class T>
for_3333(const T& /*unused*/)
{
CHECK(false);
}
int x = 0;
int y = 0;
};
template <>
inline for_3333::for_3333(const json& j)
: for_3333(j.value("x", 0), j.value("y", 0))
{}
/////////////////////////////////////////////////////////////////////
// for #3810
/////////////////////////////////////////////////////////////////////
struct Example_3810
{
int bla{};
Example_3810() = default;
};
NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE(Example_3810, bla) // NOLINT(misc-use-internal-linkage)
/////////////////////////////////////////////////////////////////////
// for #4740
/////////////////////////////////////////////////////////////////////
#ifdef JSON_HAS_CPP_17
struct Example_4740
{
std::optional<std::string> host = std::nullopt;
std::optional<int> port = std::nullopt;
NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT(Example_4740, host, port)
};
#endif
TEST_CASE("regression tests 2") TEST_CASE("regression tests 2")
{ {
SECTION("issue #1001 - Fix memory leak during parser callback") SECTION("issue #1001 - Fix memory leak during parser callback")
@@ -964,611 +761,6 @@ TEST_CASE("regression tests 2")
CHECK(j == k); CHECK(j == k);
} }
#if JSON_HAS_FILESYSTEM || JSON_HAS_EXPERIMENTAL_FILESYSTEM
// JSON_HAS_CPP_17 (do not remove; see note at top of file)
SECTION("issue #3070 - Version 3.10.3 breaks backward-compatibility with 3.10.2 ")
{
nlohmann::detail::std_fs::path text_path("/tmp/text.txt");
const json j(text_path);
const auto j_path = j.get<nlohmann::detail::std_fs::path>();
CHECK(j_path == text_path);
#if DOCTEST_CLANG || DOCTEST_GCC >= DOCTEST_COMPILER(8, 4, 0)
// only known to work on Clang and GCC >=8.4
CHECK_THROWS_WITH_AS(nlohmann::detail::std_fs::path(json(1)), "[json.exception.type_error.302] type must be string, but is number", json::type_error);
#endif
}
#endif
SECTION("issue #3077 - explicit constructor with default does not compile")
{
json j;
j[0]["value"] = true;
std::vector<FooBar> foo;
j.get_to(foo);
}
SECTION("issue #3108 - ordered_json doesn't support range based erase")
{
ordered_json j = {1, 2, 2, 4};
auto last = std::unique(j.begin(), j.end());
j.erase(last, j.end());
CHECK(j.dump() == "[1,2,4]");
j.erase(std::remove_if(j.begin(), j.end(), [](const ordered_json & val)
{
return val == 2;
}), j.end());
CHECK(j.dump() == "[1,4]");
}
SECTION("issue #3343 - json and ordered_json are not interchangeable")
{
json::object_t jobj({ { "product", "one" } });
ordered_json::object_t ojobj({{"product", "one"}});
auto jit = jobj.begin();
auto ojit = ojobj.begin();
CHECK(jit->first == ojit->first);
CHECK(jit->second.get<std::string>() == ojit->second.get<std::string>());
}
SECTION("issue #3171 - if class is_constructible from std::string wrong from_json overload is being selected, compilation failed")
{
const json j{{ "str", "value"}};
// failed with: error: no match for operator= (operand types are for_3171_derived and const nlohmann::basic_json<>::string_t
// {aka const std::__cxx11::basic_string<char>})
// s = *j.template get_ptr<const typename BasicJsonType::string_t*>();
auto td = j.get<for_3171_derived>();
CHECK(td.str == "value");
}
#ifdef JSON_HAS_CPP_20
SECTION("issue #3312 - Parse to custom class from unordered_json breaks on G++11.2.0 with C++20")
{
// see test for #3171
const ordered_json j = {{"name", "class"}};
for_3312 obj{};
j.get_to(obj);
CHECK(obj.name == "class");
}
#endif
#if defined(JSON_HAS_CPP_17) && JSON_USE_IMPLICIT_CONVERSIONS
SECTION("issue #3428 - Error occurred when converting nlohmann::json to std::any")
{
const json j;
const std::any a1 = j;
std::any&& a2 = j;
CHECK(a1.type() == typeid(j));
CHECK(a2.type() == typeid(j));
}
#endif
SECTION("issue #3204 - ambiguous regression")
{
const for_3204_bar bar_from_foo([](for_3204_foo) noexcept {}); // NOLINT(performance-unnecessary-value-param)
const for_3204_bar bar_from_json([](json) noexcept {}); // NOLINT(performance-unnecessary-value-param)
CHECK(bar_from_foo.constructed_from == for_3204_bar::constructed_from_foo);
CHECK(bar_from_json.constructed_from == for_3204_bar::constructed_from_json);
}
SECTION("issue #3333 - Ambiguous conversion from nlohmann::basic_json<> to custom class")
{
const json j
{
{"x", 1},
{"y", 2}
};
const for_3333 p = j;
CHECK(p.x == 1);
CHECK(p.y == 2);
}
SECTION("issue #3810 - ordered_json doesn't support construction from C array of custom type")
{
Example_3810 states[45]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
// fix "not used" warning
states[0].bla = 1;
const auto* const expected = R"([{"bla":1},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0}])";
// This works:
nlohmann::json j;
j["test"] = states;
CHECK(j["test"].dump() == expected);
// This doesn't compile:
nlohmann::ordered_json oj;
oj["test"] = states;
CHECK(oj["test"].dump() == expected);
}
#ifdef JSON_HAS_CPP_17
SECTION("issue #4740 - build issue with std::optional")
{
const auto t1 = Example_4740();
const auto j1 = nlohmann::json(t1);
CHECK(j1.dump() == "{\"host\":null,\"port\":null}");
const auto t2 = j1.get<Example_4740>();
CHECK(!t2.host.has_value());
CHECK(!t2.port.has_value());
// improve coverage
auto t3 = Example_4740();
t3.port = 80;
t3.host = "example.com";
const auto j2 = nlohmann::json(t3);
CHECK(j2.dump() == "{\"host\":\"example.com\",\"port\":80}");
const auto t4 = j2.get<Example_4740>();
CHECK(t4.host.has_value());
CHECK(t4.port.has_value());
}
#endif
#if !defined(_MSVC_LANG)
// MSVC returns garbage on invalid enum values, so this test is excluded
// there.
SECTION("issue #4762 - json exception 302 with unhelpful explanation : type must be number, but is number")
{
// In #4762, the main issue was that a json object with an invalid type
// returned "number" as type_name(), because this was the default case.
// This test makes sure we now return "invalid" instead.
json j;
j.m_data.m_type = static_cast<json::value_t>(100); // NOLINT(clang-analyzer-optin.core.EnumCastOutOfRange)
CHECK(j.type_name() == "invalid");
}
#endif
#ifdef JSON_HAS_CPP_17
SECTION("issue #4804: from_cbor incompatible with std::vector<std::byte> as binary_t")
{
const std::vector<std::uint8_t> data = {0x80};
const auto decoded = json_4804::from_cbor(data);
CHECK((decoded == json_4804::array()));
}
SECTION("discussion #4209 - custom BinaryType direct assignment and round-tripping")
{
// Test that assigning a custom BinaryType directly creates a binary value, not an array
const std::vector<std::byte> original{std::byte{1}, std::byte{2}, std::byte{3}};
const json_4804 j = original;
CHECK(j.is_binary());
CHECK(!j.is_array());
// Test round-tripping: extracting the binary value back as the custom container type
const auto extracted = j.get<std::vector<std::byte>>();
CHECK(extracted == original);
// Test that the default json alias behavior is unchanged: std::vector<uint8_t> -> array
const json default_json = std::vector<std::uint8_t> {1, 2, 3};
CHECK(default_json.is_array());
CHECK(!default_json.is_binary());
}
SECTION("discussion #4209 - custom BinaryType extraction from parsed array")
{
// Test that extracting a custom BinaryType from a parsed JSON array still works
// (not just from a binary-typed node)
const auto j = json_4804::parse("[1,2,3]");
CHECK(j.is_array());
CHECK(!j.is_binary());
// Extracting as custom BinaryType should work from arrays
const auto extracted = j.get<std::vector<std::byte>>();
CHECK(extracted.size() == 3);
CHECK(extracted[0] == std::byte{1});
CHECK(extracted[1] == std::byte{2});
CHECK(extracted[2] == std::byte{3});
}
SECTION("issue #5046 - implicit conversion of return json to std::optional no longer implicit")
{
const json jval{};
auto GetValue = [](const json & valRoot) -> std::optional<json>
{
if (valRoot.contains("default"))
{
return valRoot.at("default");
}
return std::nullopt;
};
auto result = GetValue(jval);
CHECK(!result.has_value());
}
#endif
#if JSON_HAS_RANGES == 1
SECTION("issue #4440 - assert when using std::views::filter and GCC 10")
{
auto noOpFilter = std::views::filter([](auto&&) noexcept
{
return true;
});
json j = {1, 2, 3};
auto filtered = j | noOpFilter;
CHECK(*filtered.begin() == 1);
}
#endif
#if JSON_HAS_RANGES && !defined(__MINGW32__)
SECTION("issue #4916 - constructing array from C++20 ranges view does not work")
{
std::vector<int> nums{1, 2, 37, 42, 21};
auto filteredNums = nums | std::views::filter([](int i)
{
return i > 10;
});
json const j(filteredNums);
CHECK(j.type() == json::value_t::array);
CHECK(j == json({37, 42, 21}));
}
#endif
// owning_view is not available in libstdc++ < 12
#if JSON_HAS_RANGES && !defined(__MINGW32__) && !(defined(__GLIBCXX__) && _GLIBCXX_RELEASE < 12)
SECTION("issue #4916 - constructing array from prvalue C++20 ranges view (owning_view)")
{
json const j(std::vector<int> {1, 2, 37, 42, 21} | std::views::filter([](int i)
{
return i > 10;
}));
CHECK(j.type() == json::value_t::array);
CHECK(j == json({37, 42, 21}));
}
#endif
#if JSON_HAS_RANGES && !defined(__MINGW32__)
SECTION("issue #4916 - constructing array from C++20 transform view (prvalue elements)")
{
std::vector<int> nums{1, 2, 3};
auto t = nums | std::views::transform([](int i) noexcept
{
return i * 2;
});
json const j(t);
CHECK(j.type() == json::value_t::array);
CHECK(j == json({2, 4, 6}));
}
#endif
}
TEST_CASE_TEMPLATE("issue #4798 - nlohmann::json::to_msgpack() encode float NaN as double", T, double, float) // NOLINT(readability-math-missing-parentheses, bugprone-throwing-static-initialization)
{
// With issue #4798, we encode NaN, infinity, and -infinity as float instead
// of double to allow for smaller encodings.
const json jx = std::numeric_limits<T>::quiet_NaN();
const json jy = std::numeric_limits<T>::infinity();
const json jz = -std::numeric_limits<T>::infinity();
/////////////////////////////////////////////////////////////////////////
// MessagePack
/////////////////////////////////////////////////////////////////////////
// expected MessagePack values
const std::vector<std::uint8_t> msgpack_x = {{0xCA, 0x7F, 0xC0, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_y = {{0xCA, 0x7F, 0x80, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_z = {{0xCA, 0xFF, 0x80, 0x00, 0x00}};
CHECK(json::to_msgpack(jx) == msgpack_x);
CHECK(json::to_msgpack(jy) == msgpack_y);
CHECK(json::to_msgpack(jz) == msgpack_z);
CHECK(std::isnan(json::from_msgpack(msgpack_x).get<T>()));
CHECK(json::from_msgpack(msgpack_y).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_msgpack(msgpack_z).get<T>() == -std::numeric_limits<T>::infinity());
// Make sure the other MessagePakc encodings for NaN, infinity, and
// -infinity are still supported.
const std::vector<std::uint8_t> msgpack_x_2 = {{0xCB, 0x7F, 0xF8, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_y_2 = {{0xCB, 0x7F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_z_2 = {{0xCB, 0xFF, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
CHECK(std::isnan(json::from_msgpack(msgpack_x_2).get<T>()));
CHECK(json::from_msgpack(msgpack_y_2).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_msgpack(msgpack_z_2).get<T>() == -std::numeric_limits<T>::infinity());
/////////////////////////////////////////////////////////////////////////
// CBOR
/////////////////////////////////////////////////////////////////////////
// expected CBOR values
const std::vector<std::uint8_t> cbor_x = {{0xF9, 0x7E, 0x00}};
const std::vector<std::uint8_t> cbor_y = {{0xF9, 0x7C, 0x00}};
const std::vector<std::uint8_t> cbor_z = {{0xF9, 0xfC, 0x00}};
CHECK(json::to_cbor(jx) == cbor_x);
CHECK(json::to_cbor(jy) == cbor_y);
CHECK(json::to_cbor(jz) == cbor_z);
CHECK(std::isnan(json::from_cbor(cbor_x).get<T>()));
CHECK(json::from_cbor(cbor_y).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_cbor(cbor_z).get<T>() == -std::numeric_limits<T>::infinity());
// Make sure the other CBOR encodings for NaN, infinity, and -infinity are
// still supported.
const std::vector<std::uint8_t> cbor_x_2 = {{0xFA, 0x7F, 0xC0, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_y_2 = {{0xFA, 0x7F, 0x80, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_z_2 = {{0xFA, 0xFF, 0x80, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_x_3 = {{0xFB, 0x7F, 0xF8, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_y_3 = {{0xFB, 0x7F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_z_3 = {{0xFB, 0xFF, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
CHECK(std::isnan(json::from_cbor(cbor_x_2).get<T>()));
CHECK(json::from_cbor(cbor_y_2).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_cbor(cbor_z_2).get<T>() == -std::numeric_limits<T>::infinity());
CHECK(std::isnan(json::from_cbor(cbor_x_3).get<T>()));
CHECK(json::from_cbor(cbor_y_3).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_cbor(cbor_z_3).get<T>() == -std::numeric_limits<T>::infinity());
}
TEST_CASE("regression test #5074 - portable workaround for single-element brace init")
{
json const j_obj = {{"key", "value"}};
json const j = json::array({j_obj});
CHECK(j.is_array());
CHECK(j.size() == 1);
CHECK(j[0] == j_obj);
}
#if defined(JSON_BRACE_INIT_COPY_SEMANTICS) && (JSON_BRACE_INIT_COPY_SEMANTICS == 1)
TEST_CASE("regression test #5074 - single-element brace init with JSON_BRACE_INIT_COPY_SEMANTICS")
{
// with JSON_BRACE_INIT_COPY_SEMANTICS: single-element brace init copies/moves
json const j_obj = {{"key", "value"}, {"num", 42}};
json const j_arr = {1, 2, 3};
// object: brace init copies instead of wrapping
json const j1{j_obj};
CHECK(j1.is_object());
CHECK(j1 == j_obj);
// array: brace init copies instead of wrapping
json const j2{j_arr};
CHECK(j2.is_array());
CHECK(j2.size() == 3);
CHECK(j2 == j_arr);
// primitives still work as initializer lists
json const j3{true};
CHECK(j3.is_boolean());
json const j4{42};
CHECK(j4.is_number_integer());
}
#endif
struct Example_5122
{
float b = 2;
nlohmann::ordered_map<std::string, std::string> c{}; // NOLINT(readability-redundant-member-init): needed for GCC -Weffc++
int a = 1;
NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT(Example_5122, b, c, a)
};
TEST_CASE("regression test #5122 - from_json into types holding nlohmann::ordered_map")
{
Example_5122 src;
src.c.emplace("first", "1");
src.c.emplace("second", "2");
ordered_json const j = src;
Example_5122 const dst = j.get<Example_5122>();
CHECK(dst.b == src.b);
CHECK(dst.a == src.a);
REQUIRE(dst.c.size() == src.c.size());
auto src_it = src.c.begin();
auto dst_it = dst.c.begin();
for (; src_it != src.c.end(); ++src_it, ++dst_it)
{
CHECK(dst_it->first == src_it->first);
CHECK(dst_it->second == src_it->second);
}
}
// -Wself-assign-overloaded was introduced in Clang 7. Gate the pragma on
// __has_warning so older Clang versions do not error with "unknown warning
// group". The __has_warning check has to stay inside the __clang__ branch
// because GCC does not provide it and would tokenize-error on the argument.
#if defined(__clang__) && defined(__has_warning)
#if __has_warning("-Wself-assign-overloaded")
DOCTEST_CLANG_SUPPRESS_WARNING_PUSH
DOCTEST_CLANG_SUPPRESS_WARNING("-Wself-assign-overloaded")
#endif
#endif
TEST_CASE("regression test #5122 - nlohmann::ordered_map copy-assignment is self-assignment safe")
{
nlohmann::ordered_map<std::string, std::string> m;
m.emplace("first", "1");
m.emplace("second", "2");
// Insertion order is preserved by ordered_map, so we can check it directly.
m = m;
REQUIRE(m.size() == 2);
auto it = m.begin();
CHECK(it->first == "first");
CHECK(it->second == "1");
++it;
CHECK(it->first == "second");
CHECK(it->second == "2");
}
#if defined(__clang__) && defined(__has_warning)
#if __has_warning("-Wself-assign-overloaded")
DOCTEST_CLANG_SUPPRESS_WARNING_POP
#endif
#endif
TEST_CASE("regression test #5122 - nlohmann::ordered_map move-assignment transfers contents")
{
nlohmann::ordered_map<std::string, std::string> src;
src.emplace("first", "1");
src.emplace("second", "2");
nlohmann::ordered_map<std::string, std::string> dst;
dst.emplace("stale", "x");
dst = std::move(src);
REQUIRE(dst.size() == 2);
auto it = dst.begin();
CHECK(it->first == "first");
CHECK(it->second == "1");
++it;
CHECK(it->first == "second");
CHECK(it->second == "2");
// Re-assigning into the moved-from object must leave it in a usable state.
src = nlohmann::ordered_map<std::string, std::string> {};
src.emplace("after-move", "3");
REQUIRE(src.size() == 1);
CHECK(src.begin()->first == "after-move");
}
// Stand-in for a third-party library (e.g., Eigen as of 3.4, which added
// STL-compatible begin()/end() to its vector types), living in its own
// namespace with its own to_json overload for its vector type.
namespace issue_4320_eigen
{
// "array-compatible" from the library's point of view (it has begin()/end()),
// but for which this (fake) third-party namespace provides its own to_json.
struct vector3
{
double v[3]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays,cppcoreguidelines-use-default-member-init,modernize-use-default-member-init)
vector3(double x, double y, double z) : v{x, y, z} {} // NOLINT(hicpp-member-init,cppcoreguidelines-pro-type-member-init)
double x() const
{
return v[0];
}
double y() const
{
return v[1];
}
double z() const
{
return v[2];
}
double* begin()
{
return v;
}
double* end()
{
return v + 3;
}
const double* begin() const
{
return v;
}
const double* end() const
{
return v + 3;
}
};
inline void to_json(json& j, const vector3& v) // NOLINT(misc-use-internal-linkage)
{
j = {{"x", v.x()}, {"y", v.y()}, {"z", v.z()}};
}
} // namespace issue_4320_eigen
// The user's own namespace, using the (fake) Eigen type as an implementation
// detail behind a payload type that has nothing to do with vectors/arrays.
namespace issue_4320
{
// Publicly derives from issue_4320_eigen::vector3 but does *not* define its
// own to_json - it is only ever used as a temporary to reach the base
// class's to_json via ADL.
struct vector3_wrapper : issue_4320_eigen::vector3
{
using issue_4320_eigen::vector3::vector3;
};
struct payload
{
double x, y, z;
};
inline vector3_wrapper to_eigen(const payload& p) // NOLINT(misc-use-internal-linkage)
{
return {p.x, p.y, p.z};
}
inline void to_json(json& j, const payload& p) // NOLINT(misc-use-internal-linkage)
{
// Unqualified call, passing a *derived* vector3_wrapper: relies on ADL
// finding issue_4320_eigen::to_json(json&, const vector3&) through the
// vector3 base class, via a derived-to-base conversion. Must NOT resolve
// to the library's own generic array-compatible to_json (an exact-match
// template for vector3_wrapper, since it also has begin()/end()), which
// would serialize this as [x, y, z] instead of {"x":x, "y":y, "z":z}.
to_json(j, to_eigen(p));
}
} // namespace issue_4320
TEST_CASE("issue #4320 - custom base class must not leak nlohmann::detail into ADL")
{
// Before the fix, basic_json unconditionally derived from a type living in
// nlohmann::detail (json_default_base), which made nlohmann::detail an
// associated namespace of every basic_json for ADL purposes. That leaked
// the library's internal generic-array to_json overload into unqualified
// to_json() calls made from user code, silently bypassing user-defined
// to_json overloads reached via a derived-to-base conversion.
const issue_4320::payload p{1.0, 2.0, 3.0};
json j;
to_json(j, p);
CHECK(j == json({{"x", 1.0}, {"y", 2.0}, {"z", 3.0}}));
}
TEST_CASE("issue #5338 - truncated CBOR tagged binary subtype is rejected")
{
const std::vector<std::vector<std::uint8_t>> truncated_tags =
{
{0xD8},
{0xD9, 0x00},
{0xDA, 0x00, 0x00, 0x00},
{0xDB, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}
};
for (const auto& data : truncated_tags)
{
CAPTURE(data);
for (const auto tag_handler :
{
json::cbor_tag_handler_t::ignore, json::cbor_tag_handler_t::store
})
{
CAPTURE(tag_handler);
const auto result = json::from_cbor(data, true, false, tag_handler);
CHECK(result.is_discarded());
}
}
}
TEST_CASE("issue #5402 - update(merge_objects=true) overwrites a primitive with an object")
{
json t = {{"k", 1}};
t.update(json{{"k", {{"x", 2}}}}, true);
CHECK(t == json({{"k", {{"x", 2}}}}));
json mixed = {{"keep", {{"a", 1}}}, {"replace", 1}};
mixed.update(json{{"keep", {{"b", 2}}}, {"replace", {{"x", 2}}}}, true);
CHECK(mixed == json({{"keep", {{"a", 1}, {"b", 2}}}, {"replace", {{"x", 2}}}}));
} }
DOCTEST_CLANG_SUPPRESS_WARNING_POP DOCTEST_CLANG_SUPPRESS_WARNING_POP
+899
View File
@@ -0,0 +1,899 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// cmake/test.cmake selects the C++ standard versions with which to build a
// unit test based on the presence of JSON_HAS_CPP_<VERSION> macros.
// When using macros that are only defined for particular versions of the standard
// (e.g., JSON_HAS_FILESYSTEM for C++17 and up), please mention the corresponding
// version macro in a comment close by, like this:
// JSON_HAS_CPP_<VERSION> (do not remove; see note at top of file)
#include "doctest_compatibility.h"
// for some reason including this after the json header leads to linker errors with VS 2017...
#include <locale>
#define JSON_TESTS_PRIVATE
#include <nlohmann/json.hpp>
using json = nlohmann::json;
using ordered_json = nlohmann::ordered_json;
#ifdef JSON_TEST_NO_GLOBAL_UDLS
using namespace nlohmann::literals; // NOLINT(google-build-using-namespace)
#endif
#include <cstdio>
#include <list>
#include <type_traits>
#include <utility>
#ifdef JSON_HAS_CPP_17
#include <any>
#include <variant>
#endif
#ifdef JSON_HAS_CPP_17
#if __has_include(<optional>)
#include <optional>
#elif __has_include(<experimental/optional>)
#endif
/////////////////////////////////////////////////////////////////////
// for #4804
/////////////////////////////////////////////////////////////////////
using json_4804 = nlohmann::basic_json<std::map, // ObjectType
std::vector, // ArrayType
std::string, // StringType
bool, // BooleanType
std::int64_t, // NumberIntegerType
std::uint64_t, // NumberUnsignedType
double, // NumberFloatType
std::allocator, // AllocatorType
nlohmann::adl_serializer, // JSONSerializer
std::vector<std::byte>, // BinaryType
void // CustomBaseClass
>;
#endif
#ifdef JSON_HAS_CPP_20
#if __has_include(<span>)
#include <span>
#endif
#endif
/////////////////////////////////////////////////////////////////////
// for #4825 - explicitly instantiating basic_json must compile; this
// forces instantiation of binary_writer::write_bjdata_ndarray, whose
// static_cast<string_t> was ambiguous under explicit instantiation on
// C++17. Merely compiling this translation unit is the regression test.
/////////////////////////////////////////////////////////////////////
template class nlohmann::basic_json<>;
/////////////////////////////////////////////////////////////////////
// for #4440
/////////////////////////////////////////////////////////////////////
#if JSON_HAS_RANGES == 1
#include <ranges>
#endif
// NLOHMANN_JSON_SERIALIZE_ENUM uses a static std::pair
DOCTEST_CLANG_SUPPRESS_WARNING_PUSH
DOCTEST_CLANG_SUPPRESS_WARNING("-Wexit-time-destructors")
/////////////////////////////////////////////////////////////////////
// for #3077
/////////////////////////////////////////////////////////////////////
class FooAlloc
{};
class Foo
{
public:
explicit Foo(const FooAlloc& /* unused */ = FooAlloc()) {}
bool value = false;
};
class FooBar
{
public:
Foo foo{}; // NOLINT(readability-redundant-member-init)
};
inline void from_json(const nlohmann::json& j, FooBar& fb) // NOLINT(misc-use-internal-linkage)
{
j.at("value").get_to(fb.foo.value);
}
/////////////////////////////////////////////////////////////////////
// for #3171
/////////////////////////////////////////////////////////////////////
struct for_3171_base // NOLINT(cppcoreguidelines-special-member-functions)
{
for_3171_base(const std::string& /*unused*/ = {}) {}
virtual ~for_3171_base();
for_3171_base(const for_3171_base& other) // NOLINT(hicpp-use-equals-default,modernize-use-equals-default)
: str(other.str)
{}
for_3171_base& operator=(const for_3171_base& other)
{
if (this != &other)
{
str = other.str;
}
return *this;
}
for_3171_base(for_3171_base&& other) noexcept
: str(std::move(other.str))
{}
for_3171_base& operator=(for_3171_base&& other) noexcept
{
if (this != &other)
{
str = std::move(other.str);
}
return *this;
}
virtual void _from_json(const json& j)
{
j.at("str").get_to(str);
}
std::string str{}; // NOLINT(readability-redundant-member-init)
};
for_3171_base::~for_3171_base() = default;
struct for_3171_derived : public for_3171_base
{
for_3171_derived() = default;
~for_3171_derived() override;
explicit for_3171_derived(const std::string& /*unused*/) { }
for_3171_derived(const for_3171_derived& other) // NOLINT(hicpp-use-equals-default,modernize-use-equals-default)
: for_3171_base(other)
{}
for_3171_derived& operator=(const for_3171_derived& other)
{
if (this != &other)
{
for_3171_base::operator=(other); // Call base class assignment operator
}
return *this;
}
for_3171_derived(for_3171_derived&& other) noexcept
: for_3171_base(std::move(other))
{}
for_3171_derived& operator=(for_3171_derived&& other) noexcept
{
if (this != &other)
{
for_3171_base::operator=(std::move(other)); // Call base class move assignment operator
}
return *this;
}
};
for_3171_derived::~for_3171_derived() = default;
inline void from_json(const json& j, for_3171_base& tb) // NOLINT(misc-use-internal-linkage)
{
tb._from_json(j);
}
/////////////////////////////////////////////////////////////////////
// for #3312
/////////////////////////////////////////////////////////////////////
#ifdef JSON_HAS_CPP_20
struct for_3312
{
std::string name;
};
inline void from_json(const json& j, for_3312& obj) // NOLINT(misc-use-internal-linkage)
{
j.at("name").get_to(obj.name);
}
#endif
/////////////////////////////////////////////////////////////////////
// for #3204
/////////////////////////////////////////////////////////////////////
struct for_3204_foo
{
for_3204_foo() = default;
explicit for_3204_foo(std::string /*unused*/) {} // NOLINT(performance-unnecessary-value-param)
};
struct for_3204_bar
{
enum constructed_from_t // NOLINT(cppcoreguidelines-use-enum-class)
{
constructed_from_none = 0,
constructed_from_foo = 1,
constructed_from_json = 2
};
explicit for_3204_bar(std::function<void(for_3204_foo)> /*unused*/) noexcept // NOLINT(performance-unnecessary-value-param)
: constructed_from(constructed_from_foo) {}
explicit for_3204_bar(std::function<void(json)> /*unused*/) noexcept // NOLINT(performance-unnecessary-value-param)
: constructed_from(constructed_from_json) {}
constructed_from_t constructed_from = constructed_from_none;
};
/////////////////////////////////////////////////////////////////////
// for #3333
/////////////////////////////////////////////////////////////////////
struct for_3333 final
{
for_3333(int x_ = 0, int y_ = 0) : x(x_), y(y_) {}
template <class T>
for_3333(const T& /*unused*/)
{
CHECK(false);
}
int x = 0;
int y = 0;
};
template <>
inline for_3333::for_3333(const json& j)
: for_3333(j.value("x", 0), j.value("y", 0))
{}
/////////////////////////////////////////////////////////////////////
// for #3810
/////////////////////////////////////////////////////////////////////
struct Example_3810
{
int bla{};
Example_3810() = default;
};
NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE(Example_3810, bla) // NOLINT(misc-use-internal-linkage)
/////////////////////////////////////////////////////////////////////
// for #4740
/////////////////////////////////////////////////////////////////////
#ifdef JSON_HAS_CPP_17
struct Example_4740
{
std::optional<std::string> host = std::nullopt;
std::optional<int> port = std::nullopt;
NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT(Example_4740, host, port)
};
#endif
TEST_CASE("regression tests 3")
{
#if JSON_HAS_FILESYSTEM || JSON_HAS_EXPERIMENTAL_FILESYSTEM
// JSON_HAS_CPP_17 (do not remove; see note at top of file)
SECTION("issue #3070 - Version 3.10.3 breaks backward-compatibility with 3.10.2 ")
{
nlohmann::detail::std_fs::path text_path("/tmp/text.txt");
const json j(text_path);
const auto j_path = j.get<nlohmann::detail::std_fs::path>();
CHECK(j_path == text_path);
#if DOCTEST_CLANG || DOCTEST_GCC >= DOCTEST_COMPILER(8, 4, 0)
// only known to work on Clang and GCC >=8.4
CHECK_THROWS_WITH_AS(nlohmann::detail::std_fs::path(json(1)), "[json.exception.type_error.302] type must be string, but is number", json::type_error);
#endif
}
#endif
SECTION("issue #3077 - explicit constructor with default does not compile")
{
json j;
j[0]["value"] = true;
std::vector<FooBar> foo;
j.get_to(foo);
}
SECTION("issue #3108 - ordered_json doesn't support range based erase")
{
ordered_json j = {1, 2, 2, 4};
auto last = std::unique(j.begin(), j.end());
j.erase(last, j.end());
CHECK(j.dump() == "[1,2,4]");
j.erase(std::remove_if(j.begin(), j.end(), [](const ordered_json & val)
{
return val == 2;
}), j.end());
CHECK(j.dump() == "[1,4]");
}
SECTION("issue #3343 - json and ordered_json are not interchangeable")
{
json::object_t jobj({ { "product", "one" } });
ordered_json::object_t ojobj({{"product", "one"}});
auto jit = jobj.begin();
auto ojit = ojobj.begin();
CHECK(jit->first == ojit->first);
CHECK(jit->second.get<std::string>() == ojit->second.get<std::string>());
}
SECTION("issue #3171 - if class is_constructible from std::string wrong from_json overload is being selected, compilation failed")
{
const json j{{ "str", "value"}};
// failed with: error: no match for operator= (operand types are for_3171_derived and const nlohmann::basic_json<>::string_t
// {aka const std::__cxx11::basic_string<char>})
// s = *j.template get_ptr<const typename BasicJsonType::string_t*>();
auto td = j.get<for_3171_derived>();
CHECK(td.str == "value");
}
#ifdef JSON_HAS_CPP_20
SECTION("issue #3312 - Parse to custom class from unordered_json breaks on G++11.2.0 with C++20")
{
// see test for #3171
const ordered_json j = {{"name", "class"}};
for_3312 obj{};
j.get_to(obj);
CHECK(obj.name == "class");
}
#endif
#if defined(JSON_HAS_CPP_17) && JSON_USE_IMPLICIT_CONVERSIONS
SECTION("issue #3428 - Error occurred when converting nlohmann::json to std::any")
{
const json j;
const std::any a1 = j;
std::any&& a2 = j;
CHECK(a1.type() == typeid(j));
CHECK(a2.type() == typeid(j));
}
#endif
SECTION("issue #3204 - ambiguous regression")
{
const for_3204_bar bar_from_foo([](for_3204_foo) noexcept {}); // NOLINT(performance-unnecessary-value-param)
const for_3204_bar bar_from_json([](json) noexcept {}); // NOLINT(performance-unnecessary-value-param)
CHECK(bar_from_foo.constructed_from == for_3204_bar::constructed_from_foo);
CHECK(bar_from_json.constructed_from == for_3204_bar::constructed_from_json);
}
SECTION("issue #3333 - Ambiguous conversion from nlohmann::basic_json<> to custom class")
{
const json j
{
{"x", 1},
{"y", 2}
};
const for_3333 p = j;
CHECK(p.x == 1);
CHECK(p.y == 2);
}
SECTION("issue #3810 - ordered_json doesn't support construction from C array of custom type")
{
Example_3810 states[45]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
// fix "not used" warning
states[0].bla = 1;
const auto* const expected = R"([{"bla":1},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0},{"bla":0}])";
// This works:
nlohmann::json j;
j["test"] = states;
CHECK(j["test"].dump() == expected);
// This doesn't compile:
nlohmann::ordered_json oj;
oj["test"] = states;
CHECK(oj["test"].dump() == expected);
}
#ifdef JSON_HAS_CPP_17
SECTION("issue #4740 - build issue with std::optional")
{
const auto t1 = Example_4740();
const auto j1 = nlohmann::json(t1);
CHECK(j1.dump() == "{\"host\":null,\"port\":null}");
const auto t2 = j1.get<Example_4740>();
CHECK(!t2.host.has_value());
CHECK(!t2.port.has_value());
// improve coverage
auto t3 = Example_4740();
t3.port = 80;
t3.host = "example.com";
const auto j2 = nlohmann::json(t3);
CHECK(j2.dump() == "{\"host\":\"example.com\",\"port\":80}");
const auto t4 = j2.get<Example_4740>();
CHECK(t4.host.has_value());
CHECK(t4.port.has_value());
}
#endif
#if !defined(_MSVC_LANG)
// MSVC returns garbage on invalid enum values, so this test is excluded
// there.
SECTION("issue #4762 - json exception 302 with unhelpful explanation : type must be number, but is number")
{
// In #4762, the main issue was that a json object with an invalid type
// returned "number" as type_name(), because this was the default case.
// This test makes sure we now return "invalid" instead.
json j;
j.m_data.m_type = static_cast<json::value_t>(100); // NOLINT(clang-analyzer-optin.core.EnumCastOutOfRange)
CHECK(j.type_name() == "invalid");
}
#endif
#ifdef JSON_HAS_CPP_17
SECTION("issue #4804: from_cbor incompatible with std::vector<std::byte> as binary_t")
{
const std::vector<std::uint8_t> data = {0x80};
const auto decoded = json_4804::from_cbor(data);
CHECK((decoded == json_4804::array()));
}
SECTION("discussion #4209 - custom BinaryType direct assignment and round-tripping")
{
// Test that assigning a custom BinaryType directly creates a binary value, not an array
const std::vector<std::byte> original{std::byte{1}, std::byte{2}, std::byte{3}};
const json_4804 j = original;
CHECK(j.is_binary());
CHECK(!j.is_array());
// Test round-tripping: extracting the binary value back as the custom container type
const auto extracted = j.get<std::vector<std::byte>>();
CHECK(extracted == original);
// Test that the default json alias behavior is unchanged: std::vector<uint8_t> -> array
const json default_json = std::vector<std::uint8_t> {1, 2, 3};
CHECK(default_json.is_array());
CHECK(!default_json.is_binary());
}
SECTION("discussion #4209 - custom BinaryType extraction from parsed array")
{
// Test that extracting a custom BinaryType from a parsed JSON array still works
// (not just from a binary-typed node)
const auto j = json_4804::parse("[1,2,3]");
CHECK(j.is_array());
CHECK(!j.is_binary());
// Extracting as custom BinaryType should work from arrays
const auto extracted = j.get<std::vector<std::byte>>();
CHECK(extracted.size() == 3);
CHECK(extracted[0] == std::byte{1});
CHECK(extracted[1] == std::byte{2});
CHECK(extracted[2] == std::byte{3});
}
SECTION("issue #5046 - implicit conversion of return json to std::optional no longer implicit")
{
const json jval{};
auto GetValue = [](const json & valRoot) -> std::optional<json>
{
if (valRoot.contains("default"))
{
return valRoot.at("default");
}
return std::nullopt;
};
auto result = GetValue(jval);
CHECK(!result.has_value());
}
#endif
#if JSON_HAS_RANGES == 1
SECTION("issue #4440 - assert when using std::views::filter and GCC 10")
{
auto noOpFilter = std::views::filter([](auto&&) noexcept
{
return true;
});
json j = {1, 2, 3};
auto filtered = j | noOpFilter;
CHECK(*filtered.begin() == 1);
}
#endif
#if JSON_HAS_RANGES && !defined(__MINGW32__)
SECTION("issue #4916 - constructing array from C++20 ranges view does not work")
{
std::vector<int> nums{1, 2, 37, 42, 21};
auto filteredNums = nums | std::views::filter([](int i)
{
return i > 10;
});
json const j(filteredNums);
CHECK(j.type() == json::value_t::array);
CHECK(j == json({37, 42, 21}));
}
#endif
// owning_view is not available in libstdc++ < 12
#if JSON_HAS_RANGES && !defined(__MINGW32__) && !(defined(__GLIBCXX__) && _GLIBCXX_RELEASE < 12)
SECTION("issue #4916 - constructing array from prvalue C++20 ranges view (owning_view)")
{
json const j(std::vector<int> {1, 2, 37, 42, 21} | std::views::filter([](int i)
{
return i > 10;
}));
CHECK(j.type() == json::value_t::array);
CHECK(j == json({37, 42, 21}));
}
#endif
#if JSON_HAS_RANGES && !defined(__MINGW32__)
SECTION("issue #4916 - constructing array from C++20 transform view (prvalue elements)")
{
std::vector<int> nums{1, 2, 3};
auto t = nums | std::views::transform([](int i) noexcept
{
return i * 2;
});
json const j(t);
CHECK(j.type() == json::value_t::array);
CHECK(j == json({2, 4, 6}));
}
#endif
}
TEST_CASE_TEMPLATE("issue #4798 - nlohmann::json::to_msgpack() encode float NaN as double", T, double, float) // NOLINT(readability-math-missing-parentheses, bugprone-throwing-static-initialization)
{
// With issue #4798, we encode NaN, infinity, and -infinity as float instead
// of double to allow for smaller encodings.
const json jx = std::numeric_limits<T>::quiet_NaN();
const json jy = std::numeric_limits<T>::infinity();
const json jz = -std::numeric_limits<T>::infinity();
/////////////////////////////////////////////////////////////////////////
// MessagePack
/////////////////////////////////////////////////////////////////////////
// expected MessagePack values
const std::vector<std::uint8_t> msgpack_x = {{0xCA, 0x7F, 0xC0, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_y = {{0xCA, 0x7F, 0x80, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_z = {{0xCA, 0xFF, 0x80, 0x00, 0x00}};
CHECK(json::to_msgpack(jx) == msgpack_x);
CHECK(json::to_msgpack(jy) == msgpack_y);
CHECK(json::to_msgpack(jz) == msgpack_z);
CHECK(std::isnan(json::from_msgpack(msgpack_x).get<T>()));
CHECK(json::from_msgpack(msgpack_y).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_msgpack(msgpack_z).get<T>() == -std::numeric_limits<T>::infinity());
// Make sure the other MessagePakc encodings for NaN, infinity, and
// -infinity are still supported.
const std::vector<std::uint8_t> msgpack_x_2 = {{0xCB, 0x7F, 0xF8, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_y_2 = {{0xCB, 0x7F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> msgpack_z_2 = {{0xCB, 0xFF, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
CHECK(std::isnan(json::from_msgpack(msgpack_x_2).get<T>()));
CHECK(json::from_msgpack(msgpack_y_2).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_msgpack(msgpack_z_2).get<T>() == -std::numeric_limits<T>::infinity());
/////////////////////////////////////////////////////////////////////////
// CBOR
/////////////////////////////////////////////////////////////////////////
// expected CBOR values
const std::vector<std::uint8_t> cbor_x = {{0xF9, 0x7E, 0x00}};
const std::vector<std::uint8_t> cbor_y = {{0xF9, 0x7C, 0x00}};
const std::vector<std::uint8_t> cbor_z = {{0xF9, 0xfC, 0x00}};
CHECK(json::to_cbor(jx) == cbor_x);
CHECK(json::to_cbor(jy) == cbor_y);
CHECK(json::to_cbor(jz) == cbor_z);
CHECK(std::isnan(json::from_cbor(cbor_x).get<T>()));
CHECK(json::from_cbor(cbor_y).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_cbor(cbor_z).get<T>() == -std::numeric_limits<T>::infinity());
// Make sure the other CBOR encodings for NaN, infinity, and -infinity are
// still supported.
const std::vector<std::uint8_t> cbor_x_2 = {{0xFA, 0x7F, 0xC0, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_y_2 = {{0xFA, 0x7F, 0x80, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_z_2 = {{0xFA, 0xFF, 0x80, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_x_3 = {{0xFB, 0x7F, 0xF8, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_y_3 = {{0xFB, 0x7F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
const std::vector<std::uint8_t> cbor_z_3 = {{0xFB, 0xFF, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}};
CHECK(std::isnan(json::from_cbor(cbor_x_2).get<T>()));
CHECK(json::from_cbor(cbor_y_2).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_cbor(cbor_z_2).get<T>() == -std::numeric_limits<T>::infinity());
CHECK(std::isnan(json::from_cbor(cbor_x_3).get<T>()));
CHECK(json::from_cbor(cbor_y_3).get<T>() == std::numeric_limits<T>::infinity());
CHECK(json::from_cbor(cbor_z_3).get<T>() == -std::numeric_limits<T>::infinity());
}
TEST_CASE("regression test #5074 - portable workaround for single-element brace init")
{
json const j_obj = {{"key", "value"}};
json const j = json::array({j_obj});
CHECK(j.is_array());
CHECK(j.size() == 1);
CHECK(j[0] == j_obj);
}
#if defined(JSON_BRACE_INIT_COPY_SEMANTICS) && (JSON_BRACE_INIT_COPY_SEMANTICS == 1)
TEST_CASE("regression test #5074 - single-element brace init with JSON_BRACE_INIT_COPY_SEMANTICS")
{
// with JSON_BRACE_INIT_COPY_SEMANTICS: single-element brace init copies/moves
json const j_obj = {{"key", "value"}, {"num", 42}};
json const j_arr = {1, 2, 3};
// object: brace init copies instead of wrapping
json const j1{j_obj};
CHECK(j1.is_object());
CHECK(j1 == j_obj);
// array: brace init copies instead of wrapping
json const j2{j_arr};
CHECK(j2.is_array());
CHECK(j2.size() == 3);
CHECK(j2 == j_arr);
// primitives still work as initializer lists
json const j3{true};
CHECK(j3.is_boolean());
json const j4{42};
CHECK(j4.is_number_integer());
}
#endif
struct Example_5122
{
float b = 2;
nlohmann::ordered_map<std::string, std::string> c{}; // NOLINT(readability-redundant-member-init): needed for GCC -Weffc++
int a = 1;
NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT(Example_5122, b, c, a)
};
TEST_CASE("regression test #5122 - from_json into types holding nlohmann::ordered_map")
{
Example_5122 src;
src.c.emplace("first", "1");
src.c.emplace("second", "2");
ordered_json const j = src;
Example_5122 const dst = j.get<Example_5122>();
CHECK(dst.b == src.b);
CHECK(dst.a == src.a);
REQUIRE(dst.c.size() == src.c.size());
auto src_it = src.c.begin();
auto dst_it = dst.c.begin();
for (; src_it != src.c.end(); ++src_it, ++dst_it)
{
CHECK(dst_it->first == src_it->first);
CHECK(dst_it->second == src_it->second);
}
}
// -Wself-assign-overloaded was introduced in Clang 7. Gate the pragma on
// __has_warning so older Clang versions do not error with "unknown warning
// group". The __has_warning check has to stay inside the __clang__ branch
// because GCC does not provide it and would tokenize-error on the argument.
#if defined(__clang__) && defined(__has_warning)
#if __has_warning("-Wself-assign-overloaded")
DOCTEST_CLANG_SUPPRESS_WARNING_PUSH
DOCTEST_CLANG_SUPPRESS_WARNING("-Wself-assign-overloaded")
#endif
#endif
TEST_CASE("regression test #5122 - nlohmann::ordered_map copy-assignment is self-assignment safe")
{
nlohmann::ordered_map<std::string, std::string> m;
m.emplace("first", "1");
m.emplace("second", "2");
// Insertion order is preserved by ordered_map, so we can check it directly.
m = m;
REQUIRE(m.size() == 2);
auto it = m.begin();
CHECK(it->first == "first");
CHECK(it->second == "1");
++it;
CHECK(it->first == "second");
CHECK(it->second == "2");
}
#if defined(__clang__) && defined(__has_warning)
#if __has_warning("-Wself-assign-overloaded")
DOCTEST_CLANG_SUPPRESS_WARNING_POP
#endif
#endif
TEST_CASE("regression test #5122 - nlohmann::ordered_map move-assignment transfers contents")
{
nlohmann::ordered_map<std::string, std::string> src;
src.emplace("first", "1");
src.emplace("second", "2");
nlohmann::ordered_map<std::string, std::string> dst;
dst.emplace("stale", "x");
dst = std::move(src);
REQUIRE(dst.size() == 2);
auto it = dst.begin();
CHECK(it->first == "first");
CHECK(it->second == "1");
++it;
CHECK(it->first == "second");
CHECK(it->second == "2");
// Re-assigning into the moved-from object must leave it in a usable state.
src = nlohmann::ordered_map<std::string, std::string> {};
src.emplace("after-move", "3");
REQUIRE(src.size() == 1);
CHECK(src.begin()->first == "after-move");
}
// Stand-in for a third-party library (e.g., Eigen as of 3.4, which added
// STL-compatible begin()/end() to its vector types), living in its own
// namespace with its own to_json overload for its vector type.
namespace issue_4320_eigen
{
// "array-compatible" from the library's point of view (it has begin()/end()),
// but for which this (fake) third-party namespace provides its own to_json.
struct vector3
{
double v[3]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays,cppcoreguidelines-use-default-member-init,modernize-use-default-member-init)
vector3(double x, double y, double z) : v{x, y, z} {} // NOLINT(hicpp-member-init,cppcoreguidelines-pro-type-member-init)
double x() const
{
return v[0];
}
double y() const
{
return v[1];
}
double z() const
{
return v[2];
}
double* begin()
{
return v;
}
double* end()
{
return v + 3;
}
const double* begin() const
{
return v;
}
const double* end() const
{
return v + 3;
}
};
inline void to_json(json& j, const vector3& v) // NOLINT(misc-use-internal-linkage)
{
j = {{"x", v.x()}, {"y", v.y()}, {"z", v.z()}};
}
} // namespace issue_4320_eigen
// The user's own namespace, using the (fake) Eigen type as an implementation
// detail behind a payload type that has nothing to do with vectors/arrays.
namespace issue_4320
{
// Publicly derives from issue_4320_eigen::vector3 but does *not* define its
// own to_json - it is only ever used as a temporary to reach the base
// class's to_json via ADL.
struct vector3_wrapper : issue_4320_eigen::vector3
{
using issue_4320_eigen::vector3::vector3;
};
struct payload
{
double x, y, z;
};
inline vector3_wrapper to_eigen(const payload& p) // NOLINT(misc-use-internal-linkage)
{
return {p.x, p.y, p.z};
}
inline void to_json(json& j, const payload& p) // NOLINT(misc-use-internal-linkage)
{
// Unqualified call, passing a *derived* vector3_wrapper: relies on ADL
// finding issue_4320_eigen::to_json(json&, const vector3&) through the
// vector3 base class, via a derived-to-base conversion. Must NOT resolve
// to the library's own generic array-compatible to_json (an exact-match
// template for vector3_wrapper, since it also has begin()/end()), which
// would serialize this as [x, y, z] instead of {"x":x, "y":y, "z":z}.
to_json(j, to_eigen(p));
}
} // namespace issue_4320
TEST_CASE("issue #4320 - custom base class must not leak nlohmann::detail into ADL")
{
// Before the fix, basic_json unconditionally derived from a type living in
// nlohmann::detail (json_default_base), which made nlohmann::detail an
// associated namespace of every basic_json for ADL purposes. That leaked
// the library's internal generic-array to_json overload into unqualified
// to_json() calls made from user code, silently bypassing user-defined
// to_json overloads reached via a derived-to-base conversion.
const issue_4320::payload p{1.0, 2.0, 3.0};
json j;
to_json(j, p);
CHECK(j == json({{"x", 1.0}, {"y", 2.0}, {"z", 3.0}}));
}
TEST_CASE("issue #5338 - truncated CBOR tagged binary subtype is rejected")
{
const std::vector<std::vector<std::uint8_t>> truncated_tags =
{
{0xD8},
{0xD9, 0x00},
{0xDA, 0x00, 0x00, 0x00},
{0xDB, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}
};
for (const auto& data : truncated_tags)
{
CAPTURE(data);
for (const auto tag_handler :
{
json::cbor_tag_handler_t::ignore, json::cbor_tag_handler_t::store
})
{
CAPTURE(tag_handler);
const auto result = json::from_cbor(data, true, false, tag_handler);
CHECK(result.is_discarded());
}
}
}
TEST_CASE("issue #5402 - update(merge_objects=true) overwrites a primitive with an object")
{
json t = {{"k", 1}};
t.update(json{{"k", {{"x", 2}}}}, true);
CHECK(t == json({{"k", {{"x", 2}}}}));
json mixed = {{"keep", {{"a", 1}}}, {"replace", 1}};
mixed.update(json{{"keep", {{"b", 2}}}, {"replace", {{"x", 2}}}}, true);
CHECK(mixed == json({{"keep", {{"a", 1}, {"b", 2}}}, {"replace", {{"x", 2}}}}));
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
-229
View File
@@ -387,232 +387,3 @@ TEST_CASE("dump for basic_json with long double number_float_t")
check_same(100.0L, 100.0); check_same(100.0L, 100.0);
} }
} }
TEST_CASE("serialization of strings (bulk fast path)")
{
// These cases exercise the SWAR bulk-copy fast path in dump_escaped and the
// internal write buffer: long runs, escapes interrupting runs, 0x7F/DEL,
// multibyte UTF-8 under both ensure_ascii settings, and payloads larger than
// the write buffer.
SECTION("long unescaped ASCII exceeds the write buffer")
{
const std::string big(3000, 'a');
const json j = big;
CHECK(j.dump() == '"' + big + '"');
CHECK(j.dump(-1, ' ', true) == '"' + big + '"');
// round-trips
CHECK(json::parse(j.dump()) == j);
}
SECTION("runs interrupted by escapes")
{
const json j = std::string(500, 'x') + "\n\"\\" + std::string(500, 'y');
const std::string out = j.dump();
CHECK(out == '"' + std::string(500, 'x') + "\\n\\\"\\\\" + std::string(500, 'y') + '"');
CHECK(json::parse(out) == j);
}
SECTION("DEL (0x7F) depends on ensure_ascii")
{
const json j = std::string("a\x7f" "b");
CHECK(j.dump(-1, ' ', false) == "\"a\x7f" "b\""); // copied verbatim
CHECK(j.dump(-1, ' ', true) == "\"a\\u007fb\""); // escaped
}
SECTION("multibyte UTF-8 under both ensure_ascii settings")
{
const json j = std::string("A\xc3\xa9\xe4\xbd\xa0\xf0\x9f\x98\x80Z"); // A é 你 😀 Z
// not escaping non-ASCII: bytes are copied through the bulk validator
CHECK(j.dump(-1, ' ', false) == "\"A\xc3\xa9\xe4\xbd\xa0\xf0\x9f\x98\x80Z\"");
// ensure_ascii: escaped (with a surrogate pair for the emoji)
CHECK(j.dump(-1, ' ', true) == "\"A\\u00e9\\u4f60\\ud83d\\ude00Z\"");
CHECK(json::parse(j.dump(-1, ' ', true)) == j);
}
SECTION("many small structural writes exceed the write buffer")
{
json arr = json::array();
for (int i = 0; i < 2000; ++i)
{
arr.push_back(i);
}
const std::string out = arr.dump();
CHECK(out.front() == '[');
CHECK(out.back() == ']');
CHECK(json::parse(out) == arr);
json obj = json::object();
for (int i = 0; i < 500; ++i)
{
obj["key" + std::to_string(i)] = i;
}
CHECK(json::parse(obj.dump()) == obj);
CHECK(json::parse(obj.dump(2)) == obj);
// an array of many empty strings emits a long run of single-character
// writes ('"', '"', ',') at shallow nesting depth, so the write buffer
// fills and flushes mid-run without the deep recursion that would
// overflow the stack on some debug builds
json many_empty = json::array();
for (int i = 0; i < 500; ++i)
{
many_empty.push_back("");
}
const std::string out2 = many_empty.dump();
CHECK(out2.size() > 1024); // spans multiple write-buffer flushes
CHECK(out2.front() == '[');
CHECK(out2.back() == ']');
CHECK(json::parse(out2) == many_empty);
}
SECTION("invalid UTF-8 handling is unaffected by the fast path")
{
const json j = std::string("valid\xff" "more");
CHECK_THROWS_WITH_AS(j.dump(), "[json.exception.type_error.316] invalid UTF-8 byte at index 5: 0xFF", json::type_error&);
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"valid\xef\xbf\xbd" "more\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"valid\\ufffdmore\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"validmore\"");
}
}
TEST_CASE("indentation is written straight into the write buffer")
{
// put_indent() memsets the indentation into the write buffer instead of
// copying it out of a pre-grown indentation string. These cases cover an
// indentation wider than the buffer, a non-space indentation character, and
// nesting deep enough that the accumulated indentation spans several
// buffer-fulls - the situations the old grow-a-string approach got wrong.
SECTION("indent_step wider than the write buffer")
{
const json j = {{"a", 1}};
// 2000 > the 1024-byte write buffer, and > the 512 the indentation
// string used to start at
CHECK(j.dump(2000) == "{\n" + std::string(2000, ' ') + "\"a\": 1\n}");
// several whole buffer-fulls, so the buffer is refilled once and then
// flushed repeatedly
CHECK(j.dump(5000) == "{\n" + std::string(5000, ' ') + "\"a\": 1\n}");
CHECK(j.dump(5000, '\t') == "{\n" + std::string(5000, '\t') + "\"a\": 1\n}");
// an exact multiple of the buffer size
CHECK(j.dump(4096) == "{\n" + std::string(4096, ' ') + "\"a\": 1\n}");
}
SECTION("a non-space indentation character is used throughout")
{
const json j = {{"a", 1}};
// 600 is past the point where the indentation used to be grown, which
// is where a hard-coded space would have shown up
CHECK(j.dump(600, '\t') == "{\n" + std::string(600, '\t') + "\"a\": 1\n}");
CHECK(j.dump(3, '.') == "{\n...\"a\": 1\n}");
}
SECTION("accumulated indentation spans several buffer-fulls")
{
// five levels deep at 400 per level: the innermost value is indented by
// 2000 characters, reached in steps that each straddle the buffer end
json j = json::array({1});
for (int i = 0; i < 4; ++i)
{
j = json::array({j});
}
const std::string out = j.dump(400);
CHECK(out.find(std::string("\n") + std::string(2000, ' ') + "1\n") != std::string::npos);
CHECK(json::parse(out) == j);
}
SECTION("indentation is unchanged for ordinary widths")
{
const json j = {{"a", {1, 2}}, {"b", nullptr}};
CHECK(j.dump(2) == "{\n \"a\": [\n 1,\n 2\n ],\n \"b\": null\n}");
CHECK(j.dump(0) == "{\n\"a\": [\n1,\n2\n],\n\"b\": null\n}");
}
}
TEST_CASE("serialization of deeply nested values")
{
// dump() descends into a bounded number of levels and writes out whatever
// is nested deeper than that without the call stack; see
// https://github.com/nlohmann/json/issues/5387
SECTION("nested deeper than the call stack could follow")
{
// parsing is iterative, so building these costs little
const std::size_t depth = 100000;
const std::string array_text = std::string(depth, '[') + '0' + std::string(depth, ']');
CHECK(json::parse(array_text).dump() == array_text);
std::string object_text;
object_text.reserve((6 * depth) + 1);
for (std::size_t i = 0; i < depth; ++i)
{
object_text += "{\"a\":";
}
object_text += '1';
object_text.append(depth, '}');
CHECK(json::parse(object_text).dump() == object_text);
}
SECTION("depths around the bound of the descent")
{
// Cover every depth around the bound, so that the two ways of writing a
// value are known to meet cleanly - wherever the bound is set.
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d);
const std::string array_text = std::string(d, '[') + '7' + std::string(d, ']');
CHECK(json::parse(array_text).dump() == array_text);
std::string object_text;
for (std::size_t i = 0; i < d; ++i)
{
object_text += "{\"k\":";
}
object_text += '7';
object_text.append(d, '}');
CHECK(json::parse(object_text).dump() == object_text);
}
}
SECTION("pretty-printing across the bound")
{
for (std::size_t d = 120; d <= 140; ++d)
{
CAPTURE(d);
const json j = json::parse(std::string(d, '[') + '7' + std::string(d, ']'));
std::string expected;
for (std::size_t i = 0; i < d; ++i)
{
expected += std::string(2 * i, ' ') + "[\n";
}
expected += std::string(2 * d, ' ') + '7';
for (std::size_t i = d; i > 0; --i)
{
expected += '\n' + std::string(2 * (i - 1), ' ') + ']';
}
CHECK(j.dump(2) == expected);
}
}
SECTION("an empty container below the bound")
{
// an empty container is written out in full and never descended into,
// so it must not gain a newline when it is reached iteratively
for (std::size_t d = 125; d <= 135; ++d)
{
CAPTURE(d);
const std::string compact = std::string(d, '[') + "[]" + std::string(d, ']');
CHECK(json::parse(compact).dump() == compact);
const std::string with_object = std::string(d, '[') + "{}" + std::string(d, ']');
CHECK(json::parse(with_object).dump() == with_object);
}
}
}
+61
View File
@@ -2149,6 +2149,67 @@ TEST_CASE("UBJSON")
} }
} }
TEST_CASE("UBJSON optimized arrays of a valueless type are bounded")
{
// An element of type 'Z', 'T' or 'F' is encoded by its marker alone, so an
// optimized array of one of those has no payload and the declared count is
// the only thing deciding how much is allocated. Ten bytes used to produce
// billions of values (#2793); every other type costs at least one byte per
// element and is bounded by the end of the input.
json _;
SECTION("an excessive count is rejected")
{
// 'l' is a big-endian int32: 0x7FFFFFFF elements, about 34 GB of value
for (const auto marker :
{'Z', 'T', 'F'
})
{
const std::vector<uint8_t> input = {'[', '$', static_cast<uint8_t>(marker), '#', 'l', 0x7F, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(input), "[json.exception.out_of_range.408] syntax error while parsing UBJSON size: excessive array size", json::out_of_range&);
CHECK(json::from_ubjson(input, true, false).is_discarded());
}
}
SECTION("ordinary counts are unaffected")
{
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'Z', '#', 'i', 3})) == json({nullptr, nullptr, nullptr}));
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'T', '#', 'i', 2})) == json({true, true}));
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'F', '#', 'i', 2})) == json({false, false}));
// 'N' is a no-op rather than a value, and still yields an empty array
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'N', '#', 'i', 2})) == json::array());
}
SECTION("a type with a payload is unaffected")
{
// A count past the limit is not rejected for 'U', which costs a byte
// per element and is bounded by the end of the input instead. The
// count is kept just past the limit rather than made huge, because a
// count that also exceeds the array's max_size() is reported as
// out_of_range before the input runs out, and max_size() depends on
// the width of std::size_t.
const std::vector<uint8_t> input = {'[', '$', 'U', '#', 'l', 0x00, 0x10, 0x00, 0x01};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(input), "[json.exception.parse_error.110] parse error at byte 10: syntax error while parsing UBJSON number: unexpected end of input", json::parse_error&);
CHECK(json::from_ubjson(input, true, false).is_discarded());
}
SECTION("the writer stays within what the reader accepts")
{
// below the limit the optimized form is used and is tiny; above it the
// writer falls back so that the result can still be read back
json const at_limit(1048576, nullptr);
const auto v_at_limit = json::to_ubjson(at_limit, true, true);
CHECK(v_at_limit.size() == 9);
CHECK(v_at_limit.at(1) == '$');
CHECK(json::from_ubjson(v_at_limit) == at_limit);
json const above_limit(1048577, nullptr);
const auto v_above_limit = json::to_ubjson(above_limit, true, true);
CHECK(v_above_limit.at(1) != '$');
CHECK(json::from_ubjson(v_above_limit) == above_limit);
}
}
TEST_CASE("Universal Binary JSON Specification Examples 1") TEST_CASE("Universal Binary JSON Specification Examples 1")
{ {
SECTION("Null Value") SECTION("Null Value")
-239
View File
@@ -18,12 +18,7 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint8_t
#include <list> #include <list>
#include <string> // string
#include <vector> // vector
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20) #if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
#include <iterator> #include <iterator>
@@ -217,66 +212,6 @@ TEST_CASE("Parse with heterogeneous iterator and sentinel types")
CHECK(j2.at(0) == 1); CHECK(j2.at(0) == 1);
} }
// A type whose data() hands out raw bytes but whose size() counts something
// else - here fixed-size records. Reading [data(), data() + size()) as bytes
// would silently truncate the input, so data() and size() alone must not be
// taken as evidence of contiguous byte storage.
struct record_buffer
{
using value_type = std::array<char, 4>;
std::string bytes;
const char* data() const noexcept
{
return bytes.data();
}
std::size_t size() const noexcept
{
return bytes.size() / sizeof(value_type);
}
const char* begin() const noexcept
{
return bytes.data();
}
const char* end() const noexcept
{
return bytes.data() + bytes.size();
}
};
TEST_CASE("Contiguous byte containers take the pointer adapter")
{
// Containers with contiguous single-byte storage are routed through the
// pointer-based adapter so the bulk fast paths apply in every standard, not
// only in C++20 where the library iterators model std::contiguous_iterator.
CHECK(nlohmann::detail::is_contiguous_byte_container<std::string>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<std::vector<char>>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<std::vector<std::uint8_t>>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<std::array<char, 4>>::value);
// input_adapter() takes its container by forwarding reference, so the trait
// is also asked about reference types
CHECK(nlohmann::detail::is_contiguous_byte_container<std::string&>::value);
CHECK(nlohmann::detail::is_contiguous_byte_container<const std::string&>::value);
// everything else keeps the iterator-based adapter
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<std::list<char>>::value);
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<std::vector<int>>::value);
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<const char*>::value);
// including a type that has data() and size() but whose size() does not
// count the units data() points at: its value_type says so
CHECK_FALSE(nlohmann::detail::is_contiguous_byte_container<record_buffer>::value);
// and such a container still parses through its iterators, in full - taking
// it for a byte container would stop after data() + size() bytes
const record_buffer buffer{"[1,2,3,4,5]"};
CHECK(buffer.data() == buffer.bytes.data());
CHECK(buffer.size() * sizeof(record_buffer::value_type) < buffer.bytes.size());
CHECK(json::parse(buffer) == json({1, 2, 3, 4, 5}));
}
#if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20) #if defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// JSON_HAS_CPP_20 (do not remove; see note at top of file) // JSON_HAS_CPP_20 (do not remove; see note at top of file)
TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t") TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t")
@@ -293,180 +228,6 @@ TEST_CASE("Parse with std::counted_iterator and std::default_sentinel_t")
const std::counted_iterator<iterator_type> first2(json_str.begin(), len); const std::counted_iterator<iterator_type> first2(json_str.begin(), len);
CHECK(json::accept(first2, std::default_sentinel)); CHECK(json::accept(first2, std::default_sentinel));
} }
TEST_CASE("std::counted_iterator reaches the contiguous fast paths")
{
// A sized sentinel makes the remaining element count computable in O(1), so
// std::counted_iterator over a contiguous iterator must reach the same bulk
// string/number scanners as a plain pointer - not just the byte-at-a-time
// fallback (see #5268 for the equivalent memcpy fast path).
#if JSON_HAS_RANGES
// JSON_HAS_RANGES is 0 on standard libraries with an incomplete <ranges>
// (libstdc++ < 11, libc++ < 16), where the adapter deliberately falls back
// to the byte-at-a-time scanner; everything below still has to work there.
using adapter_type = nlohmann::detail::iterator_input_adapter<std::counted_iterator<const char*>, std::default_sentinel_t>;
CHECK(adapter_type::supports_bulk_scan);
CHECK(adapter_type::supports_seek);
#endif
// exercise every fast path: long ASCII run, multibyte UTF-8, escapes, and
// integer/floating-point numbers
const std::string json_str =
R"({"ascii":"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",)"
"\"utf8\":\"\xe4\xb8\xad\xe6\x96\x87\xf0\x9f\x98\x80\xc3\xa9\","
R"("escaped":"aéb\n\\","ints":[0,-1,18446744073709551615,-9223372036854775808],)"
R"("floats":[1.5,-2.25e3,0.30000000000000004]})";
const auto len = static_cast<std::iter_difference_t<const char*>>(json_str.size());
const std::counted_iterator<const char*> first(json_str.data(), len);
const json j = json::parse(first, std::default_sentinel);
// parsing through the pointer adapter must give exactly the same result
CHECK(j == json::parse(json_str));
#if !defined(JSON_NOEXCEPTION)
// Diagnostics that quote the offending token are reconstructed from the
// already-consumed input (supports_seek), a path a sized sentinel only
// reaches now; check a few that include the "last read" text. Parsing
// invalid input aborts when exceptions are off, hence the guard.
// Raw strings and explicit bytes: an escaped literal and two literals
// written next to each other both read as mistakes to static analysis.
const auto byte = [](int value)
{
return std::string(1, static_cast<char>(value));
};
const std::vector<std::string> diagnostic_docs =
{
"1\nx",
"truX",
"[tru]",
R"("abc)",
R"(["\ud834"])",
R"(["a)" + byte(0x01) + R"(b"])",
R"([")" + byte(0xC3) + byte(0x28) + R"("])",
"[1e]",
R"(["aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaX)"
};
for (const auto& text : diagnostic_docs)
{
CAPTURE(text);
const std::counted_iterator<const char*> it(text.data(), static_cast<std::iter_difference_t<const char*>>(text.size()));
std::string counted_message;
std::string string_message;
try
{
const json counted_result = json::parse(it, std::default_sentinel);
static_cast<void>(counted_result);
}
catch (const json::parse_error& e)
{
counted_message = e.what();
}
try
{
const json string_result = json::parse(text);
static_cast<void>(string_result);
}
catch (const json::parse_error& e)
{
string_message = e.what();
}
CHECK_FALSE(counted_message.empty());
CHECK(counted_message == string_message);
}
// and errors must still be reported identically
const std::string bad = "[01\n]";
const std::counted_iterator<const char*> bad_first(bad.data(), static_cast<std::iter_difference_t<const char*>>(bad.size()));
std::string counted_what;
std::string string_what;
try
{
const json counted_result = json::parse(bad_first, std::default_sentinel);
static_cast<void>(counted_result);
}
catch (const json::parse_error& e)
{
counted_what = e.what();
}
try
{
const json string_result = json::parse(bad);
static_cast<void>(string_result);
}
catch (const json::parse_error& e)
{
string_what = e.what();
}
CHECK_FALSE(counted_what.empty());
CHECK(counted_what == string_what);
#endif
}
#if !defined(JSON_NOEXCEPTION)
// several cases below are truncated on purpose, and parsing invalid input
// aborts when exceptions are off
TEST_CASE("std::counted_iterator bulk scanning stops at the counted end")
{
// The count, not the size of the underlying buffer, is the end of the
// input: the bulk scanners must never look at the bytes behind it, even
// though they are readable. Each case is compared against parsing the
// equivalent prefix as a std::string.
const auto via_counted = [](const std::string & buf, std::size_t n) -> std::string
{
const std::counted_iterator<const char*> first(buf.data(), static_cast<std::iter_difference_t<const char*>>(n));
try
{
const json j = json::parse(first, std::default_sentinel);
return "OK|" + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
const auto via_prefix = [](const std::string & buf, std::size_t n) -> std::string
{
try
{
const json j = json::parse(buf.substr(0, n));
return "OK|" + j.dump();
}
catch (const json::parse_error& e)
{
return {e.what()};
}
};
struct testcase // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init)
{
const char* buffer;
std::size_t count;
};
const std::vector<testcase> cases =
{
{"[\"abc\"]____TRAILING____", 7}, // exact fit, tail hidden
{"[\"abcdefghijklmnop\"]____", 8}, // cut inside a string
{"[\"abc\"]____", 6}, // cut just before the closing quote
{"[12345]xxxxx", 4}, // cut inside a number
{"[123]999999", 5}, // number ends exactly at the count
{"[\"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\"]", 12}, // closing quote only behind the count
{"[\"aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa\"]", 19}, // cut inside an 8-byte SWAR stride
{"[\"\xe4\xb8\xad\xe6\x96\x87\"]", 5}, // cut inside a UTF-8 sequence
{"[\"\xe4\xb8\xad\xe6\x96\x87\"]____", 10}, // complete UTF-8, tail hidden
{"[1.25e3]TRAILINGDIGITS999", 7}, // number token reaches the count
};
for (const auto& tc : cases)
{
CAPTURE(tc.buffer);
CAPTURE(tc.count);
const std::string buffer = tc.buffer;
CHECK(via_counted(buffer, tc.count) == via_prefix(buffer, tc.count));
}
}
#endif
#endif #endif
} // namespace } // namespace