Compare commits

..
Author SHA1 Message Date
Niels Lohmann dd2dd45abd Fix false GCC -Warray-bounds error with JSON_DIAGNOSTICS at -O3 (#5744)
* Re-amalgamate single_include after #5737

#5737 changed sources under include/ without regenerating single_include,
so the single header on develop is out of date. This commit only runs
make amalgamate.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Suppress GCC's false -Warray-bounds error in set_parents() with JSON_DIAGNOSTICS

With JSON_DIAGNOSTICS, GCC 13 to 15 report a false -Warray-bounds error
at -O3 when set_parents() is inlined right after a non-container value
(e.g., a string) was created: the std::map access in the object branch
is checked against the string allocation although m_type rules that
branch out. GCC 15 fixed the pattern from #4819 but not this one.

Suppress -Warray-bounds around set_parents() only, and add a regression
test that is compiled with -O3 -Werror=array-bounds on GCC.

Fixes #5742

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the false -Warray-bounds error by setting the type after the value

The pragma around set_parents() is no longer needed: the base branch
(#5585) creates a value before setting its type. Before, the type of a new
string was stored first, so GCC had to assume that operator new could
change it again, kept the object branch of the inlined set_parents() alive,
and checked it against the string's allocation. Rewriting set_parents()
itself (if/else instead of switch) does not help.

Verified with Docker GCC 12.5, 13, 14, 15.3 and 16.2: the regression test
builds and passes without the pragma, and fails to build on develop.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-02 11:24:10 +02:00
Niels Lohmann 362a902170 Merge branch 'develop' into claude/fix-allocation-exception-safety
include/nlohmann/json.hpp conflicted in update(const_iterator, const_iterator, bool):
kept develop's prepare_update() helper, which already creates the object
before setting its type (the same fix this branch makes elsewhere).
single_include/nlohmann/json.hpp was regenerated via `make amalgamate`
instead of being hand-resolved.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:53:14 +02:00
Niels Lohmann e400780533 Re-amalgamate single_include (#5745)
* Re-amalgamate single_include

#5737 changed 13 headers under include/ but merged without the matching
single_include/nlohmann/json.hpp update, so the amalgamated header still
had, among others, the GCC C++20 -Wignored-attributes pragma block and
the clang -Wdocumentation push/pop that #5737 removed, the forwarding
from_json tuple/array helpers it replaced with const references, and
lacked the output_adapter char_traits changes it added.

Regenerated with `make amalgamate` (astyle 3.4.13). The diff is exactly
`git diff b54ed188e e5a89d671 -- include/` (164+/95-); json_fwd.hpp and
json_literals.hpp were already up to date.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Run the amalgamation check on pushes to develop

#5737 was merged 39 seconds after its last push, while its own
check_amalgamation run was still queued (earlier runs had been
cancelled by the concurrency group), so the stale single_include
reached develop without any failing check.

Also run the check on pushes to develop, without cancelling in-progress
develop runs. The "save" job (PR number/author for the comment
workflow) only runs for pull requests, the checkout falls back to
github.sha, and comment_check_amalgamation.yml only comments for
PR-triggered runs, since push runs have no PR and no "pr" artifact.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:09:34 +02:00
Niels Lohmann ba6bf32f6d Skip the to_json allocation-failure section on VS 2015 Debug
AppVeyor's VS 2015 Debug x86 job crashed (SIGSEGV) in "a failed
allocation leaves the value unchanged / converting into an existing
value". With iterator debugging, VS 2015's containers construct a proxy
with the allocator in constructors that cannot report its failure, so
the throwing test allocator crashes there. The existing guard covered
only the std::vector<bool> check; skip the whole section on that
toolchain instead. The Release builds and all other compilers still run
it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 18:12:52 +02:00
Niels Lohmann 056d187986 Merge branch 'develop' into claude/fix-allocation-exception-safety
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 20:58:25 +02:00
Niels Lohmann 90078597c7 Test the remaining to_json overloads with a failing allocation
The test only reached 5 of the 14 external_constructor overloads that
were reordered: reverting the fix in any of the other 9 (const string_t&,
compatible strings, binary_t&&, array_t const&/&&, std::valarray, ranges
views, object_t const&/&&) left it green. Each of them is now called with
a failing allocation, and reverting any single hunk of the fix makes the
test fail under ASan.

binary_t&& is only reached by to_json for a binary container type other
than std::vector<std::uint8_t>, so the test calls the constructor
directly.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 18:29:37 +02:00
Niels Lohmann fa28fff7d6 Skip the vector<bool> failed-allocation check for VS 2015 with iterator debugging
AppVeyor's VS 2015 Debug build terminated in "converting into an
existing value": to_json for std::vector<bool> default-constructs an
array_t first, and VS 2015's std::vector default constructor is noexcept
but, with iterator debugging, constructs its proxy through the test's
allocator, which throws.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 13:30:57 +02:00
Niels Lohmann e9c84befa1 Keep the created pointer rather than an uninitialized json_value
clang-tidy reported the json_value unions declared in the string, array,
and object constructors of to_json.hpp as uninitialized
(cppcoreguidelines-pro-type-member-init). Assign the created pointer to
the member directly once the old value is destroyed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-26 10:07:22 +02:00
Niels Lohmann 8a26f2dc8f Skip the failed-allocation test when exceptions are disabled
The no-exceptions CI job runs doctest with --no-throw, which skips every
CHECK_THROWS_AS. The test then left next_construct_fails set, and the
next allocation outside a check threw std::bad_alloc.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-26 07:41:42 +02:00
Niels Lohmann 2c108d0b56 Create a value before giving it its type
Several functions set the type of a value before creating the string,
array, object, or binary value it stands for. When that creation threw -
std::bad_alloc from the allocator, say - the value was left behind with
the new type but nothing behind it:

- json::binary() returned no value, but its destructor asserted that a
  binary value has a binary array, aborting debug builds.
- operator[], push_back, emplace_back, emplace, and update turned a null
  value into an array or object that did not exist; any later access
  dereferenced a null pointer.
- The to_json conversions destroyed the old value first. A failed
  creation of the new one left a pointer to the destroyed old value,
  which was then used and freed again (heap-use-after-free).

The value is now created first and the type set after it, and to_json
destroys the old value only once the new one exists, so a failed
allocation leaves the value unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 22:19:39 +02:00
40 changed files with 587 additions and 827 deletions
+10 -3
View File
@@ -2,16 +2,23 @@ name: "Check amalgamation"
on: on:
pull_request: pull_request:
# also check develop itself: a PR can be merged before its own run of this
# workflow completes (e.g. while it is still queued), leaving single_include
# stale on develop without any failing check
push:
branches:
- develop
concurrency: concurrency:
group: ${{ github.workflow }}-${{ github.ref || github.run_id }} group: ${{ github.workflow }}-${{ github.ref || github.run_id }}
cancel-in-progress: true cancel-in-progress: ${{ github.event_name == 'pull_request' }}
permissions: permissions:
contents: read contents: read
jobs: jobs:
save: save:
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
- name: Harden Runner - name: Harden Runner
@@ -43,11 +50,11 @@ jobs:
with: with:
egress-policy: audit egress-policy: audit
- name: Checkout pull request - name: Checkout pull request or pushed commit
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with: with:
path: main path: main
ref: ${{ github.event.pull_request.head.sha }} ref: ${{ github.event.pull_request.head.sha || github.sha }}
persist-credentials: false persist-credentials: false
- name: Checkout tools - name: Checkout tools
@@ -10,7 +10,8 @@ permissions:
jobs: jobs:
comment: comment:
if: ${{ github.event.workflow_run.conclusion == 'failure' }} # push runs on develop have no PR to comment on (and no "pr" artifact)
if: ${{ github.event.workflow_run.conclusion == 'failure' && github.event.workflow_run.event == 'pull_request' }}
runs-on: ubuntu-latest runs-on: ubuntu-latest
permissions: permissions:
contents: read contents: read
-6
View File
@@ -61,7 +61,6 @@ option(JSON_Install "Install CMake targets during install
option(JSON_MultipleHeaders "Use non-amalgamated version of the library." ON) option(JSON_MultipleHeaders "Use non-amalgamated version of the library." ON)
option(JSON_SystemInclude "Include as system headers (skip for clang-tidy)." OFF) option(JSON_SystemInclude "Include as system headers (skip for clang-tidy)." OFF)
option(JSON_StrictNulHandling "Build with strict NUL-byte handling enabled." OFF) option(JSON_StrictNulHandling "Build with strict NUL-byte handling enabled." OFF)
option(JSON_StrictBinaryUTF8 "Build with UTF-8 checks in the CBOR, UBJSON, BJData, and BSON writers enabled." OFF)
if (JSON_CI) if (JSON_CI)
include(ci) include(ci)
@@ -119,10 +118,6 @@ if (JSON_StrictNulHandling)
message(STATUS "Strict NUL-byte handling enabled (JSON_STRICT_NUL_HANDLING=1)") message(STATUS "Strict NUL-byte handling enabled (JSON_STRICT_NUL_HANDLING=1)")
endif() endif()
if (JSON_StrictBinaryUTF8)
message(STATUS "Strict UTF-8 checks in binary writers enabled (JSON_STRICT_BINARY_UTF8=1)")
endif()
if (JSON_Diagnostic_Positions) if (JSON_Diagnostic_Positions)
message(STATUS "Diagnostic positions enabled (JSON_DIAGNOSTIC_POSITIONS=1)") message(STATUS "Diagnostic positions enabled (JSON_DIAGNOSTIC_POSITIONS=1)")
endif() endif()
@@ -158,7 +153,6 @@ target_compile_definitions(
$<$<BOOL:${JSON_Diagnostic_Positions}>:JSON_DIAGNOSTIC_POSITIONS=1> $<$<BOOL:${JSON_Diagnostic_Positions}>:JSON_DIAGNOSTIC_POSITIONS=1>
$<$<BOOL:${JSON_LegacyDiscardedValueComparison}>:JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1> $<$<BOOL:${JSON_LegacyDiscardedValueComparison}>:JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1>
$<$<BOOL:${JSON_StrictNulHandling}>:JSON_STRICT_NUL_HANDLING=1> $<$<BOOL:${JSON_StrictNulHandling}>:JSON_STRICT_NUL_HANDLING=1>
$<$<BOOL:${JSON_StrictBinaryUTF8}>:JSON_STRICT_BINARY_UTF8=1>
) )
target_include_directories( target_include_directories(
+1 -1
View File
@@ -691,7 +691,7 @@ ci_get_cmake(4.0.0 CMAKE_4_0_0_BINARY)
# the tests require CMake 3.13 or later, so they are excluded for CMake 3.5.0 # the tests require CMake 3.13 or later, so they are excluded for CMake 3.5.0
set(JSON_CMAKE_FLAGS_3_5_0 JSON_Diagnostics JSON_Diagnostic_Positions JSON_GlobalUDLs JSON_ImplicitConversions JSON_DisableEnumSerialization set(JSON_CMAKE_FLAGS_3_5_0 JSON_Diagnostics JSON_Diagnostic_Positions JSON_GlobalUDLs JSON_ImplicitConversions JSON_DisableEnumSerialization
JSON_LegacyDiscardedValueComparison JSON_Install JSON_MultipleHeaders JSON_SystemInclude JSON_Valgrind JSON_LegacyDiscardedValueComparison JSON_Install JSON_MultipleHeaders JSON_SystemInclude JSON_Valgrind
JSON_StrictNulHandling JSON_StrictBinaryUTF8) JSON_StrictNulHandling)
set(JSON_CMAKE_FLAGS_3_31_6 JSON_BuildTests ${JSON_CMAKE_FLAGS_3_5_0}) set(JSON_CMAKE_FLAGS_3_31_6 JSON_BuildTests ${JSON_CMAKE_FLAGS_3_5_0})
set(JSON_CMAKE_FLAGS_4_0_0 JSON_BuildTests ${JSON_CMAKE_FLAGS_3_5_0}) set(JSON_CMAKE_FLAGS_4_0_0 JSON_BuildTests ${JSON_CMAKE_FLAGS_3_5_0})
+1 -6
View File
@@ -56,9 +56,6 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size` - Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false. is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is not
valid UTF-8 and [`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled; otherwise, the bytes are
written unchanged
## Complexity ## Complexity
@@ -92,6 +89,4 @@ Linear in the size of the JSON value `j`.
## Version history ## Version history
- Added in version 3.11.0. - Added in version 3.11.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0. - BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 if
[`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled added in version 3.13.0.
@@ -46,9 +46,6 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`out_of_range.415`](../../home/exceptions.md#jsonexceptionout_of_range415) if the subtype of a binary value - Throws [`out_of_range.415`](../../home/exceptions.md#jsonexceptionout_of_range415) if the subtype of a binary value
exceeds 255, the maximum of the BSON binary subtype; example: exceeds 255, the maximum of the BSON binary subtype; example:
`"subtype 70000 is too large for the BSON binary subtype (max 255)"` `"subtype 70000 is too large for the BSON binary subtype (max 255)"`
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key is not valid
UTF-8 and [`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled; otherwise, the bytes are
written unchanged
## Complexity ## Complexity
@@ -85,6 +82,3 @@ pass before anything is written.
- Added in version 3.4.0. - Added in version 3.4.0.
- Linear in the size of `j`, and no longer limited by the call stack for deeply nested values, since version 3.13.0. - Linear in the size of `j`, and no longer limited by the call stack for deeply nested values, since version 3.13.0.
- `out_of_range.415` is now detected before anything is written, like the other exceptions above, since version 3.13.0. - `out_of_range.415` is now detected before anything is written, like the other exceptions above, since version 3.13.0.
- Throwing `type_error.316` for a string value or object key that is not valid UTF-8 if
[`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled, detected before anything is written,
added in version 3.13.0.
@@ -35,12 +35,6 @@ The exact mapping and its limitations are described on a [dedicated page](../../
Strong guarantee: if an exception is thrown, there are no changes in the JSON value. Strong guarantee: if an exception is thrown, there are no changes in the JSON value.
## Exceptions
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is not
valid UTF-8 and [`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled; otherwise, the bytes are
written unchanged
## Complexity ## Complexity
Linear in the size of the JSON value `j`. Linear in the size of the JSON value `j`.
@@ -74,5 +68,3 @@ Linear in the size of the JSON value `j`.
- Added in version 2.0.9. - Added in version 2.0.9.
- Compact representation of floating-point numbers added in version 3.8.0. - Compact representation of floating-point numbers added in version 3.8.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 if
[`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled added in version 3.13.0.
@@ -49,9 +49,6 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size` - Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false. is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is not
valid UTF-8 and [`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled; otherwise, the bytes are
written unchanged
## Complexity ## Complexity
@@ -85,5 +82,3 @@ Linear in the size of the JSON value `j`.
## Version history ## Version history
- Added in version 3.1.0. - Added in version 3.1.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 if
[`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md) is enabled added in version 3.13.0.
-2
View File
@@ -18,8 +18,6 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_PRECISE_STREAM_POSITION**](json_precise_stream_position.md) - opt in to leaving an input stream positioned - [**JSON_PRECISE_STREAM_POSITION**](json_precise_stream_position.md) - opt in to leaving an input stream positioned
right after a parsed number right after a parsed number
- [**JSON_STRICT_BINARY_UTF8**](json_strict_binary_utf8.md) - opt in to checking strings for valid UTF-8 in the CBOR,
UBJSON, BJData, and BSON writers
- [**JSON_STRICT_NUL_HANDLING**](json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input instead of - [**JSON_STRICT_NUL_HANDLING**](json_strict_nul_handling.md) - opt in to rejecting a NUL byte in the input instead of
treating it as end of input treating it as end of input
@@ -1,97 +0,0 @@
# JSON_STRICT_BINARY_UTF8
```cpp
#define JSON_STRICT_BINARY_UTF8 /* value */
```
When defined to `1`, the binary writers [`to_cbor`](../basic_json/to_cbor.md), [`to_ubjson`](../basic_json/to_ubjson.md),
[`to_bjdata`](../basic_json/to_bjdata.md), and [`to_bson`](../basic_json/to_bson.md) check every string value and
object key for valid UTF-8 and throw [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for
ill-formed UTF-8, like [`dump`](../basic_json/dump.md) does. Without it, they write the bytes unchanged.
The macro does not affect:
- [`to_msgpack`](../basic_json/to_msgpack.md): the MessagePack specification allows a `str` value to contain bytes that
are not valid UTF-8, so it always writes them unchanged.
- [`to_bon8`](../basic_json/to_bon8.md): BON8 always checks, because the UTF-8 lead bytes mark where a string ends.
- The binary readers ([`from_cbor`](../basic_json/from_cbor.md), [`from_msgpack`](../basic_json/from_msgpack.md),
[`from_ubjson`](../basic_json/from_ubjson.md), [`from_bjdata`](../basic_json/from_bjdata.md),
[`from_bson`](../basic_json/from_bson.md)): none of these formats requires a decoder to reject ill-formed UTF-8, so
they always return the bytes unchanged.
## Default definition
The default value is `0` (disabled, the behavior of version 3.12.0 and earlier is preserved).
```cpp
#define JSON_STRICT_BINARY_UTF8 0
```
## Notes
!!! note "Background"
CBOR, UBJSON, BJData, and BSON all require strings to be UTF-8. Up to version 3.12.0, the writers did not check
this, so they could produce output that other decoders reject. Checking by default would break code that stores
other encodings (for instance ISO 8859-1) in a string and only ever writes it to a binary format, so this macro
offers the check as an opt-in ahead of version 4.0.0, where it is planned to become the default (see
[#5529](https://github.com/nlohmann/json/issues/5529) and [#5651](https://github.com/nlohmann/json/issues/5651)).
!!! warning "Opt-in only"
This macro must be defined **before** including `<nlohmann/json.hpp>`. Defining it after the include has no
effect.
!!! note "ABI compatibility"
The value of this macro is encoded in the [namespace](../../features/namespace.md) (tag `_sbu8`), resulting in
distinct symbol names. Translation units compiled with and without it can therefore be linked into the same program
without One Definition Rule (ODR) violations, but they cannot exchange instances of library types.
## Examples
??? example "Default behavior (macro not defined)"
Without the macro, the bytes are written unchanged:
```cpp
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
auto v = json::to_cbor(json("\xFF"));
// v is {0x61, 0xFF}
}
```
??? example "Opt-in check (macro defined to 1)"
With the macro, ill-formed UTF-8 is rejected:
```cpp
#define JSON_STRICT_BINARY_UTF8 1
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
auto v = json::to_cbor(json("\xFF"));
// throws type_error.316: invalid UTF-8 byte at index 0: 0xFF
}
```
## See also
- [**to_cbor**](../basic_json/to_cbor.md) - create a CBOR serialization of a JSON value
- [**to_ubjson**](../basic_json/to_ubjson.md) - create a UBJSON serialization of a JSON value
- [**to_bjdata**](../basic_json/to_bjdata.md) - create a BJData serialization of a JSON value
- [**to_bson**](../basic_json/to_bson.md) - create a BSON serialization of a JSON value
- [**error_handler_t**](../basic_json/error_handler_t.md) - how [`dump`](../basic_json/dump.md) treats ill-formed UTF-8
## Version history
- Added in version 3.13.0.
- Planned to become the default (with the macro removed) in version 4.0.0.
@@ -63,13 +63,6 @@ The library uses the following mapping from JSON values types to BJData types ac
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical) - strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
BJData strings must use UTF-8 encoding. By default, `to_bjdata()` writes the bytes of string values and object keys
unchanged, even if they are not valid UTF-8. If
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
!!! info "Unused BJData markers" !!! info "Unused BJData markers"
The following markers are not used in the conversion: The following markers are not used in the conversion:
@@ -215,15 +208,6 @@ The library maps BJData types to JSON value types as follows:
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value. The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
BJData strings must use UTF-8 encoding, but this is not enforced on read: `from_bjdata()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_bjdata()` writes such a value
back unchanged (see above).
!!! info "Round trips" !!! info "Round trips"
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
@@ -109,16 +109,14 @@ The library maps BSON record types to JSON value types as follows:
If BSON input must be validated for strict specification compliance, validate it separately before passing it to If BSON input must be validated for strict specification compliance, validate it separately before passing it to
`from_bson()`. `from_bson()`.
!!! warning "Ill-formed UTF-8 in string values" !!! warning "UTF-8 validation of string values"
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
decoder. `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged. bytes of every such string at decode time and rejects ill-formed UTF-8 with a
However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
passed that replaces or ignores the ill-formed bytes. By default, `to_bson()` writes such a string value or element (key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
(key) name unchanged; if [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it byte-by-byte as a C string, or are not required to hold text, respectively.
throws the same exception instead. Element (key) names are never validated on read, since they are read byte-by-byte
as a C string. `binary` values (type `0x05`) are unaffected, since they are not required to hold text.
??? example ??? example
@@ -189,16 +189,15 @@ The library maps CBOR types to JSON value types as follows:
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a ([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead. general-purpose CBOR library instead.
!!! warning "Ill-formed UTF-8 in text strings" !!! warning "UTF-8 validation of text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings (major [RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this. This library does not: (major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
`from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back decode time and rejects ill-formed UTF-8 with a
unchanged. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
passed that replaces or ignores the ill-formed bytes. By default, `to_cbor()` writes such a value back unchanged; if dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws the same exception text.
instead. Byte strings (major type 2) are unaffected, since they are not required to hold text.
!!! warning "Tagged items" !!! warning "Tagged items"
@@ -153,15 +153,14 @@ The library maps MessagePack types to JSON value types as follows:
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead. on. Such input needs a general-purpose MessagePack library instead.
!!! warning "Ill-formed UTF-8 in string values" !!! warning "UTF-8 validation of string values"
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged. This library validates the bytes of every such string (object keys included) at decode time and rejects
This library follows that: `from_msgpack()` reads `str` bytes (object keys included) as-is, without validating ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
them, and `to_msgpack()` writes them back as-is, so such a value round-trips through `from_msgpack(to_msgpack(j))` with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
byte for byte. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way, unless an to hold text.
error handler is passed that replaces or ignores the ill-formed bytes.
??? example ??? example
@@ -47,13 +47,6 @@ The library uses the following mapping from JSON values types to UBJSON types ac
- strings with more than 9223372036854775807 bytes (theoretical) - strings with more than 9223372036854775807 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
UBJSON's required string encoding is UTF-8. By default, `to_ubjson()` writes the bytes of string values and object
keys unchanged, even if they are not valid UTF-8. If
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
!!! info "Unused UBJSON markers" !!! info "Unused UBJSON markers"
The following markers are not used in the conversion: The following markers are not used in the conversion:
@@ -127,15 +120,6 @@ The library maps UBJSON types to JSON value types as follows:
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value. The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
UBJSON's required string encoding is UTF-8, but this is not enforced on read: `from_ubjson()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_ubjson()` writes such a value
back unchanged (see above).
??? example ??? example
```cpp ```cpp
-14
View File
@@ -138,20 +138,6 @@ using the library with compilers that do not fully support C++11 and may only wo
See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](../api/macros/json_skip_unsupported_compiler_check.md). See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](../api/macros/json_skip_unsupported_compiler_check.md).
## `JSON_STRICT_BINARY_UTF8`
When defined to `1`, [`to_cbor`](../api/basic_json/to_cbor.md), [`to_ubjson`](../api/basic_json/to_ubjson.md),
[`to_bjdata`](../api/basic_json/to_bjdata.md), and [`to_bson`](../api/basic_json/to_bson.md) throw
[`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) for a string value or object key that is not
valid UTF-8. The default value is `0`, which writes the bytes unchanged as before version 3.13.0; this is planned to
become the default in version 4.0.0.
The check can also be enabled with the CMake option
[`JSON_StrictBinaryUTF8`](../integration/cmake.md#json_strictbinaryutf8) (`OFF` by default) which sets
`JSON_STRICT_BINARY_UTF8` accordingly.
See [full documentation of `JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md).
## `JSON_STRICT_NUL_HANDLING` ## `JSON_STRICT_NUL_HANDLING`
When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other
-1
View File
@@ -20,7 +20,6 @@ The complete default namespace name is derived as follows:
`_bics`. `_bics`.
- [`JSON_PRECISE_STREAM_POSITION`](../api/macros/json_precise_stream_position.md) defined non-zero appends `_psp`. - [`JSON_PRECISE_STREAM_POSITION`](../api/macros/json_precise_stream_position.md) defined non-zero appends `_psp`.
- [`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md) defined non-zero appends `_snul`. - [`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md) defined non-zero appends `_snul`.
- [`JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md) defined non-zero appends `_sbu8`.
- The inline namespace ends with the suffix `_v` followed by the 3 components of the version number separated by - The inline namespace ends with the suffix `_v` followed by the 3 components of the version number separated by
underscores. To omit the version component, see [Disabling the version component](#disabling-the-version-component) underscores. To omit the version component, see [Disabling the version component](#disabling-the-version-component)
below. below.
+5 -8
View File
@@ -340,9 +340,8 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
### json.exception.parse_error.113 ### json.exception.parse_error.113
A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
string was read where one was required (for instance as a map key), or the string's length specification is invalid. string was read where one was required (for instance as a map key), the string's length specification is invalid, or
The bytes of a string itself are not checked for valid UTF-8 on read; see the ill-formed UTF-8 notes on the the string's bytes are not valid UTF-8.
individual [binary format](../features/binary_formats/index.md) pages for how such a string is handled afterward.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on type (for instance integers or `null`) are therefore not supported; see the notes on
@@ -365,6 +364,9 @@ type (for instance integers or `null`) are therefore not supported; see the note
``` ```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative [json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative
``` ```
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte
```
### json.exception.parse_error.114 ### json.exception.parse_error.114
@@ -747,11 +749,6 @@ The `unflatten()` function only works for an object whose keys are JSON Pointers
The `dump()` function only works with UTF-8 encoded strings; that is, if you assign a `std::string` to a JSON value, make sure it is UTF-8 encoded. The `dump()` function only works with UTF-8 encoded strings; that is, if you assign a `std::string` to a JSON value, make sure it is UTF-8 encoded.
If [`JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md) is enabled, the binary writers
[`to_cbor()`](../api/basic_json/to_cbor.md), [`to_ubjson()`](../api/basic_json/to_ubjson.md),
[`to_bjdata()`](../api/basic_json/to_bjdata.md), and [`to_bson()`](../api/basic_json/to_bson.md) throw this exception
for a string value or object key that is not valid UTF-8 as well.
!!! failure "Example message" !!! failure "Example message"
Calling `dump()` on a JSON value containing an ISO 8859-1 encoded string: Calling `dump()` on a JSON value containing an ISO 8859-1 encoded string:
-5
View File
@@ -204,11 +204,6 @@ Use the non-amalgamated version of the library. This option is `ON` by default.
Treat the library headers like system headers (i.e., adding `SYSTEM` to the [`target_include_directories`](https://cmake.org/cmake/help/latest/command/target_include_directories.html) call) to check for this library by tools like Clang-Tidy. This option is `OFF` by default. Treat the library headers like system headers (i.e., adding `SYSTEM` to the [`target_include_directories`](https://cmake.org/cmake/help/latest/command/target_include_directories.html) call) to check for this library by tools like Clang-Tidy. This option is `OFF` by default.
### `JSON_StrictBinaryUTF8`
Check string values and object keys for valid UTF-8 in the CBOR, UBJSON, BJData, and BSON writers, by defining the
macro [`JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md). This option is `OFF` by default.
### `JSON_StrictNulHandling` ### `JSON_StrictNulHandling`
Reject a `'\0'` (NUL) byte in the input instead of treating it as end of input, by defining the macro Reject a `'\0'` (NUL) byte in the input instead of treating it as end of input, by defining the macro
-1
View File
@@ -301,7 +301,6 @@ nav:
- 'JSON_PRECISE_STREAM_POSITION': api/macros/json_precise_stream_position.md - 'JSON_PRECISE_STREAM_POSITION': api/macros/json_precise_stream_position.md
- 'JSON_SKIP_LIBRARY_VERSION_CHECK': api/macros/json_skip_library_version_check.md - 'JSON_SKIP_LIBRARY_VERSION_CHECK': api/macros/json_skip_library_version_check.md
- 'JSON_SKIP_UNSUPPORTED_COMPILER_CHECK': api/macros/json_skip_unsupported_compiler_check.md - 'JSON_SKIP_UNSUPPORTED_COMPILER_CHECK': api/macros/json_skip_unsupported_compiler_check.md
- 'JSON_STRICT_BINARY_UTF8': api/macros/json_strict_binary_utf8.md
- 'JSON_STRICT_NUL_HANDLING': api/macros/json_strict_nul_handling.md - 'JSON_STRICT_NUL_HANDLING': api/macros/json_strict_nul_handling.md
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md - 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md - 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
+4 -15
View File
@@ -46,10 +46,6 @@
#define JSON_STRICT_NUL_HANDLING 0 #define JSON_STRICT_NUL_HANDLING 0
#endif #endif
#ifndef JSON_STRICT_BINARY_UTF8
#define JSON_STRICT_BINARY_UTF8 0
#endif
#if JSON_DIAGNOSTICS #if JSON_DIAGNOSTICS
#define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag #define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag
#else #else
@@ -86,20 +82,14 @@
#define NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING #define NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING
#endif #endif
#if JSON_STRICT_BINARY_UTF8
#define NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8 _sbu8
#else
#define NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8
#endif
#ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION #ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION
#define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0 #define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0
#endif #endif
// Construct the namespace ABI tags component // Construct the namespace ABI tags component
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f, g) json_abi ## a ## b ## c ## d ## e ## f ## g #define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f) json_abi ## a ## b ## c ## d ## e ## f
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e, f, g) \ #define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e, f) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f, g) NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f)
#define NLOHMANN_JSON_ABI_TAGS \ #define NLOHMANN_JSON_ABI_TAGS \
NLOHMANN_JSON_ABI_TAGS_CONCAT( \ NLOHMANN_JSON_ABI_TAGS_CONCAT( \
@@ -108,8 +98,7 @@
NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \ NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \ NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \
NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION, \ NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION, \
NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING, \ NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING)
NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8)
// Construct the namespace version component // Construct the namespace version component
#define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \ #define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \
+45 -25
View File
@@ -42,6 +42,10 @@ namespace detail
* j.m_data.m_value.destroy(j.m_data.m_type) to avoid a memory leak in case j contains an * j.m_data.m_value.destroy(j.m_data.m_type) to avoid a memory leak in case j contains an
* allocated value (e.g., a string). See bug issue * allocated value (e.g., a string). See bug issue
* https://github.com/nlohmann/json/issues/2865 for more information. * https://github.com/nlohmann/json/issues/2865 for more information.
*
* A value that has to be allocated is created before the old one is destroyed:
* were it the other way around, an exception while creating the new value would
* leave j with the type of the new value, but the pointer to the destroyed old one.
*/ */
template<value_t> struct external_constructor; template<value_t> struct external_constructor;
@@ -65,18 +69,20 @@ struct external_constructor<value_t::string>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::string_t& s) static void construct(BasicJsonType& j, const typename BasicJsonType::string_t& s)
{ {
const typename BasicJsonType::json_value value(s);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::string; j.m_data.m_type = value_t::string;
j.m_data.m_value = s; j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::string_t&& s) static void construct(BasicJsonType& j, typename BasicJsonType::string_t&& s)
{ {
const typename BasicJsonType::json_value value(std::move(s));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::string; j.m_data.m_type = value_t::string;
j.m_data.m_value = std::move(s); j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
@@ -85,9 +91,10 @@ struct external_constructor<value_t::string>
int > = 0 > int > = 0 >
static void construct(BasicJsonType& j, const CompatibleStringType& str) static void construct(BasicJsonType& j, const CompatibleStringType& str)
{ {
auto* created = j.template create<typename BasicJsonType::string_t>(str);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::string; j.m_data.m_type = value_t::string;
j.m_data.m_value.string = j.template create<typename BasicJsonType::string_t>(str); j.m_data.m_value.string = created;
j.assert_invariant(); j.assert_invariant();
} }
}; };
@@ -98,18 +105,20 @@ struct external_constructor<value_t::binary>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::binary_t& b) static void construct(BasicJsonType& j, const typename BasicJsonType::binary_t& b)
{ {
const typename BasicJsonType::json_value value(b);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::binary; j.m_data.m_type = value_t::binary;
j.m_data.m_value = typename BasicJsonType::binary_t(b); j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::binary_t&& b) static void construct(BasicJsonType& j, typename BasicJsonType::binary_t&& b)
{ {
const typename BasicJsonType::json_value value(std::move(b));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::binary; j.m_data.m_type = value_t::binary;
j.m_data.m_value = typename BasicJsonType::binary_t(std::move(b)); j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
}; };
@@ -159,9 +168,10 @@ struct external_constructor<value_t::array>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::array_t& arr) static void construct(BasicJsonType& j, const typename BasicJsonType::array_t& arr)
{ {
const typename BasicJsonType::json_value value(arr);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value = arr; j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -169,9 +179,10 @@ struct external_constructor<value_t::array>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::array_t&& arr) static void construct(BasicJsonType& j, typename BasicJsonType::array_t&& arr)
{ {
const typename BasicJsonType::json_value value(std::move(arr));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value = std::move(arr); j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -187,9 +198,10 @@ struct external_constructor<value_t::array>
using std::begin; using std::begin;
using std::end; using std::end;
auto* created = j.template create<typename BasicJsonType::array_t>(begin(arr), end(arr));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value.array = j.template create<typename BasicJsonType::array_t>(begin(arr), end(arr)); j.m_data.m_value.array = created;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -197,15 +209,17 @@ struct external_constructor<value_t::array>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const std::vector<bool>& arr) static void construct(BasicJsonType& j, const std::vector<bool>& arr)
{ {
j.m_data.m_value.destroy(j.m_data.m_type); typename BasicJsonType::array_t elements;
j.m_data.m_type = value_t::array; elements.reserve(arr.size());
j.m_data.m_value = value_t::array;
j.m_data.m_value.array->reserve(arr.size());
for (const bool x : arr) for (const bool x : arr)
{ {
j.m_data.m_value.array->push_back(x); elements.push_back(x);
j.set_parent(j.m_data.m_value.array->back());
} }
const typename BasicJsonType::json_value value(std::move(elements));
j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array;
j.m_data.m_value = value;
j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -213,11 +227,12 @@ struct external_constructor<value_t::array>
enable_if_t<std::is_convertible<T, BasicJsonType>::value, int> = 0> enable_if_t<std::is_convertible<T, BasicJsonType>::value, int> = 0>
static void construct(BasicJsonType& j, const std::valarray<T>& arr) static void construct(BasicJsonType& j, const std::valarray<T>& arr)
{ {
typename BasicJsonType::array_t elements(arr.size());
std::copy(std::begin(arr), std::end(arr), elements.begin());
const typename BasicJsonType::json_value value(std::move(elements));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value = value_t::array; j.m_data.m_value = value;
j.m_data.m_value.array->resize(arr.size());
std::copy(std::begin(arr), std::end(arr), j.m_data.m_value.array->begin());
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -229,14 +244,16 @@ struct external_constructor<value_t::array>
enable_if_t<is_compatible_range_view<std::remove_cvref_t<CompatibleArrayType>>::value, int> = 0> enable_if_t<is_compatible_range_view<std::remove_cvref_t<CompatibleArrayType>>::value, int> = 0>
static void construct(BasicJsonType& j, CompatibleArrayType && arr) static void construct(BasicJsonType& j, CompatibleArrayType && arr)
{ {
j.m_data.m_value.destroy(j.m_data.m_type); typename BasicJsonType::array_t elements;
j.m_data.m_type = value_t::array;
j.m_data.m_value = value_t::array;
for (auto&& x : std::forward<CompatibleArrayType>(arr)) for (auto&& x : std::forward<CompatibleArrayType>(arr))
{ {
j.m_data.m_value.array->push_back(x); elements.push_back(x);
j.set_parent(j.m_data.m_value.array->back());
} }
const typename BasicJsonType::json_value value(std::move(elements));
j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array;
j.m_data.m_value = value;
j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
#endif #endif
@@ -248,9 +265,10 @@ struct external_constructor<value_t::object>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::object_t& obj) static void construct(BasicJsonType& j, const typename BasicJsonType::object_t& obj)
{ {
const typename BasicJsonType::json_value value(obj);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::object; j.m_data.m_type = value_t::object;
j.m_data.m_value = obj; j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -258,9 +276,10 @@ struct external_constructor<value_t::object>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::object_t&& obj) static void construct(BasicJsonType& j, typename BasicJsonType::object_t&& obj)
{ {
const typename BasicJsonType::json_value value(std::move(obj));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::object; j.m_data.m_type = value_t::object;
j.m_data.m_value = std::move(obj); j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -272,9 +291,10 @@ struct external_constructor<value_t::object>
using std::begin; using std::begin;
using std::end; using std::end;
auto* created = j.template create<typename BasicJsonType::object_t>(begin(obj), end(obj));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::object; j.m_data.m_type = value_t::object;
j.m_data.m_value.object = j.template create<typename BasicJsonType::object_t>(begin(obj), end(obj)); j.m_data.m_value.object = created;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -4031,13 +4031,28 @@ class binary_reader
const NumberType len, const NumberType len,
string_t& result) string_t& result)
{ {
// Strings are taken as is: none of CBOR (RFC 8949 §3.1 leaves the // get_bytes() appends to result, and CBOR indefinite-length strings
// choice to the decoder), MessagePack (whose spec explicitly allows // collect all their chunks in the same result; validating only the
// a str object to contain an invalid byte sequence), UBJSON, BJData, // newly read bytes keeps the check linear in the input size
// or BSON requires a decoder to reject ill-formed UTF-8. The bytes const std::size_t old_size = result.size();
// are kept unchanged; dump() and the binary writers are the ones if (JSON_HEDLEY_UNLIKELY(!get_bytes(format, len, "string", result)))
// that check them and report type_error.316 if they are not valid. {
return get_bytes(format, len, "string", result); return false;
}
// RFC 8949 (CBOR) §3.1 and the MessagePack/BSON/UBJSON specifications
// all require text strings to be valid UTF-8; reject anything else
// right here so malformed input is caught at decode time instead of
// only surfacing later as a type_error.316 when the value is dumped
// (which would defeat allow_exceptions=false / strict discarding).
if (JSON_HEDLEY_UNLIKELY(!is_valid_utf8(result, old_size)))
{
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(113, chars_read,
exception_message(format, "invalid string: ill-formed UTF-8 byte", "string"), nullptr));
}
return true;
} }
/*! /*!
@@ -41,7 +41,6 @@
#undef JSON_BRACE_INIT_COPY_SEMANTICS #undef JSON_BRACE_INIT_COPY_SEMANTICS
#undef JSON_PRECISE_STREAM_POSITION #undef JSON_PRECISE_STREAM_POSITION
#undef JSON_STRICT_NUL_HANDLING #undef JSON_STRICT_NUL_HANDLING
#undef JSON_STRICT_BINARY_UTF8
#endif #endif
#include <nlohmann/thirdparty/hedley/hedley_undef.hpp> #include <nlohmann/thirdparty/hedley/hedley_undef.hpp>
@@ -115,8 +115,6 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or an object key is not valid UTF-8
@throw type_error.317 if @a j is not an object @throw type_error.317 if @a j is not an object
*/ */
void write_bson(const BasicJsonType& j) void write_bson(const BasicJsonType& j)
@@ -147,8 +145,6 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or an object key is not valid UTF-8
*/ */
void write_cbor(const BasicJsonType& j) void write_cbor(const BasicJsonType& j)
{ {
@@ -215,8 +211,6 @@ class binary_writer
case value_t::string: case value_t::string:
{ {
check_text_utf8(*j.m_data.m_value.string, j);
// step 1: write control byte and the string length // step 1: write control byte and the string length
write_cbor_head(0x60, j.m_data.m_value.string->size()); write_cbor_head(0x60, j.m_data.m_value.string->size());
@@ -293,11 +287,6 @@ class binary_writer
// step 2: write each element // step 2: write each element
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
// el.first is checked here, against the object as
// diagnostics context, because write_cbor(el.first)
// converts it to a temporary basic_json that would be
// used as the context instead
check_text_utf8(el.first, j);
write_cbor(el.first); write_cbor(el.first);
write_cbor(el.second); write_cbor(el.second);
} }
@@ -640,8 +629,6 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value @param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false @param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2 @param[in] bjdata_version which BJData version to use, default is draft2
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or an object key is not valid UTF-8
*/ */
void write_ubjson(const BasicJsonType& j, const bool use_count, void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true, const bool use_type, const bool add_prefix = true,
@@ -691,8 +678,6 @@ class binary_writer
case value_t::string: case value_t::string:
{ {
check_text_utf8(*j.m_data.m_value.string, j);
if (add_prefix) if (add_prefix)
{ {
oa.write_character(to_char_type('S')); oa.write_character(to_char_type('S'));
@@ -855,7 +840,6 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
check_text_utf8(el.first, j);
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata); write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
oa.write_characters( oa.write_characters(
reinterpret_cast<const CharType*>(el.first.data()), reinterpret_cast<const CharType*>(el.first.data()),
@@ -900,10 +884,6 @@ class binary_writer
/*! /*!
@return The size of a BSON document entry header, including the id marker @return The size of a BSON document entry header, including the id marker
and the entry name size (and its null-terminator). and the entry name size (and its null-terminator).
@throw out_of_range.409 if @a name contains U+0000, before anything is
written
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a name is
not valid UTF-8, before anything is written
*/ */
static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j) static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
{ {
@@ -913,8 +893,7 @@ class binary_writer
JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j)); JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j));
} }
check_text_utf8(name, j); static_cast<void>(j);
return /*id*/ 1ul + name.size() + /*zero-terminator*/1u; return /*id*/ 1ul + name.size() + /*zero-terminator*/1u;
} }
@@ -970,21 +949,9 @@ class binary_writer
/*! /*!
@return The size of the BSON-encoded string in @a value @return The size of the BSON-encoded string in @a value
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a value
is not valid UTF-8, before anything is written
@note The UTF-8 check is skipped if @a value is already too long for the
32-bit BSON length field (@ref to_bson_length rejects it later, once
the size of the whole document is known); this also keeps the check
from reading past a StringType that reports a size larger than what
it actually holds.
*/ */
static std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j) static std::size_t calc_bson_string_size(const string_t& value)
{ {
if (JSON_HEDLEY_LIKELY(value_in_range_of<std::int32_t>(value.size())))
{
check_text_utf8(value, j);
}
return sizeof(std::int32_t) + value.size() + 1ul; return sizeof(std::int32_t) + value.size() + 1ul;
} }
@@ -1113,8 +1080,6 @@ class binary_writer
is neither an object nor an array is neither an object nor an array
@throw out_of_range.415 if @a j is binary with a subtype that does not fit @throw out_of_range.415 if @a j is binary with a subtype that does not fit
into a byte, before anything is written into a byte, before anything is written
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a j is a
string that is not valid UTF-8, before anything is written
*/ */
static std::size_t calc_bson_value_size(const BasicJsonType& j) static std::size_t calc_bson_value_size(const BasicJsonType& j)
{ {
@@ -1136,7 +1101,7 @@ class binary_writer
return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned); return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned);
case value_t::string: case value_t::string:
return calc_bson_string_size(*j.m_data.m_value.string, j); return calc_bson_string_size(*j.m_data.m_value.string);
case value_t::null: case value_t::null:
return 0ul; return 0ul;
@@ -1249,8 +1214,6 @@ class binary_writer
written written
@throw out_of_range.415 if a binary value's subtype does not fit into a @throw out_of_range.415 if a binary value's subtype does not fit into a
byte, before anything is written byte, before anything is written
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or a key is not valid UTF-8, before anything is written
*/ */
static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes) static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
{ {
@@ -2129,7 +2092,7 @@ class binary_writer
*/ */
void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context) void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context)
{ {
check_utf8(s, context); check_bon8_utf8(s, context);
// a string that follows another string terminates it // a string that follows another string terminates it
if (string_open) if (string_open)
@@ -2159,7 +2122,7 @@ class binary_writer
@throw type_error.316 if @a s is not valid UTF-8; the message names the @throw type_error.316 if @a s is not valid UTF-8; the message names the
first byte of the first invalid or incomplete sequence first byte of the first invalid or incomplete sequence
*/ */
static void check_utf8(const string_t& s, const BasicJsonType& context) static void check_bon8_utf8(const string_t& s, const BasicJsonType& context)
{ {
static_cast<void>(context); // only used when exceptions are enabled static_cast<void>(context); // only used when exceptions are enabled
const auto* data = reinterpret_cast<const unsigned char*>(s.data()); const auto* data = reinterpret_cast<const unsigned char*>(s.data());
@@ -2170,29 +2133,6 @@ class binary_writer
} }
} }
/*!
@brief check a CBOR, UBJSON, BJData, or BSON text string for valid UTF-8
The check only happens if JSON_STRICT_BINARY_UTF8 is enabled. Otherwise,
the bytes are written unchanged, as before version 3.13.0. MessagePack
always writes the bytes as is, and BON8 always checks them (see
@ref check_utf8).
@param[in] s the string to check
@param[in] context the value that holds @a s (for diagnostics)
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a s is
not valid UTF-8
*/
static void check_text_utf8(const string_t& s, const BasicJsonType& context)
{
#if JSON_STRICT_BINARY_UTF8
check_utf8(s, context);
#else
static_cast<void>(s);
static_cast<void>(context);
#endif
}
/*! /*!
@brief write an integer in the shortest encoding @brief write an integer in the shortest encoding
+38 -6
View File
@@ -117,14 +117,13 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details. http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
The library checks UTF-8 well-formedness (RFC 3629, section 4) in three The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
places, which differ in speed, diagnostics, and how they read the input: places, which differ in speed, diagnostics, and how they read the input:
- decode() below: the serializer, to escape and, in strict mode, reject - decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
ill-formed UTF-8 when dumping a string. The CBOR, MessagePack, BSON, strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
UBJSON and BJData readers do not use it: none of those specs requires a MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
decoder to reject ill-formed UTF-8 in text strings, so the readers keep text strings at decode time).
the bytes as is and leave the check to dump() and the binary writers.
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a - the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error. diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's - validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
@@ -179,5 +178,38 @@ inline std::uint8_t decode(std::uint8_t& state, std::uint32_t& codep, const std:
return state; return state;
} }
/*!
@brief check whether a string consists solely of valid UTF-8
Used by the CBOR/MessagePack/BSON/UBJSON binary readers to reject text
strings that are not valid UTF-8 at decode time (RFC 8949 §3.1 and the
MessagePack/BSON specifications all require text strings to be UTF-8), so
that malformed input is caught immediately instead of only surfacing later
as a type_error.316 when the resulting value is dumped.
@param[in] s the string to check
@param[in] first index of the first byte to check; the bytes before it are
assumed to have been validated already and to end on a
code point boundary
@return whether @a s (from index @a first on) is valid UTF-8
*/
template<typename StringType>
inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noexcept
{
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
for (std::size_t i = first; i < s.size(); ++i)
{
decode(state, codepoint, static_cast<std::uint8_t>(s[i]));
if (state == UTF8_REJECT)
{
return false;
}
}
return state == UTF8_ACCEPT;
}
} // namespace detail } // namespace detail
NLOHMANN_JSON_NAMESPACE_END NLOHMANN_JSON_NAMESPACE_END
+14 -14
View File
@@ -1894,8 +1894,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
if (is_an_object) if (is_an_object)
{ {
// the initializer list is a list of pairs -> create an object // the initializer list is a list of pairs -> create an object
m_data.m_type = value_t::object;
m_data.m_value = value_t::object; m_data.m_value = value_t::object;
m_data.m_type = value_t::object;
for (auto& element_ref : init) for (auto& element_ref : init)
{ {
@@ -1917,8 +1917,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
} }
#endif #endif
// the initializer list describes an array -> create an array // the initializer list describes an array -> create an array
m_data.m_type = value_t::array;
m_data.m_value.array = create<array_t>(init.begin(), init.end()); m_data.m_value.array = create<array_t>(init.begin(), init.end());
m_data.m_type = value_t::array;
} }
set_parents(); set_parents();
@@ -1931,8 +1931,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(const typename binary_t::container_type& init) static basic_json binary(const typename binary_t::container_type& init)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = init; res.m_data.m_value = init;
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -1942,8 +1942,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(const typename binary_t::container_type& init, typename binary_t::subtype_type subtype) static basic_json binary(const typename binary_t::container_type& init, typename binary_t::subtype_type subtype)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = binary_t(init, subtype); res.m_data.m_value = binary_t(init, subtype);
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -1953,8 +1953,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(typename binary_t::container_type&& init) static basic_json binary(typename binary_t::container_type&& init)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = std::move(init); res.m_data.m_value = std::move(init);
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -1964,8 +1964,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(typename binary_t::container_type&& init, typename binary_t::subtype_type subtype) static basic_json binary(typename binary_t::container_type&& init, typename binary_t::subtype_type subtype)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = binary_t(std::move(init), subtype); res.m_data.m_value = binary_t(std::move(init), subtype);
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -3011,8 +3011,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// implicitly convert a null value to an empty array // implicitly convert a null value to an empty array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value.array = create<array_t>(); m_data.m_value.array = create<array_t>();
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -3078,8 +3078,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// implicitly convert a null value to an empty object // implicitly convert a null value to an empty object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value.object = create<object_t>(); m_data.m_value.object = create<object_t>();
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
@@ -3131,8 +3131,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// implicitly convert a null value to an empty object // implicitly convert a null value to an empty object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value.object = create<object_t>(); m_data.m_value.object = create<object_t>();
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
@@ -4082,8 +4082,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an array // transform a null object into an array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value = value_t::array; m_data.m_value = value_t::array;
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -4115,8 +4115,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an array // transform a null object into an array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value = value_t::array; m_data.m_value = value_t::array;
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -4147,8 +4147,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an object // transform a null object into an object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value = value_t::object; m_data.m_value = value_t::object;
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
@@ -4203,8 +4203,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an array // transform a null object into an array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value = value_t::array; m_data.m_value = value_t::array;
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -4228,8 +4228,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an object // transform a null object into an object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value = value_t::object; m_data.m_value = value_t::object;
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
+128 -133
View File
@@ -104,10 +104,6 @@
#define JSON_STRICT_NUL_HANDLING 0 #define JSON_STRICT_NUL_HANDLING 0
#endif #endif
#ifndef JSON_STRICT_BINARY_UTF8
#define JSON_STRICT_BINARY_UTF8 0
#endif
#if JSON_DIAGNOSTICS #if JSON_DIAGNOSTICS
#define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag #define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag
#else #else
@@ -144,20 +140,14 @@
#define NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING #define NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING
#endif #endif
#if JSON_STRICT_BINARY_UTF8
#define NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8 _sbu8
#else
#define NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8
#endif
#ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION #ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION
#define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0 #define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0
#endif #endif
// Construct the namespace ABI tags component // Construct the namespace ABI tags component
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f, g) json_abi ## a ## b ## c ## d ## e ## f ## g #define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f) json_abi ## a ## b ## c ## d ## e ## f
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e, f, g) \ #define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e, f) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f, g) NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f)
#define NLOHMANN_JSON_ABI_TAGS \ #define NLOHMANN_JSON_ABI_TAGS \
NLOHMANN_JSON_ABI_TAGS_CONCAT( \ NLOHMANN_JSON_ABI_TAGS_CONCAT( \
@@ -166,8 +156,7 @@
NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \ NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \ NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \
NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION, \ NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION, \
NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING, \ NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING)
NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8)
// Construct the namespace version component // Construct the namespace version component
#define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \ #define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \
@@ -6332,14 +6321,13 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details. http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
The library checks UTF-8 well-formedness (RFC 3629, section 4) in three The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
places, which differ in speed, diagnostics, and how they read the input: places, which differ in speed, diagnostics, and how they read the input:
- decode() below: the serializer, to escape and, in strict mode, reject - decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
ill-formed UTF-8 when dumping a string. The CBOR, MessagePack, BSON, strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
UBJSON and BJData readers do not use it: none of those specs requires a MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
decoder to reject ill-formed UTF-8 in text strings, so the readers keep text strings at decode time).
the bytes as is and leave the check to dump() and the binary writers.
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a - the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error. diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's - validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
@@ -6394,6 +6382,39 @@ inline std::uint8_t decode(std::uint8_t& state, std::uint32_t& codep, const std:
return state; return state;
} }
/*!
@brief check whether a string consists solely of valid UTF-8
Used by the CBOR/MessagePack/BSON/UBJSON binary readers to reject text
strings that are not valid UTF-8 at decode time (RFC 8949 §3.1 and the
MessagePack/BSON specifications all require text strings to be UTF-8), so
that malformed input is caught immediately instead of only surfacing later
as a type_error.316 when the resulting value is dumped.
@param[in] s the string to check
@param[in] first index of the first byte to check; the bytes before it are
assumed to have been validated already and to end on a
code point boundary
@return whether @a s (from index @a first on) is valid UTF-8
*/
template<typename StringType>
inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noexcept
{
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
for (std::size_t i = first; i < s.size(); ++i)
{
decode(state, codepoint, static_cast<std::uint8_t>(s[i]));
if (state == UTF8_REJECT)
{
return false;
}
}
return state == UTF8_ACCEPT;
}
} // namespace detail } // namespace detail
NLOHMANN_JSON_NAMESPACE_END NLOHMANN_JSON_NAMESPACE_END
@@ -6634,6 +6655,10 @@ namespace detail
* j.m_data.m_value.destroy(j.m_data.m_type) to avoid a memory leak in case j contains an * j.m_data.m_value.destroy(j.m_data.m_type) to avoid a memory leak in case j contains an
* allocated value (e.g., a string). See bug issue * allocated value (e.g., a string). See bug issue
* https://github.com/nlohmann/json/issues/2865 for more information. * https://github.com/nlohmann/json/issues/2865 for more information.
*
* A value that has to be allocated is created before the old one is destroyed:
* were it the other way around, an exception while creating the new value would
* leave j with the type of the new value, but the pointer to the destroyed old one.
*/ */
template<value_t> struct external_constructor; template<value_t> struct external_constructor;
@@ -6657,18 +6682,20 @@ struct external_constructor<value_t::string>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::string_t& s) static void construct(BasicJsonType& j, const typename BasicJsonType::string_t& s)
{ {
const typename BasicJsonType::json_value value(s);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::string; j.m_data.m_type = value_t::string;
j.m_data.m_value = s; j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::string_t&& s) static void construct(BasicJsonType& j, typename BasicJsonType::string_t&& s)
{ {
const typename BasicJsonType::json_value value(std::move(s));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::string; j.m_data.m_type = value_t::string;
j.m_data.m_value = std::move(s); j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
@@ -6677,9 +6704,10 @@ struct external_constructor<value_t::string>
int > = 0 > int > = 0 >
static void construct(BasicJsonType& j, const CompatibleStringType& str) static void construct(BasicJsonType& j, const CompatibleStringType& str)
{ {
auto* created = j.template create<typename BasicJsonType::string_t>(str);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::string; j.m_data.m_type = value_t::string;
j.m_data.m_value.string = j.template create<typename BasicJsonType::string_t>(str); j.m_data.m_value.string = created;
j.assert_invariant(); j.assert_invariant();
} }
}; };
@@ -6690,18 +6718,20 @@ struct external_constructor<value_t::binary>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::binary_t& b) static void construct(BasicJsonType& j, const typename BasicJsonType::binary_t& b)
{ {
const typename BasicJsonType::json_value value(b);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::binary; j.m_data.m_type = value_t::binary;
j.m_data.m_value = typename BasicJsonType::binary_t(b); j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::binary_t&& b) static void construct(BasicJsonType& j, typename BasicJsonType::binary_t&& b)
{ {
const typename BasicJsonType::json_value value(std::move(b));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::binary; j.m_data.m_type = value_t::binary;
j.m_data.m_value = typename BasicJsonType::binary_t(std::move(b)); j.m_data.m_value = value;
j.assert_invariant(); j.assert_invariant();
} }
}; };
@@ -6751,9 +6781,10 @@ struct external_constructor<value_t::array>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::array_t& arr) static void construct(BasicJsonType& j, const typename BasicJsonType::array_t& arr)
{ {
const typename BasicJsonType::json_value value(arr);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value = arr; j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -6761,9 +6792,10 @@ struct external_constructor<value_t::array>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::array_t&& arr) static void construct(BasicJsonType& j, typename BasicJsonType::array_t&& arr)
{ {
const typename BasicJsonType::json_value value(std::move(arr));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value = std::move(arr); j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -6779,9 +6811,10 @@ struct external_constructor<value_t::array>
using std::begin; using std::begin;
using std::end; using std::end;
auto* created = j.template create<typename BasicJsonType::array_t>(begin(arr), end(arr));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value.array = j.template create<typename BasicJsonType::array_t>(begin(arr), end(arr)); j.m_data.m_value.array = created;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -6789,15 +6822,17 @@ struct external_constructor<value_t::array>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const std::vector<bool>& arr) static void construct(BasicJsonType& j, const std::vector<bool>& arr)
{ {
j.m_data.m_value.destroy(j.m_data.m_type); typename BasicJsonType::array_t elements;
j.m_data.m_type = value_t::array; elements.reserve(arr.size());
j.m_data.m_value = value_t::array;
j.m_data.m_value.array->reserve(arr.size());
for (const bool x : arr) for (const bool x : arr)
{ {
j.m_data.m_value.array->push_back(x); elements.push_back(x);
j.set_parent(j.m_data.m_value.array->back());
} }
const typename BasicJsonType::json_value value(std::move(elements));
j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array;
j.m_data.m_value = value;
j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -6805,11 +6840,12 @@ struct external_constructor<value_t::array>
enable_if_t<std::is_convertible<T, BasicJsonType>::value, int> = 0> enable_if_t<std::is_convertible<T, BasicJsonType>::value, int> = 0>
static void construct(BasicJsonType& j, const std::valarray<T>& arr) static void construct(BasicJsonType& j, const std::valarray<T>& arr)
{ {
typename BasicJsonType::array_t elements(arr.size());
std::copy(std::begin(arr), std::end(arr), elements.begin());
const typename BasicJsonType::json_value value(std::move(elements));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array; j.m_data.m_type = value_t::array;
j.m_data.m_value = value_t::array; j.m_data.m_value = value;
j.m_data.m_value.array->resize(arr.size());
std::copy(std::begin(arr), std::end(arr), j.m_data.m_value.array->begin());
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -6821,14 +6857,16 @@ struct external_constructor<value_t::array>
enable_if_t<is_compatible_range_view<std::remove_cvref_t<CompatibleArrayType>>::value, int> = 0> enable_if_t<is_compatible_range_view<std::remove_cvref_t<CompatibleArrayType>>::value, int> = 0>
static void construct(BasicJsonType& j, CompatibleArrayType && arr) static void construct(BasicJsonType& j, CompatibleArrayType && arr)
{ {
j.m_data.m_value.destroy(j.m_data.m_type); typename BasicJsonType::array_t elements;
j.m_data.m_type = value_t::array;
j.m_data.m_value = value_t::array;
for (auto&& x : std::forward<CompatibleArrayType>(arr)) for (auto&& x : std::forward<CompatibleArrayType>(arr))
{ {
j.m_data.m_value.array->push_back(x); elements.push_back(x);
j.set_parent(j.m_data.m_value.array->back());
} }
const typename BasicJsonType::json_value value(std::move(elements));
j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::array;
j.m_data.m_value = value;
j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
#endif #endif
@@ -6840,9 +6878,10 @@ struct external_constructor<value_t::object>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, const typename BasicJsonType::object_t& obj) static void construct(BasicJsonType& j, const typename BasicJsonType::object_t& obj)
{ {
const typename BasicJsonType::json_value value(obj);
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::object; j.m_data.m_type = value_t::object;
j.m_data.m_value = obj; j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -6850,9 +6889,10 @@ struct external_constructor<value_t::object>
template<typename BasicJsonType> template<typename BasicJsonType>
static void construct(BasicJsonType& j, typename BasicJsonType::object_t&& obj) static void construct(BasicJsonType& j, typename BasicJsonType::object_t&& obj)
{ {
const typename BasicJsonType::json_value value(std::move(obj));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::object; j.m_data.m_type = value_t::object;
j.m_data.m_value = std::move(obj); j.m_data.m_value = value;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -6864,9 +6904,10 @@ struct external_constructor<value_t::object>
using std::begin; using std::begin;
using std::end; using std::end;
auto* created = j.template create<typename BasicJsonType::object_t>(begin(obj), end(obj));
j.m_data.m_value.destroy(j.m_data.m_type); j.m_data.m_value.destroy(j.m_data.m_type);
j.m_data.m_type = value_t::object; j.m_data.m_type = value_t::object;
j.m_data.m_value.object = j.template create<typename BasicJsonType::object_t>(begin(obj), end(obj)); j.m_data.m_value.object = created;
j.set_parents(); j.set_parents();
j.assert_invariant(); j.assert_invariant();
} }
@@ -17523,13 +17564,28 @@ class binary_reader
const NumberType len, const NumberType len,
string_t& result) string_t& result)
{ {
// Strings are taken as is: none of CBOR (RFC 8949 §3.1 leaves the // get_bytes() appends to result, and CBOR indefinite-length strings
// choice to the decoder), MessagePack (whose spec explicitly allows // collect all their chunks in the same result; validating only the
// a str object to contain an invalid byte sequence), UBJSON, BJData, // newly read bytes keeps the check linear in the input size
// or BSON requires a decoder to reject ill-formed UTF-8. The bytes const std::size_t old_size = result.size();
// are kept unchanged; dump() and the binary writers are the ones if (JSON_HEDLEY_UNLIKELY(!get_bytes(format, len, "string", result)))
// that check them and report type_error.316 if they are not valid. {
return get_bytes(format, len, "string", result); return false;
}
// RFC 8949 (CBOR) §3.1 and the MessagePack/BSON/UBJSON specifications
// all require text strings to be valid UTF-8; reject anything else
// right here so malformed input is caught at decode time instead of
// only surfacing later as a type_error.316 when the value is dumped
// (which would defeat allow_exceptions=false / strict discarding).
if (JSON_HEDLEY_UNLIKELY(!is_valid_utf8(result, old_size)))
{
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(113, chars_read,
exception_message(format, "invalid string: ill-formed UTF-8 byte", "string"), nullptr));
}
return true;
} }
/*! /*!
@@ -21204,8 +21260,6 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or an object key is not valid UTF-8
@throw type_error.317 if @a j is not an object @throw type_error.317 if @a j is not an object
*/ */
void write_bson(const BasicJsonType& j) void write_bson(const BasicJsonType& j)
@@ -21236,8 +21290,6 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or an object key is not valid UTF-8
*/ */
void write_cbor(const BasicJsonType& j) void write_cbor(const BasicJsonType& j)
{ {
@@ -21304,8 +21356,6 @@ class binary_writer
case value_t::string: case value_t::string:
{ {
check_text_utf8(*j.m_data.m_value.string, j);
// step 1: write control byte and the string length // step 1: write control byte and the string length
write_cbor_head(0x60, j.m_data.m_value.string->size()); write_cbor_head(0x60, j.m_data.m_value.string->size());
@@ -21382,11 +21432,6 @@ class binary_writer
// step 2: write each element // step 2: write each element
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
// el.first is checked here, against the object as
// diagnostics context, because write_cbor(el.first)
// converts it to a temporary basic_json that would be
// used as the context instead
check_text_utf8(el.first, j);
write_cbor(el.first); write_cbor(el.first);
write_cbor(el.second); write_cbor(el.second);
} }
@@ -21729,8 +21774,6 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value @param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false @param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2 @param[in] bjdata_version which BJData version to use, default is draft2
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or an object key is not valid UTF-8
*/ */
void write_ubjson(const BasicJsonType& j, const bool use_count, void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true, const bool use_type, const bool add_prefix = true,
@@ -21780,8 +21823,6 @@ class binary_writer
case value_t::string: case value_t::string:
{ {
check_text_utf8(*j.m_data.m_value.string, j);
if (add_prefix) if (add_prefix)
{ {
oa.write_character(to_char_type('S')); oa.write_character(to_char_type('S'));
@@ -21944,7 +21985,6 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
check_text_utf8(el.first, j);
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata); write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
oa.write_characters( oa.write_characters(
reinterpret_cast<const CharType*>(el.first.data()), reinterpret_cast<const CharType*>(el.first.data()),
@@ -21989,10 +22029,6 @@ class binary_writer
/*! /*!
@return The size of a BSON document entry header, including the id marker @return The size of a BSON document entry header, including the id marker
and the entry name size (and its null-terminator). and the entry name size (and its null-terminator).
@throw out_of_range.409 if @a name contains U+0000, before anything is
written
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a name is
not valid UTF-8, before anything is written
*/ */
static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j) static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
{ {
@@ -22002,8 +22038,7 @@ class binary_writer
JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j)); JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j));
} }
check_text_utf8(name, j); static_cast<void>(j);
return /*id*/ 1ul + name.size() + /*zero-terminator*/1u; return /*id*/ 1ul + name.size() + /*zero-terminator*/1u;
} }
@@ -22059,21 +22094,9 @@ class binary_writer
/*! /*!
@return The size of the BSON-encoded string in @a value @return The size of the BSON-encoded string in @a value
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a value
is not valid UTF-8, before anything is written
@note The UTF-8 check is skipped if @a value is already too long for the
32-bit BSON length field (@ref to_bson_length rejects it later, once
the size of the whole document is known); this also keeps the check
from reading past a StringType that reports a size larger than what
it actually holds.
*/ */
static std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j) static std::size_t calc_bson_string_size(const string_t& value)
{ {
if (JSON_HEDLEY_LIKELY(value_in_range_of<std::int32_t>(value.size())))
{
check_text_utf8(value, j);
}
return sizeof(std::int32_t) + value.size() + 1ul; return sizeof(std::int32_t) + value.size() + 1ul;
} }
@@ -22202,8 +22225,6 @@ class binary_writer
is neither an object nor an array is neither an object nor an array
@throw out_of_range.415 if @a j is binary with a subtype that does not fit @throw out_of_range.415 if @a j is binary with a subtype that does not fit
into a byte, before anything is written into a byte, before anything is written
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a j is a
string that is not valid UTF-8, before anything is written
*/ */
static std::size_t calc_bson_value_size(const BasicJsonType& j) static std::size_t calc_bson_value_size(const BasicJsonType& j)
{ {
@@ -22225,7 +22246,7 @@ class binary_writer
return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned); return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned);
case value_t::string: case value_t::string:
return calc_bson_string_size(*j.m_data.m_value.string, j); return calc_bson_string_size(*j.m_data.m_value.string);
case value_t::null: case value_t::null:
return 0ul; return 0ul;
@@ -22338,8 +22359,6 @@ class binary_writer
written written
@throw out_of_range.415 if a binary value's subtype does not fit into a @throw out_of_range.415 if a binary value's subtype does not fit into a
byte, before anything is written byte, before anything is written
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and a string
value or a key is not valid UTF-8, before anything is written
*/ */
static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes) static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
{ {
@@ -23218,7 +23237,7 @@ class binary_writer
*/ */
void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context) void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context)
{ {
check_utf8(s, context); check_bon8_utf8(s, context);
// a string that follows another string terminates it // a string that follows another string terminates it
if (string_open) if (string_open)
@@ -23248,7 +23267,7 @@ class binary_writer
@throw type_error.316 if @a s is not valid UTF-8; the message names the @throw type_error.316 if @a s is not valid UTF-8; the message names the
first byte of the first invalid or incomplete sequence first byte of the first invalid or incomplete sequence
*/ */
static void check_utf8(const string_t& s, const BasicJsonType& context) static void check_bon8_utf8(const string_t& s, const BasicJsonType& context)
{ {
static_cast<void>(context); // only used when exceptions are enabled static_cast<void>(context); // only used when exceptions are enabled
const auto* data = reinterpret_cast<const unsigned char*>(s.data()); const auto* data = reinterpret_cast<const unsigned char*>(s.data());
@@ -23259,29 +23278,6 @@ class binary_writer
} }
} }
/*!
@brief check a CBOR, UBJSON, BJData, or BSON text string for valid UTF-8
The check only happens if JSON_STRICT_BINARY_UTF8 is enabled. Otherwise,
the bytes are written unchanged, as before version 3.13.0. MessagePack
always writes the bytes as is, and BON8 always checks them (see
@ref check_utf8).
@param[in] s the string to check
@param[in] context the value that holds @a s (for diagnostics)
@throw type_error.316 if JSON_STRICT_BINARY_UTF8 is enabled and @a s is
not valid UTF-8
*/
static void check_text_utf8(const string_t& s, const BasicJsonType& context)
{
#if JSON_STRICT_BINARY_UTF8
check_utf8(s, context);
#else
static_cast<void>(s);
static_cast<void>(context);
#endif
}
/*! /*!
@brief write an integer in the shortest encoding @brief write an integer in the shortest encoding
@@ -28596,8 +28592,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
if (is_an_object) if (is_an_object)
{ {
// the initializer list is a list of pairs -> create an object // the initializer list is a list of pairs -> create an object
m_data.m_type = value_t::object;
m_data.m_value = value_t::object; m_data.m_value = value_t::object;
m_data.m_type = value_t::object;
for (auto& element_ref : init) for (auto& element_ref : init)
{ {
@@ -28619,8 +28615,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
} }
#endif #endif
// the initializer list describes an array -> create an array // the initializer list describes an array -> create an array
m_data.m_type = value_t::array;
m_data.m_value.array = create<array_t>(init.begin(), init.end()); m_data.m_value.array = create<array_t>(init.begin(), init.end());
m_data.m_type = value_t::array;
} }
set_parents(); set_parents();
@@ -28633,8 +28629,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(const typename binary_t::container_type& init) static basic_json binary(const typename binary_t::container_type& init)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = init; res.m_data.m_value = init;
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -28644,8 +28640,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(const typename binary_t::container_type& init, typename binary_t::subtype_type subtype) static basic_json binary(const typename binary_t::container_type& init, typename binary_t::subtype_type subtype)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = binary_t(init, subtype); res.m_data.m_value = binary_t(init, subtype);
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -28655,8 +28651,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(typename binary_t::container_type&& init) static basic_json binary(typename binary_t::container_type&& init)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = std::move(init); res.m_data.m_value = std::move(init);
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -28666,8 +28662,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json binary(typename binary_t::container_type&& init, typename binary_t::subtype_type subtype) static basic_json binary(typename binary_t::container_type&& init, typename binary_t::subtype_type subtype)
{ {
auto res = basic_json(); auto res = basic_json();
res.m_data.m_type = value_t::binary;
res.m_data.m_value = binary_t(std::move(init), subtype); res.m_data.m_value = binary_t(std::move(init), subtype);
res.m_data.m_type = value_t::binary;
return res; return res;
} }
@@ -29713,8 +29709,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// implicitly convert a null value to an empty array // implicitly convert a null value to an empty array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value.array = create<array_t>(); m_data.m_value.array = create<array_t>();
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -29780,8 +29776,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// implicitly convert a null value to an empty object // implicitly convert a null value to an empty object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value.object = create<object_t>(); m_data.m_value.object = create<object_t>();
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
@@ -29833,8 +29829,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// implicitly convert a null value to an empty object // implicitly convert a null value to an empty object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value.object = create<object_t>(); m_data.m_value.object = create<object_t>();
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
@@ -30784,8 +30780,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an array // transform a null object into an array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value = value_t::array; m_data.m_value = value_t::array;
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -30817,8 +30813,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an array // transform a null object into an array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value = value_t::array; m_data.m_value = value_t::array;
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -30849,8 +30845,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an object // transform a null object into an object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value = value_t::object; m_data.m_value = value_t::object;
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
@@ -30905,8 +30901,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an array // transform a null object into an array
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::array;
m_data.m_value = value_t::array; m_data.m_value = value_t::array;
m_data.m_type = value_t::array;
assert_invariant(); assert_invariant();
} }
@@ -30930,8 +30926,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// transform a null object into an object // transform a null object into an object
if (is_null()) if (is_null())
{ {
m_data.m_type = value_t::object;
m_data.m_value = value_t::object; m_data.m_value = value_t::object;
m_data.m_type = value_t::object;
assert_invariant(); assert_invariant();
} }
@@ -33847,7 +33843,6 @@ struct formatter<nlohmann::NLOHMANN_BASIC_JSON_TPL, char> // NOLINT(cert-dcl58-c
#undef JSON_BRACE_INIT_COPY_SEMANTICS #undef JSON_BRACE_INIT_COPY_SEMANTICS
#undef JSON_PRECISE_STREAM_POSITION #undef JSON_PRECISE_STREAM_POSITION
#undef JSON_STRICT_NUL_HANDLING #undef JSON_STRICT_NUL_HANDLING
#undef JSON_STRICT_BINARY_UTF8
#endif #endif
// #include <nlohmann/thirdparty/hedley/hedley_undef.hpp> // #include <nlohmann/thirdparty/hedley/hedley_undef.hpp>
+4 -15
View File
@@ -63,10 +63,6 @@
#define JSON_STRICT_NUL_HANDLING 0 #define JSON_STRICT_NUL_HANDLING 0
#endif #endif
#ifndef JSON_STRICT_BINARY_UTF8
#define JSON_STRICT_BINARY_UTF8 0
#endif
#if JSON_DIAGNOSTICS #if JSON_DIAGNOSTICS
#define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag #define NLOHMANN_JSON_ABI_TAG_DIAGNOSTICS _diag
#else #else
@@ -103,20 +99,14 @@
#define NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING #define NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING
#endif #endif
#if JSON_STRICT_BINARY_UTF8
#define NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8 _sbu8
#else
#define NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8
#endif
#ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION #ifndef NLOHMANN_JSON_NAMESPACE_NO_VERSION
#define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0 #define NLOHMANN_JSON_NAMESPACE_NO_VERSION 0
#endif #endif
// Construct the namespace ABI tags component // Construct the namespace ABI tags component
#define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f, g) json_abi ## a ## b ## c ## d ## e ## f ## g #define NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f) json_abi ## a ## b ## c ## d ## e ## f
#define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e, f, g) \ #define NLOHMANN_JSON_ABI_TAGS_CONCAT(a, b, c, d, e, f) \
NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f, g) NLOHMANN_JSON_ABI_TAGS_CONCAT_EX(a, b, c, d, e, f)
#define NLOHMANN_JSON_ABI_TAGS \ #define NLOHMANN_JSON_ABI_TAGS \
NLOHMANN_JSON_ABI_TAGS_CONCAT( \ NLOHMANN_JSON_ABI_TAGS_CONCAT( \
@@ -125,8 +115,7 @@
NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \ NLOHMANN_JSON_ABI_TAG_DIAGNOSTIC_POSITIONS, \
NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \ NLOHMANN_JSON_ABI_TAG_BRACE_INIT_COPY_SEMANTICS, \
NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION, \ NLOHMANN_JSON_ABI_TAG_PRECISE_STREAM_POSITION, \
NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING, \ NLOHMANN_JSON_ABI_TAG_STRICT_NUL_HANDLING)
NLOHMANN_JSON_ABI_TAG_STRICT_BINARY_UTF8)
// Construct the namespace version component // Construct the namespace version component
#define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \ #define NLOHMANN_JSON_NAMESPACE_VERSION_CONCAT_EX(major, minor, patch) \
+5
View File
@@ -139,6 +139,11 @@ json_test_set_test_options(test-unicode4 TEST_PROPERTIES TIMEOUT 3000)
# only the #972 regression test needs thirdparty/fifo_map on its include path # only the #972 regression test needs thirdparty/fifo_map on its include path
json_test_set_test_options(test-regression1 LINK_LIBRARIES fifo_map_include) json_test_set_test_options(test-regression1 LINK_LIBRARIES fifo_map_include)
# GCC's false -Warray-bounds error with JSON_DIAGNOSTICS only shows up when optimizing (#5742)
json_test_set_test_options(test-diagnostics-optimized
COMPILE_OPTIONS $<$<CXX_COMPILER_ID:GNU>:-O3 -Werror=array-bounds>
)
############################################################################# #############################################################################
# add unit tests # add unit tests
############################################################################# #############################################################################
-4
View File
@@ -44,10 +44,6 @@ TEST_CASE("default namespace")
expected += "_snul"; expected += "_snul";
#endif #endif
#if JSON_STRICT_BINARY_UTF8
expected += "_sbu8";
#endif
expected += "_v" STRINGIZE(NLOHMANN_JSON_VERSION_MAJOR); expected += "_v" STRINGIZE(NLOHMANN_JSON_VERSION_MAJOR);
expected += "_" STRINGIZE(NLOHMANN_JSON_VERSION_MINOR); expected += "_" STRINGIZE(NLOHMANN_JSON_VERSION_MINOR);
expected += "_" STRINGIZE(NLOHMANN_JSON_VERSION_PATCH) "::basic_json"; expected += "_" STRINGIZE(NLOHMANN_JSON_VERSION_PATCH) "::basic_json";
-4
View File
@@ -45,10 +45,6 @@ TEST_CASE("default namespace without version component")
expected += "_snul"; expected += "_snul";
#endif #endif
#if JSON_STRICT_BINARY_UTF8
expected += "_sbu8";
#endif
expected += "::basic_json"; expected += "::basic_json";
// fallback for Clang // fallback for Clang
+175
View File
@@ -12,6 +12,11 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <valarray>
#if JSON_HAS_RANGES
#include <ranges>
#endif
namespace namespace
{ {
// special test case to check if memory is leaked if constructor throws // special test case to check if memory is leaked if constructor throws
@@ -592,3 +597,173 @@ TEST_CASE("bad my_allocator::construct")
j["test"].push_back("should not leak"); j["test"].push_back("should not leak");
} }
} }
// the no-exceptions CI job skips every CHECK_THROWS_AS, which would leave
// next_construct_fails set for the next allocation outside a check
#if !defined(JSON_NOEXCEPTION)
TEST_CASE("a failed allocation leaves the value unchanged")
{
// create JSON type using the throwing allocator
using my_json = nlohmann::basic_json<std::map,
std::vector,
std::string,
bool,
std::int64_t,
std::uint64_t,
double,
my_allocator>;
// Each of these creates a string, array, object, or binary value. The
// value must be created before the type is changed: otherwise, a failed
// creation left a value of the new type without anything behind it (an
// assertion in its destructor, a null pointer everywhere else) or, when
// an old value was destroyed first, with a pointer to that destroyed one.
SECTION("creating a binary value")
{
const std::vector<std::uint8_t> bytes = {1, 2, 3};
my_json _;
next_construct_fails = true;
CHECK_THROWS_AS(_ = my_json::binary(bytes), std::bad_alloc&);
next_construct_fails = true;
CHECK_THROWS_AS(_ = my_json::binary(bytes, 42), std::bad_alloc&);
next_construct_fails = true;
CHECK_THROWS_AS(_ = my_json::binary(std::vector<std::uint8_t>(bytes)), std::bad_alloc&);
next_construct_fails = true;
CHECK_THROWS_AS(_ = my_json::binary(std::vector<std::uint8_t>(bytes), 42), std::bad_alloc&);
next_construct_fails = false;
}
SECTION("turning a null value into an array or object")
{
my_json j;
next_construct_fails = true;
CHECK_THROWS_AS(j[0], std::bad_alloc&);
CHECK(j.is_null());
next_construct_fails = true;
CHECK_THROWS_AS(j["key"], std::bad_alloc&);
CHECK(j.is_null());
#ifdef JSON_HAS_CPP_17
next_construct_fails = true;
CHECK_THROWS_AS(j[std::string_view("key")], std::bad_alloc&);
CHECK(j.is_null());
#endif
next_construct_fails = true;
CHECK_THROWS_AS(j.push_back(my_json(1)), std::bad_alloc&);
CHECK(j.is_null());
const my_json one = 1;
next_construct_fails = true;
CHECK_THROWS_AS(j.push_back(one), std::bad_alloc&);
CHECK(j.is_null());
next_construct_fails = true;
CHECK_THROWS_AS(j.push_back(my_json::object_t::value_type("key", 1)), std::bad_alloc&);
CHECK(j.is_null());
next_construct_fails = true;
CHECK_THROWS_AS(j.emplace_back(1), std::bad_alloc&);
CHECK(j.is_null());
next_construct_fails = true;
CHECK_THROWS_AS(j.emplace("key", 1), std::bad_alloc&);
CHECK(j.is_null());
const my_json object = {{"key", 1}};
next_construct_fails = true;
CHECK_THROWS_AS(j.update(object), std::bad_alloc&);
CHECK(j.is_null());
next_construct_fails = false;
}
// With iterator debugging, VS 2015's containers construct a proxy with the
// allocator in constructors that cannot report its failure, so a failing
// allocator crashes this section there (SIGSEGV with VS 2015 Debug x86).
#if !(defined(_MSC_VER) && _MSC_VER < 1910 && defined(_ITERATOR_DEBUG_LEVEL) && _ITERATOR_DEBUG_LEVEL > 0)
SECTION("converting into an existing value")
{
// to_json replaces the value it is given; the old one must survive a
// failed creation of the new one
my_json j = "old";
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, std::string("new")), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, std::vector<int> {1, 2}), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, std::vector<bool> {true, false}), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, std::map<std::string, int> {{"a", 1}}), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, my_json::binary_t({1, 2})), std::bad_alloc&);
CHECK(j == "old");
// the overloads for lvalues of the value types, for the value types
// themselves, and for the remaining compatible types
const std::string string = "new";
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, string), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, "new"), std::bad_alloc&);
CHECK(j == "old");
// to_json only moves a binary value that it converted from another
// container type, which my_json's std::vector<std::uint8_t> is not
using binary_constructor = nlohmann::detail::external_constructor<nlohmann::detail::value_t::binary>;
next_construct_fails = true;
CHECK_THROWS_AS(binary_constructor::construct(j, my_json::binary_t({1, 2})), std::bad_alloc&);
CHECK(j == "old");
my_json::array_t array = {1, 2};
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, array), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, std::move(array)), std::bad_alloc&);
CHECK(j == "old");
my_json::object_t object = {{"a", 1}};
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, object), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, std::move(object)), std::bad_alloc&);
CHECK(j == "old");
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, std::valarray<int> {1, 2}), std::bad_alloc&);
CHECK(j == "old");
#if JSON_HAS_RANGES && !defined(__MINGW32__)
const std::vector<int> numbers = {1, 2};
next_construct_fails = true;
CHECK_THROWS_AS(nlohmann::to_json(j, numbers | std::views::filter([](int /*unused*/)
{
return true;
})), std::bad_alloc&);
CHECK(j == "old");
#endif
next_construct_fails = false;
nlohmann::to_json(j, std::vector<int> {1, 2});
CHECK(j == my_json({1, 2}));
}
#endif
}
#endif
-110
View File
@@ -1,110 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// The binary writers check strings and object keys for valid UTF-8 only if
// JSON_STRICT_BINARY_UTF8 is enabled (planned to be the default in 4.0.0).
// Without it, they write the bytes unchanged, as before version 3.13.0; the
// tests for that are next to the other tests of each format.
#ifdef JSON_STRICT_BINARY_UTF8
#undef JSON_STRICT_BINARY_UTF8
#endif
#define JSON_STRICT_BINARY_UTF8 1
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <cstdint>
#include <vector>
TEST_CASE("JSON_STRICT_BINARY_UTF8 (see #5529, #5651)")
{
SECTION("CBOR")
{
// a string value with ill-formed UTF-8 is rejected
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_cbor(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// binary values are not text and are unaffected
CHECK_NOTHROW(json::to_cbor(json::binary(std::vector<std::uint8_t>({0xFF}))));
// a value read back from CBOR with ill-formed bytes cannot be written
// back either (the reader is lenient regardless of the macro)
const json j = json::from_cbor(std::vector<std::uint8_t>({0x62, 0xc0, 0xae}));
CHECK_THROWS_WITH_AS(json::to_cbor(j), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
}
SECTION("UBJSON")
{
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_ubjson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("BJData")
{
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_bjdata(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("BSON")
{
// to_bson() rejects the same kind of ill-formed string value, before
// any bytes reach the output adapter (the BSON document length
// prefix must be known up front, so nothing is written incrementally)
std::vector<std::uint8_t> out{0x42}; // a sentinel byte the writer must not touch
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}, nlohmann::detail::output_adapter<std::uint8_t>(out)), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
CHECK(out == std::vector<std::uint8_t> {0x42});
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC3"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xED\xA0\x80"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC0\xAF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected as well; unlike
// the reader (which never validates element names), the writer
// checks both string values and object keys
CHECK_THROWS_WITH_AS(json::to_bson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("MessagePack and BON8 are unaffected")
{
// MessagePack allows any bytes in a str, so to_msgpack() writes them as
// is; BON8 always checks, because the lead bytes mark where strings end
CHECK(json::to_msgpack(json("\xFF")) == std::vector<std::uint8_t>({0xa1, 0xff}));
CHECK_THROWS_AS(json::to_bon8(json("\xFF")), json::type_error&);
}
}
-37
View File
@@ -3907,43 +3907,6 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_bjdata(j) == v); CHECK(json::to_bjdata(j) == v);
CHECK(json::from_bjdata(v) == j); CHECK(json::from_bjdata(v) == j);
} }
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_bjdata() writes
// the bytes unchanged, as before 3.13.0, unless
// JSON_STRICT_BINARY_UTF8 is enabled (see
// unit-binary_utf8_strict.cpp)
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_bjdata(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK(json::from_bjdata(json::to_bjdata(j)) == j);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_bjdata(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_bjdata(json::to_bjdata(j_key)) == j_key);
CHECK(json::from_bjdata(json::to_bjdata(json("\xFF"))) == json("\xFF"));
// a truncated multi-byte sequence
CHECK(json::from_bjdata(json::to_bjdata(json("\xC3"))) == json("\xC3"));
// an encoded surrogate half (U+D800)
CHECK(json::from_bjdata(json::to_bjdata(json("\xED\xA0\x80"))) == json("\xED\xA0\x80"));
// an overlong encoding of '.'
CHECK(json::from_bjdata(json::to_bjdata(json("\xC0\xAF"))) == json("\xC0\xAF"));
// an object key with ill-formed UTF-8 is kept the same way
CHECK(json::from_bjdata(json::to_bjdata(json{{"\xFF", 1}})) == json{{"\xFF", 1}});
}
} }
SECTION("Array Type") SECTION("Array Type")
-37
View File
@@ -154,43 +154,6 @@ TEST_CASE("BSON")
#endif #endif
} }
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// a BSON document {"s": "\xC0\xAE"} (0xC0 0xAE is an overlong
// encoding of '.'); the BSON spec does not require a decoder to
// reject ill-formed UTF-8 in a string value, so the reader hands the
// bytes back unchanged
const std::vector<uint8_t> v =
{
0x0F, 0x00, 0x00, 0x00, // document length
0x02, 's', 0x00, // type 0x02 (string), key "s"
0x03, 0x00, 0x00, 0x00, // string length (including null)
0xc0, 0xae, 0x00, // string content and its null terminator
0x00 // document terminator
};
json j;
CHECK_NOTHROW(j = json::from_bson(v));
REQUIRE(j.is_object());
REQUIRE(j.contains("s"));
CHECK(j["s"].get_ref<const json::string_t&>() == std::string("\xc0\xae"));
// dump() still requires valid UTF-8 and throws for such a value
CHECK_THROWS_AS(j.dump(), json::type_error&);
// to_bson() writes the bytes back unchanged, as before 3.13.0,
// unless JSON_STRICT_BINARY_UTF8 is enabled (see unit-binary_utf8_strict.cpp)
CHECK(json::from_bson(json::to_bson(j)) == j);
CHECK(json::from_bson(json::to_bson(json{{"s", "\xFF"}})) == json{{"s", "\xFF"}});
// a truncated multi-byte sequence
CHECK(json::from_bson(json::to_bson(json{{"s", "\xC3"}})) == json{{"s", "\xC3"}});
// an encoded surrogate half (U+D800)
CHECK(json::from_bson(json::to_bson(json{{"s", "\xED\xA0\x80"}})) == json{{"s", "\xED\xA0\x80"}});
// an overlong encoding of '.'
CHECK(json::from_bson(json::to_bson(json{{"s", "\xC0\xAF"}})) == json{{"s", "\xC0\xAF"}});
// an object key with ill-formed UTF-8 is kept as well
CHECK(json::from_bson(json::to_bson(json{{"\xFF", 1}})) == json{{"\xFF", 1}});
}
SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON") SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON")
{ {
// out_of_range.412 is thrown from a single shared helper // out_of_range.412 is thrown from a single shared helper
+18 -67
View File
@@ -1801,41 +1801,19 @@ TEST_CASE("CBOR")
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
} }
SECTION("ill-formed UTF-8 in string (see #5529, #5651)") SECTION("invalid UTF-8 in string (see #5529)")
{ {
// RFC 8949 §3.1 leaves it up to the decoder whether to reject
// ill-formed UTF-8 in a text string; this library does not, and
// hands the original bytes back unchanged, matching the
// MessagePack reader and the behavior before #5185/#5531 (not in
// any release)
// a two-character text string (major type 3) whose bytes are not // a two-character text string (major type 3) whose bytes are not
// valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') round-trips // valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') must be
// byte for byte as a string value // rejected at decode time, matching every other kind of
const std::vector<uint8_t> ill_formed_value = {0x62, 0xc0, 0xae}; // malformed binary input, rather than only failing later when
json j_value; // the resulting value is dumped
CHECK_NOTHROW(j_value = json::from_cbor(ill_formed_value)); json _;
REQUIRE(j_value.is_string()); CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae")); CHECK(json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae}), true, false).is_discarded());
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// to_cbor() writes the bytes back unchanged, as before 3.13.0,
// unless JSON_STRICT_BINARY_UTF8 is enabled (see unit-binary_utf8_strict.cpp)
CHECK(json::from_cbor(json::to_cbor(j_value)) == j_value);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0xa1, 0x62, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_cbor(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_cbor(json::to_cbor(j_key)) == j_key);
// a CBOR byte string (major type 2) with the very same bytes is // a CBOR byte string (major type 2) with the very same bytes is
// NOT text and must still be accepted as-is // NOT text and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x42, 0xc0, 0xae}))); CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x42, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae}))); CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
@@ -1844,47 +1822,17 @@ TEST_CASE("CBOR")
CHECK(json::from_cbor(json::to_cbor(j)) == j); CHECK(json::from_cbor(json::to_cbor(j)) == j);
} }
SECTION("to_cbor keeps ill-formed UTF-8 (see #5651)") SECTION("invalid UTF-8 in indefinite-length string")
{
// to_cbor() writes the bytes unchanged, as before 3.13.0, unless
// JSON_STRICT_BINARY_UTF8 is enabled (see
// unit-binary_utf8_strict.cpp); from_cbor() reads them back as is
CHECK(json::from_cbor(json::to_cbor(json("\xFF"))) == json("\xFF"));
// a truncated multi-byte sequence
CHECK(json::from_cbor(json::to_cbor(json("\xC3"))) == json("\xC3"));
// an encoded surrogate half (U+D800)
CHECK(json::from_cbor(json::to_cbor(json("\xED\xA0\x80"))) == json("\xED\xA0\x80"));
// an overlong encoding of '.'
CHECK(json::from_cbor(json::to_cbor(json("\xC0\xAF"))) == json("\xC0\xAF"));
// an object key with ill-formed UTF-8 is kept the same way
CHECK(json::from_cbor(json::to_cbor(json{{"\xFF", 1}})) == json{{"\xFF", 1}});
// binary values are not text and are unaffected
CHECK_NOTHROW(json::to_cbor(json::binary(std::vector<std::uint8_t>({0xFF}))));
}
SECTION("ill-formed UTF-8 in indefinite-length string")
{ {
json _; json _;
// the chunks are concatenated as is, without checking that each // every chunk must be valid UTF-8 on its own (RFC 8949, Section
// chunk is valid UTF-8 on its own (RFC 8949, Section 3.2.3), so // 3.2.3), so a code point split across two chunks is rejected
// a code point split across two chunks yields a valid string CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff}))); CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff}), true, false).is_discarded());
CHECK(_ == "\xc3\xa9");
CHECK(_.dump() == "\"\xc3\xa9\"");
// a truncated code point is kept as is // an ill-formed later chunk is rejected after valid ones
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0xff}))); CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})), "[json.exception.parse_error.113] parse error at byte 7: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(_ == "\xc3");
CHECK_THROWS_AS(_.dump(), json::type_error&);
CHECK(json::from_cbor(json::to_cbor(_)) == _);
// an ill-formed later chunk is kept after valid ones
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})));
CHECK(_ == "\xc3\xa9\xc0\xae");
CHECK_THROWS_AS(_.dump(), json::type_error&);
// valid multi-byte chunks are accepted // valid multi-byte chunks are accepted
CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc3, 0xb6, 0xff})) == "\xc3\xa9\xc3\xb6"); CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc3, 0xb6, 0xff})) == "\xc3\xa9\xc3\xb6");
@@ -1892,6 +1840,9 @@ TEST_CASE("CBOR")
SECTION("many chunks in indefinite-length string") SECTION("many chunks in indefinite-length string")
{ {
// only the newly read chunk is validated, not the whole string
// collected so far; validating the latter made this input take
// quadratic time (about ten seconds for 100000 chunks)
constexpr std::size_t chunks = 100000; constexpr std::size_t chunks = 100000;
std::vector<uint8_t> v{0x7f}; std::vector<uint8_t> v{0x7f};
for (std::size_t i = 0; i < chunks; ++i) for (std::size_t i = 0; i < chunks; ++i)
+80
View File
@@ -0,0 +1,80 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Regression test for https://github.com/nlohmann/json/issues/5742: with
// JSON_DIAGNOSTICS, GCC (12 to at least 16) reported a false -Warray-bounds
// error in the inlined set_parents() at -O3. The type of a new string was set
// before the string was allocated, so GCC had to assume that operator new
// could change it again and checked the object branch of set_parents()
// against the string's allocation. Setting the type after creating the value
// avoids this. The warning depends on GCC's inlining decisions, so the
// sections cover two patterns that trigger it on different GCC versions
// (#4819 and #5742).
// On GCC, this file is compiled with -O3 -Werror=array-bounds (see
// tests/CMakeLists.txt), so the test fails to build if the warning returns.
#include "doctest_compatibility.h"
#ifdef JSON_DIAGNOSTICS
#undef JSON_DIAGNOSTICS
#endif
#define JSON_DIAGNOSTICS 1
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <iterator>
#include <utility>
#include <vector>
namespace
{
enum class diag_color
{
red,
green,
blue
};
void to_json(json& j, const diag_color& c)
{
static const std::pair<diag_color, json> m[] = // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
{
{diag_color::red, "r"},
{diag_color::green, "g"},
{diag_color::blue, "b"},
};
const auto* it = std::find_if(std::begin(m), std::end(m), [c](const std::pair<diag_color, json>& p)
{
return p.first == c;
});
j = it->second;
}
} // namespace
TEST_CASE("diagnostics with optimization")
{
SECTION("issue #4819 - object in vector")
{
std::vector<json> jsons{};
jsons.emplace_back(json({{"key", "value"}}));
CHECK(jsons.back()["key"] == "value");
}
SECTION("issue #5742 - string values from a static table")
{
json j = json::array();
j.push_back(diag_color::red);
j.push_back(diag_color::green);
j.push_back(diag_color::blue);
CHECK(j.dump() == R"(["r","g","b"])");
CHECK_THROWS_WITH_AS(j[1].get<int>(), "[json.exception.type_error.302] (/1) type must be number, but is string", json::type_error);
}
}
+8 -28
View File
@@ -1540,39 +1540,19 @@ TEST_CASE("MessagePack")
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
} }
SECTION("ill-formed UTF-8 in string (see #5529, #5651)") SECTION("invalid UTF-8 in string (see #5529)")
{ {
// the MessagePack specification explicitly allows a str object to
// contain a byte sequence that is not valid UTF-8 and expects a
// deserializer to hand the original bytes back unchanged; this
// library follows that, unlike CBOR/UBJSON/BJData/BSON, whose
// specifications require text strings to be valid UTF-8
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8 // a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
// (0xC0 0xAE is an overlong encoding of '.') round-trips byte for // (0xC0 0xAE is an overlong encoding of '.') must be rejected at
// byte as a string value // decode time, matching every other kind of malformed binary
const std::vector<uint8_t> ill_formed_value = {0xa2, 0xc0, 0xae}; // input, rather than only failing later when the resulting
json j_value; // value is dumped
CHECK_NOTHROW(j_value = json::from_msgpack(ill_formed_value)); json _;
REQUIRE(j_value.is_string()); CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing MessagePack string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae")); CHECK(json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae}), true, false).is_discarded());
CHECK(json::from_msgpack(json::to_msgpack(j_value)) == j_value);
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0x81, 0xa2, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_msgpack(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_msgpack(json::to_msgpack(j_key)) == j_key);
// a MessagePack bin8 blob with the very same bytes is NOT text // a MessagePack bin8 blob with the very same bytes is NOT text
// and must still be accepted as-is // and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_msgpack(std::vector<uint8_t>({0xc4, 0x02, 0xc0, 0xae}))); CHECK_NOTHROW(_ = json::from_msgpack(std::vector<uint8_t>({0xc4, 0x02, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae}))); CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
-37
View File
@@ -2505,43 +2505,6 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_ubjson(j) == v); CHECK(json::to_ubjson(j) == v);
CHECK(json::from_ubjson(v) == j); CHECK(json::from_ubjson(v) == j);
} }
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_ubjson() writes
// the bytes unchanged, as before 3.13.0, unless
// JSON_STRICT_BINARY_UTF8 is enabled (see
// unit-binary_utf8_strict.cpp)
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_ubjson(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK(json::from_ubjson(json::to_ubjson(j)) == j);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_ubjson(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_ubjson(json::to_ubjson(j_key)) == j_key);
CHECK(json::from_ubjson(json::to_ubjson(json("\xFF"))) == json("\xFF"));
// a truncated multi-byte sequence
CHECK(json::from_ubjson(json::to_ubjson(json("\xC3"))) == json("\xC3"));
// an encoded surrogate half (U+D800)
CHECK(json::from_ubjson(json::to_ubjson(json("\xED\xA0\x80"))) == json("\xED\xA0\x80"));
// an overlong encoding of '.'
CHECK(json::from_ubjson(json::to_ubjson(json("\xC0\xAF"))) == json("\xC0\xAF"));
// an object key with ill-formed UTF-8 is kept the same way
CHECK(json::from_ubjson(json::to_ubjson(json{{"\xFF", 1}})) == json{{"\xFF", 1}});
}
} }
SECTION("Array Type") SECTION("Array Type")