Add an error_handler parameter for UTF-8 to the binary readers and writers

Adds error_handler_t::keep (invalid UTF-8 sequences are left unchanged) as
a fourth error_handler_t value, and threads an error_handler parameter
through the binary writers and readers:

- to_cbor/to_ubjson/to_bjdata/to_bson gain a trailing error_handler
  parameter (default strict, matching their existing type_error.316
  behavior); keep writes a string value or object key's bytes as is
  instead of throwing, and replace/ignore sanitize it exactly like
  dump() would, including for the BSON length prefix. to_msgpack and
  to_bon8 are unchanged.
- from_cbor/from_msgpack/from_ubjson/from_bjdata/from_bson gain a
  trailing error_handler parameter (default keep, i.e. the lenient
  behavior every binary reader had in 3.12.0 and still has after
  #5741); strict checks every string value and object key and raises
  parse_error.113 for ill-formed UTF-8, honoring allow_exceptions;
  replace/ignore sanitize it like dump() would. from_bon8 is
  unchanged, since UTF-8 lead bytes are structural there.

dump()'s own keep support writes ill-formed bytes as is, even with
ensure_ascii, while still \u-escaping well-formed characters around
them as usual.

The UTF-8 validity check (is_valid_utf8) and the replace/ignore
sanitizing logic (sanitize_utf8) now live in string_utils.hpp, shared
by the serializer and the binary reader/writer; error_handler_t itself
moved to its own header (detail/output/error_handler.hpp) so that
string_utils.hpp does not need to depend on serializer.hpp.

See #5529 and #5741.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-10-01 08:43:15 +02:00
parent 92c729f657
commit fff221c58b
27 changed files with 1601 additions and 308 deletions
+6 -3
View File
@@ -25,10 +25,12 @@ and `ensure_ascii` parameters.
result consists of ASCII characters only.
`error_handler` (in)
: how to react on decoding errors; there are three possible values (see [`error_handler_t`](error_handler_t.md):
: how to react on decoding errors; there are four possible values (see [`error_handler_t`](error_handler_t.md):
`strict` (throws an exception in case a decoding error occurs; default), `replace` (replace invalid UTF-8 sequences
with U+FFFD), and `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped)).
with U+FFFD), `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped), and `keep` (write the ill-formed bytes to the output as is,
without escaping them, even if `ensure_ascii` is `#!cpp true`; the result is then not valid UTF-8, but equals the
input bytes exactly, and well-formed characters around the ill-formed bytes are still escaped as usual)).
## Return value
@@ -94,3 +96,4 @@ Binary values are serialized as an object containing two keys:
- Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0.
- Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0.
- Error handler `keep` added in version 3.13.0.
@@ -4,15 +4,30 @@
enum class error_handler_t {
strict,
replace,
ignore
ignore,
keep
};
```
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
`basic_json` value. Three values are differentiated:
This enumeration is used to choose how to treat ill-formed UTF-8 in a string value or object key:
- [`dump`](dump.md) uses it while serializing a `basic_json` value to text.
- [`to_cbor`](to_cbor.md), [`to_ubjson`](to_ubjson.md), [`to_bjdata`](to_bjdata.md), and [`to_bson`](to_bson.md) use it
while serializing a `basic_json` value to that binary format; none of CBOR, UBJSON, BJData, or BSON requires a
decoder to reject ill-formed UTF-8, so by default (`strict`) the library checks on write instead. `to_msgpack` and
`to_bon8` do not take this parameter: MessagePack's specification explicitly allows a string to contain ill-formed
UTF-8, so `to_msgpack` always passes it through, while BON8 always validates, since UTF-8 lead bytes are structural
to that format.
- [`from_cbor`](from_cbor.md), [`from_msgpack`](from_msgpack.md), [`from_ubjson`](from_ubjson.md),
[`from_bjdata`](from_bjdata.md), and [`from_bson`](from_bson.md) use it while parsing that binary format, to decide
whether to check a string value or object key for well-formed UTF-8 at all; by default (`keep`) they do not, as no
binary reader did before this parameter was added. `from_bon8` does not take this parameter, for the same reason
`to_bon8` does not.
Four values are differentiated:
strict
: throw a `type_error` exception in case of invalid UTF-8
: throw a `type_error`/`parse_error` exception in case of invalid UTF-8
replace
: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER)
@@ -20,6 +35,12 @@ replace
ignore
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
keep
: keep invalid UTF-8 sequences unchanged; only meaningful for the binary formats mentioned above, since [`dump`]
(dump.md) itself must produce text, and `keep` there writes the ill-formed bytes to the output as is, so the
result is then not valid UTF-8 (but still equals the input bytes exactly, including around any well-formed
characters, which are still escaped as usual)
## Examples
??? example
@@ -40,3 +61,5 @@ ignore
## Version history
- Added in version 3.4.0.
- Added `keep`, and made this enumeration apply to the binary readers and writers in addition to `dump`, in version
3.13.0.
+12 -3
View File
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_bjdata(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_bjdata(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the BJData (Binary JData) serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
BJData does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -73,7 +81,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if a parse error occurs
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string could not be parsed
successfully
successfully, or if a string value or object key is not valid UTF-8 and `error_handler` is `strict`
- Throws [out_of_range.408](../../home/exceptions.md#jsonexceptionout_of_range408) if the size of an optimized container
or n-dimensional array cannot be represented by `std::size_t`
@@ -111,3 +119,4 @@ Linear in the size of the input.
- Added in version 3.11.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+13 -2
View File
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_bson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_bson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the BSON (Binary JSON) serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
BSON does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -75,6 +83,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
invalid string or byte array length)
- Throws [`parse_error.114`](../../home/exceptions.md#jsonexceptionparse_error114) if an unsupported BSON record type is
encountered
- Throws [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) if a string value or object key is
not valid UTF-8 and `error_handler` is `strict`
## Complexity
@@ -111,6 +121,7 @@ Linear in the size of the input.
- Added in version 3.4.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+14 -4
View File
@@ -6,14 +6,16 @@ template<typename InputType>
static basic_json from_cbor(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error);
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_cbor(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error);
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the CBOR (Concise Binary Object Representation) serialization format.
@@ -65,6 +67,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: how to treat CBOR tags (optional, `error` by default); see [`cbor_tag_handler_t`](cbor_tag_handler_t.md) for more
information
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
CBOR does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -80,8 +88,9 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were
used in the given input or if the input is not valid CBOR
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of
other types are not supported, as JSON object keys are always strings), or if a string value or object key is not
valid UTF-8 and `error_handler` is `strict`
## Complexity
@@ -121,6 +130,7 @@ Linear in the size of the input.
- Added `tag_handler` parameter in version 3.9.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_msgpack(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_msgpack(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the MessagePack serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
MessagePack's specification explicitly allows ill-formed UTF-8, so checking is opt-in: the default, `keep`, does
not check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -73,8 +81,9 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from
MessagePack were used in the given input or if the input is not valid MessagePack
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of
other types are not supported, as JSON object keys are always strings), or if a string value or object key is not
valid UTF-8 and `error_handler` is `strict`
## Complexity
@@ -113,6 +122,7 @@ Linear in the size of the input.
- Added `allow_exceptions` parameter in version 3.2.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+12 -3
View File
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_ubjson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_ubjson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the UBJSON (Universal Binary JSON) serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
UBJSON does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -73,7 +81,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if a parse error occurs
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string could not be parsed
successfully
successfully, or if a string value or object key is not valid UTF-8 and `error_handler` is `strict`
- Throws [out_of_range.408](../../home/exceptions.md#jsonexceptionout_of_range408) if the size of an optimized container
or n-dimensional array cannot be represented by `std::size_t`
@@ -112,6 +120,7 @@ Linear in the size of the input.
- Added `allow_exceptions` parameter in version 3.2.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+15 -5
View File
@@ -5,15 +5,18 @@
static std::vector<std::uint8_t> to_bjdata(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2);
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_bjdata(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2);
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
static void to_bjdata(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2);
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
```
Serializes a given JSON value `j` to a byte vector using the BJData (Binary JData) serialization format. BJData aims to
@@ -43,6 +46,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: which version of BJData to use (see note on "Binary values" on [BJData](../../features/binary_formats/bjdata.md));
optional, `#!cpp bjdata_version_t::draft2` by default.
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_bjdata` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. BJData serialization as byte vector
@@ -57,7 +66,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -92,4 +101,5 @@ Linear in the size of the JSON value `j`.
- Added in version 3.11.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+14 -4
View File
@@ -2,11 +2,14 @@
```cpp
// (1)
static std::vector<std::uint8_t> to_bson(const basic_json& j);
static std::vector<std::uint8_t> to_bson(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o);
static void to_bson(const basic_json& j, detail::output_adapter<char> o);
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_bson(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict);
```
BSON (Binary JSON) is a binary format in which zero or more ordered key/value pairs are stored as a single entity (a
@@ -25,6 +28,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`o` (in)
: output adapter to write serialization to
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_bson` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. BSON serialization as a byte vector
@@ -47,7 +56,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
exceeds 255, the maximum of the BSON binary subtype; example:
`"subtype 70000 is too large for the BSON binary subtype (max 255)"`
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key is
not valid UTF-8
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -86,3 +95,4 @@ pass before anything is written.
- `out_of_range.415` is now detected before anything is written, like the other exceptions above, since version 3.13.0.
- Throwing `type_error.316` for a string value or object key that is not valid UTF-8, detected before anything is
written, added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+14 -4
View File
@@ -2,11 +2,14 @@
```cpp
// (1)
static std::vector<std::uint8_t> to_cbor(const basic_json& j);
static std::vector<std::uint8_t> to_cbor(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o);
static void to_cbor(const basic_json& j, detail::output_adapter<char> o);
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_cbor(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict);
```
Serializes a given JSON value `j` to a byte vector using the CBOR (Concise Binary Object Representation) serialization
@@ -26,6 +29,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`o` (in)
: output adapter to write serialization to
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_cbor` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. CBOR serialization as a byte vector
@@ -38,7 +47,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
## Exceptions
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -74,3 +83,4 @@ Linear in the size of the JSON value `j`.
- Added in version 2.0.9.
- Compact representation of floating-point numbers added in version 3.8.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+14 -4
View File
@@ -4,13 +4,16 @@
// (1)
static std::vector<std::uint8_t> to_ubjson(const basic_json& j,
const bool use_size = false,
const bool use_type = false);
const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_ubjson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false);
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
static void to_ubjson(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false);
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
```
Serializes a given JSON value `j` to a byte vector using the UBJSON (Universal Binary JSON) serialization format. UBJSON
@@ -36,6 +39,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: whether to add type annotations to container types (must be combined with `#!cpp use_size = true`); optional,
`#!cpp false` by default.
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_ubjson` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. UBJSON serialization as a byte vector
@@ -50,7 +59,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -85,3 +94,4 @@ Linear in the size of the JSON value `j`.
- Added in version 3.1.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
@@ -20,5 +20,6 @@ int main()
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::replace)
<< "\nstring with ignored invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
<< '\n';
<< "\nstring with the invalid byte kept as is (" << j_invalid.dump(-1, ' ', false, json::error_handler_t::keep).size()
<< " bytes, not valid UTF-8 itself)\n";
}
@@ -1,3 +1,4 @@
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
string with replaced invalid characters: "ä�ü"
string with ignored invalid characters: "äü"
string with the invalid byte kept as is (7 bytes, not valid UTF-8 itself)
@@ -216,12 +216,17 @@ The library maps BJData types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in string values and object keys"
BJData strings must use UTF-8 encoding, but this is not enforced on read: `from_bjdata()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
BJData strings must use UTF-8 encoding, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_bjdata.md) parameter left at `keep` (the default), `from_bjdata()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_bjdata()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()` is strict as well (see above), so
a value read this way cannot be written back to BJData.
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()`'s
own `error_handler` parameter defaults to `strict` (see above), so a value read this way cannot be written back
to BJData unless a non-strict handler is passed there too.
!!! info "Round trips"
@@ -112,14 +112,19 @@ The library maps BSON record types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in string values"
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
decoder. `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged.
However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_bson()` is strict as well and throws the
same exception for a string value or element (key) name that is not valid UTF-8, so an object with such a key
or value cannot be produced in the first place, even though `from_bson()` would accept it from another source.
Element (key) names are never validated on read, since they are read byte-by-byte as a C string. `binary`
values (type `0x05`) are unaffected, since they are not required to hold text.
decoder, so checking is opt-in: with the [`error_handler`](../../api/basic_json/from_bson.md) parameter left at
`keep` (the default), `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back
unchanged. Passing `error_handler_t::strict` makes `from_bson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores
the ill-formed bytes. `to_bson()`'s own `error_handler` parameter defaults to `strict` and throws the same
exception for a string value or element (key) name that is not valid UTF-8, so an object with such a key or
value cannot be produced in the first place unless a non-strict handler is passed there, even though
`from_bson()` would accept it from another source with the default `keep` handler. Element (key) names are
never validated on read, since they are read byte-by-byte as a C string. `binary` values (type `0x05`) are
unaffected, since they are not required to hold text.
??? example
@@ -192,13 +192,18 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
(major type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this. This library does
not: `from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them
back unchanged. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_cbor()` is strict as well and throws the
same exception for a string value or object key that is not valid UTF-8, so such a value cannot be written back
to CBOR. Byte strings (major type 2) are unaffected, since they are not required to hold text.
(major type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this, so checking is
opt-in: with the [`error_handler`](../../api/basic_json/from_cbor.md) parameter left at `keep` (the default),
`from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back
unchanged. Passing `error_handler_t::strict` makes `from_cbor()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores
the ill-formed bytes. `to_cbor()`'s own [`error_handler`](../../api/basic_json/to_cbor.md) parameter defaults
to `strict` and throws the same exception for a string value or object key that is not valid UTF-8, so such a
value cannot be written back to CBOR unless a non-strict handler is passed there too. Byte strings (major
type 2) are unaffected, since they are not required to hold text.
!!! warning "Tagged items"
@@ -157,11 +157,17 @@ The library maps MessagePack types to JSON value types as follows:
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
This library follows that: `from_msgpack()` reads `str` bytes (object keys included) as-is, without validating
them, and `to_msgpack()` writes them back as-is, so such a value round-trips through `from_msgpack(to_msgpack(j))`
byte for byte. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way, unless an
error handler is passed that replaces or ignores the ill-formed bytes.
This library follows that by default: with its
[`error_handler`](../../api/basic_json/from_msgpack.md) parameter left at `keep` (the default),
`from_msgpack()` reads `str` bytes (object keys included) as-is, without validating them, so such a value
round-trips through `from_msgpack(to_msgpack(j))` byte for byte. Passing `error_handler_t::strict` makes
`from_msgpack()` check anyway and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. `to_msgpack()` itself has no `error_handler`
parameter and always writes `str` bytes as-is, since the specification permits it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way with the
default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes.
??? example
@@ -128,12 +128,17 @@ The library maps UBJSON types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in string values and object keys"
UBJSON's required string encoding is UTF-8, but this is not enforced on read: `from_ubjson()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
UBJSON's required string encoding is UTF-8, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_ubjson.md) parameter left at `keep` (the default), `from_ubjson()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_ubjson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()` is strict as well (see above), so
a value read this way cannot be written back to UBJSON.
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()`'s
own `error_handler` parameter defaults to `strict` (see above), so a value read this way cannot be written back
to UBJSON unless a non-strict handler is passed there too.
??? example
+8 -3
View File
@@ -340,9 +340,11 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
### json.exception.parse_error.113
A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
string was read where one was required (for instance as a map key), or the string's length specification is invalid.
The bytes of a string itself are not checked for valid UTF-8 on read; see the ill-formed UTF-8 notes on the
individual [binary format](../features/binary_formats/index.md) pages for how such a string is handled afterward.
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8 and the `error_handler` parameter of the corresponding `from_*` function is
set to `strict`. By default (`error_handler_t::keep`), the bytes of a string are not checked for valid UTF-8 on read;
see the ill-formed UTF-8 notes on the individual [binary format](../features/binary_formats/index.md) pages for how
such a string is handled depending on `error_handler`.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
@@ -365,6 +367,9 @@ type (for instance integers or `null`) are therefore not supported; see the note
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative
```
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte
```
### json.exception.parse_error.114
+82 -32
View File
@@ -31,6 +31,7 @@
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/is_sax.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/string_utils.hpp>
#include <nlohmann/detail/value_t.hpp>
@@ -108,8 +109,16 @@ class binary_reader
@brief create a binary reader
@param[in] adapter input adapter to read from
@param[in] format the binary format to parse
@param[in] error_handler how to treat text strings and object keys that
are not well-formed UTF-8; none of the supported formats
requires a decoder to reject those, so the default is to
@ref error_handler_t::keep them unchanged, as every binary
reader did before this parameter existed
*/
explicit binary_reader(InputAdapterType&& adapter, const input_format_t format = input_format_t::json) noexcept : ia(std::move(adapter)), input_format(format)
explicit binary_reader(InputAdapterType&& adapter, const input_format_t format = input_format_t::json,
const error_handler_t error_handler = error_handler_t::keep) noexcept
: ia(std::move(adapter)), input_format(format), error_handler(error_handler)
{
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
}
@@ -428,7 +437,7 @@ class binary_reader
{
if (get_bson_cstr_bulk(result, std::integral_constant<bool, bulk_scan> {}))
{
return true;
return check_string_utf8(result, "key");
}
auto out = std::back_inserter(result);
@@ -441,7 +450,7 @@ class binary_reader
}
if (current == 0x00)
{
return true;
return check_string_utf8(result, "key");
}
*out++ = static_cast<typename string_t::value_type>(current);
}
@@ -522,7 +531,7 @@ class binary_reader
"string"), nullptr));
}
return true;
return check_string_utf8(result, "string");
}
/*!
@@ -1149,7 +1158,7 @@ class binary_reader
@return whether string creation completed
*/
bool get_cbor_string(string_t& result)
bool get_cbor_string(string_t& result, const char* context = "string")
{
// number of indefinite-length strings that have been opened and not
// closed yet. RFC 8949, Section 3.2.3 does not permit nesting them,
@@ -1179,7 +1188,7 @@ class binary_reader
{
if (--open == 0)
{
return true;
return check_string_utf8(result, context);
}
get();
continue;
@@ -1192,7 +1201,7 @@ class binary_reader
if (open == 0)
{
return true;
return check_string_utf8(result, context);
}
get();
@@ -1216,7 +1225,7 @@ class binary_reader
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
return get_cbor_string(result, "key");
}
const char* found = nullptr;
@@ -2004,7 +2013,7 @@ class binary_reader
@return whether string creation completed
*/
bool get_msgpack_string(string_t& result)
bool get_msgpack_string(string_t& result, const char* context = "string")
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::msgpack, "string")))
{
@@ -2047,25 +2056,25 @@ class binary_reader
case 0xBE:
case 0xBF:
{
return get_string(input_format_t::msgpack, static_cast<unsigned int>(current) & 0x1Fu, result);
return get_string(input_format_t::msgpack, static_cast<unsigned int>(current) & 0x1Fu, result) && check_string_utf8(result, context);
}
case 0xD9: // str 8
{
std::uint8_t len{};
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result);
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result) && check_string_utf8(result, context);
}
case 0xDA: // str 16
{
std::uint16_t len{};
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result);
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result) && check_string_utf8(result, context);
}
case 0xDB: // str 32
{
std::uint32_t len{};
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result);
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result) && check_string_utf8(result, context);
}
default:
@@ -2143,7 +2152,7 @@ class binary_reader
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
return get_msgpack_string(result, "key");
}
if (current <= 0x7F || current >= 0xE0)
{
@@ -2159,7 +2168,7 @@ class binary_reader
}
else
{
return get_msgpack_string(result);
return get_msgpack_string(result, "key");
}
break;
}
@@ -2405,7 +2414,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key, true, "key") || !sax->key(key)))
{
return false;
}
@@ -2427,7 +2436,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key, false) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key, false, "key") || !sax->key(key)))
{
return false;
}
@@ -2495,7 +2504,7 @@ class binary_reader
@return whether string creation completed
*/
bool get_ubjson_string(string_t& result, const bool get_char = true)
bool get_ubjson_string(string_t& result, const bool get_char = true, const char* context = "string")
{
if (get_char)
{
@@ -2516,31 +2525,31 @@ class binary_reader
case 'U':
{
std::uint8_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'i':
{
std::int8_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'I':
{
std::int16_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'l':
{
std::int32_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'L':
{
std::int64_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'u':
@@ -2550,7 +2559,7 @@ class binary_reader
break;
}
std::uint16_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'm':
@@ -2560,7 +2569,7 @@ class binary_reader
break;
}
std::uint32_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'M':
@@ -2570,7 +2579,7 @@ class binary_reader
break;
}
std::uint64_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
default:
@@ -4031,15 +4040,53 @@ class binary_reader
const NumberType len,
string_t& result)
{
// Strings are taken as is: none of CBOR (RFC 8949 §3.1 leaves the
// choice to the decoder), MessagePack (whose spec explicitly allows
// a str object to contain an invalid byte sequence), UBJSON, BJData,
// or BSON requires a decoder to reject ill-formed UTF-8. The bytes
// are kept unchanged; dump() and the binary writers are the ones
// that check them and report type_error.316 if they are not valid.
// Strings are taken as is by default: none of CBOR (RFC 8949 §3.1
// leaves the choice to the decoder), MessagePack (whose spec
// explicitly allows a str object to contain an invalid byte
// sequence), UBJSON, BJData, or BSON requires a decoder to reject
// ill-formed UTF-8. Checking (and, with @ref error_handler_t::strict,
// rejecting, or with `replace`/`ignore`, sanitizing) is opt-in via
// @ref error_handler, applied once the whole string (all chunks of
// an indefinite-length CBOR string included) has been assembled, by
// @ref check_string_utf8 at the call site.
return get_bytes(format, len, "string", result);
}
/*!
@brief validate a decoded text string (value or object key) against @ref error_handler
None of the binary formats requires a decoder to reject ill-formed UTF-8
in a text string (see @ref get_string), so by default
(@ref error_handler_t::keep) this does nothing. A stricter
@ref error_handler opts into the same well-formedness check @ref
serializer::dump_escaped_impl applies when dumping a string:
@ref error_handler_t::strict rejects ill-formed input with
parse_error.113 (honoring `allow_exceptions` via @a sax), while
@ref error_handler_t::replace / @ref error_handler_t::ignore sanitize
@a result in place, using the exact same rules.
@param[in,out] result the already assembled string to check
@param[in] context further context information (for diagnostics)
@return whether @a result is acceptable (always true for `keep`)
*/
bool check_string_utf8(string_t& result, const char* context)
{
if (error_handler == error_handler_t::keep || is_valid_utf8(result))
{
return true;
}
if (error_handler == error_handler_t::strict)
{
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, "invalid string: ill-formed UTF-8 byte", context), nullptr));
}
result = sanitize_utf8(result, error_handler);
return true;
}
/*!
@brief create a byte array by reading bytes from the input
@@ -4211,6 +4258,9 @@ class binary_reader
/// input format
const input_format_t input_format = input_format_t::json;
/// how to treat text strings/object keys that are not well-formed UTF-8
const error_handler_t error_handler = error_handler_t::keep;
/// the SAX parser
json_sax_t* sax = nullptr;
+133 -29
View File
@@ -26,6 +26,7 @@
#include <nlohmann/detail/input/binary_reader.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/string_utils.hpp>
@@ -93,8 +94,12 @@ class binary_writer
@param[in] sink output sink to write to (a value-type sink such as
output_vector_sink, or output_adapter_sink wrapping a
type-erased output adapter)
@param[in] error_handler_ how to treat a string value or object key that
is not valid UTF-8 (CBOR, UBJSON, BJData, and BSON only; never
consulted by @ref write_msgpack or @ref write_bon8)
*/
explicit binary_writer(OutputSinkType sink) : oa(std::move(sink))
explicit binary_writer(OutputSinkType sink, const error_handler_t error_handler_ = error_handler_t::strict)
: oa(std::move(sink)), error_handler(error_handler_)
{}
/*!
@@ -107,10 +112,14 @@ class binary_writer
from one.
@param[in] adapter output adapter to write to
@param[in] error_handler_ how to treat a string value or object key that
is not valid UTF-8 (CBOR, UBJSON, BJData, and BSON only; never
consulted by @ref write_msgpack or @ref write_bon8)
*/
template < typename SinkType = OutputSinkType,
typename std::enable_if < std::is_constructible<SinkType, output_adapter_t<CharType>>::value, int >::type = 0 >
explicit binary_writer(output_adapter_t<CharType> adapter) : oa(SinkType(std::move(adapter)))
explicit binary_writer(output_adapter_t<CharType> adapter, const error_handler_t error_handler_ = error_handler_t::strict)
: oa(SinkType(std::move(adapter))), error_handler(error_handler_)
{}
/*!
@@ -215,15 +224,16 @@ class binary_writer
case value_t::string:
{
check_utf8(*j.m_data.m_value.string, j);
string_t storage;
const string_t& value = sanitize_utf8_for_write(*j.m_data.m_value.string, j, storage);
// step 1: write control byte and the string length
write_cbor_head(0x60, j.m_data.m_value.string->size());
write_cbor_head(0x60, value.size());
// step 2: write the string
oa.write_characters(
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
reinterpret_cast<const CharType*>(value.data()),
value.size());
break;
}
@@ -296,8 +306,14 @@ class binary_writer
// el.first is checked here, against the object as
// diagnostics context, because write_cbor(el.first)
// converts it to a temporary basic_json that would be
// used as the context instead
check_utf8(el.first, j);
// used as the context instead; for error_handler_t::keep
// and ::replace/::ignore the recursive write_cbor(el.first)
// call below handles the key like any other string, so no
// separate check is needed here for those
if (error_handler == error_handler_t::strict)
{
check_utf8(el.first, j);
}
write_cbor(el.first);
write_cbor(el.second);
}
@@ -691,16 +707,17 @@ class binary_writer
case value_t::string:
{
check_utf8(*j.m_data.m_value.string, j);
string_t storage;
const string_t& value = sanitize_utf8_for_write(*j.m_data.m_value.string, j, storage);
if (add_prefix)
{
oa.write_character(to_char_type('S'));
}
write_number_with_ubjson_prefix(j.m_data.m_value.string->size(), true, use_bjdata);
write_number_with_ubjson_prefix(value.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
reinterpret_cast<const CharType*>(value.data()),
value.size());
break;
}
@@ -855,11 +872,12 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object)
{
check_utf8(el.first, j);
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
string_t storage;
const string_t& key = sanitize_utf8_for_write(el.first, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(el.first.data()),
el.first.size());
reinterpret_cast<const CharType*>(key.data()),
key.size());
write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
}
@@ -905,7 +923,7 @@ class binary_writer
@throw type_error.316 if @a name is not valid UTF-8, before anything is
written
*/
static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
{
const auto it = name.find(static_cast<typename string_t::value_type>(0));
if (JSON_HEDLEY_UNLIKELY(it != BasicJsonType::string_t::npos))
@@ -913,9 +931,10 @@ class binary_writer
JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j));
}
check_utf8(name, j);
string_t storage;
const string_t& sanitized = sanitize_utf8_for_write(name, j, storage);
return /*id*/ 1ul + name.size() + /*zero-terminator*/1u;
return /*id*/ 1ul + sanitized.size() + /*zero-terminator*/1u;
}
/*!
@@ -935,14 +954,28 @@ class binary_writer
/*!
@brief Writes the given @a element_type and @a name to the output adapter
@a name has already been validated (and, for @ref error_handler_t::strict,
found well-formed) by @ref calc_bson_entry_header_size during the earlier
size pass, so only @ref error_handler_t::replace / @ref
error_handler_t::ignore need to sanitize it again here, to actually write
the bytes that size was computed from.
*/
void write_bson_entry_header(const string_t& name,
const std::uint8_t element_type)
{
oa.write_character(to_char_type(element_type));
oa.write_characters(
reinterpret_cast<const CharType*>(name.data()),
name.size());
if (error_handler == error_handler_t::keep || error_handler == error_handler_t::strict || is_valid_utf8(name))
{
oa.write_characters(reinterpret_cast<const CharType*>(name.data()), name.size());
}
else
{
const string_t sanitized = sanitize_utf8(name, error_handler);
oa.write_characters(reinterpret_cast<const CharType*>(sanitized.data()), sanitized.size());
}
// the terminating null byte is written explicitly rather than taken
// from the buffer, so that string_t::data() need not be null-terminated
oa.write_character(to_char_type(0x00));
@@ -979,27 +1012,41 @@ class binary_writer
from reading past a StringType that reports a size larger than what
it actually holds.
*/
static std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j)
std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<std::int32_t>(value.size())))
{
check_utf8(value, j);
string_t storage;
const string_t& sanitized = sanitize_utf8_for_write(value, j, storage);
return sizeof(std::int32_t) + sanitized.size() + 1ul;
}
return sizeof(std::int32_t) + value.size() + 1ul;
}
/*!
@brief Writes a BSON element with key @a name and string value @a value
@a value has already been validated (and, for @ref error_handler_t::strict,
found well-formed) by @ref calc_bson_string_size during the earlier size
pass, so only @ref error_handler_t::replace / @ref error_handler_t::ignore
need to sanitize it again here, to actually write the bytes that size was
computed from.
*/
void write_bson_string(const string_t& name,
const string_t& value)
{
write_bson_entry_header(name, 0x02);
write_number<std::int32_t>(to_bson_length(value.size() + 1ul), true);
const bool sanitize = error_handler != error_handler_t::keep
&& error_handler != error_handler_t::strict
&& !is_valid_utf8(value);
const string_t sanitized = sanitize ? sanitize_utf8(value, error_handler) : string_t{};
const string_t& written = sanitize ? sanitized : value;
write_number<std::int32_t>(to_bson_length(written.size() + 1ul), true);
oa.write_characters(
reinterpret_cast<const CharType*>(value.data()),
value.size());
reinterpret_cast<const CharType*>(written.data()),
written.size());
// the terminating null byte is written explicitly rather than taken
// from the buffer, so that string_t::data() need not be null-terminated
oa.write_character(to_char_type(0x00));
@@ -1116,7 +1163,7 @@ class binary_writer
@throw type_error.316 if @a j is a string that is not valid UTF-8, before
anything is written
*/
static std::size_t calc_bson_value_size(const BasicJsonType& j)
std::size_t calc_bson_value_size(const BasicJsonType& j)
{
switch (j.type())
{
@@ -1252,7 +1299,7 @@ class binary_writer
@throw type_error.316 if a string value or a key is not valid UTF-8,
before anything is written
*/
static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
{
// the object or array whose entries are being sized, and the ones it
// is in; nothing is allocated unless the document nests
@@ -2170,6 +2217,59 @@ class binary_writer
}
}
/*!
@brief return @a s as it should be written, honoring @ref error_handler
Used by @ref write_cbor, @ref write_ubjson (and so @ref write_bjdata), and
the BSON writing functions for string values and object keys; never by
@ref write_msgpack or @ref write_bon8, which do not take an @ref
error_handler (MessagePack's spec allows a str object to contain
ill-formed UTF-8, and BON8 always validates, since UTF-8 lead bytes are
structural there).
- @ref error_handler_t::keep: @a s is returned unchanged, without even
checking it (the behavior of release 3.12.0 and earlier).
- @ref error_handler_t::strict: @ref check_utf8 is called, which throws
type_error.316 if @a s is not valid UTF-8.
- @ref error_handler_t::replace / @ref error_handler_t::ignore: @a s is
sanitized into @a storage with exactly the rules @ref
serializer::dump_escaped_impl uses, so that parsing what @ref
basic_json::dump produces for the same string and the same handler
yields the same result.
Well-formed input is never copied: this returns a reference to @a s
itself in every case but a sanitized `replace`/`ignore` one, so @a
storage must outlive the returned reference only then.
@param[in] s the string (value or object key) to write
@param[in] context the value @a s belongs to (for diagnostics)
@param[out] storage backing storage for a sanitized copy
@return a reference to @a s, or to @a storage once it holds a sanitized copy
*/
const string_t& sanitize_utf8_for_write(const string_t& s, const BasicJsonType& context, string_t& storage) const
{
switch (error_handler)
{
case error_handler_t::keep:
return s;
case error_handler_t::strict:
check_utf8(s, context);
return s;
case error_handler_t::replace:
case error_handler_t::ignore:
default:
if (is_valid_utf8(s))
{
return s;
}
storage = sanitize_utf8(s, error_handler);
return storage;
}
}
/*!
@brief write an integer in the shortest encoding
@@ -2498,6 +2598,10 @@ class binary_writer
/// the output
OutputSinkType oa;
/// how to treat a string value or object key that is not valid UTF-8
/// (CBOR, UBJSON, BJData, and BSON only)
const error_handler_t error_handler = error_handler_t::strict;
};
} // namespace detail
@@ -0,0 +1,38 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/// how to treat decoding errors
///
/// @ref basic_json::dump uses this to decide what to do with ill-formed
/// UTF-8 while escaping a string, and the binary writers (@ref
/// basic_json::to_cbor, @ref basic_json::to_ubjson, @ref
/// basic_json::to_bjdata, @ref basic_json::to_bson) use it the same way for
/// string values and object keys. The binary readers (@ref
/// basic_json::from_cbor, @ref basic_json::from_msgpack, @ref
/// basic_json::from_ubjson, @ref basic_json::from_bjdata, @ref
/// basic_json::from_bson) use it to decide whether to check text strings
/// and object keys for well-formed UTF-8 at all, since none of those
/// formats requires a decoder to do so.
enum class error_handler_t
{
strict, ///< throw a type_error/parse_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences unchanged
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+62 -10
View File
@@ -27,6 +27,7 @@
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/cpp_future.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/recursion_depth_limit.hpp>
#include <nlohmann/detail/string_concat.hpp>
@@ -41,14 +42,6 @@ namespace detail
// serialization //
///////////////////
/// how to treat decoding errors
enum class error_handler_t
{
strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore ///< ignore invalid UTF-8 sequences
};
template<typename BasicJsonType>
class serializer
{
@@ -839,6 +832,16 @@ class serializer
// EnsureAscii parameter is used, non-ASCII characters
if ((codepoint <= 0x1F) || (EnsureAscii && (codepoint >= 0x7F)))
{
if (EnsureAscii && error_handler == error_handler_t::keep)
{
// this character was buffered as raw bytes
// below in case it turned out to be part of
// an ill-formed sequence (which is kept as
// is); now that it decoded to a well-formed
// code point, undo that and \u-escape it
// like any other character instead
bytes = bytes_after_last_accept;
}
if (codepoint <= 0xFFFF)
{
write_u_escape(bytes, static_cast<std::uint16_t>(codepoint));
@@ -937,6 +940,44 @@ class serializer
break;
}
case error_handler_t::keep:
{
// the bytes of this (now abandoned) ill-formed
// sequence seen so far are already buffered below
// and are kept unchanged in the output
if (undumped_chars > 0)
{
// the byte that ended the sequence may be OK
// for itself (e.g., a quote that must still be
// escaped, or the lead byte of a well-formed
// code point), so read it again
--i;
}
else
{
// a byte that cannot start a sequence (e.g.,
// 0xFF or a stray continuation byte) is kept
// as well
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -945,9 +986,12 @@ class serializer
default: // decode found yet incomplete multibyte code point
{
if (!EnsureAscii)
if (!EnsureAscii || error_handler == error_handler_t::keep)
{
// code point will not be escaped - copy byte to buffer
// code point will not be escaped (or will be kept as
// is if it turns out to be ill-formed) - copy byte to
// buffer; dropped again above if it decodes to a
// well-formed code point that needs \u-escaping
string_buffer[bytes++] = s[i];
}
++undumped_chars;
@@ -998,6 +1042,14 @@ class serializer
break;
}
case error_handler_t::keep:
{
// write the ill-formed trailing bytes as is; they were
// buffered above regardless of EnsureAscii
put_buffer(string_buffer, bytes);
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
+130
View File
@@ -16,6 +16,7 @@
#include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -179,5 +180,134 @@ inline std::uint8_t decode(std::uint8_t& state, std::uint32_t& codep, const std:
return state;
}
/*!
@brief check a string for well-formed UTF-8 (RFC 3629, section 4)
Used by the binary readers (CBOR, MessagePack, UBJSON, BJData, BSON) when an
@ref error_handler_t other than `keep` is requested for a text string value
or object key: none of those formats requires a decoder to reject ill-formed
UTF-8 on its own, so the check is opt-in there, unlike the JSON lexer and the
serializer's @ref decode -based escaping, which always run it.
@param[in] s the string to check
@param[in] first the index to start checking at
@return whether `s.substr(first)` is well-formed UTF-8
@sa @ref decode
*/
template<typename StringType>
inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noexcept
{
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
for (std::size_t i = first; i < s.size(); ++i)
{
decode(state, codepoint, static_cast<std::uint8_t>(s[i]));
if (state == UTF8_REJECT)
{
return false;
}
}
return state == UTF8_ACCEPT;
}
/*!
@brief sanitize a string with ill-formed UTF-8 for @ref error_handler_t::replace or @ref error_handler_t::ignore
Replaces every maximal ill-formed subsequence with U+FFFD (`replace`) or
drops it (`ignore`), using exactly the same boundaries @ref
serializer::dump_escaped_impl uses while escaping a string: a byte that does
not extend the sequence started by the previous byte(s) is reread as the
start of a new one, instead of being swallowed along with them.
@pre @a error_handler is @ref error_handler_t::replace or @ref error_handler_t::ignore
@note Well-formed input is copied through unchanged, including bytes (e.g.
control characters or quotes) that @ref serializer::dump_escaped_impl
would itself escape; this function only concerns itself with
well-formedness, not with producing valid JSON text.
@param[in] s the string to sanitize
@param[in] error_handler @ref error_handler_t::replace or @ref error_handler_t::ignore
@return @a s with every ill-formed subsequence replaced or removed
@sa @ref decode
*/
template<typename StringType>
inline StringType sanitize_utf8(const StringType& s, const error_handler_t error_handler)
{
JSON_ASSERT(error_handler == error_handler_t::replace || error_handler == error_handler_t::ignore);
StringType result;
result.reserve(s.size());
std::uint32_t codepoint = 0;
std::uint8_t state = UTF8_ACCEPT;
// length of result after the last accepted code point
std::size_t result_len_after_last_accept = 0;
// whether bytes of an as yet unresolved sequence were already appended
bool pending = false;
for (std::size_t i = 0; i < s.size(); ++i)
{
switch (decode(state, codepoint, static_cast<std::uint8_t>(s[i])))
{
case UTF8_ACCEPT: // decode found a well-formed code point
{
result.push_back(s[i]);
result_len_after_last_accept = result.size();
pending = false;
break;
}
case UTF8_REJECT: // decode found an ill-formed byte
{
// in case we saw this byte for the first time, read it again,
// because it may be fine for itself, just not for the
// sequence that came before it
if (pending)
{
--i;
}
// drop the bytes of the ill-formed sequence buffered below
result.resize(result_len_after_last_accept);
if (error_handler == error_handler_t::replace)
{
result.append("\xEF\xBF\xBD");
result_len_after_last_accept = result.size();
}
pending = false;
state = UTF8_ACCEPT;
break;
}
default: // decode found yet incomplete multibyte code point
{
result.push_back(s[i]);
pending = true;
break;
}
}
}
// the string ended with an incomplete sequence
if (state != UTF8_ACCEPT)
{
result.resize(result_len_after_last_accept);
if (error_handler == error_handler_t::replace)
{
result.append("\xEF\xBF\xBD");
}
}
return result;
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+69 -46
View File
@@ -201,9 +201,10 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// used by the vector-returning to_* overloads
template<typename CharType> using vector_binary_writer =
::nlohmann::detail::binary_writer<basic_json, CharType, ::nlohmann::detail::output_vector_sink<CharType>>;
template<typename CharType> static vector_binary_writer<CharType> vector_writer(std::vector<CharType>& v)
template<typename CharType> static vector_binary_writer<CharType> vector_writer(
std::vector<CharType>& v, const ::nlohmann::detail::error_handler_t error_handler = ::nlohmann::detail::error_handler_t::strict)
{
return vector_binary_writer<CharType>(::nlohmann::detail::output_vector_sink<CharType>(v));
return vector_binary_writer<CharType>(::nlohmann::detail::output_vector_sink<CharType>(v), error_handler);
}
JSON_PRIVATE_UNLESS_TESTED:
@@ -5445,26 +5446,29 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
public:
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static std::vector<std::uint8_t> to_cbor(const basic_json& j)
static std::vector<std::uint8_t> to_cbor(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_cbor(j);
vector_writer(result, error_handler).write_cbor(j);
return result;
}
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o)
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_cbor(j);
binary_writer<std::uint8_t>(o, error_handler).write_cbor(j);
}
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static void to_cbor(const basic_json& j, detail::output_adapter<char> o)
static void to_cbor(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_cbor(j);
binary_writer<char>(o, error_handler).write_cbor(j);
}
/// @brief create a MessagePack serialization of a given JSON value
@@ -5495,28 +5499,31 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static std::vector<std::uint8_t> to_ubjson(const basic_json& j,
const bool use_size = false,
const bool use_type = false)
const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_ubjson(j, use_size, use_type);
vector_writer(result, error_handler).write_ubjson(j, use_size, use_type);
return result;
}
/// @brief create a UBJSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static void to_ubjson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false)
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_ubjson(j, use_size, use_type);
binary_writer<std::uint8_t>(o, error_handler).write_ubjson(j, use_size, use_type);
}
/// @brief create a UBJSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static void to_ubjson(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false)
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_ubjson(j, use_size, use_type);
binary_writer<char>(o, error_handler).write_ubjson(j, use_size, use_type);
}
/// @brief create a BJData serialization of a given JSON value
@@ -5524,11 +5531,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static std::vector<std::uint8_t> to_bjdata(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2)
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_ubjson(j, use_size, use_type, true, true, version);
vector_writer(result, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
return result;
}
@@ -5536,42 +5544,47 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/to_bjdata/
static void to_bjdata(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2)
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_ubjson(j, use_size, use_type, true, true, version);
binary_writer<std::uint8_t>(o, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
}
/// @brief create a BJData serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bjdata/
static void to_bjdata(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2)
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_ubjson(j, use_size, use_type, true, true, version);
binary_writer<char>(o, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static std::vector<std::uint8_t> to_bson(const basic_json& j)
static std::vector<std::uint8_t> to_bson(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_bson(j);
vector_writer(result, error_handler).write_bson(j);
return result;
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o)
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_bson(j);
binary_writer<std::uint8_t>(o, error_handler).write_bson(j);
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static void to_bson(const basic_json& j, detail::output_adapter<char> o)
static void to_bson(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_bson(j);
binary_writer<char>(o, error_handler).write_bson(j);
}
/// @brief create a BON8 serialization of a given JSON value
@@ -5605,12 +5618,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json from_cbor(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor, error_handler).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5625,12 +5639,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json from_cbor(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor, error_handler).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5672,12 +5687,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_msgpack(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5691,12 +5707,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_msgpack(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5736,12 +5753,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_ubjson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5755,12 +5773,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_ubjson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5800,12 +5819,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bjdata(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5819,12 +5839,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bjdata(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5874,12 +5895,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5893,12 +5915,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,346 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <string>
#include <vector>
namespace
{
struct ill_formed_case
{
const char* name;
std::string bytes;
};
// RFC 3629 ill-formed sequences used throughout this file, plus one
// well-formed sequence for contrast
const std::vector<ill_formed_case> ill_formed_cases =
{
{"overlong", "\xC0\xAE"},
{"lone_0xFF", "\xFF"},
{"truncated", "\xE2\x82"},
{"surrogate", "\xED\xA0\x80"},
};
const std::string valid_sequence = "\xC3\xA9"; // U+00E9, "é"
using eh = json::error_handler_t;
const std::vector<eh> all_handlers = {eh::strict, eh::replace, eh::ignore, eh::keep};
// what dump()+parse() produces for a sanitizing error_handler; this is the
// ground truth every binary writer/reader is checked against
std::string dump_and_parse(const std::string& raw, eh error_handler)
{
return json::parse(json(raw).dump(-1, ' ', false, error_handler)).get<std::string>();
}
} // namespace
TEST_CASE("UTF-8 error_handler for the binary readers and writers")
{
SECTION("writers: string value")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
const json jval = c.bytes;
CHECK_THROWS_AS(json::to_cbor(jval, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jval, false, false, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::strict), json::type_error&);
{
json jobj;
jobj["k"] = jval;
CHECK_THROWS_AS(json::to_bson(jobj, eh::strict), json::type_error&);
}
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(json::to_cbor(jval, h)).get<std::string>() == expected);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, h)).get<std::string>() == expected);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, h)).get<std::string>() == expected);
{
json jobj;
jobj["k"] = jval;
const auto bytes = json::to_bson(jobj, h);
CHECK(json::from_bson(bytes)["k"].get<std::string>() == expected);
}
}
// keep: the writer passes the ill-formed bytes through unchanged,
// exactly as every binary writer did before this parameter existed
CHECK(json::from_cbor(json::to_cbor(jval, eh::keep)).get<std::string>() == c.bytes);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, eh::keep)).get<std::string>() == c.bytes);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::keep)).get<std::string>() == c.bytes);
{
json jobj;
jobj["k"] = jval;
const auto bytes = json::to_bson(jobj, eh::keep);
CHECK(json::from_bson(bytes)["k"].get<std::string>() == c.bytes);
}
}
}
SECTION("writers: object key")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
json jobj;
jobj[c.bytes] = 1;
CHECK_THROWS_AS(json::to_cbor(jobj, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jobj, false, false, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bson(jobj, eh::strict), json::type_error&);
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(json::to_cbor(jobj, h)).begin().key() == expected);
CHECK(json::from_ubjson(json::to_ubjson(jobj, false, false, h)).begin().key() == expected);
CHECK(json::from_bjdata(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, h)).begin().key() == expected);
CHECK(json::from_bson(json::to_bson(jobj, h)).begin().key() == expected);
}
// keep: object keys round-trip unchanged too
CHECK(json::from_cbor(json::to_cbor(jobj, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_ubjson(json::to_ubjson(jobj, false, false, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_bjdata(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_bson(json::to_bson(jobj, eh::keep)).begin().key() == c.bytes);
}
}
SECTION("readers: string value")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
// bytes produced the lenient (keep) way, as any binary reader
// accepted them before this parameter existed
const auto cbor_bytes = json::to_cbor(json(c.bytes), eh::keep);
const auto msgpack_bytes = json::to_msgpack(json(c.bytes)); // to_msgpack has no error_handler; always pass-through
const auto ubjson_bytes = json::to_ubjson(json(c.bytes), false, false, eh::keep);
const auto bjdata_bytes = json::to_bjdata(json(c.bytes), false, false, json::bjdata_version_t::draft2, eh::keep);
const auto bson_bytes = [&c]
{
json jobj;
jobj["k"] = c.bytes;
return json::to_bson(jobj, eh::keep);
}();
// keep (the default): bytes are kept unchanged
CHECK(json::from_cbor(cbor_bytes).get<std::string>() == c.bytes);
CHECK(json::from_msgpack(msgpack_bytes).get<std::string>() == c.bytes);
CHECK(json::from_ubjson(ubjson_bytes).get<std::string>() == c.bytes);
CHECK(json::from_bjdata(bjdata_bytes).get<std::string>() == c.bytes);
CHECK(json::from_bson(bson_bytes)["k"].get<std::string>() == c.bytes);
// strict: parse_error.113, discarded (not thrown) when allow_exceptions is false
CHECK_THROWS_AS(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, eh::strict), json::parse_error&);
CHECK(json::from_cbor(cbor_bytes, true, false, json::cbor_tag_handler_t::error, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_msgpack(msgpack_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_msgpack(msgpack_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_ubjson(ubjson_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_ubjson(ubjson_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_bjdata(bjdata_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_bjdata(bjdata_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_bson(bson_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_bson(bson_bytes, true, false, eh::strict).is_discarded());
// replace / ignore: match what dump() would have sanitized the same bytes to
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, h).get<std::string>() == expected);
CHECK(json::from_msgpack(msgpack_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_ubjson(ubjson_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_bjdata(bjdata_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_bson(bson_bytes, true, true, h)["k"].get<std::string>() == expected);
}
}
}
SECTION("readers: object key")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
json jobj;
jobj[c.bytes] = 1;
const auto cbor_bytes = json::to_cbor(jobj, eh::keep);
const auto msgpack_bytes = json::to_msgpack(jobj);
const auto ubjson_bytes = json::to_ubjson(jobj, false, false, eh::keep);
const auto bjdata_bytes = json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::keep);
const auto bson_bytes = json::to_bson(jobj, eh::keep);
CHECK(json::from_cbor(cbor_bytes).begin().key() == c.bytes);
CHECK(json::from_msgpack(msgpack_bytes).begin().key() == c.bytes);
CHECK(json::from_ubjson(ubjson_bytes).begin().key() == c.bytes);
CHECK(json::from_bjdata(bjdata_bytes).begin().key() == c.bytes);
CHECK(json::from_bson(bson_bytes).begin().key() == c.bytes);
CHECK_THROWS_AS(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_msgpack(msgpack_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_ubjson(ubjson_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_bjdata(bjdata_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_bson(bson_bytes, true, true, eh::strict), json::parse_error&);
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, h).begin().key() == expected);
CHECK(json::from_msgpack(msgpack_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_ubjson(ubjson_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_bjdata(bjdata_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_bson(bson_bytes, true, true, h).begin().key() == expected);
}
}
}
SECTION("well-formed UTF-8 is unaffected by error_handler")
{
const json jval = valid_sequence;
json jobj;
jobj[valid_sequence] = valid_sequence;
for (const auto h : all_handlers)
{
CAPTURE(static_cast<int>(h));
CHECK(json::from_cbor(json::to_cbor(jval, h)).get<std::string>() == valid_sequence);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, h)).get<std::string>() == valid_sequence);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, h)).get<std::string>() == valid_sequence);
CHECK(json::from_bson(json::to_bson(jobj, h)).begin().key() == valid_sequence);
CHECK(json::from_cbor(json::to_cbor(jval, eh::keep), true, true, json::cbor_tag_handler_t::error, h).get<std::string>() == valid_sequence);
CHECK(json::from_msgpack(json::to_msgpack(jval), true, true, h).get<std::string>() == valid_sequence);
}
}
SECTION("dump() with error_handler_t::keep writes raw bytes as is")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
const json jval = c.bytes;
const std::string dumped = jval.dump(-1, ' ', false, eh::keep);
CHECK(dumped.find(c.bytes) != std::string::npos);
// even with ensure_ascii, the ill-formed bytes are written as is
const std::string dumped_ascii = jval.dump(-1, ' ', true, eh::keep);
CHECK(dumped_ascii.find(c.bytes) != std::string::npos);
}
// well-formed characters around an ill-formed sequence are still
// escaped as usual under ensure_ascii
const json mixed = valid_sequence + ill_formed_cases[1].bytes; // "é" + lone 0xFF
const std::string dumped_mixed = mixed.dump(-1, ' ', true, eh::keep);
CHECK(dumped_mixed.find("\\u00e9") != std::string::npos);
CHECK(dumped_mixed.find(ill_formed_cases[1].bytes) != std::string::npos);
// the byte that ends an ill-formed sequence is read again, so a quote,
// a backslash, or a control character after it is still escaped, and
// a well-formed code point after it is escaped under ensure_ascii
for (const bool ensure_ascii :
{
false, true
})
{
CAPTURE(ensure_ascii);
CHECK(json("\xC3\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\\"\"");
CHECK(json("\xC3\\").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\\\\"");
CHECK(json("\xC3\n").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\n\"");
CHECK(json("\xE2\x82\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xE2\x82\\\"\"");
CHECK(json("\xFF\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xFF\\\"\"");
CHECK(json("a\xE2\x82").dump(-1, ' ', ensure_ascii, eh::keep) == "\"a\xE2\x82\"");
}
CHECK(json("\xC3\xC3\xA9").dump(-1, ' ', false, eh::keep) == "\"\xC3\xC3\xA9\"");
CHECK(json("\xC3\xC3\xA9").dump(-1, ' ', true, eh::keep) == "\"\xC3\\u00e9\"");
}
SECTION("to_msgpack and to_bon8 are not affected by error_handler")
{
const json jval = ill_formed_cases[1].bytes; // lone 0xFF
// to_msgpack has no error_handler parameter; the bytes are always
// passed through, as MessagePack's spec allows
CHECK(json::from_msgpack(json::to_msgpack(jval)).get<std::string>() == ill_formed_cases[1].bytes);
// to_bon8 has no error_handler parameter either; UTF-8 is structural
// for BON8, so it always rejects ill-formed input
CHECK_THROWS_AS(json::to_bon8(jval), json::type_error&);
}
SECTION("allow_exceptions=false with error_handler_t::strict discards the value")
{
const auto bytes = json::to_cbor(json(ill_formed_cases[0].bytes), eh::keep);
const json result = json::from_cbor(bytes, true, false, json::cbor_tag_handler_t::error, eh::strict);
CHECK(result.is_discarded());
}
SECTION("default parameters are unchanged")
{
const json jval = ill_formed_cases[0].bytes;
// to_*: the default error_handler is strict, so ill-formed input still throws
CHECK_THROWS_AS(json::to_cbor(jval), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jval), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jval), json::type_error&);
{
json jobj;
jobj["k"] = jval;
CHECK_THROWS_AS(json::to_bson(jobj), json::type_error&);
}
// from_*: the default error_handler is keep, so ill-formed bytes are
// still accepted unchanged, exactly as in release 3.12.0
const auto cbor_bytes = json::to_cbor(jval, eh::keep);
CHECK(json::from_cbor(cbor_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto ubjson_bytes = json::to_ubjson(jval, false, false, eh::keep);
CHECK(json::from_ubjson(ubjson_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto bjdata_bytes = json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::keep);
CHECK(json::from_bjdata(bjdata_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto msgpack_bytes = json::to_msgpack(jval);
CHECK(json::from_msgpack(msgpack_bytes).get<std::string>() == ill_formed_cases[0].bytes);
json bson_obj;
bson_obj["k"] = jval;
const auto bson_bytes = json::to_bson(bson_obj, eh::keep);
CHECK(json::from_bson(bson_bytes)["k"].get<std::string>() == ill_formed_cases[0].bytes);
}
}