Compare commits

..
Author SHA1 Message Date
Niels Lohmann fff221c58b Add an error_handler parameter for UTF-8 to the binary readers and writers
Adds error_handler_t::keep (invalid UTF-8 sequences are left unchanged) as
a fourth error_handler_t value, and threads an error_handler parameter
through the binary writers and readers:

- to_cbor/to_ubjson/to_bjdata/to_bson gain a trailing error_handler
  parameter (default strict, matching their existing type_error.316
  behavior); keep writes a string value or object key's bytes as is
  instead of throwing, and replace/ignore sanitize it exactly like
  dump() would, including for the BSON length prefix. to_msgpack and
  to_bon8 are unchanged.
- from_cbor/from_msgpack/from_ubjson/from_bjdata/from_bson gain a
  trailing error_handler parameter (default keep, i.e. the lenient
  behavior every binary reader had in 3.12.0 and still has after
  #5741); strict checks every string value and object key and raises
  parse_error.113 for ill-formed UTF-8, honoring allow_exceptions;
  replace/ignore sanitize it like dump() would. from_bon8 is
  unchanged, since UTF-8 lead bytes are structural there.

dump()'s own keep support writes ill-formed bytes as is, even with
ensure_ascii, while still \u-escaping well-formed characters around
them as usual.

The UTF-8 validity check (is_valid_utf8) and the replace/ignore
sanitizing logic (sanitize_utf8) now live in string_utils.hpp, shared
by the serializer and the binary reader/writer; error_handler_t itself
moved to its own header (detail/output/error_handler.hpp) so that
string_utils.hpp does not need to depend on serializer.hpp.

See #5529 and #5741.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 08:43:15 +02:00
Niels Lohmann 92c729f657 Accept ill-formed UTF-8 in all binary readers again
RFC 8949 and the MessagePack/BSON/UBJSON/BJData specs leave UTF-8
well-formedness checking up to the decoder, so following #5529 the
binary readers are lenient by default again, as in release 3.12.0
(the reader-side check was added by #5185/#5531, not in any release);
reader-side validation becomes opt-in in a follow-up PR. The writers
stay strict and throw type_error.316 for ill-formed UTF-8. BON8 is
unchanged, since UTF-8 lead bytes are structural there.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:59:14 +02:00
Niels Lohmann f3e30fa5b7 Merge branch 'develop' into claude/binary-utf8-roundtrip-5651
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:44:29 +02:00
Niels Lohmann 73d116547a Reject ill-formed UTF-8 in CBOR/UBJSON/BJData/BSON writers, accept it in MessagePack reader
to_cbor(), to_ubjson(), to_bjdata(), and to_bson() wrote a string or object
key with ill-formed UTF-8 byte for byte, but from_cbor()/from_ubjson()/
from_bjdata()/from_bson() reject such strings with parse_error.113 since
d19f7f5dc (#5185, not yet released): a value that serialized without error
could not be read back by the same library (#5651).

CBOR (RFC 8949 Section 5), UBJSON, BJData, and BSON all require text strings
and object keys/element names to be valid UTF-8, so their writers now
validate and throw type_error.316, like to_bon8() already does. For BSON,
the check runs in the size-computation pass, before any byte is written, the
same way the binary subtype check works. check_bon8_utf8() is renamed to
check_utf8() since it is now shared by all of these writers.

MessagePack's specification explicitly allows a str object to contain an
invalid byte sequence and expects a deserializer to hand the bytes back
unchanged, so its writer is unaffected and from_msgpack() (object keys
included) no longer validates UTF-8, restoring its pre-#5185 behavior.
dump() still rejects ill-formed UTF-8 with type_error.316 unless an error
handler is passed.

Docs: add the type_error.316 exception to to_cbor/to_ubjson/to_bjdata/to_bson,
document the writer-side check on the cbor/bson/ubjson/bjdata format pages,
and rewrite the MessagePack UTF-8 warning to describe the round-trip and
dump() behavior instead of a validation requirement the spec does not have.

Tests: add ill-formed value/key cases (invalid byte, truncated sequence,
encoded surrogate, overlong encoding) expecting type_error.316 to
unit-cbor.cpp, unit-ubjson.cpp, unit-bjdata.cpp, and unit-bson.cpp (which
also checks the output vector stays empty), and turn unit-msgpack.cpp's
former parse_error.113 ill-formed-UTF-8 tests into byte-for-byte round-trip
tests for both values and keys.

This text was written by Claude Code.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:19:11 +02:00
44 changed files with 2807 additions and 2757 deletions
+3 -10
View File
@@ -2,23 +2,16 @@ name: "Check amalgamation"
on:
pull_request:
# also check develop itself: a PR can be merged before its own run of this
# workflow completes (e.g. while it is still queued), leaving single_include
# stale on develop without any failing check
push:
branches:
- develop
concurrency:
group: ${{ github.workflow }}-${{ github.ref || github.run_id }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
cancel-in-progress: true
permissions:
contents: read
jobs:
save:
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest
steps:
- name: Harden Runner
@@ -50,11 +43,11 @@ jobs:
with:
egress-policy: audit
- name: Checkout pull request or pushed commit
- name: Checkout pull request
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
path: main
ref: ${{ github.event.pull_request.head.sha || github.sha }}
ref: ${{ github.event.pull_request.head.sha }}
persist-credentials: false
- name: Checkout tools
@@ -10,8 +10,7 @@ permissions:
jobs:
comment:
# push runs on develop have no PR to comment on (and no "pr" artifact)
if: ${{ github.event.workflow_run.conclusion == 'failure' && github.event.workflow_run.event == 'pull_request' }}
if: ${{ github.event.workflow_run.conclusion == 'failure' }}
runs-on: ubuntu-latest
permissions:
contents: read
+1 -1
View File
@@ -1394,7 +1394,7 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I
- The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/)
- The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
- The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0).
- The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
- The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
<img align="right" src="https://git.fsfe.org/reuse/reuse-ci/raw/branch/master/reuse-horizontal.png" alt="REUSE Software">
+6 -3
View File
@@ -25,10 +25,12 @@ and `ensure_ascii` parameters.
result consists of ASCII characters only.
`error_handler` (in)
: how to react on decoding errors; there are three possible values (see [`error_handler_t`](error_handler_t.md):
: how to react on decoding errors; there are four possible values (see [`error_handler_t`](error_handler_t.md):
`strict` (throws an exception in case a decoding error occurs; default), `replace` (replace invalid UTF-8 sequences
with U+FFFD), and `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped)).
with U+FFFD), `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped), and `keep` (write the ill-formed bytes to the output as is,
without escaping them, even if `ensure_ascii` is `#!cpp true`; the result is then not valid UTF-8, but equals the
input bytes exactly, and well-formed characters around the ill-formed bytes are still escaped as usual)).
## Return value
@@ -94,3 +96,4 @@ Binary values are serialized as an object containing two keys:
- Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0.
- Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0.
- Error handler `keep` added in version 3.13.0.
@@ -4,15 +4,30 @@
enum class error_handler_t {
strict,
replace,
ignore
ignore,
keep
};
```
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
`basic_json` value. Three values are differentiated:
This enumeration is used to choose how to treat ill-formed UTF-8 in a string value or object key:
- [`dump`](dump.md) uses it while serializing a `basic_json` value to text.
- [`to_cbor`](to_cbor.md), [`to_ubjson`](to_ubjson.md), [`to_bjdata`](to_bjdata.md), and [`to_bson`](to_bson.md) use it
while serializing a `basic_json` value to that binary format; none of CBOR, UBJSON, BJData, or BSON requires a
decoder to reject ill-formed UTF-8, so by default (`strict`) the library checks on write instead. `to_msgpack` and
`to_bon8` do not take this parameter: MessagePack's specification explicitly allows a string to contain ill-formed
UTF-8, so `to_msgpack` always passes it through, while BON8 always validates, since UTF-8 lead bytes are structural
to that format.
- [`from_cbor`](from_cbor.md), [`from_msgpack`](from_msgpack.md), [`from_ubjson`](from_ubjson.md),
[`from_bjdata`](from_bjdata.md), and [`from_bson`](from_bson.md) use it while parsing that binary format, to decide
whether to check a string value or object key for well-formed UTF-8 at all; by default (`keep`) they do not, as no
binary reader did before this parameter was added. `from_bon8` does not take this parameter, for the same reason
`to_bon8` does not.
Four values are differentiated:
strict
: throw a `type_error` exception in case of invalid UTF-8
: throw a `type_error`/`parse_error` exception in case of invalid UTF-8
replace
: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER)
@@ -20,6 +35,12 @@ replace
ignore
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
keep
: keep invalid UTF-8 sequences unchanged; only meaningful for the binary formats mentioned above, since [`dump`]
(dump.md) itself must produce text, and `keep` there writes the ill-formed bytes to the output as is, so the
result is then not valid UTF-8 (but still equals the input bytes exactly, including around any well-formed
characters, which are still escaped as usual)
## Examples
??? example
@@ -40,3 +61,5 @@ ignore
## Version history
- Added in version 3.4.0.
- Added `keep`, and made this enumeration apply to the binary readers and writers in addition to `dump`, in version
3.13.0.
+12 -3
View File
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_bjdata(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_bjdata(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the BJData (Binary JData) serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
BJData does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -73,7 +81,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if a parse error occurs
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string could not be parsed
successfully
successfully, or if a string value or object key is not valid UTF-8 and `error_handler` is `strict`
- Throws [out_of_range.408](../../home/exceptions.md#jsonexceptionout_of_range408) if the size of an optimized container
or n-dimensional array cannot be represented by `std::size_t`
@@ -111,3 +119,4 @@ Linear in the size of the input.
- Added in version 3.11.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+13 -2
View File
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_bson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_bson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the BSON (Binary JSON) serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
BSON does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -75,6 +83,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
invalid string or byte array length)
- Throws [`parse_error.114`](../../home/exceptions.md#jsonexceptionparse_error114) if an unsupported BSON record type is
encountered
- Throws [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) if a string value or object key is
not valid UTF-8 and `error_handler` is `strict`
## Complexity
@@ -111,6 +121,7 @@ Linear in the size of the input.
- Added in version 3.4.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+14 -4
View File
@@ -6,14 +6,16 @@ template<typename InputType>
static basic_json from_cbor(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error);
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_cbor(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error);
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the CBOR (Concise Binary Object Representation) serialization format.
@@ -65,6 +67,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: how to treat CBOR tags (optional, `error` by default); see [`cbor_tag_handler_t`](cbor_tag_handler_t.md) for more
information
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
CBOR does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -80,8 +88,9 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were
used in the given input or if the input is not valid CBOR
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of
other types are not supported, as JSON object keys are always strings), or if a string value or object key is not
valid UTF-8 and `error_handler` is `strict`
## Complexity
@@ -121,6 +130,7 @@ Linear in the size of the input.
- Added `tag_handler` parameter in version 3.9.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_msgpack(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_msgpack(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the MessagePack serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
MessagePack's specification explicitly allows ill-formed UTF-8, so checking is opt-in: the default, `keep`, does
not check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -73,8 +81,9 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from
MessagePack were used in the given input or if the input is not valid MessagePack
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of
other types are not supported, as JSON object keys are always strings), or if a string value or object key is not
valid UTF-8 and `error_handler` is `strict`
## Complexity
@@ -113,6 +122,7 @@ Linear in the size of the input.
- Added `allow_exceptions` parameter in version 3.2.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
+12 -3
View File
@@ -5,12 +5,14 @@
template<typename InputType>
static basic_json from_ubjson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
// (2)
template<typename IteratorType, typename SentinelType = IteratorType>
static basic_json from_ubjson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true);
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep);
```
Deserializes a given input to a JSON value using the UBJSON (Universal Binary JSON) serialization format.
@@ -58,6 +60,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`allow_exceptions` (in)
: whether to throw exceptions in case of a parse error (optional, `#!cpp true` by default)
`error_handler` (in)
: how to treat a string value or object key that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
UBJSON does not require a decoder to reject ill-formed UTF-8, so checking is opt-in: the default, `keep`, does not
check at all, as every binary reader did before this parameter was added; `strict` checks and throws;
`replace`/`ignore` sanitize the string the same way [`dump`](dump.md) would
## Return value
deserialized JSON value; in case of a parse error and `allow_exceptions` set to `#!cpp false`, the return value will be
@@ -73,7 +81,7 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if a parse error occurs
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string could not be parsed
successfully
successfully, or if a string value or object key is not valid UTF-8 and `error_handler` is `strict`
- Throws [out_of_range.408](../../home/exceptions.md#jsonexceptionout_of_range408) if the size of an optimized container
or n-dimensional array cannot be represented by `std::size_t`
@@ -112,6 +120,7 @@ Linear in the size of the input.
- Added `allow_exceptions` parameter in version 3.2.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
!!! warning "Deprecation"
@@ -23,10 +23,9 @@ type to use.
## Template parameters
`NumberFloatType`
: the type to store floating-point numbers. The parser converts `#!cpp float`, `#!cpp double`, and a
`#!cpp long double` that is IEEE 754 binary64 itself and other `#!cpp long double` formats with
`#!cpp std::from_chars` or `#!cpp std::strtold`, and serialization falls back to `#!cpp std::snprintf`, so the
type must be `#!cpp float`, `#!cpp double`, or `#!cpp long double`. The
: the type to store floating-point numbers. Parsing and serialization are implemented in terms of
`#!cpp std::strtof`/`#!cpp std::strtod`/`#!cpp std::strtold` and `#!cpp std::snprintf`, so the type must be
`#!cpp float`, `#!cpp double`, or `#!cpp long double`. The
[binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`,
because they have no encoding for `#!cpp long double`. See
[Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype).
+17 -4
View File
@@ -5,15 +5,18 @@
static std::vector<std::uint8_t> to_bjdata(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2);
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_bjdata(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2);
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
static void to_bjdata(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2);
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict);
```
Serializes a given JSON value `j` to a byte vector using the BJData (Binary JData) serialization format. BJData aims to
@@ -43,6 +46,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: which version of BJData to use (see note on "Binary values" on [BJData](../../features/binary_formats/bjdata.md));
optional, `#!cpp bjdata_version_t::draft2` by default.
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_bjdata` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. BJData serialization as byte vector
@@ -56,6 +65,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -89,4 +100,6 @@ Linear in the size of the JSON value `j`.
## Version history
- Added in version 3.11.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+17 -3
View File
@@ -2,11 +2,14 @@
```cpp
// (1)
static std::vector<std::uint8_t> to_bson(const basic_json& j);
static std::vector<std::uint8_t> to_bson(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o);
static void to_bson(const basic_json& j, detail::output_adapter<char> o);
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_bson(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict);
```
BSON (Binary JSON) is a binary format in which zero or more ordered key/value pairs are stored as a single entity (a
@@ -25,6 +28,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`o` (in)
: output adapter to write serialization to
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_bson` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. BSON serialization as a byte vector
@@ -46,6 +55,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`out_of_range.415`](../../home/exceptions.md#jsonexceptionout_of_range415) if the subtype of a binary value
exceeds 255, the maximum of the BSON binary subtype; example:
`"subtype 70000 is too large for the BSON binary subtype (max 255)"`
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -82,3 +93,6 @@ pass before anything is written.
- Added in version 3.4.0.
- Linear in the size of `j`, and no longer limited by the call stack for deeply nested values, since version 3.13.0.
- `out_of_range.415` is now detected before anything is written, like the other exceptions above, since version 3.13.0.
- Throwing `type_error.316` for a string value or object key that is not valid UTF-8, detected before anything is
written, added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+19 -3
View File
@@ -2,11 +2,14 @@
```cpp
// (1)
static std::vector<std::uint8_t> to_cbor(const basic_json& j);
static std::vector<std::uint8_t> to_cbor(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o);
static void to_cbor(const basic_json& j, detail::output_adapter<char> o);
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict);
static void to_cbor(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict);
```
Serializes a given JSON value `j` to a byte vector using the CBOR (Concise Binary Object Representation) serialization
@@ -26,6 +29,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
`o` (in)
: output adapter to write serialization to
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_cbor` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. CBOR serialization as a byte vector
@@ -35,6 +44,11 @@ The exact mapping and its limitations are described on a [dedicated page](../../
Strong guarantee: if an exception is thrown, there are no changes in the JSON value.
## Exceptions
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
Linear in the size of the JSON value `j`.
@@ -68,3 +82,5 @@ Linear in the size of the JSON value `j`.
- Added in version 2.0.9.
- Compact representation of floating-point numbers added in version 3.8.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
+16 -3
View File
@@ -4,13 +4,16 @@
// (1)
static std::vector<std::uint8_t> to_ubjson(const basic_json& j,
const bool use_size = false,
const bool use_type = false);
const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
// (2)
static void to_ubjson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false);
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
static void to_ubjson(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false);
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict);
```
Serializes a given JSON value `j` to a byte vector using the UBJSON (Universal Binary JSON) serialization format. UBJSON
@@ -36,6 +39,12 @@ The exact mapping and its limitations are described on a [dedicated page](../../
: whether to add type annotations to container types (must be combined with `#!cpp use_size = true`); optional,
`#!cpp false` by default.
`error_handler` (in)
: how to treat a string or object key in `j` that is not valid UTF-8; see [`error_handler_t`](error_handler_t.md).
The default, `strict`, throws; `keep` writes the ill-formed bytes to the output as is, as every version of
`to_ubjson` did before this parameter was added; `replace`/`ignore` sanitize it the same way
[`dump`](dump.md) would
## Return value
1. UBJSON serialization as a byte vector
@@ -49,6 +58,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8 and `error_handler` is `strict` (the default)
## Complexity
@@ -82,3 +93,5 @@ Linear in the size of the JSON value `j`.
## Version history
- Added in version 3.1.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
- Added `error_handler` parameter in version 3.13.0.
@@ -20,5 +20,6 @@ int main()
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::replace)
<< "\nstring with ignored invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
<< '\n';
<< "\nstring with the invalid byte kept as is (" << j_invalid.dump(-1, ' ', false, json::error_handler_t::keep).size()
<< " bytes, not valid UTF-8 itself)\n";
}
@@ -1,3 +1,4 @@
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
string with replaced invalid characters: "ä�ü"
string with ignored invalid characters: "äü"
string with the invalid byte kept as is (7 bytes, not valid UTF-8 itself)
@@ -63,6 +63,12 @@ The library uses the following mapping from JSON values types to BJData types ac
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
BJData strings must use UTF-8 encoding. `to_bjdata()` validates the bytes of every string value and object key
and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8, so a
value with such a string cannot be serialized in the first place.
!!! info "Unused BJData markers"
The following markers are not used in the conversion:
@@ -208,6 +214,20 @@ The library maps BJData types to JSON value types as follows:
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
BJData strings must use UTF-8 encoding, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_bjdata.md) parameter left at `keep` (the default), `from_bjdata()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_bjdata()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()`'s
own `error_handler` parameter defaults to `strict` (see above), so a value read this way cannot be written back
to BJData unless a non-strict handler is passed there too.
!!! info "Round trips"
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
@@ -109,14 +109,22 @@ The library maps BSON record types to JSON value types as follows:
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
`from_bson()`.
!!! warning "UTF-8 validation of string values"
!!! warning "Ill-formed UTF-8 in string values"
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
bytes of every such string at decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
(key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
byte-by-byte as a C string, or are not required to hold text, respectively.
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
decoder, so checking is opt-in: with the [`error_handler`](../../api/basic_json/from_bson.md) parameter left at
`keep` (the default), `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back
unchanged. Passing `error_handler_t::strict` makes `from_bson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores
the ill-formed bytes. `to_bson()`'s own `error_handler` parameter defaults to `strict` and throws the same
exception for a string value or element (key) name that is not valid UTF-8, so an object with such a key or
value cannot be produced in the first place unless a non-strict handler is passed there, even though
`from_bson()` would accept it from another source with the default `keep` handler. Element (key) names are
never validated on read, since they are read byte-by-byte as a C string. `binary` values (type `0x05`) are
unaffected, since they are not required to hold text.
??? example
@@ -189,15 +189,21 @@ The library maps CBOR types to JSON value types as follows:
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "UTF-8 validation of text strings"
!!! warning "Ill-formed UTF-8 in text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
text.
(major type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this, so checking is
opt-in: with the [`error_handler`](../../api/basic_json/from_cbor.md) parameter left at `keep` (the default),
`from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back
unchanged. Passing `error_handler_t::strict` makes `from_cbor()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)
for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores
the ill-formed bytes. `to_cbor()`'s own [`error_handler`](../../api/basic_json/to_cbor.md) parameter defaults
to `strict` and throws the same exception for a string value or object key that is not valid UTF-8, so such a
value cannot be written back to CBOR unless a non-strict handler is passed there too. Byte strings (major
type 2) are unaffected, since they are not required to hold text.
!!! warning "Tagged items"
@@ -153,14 +153,21 @@ The library maps MessagePack types to JSON value types as follows:
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values"
!!! warning "Ill-formed UTF-8 in string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
This library validates the bytes of every such string (object keys included) at decode time and rejects
ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
to hold text.
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
This library follows that by default: with its
[`error_handler`](../../api/basic_json/from_msgpack.md) parameter left at `keep` (the default),
`from_msgpack()` reads `str` bytes (object keys included) as-is, without validating them, so such a value
round-trips through `from_msgpack(to_msgpack(j))` byte for byte. Passing `error_handler_t::strict` makes
`from_msgpack()` check anyway and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. `to_msgpack()` itself has no `error_handler`
parameter and always writes `str` bytes as-is, since the specification permits it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way with the
default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes.
??? example
@@ -47,6 +47,12 @@ The library uses the following mapping from JSON values types to UBJSON types ac
- strings with more than 9223372036854775807 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
UBJSON's required string encoding is UTF-8. `to_ubjson()` validates the bytes of every string value and object
key and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8, so
a value with such a string cannot be serialized in the first place.
!!! info "Unused UBJSON markers"
The following markers are not used in the conversion:
@@ -120,6 +126,20 @@ The library maps UBJSON types to JSON value types as follows:
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
UBJSON's required string encoding is UTF-8, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_ubjson.md) parameter left at `keep` (the default), `from_ubjson()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_ubjson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()`'s
own `error_handler` parameter defaults to `strict` (see above), so a value read this way cannot be written back
to UBJSON unless a non-strict handler is passed there too.
??? example
```cpp
@@ -71,13 +71,10 @@ otherwise, it uses unsigned integer storage.
- Numbers with a decimal digit or scientific notation are always stored as `#!c double`.
- The number types can be changed, see [Template number types](#template-number-types).
- The library converts integers and floating-point numbers itself, independent of the locale. Floating-point
numbers are correctly rounded (to nearest, ties to even). Only a `#!c long double` that is not IEEE 754 binary64
(e.g., the 80-bit x87 format) is converted with `#!cpp std::from_chars` where available, or else with
[`std::strtold`](https://en.cppreference.com/w/cpp/string/byte/strtof). For that call, the library temporarily
replaces the `.` with the decimal point of the current locale (which may be longer than one byte, e.g., in
`fa_IR.UTF-8`), so the result does not depend on the locale either. Changing the locale in another thread during
parsing is undefined behavior of the C library, though.
- As of version 3.9.1, the conversion is realized by
[`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul),
[`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof), respectively.
!!! example "Examples"
@@ -88,10 +85,10 @@ otherwise, it uses unsigned integer storage.
### Number limits
- Any 64-bit signed or unsigned integer can be stored without loss of precision.
- Numbers exceeding the limits of `#!c double` (i.e., numbers whose rounded value is not satisfying
- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion via
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) are not satisfying
[`std::isfinite`](https://en.cppreference.com/w/cpp/numeric/math/isfinite) such as `#!c 1E400`) will throw exception
[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing. Numbers too
small for `#!c double` (such as `#!c 1E-400`) become zero, with the sign of the number.
[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing.
- Floating-point numbers are rounded to the next number representable as `double`. For instance
`#!c 3.141592653589793238462643383279` is stored as [`0x400921fb54442d18`](https://float.exposed/0x400921fb54442d18).
This is the same behavior as the code `#!c double x = 3.141592653589793238462643383279;`.
@@ -26,9 +26,8 @@ Requirements are split into two groups:
diagnosed with dedicated error messages, and violating most of them results in a compiler error somewhere inside
the library. Four violations are not caught at compile time at all:
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers
stored as a `#!cpp long double` that is not IEEE 754 binary64 (e.g., the 80-bit x87 format), because the lexer
hands the buffer to `#!cpp std::strtold`.
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers,
because the lexer hands the buffer to `#!cpp std::strtoull`/`#!cpp std::strtoll`/`#!cpp std::strtod`.
- A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation,
and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance.
- The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion
@@ -536,10 +535,8 @@ therefore silently changes parse results rather than raising an error. See
`NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`:
- The [parser](../parsing/index.md) converts number literals to `#!cpp float`, `#!cpp double`, and a
`#!cpp long double` that is IEEE 754 binary64 itself; other `#!cpp long double` formats are converted with
`#!cpp std::from_chars` where available, or with `#!cpp std::strtold`. The library provides overloads for exactly
these three types.
- The [parser](../parsing/index.md) converts number literals with `#!cpp std::strtof`, `#!cpp std::strtod`, or
`#!cpp std::strtold`; the library provides overloads for exactly these three types.
- [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
(`#!cpp float` is promoted to `#!cpp double`).
+4 -1
View File
@@ -341,7 +341,10 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8.
the string's bytes are not valid UTF-8 and the `error_handler` parameter of the corresponding `from_*` function is
set to `strict`. By default (`error_handler_t::keep`), the bytes of a string are not checked for valid UTF-8 on read;
see the ill-formed UTF-8 notes on the individual [binary format](../features/binary_formats/index.md) pages for how
such a string is handled depending on `error_handler`.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
+1 -1
View File
@@ -20,4 +20,4 @@ The class contains a slightly modified version of the Grisu2 algorithm from Flor
The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
+76 -41
View File
@@ -31,6 +31,7 @@
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/is_sax.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/string_utils.hpp>
#include <nlohmann/detail/value_t.hpp>
@@ -108,8 +109,16 @@ class binary_reader
@brief create a binary reader
@param[in] adapter input adapter to read from
@param[in] format the binary format to parse
@param[in] error_handler how to treat text strings and object keys that
are not well-formed UTF-8; none of the supported formats
requires a decoder to reject those, so the default is to
@ref error_handler_t::keep them unchanged, as every binary
reader did before this parameter existed
*/
explicit binary_reader(InputAdapterType&& adapter, const input_format_t format = input_format_t::json) noexcept : ia(std::move(adapter)), input_format(format)
explicit binary_reader(InputAdapterType&& adapter, const input_format_t format = input_format_t::json,
const error_handler_t error_handler = error_handler_t::keep) noexcept
: ia(std::move(adapter)), input_format(format), error_handler(error_handler)
{
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
}
@@ -428,7 +437,7 @@ class binary_reader
{
if (get_bson_cstr_bulk(result, std::integral_constant<bool, bulk_scan> {}))
{
return true;
return check_string_utf8(result, "key");
}
auto out = std::back_inserter(result);
@@ -441,7 +450,7 @@ class binary_reader
}
if (current == 0x00)
{
return true;
return check_string_utf8(result, "key");
}
*out++ = static_cast<typename string_t::value_type>(current);
}
@@ -522,7 +531,7 @@ class binary_reader
"string"), nullptr));
}
return true;
return check_string_utf8(result, "string");
}
/*!
@@ -1149,7 +1158,7 @@ class binary_reader
@return whether string creation completed
*/
bool get_cbor_string(string_t& result)
bool get_cbor_string(string_t& result, const char* context = "string")
{
// number of indefinite-length strings that have been opened and not
// closed yet. RFC 8949, Section 3.2.3 does not permit nesting them,
@@ -1179,7 +1188,7 @@ class binary_reader
{
if (--open == 0)
{
return true;
return check_string_utf8(result, context);
}
get();
continue;
@@ -1192,7 +1201,7 @@ class binary_reader
if (open == 0)
{
return true;
return check_string_utf8(result, context);
}
get();
@@ -1216,7 +1225,7 @@ class binary_reader
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
return get_cbor_string(result, "key");
}
const char* found = nullptr;
@@ -2004,7 +2013,7 @@ class binary_reader
@return whether string creation completed
*/
bool get_msgpack_string(string_t& result)
bool get_msgpack_string(string_t& result, const char* context = "string")
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::msgpack, "string")))
{
@@ -2047,25 +2056,25 @@ class binary_reader
case 0xBE:
case 0xBF:
{
return get_string(input_format_t::msgpack, static_cast<unsigned int>(current) & 0x1Fu, result);
return get_string(input_format_t::msgpack, static_cast<unsigned int>(current) & 0x1Fu, result) && check_string_utf8(result, context);
}
case 0xD9: // str 8
{
std::uint8_t len{};
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result);
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result) && check_string_utf8(result, context);
}
case 0xDA: // str 16
{
std::uint16_t len{};
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result);
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result) && check_string_utf8(result, context);
}
case 0xDB: // str 32
{
std::uint32_t len{};
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result);
return get_number(input_format_t::msgpack, len) && get_string(input_format_t::msgpack, len, result) && check_string_utf8(result, context);
}
default:
@@ -2143,7 +2152,7 @@ class binary_reader
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
return get_msgpack_string(result, "key");
}
if (current <= 0x7F || current >= 0xE0)
{
@@ -2159,7 +2168,7 @@ class binary_reader
}
else
{
return get_msgpack_string(result);
return get_msgpack_string(result, "key");
}
break;
}
@@ -2405,7 +2414,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key, true, "key") || !sax->key(key)))
{
return false;
}
@@ -2427,7 +2436,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key, false) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_ubjson_string(key, false, "key") || !sax->key(key)))
{
return false;
}
@@ -2495,7 +2504,7 @@ class binary_reader
@return whether string creation completed
*/
bool get_ubjson_string(string_t& result, const bool get_char = true)
bool get_ubjson_string(string_t& result, const bool get_char = true, const char* context = "string")
{
if (get_char)
{
@@ -2516,31 +2525,31 @@ class binary_reader
case 'U':
{
std::uint8_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'i':
{
std::int8_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'I':
{
std::int16_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'l':
{
std::int32_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'L':
{
std::int64_t len{};
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'u':
@@ -2550,7 +2559,7 @@ class binary_reader
break;
}
std::uint16_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'm':
@@ -2560,7 +2569,7 @@ class binary_reader
break;
}
std::uint32_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
case 'M':
@@ -2570,7 +2579,7 @@ class binary_reader
break;
}
std::uint64_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && get_string(input_format, len, result) && check_string_utf8(result, context);
}
default:
@@ -4031,27 +4040,50 @@ class binary_reader
const NumberType len,
string_t& result)
{
// get_bytes() appends to result, and CBOR indefinite-length strings
// collect all their chunks in the same result; validating only the
// newly read bytes keeps the check linear in the input size
const std::size_t old_size = result.size();
if (JSON_HEDLEY_UNLIKELY(!get_bytes(format, len, "string", result)))
// Strings are taken as is by default: none of CBOR (RFC 8949 §3.1
// leaves the choice to the decoder), MessagePack (whose spec
// explicitly allows a str object to contain an invalid byte
// sequence), UBJSON, BJData, or BSON requires a decoder to reject
// ill-formed UTF-8. Checking (and, with @ref error_handler_t::strict,
// rejecting, or with `replace`/`ignore`, sanitizing) is opt-in via
// @ref error_handler, applied once the whole string (all chunks of
// an indefinite-length CBOR string included) has been assembled, by
// @ref check_string_utf8 at the call site.
return get_bytes(format, len, "string", result);
}
/*!
@brief validate a decoded text string (value or object key) against @ref error_handler
None of the binary formats requires a decoder to reject ill-formed UTF-8
in a text string (see @ref get_string), so by default
(@ref error_handler_t::keep) this does nothing. A stricter
@ref error_handler opts into the same well-formedness check @ref
serializer::dump_escaped_impl applies when dumping a string:
@ref error_handler_t::strict rejects ill-formed input with
parse_error.113 (honoring `allow_exceptions` via @a sax), while
@ref error_handler_t::replace / @ref error_handler_t::ignore sanitize
@a result in place, using the exact same rules.
@param[in,out] result the already assembled string to check
@param[in] context further context information (for diagnostics)
@return whether @a result is acceptable (always true for `keep`)
*/
bool check_string_utf8(string_t& result, const char* context)
{
if (error_handler == error_handler_t::keep || is_valid_utf8(result))
{
return false;
return true;
}
// RFC 8949 (CBOR) §3.1 and the MessagePack/BSON/UBJSON specifications
// all require text strings to be valid UTF-8; reject anything else
// right here so malformed input is caught at decode time instead of
// only surfacing later as a type_error.316 when the value is dumped
// (which would defeat allow_exceptions=false / strict discarding).
if (JSON_HEDLEY_UNLIKELY(!is_valid_utf8(result, old_size)))
if (error_handler == error_handler_t::strict)
{
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(113, chars_read,
exception_message(format, "invalid string: ill-formed UTF-8 byte", "string"), nullptr));
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format, "invalid string: ill-formed UTF-8 byte", context), nullptr));
}
result = sanitize_utf8(result, error_handler);
return true;
}
@@ -4226,6 +4258,9 @@ class binary_reader
/// input format
const input_format_t input_format = input_format_t::json;
/// how to treat text strings/object keys that are not well-formed UTF-8
const error_handler_t error_handler = error_handler_t::keep;
/// the SAX parser
json_sax_t* sax = nullptr;
+10 -13
View File
@@ -1039,11 +1039,9 @@ class lexer : public lexer_base<BasicJsonType>
token_type::parse_error otherwise
@note The scanner is independent of the current locale: token_buffer
always holds `.`. The conversion of float and double does not use
the locale either. Only the std::strtold fallback of
convert_number() for long double formats other than binary64
depends on it, and it looks up the decimal point right before
converting (see detail::convert_float_locale_aware()).
always holds `.`. Only the std::strtod fallback of convert_number()
depends on the locale, and it looks up the decimal point right
before converting (see detail::convert_float_locale_aware()).
*/
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{
@@ -1056,7 +1054,7 @@ class lexer : public lexer_base<BasicJsonType>
// offset just past the last mantissa byte in token_buffer (i.e. the
// index of 'e'/'E', or the whole token when there is no exponent).
// convert_number() uses it to split the token; npos means
// convert_number() uses it to count significant digits; npos means
// "not seen an exponent yet" and is resolved at scan_number_done
std::size_t mantissa_end = std::string::npos;
@@ -1386,8 +1384,8 @@ scan_number_done:
@param[in] mantissa_end offset just past the last mantissa byte in
token_buffer (the index of 'e'/'E', or
token_buffer.size() when there is no exponent);
with decimal_point_position, it locates the parts
of a float token without scanning it again
used to skip Clinger's fast path when it cannot
possibly succeed - see detail::mantissa_fits_clinger()
*/
token_type convert_number(token_type number_type, std::size_t mantissa_end)
{
@@ -1441,11 +1439,10 @@ scan_number_done:
}
// this code is reached if we parse a floating-point number or if an
// integer conversion above overflowed. float and double (and long
// double where it is binary64) are converted by the library itself,
// correctly rounded and independent of the locale; other long double
// formats use std::from_chars when available, otherwise the
// locale-aware strtold.
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod/strtold.
if (convert_float_fast(num_begin, num_end, decimal_point_position, mantissa_end, value_float))
{
return token_type::value_float;
File diff suppressed because it is too large Load Diff
+167 -26
View File
@@ -26,6 +26,7 @@
#include <nlohmann/detail/input/binary_reader.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/string_utils.hpp>
@@ -93,8 +94,12 @@ class binary_writer
@param[in] sink output sink to write to (a value-type sink such as
output_vector_sink, or output_adapter_sink wrapping a
type-erased output adapter)
@param[in] error_handler_ how to treat a string value or object key that
is not valid UTF-8 (CBOR, UBJSON, BJData, and BSON only; never
consulted by @ref write_msgpack or @ref write_bon8)
*/
explicit binary_writer(OutputSinkType sink) : oa(std::move(sink))
explicit binary_writer(OutputSinkType sink, const error_handler_t error_handler_ = error_handler_t::strict)
: oa(std::move(sink)), error_handler(error_handler_)
{}
/*!
@@ -107,14 +112,20 @@ class binary_writer
from one.
@param[in] adapter output adapter to write to
@param[in] error_handler_ how to treat a string value or object key that
is not valid UTF-8 (CBOR, UBJSON, BJData, and BSON only; never
consulted by @ref write_msgpack or @ref write_bon8)
*/
template < typename SinkType = OutputSinkType,
typename std::enable_if < std::is_constructible<SinkType, output_adapter_t<CharType>>::value, int >::type = 0 >
explicit binary_writer(output_adapter_t<CharType> adapter) : oa(SinkType(std::move(adapter)))
explicit binary_writer(output_adapter_t<CharType> adapter, const error_handler_t error_handler_ = error_handler_t::strict)
: oa(SinkType(std::move(adapter))), error_handler(error_handler_)
{}
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
@throw type_error.317 if @a j is not an object
*/
void write_bson(const BasicJsonType& j)
@@ -145,6 +156,8 @@ class binary_writer
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_cbor(const BasicJsonType& j)
{
@@ -211,13 +224,16 @@ class binary_writer
case value_t::string:
{
string_t storage;
const string_t& value = sanitize_utf8_for_write(*j.m_data.m_value.string, j, storage);
// step 1: write control byte and the string length
write_cbor_head(0x60, j.m_data.m_value.string->size());
write_cbor_head(0x60, value.size());
// step 2: write the string
oa.write_characters(
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
reinterpret_cast<const CharType*>(value.data()),
value.size());
break;
}
@@ -287,6 +303,17 @@ class binary_writer
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object)
{
// el.first is checked here, against the object as
// diagnostics context, because write_cbor(el.first)
// converts it to a temporary basic_json that would be
// used as the context instead; for error_handler_t::keep
// and ::replace/::ignore the recursive write_cbor(el.first)
// call below handles the key like any other string, so no
// separate check is needed here for those
if (error_handler == error_handler_t::strict)
{
check_utf8(el.first, j);
}
write_cbor(el.first);
write_cbor(el.second);
}
@@ -629,6 +656,8 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true,
@@ -678,14 +707,17 @@ class binary_writer
case value_t::string:
{
string_t storage;
const string_t& value = sanitize_utf8_for_write(*j.m_data.m_value.string, j, storage);
if (add_prefix)
{
oa.write_character(to_char_type('S'));
}
write_number_with_ubjson_prefix(j.m_data.m_value.string->size(), true, use_bjdata);
write_number_with_ubjson_prefix(value.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(j.m_data.m_value.string->data()),
j.m_data.m_value.string->size());
reinterpret_cast<const CharType*>(value.data()),
value.size());
break;
}
@@ -840,10 +872,12 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object)
{
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
string_t storage;
const string_t& key = sanitize_utf8_for_write(el.first, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(el.first.data()),
el.first.size());
reinterpret_cast<const CharType*>(key.data()),
key.size());
write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
}
@@ -884,8 +918,12 @@ class binary_writer
/*!
@return The size of a BSON document entry header, including the id marker
and the entry name size (and its null-terminator).
@throw out_of_range.409 if @a name contains U+0000, before anything is
written
@throw type_error.316 if @a name is not valid UTF-8, before anything is
written
*/
static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
{
const auto it = name.find(static_cast<typename string_t::value_type>(0));
if (JSON_HEDLEY_UNLIKELY(it != BasicJsonType::string_t::npos))
@@ -893,8 +931,10 @@ class binary_writer
JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j));
}
static_cast<void>(j);
return /*id*/ 1ul + name.size() + /*zero-terminator*/1u;
string_t storage;
const string_t& sanitized = sanitize_utf8_for_write(name, j, storage);
return /*id*/ 1ul + sanitized.size() + /*zero-terminator*/1u;
}
/*!
@@ -914,14 +954,28 @@ class binary_writer
/*!
@brief Writes the given @a element_type and @a name to the output adapter
@a name has already been validated (and, for @ref error_handler_t::strict,
found well-formed) by @ref calc_bson_entry_header_size during the earlier
size pass, so only @ref error_handler_t::replace / @ref
error_handler_t::ignore need to sanitize it again here, to actually write
the bytes that size was computed from.
*/
void write_bson_entry_header(const string_t& name,
const std::uint8_t element_type)
{
oa.write_character(to_char_type(element_type));
oa.write_characters(
reinterpret_cast<const CharType*>(name.data()),
name.size());
if (error_handler == error_handler_t::keep || error_handler == error_handler_t::strict || is_valid_utf8(name))
{
oa.write_characters(reinterpret_cast<const CharType*>(name.data()), name.size());
}
else
{
const string_t sanitized = sanitize_utf8(name, error_handler);
oa.write_characters(reinterpret_cast<const CharType*>(sanitized.data()), sanitized.size());
}
// the terminating null byte is written explicitly rather than taken
// from the buffer, so that string_t::data() need not be null-terminated
oa.write_character(to_char_type(0x00));
@@ -949,24 +1003,50 @@ class binary_writer
/*!
@return The size of the BSON-encoded string in @a value
@throw type_error.316 if @a value is not valid UTF-8, before anything is
written
@note The UTF-8 check is skipped if @a value is already too long for the
32-bit BSON length field (@ref to_bson_length rejects it later, once
the size of the whole document is known); this also keeps the check
from reading past a StringType that reports a size larger than what
it actually holds.
*/
static std::size_t calc_bson_string_size(const string_t& value)
std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<std::int32_t>(value.size())))
{
string_t storage;
const string_t& sanitized = sanitize_utf8_for_write(value, j, storage);
return sizeof(std::int32_t) + sanitized.size() + 1ul;
}
return sizeof(std::int32_t) + value.size() + 1ul;
}
/*!
@brief Writes a BSON element with key @a name and string value @a value
@a value has already been validated (and, for @ref error_handler_t::strict,
found well-formed) by @ref calc_bson_string_size during the earlier size
pass, so only @ref error_handler_t::replace / @ref error_handler_t::ignore
need to sanitize it again here, to actually write the bytes that size was
computed from.
*/
void write_bson_string(const string_t& name,
const string_t& value)
{
write_bson_entry_header(name, 0x02);
write_number<std::int32_t>(to_bson_length(value.size() + 1ul), true);
const bool sanitize = error_handler != error_handler_t::keep
&& error_handler != error_handler_t::strict
&& !is_valid_utf8(value);
const string_t sanitized = sanitize ? sanitize_utf8(value, error_handler) : string_t{};
const string_t& written = sanitize ? sanitized : value;
write_number<std::int32_t>(to_bson_length(written.size() + 1ul), true);
oa.write_characters(
reinterpret_cast<const CharType*>(value.data()),
value.size());
reinterpret_cast<const CharType*>(written.data()),
written.size());
// the terminating null byte is written explicitly rather than taken
// from the buffer, so that string_t::data() need not be null-terminated
oa.write_character(to_char_type(0x00));
@@ -1080,8 +1160,10 @@ class binary_writer
is neither an object nor an array
@throw out_of_range.415 if @a j is binary with a subtype that does not fit
into a byte, before anything is written
@throw type_error.316 if @a j is a string that is not valid UTF-8, before
anything is written
*/
static std::size_t calc_bson_value_size(const BasicJsonType& j)
std::size_t calc_bson_value_size(const BasicJsonType& j)
{
switch (j.type())
{
@@ -1101,7 +1183,7 @@ class binary_writer
return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned);
case value_t::string:
return calc_bson_string_size(*j.m_data.m_value.string);
return calc_bson_string_size(*j.m_data.m_value.string, j);
case value_t::null:
return 0ul;
@@ -1214,8 +1296,10 @@ class binary_writer
written
@throw out_of_range.415 if a binary value's subtype does not fit into a
byte, before anything is written
@throw type_error.316 if a string value or a key is not valid UTF-8,
before anything is written
*/
static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
{
// the object or array whose entries are being sized, and the ones it
// is in; nothing is allocated unless the document nests
@@ -2092,7 +2176,7 @@ class binary_writer
*/
void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context)
{
check_bon8_utf8(s, context);
check_utf8(s, context);
// a string that follows another string terminates it
if (string_open)
@@ -2122,7 +2206,7 @@ class binary_writer
@throw type_error.316 if @a s is not valid UTF-8; the message names the
first byte of the first invalid or incomplete sequence
*/
static void check_bon8_utf8(const string_t& s, const BasicJsonType& context)
static void check_utf8(const string_t& s, const BasicJsonType& context)
{
static_cast<void>(context); // only used when exceptions are enabled
const auto* data = reinterpret_cast<const unsigned char*>(s.data());
@@ -2133,6 +2217,59 @@ class binary_writer
}
}
/*!
@brief return @a s as it should be written, honoring @ref error_handler
Used by @ref write_cbor, @ref write_ubjson (and so @ref write_bjdata), and
the BSON writing functions for string values and object keys; never by
@ref write_msgpack or @ref write_bon8, which do not take an @ref
error_handler (MessagePack's spec allows a str object to contain
ill-formed UTF-8, and BON8 always validates, since UTF-8 lead bytes are
structural there).
- @ref error_handler_t::keep: @a s is returned unchanged, without even
checking it (the behavior of release 3.12.0 and earlier).
- @ref error_handler_t::strict: @ref check_utf8 is called, which throws
type_error.316 if @a s is not valid UTF-8.
- @ref error_handler_t::replace / @ref error_handler_t::ignore: @a s is
sanitized into @a storage with exactly the rules @ref
serializer::dump_escaped_impl uses, so that parsing what @ref
basic_json::dump produces for the same string and the same handler
yields the same result.
Well-formed input is never copied: this returns a reference to @a s
itself in every case but a sanitized `replace`/`ignore` one, so @a
storage must outlive the returned reference only then.
@param[in] s the string (value or object key) to write
@param[in] context the value @a s belongs to (for diagnostics)
@param[out] storage backing storage for a sanitized copy
@return a reference to @a s, or to @a storage once it holds a sanitized copy
*/
const string_t& sanitize_utf8_for_write(const string_t& s, const BasicJsonType& context, string_t& storage) const
{
switch (error_handler)
{
case error_handler_t::keep:
return s;
case error_handler_t::strict:
check_utf8(s, context);
return s;
case error_handler_t::replace:
case error_handler_t::ignore:
default:
if (is_valid_utf8(s))
{
return s;
}
storage = sanitize_utf8(s, error_handler);
return storage;
}
}
/*!
@brief write an integer in the shortest encoding
@@ -2461,6 +2598,10 @@ class binary_writer
/// the output
OutputSinkType oa;
/// how to treat a string value or object key that is not valid UTF-8
/// (CBOR, UBJSON, BJData, and BSON only)
const error_handler_t error_handler = error_handler_t::strict;
};
} // namespace detail
@@ -0,0 +1,38 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/// how to treat decoding errors
///
/// @ref basic_json::dump uses this to decide what to do with ill-formed
/// UTF-8 while escaping a string, and the binary writers (@ref
/// basic_json::to_cbor, @ref basic_json::to_ubjson, @ref
/// basic_json::to_bjdata, @ref basic_json::to_bson) use it the same way for
/// string values and object keys. The binary readers (@ref
/// basic_json::from_cbor, @ref basic_json::from_msgpack, @ref
/// basic_json::from_ubjson, @ref basic_json::from_bjdata, @ref
/// basic_json::from_bson) use it to decide whether to check text strings
/// and object keys for well-formed UTF-8 at all, since none of those
/// formats requires a decoder to do so.
enum class error_handler_t
{
strict, ///< throw a type_error/parse_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences unchanged
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+62 -10
View File
@@ -27,6 +27,7 @@
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/cpp_future.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/recursion_depth_limit.hpp>
#include <nlohmann/detail/string_concat.hpp>
@@ -41,14 +42,6 @@ namespace detail
// serialization //
///////////////////
/// how to treat decoding errors
enum class error_handler_t
{
strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore ///< ignore invalid UTF-8 sequences
};
template<typename BasicJsonType>
class serializer
{
@@ -839,6 +832,16 @@ class serializer
// EnsureAscii parameter is used, non-ASCII characters
if ((codepoint <= 0x1F) || (EnsureAscii && (codepoint >= 0x7F)))
{
if (EnsureAscii && error_handler == error_handler_t::keep)
{
// this character was buffered as raw bytes
// below in case it turned out to be part of
// an ill-formed sequence (which is kept as
// is); now that it decoded to a well-formed
// code point, undo that and \u-escape it
// like any other character instead
bytes = bytes_after_last_accept;
}
if (codepoint <= 0xFFFF)
{
write_u_escape(bytes, static_cast<std::uint16_t>(codepoint));
@@ -937,6 +940,44 @@ class serializer
break;
}
case error_handler_t::keep:
{
// the bytes of this (now abandoned) ill-formed
// sequence seen so far are already buffered below
// and are kept unchanged in the output
if (undumped_chars > 0)
{
// the byte that ended the sequence may be OK
// for itself (e.g., a quote that must still be
// escaped, or the lead byte of a well-formed
// code point), so read it again
--i;
}
else
{
// a byte that cannot start a sequence (e.g.,
// 0xFF or a stray continuation byte) is kept
// as well
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -945,9 +986,12 @@ class serializer
default: // decode found yet incomplete multibyte code point
{
if (!EnsureAscii)
if (!EnsureAscii || error_handler == error_handler_t::keep)
{
// code point will not be escaped - copy byte to buffer
// code point will not be escaped (or will be kept as
// is if it turns out to be ill-formed) - copy byte to
// buffer; dropped again above if it decodes to a
// well-formed code point that needs \u-escaping
string_buffer[bytes++] = s[i];
}
++undumped_chars;
@@ -998,6 +1042,14 @@ class serializer
break;
}
case error_handler_t::keep:
{
// write the ill-formed trailing bytes as is; they were
// buffered above regardless of EnsureAscii
put_buffer(string_buffer, bytes);
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
+113 -15
View File
@@ -16,6 +16,7 @@
#include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/error_handler.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -117,13 +118,14 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
The library checks UTF-8 well-formedness (RFC 3629, section 4) in three
places, which differ in speed, diagnostics, and how they read the input:
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
text strings at decode time).
- decode() below: the serializer, to escape and, in strict mode, reject
ill-formed UTF-8 when dumping a string. The CBOR, MessagePack, BSON,
UBJSON and BJData readers do not use it: none of those specs requires a
decoder to reject ill-formed UTF-8 in text strings, so the readers keep
the bytes as is and leave the check to dump() and the binary writers.
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
@@ -179,19 +181,19 @@ inline std::uint8_t decode(std::uint8_t& state, std::uint32_t& codep, const std:
}
/*!
@brief check whether a string consists solely of valid UTF-8
@brief check a string for well-formed UTF-8 (RFC 3629, section 4)
Used by the CBOR/MessagePack/BSON/UBJSON binary readers to reject text
strings that are not valid UTF-8 at decode time (RFC 8949 §3.1 and the
MessagePack/BSON specifications all require text strings to be UTF-8), so
that malformed input is caught immediately instead of only surfacing later
as a type_error.316 when the resulting value is dumped.
Used by the binary readers (CBOR, MessagePack, UBJSON, BJData, BSON) when an
@ref error_handler_t other than `keep` is requested for a text string value
or object key: none of those formats requires a decoder to reject ill-formed
UTF-8 on its own, so the check is opt-in there, unlike the JSON lexer and the
serializer's @ref decode -based escaping, which always run it.
@param[in] s the string to check
@param[in] first index of the first byte to check; the bytes before it are
assumed to have been validated already and to end on a
code point boundary
@return whether @a s (from index @a first on) is valid UTF-8
@param[in] first the index to start checking at
@return whether `s.substr(first)` is well-formed UTF-8
@sa @ref decode
*/
template<typename StringType>
inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noexcept
@@ -211,5 +213,101 @@ inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noex
return state == UTF8_ACCEPT;
}
/*!
@brief sanitize a string with ill-formed UTF-8 for @ref error_handler_t::replace or @ref error_handler_t::ignore
Replaces every maximal ill-formed subsequence with U+FFFD (`replace`) or
drops it (`ignore`), using exactly the same boundaries @ref
serializer::dump_escaped_impl uses while escaping a string: a byte that does
not extend the sequence started by the previous byte(s) is reread as the
start of a new one, instead of being swallowed along with them.
@pre @a error_handler is @ref error_handler_t::replace or @ref error_handler_t::ignore
@note Well-formed input is copied through unchanged, including bytes (e.g.
control characters or quotes) that @ref serializer::dump_escaped_impl
would itself escape; this function only concerns itself with
well-formedness, not with producing valid JSON text.
@param[in] s the string to sanitize
@param[in] error_handler @ref error_handler_t::replace or @ref error_handler_t::ignore
@return @a s with every ill-formed subsequence replaced or removed
@sa @ref decode
*/
template<typename StringType>
inline StringType sanitize_utf8(const StringType& s, const error_handler_t error_handler)
{
JSON_ASSERT(error_handler == error_handler_t::replace || error_handler == error_handler_t::ignore);
StringType result;
result.reserve(s.size());
std::uint32_t codepoint = 0;
std::uint8_t state = UTF8_ACCEPT;
// length of result after the last accepted code point
std::size_t result_len_after_last_accept = 0;
// whether bytes of an as yet unresolved sequence were already appended
bool pending = false;
for (std::size_t i = 0; i < s.size(); ++i)
{
switch (decode(state, codepoint, static_cast<std::uint8_t>(s[i])))
{
case UTF8_ACCEPT: // decode found a well-formed code point
{
result.push_back(s[i]);
result_len_after_last_accept = result.size();
pending = false;
break;
}
case UTF8_REJECT: // decode found an ill-formed byte
{
// in case we saw this byte for the first time, read it again,
// because it may be fine for itself, just not for the
// sequence that came before it
if (pending)
{
--i;
}
// drop the bytes of the ill-formed sequence buffered below
result.resize(result_len_after_last_accept);
if (error_handler == error_handler_t::replace)
{
result.append("\xEF\xBF\xBD");
result_len_after_last_accept = result.size();
}
pending = false;
state = UTF8_ACCEPT;
break;
}
default: // decode found yet incomplete multibyte code point
{
result.push_back(s[i]);
pending = true;
break;
}
}
}
// the string ended with an incomplete sequence
if (state != UTF8_ACCEPT)
{
result.resize(result_len_after_last_accept);
if (error_handler == error_handler_t::replace)
{
result.append("\xEF\xBF\xBD");
}
}
return result;
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+69 -46
View File
@@ -201,9 +201,10 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// used by the vector-returning to_* overloads
template<typename CharType> using vector_binary_writer =
::nlohmann::detail::binary_writer<basic_json, CharType, ::nlohmann::detail::output_vector_sink<CharType>>;
template<typename CharType> static vector_binary_writer<CharType> vector_writer(std::vector<CharType>& v)
template<typename CharType> static vector_binary_writer<CharType> vector_writer(
std::vector<CharType>& v, const ::nlohmann::detail::error_handler_t error_handler = ::nlohmann::detail::error_handler_t::strict)
{
return vector_binary_writer<CharType>(::nlohmann::detail::output_vector_sink<CharType>(v));
return vector_binary_writer<CharType>(::nlohmann::detail::output_vector_sink<CharType>(v), error_handler);
}
JSON_PRIVATE_UNLESS_TESTED:
@@ -5445,26 +5446,29 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
public:
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static std::vector<std::uint8_t> to_cbor(const basic_json& j)
static std::vector<std::uint8_t> to_cbor(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_cbor(j);
vector_writer(result, error_handler).write_cbor(j);
return result;
}
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o)
static void to_cbor(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_cbor(j);
binary_writer<std::uint8_t>(o, error_handler).write_cbor(j);
}
/// @brief create a CBOR serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_cbor/
static void to_cbor(const basic_json& j, detail::output_adapter<char> o)
static void to_cbor(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_cbor(j);
binary_writer<char>(o, error_handler).write_cbor(j);
}
/// @brief create a MessagePack serialization of a given JSON value
@@ -5495,28 +5499,31 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static std::vector<std::uint8_t> to_ubjson(const basic_json& j,
const bool use_size = false,
const bool use_type = false)
const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_ubjson(j, use_size, use_type);
vector_writer(result, error_handler).write_ubjson(j, use_size, use_type);
return result;
}
/// @brief create a UBJSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static void to_ubjson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false)
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_ubjson(j, use_size, use_type);
binary_writer<std::uint8_t>(o, error_handler).write_ubjson(j, use_size, use_type);
}
/// @brief create a UBJSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_ubjson/
static void to_ubjson(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false)
const bool use_size = false, const bool use_type = false,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_ubjson(j, use_size, use_type);
binary_writer<char>(o, error_handler).write_ubjson(j, use_size, use_type);
}
/// @brief create a BJData serialization of a given JSON value
@@ -5524,11 +5531,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static std::vector<std::uint8_t> to_bjdata(const basic_json& j,
const bool use_size = false,
const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2)
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_ubjson(j, use_size, use_type, true, true, version);
vector_writer(result, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
return result;
}
@@ -5536,42 +5544,47 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/to_bjdata/
static void to_bjdata(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2)
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_ubjson(j, use_size, use_type, true, true, version);
binary_writer<std::uint8_t>(o, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
}
/// @brief create a BJData serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bjdata/
static void to_bjdata(const basic_json& j, detail::output_adapter<char> o,
const bool use_size = false, const bool use_type = false,
const bjdata_version_t version = bjdata_version_t::draft2)
const bjdata_version_t version = bjdata_version_t::draft2,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_ubjson(j, use_size, use_type, true, true, version);
binary_writer<char>(o, error_handler).write_ubjson(j, use_size, use_type, true, true, version);
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static std::vector<std::uint8_t> to_bson(const basic_json& j)
static std::vector<std::uint8_t> to_bson(const basic_json& j,
const error_handler_t error_handler = error_handler_t::strict)
{
std::vector<std::uint8_t> result;
result.reserve(detail::binary_reserve_hint(j));
vector_writer(result).write_bson(j);
vector_writer(result, error_handler).write_bson(j);
return result;
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o)
static void to_bson(const basic_json& j, detail::output_adapter<std::uint8_t> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<std::uint8_t>(o).write_bson(j);
binary_writer<std::uint8_t>(o, error_handler).write_bson(j);
}
/// @brief create a BSON serialization of a given JSON value
/// @sa https://json.nlohmann.me/api/basic_json/to_bson/
static void to_bson(const basic_json& j, detail::output_adapter<char> o)
static void to_bson(const basic_json& j, detail::output_adapter<char> o,
const error_handler_t error_handler = error_handler_t::strict)
{
binary_writer<char>(o).write_bson(j);
binary_writer<char>(o, error_handler).write_bson(j);
}
/// @brief create a BON8 serialization of a given JSON value
@@ -5605,12 +5618,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json from_cbor(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor, error_handler).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5625,12 +5639,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static basic_json from_cbor(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true,
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor, error_handler).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5672,12 +5687,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_msgpack(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5691,12 +5707,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_msgpack(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5736,12 +5753,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_ubjson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5755,12 +5773,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_ubjson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5800,12 +5819,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bjdata(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5819,12 +5839,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bjdata(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5874,12 +5895,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bson(InputType&& i,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
@@ -5893,12 +5915,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
JSON_HEDLEY_WARN_UNUSED_RESULT
static basic_json from_bson(IteratorType first, SentinelType last,
const bool strict = true,
const bool allow_exceptions = true)
const bool allow_exceptions = true,
const error_handler_t error_handler = error_handler_t::keep)
{
basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson, error_handler).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{
result = value_t::discarded;
}
File diff suppressed because it is too large Load Diff
-599
View File
@@ -1,599 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstdint> // uint32_t, uint64_t
// Number tokens that are hard to round correctly, with the IEEE-754 binary64
// and binary32 bits of their correctly rounded values (ties to even; infinity
// for an overflow, a signed zero for an underflow).
//
// For doubles and floats around 0, the smallest normal number, 1, 2^24, 2^53,
// 0.1, and the largest finite number, and for random ones, the exact midpoint
// m to the next number gives: m, m with one unit more and less in the last
// digit, m with "01" and "0...01" appended, m with trailing zeros, and m cut
// after 17 to 30 digits (rounded down and up, so that the rounding is decided
// after the 19th digit), in fixed and exponent notation, 30% of them negative.
// Tokens longer than 80 characters are left out, except for four of 700 digits
// and more. Zeros, underflow, overflow, huge exponents, and integers beyond 64
// bits complete the set. Of the 508 tokens, 134 (as double) and 150 (as
// float) need the exact comparison with the midpoint (detail::digit_comparison()).
//
// The expected bits were computed with exact rational arithmetic in Python
// (fractions.Fraction) and cross-checked with Python's float(); strtod_l and
// strtof_l of Apple's libc and of glibc agree. Generated by
// compact_hard_cases.py 5 (with hard_cases.py), see the pull request that
// added this file.
namespace float_hard_cases
{
struct hard_case
{
const char* token;
std::uint64_t bits64;
std::uint32_t bits32;
};
inline const std::array<hard_case, 508>& cases()
{
static const std::array<hard_case, 508> table =
{
{
{"-2.4703282292062327e-324", 0x8000000000000000u, 0x80000000u},
{"24703282292062328e-340", 0x0000000000000001u, 0x00000000u},
{"247032822920623272e-341", 0x0000000000000000u, 0x00000000u},
{"-0.2470328229206232721e-323", 0x8000000000000001u, 0x80000000u},
{"-0.24703282292062327208e-323", 0x8000000000000000u, 0x80000000u},
{"-2.4703282292062327209e-324", 0x8000000000000001u, 0x80000000u},
{"2.47032822920623272088e-324", 0x0000000000000000u, 0x00000000u},
{"247032822920623272089e-344", 0x0000000000000001u, 0x00000000u},
{"-247032822920623272088284396434e-353", 0x8000000000000000u, 0x80000000u},
{"0.247032822920623272088284396435e-323", 0x0000000000000001u, 0x00000000u},
{"-74109846876186981e-340", 0x8000000000000001u, 0x80000000u},
{"0.74109846876186982e-323", 0x0000000000000002u, 0x00000000u},
{"-0.7410984687618698162e-323", 0x8000000000000001u, 0x80000000u},
{"-7.410984687618698163e-324", 0x8000000000000002u, 0x80000000u},
{"7.4109846876186981626e-324", 0x0000000000000001u, 0x00000000u},
{"-74109846876186981627e-343", 0x8000000000000002u, 0x80000000u},
{"-741098468761869816264e-344", 0x8000000000000001u, 0x80000000u},
{"0.741098468761869816265e-323", 0x0000000000000002u, 0x00000000u},
{"0.741098468761869816264853189302e-323", 0x0000000000000001u, 0x00000000u},
{"-7.41098468761869816264853189303e-324", 0x8000000000000002u, 0x80000000u},
{"0.22250738585072006e-307", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"2.2250738585072007e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"2.225073858507200641e-308", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"-2225073858507200642e-326", 0x800FFFFFFFFFFFFFu, 0x80000000u},
{"22250738585072006419e-327", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"0.2225073858507200642e-307", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"0.222507385850720064199e-307", 0x000FFFFFFFFFFFFEu, 0x00000000u},
{"2.225073858507200642e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"-2.22507385850720064199176395546e-308", 0x800FFFFFFFFFFFFEu, 0x80000000u},
{"222507385850720064199176395547e-337", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"-2.2250738585072011e-308", 0x800FFFFFFFFFFFFFu, 0x80000000u},
{"-22250738585072012e-324", 0x8010000000000000u, 0x80000000u},
{"-2225073858507201136e-326", 0x800FFFFFFFFFFFFFu, 0x80000000u},
{"0.2225073858507201137e-307", 0x0010000000000000u, 0x00000000u},
{"0.2225073858507201136e-307", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"-2.2250738585072011361e-308", 0x8010000000000000u, 0x80000000u},
{"2.22507385850720113605e-308", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"222507385850720113606e-328", 0x0010000000000000u, 0x00000000u},
{"22250738585072011360574097967e-336", 0x000FFFFFFFFFFFFFu, 0x00000000u},
{"0.222507385850720113605740979671e-307", 0x0010000000000000u, 0x00000000u},
{"22250738585072016e-324", 0x0010000000000000u, 0x00000000u},
{"0.22250738585072017e-307", 0x0010000000000001u, 0x00000000u},
{"0.222507385850720163e-307", 0x0010000000000000u, 0x00000000u},
{"2.225073858507201631e-308", 0x0010000000000001u, 0x00000000u},
{"-2.2250738585072016301e-308", 0x8010000000000000u, 0x80000000u},
{"22250738585072016302e-327", 0x0010000000000001u, 0x00000000u},
{"-222507385850720163012e-328", 0x8010000000000000u, 0x80000000u},
{"0.222507385850720163013e-307", 0x0010000000000001u, 0x00000000u},
{"0.222507385850720163012305563795e-307", 0x0010000000000000u, 0x00000000u},
{"-2.22507385850720163012305563796e-308", 0x8010000000000001u, 0x80000000u},
{"0.17976931348623156E+309", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"1.7976931348623157e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"1.797693134862315608e308", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"-1797693134862315609e290", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u},
{"-17976931348623156083e289", 0xFFEFFFFFFFFFFFFEu, 0xFF800000u},
{"-0.17976931348623156084E+309", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u},
{"0.179769313486231560835E+309", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"-1.79769313486231560836e308", 0xFFEFFFFFFFFFFFFFu, 0xFF800000u},
{"1.79769313486231560835325876058e308", 0x7FEFFFFFFFFFFFFEu, 0x7F800000u},
{"179769313486231560835325876059e279", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"1.7976931348623158e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"17976931348623159e292", 0x7FF0000000000000u, 0x7F800000u},
{"1797693134862315807e290", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"0.1797693134862315808E+309", 0x7FF0000000000000u, 0x7F800000u},
{"0.17976931348623158079E+309", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"-1.797693134862315808e308", 0xFFF0000000000000u, 0xFF800000u},
{"1.79769313486231580793e308", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"179769313486231580794e288", 0x7FF0000000000000u, 0x7F800000u},
{"179769313486231580793728971405e279", 0x7FEFFFFFFFFFFFFFu, 0x7F800000u},
{"-0.179769313486231580793728971406E+309", 0xFFF0000000000000u, 0xFF800000u},
{"100000000000000011102230246251565404236316680908203125e-53", 0x3FF0000000000000u, 0x3F800000u},
{"-1.00000000000000011102230246251565404236316680908203126", 0xBFF0000000000001u, 0xBF800000u},
{"1.00000000000000011102230246251565404236316680908203124e0", 0x3FF0000000000000u, 0x3F800000u},
{"10000000000000001110223024625156540423631668090820312501e-55", 0x3FF0000000000001u, 0x3F800000u},
{"1.00000000000000011102230246251565404236316680908203125000000000000000000001", 0x3FF0000000000001u, 0x3F800000u},
{"10000000000000001e-16", 0x3FF0000000000000u, 0x3F800000u},
{"1.0000000000000002", 0x3FF0000000000001u, 0x3F800000u},
{"1.000000000000000111", 0x3FF0000000000000u, 0x3F800000u},
{"1.000000000000000112e0", 0x3FF0000000000001u, 0x3F800000u},
{"1.000000000000000111e0", 0x3FF0000000000000u, 0x3F800000u},
{"-10000000000000001111e-19", 0xBFF0000000000001u, 0xBF800000u},
{"-100000000000000011102e-20", 0xBFF0000000000000u, 0xBF800000u},
{"-1.00000000000000011103", 0xBFF0000000000001u, 0xBF800000u},
{"1.00000000000000011102230246251", 0x3FF0000000000000u, 0x3F800000u},
{"1.00000000000000011102230246252e0", 0x3FF0000000000001u, 0x3F800000u},
{"-0.999999999999999944488848768742172978818416595458984375", 0xBFF0000000000000u, 0xBF800000u},
{"-9.99999999999999944488848768742172978818416595458984376e-1", 0xBFF0000000000000u, 0xBF800000u},
{"999999999999999944488848768742172978818416595458984374e-54", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"0.99999999999999994448884876874217297881841659545898437501", 0x3FF0000000000000u, 0x3F800000u},
{"9.99999999999999944488848768742172978818416595458984375000000000000000000001e-1", 0x3FF0000000000000u, 0x3F800000u},
{"-0.99999999999999994", 0xBFEFFFFFFFFFFFFFu, 0xBF800000u},
{"9.9999999999999995e-1", 0x3FF0000000000000u, 0x3F800000u},
{"9.999999999999999444e-1", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"9999999999999999445e-19", 0x3FF0000000000000u, 0x3F800000u},
{"99999999999999994448e-20", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"-0.99999999999999994449", 0xBFF0000000000000u, 0xBF800000u},
{"0.999999999999999944488", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"-9.99999999999999944489e-1", 0xBFF0000000000000u, 0xBF800000u},
{"9.99999999999999944488848768742e-1", 0x3FEFFFFFFFFFFFFFu, 0x3F800000u},
{"999999999999999944488848768743e-30", 0x3FF0000000000000u, 0x3F800000u},
{"-9.007199254740993e15", 0xC340000000000000u, 0xDA000000u},
{"9007199254740994e0", 0x4340000000000001u, 0x5A000000u},
{"9007199254740992", 0x4340000000000000u, 0x5A000000u},
{"9.00719925474099301e15", 0x4340000000000001u, 0x5A000000u},
{"9007199254740993000000000000000000001e-21", 0x4340000000000001u, 0x5A000000u},
{"-9007199254740993.000000000000000000000000000000", 0xC340000000000000u, 0xDA000000u},
{"90071992547409915e-1", 0x4340000000000000u, 0x5A000000u},
{"-9007199254740991.6", 0xC340000000000000u, 0xDA000000u},
{"9.0071992547409914e15", 0x433FFFFFFFFFFFFFu, 0x5A000000u},
{"-9007199254740991501e-3", 0xC340000000000000u, 0xDA000000u},
{"9007199254740991.5000000000000000000001", 0x4340000000000000u, 0x5A000000u},
{"9.0071992547409915000000000000000000000000000000e15", 0x4340000000000000u, 0x5A000000u},
{"0.100000000000000012490009027033011079765856266021728515625", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"1.00000000000000012490009027033011079765856266021728515626e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"100000000000000012490009027033011079765856266021728515624e-57", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"0.10000000000000001249000902703301107976585626602172851562501", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"0.10000000000000001", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"1.0000000000000002e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"-1.000000000000000124e-1", 0xBFB999999999999Au, 0xBDCCCCCDu},
{"1000000000000000125e-19", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"-10000000000000001249e-20", 0xBFB999999999999Au, 0xBDCCCCCDu},
{"0.1000000000000000125", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"0.10000000000000001249", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"1.00000000000000012491e-1", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"1.00000000000000012490009027033e-1", 0x3FB999999999999Au, 0x3DCCCCCDu},
{"100000000000000012490009027034e-30", 0x3FB999999999999Bu, 0x3DCCCCCDu},
{"2.45134755833537796875e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"-245134755833537796876e-6", 0xC2EBDE5C4164D83Au, 0xD75EF2E2u},
{"-245134755833537.796874", 0xC2EBDE5C4164D839u, 0xD75EF2E2u},
{"2.4513475583353779687501e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"245134755833537796875000000000000000000001e-27", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"245134755833537.796875000000000000000000000000000000", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"2.4513475583353779e14", 0x42EBDE5C4164D839u, 0x575EF2E2u},
{"2451347558335378e-1", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"2451347558335377968e-4", 0x42EBDE5C4164D839u, 0x575EF2E2u},
{"245134755833537.7969", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"245134755833537.79687", 0x42EBDE5C4164D839u, 0x575EF2E2u},
{"2.4513475583353779688e14", 0x42EBDE5C4164D83Au, 0x575EF2E2u},
{"181510327827821147441864013671875e-23", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"-1815103278.27821147441864013671876", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"1.81510327827821147441864013671874e9", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"18151032782782114744186401367187501e-25", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"1815103278.27821147441864013671875000000000000000000001", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"-1.81510327827821147441864013671875000000000000000000000000000000e9", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"18151032782782114e-7", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"1815103278.2782115", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"1815103278.278211474", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"-1.815103278278211475e9", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"1.8151032782782114744e9", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"-18151032782782114745e-10", 0xC1DB0C11CB91CE38u, 0xCED8608Eu},
{"181510327827821147441e-11", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"1815103278.27821147442", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"1815103278.27821147441864013671", 0x41DB0C11CB91CE37u, 0x4ED8608Eu},
{"1.81510327827821147441864013672e9", 0x41DB0C11CB91CE38u, 0x4ED8608Eu},
{"3809325632181785344", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3.809325632181785345e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3809325632181785343e0", 0x43CA6EB8BD69FE29u, 0x5E5375C6u},
{"3809325632181785344.01", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3.809325632181785344000000000000000000001e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3809325632181785344000000000000000000000000000000e-30", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"3809325632181785300", 0x43CA6EB8BD69FE29u, 0x5E5375C6u},
{"3.8093256321817854e18", 0x43CA6EB8BD69FE2Au, 0x5E5375C6u},
{"4.046966549916366943359375e12", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916366943359376e-12", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916.366943359374", 0x428D7210076CE2EFu, 0x546B9080u},
{"4.04696654991636694335937501e12", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916366943359375000000000000000000001e-33", 0x428D7210076CE2F0u, 0x546B9080u},
{"-4046966549916.366943359375000000000000000000000000000000", 0xC28D7210076CE2F0u, 0xD46B9080u},
{"4.0469665499163669e12", 0x428D7210076CE2EFu, 0x546B9080u},
{"4046966549916367e-3", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916366943e-6", 0x428D7210076CE2EFu, 0x546B9080u},
{"4046966549916.366944", 0x428D7210076CE2F0u, 0x546B9080u},
{"4046966549916.3669433", 0x428D7210076CE2EFu, 0x546B9080u},
{"4.0469665499163669434e12", 0x428D7210076CE2F0u, 0x546B9080u},
{"-4.04696654991636694335e12", 0xC28D7210076CE2EFu, 0xD46B9080u},
{"404696654991636694336e-8", 0x428D7210076CE2F0u, 0x546B9080u},
{"28093802557000874e154", 0x63529C3B77330BDBu, 0x7F800000u},
{"0.28093802557000875E+171", 0x63529C3B77330BDCu, 0x7F800000u},
{"0.2809380255700087447E+171", 0x63529C3B77330BDBu, 0x7F800000u},
{"2.809380255700087448e170", 0x63529C3B77330BDCu, 0x7F800000u},
{"2.8093802557000874472e170", 0x63529C3B77330BDBu, 0x7F800000u},
{"28093802557000874473e151", 0x63529C3B77330BDCu, 0x7F800000u},
{"280938025570008744728e150", 0x63529C3B77330BDBu, 0x7F800000u},
{"-0.280938025570008744729E+171", 0xE3529C3B77330BDCu, 0xFF800000u},
{"-0.280938025570008744728403667979E+171", 0xE3529C3B77330BDBu, 0xFF800000u},
{"2.8093802557000874472840366798e170", 0x63529C3B77330BDCu, 0x7F800000u},
{"0.39523280297734525e-154", 0x1FE0F51BF17FD374u, 0x00000000u},
{"-3.9523280297734526e-155", 0x9FE0F51BF17FD375u, 0x80000000u},
{"3.952328029773452547e-155", 0x1FE0F51BF17FD374u, 0x00000000u},
{"3952328029773452548e-173", 0x1FE0F51BF17FD375u, 0x00000000u},
{"-39523280297734525478e-174", 0x9FE0F51BF17FD374u, 0x80000000u},
{"0.39523280297734525479e-154", 0x1FE0F51BF17FD375u, 0x00000000u},
{"-0.395232802977345254787e-154", 0x9FE0F51BF17FD374u, 0x80000000u},
{"-3.95232802977345254788e-155", 0x9FE0F51BF17FD375u, 0x80000000u},
{"-3.95232802977345254787245825501e-155", 0x9FE0F51BF17FD374u, 0x80000000u},
{"395232802977345254787245825502e-184", 0x1FE0F51BF17FD375u, 0x00000000u},
{"-1.0790205420931879e-276", 0x86A3209CA6233255u, 0x80000000u},
{"-1079020542093188e-291", 0x86A3209CA6233256u, 0x80000000u},
{"-1079020542093187947e-294", 0x86A3209CA6233255u, 0x80000000u},
{"0.1079020542093187948e-275", 0x06A3209CA6233256u, 0x00000000u},
{"0.1079020542093187947e-275", 0x06A3209CA6233255u, 0x00000000u},
{"1.0790205420931879471e-276", 0x06A3209CA6233256u, 0x00000000u},
{"1.07902054209318794701e-276", 0x06A3209CA6233255u, 0x00000000u},
{"-107902054209318794702e-296", 0x86A3209CA6233256u, 0x80000000u},
{"107902054209318794701153285302e-305", 0x06A3209CA6233255u, 0x00000000u},
{"-0.107902054209318794701153285303e-275", 0x86A3209CA6233256u, 0x80000000u},
{"58530471071351308e-228", 0x1413B446E6A16A3Bu, 0x00000000u},
{"-0.58530471071351309e-211", 0x9413B446E6A16A3Cu, 0x80000000u},
{"0.5853047107135130893e-211", 0x1413B446E6A16A3Bu, 0x00000000u},
{"-5.853047107135130894e-212", 0x9413B446E6A16A3Cu, 0x80000000u},
{"5.853047107135130893e-212", 0x1413B446E6A16A3Bu, 0x00000000u},
{"58530471071351308931e-231", 0x1413B446E6A16A3Cu, 0x00000000u},
{"-585304710713513089304e-232", 0x9413B446E6A16A3Bu, 0x80000000u},
{"0.585304710713513089305e-211", 0x1413B446E6A16A3Cu, 0x00000000u},
{"-0.585304710713513089304248824438e-211", 0x9413B446E6A16A3Bu, 0x80000000u},
{"5.85304710713513089304248824439e-212", 0x1413B446E6A16A3Cu, 0x00000000u},
{"0.19334214893983531e-78", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"1.9334214893983532e-79", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"1.933421489398353102e-79", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"1933421489398353103e-97", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"19334214893983531023e-98", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"0.19334214893983531024e-78", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"0.193342148939835310231e-78", 0x2F96ECBF1CFB10F6u, 0x00000000u},
{"1.93342148939835310232e-79", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"-1.93342148939835310231359014704e-79", 0xAF96ECBF1CFB10F6u, 0x80000000u},
{"193342148939835310231359014705e-108", 0x2F96ECBF1CFB10F7u, 0x00000000u},
{"2.9873358928024455e227", 0x6F2938807814E8A2u, 0x7F800000u},
{"29873358928024456e211", 0x6F2938807814E8A3u, 0x7F800000u},
{"298733589280244551e210", 0x6F2938807814E8A2u, 0x7F800000u},
{"0.2987335892802445511E+228", 0x6F2938807814E8A3u, 0x7F800000u},
{"0.29873358928024455109E+228", 0x6F2938807814E8A2u, 0x7F800000u},
{"2.987335892802445511e227", 0x6F2938807814E8A3u, 0x7F800000u},
{"-2.98733589280244551098e227", 0xEF2938807814E8A2u, 0xFF800000u},
{"-298733589280244551099e207", 0xEF2938807814E8A3u, 0xFF800000u},
{"298733589280244551098081559931e198", 0x6F2938807814E8A2u, 0x7F800000u},
{"0.298733589280244551098081559932E+228", 0x6F2938807814E8A3u, 0x7F800000u},
{"7.0064923216240853e-46", 0x3690000000000000u, 0x00000000u},
{"70064923216240854e-62", 0x3690000000000000u, 0x00000001u},
{"7006492321624085354e-64", 0x3690000000000000u, 0x00000000u},
{"-0.7006492321624085355e-45", 0xB690000000000000u, 0x80000001u},
{"0.70064923216240853546e-45", 0x3690000000000000u, 0x00000000u},
{"7.0064923216240853547e-46", 0x3690000000000000u, 0x00000001u},
{"7.00649232162408535461e-46", 0x3690000000000000u, 0x00000000u},
{"-700649232162408535462e-66", 0xB690000000000000u, 0x80000001u},
{"-700649232162408535461864791644e-75", 0xB690000000000000u, 0x80000000u},
{"0.700649232162408535461864791645e-45", 0x3690000000000000u, 0x00000001u},
{"21019476964872256e-61", 0x36A8000000000000u, 0x00000001u},
{"0.21019476964872257e-44", 0x36A8000000000000u, 0x00000002u},
{"-0.2101947696487225606e-44", 0xB6A8000000000000u, 0x80000001u},
{"-2.101947696487225607e-45", 0xB6A8000000000000u, 0x80000002u},
{"2.1019476964872256063e-45", 0x36A8000000000000u, 0x00000001u},
{"21019476964872256064e-64", 0x36A8000000000000u, 0x00000002u},
{"-210194769648722560638e-65", 0xB6A8000000000000u, 0x80000001u},
{"-0.210194769648722560639e-44", 0xB6A8000000000000u, 0x80000002u},
{"-0.210194769648722560638559437493e-44", 0xB6A8000000000000u, 0x80000001u},
{"2.10194769648722560638559437494e-45", 0x36A8000000000000u, 0x00000002u},
{"0.11754941406275178e-37", 0x380FFFFFA0000000u, 0x007FFFFEu},
{"1.1754941406275179e-38", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"-1.175494140627517859e-38", 0xB80FFFFFA0000000u, 0x807FFFFEu},
{"117549414062751786e-55", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"-11754941406275178592e-57", 0xB80FFFFFA0000000u, 0x807FFFFEu},
{"0.11754941406275178593e-37", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"0.117549414062751785924e-37", 0x380FFFFFA0000000u, 0x007FFFFEu},
{"1.17549414062751785925e-38", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"-1.17549414062751785924617589866e-38", 0xB80FFFFFA0000000u, 0x807FFFFEu},
{"117549414062751785924617589867e-67", 0x380FFFFFA0000000u, 0x007FFFFFu},
{"1.1754942807573642e-38", 0x380FFFFFDFFFFFFFu, 0x007FFFFFu},
{"11754942807573643e-54", 0x380FFFFFE0000000u, 0x00800000u},
{"1175494280757364291e-56", 0x380FFFFFE0000000u, 0x007FFFFFu},
{"0.1175494280757364292e-37", 0x380FFFFFE0000000u, 0x00800000u},
{"0.11754942807573642917e-37", 0x380FFFFFE0000000u, 0x007FFFFFu},
{"-1.1754942807573642918e-38", 0xB80FFFFFE0000000u, 0x80800000u},
{"-1.17549428075736429172e-38", 0xB80FFFFFE0000000u, 0x807FFFFFu},
{"-117549428075736429173e-58", 0xB80FFFFFE0000000u, 0x80800000u},
{"117549428075736429172788299103e-67", 0x380FFFFFE0000000u, 0x007FFFFFu},
{"0.117549428075736429172788299104e-37", 0x380FFFFFE0000000u, 0x00800000u},
{"11754944208872107e-54", 0x3810000010000000u, 0x00800000u},
{"-0.11754944208872108e-37", 0xB810000010000000u, 0x80800001u},
{"0.1175494420887210724e-37", 0x3810000010000000u, 0x00800000u},
{"1.175494420887210725e-38", 0x3810000010000000u, 0x00800001u},
{"1.1754944208872107242e-38", 0x3810000010000000u, 0x00800000u},
{"-11754944208872107243e-57", 0xB810000010000000u, 0x80800001u},
{"-11754944208872107242e-57", 0xB810000010000000u, 0x80800000u},
{"0.117549442088721072421e-37", 0x3810000010000000u, 0x00800001u},
{"0.11754944208872107242095900834e-37", 0x3810000010000000u, 0x00800000u},
{"-1.17549442088721072420959008341e-38", 0xB810000010000000u, 0x80800001u},
{"340282336497324057985868971510891282432", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"-3.40282336497324057985868971510891282433e38", 0xC7EFFFFFD0000000u, 0xFF7FFFFFu},
{"340282336497324057985868971510891282431e0", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324057985868971510891282432.01", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.40282336497324057985868971510891282432000000000000000000001e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"340282336497324057985868971510891282432000000000000000000000000000000e-30", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324050000000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"3.4028233649732406e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.402823364973240579e38", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324058e21", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"34028233649732405798e19", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324057990000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"340282336497324057985000000000000000000", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"3.40282336497324057986e38", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.4028233649732405798586897151e38", 0x47EFFFFFD0000000u, 0x7F7FFFFEu},
{"340282336497324057985868971511e9", 0x47EFFFFFD0000000u, 0x7F7FFFFFu},
{"3.40282356779733661637539395458142568448e38", 0x47EFFFFFF0000000u, 0x7F800000u},
{"-340282356779733661637539395458142568449e0", 0xC7EFFFFFF0000000u, 0xFF800000u},
{"340282356779733661637539395458142568447", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"3.4028235677973366163753939545814256844801e38", 0x47EFFFFFF0000000u, 0x7F800000u},
{"340282356779733661637539395458142568448000000000000000000001e-21", 0x47EFFFFFF0000000u, 0x7F800000u},
{"340282356779733661637539395458142568448.000000000000000000000000000000", 0x47EFFFFFF0000000u, 0x7F800000u},
{"3.4028235677973366e38", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"-34028235677973367e22", 0xC7EFFFFFF0000000u, 0xFF800000u},
{"3402823567797336616e20", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"340282356779733661700000000000000000000", 0x47EFFFFFF0000000u, 0x7F800000u},
{"340282356779733661630000000000000000000", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"3.4028235677973366164e38", 0x47EFFFFFF0000000u, 0x7F800000u},
{"3.40282356779733661637e38", 0x47EFFFFFF0000000u, 0x7F7FFFFFu},
{"340282356779733661638e18", 0x47EFFFFFF0000000u, 0x7F800000u},
{"-340282356779733661637539395458e9", 0xC7EFFFFFF0000000u, 0xFF7FFFFFu},
{"340282356779733661637539395459000000000", 0x47EFFFFFF0000000u, 0x7F800000u},
{"-1000000059604644775390625e-24", 0xBFF0000010000000u, 0xBF800000u},
{"-1.000000059604644775390626", 0xBFF0000010000000u, 0xBF800001u},
{"-1.000000059604644775390624e0", 0xBFF0000010000000u, 0xBF800000u},
{"100000005960464477539062501e-26", 0x3FF0000010000000u, 0x3F800001u},
{"1.000000059604644775390625000000000000000000001", 0x3FF0000010000000u, 0x3F800001u},
{"1.000000059604644775390625000000000000000000000000000000e0", 0x3FF0000010000000u, 0x3F800000u},
{"-10000000596046447e-16", 0xBFF0000010000000u, 0xBF800000u},
{"1.0000000596046448", 0x3FF0000010000000u, 0x3F800001u},
{"1.000000059604644775", 0x3FF0000010000000u, 0x3F800000u},
{"-1.000000059604644776e0", 0xBFF0000010000000u, 0xBF800001u},
{"-1.0000000596046447753e0", 0xBFF0000010000000u, 0xBF800000u},
{"10000000596046447754e-19", 0x3FF0000010000000u, 0x3F800001u},
{"100000005960464477539e-20", 0x3FF0000010000000u, 0x3F800000u},
{"-1.0000000596046447754", 0xBFF0000010000000u, 0xBF800001u},
{"0.9999999701976776123046875", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"9.999999701976776123046876e-1", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"9999999701976776123046874e-25", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu},
{"0.999999970197677612304687501", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"-9.999999701976776123046875000000000000000000001e-1", 0xBFEFFFFFF0000000u, 0xBF800000u},
{"-9999999701976776123046875000000000000000000000000000000e-55", 0xBFEFFFFFF0000000u, 0xBF800000u},
{"0.99999997019767761", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu},
{"-9.9999997019767762e-1", 0xBFEFFFFFF0000000u, 0xBF800000u},
{"-9.999999701976776123e-1", 0xBFEFFFFFF0000000u, 0xBF7FFFFFu},
{"9999999701976776124e-19", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"9999999701976776123e-19", 0x3FEFFFFFF0000000u, 0x3F7FFFFFu},
{"0.99999997019767761231", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"-0.999999970197677612304", 0xBFEFFFFFF0000000u, 0xBF7FFFFFu},
{"9.99999970197677612305e-1", 0x3FEFFFFFF0000000u, 0x3F800000u},
{"-1.6777217e7", 0xC170000010000000u, 0xCB800000u},
{"16777218e0", 0x4170000020000000u, 0x4B800001u},
{"16777216", 0x4170000000000000u, 0x4B800000u},
{"-1.677721701e7", 0xC17000001028F5C3u, 0xCB800001u},
{"-16777217000000000000000000001e-21", 0xC170000010000000u, 0xCB800001u},
{"16777217.000000000000000000000000000000", 0x4170000010000000u, 0x4B800000u},
{"167772155e-1", 0x416FFFFFF0000000u, 0x4B800000u},
{"16777215.6", 0x416FFFFFF3333333u, 0x4B800000u},
{"1.67772154e7", 0x416FFFFFECCCCCCDu, 0x4B7FFFFFu},
{"16777215501e-3", 0x416FFFFFF0083127u, 0x4B800000u},
{"-16777215.5000000000000000000001", 0xC16FFFFFF0000000u, 0xCB800000u},
{"-1.67772155000000000000000000000000000000e7", 0xC16FFFFFF0000000u, 0xCB800000u},
{"0.1000000052154064178466796875", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"1.000000052154064178466796876e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"-1000000052154064178466796874e-28", 0xBFB99999B0000000u, 0xBDCCCCCDu},
{"0.100000005215406417846679687501", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"1.000000052154064178466796875000000000000000000001e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"-1000000052154064178466796875000000000000000000000000000000e-58", 0xBFB99999B0000000u, 0xBDCCCCCEu},
{"-0.10000000521540641", 0xBFB99999AFFFFFFFu, 0xBDCCCCCDu},
{"-1.0000000521540642e-1", 0xBFB99999B0000000u, 0xBDCCCCCEu},
{"1.000000052154064178e-1", 0x3FB99999B0000000u, 0x3DCCCCCDu},
{"-1000000052154064179e-19", 0xBFB99999B0000000u, 0xBDCCCCCEu},
{"10000000521540641784e-20", 0x3FB99999B0000000u, 0x3DCCCCCDu},
{"0.10000000521540641785", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"0.100000005215406417846", 0x3FB99999B0000000u, 0x3DCCCCCDu},
{"1.00000005215406417847e-1", 0x3FB99999B0000000u, 0x3DCCCCCEu},
{"5.429001220703125e3", 0x40B5350050000000u, 0x45A9A802u},
{"-5429001220703126e-12", 0xC0B5350050000001u, 0xC5A9A803u},
{"-5429.001220703124", 0xC0B535004FFFFFFFu, 0xC5A9A802u},
{"5.42900122070312501e3", 0x40B5350050000000u, 0x45A9A803u},
{"5429001220703125000000000000000000001e-33", 0x40B5350050000000u, 0x45A9A803u},
{"-5429.001220703125000000000000000000000000000000", 0xC0B5350050000000u, 0xC5A9A802u},
{"503719056e0", 0x41BE062490000000u, 0x4DF03124u},
{"503719057", 0x41BE062491000000u, 0x4DF03125u},
{"5.03719055e8", 0x41BE06248F000000u, 0x4DF03124u},
{"50371905601e-2", 0x41BE062490028F5Cu, 0x4DF03125u},
{"503719056.000000000000000000001", 0x41BE062490000000u, 0x4DF03125u},
{"5.03719056000000000000000000000000000000e8", 0x41BE062490000000u, 0x4DF03124u},
{"-92331620", 0xC196037990000000u, 0xCCB01BCCu},
{"9.233163e7", 0x41960379B8000000u, 0x4CB01BCEu},
{"9233161e1", 0x4196037968000000u, 0x4CB01BCBu},
{"92331620.1", 0x4196037990666666u, 0x4CB01BCDu},
{"9.233162000000000000000000001e7", 0x4196037990000000u, 0x4CB01BCDu},
{"9233162000000000000000000000000000000e-29", 0x4196037990000000u, 0x4CB01BCCu},
{"3.002458625e6", 0x4146E82D50000000u, 0x4A37416Au},
{"3002458626e-3", 0x4146E82D5020C49Cu, 0x4A37416Bu},
{"3002458.624", 0x4146E82D4FDF3B64u, 0x4A37416Au},
{"-3.00245862501e6", 0xC146E82D500053E3u, 0xCA37416Bu},
{"3002458625000000000000000000001e-24", 0x4146E82D50000000u, 0x4A37416Bu},
{"3002458.625000000000000000000000000000000", 0x4146E82D50000000u, 0x4A37416Au},
{"-1095485584696182596504479582065262592e1", 0xC7A07BA830000000u, 0xFD03DD42u},
{"10954855846961825965044795820652625930", 0x47A07BA830000000u, 0x7D03DD42u},
{"1.095485584696182596504479582065262591e37", 0x47A07BA830000000u, 0x7D03DD41u},
{"109548558469618259650447958206526259201e-1", 0x47A07BA830000000u, 0x7D03DD42u},
{"-10954855846961825965044795820652625920.00000000000000000001", 0xC7A07BA830000000u, 0xFD03DD42u},
{"1.095485584696182596504479582065262592000000000000000000000000000000e37", 0x47A07BA830000000u, 0x7D03DD42u},
{"10954855846961825e21", 0x47A07BA830000000u, 0x7D03DD41u},
{"10954855846961826000000000000000000000", 0x47A07BA830000000u, 0x7D03DD42u},
{"10954855846961825960000000000000000000", 0x47A07BA830000000u, 0x7D03DD41u},
{"1.095485584696182597e37", 0x47A07BA830000000u, 0x7D03DD42u},
{"1.0954855846961825965e37", 0x47A07BA830000000u, 0x7D03DD41u},
{"-10954855846961825966e18", 0xC7A07BA830000000u, 0xFD03DD42u},
{"10954855846961825965e18", 0x47A07BA830000000u, 0x7D03DD41u},
{"10954855846961825965100000000000000000", 0x47A07BA830000000u, 0x7D03DD42u},
{"10954855846961825965044795820600000000", 0x47A07BA830000000u, 0x7D03DD41u},
{"-1.09548558469618259650447958207e37", 0xC7A07BA830000000u, 0xFD03DD42u},
{"1.6449216019182103706535606608388384863861375606575165875256061553955078126e-21", 0x3B9F125A50000000u, 0x1CF892D3u},
{"-16449216019182103706535606608388384863861375606575165875256061553955078124e-94", 0xBB9F125A50000000u, 0x9CF892D2u},
{"-0.0000000000000000000016449216019182103", 0xBB9F125A50000000u, 0x9CF892D2u},
{"-1.6449216019182104e-21", 0xBB9F125A50000000u, 0x9CF892D3u},
{"-1.64492160191821037e-21", 0xBB9F125A50000000u, 0x9CF892D2u},
{"-1644921601918210371e-39", 0xBB9F125A50000000u, 0x9CF892D3u},
{"-16449216019182103706e-40", 0xBB9F125A50000000u, 0x9CF892D2u},
{"0.0000000000000000000016449216019182103707", 0x3B9F125A50000000u, 0x1CF892D3u},
{"0.00000000000000000000164492160191821037065", 0x3B9F125A50000000u, 0x1CF892D2u},
{"1.64492160191821037066e-21", 0x3B9F125A50000000u, 0x1CF892D3u},
{"1.64492160191821037065356066083e-21", 0x3B9F125A50000000u, 0x1CF892D2u},
{"164492160191821037065356066084e-50", 0x3B9F125A50000000u, 0x1CF892D3u},
{"6.565061509609222412109375e-1", 0x3FE5021930000000u, 0x3F2810CAu},
{"6565061509609222412109376e-25", 0x3FE5021930000000u, 0x3F2810CAu},
{"0.6565061509609222412109374", 0x3FE5021930000000u, 0x3F2810C9u},
{"6.56506150960922241210937501e-1", 0x3FE5021930000000u, 0x3F2810CAu},
{"6565061509609222412109375000000000000000000001e-46", 0x3FE5021930000000u, 0x3F2810CAu},
{"-0.6565061509609222412109375000000000000000000000000000000", 0xBFE5021930000000u, 0xBF2810CAu},
{"6.5650615096092224e-1", 0x3FE5021930000000u, 0x3F2810C9u},
{"-65650615096092225e-17", 0xBFE5021930000000u, 0xBF2810CAu},
{"6565061509609222412e-19", 0x3FE5021930000000u, 0x3F2810C9u},
{"0.6565061509609222413", 0x3FE5021930000000u, 0x3F2810CAu},
{"0.65650615096092224121", 0x3FE5021930000000u, 0x3F2810C9u},
{"6.5650615096092224122e-1", 0x3FE5021930000000u, 0x3F2810CAu},
{"-6.5650615096092224121e-1", 0xBFE5021930000000u, 0xBF2810C9u},
{"656506150960922241211e-21", 0x3FE5021930000000u, 0x3F2810CAu},
{"18014627239033005156980393746124491372029297053813934326171875e-77", 0x3CA9F63970000000u, 0x254FB1CCu},
{"0.00000000000000018014627239033005156980393746124491372029297053813934326171876", 0x3CA9F63970000000u, 0x254FB1CCu},
{"-1.8014627239033005156980393746124491372029297053813934326171874e-16", 0xBCA9F63970000000u, 0xA54FB1CBu},
{"1801462723903300515698039374612449137202929705381393432617187501e-79", 0x3CA9F63970000000u, 0x254FB1CCu},
{"18014627239033005e-32", 0x3CA9F63970000000u, 0x254FB1CBu},
{"0.00000000000000018014627239033006", 0x3CA9F63970000000u, 0x254FB1CCu},
{"0.0000000000000001801462723903300515", 0x3CA9F63970000000u, 0x254FB1CBu},
{"-1.801462723903300516e-16", 0xBCA9F63970000000u, 0xA54FB1CCu},
{"-1.8014627239033005156e-16", 0xBCA9F63970000000u, 0xA54FB1CBu},
{"18014627239033005157e-35", 0x3CA9F63970000000u, 0x254FB1CCu},
{"180146272390330051569e-36", 0x3CA9F63970000000u, 0x254FB1CBu},
{"-0.00000000000000018014627239033005157", 0xBCA9F63970000000u, 0xA54FB1CCu},
{"0.000000000000000180146272390330051569803937461", 0x3CA9F63970000000u, 0x254FB1CBu},
{"1.80146272390330051569803937462e-16", 0x3CA9F63970000000u, 0x254FB1CCu},
{"0.05534819327294826507568359375", 0x3FAC569930000000u, 0x3D62B4CAu},
{"-5.534819327294826507568359376e-2", 0xBFAC569930000000u, 0xBD62B4CAu},
{"5534819327294826507568359374e-29", 0x3FAC569930000000u, 0x3D62B4C9u},
{"-0.0553481932729482650756835937501", 0xBFAC569930000000u, 0xBD62B4CAu},
{"5.534819327294826507568359375000000000000000000001e-2", 0x3FAC569930000000u, 0x3D62B4CAu},
{"5534819327294826507568359375000000000000000000000000000000e-59", 0x3FAC569930000000u, 0x3D62B4CAu},
{"0.055348193272948265", 0x3FAC569930000000u, 0x3D62B4C9u},
{"5.5348193272948266e-2", 0x3FAC569930000000u, 0x3D62B4CAu},
{"5.534819327294826507e-2", 0x3FAC569930000000u, 0x3D62B4C9u},
{"5534819327294826508e-20", 0x3FAC569930000000u, 0x3D62B4CAu},
{"55348193272948265075e-21", 0x3FAC569930000000u, 0x3D62B4C9u},
{"-0.055348193272948265076", 0xBFAC569930000000u, 0xBD62B4CAu},
{"0.0553481932729482650756", 0x3FAC569930000000u, 0x3D62B4C9u},
{"-5.53481932729482650757e-2", 0xBFAC569930000000u, 0xBD62B4CAu},
{"5.179692133247783258005389047985340416e36", 0x478F2C9450000000u, 0x7C7964A2u},
{"5179692133247783258005389047985340417e0", 0x478F2C9450000000u, 0x7C7964A3u},
{"5179692133247783258005389047985340415", 0x478F2C9450000000u, 0x7C7964A2u},
{"-5.17969213324778325800538904798534041601e36", 0xC78F2C9450000000u, 0xFC7964A3u},
{"-5179692133247783258005389047985340416000000000000000000001e-21", 0xC78F2C9450000000u, 0xFC7964A3u},
{"-5179692133247783258005389047985340416.000000000000000000000000000000", 0xC78F2C9450000000u, 0xFC7964A2u},
{"5.1796921332477832e36", 0x478F2C9450000000u, 0x7C7964A2u},
{"51796921332477833e20", 0x478F2C9450000000u, 0x7C7964A3u},
{"-5179692133247783258e18", 0xC78F2C9450000000u, 0xFC7964A2u},
{"5179692133247783259000000000000000000", 0x478F2C9450000000u, 0x7C7964A3u},
{"5179692133247783258000000000000000000", 0x478F2C9450000000u, 0x7C7964A2u},
{"5.1796921332477832581e36", 0x478F2C9450000000u, 0x7C7964A3u},
{"-5.179692133247783258e36", 0xC78F2C9450000000u, 0xFC7964A2u},
{"-517969213324778325801e16", 0xC78F2C9450000000u, 0xFC7964A3u},
{"517969213324778325800538904798e7", 0x478F2C9450000000u, 0x7C7964A2u},
{"5179692133247783258005389047990000000", 0x478F2C9450000000u, 0x7C7964A3u},
{
"0.22250738585072011360574097967091319759348195463516456480234261097248222220210769455165295239081350"
"8791414915891303962110687008643869459464552765720740782062174337998814106326732925355228688137214901"
"2981122451451889849057222307285255133155755015914397476397983411801999323962548289017107081850690630"
"6666559949382757725720157630626906633326475653000092458883164330377797918696120494973903778297049050"
"5108060994073026293712895895000358379996720725430436028407889577179615094551674824347103070260914462"
"1572289880258182545180325707018860872113128079512233426288368622321503775666622503982534335974568884"
"4239002654981983854879482922068947216898310996983658468140228542433306603398508864458040010349339704"
"2756718644338377048603786162277173854562306587467901408672332763671875e-307", 0x0010000000000000u, 0x00000000u
},
{
"2.22507385850720113605740979670913197593481954635164564802342610972482222202107694551652952390813508"
"7914149158913039621106870086438694594645527657207407820621743379988141063267329253552286881372149012"
"9811224514518898490572223072852551331557550159143974763979834118019993239625482890171070818506906306"
"6665599493827577257201576306269066333264756530000924588831643303777979186961204949739037782970490505"
"1080609940730262937128958950003583799967207254304360284078895771796150945516748243471030702609144621"
"5722898802581825451803257070188608721131280795122334262883686223215037756666225039825343359745688844"
"2390026549819838548794829220689472168983109969836584681402285424333066033985088644580400103493397042"
"756718644338377048603786162277173854562306587467901408672332763671875000000000000000000001e-308", 0x0010000000000000u, 0x00000000u
},
{
"0.11754942807573642917278829910357665133228589927589904276829631184250030649651730385585324256680905"
"8189392089843750000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"00000000000000000000000000000000000000000000000000000000000000000e-37", 0x380FFFFFE0000000u, 0x00800000u
},
{
"1175494280757364291727882991035766513322858992758990427682963118425003064965173038558532425668090581"
"8939208984375000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000"
"00000000000001e-751", 0x380FFFFFE0000000u, 0x00800000u
},
{"0", 0x0000000000000000u, 0x00000000u},
{"-0", 0x8000000000000000u, 0x80000000u},
{"0.0", 0x0000000000000000u, 0x00000000u},
{"-0.0", 0x8000000000000000u, 0x80000000u},
{"0e999999999999999999999", 0x0000000000000000u, 0x00000000u},
{"-0.000e-99999", 0x8000000000000000u, 0x80000000u},
{"1e-400", 0x0000000000000000u, 0x00000000u},
{"-1e-400", 0x8000000000000000u, 0x80000000u},
{"1e400", 0x7FF0000000000000u, 0x7F800000u},
{"-1e400", 0xFFF0000000000000u, 0xFF800000u},
{"1e-50", 0x358DEE7A4AD4B81Fu, 0x00000000u},
{"-1e-50", 0xB58DEE7A4AD4B81Fu, 0x80000000u},
{"1e39", 0x48078287F49C4A1Du, 0x7F800000u},
{"-1e39", 0xC8078287F49C4A1Du, 0xFF800000u},
{"1e99999999999999999999999999", 0x7FF0000000000000u, 0x7F800000u},
{"1e-99999999999999999999999999", 0x0000000000000000u, 0x00000000u},
{"1e0000000000000000000000000000000000000000308", 0x7FE1CCF385EBC8A0u, 0x7F800000u},
{"123456789012345678901234567890e-30", 0x3FBF9ADD3746F65Fu, 0x3DFCD6EAu},
{"18446744073709551615", 0x43F0000000000000u, 0x5F800000u},
{"18446744073709551616", 0x43F0000000000000u, 0x5F800000u},
{"-9223372036854775808", 0xC3E0000000000000u, 0xDF000000u},
{"-9223372036854775809", 0xC3E0000000000000u, 0xDF000000u},
}
};
return table;
}
} // namespace float_hard_cases
@@ -0,0 +1,346 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <string>
#include <vector>
namespace
{
struct ill_formed_case
{
const char* name;
std::string bytes;
};
// RFC 3629 ill-formed sequences used throughout this file, plus one
// well-formed sequence for contrast
const std::vector<ill_formed_case> ill_formed_cases =
{
{"overlong", "\xC0\xAE"},
{"lone_0xFF", "\xFF"},
{"truncated", "\xE2\x82"},
{"surrogate", "\xED\xA0\x80"},
};
const std::string valid_sequence = "\xC3\xA9"; // U+00E9, "é"
using eh = json::error_handler_t;
const std::vector<eh> all_handlers = {eh::strict, eh::replace, eh::ignore, eh::keep};
// what dump()+parse() produces for a sanitizing error_handler; this is the
// ground truth every binary writer/reader is checked against
std::string dump_and_parse(const std::string& raw, eh error_handler)
{
return json::parse(json(raw).dump(-1, ' ', false, error_handler)).get<std::string>();
}
} // namespace
TEST_CASE("UTF-8 error_handler for the binary readers and writers")
{
SECTION("writers: string value")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
const json jval = c.bytes;
CHECK_THROWS_AS(json::to_cbor(jval, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jval, false, false, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::strict), json::type_error&);
{
json jobj;
jobj["k"] = jval;
CHECK_THROWS_AS(json::to_bson(jobj, eh::strict), json::type_error&);
}
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(json::to_cbor(jval, h)).get<std::string>() == expected);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, h)).get<std::string>() == expected);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, h)).get<std::string>() == expected);
{
json jobj;
jobj["k"] = jval;
const auto bytes = json::to_bson(jobj, h);
CHECK(json::from_bson(bytes)["k"].get<std::string>() == expected);
}
}
// keep: the writer passes the ill-formed bytes through unchanged,
// exactly as every binary writer did before this parameter existed
CHECK(json::from_cbor(json::to_cbor(jval, eh::keep)).get<std::string>() == c.bytes);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, eh::keep)).get<std::string>() == c.bytes);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::keep)).get<std::string>() == c.bytes);
{
json jobj;
jobj["k"] = jval;
const auto bytes = json::to_bson(jobj, eh::keep);
CHECK(json::from_bson(bytes)["k"].get<std::string>() == c.bytes);
}
}
}
SECTION("writers: object key")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
json jobj;
jobj[c.bytes] = 1;
CHECK_THROWS_AS(json::to_cbor(jobj, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jobj, false, false, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::strict), json::type_error&);
CHECK_THROWS_AS(json::to_bson(jobj, eh::strict), json::type_error&);
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(json::to_cbor(jobj, h)).begin().key() == expected);
CHECK(json::from_ubjson(json::to_ubjson(jobj, false, false, h)).begin().key() == expected);
CHECK(json::from_bjdata(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, h)).begin().key() == expected);
CHECK(json::from_bson(json::to_bson(jobj, h)).begin().key() == expected);
}
// keep: object keys round-trip unchanged too
CHECK(json::from_cbor(json::to_cbor(jobj, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_ubjson(json::to_ubjson(jobj, false, false, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_bjdata(json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::keep)).begin().key() == c.bytes);
CHECK(json::from_bson(json::to_bson(jobj, eh::keep)).begin().key() == c.bytes);
}
}
SECTION("readers: string value")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
// bytes produced the lenient (keep) way, as any binary reader
// accepted them before this parameter existed
const auto cbor_bytes = json::to_cbor(json(c.bytes), eh::keep);
const auto msgpack_bytes = json::to_msgpack(json(c.bytes)); // to_msgpack has no error_handler; always pass-through
const auto ubjson_bytes = json::to_ubjson(json(c.bytes), false, false, eh::keep);
const auto bjdata_bytes = json::to_bjdata(json(c.bytes), false, false, json::bjdata_version_t::draft2, eh::keep);
const auto bson_bytes = [&c]
{
json jobj;
jobj["k"] = c.bytes;
return json::to_bson(jobj, eh::keep);
}();
// keep (the default): bytes are kept unchanged
CHECK(json::from_cbor(cbor_bytes).get<std::string>() == c.bytes);
CHECK(json::from_msgpack(msgpack_bytes).get<std::string>() == c.bytes);
CHECK(json::from_ubjson(ubjson_bytes).get<std::string>() == c.bytes);
CHECK(json::from_bjdata(bjdata_bytes).get<std::string>() == c.bytes);
CHECK(json::from_bson(bson_bytes)["k"].get<std::string>() == c.bytes);
// strict: parse_error.113, discarded (not thrown) when allow_exceptions is false
CHECK_THROWS_AS(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, eh::strict), json::parse_error&);
CHECK(json::from_cbor(cbor_bytes, true, false, json::cbor_tag_handler_t::error, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_msgpack(msgpack_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_msgpack(msgpack_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_ubjson(ubjson_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_ubjson(ubjson_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_bjdata(bjdata_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_bjdata(bjdata_bytes, true, false, eh::strict).is_discarded());
CHECK_THROWS_AS(json::from_bson(bson_bytes, true, true, eh::strict), json::parse_error&);
CHECK(json::from_bson(bson_bytes, true, false, eh::strict).is_discarded());
// replace / ignore: match what dump() would have sanitized the same bytes to
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, h).get<std::string>() == expected);
CHECK(json::from_msgpack(msgpack_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_ubjson(ubjson_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_bjdata(bjdata_bytes, true, true, h).get<std::string>() == expected);
CHECK(json::from_bson(bson_bytes, true, true, h)["k"].get<std::string>() == expected);
}
}
}
SECTION("readers: object key")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
json jobj;
jobj[c.bytes] = 1;
const auto cbor_bytes = json::to_cbor(jobj, eh::keep);
const auto msgpack_bytes = json::to_msgpack(jobj);
const auto ubjson_bytes = json::to_ubjson(jobj, false, false, eh::keep);
const auto bjdata_bytes = json::to_bjdata(jobj, false, false, json::bjdata_version_t::draft2, eh::keep);
const auto bson_bytes = json::to_bson(jobj, eh::keep);
CHECK(json::from_cbor(cbor_bytes).begin().key() == c.bytes);
CHECK(json::from_msgpack(msgpack_bytes).begin().key() == c.bytes);
CHECK(json::from_ubjson(ubjson_bytes).begin().key() == c.bytes);
CHECK(json::from_bjdata(bjdata_bytes).begin().key() == c.bytes);
CHECK(json::from_bson(bson_bytes).begin().key() == c.bytes);
CHECK_THROWS_AS(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_msgpack(msgpack_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_ubjson(ubjson_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_bjdata(bjdata_bytes, true, true, eh::strict), json::parse_error&);
CHECK_THROWS_AS(json::from_bson(bson_bytes, true, true, eh::strict), json::parse_error&);
for (const auto h :
{
eh::replace, eh::ignore
})
{
CAPTURE(static_cast<int>(h));
const std::string expected = dump_and_parse(c.bytes, h);
CHECK(json::from_cbor(cbor_bytes, true, true, json::cbor_tag_handler_t::error, h).begin().key() == expected);
CHECK(json::from_msgpack(msgpack_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_ubjson(ubjson_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_bjdata(bjdata_bytes, true, true, h).begin().key() == expected);
CHECK(json::from_bson(bson_bytes, true, true, h).begin().key() == expected);
}
}
}
SECTION("well-formed UTF-8 is unaffected by error_handler")
{
const json jval = valid_sequence;
json jobj;
jobj[valid_sequence] = valid_sequence;
for (const auto h : all_handlers)
{
CAPTURE(static_cast<int>(h));
CHECK(json::from_cbor(json::to_cbor(jval, h)).get<std::string>() == valid_sequence);
CHECK(json::from_ubjson(json::to_ubjson(jval, false, false, h)).get<std::string>() == valid_sequence);
CHECK(json::from_bjdata(json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, h)).get<std::string>() == valid_sequence);
CHECK(json::from_bson(json::to_bson(jobj, h)).begin().key() == valid_sequence);
CHECK(json::from_cbor(json::to_cbor(jval, eh::keep), true, true, json::cbor_tag_handler_t::error, h).get<std::string>() == valid_sequence);
CHECK(json::from_msgpack(json::to_msgpack(jval), true, true, h).get<std::string>() == valid_sequence);
}
}
SECTION("dump() with error_handler_t::keep writes raw bytes as is")
{
for (const auto& c : ill_formed_cases)
{
CAPTURE(c.name);
const json jval = c.bytes;
const std::string dumped = jval.dump(-1, ' ', false, eh::keep);
CHECK(dumped.find(c.bytes) != std::string::npos);
// even with ensure_ascii, the ill-formed bytes are written as is
const std::string dumped_ascii = jval.dump(-1, ' ', true, eh::keep);
CHECK(dumped_ascii.find(c.bytes) != std::string::npos);
}
// well-formed characters around an ill-formed sequence are still
// escaped as usual under ensure_ascii
const json mixed = valid_sequence + ill_formed_cases[1].bytes; // "é" + lone 0xFF
const std::string dumped_mixed = mixed.dump(-1, ' ', true, eh::keep);
CHECK(dumped_mixed.find("\\u00e9") != std::string::npos);
CHECK(dumped_mixed.find(ill_formed_cases[1].bytes) != std::string::npos);
// the byte that ends an ill-formed sequence is read again, so a quote,
// a backslash, or a control character after it is still escaped, and
// a well-formed code point after it is escaped under ensure_ascii
for (const bool ensure_ascii :
{
false, true
})
{
CAPTURE(ensure_ascii);
CHECK(json("\xC3\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\\"\"");
CHECK(json("\xC3\\").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\\\\"");
CHECK(json("\xC3\n").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xC3\\n\"");
CHECK(json("\xE2\x82\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xE2\x82\\\"\"");
CHECK(json("\xFF\"").dump(-1, ' ', ensure_ascii, eh::keep) == "\"\xFF\\\"\"");
CHECK(json("a\xE2\x82").dump(-1, ' ', ensure_ascii, eh::keep) == "\"a\xE2\x82\"");
}
CHECK(json("\xC3\xC3\xA9").dump(-1, ' ', false, eh::keep) == "\"\xC3\xC3\xA9\"");
CHECK(json("\xC3\xC3\xA9").dump(-1, ' ', true, eh::keep) == "\"\xC3\\u00e9\"");
}
SECTION("to_msgpack and to_bon8 are not affected by error_handler")
{
const json jval = ill_formed_cases[1].bytes; // lone 0xFF
// to_msgpack has no error_handler parameter; the bytes are always
// passed through, as MessagePack's spec allows
CHECK(json::from_msgpack(json::to_msgpack(jval)).get<std::string>() == ill_formed_cases[1].bytes);
// to_bon8 has no error_handler parameter either; UTF-8 is structural
// for BON8, so it always rejects ill-formed input
CHECK_THROWS_AS(json::to_bon8(jval), json::type_error&);
}
SECTION("allow_exceptions=false with error_handler_t::strict discards the value")
{
const auto bytes = json::to_cbor(json(ill_formed_cases[0].bytes), eh::keep);
const json result = json::from_cbor(bytes, true, false, json::cbor_tag_handler_t::error, eh::strict);
CHECK(result.is_discarded());
}
SECTION("default parameters are unchanged")
{
const json jval = ill_formed_cases[0].bytes;
// to_*: the default error_handler is strict, so ill-formed input still throws
CHECK_THROWS_AS(json::to_cbor(jval), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(jval), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(jval), json::type_error&);
{
json jobj;
jobj["k"] = jval;
CHECK_THROWS_AS(json::to_bson(jobj), json::type_error&);
}
// from_*: the default error_handler is keep, so ill-formed bytes are
// still accepted unchanged, exactly as in release 3.12.0
const auto cbor_bytes = json::to_cbor(jval, eh::keep);
CHECK(json::from_cbor(cbor_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto ubjson_bytes = json::to_ubjson(jval, false, false, eh::keep);
CHECK(json::from_ubjson(ubjson_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto bjdata_bytes = json::to_bjdata(jval, false, false, json::bjdata_version_t::draft2, eh::keep);
CHECK(json::from_bjdata(bjdata_bytes).get<std::string>() == ill_formed_cases[0].bytes);
const auto msgpack_bytes = json::to_msgpack(jval);
CHECK(json::from_msgpack(msgpack_bytes).get<std::string>() == ill_formed_cases[0].bytes);
json bson_obj;
bson_obj["k"] = jval;
const auto bson_bytes = json::to_bson(bson_obj, eh::keep);
CHECK(json::from_bson(bson_bytes)["k"].get<std::string>() == ill_formed_cases[0].bytes);
}
}
+35
View File
@@ -3907,6 +3907,41 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_bjdata(j) == v);
CHECK(json::from_bjdata(v) == j);
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_bjdata() is
// strict, so such a value cannot be written back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_bjdata(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(j), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_bjdata(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_bjdata(j_key), json::type_error&);
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_bjdata(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")
+45
View File
@@ -154,6 +154,51 @@ TEST_CASE("BSON")
#endif
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// a BSON document {"s": "\xC0\xAE"} (0xC0 0xAE is an overlong
// encoding of '.'); the BSON spec does not require a decoder to
// reject ill-formed UTF-8 in a string value, so the reader hands the
// bytes back unchanged
const std::vector<uint8_t> v =
{
0x0F, 0x00, 0x00, 0x00, // document length
0x02, 's', 0x00, // type 0x02 (string), key "s"
0x03, 0x00, 0x00, 0x00, // string length (including null)
0xc0, 0xae, 0x00, // string content and its null terminator
0x00 // document terminator
};
json j;
CHECK_NOTHROW(j = json::from_bson(v));
REQUIRE(j.is_object());
REQUIRE(j.contains("s"));
CHECK(j["s"].get_ref<const json::string_t&>() == std::string("\xc0\xae"));
// dump() still requires valid UTF-8 and throws for such a value
CHECK_THROWS_AS(j.dump(), json::type_error&);
// to_bson() is strict as well, so the value cannot be written back
CHECK_THROWS_AS(json::to_bson(j), json::type_error&);
// to_bson() rejects the same kind of ill-formed string value, before
// any bytes reach the output adapter (the BSON document length
// prefix must be known up front, so nothing is written incrementally)
std::vector<std::uint8_t> out{0x42}; // a sentinel byte the writer must not touch
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}, nlohmann::detail::output_adapter<std::uint8_t>(out)), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
CHECK(out == std::vector<std::uint8_t> {0x42});
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC3"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xED\xA0\x80"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC0\xAF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected as well; unlike
// the reader (which never validates element names), the writer
// checks both string values and object keys
CHECK_THROWS_WITH_AS(json::to_bson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON")
{
// out_of_range.412 is thrown from a single shared helper
+65 -18
View File
@@ -1801,19 +1801,40 @@ TEST_CASE("CBOR")
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
SECTION("ill-formed UTF-8 in string (see #5529, #5651)")
{
// RFC 8949 §3.1 leaves it up to the decoder whether to reject
// ill-formed UTF-8 in a text string; this library does not, and
// hands the original bytes back unchanged, matching the
// MessagePack reader and the behavior before #5185/#5531 (not in
// any release)
// a two-character text string (major type 3) whose bytes are not
// valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') must be
// rejected at decode time, matching every other kind of
// malformed binary input, rather than only failing later when
// the resulting value is dumped
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae}), true, false).is_discarded());
// valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') round-trips
// byte for byte as a string value
const std::vector<uint8_t> ill_formed_value = {0x62, 0xc0, 0xae};
json j_value;
CHECK_NOTHROW(j_value = json::from_cbor(ill_formed_value));
REQUIRE(j_value.is_string());
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// to_cbor() is strict as well, so the value cannot be written back
CHECK_THROWS_AS(json::to_cbor(j_value), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0xa1, 0x62, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_cbor(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_cbor(j_key), json::type_error&);
// a CBOR byte string (major type 2) with the very same bytes is
// NOT text and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x42, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
@@ -1822,17 +1843,46 @@ TEST_CASE("CBOR")
CHECK(json::from_cbor(json::to_cbor(j)) == j);
}
SECTION("invalid UTF-8 in indefinite-length string")
SECTION("to_cbor rejects ill-formed UTF-8 (see #5651)")
{
// to_cbor() must reject the same ill-formed strings from_cbor()
// rejects, so a value it accepts can always be read back
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_cbor(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// binary values are not text and are unaffected
CHECK_NOTHROW(json::to_cbor(json::binary(std::vector<std::uint8_t>({0xFF}))));
}
SECTION("ill-formed UTF-8 in indefinite-length string")
{
json _;
// every chunk must be valid UTF-8 on its own (RFC 8949, Section
// 3.2.3), so a code point split across two chunks is rejected
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff}), true, false).is_discarded());
// the chunks are concatenated as is, without checking that each
// chunk is valid UTF-8 on its own (RFC 8949, Section 3.2.3), so
// a code point split across two chunks yields a valid string
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff})));
CHECK(_ == "\xc3\xa9");
CHECK(_.dump() == "\"\xc3\xa9\"");
// an ill-formed later chunk is rejected after valid ones
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})), "[json.exception.parse_error.113] parse error at byte 7: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
// a truncated code point is kept as is
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0xff})));
CHECK(_ == "\xc3");
CHECK_THROWS_AS(_.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_cbor(_), json::type_error&);
// an ill-formed later chunk is kept after valid ones
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})));
CHECK(_ == "\xc3\xa9\xc0\xae");
CHECK_THROWS_AS(_.dump(), json::type_error&);
// valid multi-byte chunks are accepted
CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc3, 0xb6, 0xff})) == "\xc3\xa9\xc3\xb6");
@@ -1840,9 +1890,6 @@ TEST_CASE("CBOR")
SECTION("many chunks in indefinite-length string")
{
// only the newly read chunk is validated, not the whole string
// collected so far; validating the latter made this input take
// quadratic time (about ten seconds for 100000 chunks)
constexpr std::size_t chunks = 100000;
std::vector<uint8_t> v{0x7f};
for (std::size_t i = 0; i < chunks; ++i)
+108 -275
View File
@@ -13,18 +13,15 @@
using nlohmann::json;
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <cstdint> // uint32_t, uint64_t
#include <cstdio> // snprintf
#include <cstdlib> // strtod
#include <cstring> // memcpy
#include <map> // map
#include <sstream> // stringstream
#include <string> // string
#include <utility> // pair
#include <vector> // vector
#include "float_hard_cases.hpp"
namespace
{
// shortcut to scan a string literal
@@ -260,7 +257,7 @@ TEST_CASE("lexer number fast path")
"123456789012345678901234567890", // huge -> float
"0.30000000000000004", "2.2250738585072014e-308", "1e308",
// high-precision / wide-exponent values that exercise the
// Eisel-Lemire path beyond the Clinger subset
// std::from_chars (Eisel-Lemire) path beyond the Clinger subset
"1.7976931348623157e308", "1.2345678901234567e-250",
"9007199254740993", "5e-324", "1e-320"
};
@@ -282,18 +279,20 @@ TEST_CASE("lexer number fast path")
}
}
SECTION("significant digits around Clinger's fast path")
SECTION("significant-digit gate for the Clinger fast path")
{
// Clinger's fast path needs a significand of at most 2^53, which
// tokens with 17 or more significant digits exceed. The conversion
// splits the token at the positions the scanners recorded, so leading
// zeros must not count as digits - "0.1234567890123456" has 16
// significant digits, not 17 - and both scanners must agree.
// Clinger's fast path needs a significand below 2^53, so it cannot
// succeed once the mantissa has 17 or more significant digits (the
// significand would be at least 10^16). The lexer skips the attempt
// there. That is only allowed to save work: every value must still come
// out bit-exactly, and both scanners must agree. In particular the gate
// must not fire for tokens whose leading zeros merely look like extra
// digits - "0.1234567890123456" has 16 significant digits, not 17.
const std::vector<std::string> numbers =
{
"1234567890123456", // 16 significant digits
"12345678901234567", // 17
"123456789012345678", // 18
"12345678901234567", // 17 -> attempt skipped
"123456789012345678", // 18 -> attempt skipped
"0.1234567890123456", // 16: the leading "0" is not significant
"0.12345678901234567", // 17
"0.00000000000000001", // 1, in a long token
@@ -664,145 +663,46 @@ TEST_CASE("lexer string fast path")
}
}
namespace
TEST_CASE("parse_float_fast declines what it cannot convert exactly")
{
// the index of the decimal point (or npos) and of the end of the mantissa of a
// number token, which the lexer records while scanning it
std::pair<std::size_t, std::size_t> float_token_layout(const std::string& s)
{
std::size_t dot = std::string::npos;
std::size_t mantissa_end = s.size();
for (std::size_t i = 0; i < s.size(); ++i)
// The lexer only hands well-formed numbers to parse_float_fast, so the
// malformed ones below can only be passed to it directly. Declining is
// always safe: the caller then falls back to a slower, exact conversion.
const auto fast = [](const std::string & s, double & out)
{
if (s[i] == '.')
{
dot = i;
}
else if (s[i] == 'e' || s[i] == 'E')
{
mantissa_end = i;
break;
}
}
return {dot, mantissa_end};
}
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out);
};
double out = 0;
template<typename FloatType>
FloatType parse_native(const std::string& s)
{
const auto layout = float_token_layout(s);
return nlohmann::detail::parse_float_native<FloatType>(s.data(), s.data() + s.size(), layout.first, layout.second);
}
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// without true double precision, the fast path declines everything
CHECK_FALSE(fast("1.5", out));
#else
CHECK(fast("1.5", out));
CHECK(out == 1.5);
CHECK(fast("+2.5e1", out));
CHECK(out == 25.0);
CHECK(fast("-25E-1", out));
CHECK(out == -2.5);
CHECK(fast("1e", out));
CHECK(out == 1.0);
#endif
std::uint64_t bits_of(double d)
{
std::uint64_t b = 0;
std::memcpy(&b, &d, sizeof(b));
return b;
}
// not a number
CHECK_FALSE(fast("", out));
CHECK_FALSE(fast("-", out));
CHECK_FALSE(fast(".", out));
CHECK_FALSE(fast("1.2.3", out));
CHECK_FALSE(fast("1x", out));
CHECK_FALSE(fast("1e+", out));
CHECK_FALSE(fast("1e1x", out));
std::uint32_t bits_of(float f)
{
std::uint32_t b = 0;
std::memcpy(&b, &f, sizeof(b));
return b;
}
std::uint64_t native_bits64(const std::string& s)
{
return bits_of(parse_native<double>(s));
}
std::uint32_t native_bits32(const std::string& s)
{
return bits_of(parse_native<float>(s));
}
} // namespace
TEST_CASE("parse_float_native rounds correctly")
{
SECTION("double")
{
CHECK(native_bits64("1.5") == 0x3FF8000000000000u);
CHECK(native_bits64("0.1") == 0x3FB999999999999Au);
CHECK(native_bits64("-0.0") == 0x8000000000000000u);
CHECK(native_bits64("0e999999999999999999999") == 0u);
// 2^53 + 1 is exactly between two doubles: ties to even, unless more digits follow
CHECK(native_bits64("9007199254740993") == 0x4340000000000000u);
CHECK(native_bits64("9007199254740993.0000000000000000001") == 0x4340000000000001u);
CHECK(native_bits64("9007199254740992.9999999999999999999") == 0x4340000000000000u);
// 1 + 2^-53 exactly (a tie), and one unit in the 55th digit around it
CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203125") == 0x3FF0000000000000u);
CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203126") == 0x3FF0000000000001u);
CHECK(native_bits64("1.00000000000000011102230246251565404236316680908203124") == 0x3FF0000000000000u);
// subnormal and overflow boundaries
CHECK(native_bits64("2.4703282292062327e-324") == 0u);
CHECK(native_bits64("2.4703282292062328e-324") == 1u);
CHECK(native_bits64("2.2250738585072011e-308") == 0x000FFFFFFFFFFFFFu);
CHECK(native_bits64("2.2250738585072012e-308") == 0x0010000000000000u);
CHECK(native_bits64("1.7976931348623157e308") == 0x7FEFFFFFFFFFFFFFu);
CHECK(native_bits64("1.7976931348623159e308") == 0x7FF0000000000000u);
CHECK(native_bits64("-1e400") == 0xFFF0000000000000u);
CHECK(native_bits64("-1e-400") == 0x8000000000000000u);
// exponents and zeros far beyond the range cancel out
CHECK(native_bits64("0." + std::string(1000, '0') + "1e1001") == 0x3FF0000000000000u);
CHECK(native_bits64("1" + std::string(1000, '0') + "e-1000") == 0x3FF0000000000000u);
CHECK(native_bits64("1e-99999999999999999999999") == 0u);
CHECK(native_bits64("1E+99999999999999999999999") == 0x7FF0000000000000u);
// more digits than any midpoint has (769): only whether a nonzero digit follows matters
const std::string tie = "1.00000000000000011102230246251565404236316680908203125";
CHECK(native_bits64(tie + std::string(800, '0')) == 0x3FF0000000000000u);
CHECK(native_bits64(tie + std::string(800, '0') + "1") == 0x3FF0000000000001u);
}
SECTION("float")
{
CHECK(native_bits32("1.5") == 0x3FC00000u);
CHECK(native_bits32("0.1") == 0x3DCCCCCDu);
CHECK(native_bits32("-0.0") == 0x80000000u);
// 2^24 + 1 is exactly between two floats
CHECK(native_bits32("16777217") == 0x4B800000u);
CHECK(native_bits32("16777217.000000000000000000001") == 0x4B800001u);
CHECK(native_bits32("16777218.999999999999999999999") == 0x4B800001u);
CHECK(native_bits32("16777219") == 0x4B800002u);
// subnormal and overflow boundaries
CHECK(native_bits32("3.4028235677973366e38") == 0x7F7FFFFFu);
CHECK(native_bits32("3.4028235677973367e38") == 0x7F800000u);
CHECK(native_bits32("7.006492321624085e-46") == 0u);
CHECK(native_bits32("7.006492321624086e-46") == 1u);
CHECK(native_bits32("1.1754942e-38") == 0x007FFFFFu);
CHECK(native_bits32("-1.17549435e-38") == 0x80800000u);
CHECK(native_bits32("1e39") == 0x7F800000u);
CHECK(native_bits32("-1e-50") == 0x80000000u);
// not rounded through double: its double would round to another float
CHECK(native_bits32("1.00000005960464477539062500000000001") == 0x3F800001u);
CHECK(native_bits32("9007199254740993") == 0x5A000000u);
}
SECTION("the conversion shared with other parsers")
{
// convert_float() gives the lexer's results, for every type
const std::vector<std::string> tokens =
{
"0", "-0.0", "1.5", "0.1", "1e-400", "-2.5E+3", "123456789012345678901234567890",
"9007199254740993.0000000000000000001", "4.9406564584124654e-324"
};
using float_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
for (const auto& t : tokens)
{
CAPTURE(t);
const auto layout = float_token_layout(t);
const char* const first = t.data();
const char* const last = first + t.size();
const auto d = nlohmann::detail::convert_float<double>(first, last, layout.first, layout.second);
const auto f = nlohmann::detail::convert_float<float>(first, last, layout.first, layout.second);
const auto ld = nlohmann::detail::convert_float<long double>(first, last, layout.first, layout.second);
CHECK(bits_of(d) == bits_of(json::parse(t).get<double>()));
CHECK(bits_of(f) == bits_of(float_json::parse(t).get<float>()));
CHECK(ld == long_double_json::parse(t).get<long double>());
}
}
// numbers that are not represented exactly on the fast path
CHECK_FALSE(fast("12345678901234567890", out));
CHECK_FALSE(fast("1e10000", out));
CHECK_FALSE(fast("9007199254740993", out));
CHECK_FALSE(fast("1e23", out));
CHECK_FALSE(fast("1e-23", out));
}
namespace
@@ -906,6 +806,40 @@ std::size_t big_bit_length(const big_uint& a)
}
return n;
}
std::uint64_t bits_of(double d)
{
std::uint64_t b = 0;
std::memcpy(&b, &d, sizeof(b));
return b;
}
bool eisel_lemire(const std::string& s, double& out)
{
return nlohmann::detail::parse_float_eisel_lemire(s.data(), s.data() + s.size(), out);
}
// significant digits of a token, without trailing zeros
std::size_t significant_digits(const std::string& s)
{
std::string digits;
for (const char c : s)
{
if (c == 'e' || c == 'E')
{
break;
}
if (c >= '0' && c <= '9' && !(digits.empty() && c == '0'))
{
digits += c;
}
}
while (!digits.empty() && digits.back() == '0')
{
digits.pop_back();
}
return digits.size();
}
} // namespace
TEST_CASE("Eisel-Lemire float conversion")
@@ -1303,33 +1237,26 @@ TEST_CASE("Eisel-Lemire float conversion")
for (const auto& c : known)
{
CAPTURE(c.first)
CHECK(native_bits64(c.first) == c.second);
double out = 0;
if (eisel_lemire(c.first, out))
{
CHECK(bits_of(out) == c.second);
}
else
{
// only tokens with more than 19 significant digits are left to
// strtod: those whose value lies too close to a tie
CHECK(significant_digits(c.first) > 19);
}
}
}
SECTION("binary32")
{
using binary32 = nlohmann::detail::ieee_binary_format<24>;
CHECK(nlohmann::detail::eisel_lemire<binary32>(0, 1) == 0x3F800000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-1, 1) == 0x3DCCCCCDu);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-1, 15) == 0x3FC00000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(0, 16777217) == 0x4B800000u); // tie, to even
CHECK(nlohmann::detail::eisel_lemire<binary32>(0, 16777219) == 0x4B800002u); // tie, to even
CHECK(nlohmann::detail::eisel_lemire<binary32>(-45, 1) == 0x00000001u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-46, 7) == 0x00000000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-46, 8) == 0x00000001u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-65, 9999999999999999999u) == 0x00000000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(20, 3402823466385288598u) == 0x7F7FFFFFu);
CHECK(nlohmann::detail::eisel_lemire<binary32>(20, 3402823669209384635u) == 0x7F800000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(39, 1) == 0x7F800000u);
CHECK(nlohmann::detail::eisel_lemire<binary32>(-5, 0) == 0x00000000u);
}
SECTION("round trip")
{
// every double written by to_chars and read back, and its 17-digit
// form with trailing digits that make the token longer than 19 digits
// every double written by to_chars and read back, also with trailing
// digits that make the token longer than 19 digits
std::uint64_t state = 5295;
std::size_t declined = 0;
for (int i = 0; i < 200000; ++i)
{
state ^= state << 13u;
@@ -1351,51 +1278,30 @@ TEST_CASE("Eisel-Lemire float conversion")
const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), d);
const std::string token(buffer.data(), static_cast<std::size_t>(end - buffer.data()));
CAPTURE(token)
CHECK(native_bits64(token) == b);
double out = 0;
REQUIRE(eisel_lemire(token, out));
CHECK(bits_of(out) == b);
// insert digits before the exponent of the 17-digit form: that
// form lies strictly inside the rounding interval of the double
// (the shortest one may lie on its boundary), and the digits move
// it by far less than the distance to the boundary, so the value
// must not change
std::array<char, 64> digits17{};
static_cast<void>(std::snprintf(digits17.data(), digits17.size(), "%.17g", d)); // NOLINT(cppcoreguidelines-pro-type-vararg,hicpp-vararg)
std::string longer = digits17.data();
// insert digits before the exponent: the value moves by far less
// than the distance to the rounding boundary, so it must not change
std::string longer = token;
const std::size_t e = longer.find('e');
const std::size_t dot = longer.find('.');
const std::string extra = dot == std::string::npos ? ".000000000000000000001" : "000000000000000000001";
longer.insert(e == std::string::npos ? longer.size() : e, extra);
CAPTURE(longer)
CHECK(native_bits64(longer) == b);
}
}
SECTION("round trip, binary32")
{
std::uint32_t state = 5295;
for (int i = 0; i < 100000; ++i)
{
state ^= state << 13u;
state ^= state >> 17u;
state ^= state << 5u;
std::uint32_t b = state;
if ((b & 0x7F800000u) == 0x7F800000u)
if (eisel_lemire(longer, out))
{
continue; // infinity or NaN
CHECK(bits_of(out) == b);
}
if (i % 4 == 0)
else
{
b &= 0x807FFFFFu; // subnormals
// w and w + 1 round differently: only when the value is very
// close to a rounding boundary
++declined;
}
float f = 0;
std::memcpy(&f, &b, sizeof(f));
std::array<char, 64> buffer{};
const char* end = nlohmann::detail::to_chars(buffer.data(), buffer.data() + buffer.size(), f);
const std::string token(buffer.data(), static_cast<std::size_t>(end - buffer.data()));
CAPTURE(token);
CHECK(native_bits32(token) == b);
}
CHECK(declined < 1000); // 107 of the 200,000
}
SECTION("used by the lexer")
@@ -1409,76 +1315,3 @@ TEST_CASE("Eisel-Lemire float conversion")
"[json.exception.out_of_range.406] number overflow parsing '1.7976931348623159e308'", json::out_of_range&);
}
}
namespace
{
using float_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
// the bits of the float that parse() gives for a token, via both scanners;
// the value must be the same for both
template<typename Json, typename Bits>
void check_parse(const std::string& token, Bits expected, Bits infinity)
{
std::stringstream stream(token);
if ((expected & ~(Bits{1} << (8 * sizeof(Bits) - 1))) == infinity)
{
Json _;
CHECK_THROWS_WITH_AS(_ = Json::parse(token), ("[json.exception.out_of_range.406] number overflow parsing '" + token + "'").c_str(), typename Json::out_of_range&);
CHECK_THROWS_WITH_AS(_ = Json::parse(stream), ("[json.exception.out_of_range.406] number overflow parsing '" + token + "'").c_str(), typename Json::out_of_range&);
return;
}
const Json contiguous = Json::parse(token);
const Json streamed = Json::parse(stream);
if (contiguous.is_number_float()) // not an integer that fits
{
CHECK(bits_of(contiguous.template get<typename Json::number_float_t>()) == expected);
CHECK(bits_of(streamed.template get<typename Json::number_float_t>()) == expected);
}
else
{
CHECK(streamed.is_number_integer());
}
}
} // namespace
TEST_CASE("float conversion of hard cases")
{
// see float_hard_cases.hpp
for (const auto& c : float_hard_cases::cases())
{
const std::string token = c.token;
CAPTURE(token);
CHECK(native_bits64(token) == c.bits64);
CHECK(native_bits32(token) == c.bits32);
check_parse<json>(token, c.bits64, std::uint64_t{0x7FF0000000000000u});
check_parse<float_json>(token, c.bits32, std::uint32_t{0x7F800000u});
}
}
TEST_CASE("float overflow and underflow in the parser")
{
SECTION("double")
{
check_parse<json>("1.7976931348623157e308", std::uint64_t{0x7FEFFFFFFFFFFFFFu}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("1.7976931348623159e308", std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("-1e309", std::uint64_t{0xFFF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("1" + std::string(400, '0'), std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("1e99999999999999999999", std::uint64_t{0x7FF0000000000000u}, std::uint64_t{0x7FF0000000000000u});
// an underflow gives a zero with the sign of the token
check_parse<json>("1e-400", std::uint64_t{0}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("-1e-400", std::uint64_t{0x8000000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("-2.4703282292062327e-324", std::uint64_t{0x8000000000000000u}, std::uint64_t{0x7FF0000000000000u});
check_parse<json>("0." + std::string(400, '0') + "1", std::uint64_t{0}, std::uint64_t{0x7FF0000000000000u});
}
SECTION("float")
{
check_parse<float_json>("3.4028234e38", std::uint32_t{0x7F7FFFFFu}, std::uint32_t{0x7F800000u});
check_parse<float_json>("3.4028236e38", std::uint32_t{0x7F800000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-1e39", std::uint32_t{0xFF800000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("1e-46", std::uint32_t{0}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-1e-46", std::uint32_t{0x80000000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-7.006492321624085e-46", std::uint32_t{0x80000000u}, std::uint32_t{0x7F800000u});
check_parse<float_json>("-7.006492321624086e-46", std::uint32_t{0x80000001u}, std::uint32_t{0x7F800000u});
}
}
+8 -25
View File
@@ -260,11 +260,10 @@ struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
{
// float and double are converted without the locale. A long double that
// is not binary64 can take the strtold fallback, which honors the locale
// that is current at conversion time. The numbers are chosen so that it
// does: too many significant digits for Clinger's fast path, an underflow
// that std::from_chars rejects, and a plain value.
// The numbers are chosen so that the conversion also takes the strtod
// fallback, which honors the locale that is current at conversion time:
// too many significant digits for Clinger's fast path, an underflow that
// std::from_chars rejects, and a plain value.
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
std::string text = "[";
for (const auto& n : numbers)
@@ -328,8 +327,7 @@ TEST_CASE("locale changes between lexer construction and number conversion (#519
}
}
// a long double goes through std::strtold unless it is binary64 or
// std::from_chars supports it
// a long double goes through std::strtold unless std::from_chars supports it
{
bool switched = false;
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept
@@ -355,15 +353,8 @@ TEST_CASE("locale with a multi-byte decimal point")
{
// Some locales use a decimal point that is not a single character, e.g.
// U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be
// substituted in place for '.', so the strtold fallback (only for long
// double formats other than binary64) converts a copy of the token with
// the whole decimal point instead (#5660). The values must be those of the
// "C" locale.
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
const char* const long_double_numbers = "[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]";
REQUIRE(std::setlocale(LC_NUMERIC, "C") != nullptr);
const long_double_json expected_long_double = long_double_json::parse(long_double_numbers);
// substituted in place for '.', so the strtod fallback stops early. The
// conversion must still terminate rather than retry forever.
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
bool tested = false;
for (const char* name : names)
@@ -381,20 +372,12 @@ TEST_CASE("locale with a multi-byte decimal point")
tested = true;
// too many significant digits for Clinger's fast path, and an underflow
// that std::from_chars rejects: double does not depend on the locale
// that std::from_chars rejects: both reach the strtod fallback
json j;
CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
CHECK(j.is_array());
CHECK(j[0] == 3.14159265358979323846);
CHECK(j[1] == 0.0);
CHECK(j[2] == -0.000123456789012345678);
CHECK(json::accept("3.14159265358979323846"));
// a long double that reaches the strtold fallback is not truncated
long_double_json ld;
CHECK_NOTHROW(ld = long_double_json::parse(long_double_numbers));
CHECK(ld == expected_long_double);
// a value the locale-independent paths convert is not affected
CHECK(json::parse("12.5") == 12.5);
}
+28 -8
View File
@@ -1540,19 +1540,39 @@ TEST_CASE("MessagePack")
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
SECTION("ill-formed UTF-8 in string (see #5529, #5651)")
{
// the MessagePack specification explicitly allows a str object to
// contain a byte sequence that is not valid UTF-8 and expects a
// deserializer to hand the original bytes back unchanged; this
// library follows that, unlike CBOR/UBJSON/BJData/BSON, whose
// specifications require text strings to be valid UTF-8
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
// (0xC0 0xAE is an overlong encoding of '.') must be rejected at
// decode time, matching every other kind of malformed binary
// input, rather than only failing later when the resulting
// value is dumped
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing MessagePack string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae}), true, false).is_discarded());
// (0xC0 0xAE is an overlong encoding of '.') round-trips byte for
// byte as a string value
const std::vector<uint8_t> ill_formed_value = {0xa2, 0xc0, 0xae};
json j_value;
CHECK_NOTHROW(j_value = json::from_msgpack(ill_formed_value));
REQUIRE(j_value.is_string());
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK(json::from_msgpack(json::to_msgpack(j_value)) == j_value);
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0x81, 0xa2, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_msgpack(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_msgpack(json::to_msgpack(j_key)) == j_key);
// a MessagePack bin8 blob with the very same bytes is NOT text
// and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_msgpack(std::vector<uint8_t>({0xc4, 0x02, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
+35
View File
@@ -2505,6 +2505,41 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_ubjson(j) == v);
CHECK(json::from_ubjson(v) == j);
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_ubjson() is
// strict, so such a value cannot be written back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_ubjson(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(j), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_ubjson(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_ubjson(j_key), json::type_error&);
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_ubjson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")