Compare commits

..
Author SHA1 Message Date
Niels Lohmann 9505be15fd Merge branch 'develop' into claude/binary-reader-narrow-numbers
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:22:22 +02:00
Niels Lohmann 633de8e44b Fix CI: clang-tidy and GCC -Wnoexcept in the locale test (#5613)
#5597 was merged before all of its CI jobs had run, and two of them fail
on develop now, and so on every pull request:

- ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C")
  calls whose result was discarded. Check the result, like the other
  resets in the file.
- ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser
  callbacks, which cannot throw but were not declared noexcept.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:20:43 +02:00
Niels Lohmann 44325873ea Check integer-to-float fallbacks for overflow in the binary readers
emit_signed, emit_unsigned, and the CBOR negative integer fallback now
pass their number_float_t fallback through emit_float, so a value that
overflows number_float_t is rejected with out_of_range.406 like a
floating-point value, instead of silently becoming infinity. This only
matters for a number_float_t that cannot represent 2^64, such as a
half-precision type. The CBOR value -1 - n is computed as long double so
that emit_float sees a finite value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 20:19:28 +02:00
Niels Lohmann 0c630d4c30 Fix MSVC and clang 3.5 in the narrow number type test
MSVC types 3000000000 and 5000000000 as unsigned long, so
json(-3000000000) triggered C4146 (unary minus on an unsigned type),
which /WX turns into an error. Use LL literals, as elsewhere in the
tests.

clang 3.5 cannot convert the lambdas in the braced initializer of the
format table to function pointers. Use named functions instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 18:39:45 +02:00
Niels Lohmann fc03b9912e Look up the locale decimal point at conversion time, not lexer construction (#5597)
* Look up the locale decimal point at conversion time, not lexer construction

The lexer read localeconv()->decimal_point once in its constructor and wrote
that character into token_buffer in place of '.'. The strtod fallback then
used the locale current at conversion time, so an LC_NUMERIC change in
between (parser callback, SAX handler, another thread) truncated the value
in release builds and fired the endptr assertion in debug builds.

token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback
depends on the locale: it looks up the decimal point right before the call,
restores '.' afterwards, and repeats the conversion if the locale changed in
between. As a side effect, std::from_chars and Clinger's fast path now also
apply under locales whose decimal point is not '.'.

Fixes #5198

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the strtod retry loop when the decimal point is unchanged

convert_float_locale_aware() repeated the conversion until strtod
consumed the whole token, assuming an early stop can only mean a locale
change. Under a locale whose decimal point is not a single character
(e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8,
all available on macOS), the in-place substitution can never succeed,
so parsing any float that reaches the strtod fallback (for example
3.14159265358979323846 at C++11) hung forever. Before this branch, the
same input was truncated.

Retry only if the decimal point changed since the previous attempt;
otherwise keep the value strtod parsed so far, as before. Add a test
that parses such numbers under a multi-byte decimal point locale; it
hangs without this change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix -Weffc++ errors in the #5198 locale test

GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc)
rejected LocaleSwitchingSax: it has a pointer data member but does not
declare its copy operations, and its vectors are not initialized in the
member initializer list. Store the locale name as a std::string and give
the vectors brace initializers, like SaxEventLogger in
unit-deserialization.cpp.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:56:11 +02:00
Niels Lohmann 9e1a09eec0 Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys

CBOR and MessagePack allow map keys of any type, but JSON object keys
are always strings, so such maps are rejected. The error so far was the
one for a malformed string (e.g. "expected length specification
(0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not
tell the user what went wrong. Report the type of the key instead:

  syntax error while parsing MessagePack object key: only string keys
  are supported, but found nil; last byte: 0xC0

The exception id (parse_error.113) and type are unchanged. Malformed
string keys and a missing key keep their previous messages. Document
the restriction on the CBOR and MessagePack pages.

Refs #2766, #3381

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Point the MessagePack key note to the spec's profile section

The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 17:51:07 +02:00
Niels Lohmann 04ed4f593c Handle numbers that do not fit narrow number types in the binary readers
With custom number types narrower than the values in a binary document,
for example basic_json<..., std::int32_t, std::uint32_t, float>, every
binary reader (CBOR, MessagePack, UBJSON, BJData, BSON, BON8) passed the
decoded number to the SAX interface with an implicit conversion: the
integer 5000000000 silently became 705032704, and a finite double such as
1e300 became infinity. The lexer handles the same values in JSON text: an
integer that fits neither integer type is stored as number_float_t, and a
finite number that overflows number_float_t is rejected with
out_of_range.406.

Pass every number read from binary input through three helpers that
apply the lexer's rules:
- emit_signed(): number_integer_t, else number_unsigned_t for a
  non-negative value, else number_float_t
- emit_unsigned(): number_unsigned_t, else number_float_t
- emit_float(): out_of_range.406 if a finite value overflows
  number_float_t; infinity and NaN are passed on

For consistency, a CBOR negative integer below the range of
number_integer_t is now stored as number_float_t, like a too small
integer in JSON text, instead of being rejected with parse_error.112.
With the default number types, this is the only change in behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 23:23:16 +02:00
20 changed files with 1642 additions and 1203 deletions
+2 -2
View File
@@ -80,8 +80,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were - Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were
used in the given input or if the input is not valid CBOR used in the given input or if the input is not valid CBOR
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key, - Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
but not found types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity ## Complexity
@@ -73,8 +73,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from - Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from
MessagePack were used in the given input or if the input is not valid MessagePack MessagePack were used in the given input or if the input is not valid MessagePack
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key, - Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
but not found types are not supported, as JSON object keys are always strings) or a string is malformed
## Complexity ## Complexity
@@ -55,6 +55,10 @@ This implementation does exactly follow this approach, as it uses double precisi
smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally
and be serialized to `null`. and be serialized to `null`.
During deserialization (from JSON text or any of the binary formats), a finite number that does not fit into
`number_float_t` is rejected with [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406), for
example a double-precision number in a binary format when `number_float_t` is `#!cpp float`.
#### Storage #### Storage
Floating-point number values are stored directly inside a `basic_json` type. Floating-point number values are stored directly inside a `basic_json` type.
@@ -47,8 +47,9 @@ With the default values for `NumberIntegerType` (`std::int64_t`), the default va
When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and
the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of
range will yield over/underflow when used in a constructor. During deserialization, too large or small integer numbers range will yield over/underflow when used in a constructor. During deserialization (from JSON text or any of the binary
will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md). formats), too large or small integer numbers will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md)
or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states: [RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are > Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
@@ -48,8 +48,9 @@ With the default values for `NumberUnsignedType` (`std::uint64_t`), the default
When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and
the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow
when used in a constructor. During deserialization, too large or small integer numbers will automatically be stored when used in a constructor. During deserialization (from JSON text or any of the binary formats), too large or small
as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md). integer numbers will automatically be stored as [`number_integer_t`](number_integer_t.md) or
[`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states: [RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are > Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
@@ -168,13 +168,26 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Negative integer overflow" !!! warning "Negative integer overflow"
CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the
result to fit into `number_integer_t` (`std::int64_t` by default), parsing fails with a result to fit into `number_integer_t` (`std::int64_t` by default), the result is stored as `number_float_t`, like
[`parse_error.112`](../../home/exceptions.md#jsonexceptionparse_error112) exception rather than overflowing a too small integer in JSON text. For example, `-18446744073709551616` (`0x3B` followed by eight `0xFF` bytes) is
silently. stored as `-1.8446744073709552e+19`.
!!! warning "Object keys" !!! warning "Object keys"
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps with keys other than UTF-8 strings are rejected. CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps
with keys other than text strings (major type 3) are rejected with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set
to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. This is a deliberate restriction of the library's JSON value model, not an oversight: formats built on CBOR
maps with integer keys, such as COSE ([RFC 9052](https://www.rfc-editor.org/rfc/rfc9052.html)) or CWT
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "UTF-8 validation of text strings" !!! warning "UTF-8 validation of text strings"
@@ -138,6 +138,21 @@ The library maps MessagePack types to JSON value types as follows:
Any MessagePack output created by `to_msgpack` can be successfully parsed by `from_msgpack`. Any MessagePack output created by `to_msgpack` can be successfully parsed by `from_msgpack`.
!!! warning "Object keys"
MessagePack allows map keys of any type, whereas JSON only allows strings as keys in object values. Like the
JSON-compatible [profile](https://github.com/msgpack/msgpack/blob/master/spec.md#profile) sketched in the
MessagePack specification, this library restricts map keys to `str` values. Maps with keys of any other type are
rejected with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values" !!! warning "UTF-8 validation of string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8. The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
+16 -7
View File
@@ -331,9 +331,6 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1 [json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1
``` ```
``` ```
[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow
```
```
[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5) [json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)
``` ```
@@ -343,13 +340,20 @@ A string could not be read from a [binary format](../features/binary_formats/ind
string was read where one was required (for instance as a map key), the string's length specification is invalid, or string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8. the string's bytes are not valid UTF-8.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
[CBOR](../features/binary_formats/cbor.md) and [MessagePack](../features/binary_formats/messagepack.md).
!!! failure "Example messages" !!! failure "Example messages"
``` ```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF [json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
``` ```
``` ```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF [json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C
``` ```
``` ```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82 [json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82
@@ -847,13 +851,18 @@ The JSON Patch operations 'remove' and 'add' cannot be applied to the root eleme
### json.exception.out_of_range.406 ### json.exception.out_of_range.406
A parsed number could not be stored as without changing it to NaN or INF. A parsed number could not be stored without changing it to NaN or INF. For the binary formats, this happens when a
finite floating-point number does not fit into [`number_float_t`](../api/basic_json/number_float_t.md), for example a
double-precision number when `number_float_t` is `#!cpp float`.
!!! failure "Example message" !!! failure "Example messages"
``` ```
number overflow parsing '10E1000' number overflow parsing '10E1000'
``` ```
```
[json.exception.out_of_range.406] syntax error while parsing CBOR value: number overflow
```
### json.exception.out_of_range.407 ### json.exception.out_of_range.407
+301 -47
View File
@@ -559,7 +559,7 @@ class binary_reader
case 0x01: // double case 0x01: // double
{ {
double number{}; double number{};
return get_number<double, true>(input_format_t::bson, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number<double, true>(input_format_t::bson, number) && emit_float(input_format_t::bson, number);
} }
case 0x02: // string case 0x02: // string
@@ -600,19 +600,19 @@ class binary_reader
case 0x10: // int32 case 0x10: // int32
{ {
std::int32_t value{}; std::int32_t value{};
return get_number<std::int32_t, true>(input_format_t::bson, value) && sax->number_integer(value); return get_number<std::int32_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value);
} }
case 0x12: // int64 case 0x12: // int64
{ {
std::int64_t value{}; std::int64_t value{};
return get_number<std::int64_t, true>(input_format_t::bson, value) && sax->number_integer(value); return get_number<std::int64_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value);
} }
case 0x11: // uint64 case 0x11: // uint64
{ {
std::uint64_t value{}; std::uint64_t value{};
return get_number<std::uint64_t, true>(input_format_t::bson, value) && sax->number_unsigned(value); return get_number<std::uint64_t, true>(input_format_t::bson, value) && emit_unsigned(input_format_t::bson, value);
} }
default: // anything else is not supported (yet) default: // anything else is not supported (yet)
@@ -638,14 +638,19 @@ class binary_reader
{ {
return false; return false;
} }
const auto max_val = static_cast<NumberType>((std::numeric_limits<number_integer_t>::max)());
if (number > max_val) // the value is -1 - number, which fits into number_integer_t
// whenever number does
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{ {
return sax->parse_error(chars_read, get_token_string(), return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
parse_error::create(112, chars_read,
exception_message(input_format_t::cbor, "negative integer overflow", "value"), nullptr));
} }
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
// like the lexer does for JSON text, store a value too small for
// number_integer_t as number_float_t; compute it as long double so
// that emit_float sees a finite value and can detect an overflow of
// number_float_t
return emit_float(input_format_t::cbor, static_cast<long double>(-1) - static_cast<long double>(number));
} }
/*! /*!
@@ -702,25 +707,25 @@ class binary_reader
case 0x18: // Unsigned integer (one-byte uint8_t follows) case 0x18: // Unsigned integer (one-byte uint8_t follows)
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number); return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
} }
case 0x19: // Unsigned integer (two-byte uint16_t follows) case 0x19: // Unsigned integer (two-byte uint16_t follows)
{ {
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number); return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
} }
case 0x1A: // Unsigned integer (four-byte uint32_t follows) case 0x1A: // Unsigned integer (four-byte uint32_t follows)
{ {
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number); return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
} }
case 0x1B: // Unsigned integer (eight-byte uint64_t follows) case 0x1B: // Unsigned integer (eight-byte uint64_t follows)
{ {
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number); return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
} }
// Negative integer -1-0x00..-1-0x17 (-1..-24) // Negative integer -1-0x00..-1-0x17 (-1..-24)
@@ -1165,13 +1170,13 @@ class binary_reader
case 0xFA: // Single-Precision Float (four-byte IEEE 754) case 0xFA: // Single-Precision Float (four-byte IEEE 754)
{ {
float number{}; float number{};
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
} }
case 0xFB: // Double-Precision Float (eight-byte IEEE 754) case 0xFB: // Double-Precision Float (eight-byte IEEE 754)
{ {
double number{}; double number{};
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
} }
default: // anything else (0xFF is handled inside the other types) default: // anything else (0xFF is handled inside the other types)
@@ -1324,6 +1329,80 @@ class binary_reader
} }
} }
/*!
@brief reads a CBOR object key
RFC 8949 allows any data item as a map key, but only strings have a
counterpart in JSON. A key of any other type is rejected with a message
naming that type, rather than the one @ref get_cbor_string gives for a
malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_cbor_object_key(string_t& result)
{
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
}
const char* found = nullptr;
switch (static_cast<unsigned int>(current) >> 5u)
{
case 0:
found = "an unsigned integer";
break;
case 1:
found = "a negative integer";
break;
case 2:
found = "a byte string";
break;
case 4:
found = "an array";
break;
case 5:
found = "a map";
break;
case 6:
found = "a tag";
break;
default: // major type 7
switch (current)
{
case 0xF4:
case 0xF5:
found = "a boolean";
break;
case 0xF6:
found = "null";
break;
case 0xF7:
found = "undefined";
break;
case 0xF9:
case 0xFA:
case 0xFB:
found = "a floating-point number";
break;
case 0xFF:
found = "a break stop code";
break;
default:
found = "a simple value";
break;
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::cbor, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*! /*!
@brief reads a definite-length CBOR byte array @brief reads a definite-length CBOR byte array
@@ -1568,7 +1647,7 @@ class binary_reader
if (top.is_object) if (top.is_object)
{ {
key.clear(); key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string(key) || !sax->key(key))) if (JSON_HEDLEY_UNLIKELY(!get_cbor_object_key(key) || !sax->key(key)))
{ {
return false; return false;
} }
@@ -1861,61 +1940,61 @@ class binary_reader
case 0xCA: // float 32 case 0xCA: // float 32
{ {
float number{}; float number{};
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
} }
case 0xCB: // float 64 case 0xCB: // float 64
{ {
double number{}; double number{};
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
} }
case 0xCC: // uint 8 case 0xCC: // uint 8
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number); return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
} }
case 0xCD: // uint 16 case 0xCD: // uint 16
{ {
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number); return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
} }
case 0xCE: // uint 32 case 0xCE: // uint 32
{ {
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number); return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
} }
case 0xCF: // uint 64 case 0xCF: // uint 64
{ {
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number); return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
} }
case 0xD0: // int 8 case 0xD0: // int 8
{ {
std::int8_t number{}; std::int8_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number); return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
} }
case 0xD1: // int 16 case 0xD1: // int 16
{ {
std::int16_t number{}; std::int16_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number); return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
} }
case 0xD2: // int 32 case 0xD2: // int 32
{ {
std::int32_t number{}; std::int32_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number); return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
} }
case 0xD3: // int 64 case 0xD3: // int 64
{ {
std::int64_t number{}; std::int64_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number); return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
} }
case 0xDC: // array 16 case 0xDC: // array 16
@@ -2069,6 +2148,98 @@ class binary_reader
} }
} }
/*!
@brief reads a MessagePack object key
The MessagePack specification allows any type as a map key, but only
strings have a counterpart in JSON. A key of any other type is rejected
with a message naming that type, rather than the one @ref
get_msgpack_string gives for a malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_msgpack_object_key(string_t& result)
{
const char* found = nullptr;
switch (current)
{
case 0xC0:
found = "nil";
break;
case 0xC2:
case 0xC3:
found = "a boolean";
break;
case 0xCA:
case 0xCB:
found = "a float";
break;
case 0xC4:
case 0xC5:
case 0xC6:
found = "a bin";
break;
case 0xC7:
case 0xC8:
case 0xC9:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
found = "an ext";
break;
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
found = "an integer";
break;
case 0xDC:
case 0xDD:
found = "an array";
break;
case 0xDE:
case 0xDF:
found = "a map";
break;
default:
// fixint, fixmap, and fixarray; strings, EOF, and the unused
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
}
if (current <= 0x7F || current >= 0xE0)
{
found = "an integer";
}
else if (current <= 0x8F)
{
found = "a map";
}
else if (current <= 0x9F)
{
found = "an array";
}
else
{
return get_msgpack_string(result);
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::msgpack, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*! /*!
@brief reads a MessagePack byte array @brief reads a MessagePack byte array
@@ -2231,7 +2402,7 @@ class binary_reader
{ {
get(); get();
key.clear(); key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key))) if (JSON_HEDLEY_UNLIKELY(!get_msgpack_object_key(key) || !sax->key(key)))
{ {
return false; return false;
} }
@@ -2756,7 +2927,7 @@ class binary_reader
{ {
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr)); return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr));
} }
if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(static_cast<number_unsigned_t>(i)))) if (JSON_HEDLEY_UNLIKELY(!emit_unsigned(input_format, i)))
{ {
return false; return false;
} }
@@ -2888,37 +3059,37 @@ class binary_reader
break; break;
} }
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format, number) && sax->number_unsigned(number); return get_number(input_format, number) && emit_unsigned(input_format, number);
} }
case 'U': case 'U':
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format, number) && sax->number_unsigned(number); return get_number(input_format, number) && emit_unsigned(input_format, number);
} }
case 'i': case 'i':
{ {
std::int8_t number{}; std::int8_t number{};
return get_number(input_format, number) && sax->number_integer(number); return get_number(input_format, number) && emit_signed(input_format, number);
} }
case 'I': case 'I':
{ {
std::int16_t number{}; std::int16_t number{};
return get_number(input_format, number) && sax->number_integer(number); return get_number(input_format, number) && emit_signed(input_format, number);
} }
case 'l': case 'l':
{ {
std::int32_t number{}; std::int32_t number{};
return get_number(input_format, number) && sax->number_integer(number); return get_number(input_format, number) && emit_signed(input_format, number);
} }
case 'L': case 'L':
{ {
std::int64_t number{}; std::int64_t number{};
return get_number(input_format, number) && sax->number_integer(number); return get_number(input_format, number) && emit_signed(input_format, number);
} }
case 'u': case 'u':
@@ -2928,7 +3099,7 @@ class binary_reader
break; break;
} }
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format, number) && sax->number_unsigned(number); return get_number(input_format, number) && emit_unsigned(input_format, number);
} }
case 'm': case 'm':
@@ -2938,7 +3109,7 @@ class binary_reader
break; break;
} }
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format, number) && sax->number_unsigned(number); return get_number(input_format, number) && emit_unsigned(input_format, number);
} }
case 'M': case 'M':
@@ -2948,7 +3119,7 @@ class binary_reader
break; break;
} }
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format, number) && sax->number_unsigned(number); return get_number(input_format, number) && emit_unsigned(input_format, number);
} }
case 'h': case 'h':
@@ -3006,13 +3177,13 @@ class binary_reader
case 'd': case 'd':
{ {
float number{}; float number{};
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format, number) && emit_float(input_format, number);
} }
case 'D': case 'D':
{ {
double number{}; double number{};
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format, number) && emit_float(input_format, number);
} }
case 'H': case 'H':
@@ -3479,13 +3650,13 @@ class binary_reader
case 0x8E: // binary32 case 0x8E: // binary32
{ {
float number{}; float number{};
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
} }
case 0x8F: // binary64 case 0x8F: // binary64
{ {
double number{}; double number{};
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), ""); return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
} }
case 0xF8: case 0xF8:
@@ -3551,7 +3722,9 @@ class binary_reader
@brief pass an integer to the SAX parser @brief pass an integer to the SAX parser
Non-negative integers are passed as unsigned, negative integers as signed Non-negative integers are passed as unsigned, negative integers as signed
numbers, like the other binary formats do. numbers, like the other binary formats do. A value that does not fit the
number type is passed as described for @ref emit_unsigned and
@ref emit_signed.
@param[in] number the integer @param[in] number the integer
@return whether the SAX parser accepted the value @return whether the SAX parser accepted the value
@@ -3560,9 +3733,9 @@ class binary_reader
{ {
if (number >= 0) if (number >= 0)
{ {
return sax->number_unsigned(static_cast<number_unsigned_t>(number)); return emit_unsigned(input_format_t::bon8, static_cast<std::uint64_t>(number));
} }
return sax->number_integer(static_cast<number_integer_t>(number)); return emit_signed(input_format_t::bon8, number);
} }
/*! /*!
@@ -3619,8 +3792,7 @@ class binary_reader
value = (value << 8) | static_cast<std::int64_t>(current); value = (value << 8) | static_cast<std::int64_t>(current);
} }
return negative ? sax->number_integer(static_cast<number_integer_t>(-(value + offset))) return emit_bon8_integer(negative ? -(value + offset) : value + offset);
: sax->number_unsigned(static_cast<number_unsigned_t>(value + offset));
} }
/*! /*!
@@ -3917,6 +4089,88 @@ class binary_reader
return true; return true;
} }
/*!
@brief pass a signed integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_integer_t is passed as number_unsigned_t if it is non-negative and
fits there, and as number_float_t otherwise. With the default number
types, every integer the binary formats can encode fits, so this only
matters for narrower custom number types.
@tparam NumberType a signed integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_signed(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->number_integer(static_cast<number_integer_t>(number));
}
if (value_in_range_of<number_unsigned_t>(number))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass an unsigned integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_unsigned_t is passed as number_float_t.
@tparam NumberType an unsigned integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_unsigned(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_unsigned_t>(number)))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass a floating-point number read from the input to the SAX parser
Like the lexer does for JSON text, a finite value that overflows
number_float_t is rejected instead of silently becoming infinity. Infinity
and NaN in the input are passed on unchanged. Integers only overflow if
number_float_t cannot represent 2^64, e.g., a half-precision type.
@tparam NumberType a floating-point or integer type
@param[in] format the current format (for diagnostics)
@param[in] number the number
@return whether the SAX parser accepted the value
@throw out_of_range.406 if a finite @a number overflows number_float_t
*/
template<typename NumberType>
bool emit_float(const input_format_t format, const NumberType number)
{
const auto result = static_cast<number_float_t>(number);
if (JSON_HEDLEY_UNLIKELY(std::isfinite(number) && !std::isfinite(result)))
{
return sax->parse_error(chars_read, get_token_string(),
out_of_range::create(406, exception_message(format, "number overflow", "value"), nullptr));
}
return sax->number_float(result, "");
}
/*! /*!
@brief create a string by reading characters from the input @brief create a string by reading characters from the input
+76 -39
View File
@@ -206,7 +206,6 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter)) : ia(std::move(adapter))
, ignore_comments(ignore_comments_) , ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_) , discard_number_values(discard_number_values_)
{} {}
@@ -222,8 +221,7 @@ class lexer : public lexer_base<BasicJsonType>
// locales // locales
///////////////////// /////////////////////
/// return the locale-dependent decimal point /// return the decimal point of the current locale
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept static char get_decimal_point() noexcept
{ {
const auto* loc = localeconv(); const auto* loc = localeconv();
@@ -1092,9 +1090,10 @@ class lexer : public lexer_base<BasicJsonType>
token_type::value_float if number could be successfully scanned, token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise token_type::parse_error otherwise
@note The scanner is independent of the current locale. Internally, the @note The scanner is independent of the current locale: token_buffer
locale's decimal point is used instead of `.` to work with the always holds `.`. Only the std::strtod fallback of convert_number()
locale-dependent converters. depends on the locale, and it looks up the decimal point right
before converting (see convert_float_locale_aware()).
*/ */
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated. token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{ {
@@ -1183,7 +1182,7 @@ scan_number_zero:
{ {
case '.': case '.':
{ {
add(decimal_point_char); add(current);
decimal_point_position = token_buffer.size() - 1; decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1; goto scan_number_decimal1;
} }
@@ -1220,7 +1219,7 @@ scan_number_any1:
case '.': case '.':
{ {
add(decimal_point_char); add(current);
decimal_point_position = token_buffer.size() - 1; decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1; goto scan_number_decimal1;
} }
@@ -1462,9 +1461,9 @@ scan_number_done:
// Only a number below 1 can carry further insignificant zeros, and only // Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the // while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. Note // answer - so this loop is skipped for all but a few tokens. The
// token_buffer holds the locale's decimal point, so the fraction is // fraction is located through decimal_point_position rather than by
// located through decimal_point_position rather than by searching '.'. // searching '.'.
if (lead_zero != 0) if (lead_zero != 0)
{ {
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
@@ -1482,8 +1481,8 @@ scan_number_done:
@brief convert the number text in token_buffer to its value and token type @brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds the scan_number() state machine or by the contiguous fast path) and holds '.'
locale decimal point in place of '.'. Integers are parsed first and fall as decimal point, independent of the locale. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce back to floating point on overflow. This is shared so both scanners produce
identical results. identical results.
@@ -1563,7 +1562,7 @@ scan_number_done:
// integer conversion above overflowed. Prefer std::from_chars // integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available; // (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the // otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod. // locale-aware strtof/strtod/strtold.
if (parse_float_from_chars(num_begin, num_end, value_float)) if (parse_float_from_chars(num_begin, num_end, value_float))
{ {
return token_type::value_float; return token_type::value_float;
@@ -1572,26 +1571,75 @@ scan_number_done:
// extra pass over the token's bytes, which otherwise shows up on // extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json // high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end) if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float)) && parse_float_fast(num_begin, num_end, value_float))
{ {
return token_type::value_float; return token_type::value_float;
} }
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg) convert_float_locale_aware();
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
return token_type::value_float; return token_type::value_float;
} }
/*!
@brief convert the float in token_buffer with strtof/strtod/strtold
These functions expect the decimal point of the *current* locale, so it is
looked up right before the conversion instead of once when the lexer is
constructed: a locale change in between (by a parser callback, a SAX
handler, or another thread) must not truncate the value (#5198). The
token has been validated before, so if the conversion stops early and the
decimal point changed in the meantime, the locale changed between the
lookup and the call, and the conversion is repeated with the new decimal
point. If the decimal point did not change, a retry cannot succeed: the
locale's decimal point is not a single character (e.g., the two-byte
U+066B of ar_EG.UTF-8 or fa_IR.UTF-8) and cannot be substituted in place.
The value strtod parsed up to that point is kept, as before this change.
Note that changing the locale in another thread *while* strtod runs is
undefined behavior of the C library, which this function cannot prevent.
*/
void convert_float_locale_aware()
{
const bool has_dot = decimal_point_position != std::string::npos;
char decimal_point = get_decimal_point();
for (;;)
{
const bool substitute = has_dot && decimal_point != '.';
if (substitute)
{
token_buffer[decimal_point_position] = static_cast<typename string_t::value_type>(decimal_point);
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
if (substitute)
{
// get_string() hands the token to the SAX interface with '.'
token_buffer[decimal_point_position] = '.';
}
if (JSON_HEDLEY_LIKELY(endptr == token_buffer.data() + token_buffer.size()))
{
return;
}
// retry only if the locale changed; otherwise, this would loop forever
const char current_decimal_point = get_decimal_point();
if (current_decimal_point == decimal_point)
{
return;
}
decimal_point = current_decimal_point;
}
}
/*! /*!
@brief contiguous fast path for scanning a number @brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer per-character get()/add() of scan_number(). On success it fills token_buffer
(with the locale decimal point substituted, as scan_number() does) and (as scan_number() does) and
returns the token type. On anything it does not fully recognize as a returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which token_type::uninitialized, so the caller falls back to scan_number(), which
@@ -1707,16 +1755,11 @@ scan_number_done:
} }
#endif #endif
// materialize the token exactly as scan_number() would, substituting the // materialize the token exactly as scan_number() would. reset() already
// locale decimal point so convert_number()'s strtof fallback stays valid. // cleared token_buffer, so append() fills it (assign() is avoided
// reset() already cleared token_buffer, so append() fills it (assign() is // because custom string_t types need not provide it)
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len); token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
if (dot_index != std::string::npos) decimal_point_position = dot_index;
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1); ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1); position.chars_read_total += (len - 1);
@@ -1983,11 +2026,7 @@ scan_number_done:
/// return current string value (implicitly resets the token; useful only once) /// return current string value (implicitly resets the token; useful only once)
string_t& get_string() string_t& get_string()
{ {
// translate decimal points from locale back to '.' (#4084) // a number token holds '.' regardless of the locale (#4084)
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
{
token_buffer[decimal_point_position] = '.';
}
return token_buffer; return token_buffer;
} }
@@ -2283,9 +2322,7 @@ scan_number_done:
number_unsigned_t value_unsigned = 0; number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0; number_float_t value_float = 0;
/// the decimal point /// the position of the decimal point in token_buffer
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
std::size_t decimal_point_position = std::string::npos; std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the /// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
+8 -13
View File
@@ -118,14 +118,12 @@ std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload long double keep the std::strtof/std::strtold paths (see the templated overload
below). below).
@param[in] first pointer to the first character of the number @param[in] first pointer to the first character of the number
@param[in] last pointer past the last character @param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character @param[out] out the parsed value on success
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod @return true if the value was parsed exactly; false to fall back to strtod
*/ */
template<typename DecimalPointType> inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{ {
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0 #if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in // Clinger's fast path is only exact when double operations are evaluated in
@@ -136,7 +134,6 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
// std::from_chars / std::strtod path. // std::from_chars / std::strtod path.
static_cast<void>(first); static_cast<void>(first);
static_cast<void>(last); static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out); static_cast<void>(out);
return false; return false;
#else #else
@@ -175,7 +172,7 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
++num_digits; ++num_digits;
fractional_digits += static_cast<int>(seen_dot); fractional_digits += static_cast<int>(seen_dot);
} }
else if (static_cast<DecimalPointType>(c) == decimal_point) else if (c == '.')
{ {
if (JSON_HEDLEY_UNLIKELY(seen_dot)) if (JSON_HEDLEY_UNLIKELY(seen_dot))
{ {
@@ -260,8 +257,8 @@ bool parse_float_fast(const char* first, const char* last, DecimalPointType deci
} }
/// fast float path is only exact for `double`; decline for float/long double /// fast float path is only exact for `double`; decline for float/long double
template<typename DecimalPointType, typename FloatType> template<typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
{ {
return false; return false;
} }
@@ -273,9 +270,7 @@ std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it __cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)); a partial parse means the buffer consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations). expects (side-stepping the P4168 divergence between implementations).
+168 -406
View File
@@ -6061,240 +6061,21 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
{ {
// the patch // the patch
basic_json result(value_t::array); basic_json result(value_t::array);
diff_recursively(result, source, target, path, 0);
return result;
}
private: // if the values are the same, return an empty patch
/// @brief two arrays or two objects @ref diff_iteratively is diffing
struct diff_frame
{
diff_frame(const basic_json* source_, const basic_json* target_, const std::size_t path_length_) noexcept
: source(source_), target(target_), path_length(path_length_)
{}
// declared for GCC's -Weffc++, which asks for them in a class with
// pointer members and a non-trivial destructor; the exception
// specifications are left implicit, as GCC 4.8 rejects explicit ones
// that differ from them
diff_frame(const diff_frame&) = default;
diff_frame(diff_frame&&) = default;
diff_frame& operator=(const diff_frame&) = default;
diff_frame& operator=(diff_frame&&) = default;
~diff_frame() = default;
/// the values being diffed, both arrays or both objects
const basic_json* source;
const basic_json* target;
/// the length of their path in `current_path`
std::size_t path_length;
/// arrays: the next index to diff
std::size_t index = 0;
/// objects: the next member of source to look at
const_iterator member{}; // NOLINT(readability-redundant-member-init)
/// objects: the keys common to both, in source's order
std::vector<typename object_t::key_type> common_keys{}; // NOLINT(readability-redundant-member-init)
/// objects: the next entry of common_keys
std::size_t next_common = 0;
/// objects: the "add" operations for keys only target has
basic_json added_ops{}; // NOLINT(readability-redundant-member-init)
};
// The operations of a diff are built by the functions below rather than
// where they are needed: building one takes several temporaries, and
// unoptimized builds give each temporary a stack slot of its own in the
// function it appears in. In diff_recursively, which is on the call stack
// once per nesting level, that made every level cost kilobytes of stack.
/// @brief append a "replace" operation for @a path with @a value to @a result
static void diff_replace(basic_json& result, const string_t& path, const basic_json& value)
{
result.push_back(
{
{"op", "replace"}, {"path", path}, {"value", value}
});
}
/// @brief append a "remove" operation for @a path to @a result
static void diff_remove(basic_json& result, const string_t& path)
{
result.push_back(object(
{
{"op", "remove"}, {"path", path}
}));
}
/// @brief append an "add" operation for @a path with @a value to @a result
static void diff_add(basic_json& result, const string_t& path, const basic_json& value)
{
result.push_back(
{
{"op", "add"}, {"path", path}, {"value", value}
});
}
/// @brief append the "remove" operations for the elements of array
/// @a source from @a index on, and the "add" operations for the
/// elements of array @a target from source's size on, to @a result
static void diff_array_tails(basic_json& result, const basic_json& source, const basic_json& target,
const string_t& path, const std::size_t index)
{
// remove my remaining elements, highest index first; appending
// in that order avoids the quadratic reinsertion done before
for (std::size_t j = source.size(); j > index; --j)
{
diff_remove(result, detail::concat<string_t>(path, '/', detail::to_string<string_t>(j - 1)));
}
// add other remaining elements
for (std::size_t i = source.size(); i < target.size(); ++i)
{
diff_add(result, detail::concat<string_t>(path, "/-"), target[i]);
}
}
/*!
@brief compare the keys of objects @a source and @a target
If the keys both objects have are in the same order in both, and the keys
only @a target has come after them, stores the keys common to both in
source's order in @a common_keys, stores the "add" operations for the keys
only @a target has in @a added_ops, and returns true: the caller then diffs
the objects member by member. Otherwise, appends operations that remove
every member of @a source and add every member of @a target to @a result,
and returns false.
*/
static bool diff_object_keys(basic_json& result, const basic_json& source, const basic_json& target,
const string_t& path, std::vector<typename object_t::key_type>& common_keys,
basic_json& added_ops)
{
// first pass: record, for every source key, whether it is
// common to both objects (in source's iteration order) or
// was deleted (i.e., in source but not in target) -- this is
// a by-product of the target.find() call already needed to
// tell the two cases apart, so it adds no extra lookups. The
// "remove" ops themselves are emitted later, interleaved
// with the per-key diffs in the caller's fast path, to match
// source's original iteration order (as the original,
// pre-reordering-aware implementation did) instead of
// grouping all removes before all per-key diffs.
std::vector<typename object_t::key_type> common_keys_source_order;
for (auto it = source.cbegin(); it != source.cend(); ++it)
{
if (target.find(it.key()) != target.end())
{
common_keys_source_order.push_back(it.key());
}
}
// second pass: find keys that were added (i.e., in target but
// not in source), and record the keys common to both, in
// target's iteration order -- again a by-product of the
// source.find() call already needed to detect added keys. At
// the same time, determine whether every added key comes
// after every common key in target's order (a precondition
// for the fast path, which only ever appends new keys
// at the very end): for an object_t whose iteration order is
// a pure function of the key set (e.g. the default std::map,
// which always iterates in sorted key order), the order
// check further below is always true and this whole
// mechanism is effectively a no-op; it only matters for a
// reorderable object_t such as the one backing `ordered_json`.
// The patch ops for keys that were added (i.e., in target but not
// in source) are built here so the fast path can reuse
// them without a second source.find() per target key. Only
// used by the fast path -- the slow (reordering) path
// rebuilds "add" ops for every key itself.
std::vector<typename object_t::key_type> common_keys_target_order;
bool new_keys_form_suffix = true;
bool seen_new_key = false;
for (auto it = target.cbegin(); it != target.cend(); ++it)
{
if (source.find(it.key()) == source.end())
{
seen_new_key = true;
diff_add(added_ops, detail::concat<string_t>(path, '/', detail::escape(it.key())), it.value());
}
else
{
common_keys_target_order.push_back(it.key());
if (seen_new_key)
{
new_keys_form_suffix = false;
}
}
}
if (common_keys_source_order == common_keys_target_order && new_keys_form_suffix)
{
// fast path: order of common keys already matches (or the
// object_t's iteration order does not depend on
// insertion history), so a plain per-key diff is correct
// and minimal, as before
common_keys = std::move(common_keys_source_order);
return true;
}
// slow path: the common keys are in a different relative
// order in source and target (only possible for a
// reorderable object_t like ordered_map). Building a
// minimal reordering patch is a nontrivial (LCS-like)
// problem; instead, remove every source key -- both
// deleted keys (which must be removed regardless) and
// common keys (removed so they can be re-added in
// target's order) -- and re-add every key that should
// remain, with its final target value, in target's
// order. basic_json::patch()'s "add" operation on an
// object uses operator[], which appends at the end for a
// vector-backed insertion-ordered map when the key does
// not already exist -- so removing a key and then adding
// it moves it to the end, fixing its position.
for (auto it = source.cbegin(); it != source.cend(); ++it)
{
diff_remove(result, detail::concat<string_t>(path, '/', detail::escape(it.key())));
}
// add every key that is either common (just removed
// above) or brand new, in target's iteration order, so
// that the final order after applying the patch matches
// target exactly
for (auto it = target.cbegin(); it != target.cend(); ++it)
{
diff_add(result, detail::concat<string_t>(path, '/', detail::escape(it.key())), it.value());
}
return false;
}
/*!
@brief @ref diff, for values at nesting level @a depth, appending the
operations to @a result
Diffing two arrays or objects calls this function again, once per nesting
level, so values nested deeply enough used to exhaust the call stack and
terminate the process. The descent is bounded here: once @ref
detail::recursion_depth_limit levels have been entered, @ref
diff_iteratively diffs what is left without the call stack.
*/
static void diff_recursively(basic_json& result, const basic_json& source, const basic_json& target,
const string_t& path, const std::size_t depth)
{
// if the values are the same, there is nothing to do
if (source == target) if (source == target)
{ {
return; return result;
}
if (JSON_HEDLEY_UNLIKELY(depth >= detail::recursion_depth_limit()))
{
diff_iteratively(result, source, target, path);
return;
} }
if (source.type() != target.type()) if (source.type() != target.type())
{ {
// different types: replace value // different types: replace value
diff_replace(result, path, target); result.push_back(
return; {
{"op", "replace"}, {"path", path}, {"value", target}
});
return result;
} }
switch (source.type()) switch (source.type())
@@ -6306,50 +6087,185 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
while (i < source.size() && i < target.size()) while (i < source.size() && i < target.size())
{ {
// recursive call to compare array values at index i // recursive call to compare array values at index i
diff_recursively(result, source[i], target[i], detail::concat<string_t>(path, '/', detail::to_string<string_t>(i)), depth + 1); auto temp_diff = diff(source[i], target[i], detail::concat<string_t>(path, '/', detail::to_string<string_t>(i)));
result.insert(result.end(), temp_diff.begin(), temp_diff.end());
++i; ++i;
} }
// We now reached the end of at least one array // We now reached the end of at least one array
// in a second pass, traverse the remaining elements // in a second pass, traverse the remaining elements
diff_array_tails(result, source, target, path, i);
// remove my remaining elements, highest index first; appending
// in that order avoids the quadratic reinsertion done before
for (std::size_t j = source.size(); j > i; --j)
{
result.push_back(object(
{
{"op", "remove"},
{"path", detail::concat<string_t>(path, '/', detail::to_string<string_t>(j - 1))}
}));
}
i = source.size();
// add other remaining elements
while (i < target.size())
{
result.push_back(
{
{"op", "add"},
{"path", detail::concat<string_t>(path, "/-")},
{"value", target[i]}
});
++i;
}
break; break;
} }
case value_t::object: case value_t::object:
{ {
std::vector<typename object_t::key_type> common_keys; // first pass: record, for every source key, whether it is
basic_json added_ops(value_t::array); // common to both objects (in source's iteration order) or
if (diff_object_keys(result, source, target, path, common_keys, added_ops)) // was deleted (i.e., in source but not in target) -- this is
// a by-product of the target.find() call already needed to
// tell the two cases apart, so it adds no extra lookups. The
// "remove" ops themselves are emitted later, interleaved
// with the recursive per-key diffs in the fast path below,
// to match source's original iteration order (as the
// original, pre-reordering-aware implementation did) instead
// of grouping all removes before all recursive diffs.
std::vector<typename object_t::key_type> common_keys_source_order;
for (auto it = source.cbegin(); it != source.cend(); ++it)
{ {
// fast path: common_keys is, by construction, the if (target.find(it.key()) != target.end())
// subsequence of source's keys that are common to both {
// objects, in source's iteration order -- so it can be common_keys_source_order.push_back(it.key());
// walked in lockstep with `source` using a cheap key }
// comparison instead of another lookup. Deleted keys }
// (those source keys not in common_keys) are interleaved
// here too, in source's original order, to match the // second pass: find keys that were added (i.e., in target but
// historical (pre-reordering-aware) output order. // not in source), and record the keys common to both, in
auto common_it = common_keys.cbegin(); // target's iteration order -- again a by-product of the
// source.find() call already needed to detect added keys. At
// the same time, determine whether every added key comes
// after every common key in target's order (a precondition
// for the fast path below, which only ever appends new keys
// at the very end): for an object_t whose iteration order is
// a pure function of the key set (e.g. the default std::map,
// which always iterates in sorted key order), the order
// check further below is always true and this whole
// mechanism is effectively a no-op; it only matters for a
// reorderable object_t such as the one backing `ordered_json`.
// patch ops for keys that were added (i.e., in target but not
// in source); built here so the fast path below can reuse
// them without a second source.find() per target key. Only
// used by the fast path -- the slow (reordering) path
// rebuilds "add" ops for every key itself.
std::vector<typename object_t::key_type> common_keys_target_order;
basic_json added_ops(value_t::array);
bool new_keys_form_suffix = true;
bool seen_new_key = false;
for (auto it = target.cbegin(); it != target.cend(); ++it)
{
if (source.find(it.key()) == source.end())
{
seen_new_key = true;
const auto path_key = detail::concat<string_t>(path, '/', detail::escape(it.key()));
added_ops.push_back(
{
{"op", "add"}, {"path", path_key},
{"value", it.value()}
});
}
else
{
common_keys_target_order.push_back(it.key());
if (seen_new_key)
{
new_keys_form_suffix = false;
}
}
}
if (common_keys_source_order == common_keys_target_order && new_keys_form_suffix)
{
// fast path: order of common keys already matches (or the
// object_t's iteration order does not depend on
// insertion history), so a plain per-key recursive diff
// is correct and minimal, as before. common_keys_source_order
// is, by construction, the subsequence of source's keys
// that are common to both objects, in source's iteration
// order -- so it can be walked in lockstep with `source`
// using a cheap key comparison instead of another lookup.
// Deleted keys (those source keys not in common_keys_source_order)
// are interleaved here too, in source's original order, to
// match the historical (pre-reordering-aware) output order.
auto common_it = common_keys_source_order.cbegin();
for (auto it = source.cbegin(); it != source.cend(); ++it) for (auto it = source.cbegin(); it != source.cend(); ++it)
{ {
if (common_it != common_keys.cend() && it.key() == *common_it) if (common_it != common_keys_source_order.cend() && it.key() == *common_it)
{ {
diff_recursively(result, it.value(), target[it.key()], detail::concat<string_t>(path, '/', detail::escape(it.key())), depth + 1); const auto path_key = detail::concat<string_t>(path, '/', detail::escape(it.key()));
auto temp_diff = diff(it.value(), target[it.key()], path_key);
result.insert(result.end(), temp_diff.begin(), temp_diff.end());
++common_it; ++common_it;
} }
else else
{ {
// found a key that is not in target -> remove it // found a key that is not in target -> remove it
diff_remove(result, detail::concat<string_t>(path, '/', detail::escape(it.key()))); const auto path_key = detail::concat<string_t>(path, '/', detail::escape(it.key()));
result.push_back(object(
{
{"op", "remove"}, {"path", path_key}
}));
} }
} }
// append the "add" ops for brand-new keys collected by // append the "add" ops for brand-new keys collected above
// diff_object_keys -- no second source.find() per target // during the pass over target -- no second source.find()
// key needed // per target key needed
result.insert(result.end(), added_ops.begin(), added_ops.end()); result.insert(result.end(), added_ops.begin(), added_ops.end());
} }
else
{
// slow path: the common keys are in a different relative
// order in source and target (only possible for a
// reorderable object_t like ordered_map). Building a
// minimal reordering patch is a nontrivial (LCS-like)
// problem; instead, remove every source key -- both
// deleted keys (which must be removed regardless) and
// common keys (removed so they can be re-added in
// target's order) -- and re-add every key that should
// remain, with its final target value, in target's
// order. basic_json::patch()'s "add" operation on an
// object uses operator[], which appends at the end for a
// vector-backed insertion-ordered map when the key does
// not already exist -- so removing a key and then adding
// it moves it to the end, fixing its position.
for (auto it = source.cbegin(); it != source.cend(); ++it)
{
const auto path_key = detail::concat<string_t>(path, '/', detail::escape(it.key()));
result.push_back(object(
{
{"op", "remove"}, {"path", path_key}
}));
}
// add every key that is either common (just removed
// above) or brand new, in target's iteration order, so
// that the final order after applying the patch matches
// target exactly
for (auto it = target.cbegin(); it != target.cend(); ++it)
{
const auto path_key = detail::concat<string_t>(path, '/', detail::escape(it.key()));
result.push_back(
{
{"op", "add"}, {"path", path_key},
{"value", it.value()}
});
}
}
break; break;
} }
@@ -6364,170 +6280,16 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
default: default:
{ {
// both primitive types: replace value // both primitive types: replace value
diff_replace(result, path, target); result.push_back(
{
{"op", "replace"}, {"path", path}, {"value", target}
});
break; break;
} }
} }
return result;
} }
/*!
@brief @ref diff without the call stack, appending the operations to
@a result
Produces the same operations as @ref diff_recursively. Only reached for
values nested more deeply than @ref detail::recursion_depth_limit.
*/
static void diff_iteratively(basic_json& result, const basic_json& source, const basic_json& target,
const string_t& path)
{
// The arrays and objects being diffed are kept on an explicit stack,
// and every pair of elements is still diffed completely before the
// next one, so the operations come out in the same order as in
// diff_recursively. The path of the values being diffed is kept in
// one buffer that grows and shrinks with the stack, rather than in a
// new string per level.
std::vector<diff_frame> stack;
string_t current_path = path;
// diff `s` against `t`, whose path is current_path: primitives,
// values of different types, and objects whose members were reordered
// are handled right away; arrays and other objects get a frame
const auto enter = [&result, &stack, &current_path](const basic_json & s, const basic_json & t)
{
// if the values are the same, there is nothing to do. Arrays and
// objects are not compared up front: comparing them visits
// everything below them, so doing that at every level would take
// quadratic time in the nesting depth - equal ones yield no
// operations anyway.
if ((!s.is_structured() || !t.is_structured()) && s == t)
{
return;
}
if (s.type() != t.type())
{
// different types: replace value
diff_replace(result, current_path, t);
return;
}
switch (s.type())
{
case value_t::array:
{
stack.emplace_back(&s, &t, current_path.size());
return;
}
case value_t::object:
{
std::vector<typename object_t::key_type> common_keys;
basic_json added_ops(value_t::array);
if (diff_object_keys(result, s, t, current_path, common_keys, added_ops))
{
// fast path: the frame walks source in lockstep with
// common_keys, as diff_recursively does, and appends
// added_ops once all members are done
stack.emplace_back(&s, &t, current_path.size());
stack.back().member = s.cbegin();
stack.back().common_keys = std::move(common_keys);
stack.back().added_ops = std::move(added_ops);
}
return;
}
case value_t::null:
case value_t::string:
case value_t::boolean:
case value_t::number_integer:
case value_t::number_unsigned:
case value_t::number_float:
case value_t::binary:
case value_t::discarded:
default:
{
// both primitive types: replace value
diff_replace(result, current_path, t);
return;
}
}
};
enter(source, target);
while (!stack.empty())
{
// the frame is copied out member by member and changed through
// stack.back(): enter() may push a frame and the end of the loop
// pops it, either of which would invalidate a reference to it
const basic_json* const s = stack.back().source;
const basic_json* const t = stack.back().target;
const std::size_t path_length = stack.back().path_length;
const std::size_t depth = stack.size();
if (s->is_array())
{
const auto& source_array = *s->m_data.m_value.array;
const auto& target_array = *t->m_data.m_value.array;
// first pass: traverse common elements
const std::size_t i = stack.back().index;
if (i < source_array.size() && i < target_array.size())
{
++stack.back().index;
detail::concat_into(current_path, '/', detail::to_string<string_t>(i));
enter(source_array[i], target_array[i]);
if (stack.size() == depth)
{
current_path.resize(path_length);
}
continue;
}
// We now reached the end of at least one array
// in a second pass, traverse the remaining elements
diff_array_tails(result, *s, *t, current_path, i);
}
else
{
const const_iterator it = stack.back().member;
if (it != s->cend())
{
++stack.back().member;
const std::size_t next_common = stack.back().next_common;
if (next_common < stack.back().common_keys.size() && it.key() == stack.back().common_keys[next_common])
{
++stack.back().next_common;
const basic_json& target_value = (*t)[it.key()];
detail::concat_into(current_path, '/', detail::escape(it.key()));
enter(it.value(), target_value);
if (stack.size() == depth)
{
current_path.resize(path_length);
}
}
else
{
// found a key that is not in target -> remove it
diff_remove(result, detail::concat<string_t>(current_path, '/', detail::escape(it.key())));
}
continue;
}
// append the "add" ops for brand-new keys collected when the
// object was entered
result.insert(result.end(), stack.back().added_ops.begin(), stack.back().added_ops.end());
}
// this array or object is done: continue with the one it is in
stack.pop_back();
if (!stack.empty())
{
current_path.resize(stack.back().path_length);
}
}
}
public:
/// @} /// @}
//////////////////////////////// ////////////////////////////////
File diff suppressed because it is too large Load Diff
+141
View File
@@ -11,7 +11,12 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <cmath>
#include <fstream> #include <fstream>
#include <limits>
#include <map>
#include <string>
#include <vector>
#include "make_test_data_available.hpp" #include "make_test_data_available.hpp"
TEST_CASE("Binary Formats" * doctest::skip()) TEST_CASE("Binary Formats" * doctest::skip())
@@ -224,3 +229,139 @@ TEST_CASE("Binary Formats" * doctest::skip())
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450)); CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450));
} }
} }
namespace
{
// the binary formats as function pointers for "Binary formats with narrow number types";
// named functions rather than lambdas, because clang 3.5 cannot convert a lambda
// to a function pointer in the braced initializer of the format table
using narrow_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int32_t, std::uint32_t, float>;
using bytes = std::vector<std::uint8_t>;
bytes encode_cbor(const json& j)
{
return json::to_cbor(j);
}
narrow_json decode_cbor(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_cbor(v, true, allow_exceptions);
}
bytes encode_msgpack(const json& j)
{
return json::to_msgpack(j);
}
narrow_json decode_msgpack(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_msgpack(v, true, allow_exceptions);
}
bytes encode_ubjson(const json& j)
{
return json::to_ubjson(j);
}
narrow_json decode_ubjson(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_ubjson(v, true, allow_exceptions);
}
bytes encode_bjdata(const json& j)
{
return json::to_bjdata(j);
}
narrow_json decode_bjdata(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_bjdata(v, true, allow_exceptions);
}
// BSON can only store numbers as object members
bytes encode_bson(const json& j)
{
return json::to_bson(json{{"a", j}});
}
narrow_json decode_bson(const bytes& v, bool allow_exceptions)
{
const auto result = narrow_json::from_bson(v, true, allow_exceptions);
return result.is_discarded() ? result : result.at("a");
}
bytes encode_bon8(const json& j)
{
return json::to_bon8(j);
}
narrow_json decode_bon8(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_bon8(v, true, allow_exceptions);
}
} // namespace
TEST_CASE("Binary formats with narrow number types")
{
// Numbers that do not fit the number types are handled like the lexer
// handles them in JSON text: an integer that fits neither integer type is
// stored as a floating-point number, and a finite floating-point number
// that overflows number_float_t is rejected with out_of_range.406.
struct binary_format
{
const char* name;
bytes (*encode)(const json&);
narrow_json (*decode)(const bytes&, bool);
};
const std::vector<binary_format> formats =
{
{"CBOR", encode_cbor, decode_cbor},
{"MessagePack", encode_msgpack, decode_msgpack},
{"UBJSON", encode_ubjson, decode_ubjson},
{"BJData", encode_bjdata, decode_bjdata},
{"BSON", encode_bson, decode_bson},
{"BON8", encode_bon8, decode_bon8},
};
for (const auto& format : formats)
{
const std::string name = format.name;
INFO("format := ", name);
const auto roundtrip = [&format](const json & j)
{
return format.decode(format.encode(j), true);
};
// integers that fit keep their type
CHECK(roundtrip(json(-5)).is_number_integer());
CHECK(roundtrip(json(-5)).get<std::int32_t>() == -5);
CHECK(roundtrip(json(3000000000u)).is_number_unsigned());
CHECK(roundtrip(json(3000000000u)).get<std::uint32_t>() == 3000000000u);
// integers that fit neither integer type are stored as float
CHECK(roundtrip(json(5000000000u)).is_number_float());
CHECK(roundtrip(json(5000000000u)).get<float>() == 5000000000.0f);
if (name != "BON8") // BON8 cannot encode integers above INT64_MAX
{
CHECK(roundtrip(json(10000000000000000000u)).is_number_float());
CHECK(roundtrip(json(10000000000000000000u)).get<float>() == 10000000000000000000.0f);
}
CHECK(roundtrip(json(-3000000000LL)).is_number_float());
CHECK(roundtrip(json(-3000000000LL)).get<float>() == -3000000000.0f);
CHECK(roundtrip(json(-5000000000LL)).is_number_float());
CHECK(roundtrip(json(-5000000000LL)).get<float>() == -5000000000.0f);
// floating-point numbers that fit
CHECK(roundtrip(json(1.5)).get<float>() == 1.5f);
const auto just_above_max = std::nextafter(static_cast<double>((std::numeric_limits<float>::max)()),
std::numeric_limits<double>::infinity());
CHECK(roundtrip(json(just_above_max)).get<float>() == (std::numeric_limits<float>::max)());
// infinity and NaN are passed on
CHECK(std::isinf(roundtrip(json(std::numeric_limits<double>::infinity())).get<float>()));
CHECK(std::isnan(roundtrip(json(std::numeric_limits<double>::quiet_NaN())).get<float>()));
// finite floating-point numbers that overflow number_float_t are rejected
const std::string message = "[json.exception.out_of_range.406] syntax error while parsing " + name
+ " value: number overflow";
CHECK_THROWS_WITH_AS(roundtrip(json(1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK_THROWS_WITH_AS(roundtrip(json(-1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK(format.decode(format.encode(json(1e300)), false).is_discarded());
}
}
+59 -16
View File
@@ -1830,10 +1830,51 @@ TEST_CASE("CBOR")
SECTION("invalid string in map") SECTION("invalid string in map")
{ {
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01}), true, false).is_discarded()); CHECK(json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01}), true, false).is_discarded());
} }
SECTION("non-string key (see #2766 and #3381)")
{
// only text strings map to JSON object keys; any other key is
// rejected with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0xA1, 0x01, 0x01}, "an unsigned integer; last byte: 0x01"},
{{0xA1, 0x20, 0x01}, "a negative integer; last byte: 0x20"},
{{0xA1, 0x41, 0x61, 0x01}, "a byte string; last byte: 0x41"},
{{0xA1, 0x80, 0x01}, "an array; last byte: 0x80"},
{{0xA1, 0xA0, 0x01}, "a map; last byte: 0xA0"},
{{0xA1, 0xC0, 0x61, 0x61, 0x01}, "a tag; last byte: 0xC0"},
{{0xA1, 0xF4, 0x01}, "a boolean; last byte: 0xF4"},
{{0xA1, 0xF5, 0x01}, "a boolean; last byte: 0xF5"},
{{0xA1, 0xF6, 0x01}, "null; last byte: 0xF6"},
{{0xA1, 0xF7, 0x01}, "undefined; last byte: 0xF7"},
{{0xA1, 0xF9, 0x3C, 0x00, 0x01}, "a floating-point number; last byte: 0xF9"},
{{0xA1, 0xFA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFA"},
{{0xA1, 0xFB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFB"},
{{0xA1, 0xE0, 0x01}, "a simple value; last byte: 0xE0"},
{{0xA1, 0xF8, 0x20, 0x01}, "a simple value; last byte: 0xF8"},
// indefinite-length map
{{0xBF, 0x01, 0x01, 0xFF}, "an unsigned integer; last byte: 0x01"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_cbor(c.first, true, false).is_discarded());
}
// a key of major type 3 with a reserved length is still reported as
// a malformed string, and a missing key as the end of input
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing CBOR string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)") SECTION("invalid UTF-8 in string (see #5529)")
{ {
// a two-character text string (major type 3) whose bytes are not // a two-character text string (major type 3) whose bytes are not
@@ -2284,7 +2325,7 @@ TEST_CASE("CBOR indefinite-length strings do not recurse per chunk")
SECTION("a break marker outside an indefinite-length string is not a string") SECTION("a break marker outside an indefinite-length string is not a string")
{ {
// 0xFF only closes a string that was opened; on its own it is not one // 0xFF only closes a string that was opened; on its own it is not one
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
} }
} }
@@ -3146,7 +3187,8 @@ TEST_CASE("Tagged values")
// CBOR encodes negative integers as: result = -1 - n // CBOR encodes negative integers as: result = -1 - n
// For type 0x3B, n is an 8-byte uint64_t. Valid range for n with // For type 0x3B, n is an 8-byte uint64_t. Valid range for n with
// the default int64_t is [0, INT64_MAX], producing results in [INT64_MIN, -1]. // the default int64_t is [0, INT64_MAX], producing results in [INT64_MIN, -1].
// When n > INT64_MAX, the result exceeds int64_t range and is rejected. // When n > INT64_MAX, the result exceeds int64_t range and is stored
// as a floating-point number, as the lexer does for JSON text.
SECTION("n = 0 is valid (result = -1)") SECTION("n = 0 is valid (result = -1)")
{ {
@@ -3167,33 +3209,34 @@ TEST_CASE("Tagged values")
CHECK(result.get<int64_t>() == (std::numeric_limits<int64_t>::min)()); CHECK(result.get<int64_t>() == (std::numeric_limits<int64_t>::min)());
} }
SECTION("n = INT64_MAX + 1 is rejected (overflow)") SECTION("n = INT64_MAX + 1 is stored as float")
{ {
// n = INT64_MAX + 1 (0x8000000000000000) // n = INT64_MAX + 1 (0x8000000000000000)
// result = -1 - n = -9223372036854775809, which exceeds int64_t range // result = -1 - n = -9223372036854775809, which exceeds int64_t range;
// the nearest double is -9223372036854775808.0
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}; const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
json _; const auto result = json::from_cbor(input);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), CHECK(result.is_number_float());
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow", CHECK(result.get<double>() == -9223372036854775808.0);
json::parse_error); CHECK(result == json::parse("-9223372036854775809"));
} }
SECTION("n = UINT64_MAX is rejected (overflow)") SECTION("n = UINT64_MAX is stored as float")
{ {
// n = UINT64_MAX (0xFFFFFFFFFFFFFFFF) // n = UINT64_MAX (0xFFFFFFFFFFFFFFFF)
// result = -1 - n = -18446744073709551616, which exceeds int64_t range // result = -1 - n = -18446744073709551616, which exceeds int64_t range
const std::vector<uint8_t> input = {0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF}; const std::vector<uint8_t> input = {0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
json _; const auto result = json::from_cbor(input);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), CHECK(result.is_number_float());
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow", CHECK(result.get<double>() == -18446744073709551616.0);
json::parse_error); CHECK(result == json::parse("-18446744073709551616"));
} }
SECTION("overflow with allow_exceptions=false returns discarded") SECTION("overflow with allow_exceptions=false is not an error")
{ {
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}; const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
const auto result = json::from_cbor(input, true, false); const auto result = json::from_cbor(input, true, false);
CHECK(result.is_discarded()); CHECK(result.is_number_float());
} }
} }
+1 -1
View File
@@ -666,7 +666,7 @@ TEST_CASE("parse_float_fast declines what it cannot convert exactly")
// always safe: the caller then falls back to a slower, exact conversion. // always safe: the caller then falls back to a slower, exact conversion.
const auto fast = [](const std::string & s, double & out) const auto fast = [](const std::string & s, double & out)
{ {
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), '.', out); return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out);
}; };
double out = 0; double out = 0;
-153
View File
@@ -15,65 +15,8 @@ using nlohmann::json;
#endif #endif
#include <fstream> #include <fstream>
#include <string>
#include <vector>
#include "make_test_data_available.hpp" #include "make_test_data_available.hpp"
namespace
{
// alternating objects and arrays nested `depth` levels deep, with members that
// depend on `variant` at some levels, so diffing two variants yields
// operations on many levels: replacing the innermost value, adding, removing,
// and (for ordered_json) reordering members, and changing array lengths
template<typename BasicJsonType>
BasicJsonType nested(const std::size_t depth, const int variant)
{
BasicJsonType value = variant;
for (std::size_t i = 0; i < depth; ++i)
{
if (i % 2 == 0)
{
BasicJsonType object = BasicJsonType::object();
if ((i + static_cast<std::size_t>(variant)) % 7 == 0)
{
object["x"] = i;
}
if (variant == 2 && i % 11 == 0)
{
object["z"] = "z";
}
object["a"] = std::move(value);
if (variant == 1 && i % 5 == 0)
{
object["y"] = 1;
}
value = std::move(object);
}
else
{
BasicJsonType array = BasicJsonType::array({std::move(value)});
if ((i + static_cast<std::size_t>(variant)) % 3 == 0)
{
array.push_back(i);
}
value = std::move(array);
}
}
return value;
}
// a path of `depth` reference tokens, as nested() nests its values
std::string nested_path(const std::size_t depth)
{
std::string path;
for (std::size_t i = depth; i > 0; --i)
{
path += (i - 1) % 2 == 0 ? "/a" : "/0";
}
return path;
}
} // namespace
TEST_CASE("JSON patch") TEST_CASE("JSON patch")
{ {
SECTION("examples from RFC 6902") SECTION("examples from RFC 6902")
@@ -1809,102 +1752,6 @@ TEST_CASE("JSON patch - diff emits array removals in descending index order")
} }
} }
TEST_CASE("JSON patch: diff of deeply nested values")
{
SECTION("the diff reproduces the target at every depth")
{
// depths on either side of the nesting depth up to which diff()
// recurses (detail::recursion_depth_limit(), 128); not every depth up
// to 300, as the test would then time out under Valgrind
std::vector<std::size_t> depths;
for (std::size_t depth = 0; depth <= 16; ++depth)
{
depths.push_back(depth);
}
for (std::size_t depth = 120; depth <= 136; ++depth)
{
depths.push_back(depth);
}
depths.push_back(300);
for (const auto depth : depths)
{
CAPTURE(depth);
for (int from = 0; from < 3; ++from)
{
for (int to = 0; to < 3; ++to)
{
CAPTURE(from);
CAPTURE(to);
const auto source = nested<json>(depth, from);
const auto target = nested<json>(depth, to);
const auto patch = json::diff(source, target);
CHECK(source.patch(patch) == target);
CHECK(patch.empty() == (from == to));
const auto ordered_source = nested<nlohmann::ordered_json>(depth, from);
const auto ordered_target = nested<nlohmann::ordered_json>(depth, to);
CHECK(ordered_source.patch(nlohmann::ordered_json::diff(ordered_source, ordered_target)) == ordered_target);
}
}
}
}
SECTION("a difference only in the innermost value is one replace operation")
{
for (std::size_t depth = 0; depth <= 300; ++depth)
{
CAPTURE(depth);
json source = 1;
json target = 2;
for (std::size_t i = 0; i < depth; ++i)
{
source = i % 2 == 0 ? json::object({{"a", std::move(source)}}) : json::array({std::move(source)});
target = i % 2 == 0 ? json::object({{"a", std::move(target)}}) : json::array({std::move(target)});
}
CHECK(json::diff(source, target, "/root") == json::array({{{"op", "replace"}, {"path", "/root" + nested_path(depth)}, {"value", 2}}}));
}
}
SECTION("values nested too deeply for the call stack (#5393)")
{
// diff() used to recurse once per nesting level, and compared the
// values with operator== on every level. The values are only
// parsed and diffed, never copied or compared, since those recurse
// too.
const std::size_t depth = 100000;
for (const bool objects :
{
false, true
})
{
CAPTURE(objects);
std::string source_text;
std::string target_text;
std::string equal_text;
std::string path;
for (std::size_t i = 0; i < depth; ++i)
{
source_text += objects ? "{\"a\":" : "[";
path += objects ? "/a" : "/0";
}
target_text = source_text + "2";
equal_text = source_text + "1";
source_text += "1";
const std::string closing(depth, objects ? '}' : ']');
const auto source = json::parse(source_text + closing);
const auto patch = json::diff(source, json::parse(target_text + closing));
REQUIRE(patch.size() == 1);
CHECK(patch[0]["op"] == "replace");
CHECK(patch[0]["path"] == path);
CHECK(patch[0]["value"] == 2);
CHECK(json::diff(source, json::parse(equal_text + closing)).empty());
}
}
}
TEST_CASE("JSON patch - every operation on ordered_json") TEST_CASE("JSON patch - every operation on ordered_json")
{ {
using nlohmann::ordered_json; using nlohmann::ordered_json;
+210
View File
@@ -12,7 +12,12 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <array>
#include <clocale> #include <clocale>
#include <map>
#include <string>
#include <utility>
#include <vector>
struct ParserImpl final: public nlohmann::json_sax<json> struct ParserImpl final: public nlohmann::json_sax<json>
{ {
@@ -175,3 +180,208 @@ TEST_CASE("locale-dependent test (LC_NUMERIC=de_DE)")
MESSAGE("locale de_DE is not usable"); MESSAGE("locale de_DE is not usable");
} }
} }
namespace
{
// records the numbers of a flat array and switches LC_NUMERIC to the given
// locale once the array opens - after the lexer was constructed, but before
// any number in the array is lexed
struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
{
explicit LocaleSwitchingSax(const char* switch_to)
: locale_after_open(switch_to)
{}
bool null() override
{
return true;
}
bool boolean(bool /*val*/) override
{
return true;
}
bool number_integer(json::number_integer_t /*val*/) override
{
return true;
}
bool number_unsigned(json::number_unsigned_t /*val*/) override
{
return true;
}
bool number_float(json::number_float_t val, const json::string_t& s) override
{
values.push_back(val);
strings.push_back(s);
return true;
}
bool string(json::string_t& /*val*/) override
{
return true;
}
bool binary(json::binary_t& /*val*/) override
{
return true;
}
bool start_object(std::size_t /*val*/) override
{
return true;
}
bool key(json::string_t& /*val*/) override
{
return true;
}
bool end_object() override
{
return true;
}
bool start_array(std::size_t /*val*/) override
{
switched = std::setlocale(LC_NUMERIC, locale_after_open.c_str()) != nullptr;
return true;
}
bool end_array() override
{
return true;
}
bool parse_error(std::size_t /*val*/, const std::string& /*val*/, const nlohmann::detail::exception& /*val*/) override
{
return false;
}
std::string locale_after_open;
bool switched = false;
std::vector<json::number_float_t> values {}; // NOLINT(readability-redundant-member-init)
std::vector<json::string_t> strings {}; // NOLINT(readability-redundant-member-init)
};
} // namespace
TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
{
// The numbers are chosen so that the conversion also takes the strtod
// fallback, which honors the locale that is current at conversion time:
// too many significant digits for Clinger's fast path, an underflow that
// std::from_chars rejects, and a plain value.
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
std::string text = "[";
for (const auto& n : numbers)
{
text += (text.size() == 1 ? "" : ",") + n;
}
text += "]";
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
// reference values, parsed without a locale switch
REQUIRE(std::setlocale(LC_NUMERIC, "C") != nullptr);
const json expected = json::parse(text);
const long_double_json expected_ld = long_double_json::parse(text);
const std::array<std::pair<const char*, const char*>, 2> transitions =
{
{
{"C", "de_DE"},
{"de_DE", "C"}
}
};
for (const auto& transition : transitions)
{
CAPTURE(transition.first);
CAPTURE(transition.second);
if (std::setlocale(LC_NUMERIC, transition.first) == nullptr)
{
MESSAGE("locale is not usable");
continue;
}
// SAX parsing
{
LocaleSwitchingSax sax(transition.second);
CHECK(json::sax_parse(text, &sax));
if (sax.switched)
{
CHECK(sax.values == expected.get<std::vector<json::number_float_t>>());
CHECK(sax.strings == numbers);
}
}
// DOM parsing with a callback
{
bool switched = false;
const auto cb = [&](int /*depth*/, json::parse_event_t event, json& /*parsed*/) noexcept
{
if (event == json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const json j = json::parse(text, cb);
if (switched)
{
CHECK(j == expected);
}
}
// a long double goes through std::strtold unless std::from_chars supports it
{
bool switched = false;
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept
{
if (event == long_double_json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const long_double_json j = long_double_json::parse(text, cb);
if (switched)
{
CHECK(j == expected_ld);
}
}
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
TEST_CASE("locale with a multi-byte decimal point")
{
// Some locales use a decimal point that is not a single character, e.g.
// U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be
// substituted in place for '.', so the strtod fallback stops early. The
// conversion must still terminate rather than retry forever.
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
bool tested = false;
for (const char* name : names)
{
if (std::setlocale(LC_NUMERIC, name) == nullptr)
{
continue;
}
const std::string decimal_point = std::localeconv()->decimal_point;
if (decimal_point.size() < 2)
{
continue;
}
CAPTURE(name);
tested = true;
// too many significant digits for Clinger's fast path, and an underflow
// that std::from_chars rejects: both reach the strtod fallback
json j;
CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
CHECK(j.is_array());
CHECK(json::accept("3.14159265358979323846"));
// a value the locale-independent paths convert is not affected
CHECK(json::parse("12.5") == 12.5);
}
if (!tested)
{
MESSAGE("no locale with a multi-byte decimal point is usable");
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
+60 -1
View File
@@ -1551,10 +1551,69 @@ TEST_CASE("MessagePack")
SECTION("invalid string in map") SECTION("invalid string in map")
{ {
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found an integer; last byte: 0xFF", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01}), true, false).is_discarded()); CHECK(json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01}), true, false).is_discarded());
} }
SECTION("non-string key (see #3381)")
{
// only strings map to JSON object keys; any other key is rejected
// with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0x81, 0xC0, 0x01}, "nil; last byte: 0xC0"},
{{0x81, 0xC2, 0x01}, "a boolean; last byte: 0xC2"},
{{0x81, 0xC3, 0x01}, "a boolean; last byte: 0xC3"},
{{0x81, 0xCA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a float; last byte: 0xCA"},
{{0x81, 0xCB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a float; last byte: 0xCB"},
{{0x81, 0xC4, 0x00, 0x01}, "a bin; last byte: 0xC4"},
{{0x81, 0xC5, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC5"},
{{0x81, 0xC6, 0x00, 0x00, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC6"},
{{0x81, 0xC7, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC7"},
{{0x81, 0xC8, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC8"},
{{0x81, 0xC9, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC9"},
{{0x81, 0xD4, 0x01, 0x00, 0x01}, "an ext; last byte: 0xD4"},
{{0x81, 0xD5, 0x01, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD5"},
{{0x81, 0xD6, 0x01, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD6"},
{{0x81, 0xD7, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD7"},
{{0x81, 0xD8, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD8"},
{{0x81, 0xCC, 0x01, 0x01}, "an integer; last byte: 0xCC"},
{{0x81, 0xCD, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCD"},
{{0x81, 0xCE, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCE"},
{{0x81, 0xCF, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCF"},
{{0x81, 0xD0, 0x01, 0x01}, "an integer; last byte: 0xD0"},
{{0x81, 0xD1, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD1"},
{{0x81, 0xD2, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD2"},
{{0x81, 0xD3, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD3"},
{{0x81, 0x00, 0x01}, "an integer; last byte: 0x00"},
{{0x81, 0x7F, 0x01}, "an integer; last byte: 0x7F"},
{{0x81, 0xE0, 0x01}, "an integer; last byte: 0xE0"},
{{0x81, 0x80, 0x01}, "a map; last byte: 0x80"},
{{0x81, 0x8F, 0x01}, "a map; last byte: 0x8F"},
{{0x81, 0xDE, 0x00, 0x00, 0x01}, "a map; last byte: 0xDE"},
{{0x81, 0xDF, 0x00, 0x00, 0x00, 0x00, 0x01}, "a map; last byte: 0xDF"},
{{0x81, 0x90, 0x01}, "an array; last byte: 0x90"},
{{0x81, 0x9F, 0x01}, "an array; last byte: 0x9F"},
{{0x81, 0xDC, 0x00, 0x00, 0x01}, "an array; last byte: 0xDC"},
{{0x81, 0xDD, 0x00, 0x00, 0x00, 0x00, 0x01}, "an array; last byte: 0xDD"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_msgpack(c.first, true, false).is_discarded());
}
json _;
// the unused byte 0xC1 is still reported as a malformed string
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xC1, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xC1", json::parse_error&);
// a missing key is still reported as the end of input
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)") SECTION("invalid UTF-8 in string (see #5529)")
{ {
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8 // a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
+3 -3
View File
@@ -1018,7 +1018,7 @@ TEST_CASE("regression tests 1")
}; };
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x98", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an array; last byte: 0x98", json::parse_error&);
// related test case: nonempty UTF-8 string (indefinite length) // related test case: nonempty UTF-8 string (indefinite length)
std::vector<uint8_t> const vec1 {0x7f, 0x61, 0x61}; std::vector<uint8_t> const vec1 {0x7f, 0x61, 0x61};
@@ -1065,7 +1065,7 @@ TEST_CASE("regression tests 1")
}; };
json _; json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
// related test case: double-precision // related test case: double-precision
std::vector<uint8_t> const vec2 std::vector<uint8_t> const vec2
@@ -1077,7 +1077,7 @@ TEST_CASE("regression tests 1")
0x96, 0x96, 0xb4, 0xb4, 0xfa, 0x94, 0x94, 0x61, 0x96, 0x96, 0xb4, 0xb4, 0xfa, 0x94, 0x94, 0x61,
0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0xfb 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0xfb
}; };
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&); CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
} }
SECTION("issue #452 - Heap-buffer-overflow (OSS-Fuzz issue 585)") SECTION("issue #452 - Heap-buffer-overflow (OSS-Fuzz issue 585)")