Compare commits

..
Author SHA1 Message Date
Niels Lohmann b98aef8a07 Round-trip BJData ND-array annotations exactly (single precision, key order)
to_bjdata() encoded a JData-annotated object as a BJData ND-array in two
cases where from_bjdata() then returned a different value, breaking the
documented round-trip guarantee:

1. A "single" element that is finite and in range but not exactly
   representable as float (e.g. 0.1) or that underflows to 0 (e.g. 1e-300)
   was silently narrowed instead of falling back to a plain object, unlike
   out-of-range integer elements. write_bjdata_ndarray() now only accepts a
   "single" element if it survives the narrowing to float and back, the
   same criterion write_compact_float() already uses for CBOR/MessagePack.

2. from_bjdata() emitted the annotation keys as _ArraySize_, _ArrayType_,
   _ArrayData_ instead of the documented _ArrayType_, _ArraySize_,
   _ArrayData_, because the size key is written while the dimension vector
   is read, before the type key. For ordered_json, whose comparison takes
   key order into account, this made a round trip of the documented example
   compare unequal. The element type marker is known before the dimension
   vector is read (it precedes '#'), so it is now passed down and the
   "_ArrayType_" key is emitted first.

Fixes #5661.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 23:52:49 +02:00
12 changed files with 278 additions and 495 deletions
@@ -55,10 +55,6 @@ This implementation does exactly follow this approach, as it uses double precisi
smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally
and be serialized to `null`. and be serialized to `null`.
During deserialization (from JSON text or any of the binary formats), a finite number that does not fit into
`number_float_t` is rejected with [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406), for
example a double-precision number in a binary format when `number_float_t` is `#!cpp float`.
#### Storage #### Storage
Floating-point number values are stored directly inside a `basic_json` type. Floating-point number values are stored directly inside a `basic_json` type.
@@ -47,9 +47,8 @@ With the default values for `NumberIntegerType` (`std::int64_t`), the default va
When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and
the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of
range will yield over/underflow when used in a constructor. During deserialization (from JSON text or any of the binary range will yield over/underflow when used in a constructor. During deserialization, too large or small integer numbers
formats), too large or small integer numbers will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md).
or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states: [RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are > Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
@@ -48,9 +48,8 @@ With the default values for `NumberUnsignedType` (`std::uint64_t`), the default
When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and
the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow
when used in a constructor. During deserialization (from JSON text or any of the binary formats), too large or small when used in a constructor. During deserialization, too large or small integer numbers will automatically be stored
integer numbers will automatically be stored as [`number_integer_t`](number_integer_t.md) or as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md).
[`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states: [RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are > Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
@@ -132,8 +132,14 @@ The library uses the following mapping from JSON values types to BJData types ac
parsed back as a regular array, parsed back as a regular array,
- every entry of `"_ArraySize_"` is a positive integer, and their product is representable as a `std::size_t`, - every entry of `"_ArraySize_"` is a positive integer, and their product is representable as a `std::size_t`,
- `"_ArrayData_"` is an array holding exactly that many elements, and - `"_ArrayData_"` is an array holding exactly that many elements, and
- every element of `"_ArrayData_"` is a number of the kind named by `"_ArrayType_"` (a floating-point number for - every element of `"_ArrayData_"` is a number of the kind named by `"_ArrayType_"`: for the integer types, a
`single` and `double`, an integer otherwise). value that fits the named width; for `double`, any value; for `single`, a value that survives narrowing to
`float` and back without change (for instance, `0.1` does not, since it is not exactly representable as
`float`).
An annotated object is always read back with its keys in the order shown above, `"_ArrayType_"`, `"_ArraySize_"`,
`"_ArrayData_"`, regardless of the order the ND-array's header stores them in on the wire. This matters for
`ordered_json`, whose comparison takes key order into account.
The current version of this library does not yet support automatic detection of and conversion from a nested JSON The current version of this library does not yet support automatic detection of and conversion from a nested JSON
array input to a BJData ND-array. array input to a BJData ND-array.
@@ -168,9 +168,9 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Negative integer overflow" !!! warning "Negative integer overflow"
CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the
result to fit into `number_integer_t` (`std::int64_t` by default), the result is stored as `number_float_t`, like result to fit into `number_integer_t` (`std::int64_t` by default), parsing fails with a
a too small integer in JSON text. For example, `-18446744073709551616` (`0x3B` followed by eight `0xFF` bytes) is [`parse_error.112`](../../home/exceptions.md#jsonexceptionparse_error112) exception rather than overflowing
stored as `-1.8446744073709552e+19`. silently.
!!! warning "Object keys" !!! warning "Object keys"
+5 -7
View File
@@ -331,6 +331,9 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1 [json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1
``` ```
``` ```
[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow
```
```
[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5) [json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)
``` ```
@@ -851,18 +854,13 @@ The JSON Patch operations 'remove' and 'add' cannot be applied to the root eleme
### json.exception.out_of_range.406 ### json.exception.out_of_range.406
A parsed number could not be stored without changing it to NaN or INF. For the binary formats, this happens when a A parsed number could not be stored as without changing it to NaN or INF.
finite floating-point number does not fit into [`number_float_t`](../api/basic_json/number_float_t.md), for example a
double-precision number when `number_float_t` is `#!cpp float`.
!!! failure "Example messages" !!! failure "Example message"
``` ```
number overflow parsing '10E1000' number overflow parsing '10E1000'
``` ```
```
[json.exception.out_of_range.406] syntax error while parsing CBOR value: number overflow
```
### json.exception.out_of_range.407 ### json.exception.out_of_range.407
+87 -154
View File
@@ -559,7 +559,7 @@ class binary_reader
case 0x01: // double case 0x01: // double
{ {
double number{}; double number{};
return get_number<double, true>(input_format_t::bson, number) && emit_float(input_format_t::bson, number); return get_number<double, true>(input_format_t::bson, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0x02: // string case 0x02: // string
@@ -600,19 +600,19 @@ class binary_reader
case 0x10: // int32 case 0x10: // int32
{ {
std::int32_t value{}; std::int32_t value{};
return get_number<std::int32_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value); return get_number<std::int32_t, true>(input_format_t::bson, value) && sax->number_integer(value);
} }
case 0x12: // int64 case 0x12: // int64
{ {
std::int64_t value{}; std::int64_t value{};
return get_number<std::int64_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value); return get_number<std::int64_t, true>(input_format_t::bson, value) && sax->number_integer(value);
} }
case 0x11: // uint64 case 0x11: // uint64
{ {
std::uint64_t value{}; std::uint64_t value{};
return get_number<std::uint64_t, true>(input_format_t::bson, value) && emit_unsigned(input_format_t::bson, value); return get_number<std::uint64_t, true>(input_format_t::bson, value) && sax->number_unsigned(value);
} }
default: // anything else is not supported (yet) default: // anything else is not supported (yet)
@@ -638,19 +638,14 @@ class binary_reader
{ {
return false; return false;
} }
const auto max_val = static_cast<NumberType>((std::numeric_limits<number_integer_t>::max)());
// the value is -1 - number, which fits into number_integer_t if (number > max_val)
// whenever number does
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{ {
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number)); return sax->parse_error(chars_read, get_token_string(),
parse_error::create(112, chars_read,
exception_message(input_format_t::cbor, "negative integer overflow", "value"), nullptr));
} }
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
// like the lexer does for JSON text, store a value too small for
// number_integer_t as number_float_t; compute it as long double so
// that emit_float sees a finite value and can detect an overflow of
// number_float_t
return emit_float(input_format_t::cbor, static_cast<long double>(-1) - static_cast<long double>(number));
} }
/*! /*!
@@ -707,25 +702,25 @@ class binary_reader
case 0x18: // Unsigned integer (one-byte uint8_t follows) case 0x18: // Unsigned integer (one-byte uint8_t follows)
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
case 0x19: // Unsigned integer (two-byte uint16_t follows) case 0x19: // Unsigned integer (two-byte uint16_t follows)
{ {
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
case 0x1A: // Unsigned integer (four-byte uint32_t follows) case 0x1A: // Unsigned integer (four-byte uint32_t follows)
{ {
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
case 0x1B: // Unsigned integer (eight-byte uint64_t follows) case 0x1B: // Unsigned integer (eight-byte uint64_t follows)
{ {
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
// Negative integer -1-0x00..-1-0x17 (-1..-24) // Negative integer -1-0x00..-1-0x17 (-1..-24)
@@ -1170,13 +1165,13 @@ class binary_reader
case 0xFA: // Single-Precision Float (four-byte IEEE 754) case 0xFA: // Single-Precision Float (four-byte IEEE 754)
{ {
float number{}; float number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xFB: // Double-Precision Float (eight-byte IEEE 754) case 0xFB: // Double-Precision Float (eight-byte IEEE 754)
{ {
double number{}; double number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
default: // anything else (0xFF is handled inside the other types) default: // anything else (0xFF is handled inside the other types)
@@ -1940,61 +1935,61 @@ class binary_reader
case 0xCA: // float 32 case 0xCA: // float 32
{ {
float number{}; float number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xCB: // float 64 case 0xCB: // float 64
{ {
double number{}; double number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xCC: // uint 8 case 0xCC: // uint 8
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xCD: // uint 16 case 0xCD: // uint 16
{ {
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xCE: // uint 32 case 0xCE: // uint 32
{ {
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xCF: // uint 64 case 0xCF: // uint 64
{ {
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xD0: // int 8 case 0xD0: // int 8
{ {
std::int8_t number{}; std::int8_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xD1: // int 16 case 0xD1: // int 16
{ {
std::int16_t number{}; std::int16_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xD2: // int 32 case 0xD2: // int 32
{ {
std::int32_t number{}; std::int32_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xD3: // int 64 case 0xD3: // int 64
{ {
std::int64_t number{}; std::int64_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xDC: // array 16 case 0xDC: // array 16
@@ -2733,10 +2728,15 @@ class binary_reader
is_ndarray can only return `true` when its initial value is_ndarray can only return `true` when its initial value
is `false` is `false`
@param[in] prefix type marker if already read, otherwise set to 0 @param[in] prefix type marker if already read, otherwise set to 0
@param[in] ndarray_dtype the element type marker of the enclosing bjdata ndarray if
already known (it precedes the dimension vector read here),
otherwise 0; used to emit the "_ArrayType_" annotation key
before "_ArraySize_" if a dimension vector turns out to
describe an ndarray
@return whether size determination completed @return whether size determination completed
*/ */
bool get_ubjson_size_value(std::size_t& result, bool& is_ndarray, char_int_type prefix = 0) bool get_ubjson_size_value(std::size_t& result, bool& is_ndarray, char_int_type prefix = 0, char_int_type ndarray_dtype = 0)
{ {
if (prefix == 0) if (prefix == 0)
{ {
@@ -2906,8 +2906,37 @@ class binary_reader
} }
} }
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(3)))
{
return false;
}
// the element type precedes the dimension vector (see get_ubjson_size_type)
// and is passed down as ndarray_dtype; emit it here so the annotation keys
// follow the documented _ArrayType_, _ArraySize_, _ArrayData_ order
if (ndarray_dtype != 0)
{
auto it = std::lower_bound(bjd_types_map.begin(), bjd_types_map.end(), ndarray_dtype, [](const bjd_type & p, char_int_type t)
{
return p.first < t;
});
if (JSON_HEDLEY_UNLIKELY(it == bjd_types_map.end() || it->first != ndarray_dtype))
{
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(112, chars_read,
exception_message(input_format, "invalid byte: 0x" + last_token, "type"), nullptr));
}
string_t type_key = "_ArrayType_";
string_t type = it->second; // sax->string() takes a reference
if (JSON_HEDLEY_UNLIKELY(!sax->key(type_key) || !sax->string(type)))
{
return false;
}
}
string_t key = "_ArraySize_"; string_t key = "_ArraySize_";
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(3) || !sax->key(key) || !sax->start_array(dim.size()))) if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->start_array(dim.size())))
{ {
return false; return false;
} }
@@ -2927,7 +2956,7 @@ class binary_reader
{ {
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr)); return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr));
} }
if (JSON_HEDLEY_UNLIKELY(!emit_unsigned(input_format, i))) if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(static_cast<number_unsigned_t>(i))))
{ {
return false; return false;
} }
@@ -3008,7 +3037,7 @@ class binary_reader
exception_message(input_format, concat("expected '#' after type information; last byte: 0x", last_token), "size"), nullptr)); exception_message(input_format, concat("expected '#' after type information; last byte: 0x", last_token), "size"), nullptr));
} }
const bool is_error = get_ubjson_size_value(result.first, is_ndarray); const bool is_error = get_ubjson_size_value(result.first, is_ndarray, 0, result.second);
// an ndarray was read here only if the flag flipped; when it was // an ndarray was read here only if the flag flipped; when it was
// seeded true, get_ubjson_size_value() already rejected the nested // seeded true, get_ubjson_size_value() already rejected the nested
// dimension vector // dimension vector
@@ -3059,37 +3088,37 @@ class binary_reader
break; break;
} }
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'U': case 'U':
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'i': case 'i':
{ {
std::int8_t number{}; std::int8_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'I': case 'I':
{ {
std::int16_t number{}; std::int16_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'l': case 'l':
{ {
std::int32_t number{}; std::int32_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'L': case 'L':
{ {
std::int64_t number{}; std::int64_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'u': case 'u':
@@ -3099,7 +3128,7 @@ class binary_reader
break; break;
} }
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'm': case 'm':
@@ -3109,7 +3138,7 @@ class binary_reader
break; break;
} }
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'M': case 'M':
@@ -3119,7 +3148,7 @@ class binary_reader
break; break;
} }
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'h': case 'h':
@@ -3177,13 +3206,13 @@ class binary_reader
case 'd': case 'd':
{ {
float number{}; float number{};
return get_number(input_format, number) && emit_float(input_format, number); return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 'D': case 'D':
{ {
double number{}; double number{};
return get_number(input_format, number) && emit_float(input_format, number); return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 'H': case 'H':
@@ -3244,30 +3273,17 @@ class binary_reader
if (input_format == input_format_t::bjdata && size_and_type.first != npos && (size_and_type.second & (1 << 8)) != 0) if (input_format == input_format_t::bjdata && size_and_type.first != npos && (size_and_type.second & (1 << 8)) != 0)
{ {
size_and_type.second &= ~(static_cast<char_int_type>(1) << 8); // use bit 8 to indicate ndarray, here we remove the bit to restore the type marker size_and_type.second &= ~(static_cast<char_int_type>(1) << 8); // use bit 8 to indicate ndarray, here we remove the bit to restore the type marker
auto it = std::lower_bound(bjd_types_map.begin(), bjd_types_map.end(), size_and_type.second, [](const bjd_type & p, char_int_type t)
{
return p.first < t;
});
string_t key = "_ArrayType_";
if (JSON_HEDLEY_UNLIKELY(it == bjd_types_map.end() || it->first != size_and_type.second))
{
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(112, chars_read,
exception_message(input_format, "invalid byte: 0x" + last_token, "type"), nullptr));
}
string_t type = it->second; // sax->string() takes a reference
if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->string(type)))
{
return false;
}
// the "_ArrayType_" and "_ArraySize_" annotation keys were already emitted by
// get_ubjson_size_value() (the type marker is known before the dimension vector
// that determines size_and_type.first is read, so it is emitted first there to
// match the documented _ArrayType_, _ArraySize_, _ArrayData_ key order)
if (size_and_type.second == 'C' || size_and_type.second == 'B') if (size_and_type.second == 'C' || size_and_type.second == 'B')
{ {
size_and_type.second = 'U'; size_and_type.second = 'U';
} }
key = "_ArrayData_"; string_t key = "_ArrayData_";
if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->start_array(size_and_type.first) )) if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->start_array(size_and_type.first) ))
{ {
return false; return false;
@@ -3650,13 +3666,13 @@ class binary_reader
case 0x8E: // binary32 case 0x8E: // binary32
{ {
float number{}; float number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number); return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0x8F: // binary64 case 0x8F: // binary64
{ {
double number{}; double number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number); return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xF8: case 0xF8:
@@ -3722,9 +3738,7 @@ class binary_reader
@brief pass an integer to the SAX parser @brief pass an integer to the SAX parser
Non-negative integers are passed as unsigned, negative integers as signed Non-negative integers are passed as unsigned, negative integers as signed
numbers, like the other binary formats do. A value that does not fit the numbers, like the other binary formats do.
number type is passed as described for @ref emit_unsigned and
@ref emit_signed.
@param[in] number the integer @param[in] number the integer
@return whether the SAX parser accepted the value @return whether the SAX parser accepted the value
@@ -3733,9 +3747,9 @@ class binary_reader
{ {
if (number >= 0) if (number >= 0)
{ {
return emit_unsigned(input_format_t::bon8, static_cast<std::uint64_t>(number)); return sax->number_unsigned(static_cast<number_unsigned_t>(number));
} }
return emit_signed(input_format_t::bon8, number); return sax->number_integer(static_cast<number_integer_t>(number));
} }
/*! /*!
@@ -3792,7 +3806,8 @@ class binary_reader
value = (value << 8) | static_cast<std::int64_t>(current); value = (value << 8) | static_cast<std::int64_t>(current);
} }
return emit_bon8_integer(negative ? -(value + offset) : value + offset); return negative ? sax->number_integer(static_cast<number_integer_t>(-(value + offset)))
: sax->number_unsigned(static_cast<number_unsigned_t>(value + offset));
} }
/*! /*!
@@ -4089,88 +4104,6 @@ class binary_reader
return true; return true;
} }
/*!
@brief pass a signed integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_integer_t is passed as number_unsigned_t if it is non-negative and
fits there, and as number_float_t otherwise. With the default number
types, every integer the binary formats can encode fits, so this only
matters for narrower custom number types.
@tparam NumberType a signed integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_signed(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->number_integer(static_cast<number_integer_t>(number));
}
if (value_in_range_of<number_unsigned_t>(number))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass an unsigned integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_unsigned_t is passed as number_float_t.
@tparam NumberType an unsigned integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_unsigned(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_unsigned_t>(number)))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass a floating-point number read from the input to the SAX parser
Like the lexer does for JSON text, a finite value that overflows
number_float_t is rejected instead of silently becoming infinity. Infinity
and NaN in the input are passed on unchanged. Integers only overflow if
number_float_t cannot represent 2^64, e.g., a half-precision type.
@tparam NumberType a floating-point or integer type
@param[in] format the current format (for diagnostics)
@param[in] number the number
@return whether the SAX parser accepted the value
@throw out_of_range.406 if a finite @a number overflows number_float_t
*/
template<typename NumberType>
bool emit_float(const input_format_t format, const NumberType number)
{
const auto result = static_cast<number_float_t>(number);
if (JSON_HEDLEY_UNLIKELY(std::isfinite(number) && !std::isfinite(result)))
{
return sax->parse_error(chars_read, get_token_string(),
out_of_range::create(406, exception_message(format, "number overflow", "value"), nullptr));
}
return sax->number_float(result, "");
}
/*! /*!
@brief create a string by reading characters from the input @brief create a string by reading characters from the input
@@ -1991,9 +1991,21 @@ class binary_writer
case 'd': case 'd':
{ {
const auto dval = el.template get<double>(); const auto dval = el.template get<double>();
in_range = !std::isfinite(dval) || #ifdef __GNUC__
JSON_HEDLEY_DIAGNOSTIC_PUSH
JSON_HEDLEY_PRAGMA(GCC diagnostic ignored "-Wfloat-equal")
#endif
// a value that would be rounded (rather than exactly represented) by the
// narrowing to float is treated like an out-of-range integer element above;
// this is the same criterion write_compact_float() uses for CBOR/MessagePack
in_range = std::isnan(dval) ||
(dval >= static_cast<double>(std::numeric_limits<float>::lowest()) && (dval >= static_cast<double>(std::numeric_limits<float>::lowest()) &&
dval <= static_cast<double>((std::numeric_limits<float>::max)())); dval <= static_cast<double>((std::numeric_limits<float>::max)()) &&
static_cast<double>(static_cast<float>(dval)) == dval) ||
std::isinf(dval);
#ifdef __GNUC__
JSON_HEDLEY_DIAGNOSTIC_POP
#endif
break; break;
} }
default: default:
+101 -156
View File
@@ -13326,7 +13326,7 @@ class binary_reader
case 0x01: // double case 0x01: // double
{ {
double number{}; double number{};
return get_number<double, true>(input_format_t::bson, number) && emit_float(input_format_t::bson, number); return get_number<double, true>(input_format_t::bson, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0x02: // string case 0x02: // string
@@ -13367,19 +13367,19 @@ class binary_reader
case 0x10: // int32 case 0x10: // int32
{ {
std::int32_t value{}; std::int32_t value{};
return get_number<std::int32_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value); return get_number<std::int32_t, true>(input_format_t::bson, value) && sax->number_integer(value);
} }
case 0x12: // int64 case 0x12: // int64
{ {
std::int64_t value{}; std::int64_t value{};
return get_number<std::int64_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value); return get_number<std::int64_t, true>(input_format_t::bson, value) && sax->number_integer(value);
} }
case 0x11: // uint64 case 0x11: // uint64
{ {
std::uint64_t value{}; std::uint64_t value{};
return get_number<std::uint64_t, true>(input_format_t::bson, value) && emit_unsigned(input_format_t::bson, value); return get_number<std::uint64_t, true>(input_format_t::bson, value) && sax->number_unsigned(value);
} }
default: // anything else is not supported (yet) default: // anything else is not supported (yet)
@@ -13405,19 +13405,14 @@ class binary_reader
{ {
return false; return false;
} }
const auto max_val = static_cast<NumberType>((std::numeric_limits<number_integer_t>::max)());
// the value is -1 - number, which fits into number_integer_t if (number > max_val)
// whenever number does
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{ {
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number)); return sax->parse_error(chars_read, get_token_string(),
parse_error::create(112, chars_read,
exception_message(input_format_t::cbor, "negative integer overflow", "value"), nullptr));
} }
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
// like the lexer does for JSON text, store a value too small for
// number_integer_t as number_float_t; compute it as long double so
// that emit_float sees a finite value and can detect an overflow of
// number_float_t
return emit_float(input_format_t::cbor, static_cast<long double>(-1) - static_cast<long double>(number));
} }
/*! /*!
@@ -13474,25 +13469,25 @@ class binary_reader
case 0x18: // Unsigned integer (one-byte uint8_t follows) case 0x18: // Unsigned integer (one-byte uint8_t follows)
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
case 0x19: // Unsigned integer (two-byte uint16_t follows) case 0x19: // Unsigned integer (two-byte uint16_t follows)
{ {
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
case 0x1A: // Unsigned integer (four-byte uint32_t follows) case 0x1A: // Unsigned integer (four-byte uint32_t follows)
{ {
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
case 0x1B: // Unsigned integer (eight-byte uint64_t follows) case 0x1B: // Unsigned integer (eight-byte uint64_t follows)
{ {
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
} }
// Negative integer -1-0x00..-1-0x17 (-1..-24) // Negative integer -1-0x00..-1-0x17 (-1..-24)
@@ -13937,13 +13932,13 @@ class binary_reader
case 0xFA: // Single-Precision Float (four-byte IEEE 754) case 0xFA: // Single-Precision Float (four-byte IEEE 754)
{ {
float number{}; float number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xFB: // Double-Precision Float (eight-byte IEEE 754) case 0xFB: // Double-Precision Float (eight-byte IEEE 754)
{ {
double number{}; double number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number); return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
default: // anything else (0xFF is handled inside the other types) default: // anything else (0xFF is handled inside the other types)
@@ -14707,61 +14702,61 @@ class binary_reader
case 0xCA: // float 32 case 0xCA: // float 32
{ {
float number{}; float number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xCB: // float 64 case 0xCB: // float 64
{ {
double number{}; double number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xCC: // uint 8 case 0xCC: // uint 8
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xCD: // uint 16 case 0xCD: // uint 16
{ {
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xCE: // uint 32 case 0xCE: // uint 32
{ {
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xCF: // uint 64 case 0xCF: // uint 64
{ {
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
} }
case 0xD0: // int 8 case 0xD0: // int 8
{ {
std::int8_t number{}; std::int8_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xD1: // int 16 case 0xD1: // int 16
{ {
std::int16_t number{}; std::int16_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xD2: // int 32 case 0xD2: // int 32
{ {
std::int32_t number{}; std::int32_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xD3: // int 64 case 0xD3: // int 64
{ {
std::int64_t number{}; std::int64_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number); return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
} }
case 0xDC: // array 16 case 0xDC: // array 16
@@ -15500,10 +15495,15 @@ class binary_reader
is_ndarray can only return `true` when its initial value is_ndarray can only return `true` when its initial value
is `false` is `false`
@param[in] prefix type marker if already read, otherwise set to 0 @param[in] prefix type marker if already read, otherwise set to 0
@param[in] ndarray_dtype the element type marker of the enclosing bjdata ndarray if
already known (it precedes the dimension vector read here),
otherwise 0; used to emit the "_ArrayType_" annotation key
before "_ArraySize_" if a dimension vector turns out to
describe an ndarray
@return whether size determination completed @return whether size determination completed
*/ */
bool get_ubjson_size_value(std::size_t& result, bool& is_ndarray, char_int_type prefix = 0) bool get_ubjson_size_value(std::size_t& result, bool& is_ndarray, char_int_type prefix = 0, char_int_type ndarray_dtype = 0)
{ {
if (prefix == 0) if (prefix == 0)
{ {
@@ -15673,8 +15673,37 @@ class binary_reader
} }
} }
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(3)))
{
return false;
}
// the element type precedes the dimension vector (see get_ubjson_size_type)
// and is passed down as ndarray_dtype; emit it here so the annotation keys
// follow the documented _ArrayType_, _ArraySize_, _ArrayData_ order
if (ndarray_dtype != 0)
{
auto it = std::lower_bound(bjd_types_map.begin(), bjd_types_map.end(), ndarray_dtype, [](const bjd_type & p, char_int_type t)
{
return p.first < t;
});
if (JSON_HEDLEY_UNLIKELY(it == bjd_types_map.end() || it->first != ndarray_dtype))
{
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(112, chars_read,
exception_message(input_format, "invalid byte: 0x" + last_token, "type"), nullptr));
}
string_t type_key = "_ArrayType_";
string_t type = it->second; // sax->string() takes a reference
if (JSON_HEDLEY_UNLIKELY(!sax->key(type_key) || !sax->string(type)))
{
return false;
}
}
string_t key = "_ArraySize_"; string_t key = "_ArraySize_";
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(3) || !sax->key(key) || !sax->start_array(dim.size()))) if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->start_array(dim.size())))
{ {
return false; return false;
} }
@@ -15694,7 +15723,7 @@ class binary_reader
{ {
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr)); return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr));
} }
if (JSON_HEDLEY_UNLIKELY(!emit_unsigned(input_format, i))) if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(static_cast<number_unsigned_t>(i))))
{ {
return false; return false;
} }
@@ -15775,7 +15804,7 @@ class binary_reader
exception_message(input_format, concat("expected '#' after type information; last byte: 0x", last_token), "size"), nullptr)); exception_message(input_format, concat("expected '#' after type information; last byte: 0x", last_token), "size"), nullptr));
} }
const bool is_error = get_ubjson_size_value(result.first, is_ndarray); const bool is_error = get_ubjson_size_value(result.first, is_ndarray, 0, result.second);
// an ndarray was read here only if the flag flipped; when it was // an ndarray was read here only if the flag flipped; when it was
// seeded true, get_ubjson_size_value() already rejected the nested // seeded true, get_ubjson_size_value() already rejected the nested
// dimension vector // dimension vector
@@ -15826,37 +15855,37 @@ class binary_reader
break; break;
} }
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'U': case 'U':
{ {
std::uint8_t number{}; std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'i': case 'i':
{ {
std::int8_t number{}; std::int8_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'I': case 'I':
{ {
std::int16_t number{}; std::int16_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'l': case 'l':
{ {
std::int32_t number{}; std::int32_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'L': case 'L':
{ {
std::int64_t number{}; std::int64_t number{};
return get_number(input_format, number) && emit_signed(input_format, number); return get_number(input_format, number) && sax->number_integer(number);
} }
case 'u': case 'u':
@@ -15866,7 +15895,7 @@ class binary_reader
break; break;
} }
std::uint16_t number{}; std::uint16_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'm': case 'm':
@@ -15876,7 +15905,7 @@ class binary_reader
break; break;
} }
std::uint32_t number{}; std::uint32_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'M': case 'M':
@@ -15886,7 +15915,7 @@ class binary_reader
break; break;
} }
std::uint64_t number{}; std::uint64_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number); return get_number(input_format, number) && sax->number_unsigned(number);
} }
case 'h': case 'h':
@@ -15944,13 +15973,13 @@ class binary_reader
case 'd': case 'd':
{ {
float number{}; float number{};
return get_number(input_format, number) && emit_float(input_format, number); return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 'D': case 'D':
{ {
double number{}; double number{};
return get_number(input_format, number) && emit_float(input_format, number); return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 'H': case 'H':
@@ -16011,30 +16040,17 @@ class binary_reader
if (input_format == input_format_t::bjdata && size_and_type.first != npos && (size_and_type.second & (1 << 8)) != 0) if (input_format == input_format_t::bjdata && size_and_type.first != npos && (size_and_type.second & (1 << 8)) != 0)
{ {
size_and_type.second &= ~(static_cast<char_int_type>(1) << 8); // use bit 8 to indicate ndarray, here we remove the bit to restore the type marker size_and_type.second &= ~(static_cast<char_int_type>(1) << 8); // use bit 8 to indicate ndarray, here we remove the bit to restore the type marker
auto it = std::lower_bound(bjd_types_map.begin(), bjd_types_map.end(), size_and_type.second, [](const bjd_type & p, char_int_type t)
{
return p.first < t;
});
string_t key = "_ArrayType_";
if (JSON_HEDLEY_UNLIKELY(it == bjd_types_map.end() || it->first != size_and_type.second))
{
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(112, chars_read,
exception_message(input_format, "invalid byte: 0x" + last_token, "type"), nullptr));
}
string_t type = it->second; // sax->string() takes a reference
if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->string(type)))
{
return false;
}
// the "_ArrayType_" and "_ArraySize_" annotation keys were already emitted by
// get_ubjson_size_value() (the type marker is known before the dimension vector
// that determines size_and_type.first is read, so it is emitted first there to
// match the documented _ArrayType_, _ArraySize_, _ArrayData_ key order)
if (size_and_type.second == 'C' || size_and_type.second == 'B') if (size_and_type.second == 'C' || size_and_type.second == 'B')
{ {
size_and_type.second = 'U'; size_and_type.second = 'U';
} }
key = "_ArrayData_"; string_t key = "_ArrayData_";
if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->start_array(size_and_type.first) )) if (JSON_HEDLEY_UNLIKELY(!sax->key(key) || !sax->start_array(size_and_type.first) ))
{ {
return false; return false;
@@ -16417,13 +16433,13 @@ class binary_reader
case 0x8E: // binary32 case 0x8E: // binary32
{ {
float number{}; float number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number); return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0x8F: // binary64 case 0x8F: // binary64
{ {
double number{}; double number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number); return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
} }
case 0xF8: case 0xF8:
@@ -16489,9 +16505,7 @@ class binary_reader
@brief pass an integer to the SAX parser @brief pass an integer to the SAX parser
Non-negative integers are passed as unsigned, negative integers as signed Non-negative integers are passed as unsigned, negative integers as signed
numbers, like the other binary formats do. A value that does not fit the numbers, like the other binary formats do.
number type is passed as described for @ref emit_unsigned and
@ref emit_signed.
@param[in] number the integer @param[in] number the integer
@return whether the SAX parser accepted the value @return whether the SAX parser accepted the value
@@ -16500,9 +16514,9 @@ class binary_reader
{ {
if (number >= 0) if (number >= 0)
{ {
return emit_unsigned(input_format_t::bon8, static_cast<std::uint64_t>(number)); return sax->number_unsigned(static_cast<number_unsigned_t>(number));
} }
return emit_signed(input_format_t::bon8, number); return sax->number_integer(static_cast<number_integer_t>(number));
} }
/*! /*!
@@ -16559,7 +16573,8 @@ class binary_reader
value = (value << 8) | static_cast<std::int64_t>(current); value = (value << 8) | static_cast<std::int64_t>(current);
} }
return emit_bon8_integer(negative ? -(value + offset) : value + offset); return negative ? sax->number_integer(static_cast<number_integer_t>(-(value + offset)))
: sax->number_unsigned(static_cast<number_unsigned_t>(value + offset));
} }
/*! /*!
@@ -16856,88 +16871,6 @@ class binary_reader
return true; return true;
} }
/*!
@brief pass a signed integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_integer_t is passed as number_unsigned_t if it is non-negative and
fits there, and as number_float_t otherwise. With the default number
types, every integer the binary formats can encode fits, so this only
matters for narrower custom number types.
@tparam NumberType a signed integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_signed(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->number_integer(static_cast<number_integer_t>(number));
}
if (value_in_range_of<number_unsigned_t>(number))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass an unsigned integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_unsigned_t is passed as number_float_t.
@tparam NumberType an unsigned integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_unsigned(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_unsigned_t>(number)))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass a floating-point number read from the input to the SAX parser
Like the lexer does for JSON text, a finite value that overflows
number_float_t is rejected instead of silently becoming infinity. Infinity
and NaN in the input are passed on unchanged. Integers only overflow if
number_float_t cannot represent 2^64, e.g., a half-precision type.
@tparam NumberType a floating-point or integer type
@param[in] format the current format (for diagnostics)
@param[in] number the number
@return whether the SAX parser accepted the value
@throw out_of_range.406 if a finite @a number overflows number_float_t
*/
template<typename NumberType>
bool emit_float(const input_format_t format, const NumberType number)
{
const auto result = static_cast<number_float_t>(number);
if (JSON_HEDLEY_UNLIKELY(std::isfinite(number) && !std::isfinite(result)))
{
return sax->parse_error(chars_read, get_token_string(),
out_of_range::create(406, exception_message(format, "number overflow", "value"), nullptr));
}
return sax->number_float(result, "");
}
/*! /*!
@brief create a string by reading characters from the input @brief create a string by reading characters from the input
@@ -22409,9 +22342,21 @@ class binary_writer
case 'd': case 'd':
{ {
const auto dval = el.template get<double>(); const auto dval = el.template get<double>();
in_range = !std::isfinite(dval) || #ifdef __GNUC__
JSON_HEDLEY_DIAGNOSTIC_PUSH
JSON_HEDLEY_PRAGMA(GCC diagnostic ignored "-Wfloat-equal")
#endif
// a value that would be rounded (rather than exactly represented) by the
// narrowing to float is treated like an out-of-range integer element above;
// this is the same criterion write_compact_float() uses for CBOR/MessagePack
in_range = std::isnan(dval) ||
(dval >= static_cast<double>(std::numeric_limits<float>::lowest()) && (dval >= static_cast<double>(std::numeric_limits<float>::lowest()) &&
dval <= static_cast<double>((std::numeric_limits<float>::max)())); dval <= static_cast<double>((std::numeric_limits<float>::max)()) &&
static_cast<double>(static_cast<float>(dval)) == dval) ||
std::isinf(dval);
#ifdef __GNUC__
JSON_HEDLEY_DIAGNOSTIC_POP
#endif
break; break;
} }
default: default:
-141
View File
@@ -11,12 +11,7 @@
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
#include <cmath>
#include <fstream> #include <fstream>
#include <limits>
#include <map>
#include <string>
#include <vector>
#include "make_test_data_available.hpp" #include "make_test_data_available.hpp"
TEST_CASE("Binary Formats" * doctest::skip()) TEST_CASE("Binary Formats" * doctest::skip())
@@ -229,139 +224,3 @@ TEST_CASE("Binary Formats" * doctest::skip())
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450)); CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450));
} }
} }
namespace
{
// the binary formats as function pointers for "Binary formats with narrow number types";
// named functions rather than lambdas, because clang 3.5 cannot convert a lambda
// to a function pointer in the braced initializer of the format table
using narrow_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int32_t, std::uint32_t, float>;
using bytes = std::vector<std::uint8_t>;
bytes encode_cbor(const json& j)
{
return json::to_cbor(j);
}
narrow_json decode_cbor(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_cbor(v, true, allow_exceptions);
}
bytes encode_msgpack(const json& j)
{
return json::to_msgpack(j);
}
narrow_json decode_msgpack(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_msgpack(v, true, allow_exceptions);
}
bytes encode_ubjson(const json& j)
{
return json::to_ubjson(j);
}
narrow_json decode_ubjson(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_ubjson(v, true, allow_exceptions);
}
bytes encode_bjdata(const json& j)
{
return json::to_bjdata(j);
}
narrow_json decode_bjdata(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_bjdata(v, true, allow_exceptions);
}
// BSON can only store numbers as object members
bytes encode_bson(const json& j)
{
return json::to_bson(json{{"a", j}});
}
narrow_json decode_bson(const bytes& v, bool allow_exceptions)
{
const auto result = narrow_json::from_bson(v, true, allow_exceptions);
return result.is_discarded() ? result : result.at("a");
}
bytes encode_bon8(const json& j)
{
return json::to_bon8(j);
}
narrow_json decode_bon8(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_bon8(v, true, allow_exceptions);
}
} // namespace
TEST_CASE("Binary formats with narrow number types")
{
// Numbers that do not fit the number types are handled like the lexer
// handles them in JSON text: an integer that fits neither integer type is
// stored as a floating-point number, and a finite floating-point number
// that overflows number_float_t is rejected with out_of_range.406.
struct binary_format
{
const char* name;
bytes (*encode)(const json&);
narrow_json (*decode)(const bytes&, bool);
};
const std::vector<binary_format> formats =
{
{"CBOR", encode_cbor, decode_cbor},
{"MessagePack", encode_msgpack, decode_msgpack},
{"UBJSON", encode_ubjson, decode_ubjson},
{"BJData", encode_bjdata, decode_bjdata},
{"BSON", encode_bson, decode_bson},
{"BON8", encode_bon8, decode_bon8},
};
for (const auto& format : formats)
{
const std::string name = format.name;
INFO("format := ", name);
const auto roundtrip = [&format](const json & j)
{
return format.decode(format.encode(j), true);
};
// integers that fit keep their type
CHECK(roundtrip(json(-5)).is_number_integer());
CHECK(roundtrip(json(-5)).get<std::int32_t>() == -5);
CHECK(roundtrip(json(3000000000u)).is_number_unsigned());
CHECK(roundtrip(json(3000000000u)).get<std::uint32_t>() == 3000000000u);
// integers that fit neither integer type are stored as float
CHECK(roundtrip(json(5000000000u)).is_number_float());
CHECK(roundtrip(json(5000000000u)).get<float>() == 5000000000.0f);
if (name != "BON8") // BON8 cannot encode integers above INT64_MAX
{
CHECK(roundtrip(json(10000000000000000000u)).is_number_float());
CHECK(roundtrip(json(10000000000000000000u)).get<float>() == 10000000000000000000.0f);
}
CHECK(roundtrip(json(-3000000000LL)).is_number_float());
CHECK(roundtrip(json(-3000000000LL)).get<float>() == -3000000000.0f);
CHECK(roundtrip(json(-5000000000LL)).is_number_float());
CHECK(roundtrip(json(-5000000000LL)).get<float>() == -5000000000.0f);
// floating-point numbers that fit
CHECK(roundtrip(json(1.5)).get<float>() == 1.5f);
const auto just_above_max = std::nextafter(static_cast<double>((std::numeric_limits<float>::max)()),
std::numeric_limits<double>::infinity());
CHECK(roundtrip(json(just_above_max)).get<float>() == (std::numeric_limits<float>::max)());
// infinity and NaN are passed on
CHECK(std::isinf(roundtrip(json(std::numeric_limits<double>::infinity())).get<float>()));
CHECK(std::isnan(roundtrip(json(std::numeric_limits<double>::quiet_NaN())).get<float>()));
// finite floating-point numbers that overflow number_float_t are rejected
const std::string message = "[json.exception.out_of_range.406] syntax error while parsing " + name
+ " value: number overflow";
CHECK_THROWS_WITH_AS(roundtrip(json(1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK_THROWS_WITH_AS(roundtrip(json(-1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK(format.decode(format.encode(json(1e300)), false).is_discarded());
}
}
+42 -4
View File
@@ -11,6 +11,7 @@
#define JSON_TESTS_PRIVATE #define JSON_TESTS_PRIVATE
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using nlohmann::json; using nlohmann::json;
using ordered_json = nlohmann::ordered_json;
#include <algorithm> #include <algorithm>
#include <climits> #include <climits>
@@ -2294,29 +2295,33 @@ TEST_CASE("BJData")
SECTION("start_array() in ndarray _ArraySize_") SECTION("start_array() in ndarray _ArraySize_")
{ {
// _ArrayType_ (2 events: key + string) is now emitted before
// _ArraySize_ (see GitHub issue #5661), which shifts the events
// below later by the same 2 events
std::vector<uint8_t> const v = {'[', '$', 'i', '#', '[', '$', 'i', '#', 'i', 2, 2, 1, 1, 2}; std::vector<uint8_t> const v = {'[', '$', 'i', '#', '[', '$', 'i', '#', 'i', 2, 2, 1, 1, 2};
SaxCountdown scp(2); SaxCountdown scp(4);
CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata)); CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata));
} }
SECTION("number_integer() in ndarray _ArraySize_") SECTION("number_integer() in ndarray _ArraySize_")
{ {
std::vector<uint8_t> const v = {'[', '$', 'U', '#', '[', '$', 'i', '#', 'i', 2, 2, 1, 1, 2}; std::vector<uint8_t> const v = {'[', '$', 'U', '#', '[', '$', 'i', '#', 'i', 2, 2, 1, 1, 2};
SaxCountdown scp(3); SaxCountdown scp(5);
CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata)); CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata));
} }
SECTION("key() in ndarray _ArrayType_") SECTION("key() in ndarray _ArrayType_")
{ {
// _ArrayType_ is emitted right after start_object(), before _ArraySize_
std::vector<uint8_t> const v = {'[', '$', 'U', '#', '[', '$', 'U', '#', 'i', 2, 2, 2, 1, 2, 3, 4}; std::vector<uint8_t> const v = {'[', '$', 'U', '#', '[', '$', 'U', '#', 'i', 2, 2, 2, 1, 2, 3, 4};
SaxCountdown scp(6); SaxCountdown scp(1);
CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata)); CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata));
} }
SECTION("string() in ndarray _ArrayType_") SECTION("string() in ndarray _ArrayType_")
{ {
std::vector<uint8_t> const v = {'[', '$', 'U', '#', '[', '$', 'U', '#', 'i', 2, 2, 2, 1, 2, 3, 4}; std::vector<uint8_t> const v = {'[', '$', 'U', '#', '[', '$', 'U', '#', 'i', 2, 2, 2, 1, 2, 3, 4};
SaxCountdown scp(7); SaxCountdown scp(2);
CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata)); CHECK_FALSE(json::sax_parse(v, &scp, json::input_format_t::bjdata));
} }
@@ -2919,6 +2924,22 @@ TEST_CASE("BJData")
CHECK(out_single.at(0) == '{'); CHECK(out_single.at(0) == '{');
CHECK(json::from_bjdata(out_single) == j_single); CHECK(json::from_bjdata(out_single) == j_single);
// a double element that is finite and within the range of "single"
// but is not exactly representable as a float, so narrowing it would
// silently round it (0.1 is read back as 0.10000000149011612); this,
// like the overflow case above, falls back to a plain object (see
// GitHub issue #5661)
json const j_single_rounded = json({{"_ArrayType_", "single"}, {"_ArraySize_", {2, 1}}, {"_ArrayData_", {1.5, 0.1}}});
const auto out_single_rounded = json::to_bjdata(j_single_rounded);
CHECK(out_single_rounded.at(0) == '{');
CHECK(json::from_bjdata(out_single_rounded) == j_single_rounded);
// a double element that underflows to 0 when narrowed to "single"
json const j_single_underflow = json({{"_ArrayType_", "single"}, {"_ArraySize_", {2, 1}}, {"_ArrayData_", {1.5, 1e-300}}});
const auto out_single_underflow = json::to_bjdata(j_single_underflow);
CHECK(out_single_underflow.at(0) == '{');
CHECK(json::from_bjdata(out_single_underflow) == j_single_underflow);
// in-range boundary values still use the compact ndarray encoding // in-range boundary values still use the compact ndarray encoding
json const j_uint8_ok = json({{"_ArrayType_", "uint8"}, {"_ArraySize_", {2, 1}}, {"_ArrayData_", {0, 255}}}); json const j_uint8_ok = json({{"_ArrayType_", "uint8"}, {"_ArraySize_", {2, 1}}, {"_ArrayData_", {0, 255}}});
CHECK(json::to_bjdata(j_uint8_ok) == std::vector<uint8_t>({'[', '$', 'U', '#', '[', 'i', 2, 'i', 1, ']', 0, 255})); CHECK(json::to_bjdata(j_uint8_ok) == std::vector<uint8_t>({'[', '$', 'U', '#', '[', 'i', 2, 'i', 1, ']', 0, 255}));
@@ -2932,6 +2953,23 @@ TEST_CASE("BJData")
CHECK(json::from_bjdata(out_single_ok) == json({{"_ArrayType_", "single"}, {"_ArraySize_", {2, 1}}, {"_ArrayData_", {1.5f, -1.5f}}})); CHECK(json::from_bjdata(out_single_ok) == json({{"_ArrayType_", "single"}, {"_ArraySize_", {2, 1}}, {"_ArrayData_", {1.5f, -1.5f}}}));
} }
SECTION("ndarray annotation keys are read back in the documented order")
{
// from_bjdata() must emit the annotation object's keys in the order
// used throughout the documentation, _ArrayType_, _ArraySize_,
// _ArrayData_: the type marker precedes the dimension vector on the
// wire (see get_ubjson_size_type()), so it is known, and emitted,
// before _ArraySize_. For a plain json this key order is invisible
// (its comparison ignores it), but for an ordered_json it is not (see
// GitHub issue #5661).
const ordered_json o = ordered_json::parse(R"({"_ArrayType_":"uint8","_ArraySize_":[2,2],"_ArrayData_":[1,2,3,4]})");
const auto packed = ordered_json::to_bjdata(o);
CHECK(packed.at(0) == '[');
const ordered_json o_back = ordered_json::from_bjdata(packed);
CHECK(o_back == o);
CHECK(o_back.dump() == o.dump());
}
SECTION("ndarray that would not be read back as an annotated object stays as object") SECTION("ndarray that would not be read back as an annotated object stays as object")
{ {
// the reader only restores an annotated object from an ND-array // the reader only restores an annotated object from an ND-array
+14 -16
View File
@@ -3187,8 +3187,7 @@ TEST_CASE("Tagged values")
// CBOR encodes negative integers as: result = -1 - n // CBOR encodes negative integers as: result = -1 - n
// For type 0x3B, n is an 8-byte uint64_t. Valid range for n with // For type 0x3B, n is an 8-byte uint64_t. Valid range for n with
// the default int64_t is [0, INT64_MAX], producing results in [INT64_MIN, -1]. // the default int64_t is [0, INT64_MAX], producing results in [INT64_MIN, -1].
// When n > INT64_MAX, the result exceeds int64_t range and is stored // When n > INT64_MAX, the result exceeds int64_t range and is rejected.
// as a floating-point number, as the lexer does for JSON text.
SECTION("n = 0 is valid (result = -1)") SECTION("n = 0 is valid (result = -1)")
{ {
@@ -3209,34 +3208,33 @@ TEST_CASE("Tagged values")
CHECK(result.get<int64_t>() == (std::numeric_limits<int64_t>::min)()); CHECK(result.get<int64_t>() == (std::numeric_limits<int64_t>::min)());
} }
SECTION("n = INT64_MAX + 1 is stored as float") SECTION("n = INT64_MAX + 1 is rejected (overflow)")
{ {
// n = INT64_MAX + 1 (0x8000000000000000) // n = INT64_MAX + 1 (0x8000000000000000)
// result = -1 - n = -9223372036854775809, which exceeds int64_t range; // result = -1 - n = -9223372036854775809, which exceeds int64_t range
// the nearest double is -9223372036854775808.0
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}; const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
const auto result = json::from_cbor(input); json _;
CHECK(result.is_number_float()); CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
CHECK(result.get<double>() == -9223372036854775808.0); "[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
CHECK(result == json::parse("-9223372036854775809")); json::parse_error);
} }
SECTION("n = UINT64_MAX is stored as float") SECTION("n = UINT64_MAX is rejected (overflow)")
{ {
// n = UINT64_MAX (0xFFFFFFFFFFFFFFFF) // n = UINT64_MAX (0xFFFFFFFFFFFFFFFF)
// result = -1 - n = -18446744073709551616, which exceeds int64_t range // result = -1 - n = -18446744073709551616, which exceeds int64_t range
const std::vector<uint8_t> input = {0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF}; const std::vector<uint8_t> input = {0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
const auto result = json::from_cbor(input); json _;
CHECK(result.is_number_float()); CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
CHECK(result.get<double>() == -18446744073709551616.0); "[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
CHECK(result == json::parse("-18446744073709551616")); json::parse_error);
} }
SECTION("overflow with allow_exceptions=false is not an error") SECTION("overflow with allow_exceptions=false returns discarded")
{ {
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00}; const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
const auto result = json::from_cbor(input, true, false); const auto result = json::from_cbor(input, true, false);
CHECK(result.is_number_float()); CHECK(result.is_discarded());
} }
} }