Compare commits

..
Author SHA1 Message Date
Niels Lohmann 04ed4f593c Handle numbers that do not fit narrow number types in the binary readers
With custom number types narrower than the values in a binary document,
for example basic_json<..., std::int32_t, std::uint32_t, float>, every
binary reader (CBOR, MessagePack, UBJSON, BJData, BSON, BON8) passed the
decoded number to the SAX interface with an implicit conversion: the
integer 5000000000 silently became 705032704, and a finite double such as
1e300 became infinity. The lexer handles the same values in JSON text: an
integer that fits neither integer type is stored as number_float_t, and a
finite number that overflows number_float_t is rejected with
out_of_range.406.

Pass every number read from binary input through three helpers that
apply the lexer's rules:
- emit_signed(): number_integer_t, else number_unsigned_t for a
  non-negative value, else number_float_t
- emit_unsigned(): number_unsigned_t, else number_float_t
- emit_float(): out_of_range.406 if a finite value overflows
  number_float_t; infinity and NaN are passed on

For consistency, a CBOR negative integer below the range of
number_integer_t is now stored as number_float_t, like a too small
integer in JSON text, instead of being rejected with parse_error.112.
With the default number types, this is the only change in behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 23:23:16 +02:00
23 changed files with 411 additions and 401 deletions
+2 -2
View File
@@ -363,7 +363,7 @@ std::cout << j_string << " == " << serialized_string << std::endl;
[`.dump()`](https://json.nlohmann.me/api/basic_json/dump/) returns the originally stored string value.
Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
Note the library only supports UTF-8. When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
#### To/from streams (e.g., files, string streams)
@@ -1915,7 +1915,7 @@ The library supports **Unicode input** as follows:
- [Unicode noncharacters](https://www.unicode.org/faq/private_use.html#nonchar1) will not be replaced by the library.
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
- When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
- When you store strings with different encodings in the library, calling [`dump()`](https://json.nlohmann.me/api/basic_json/dump/) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
- To store wide strings (e.g., `std::wstring`), you need to convert them to a UTF-8 encoded `std::string` before, see [an example](https://json.nlohmann.me/home/faq/#wide-string-handling).
### Comments in JSON
+4 -10
View File
@@ -25,15 +25,10 @@ and `ensure_ascii` parameters.
result consists of ASCII characters only.
`error_handler` (in)
: how to react on decoding errors; there are four possible values (see [`error_handler_t`](error_handler_t.md)):
- `strict`: throw a [`type_error`](../../home/exceptions.md#type-errors) exception in case a decoding error occurs
(default),
- `replace`: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER),
- `ignore`: ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the output unchanged,
and invalid bytes are dropped, and
- `keep`: keep invalid UTF-8 sequences during serialization; all bytes are copied to the output unchanged, so the
result is not valid UTF-8.
: how to react on decoding errors; there are three possible values (see [`error_handler_t`](error_handler_t.md):
`strict` (throws an exception in case a decoding error occurs; default), `replace` (replace invalid UTF-8 sequences
with U+FFFD), and `ignore` (ignore invalid UTF-8 sequences during serialization; all valid bytes are copied to the
output unchanged, and invalid bytes are dropped)).
## Return value
@@ -99,4 +94,3 @@ Binary values are serialized as an object containing two keys:
- Indentation character `indent_char`, option `ensure_ascii` and exceptions added in version 3.0.0.
- Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0.
- Error handler value `keep` added in version 3.13.0.
@@ -4,13 +4,12 @@
enum class error_handler_t {
strict,
replace,
ignore,
keep
ignore
};
```
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
`basic_json` value. Four values are differentiated:
`basic_json` value. Three values are differentiated:
strict
: throw a `type_error` exception in case of invalid UTF-8
@@ -21,12 +20,6 @@ replace
ignore
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
keep
: keep invalid UTF-8 sequences; all bytes are copied to the output unchanged. Valid characters are still escaped as
usual (e.g., `"`, `\\`, and control characters), so the result has valid JSON syntax, but it is not valid UTF-8.
In particular, [`parse`](parse.md) rejects it, and with `ensure_ascii` set to `true`, the invalid bytes are the
only non-ASCII bytes of the output.
## Examples
??? example
@@ -47,4 +40,3 @@ keep
## Version history
- Added in version 3.4.0.
- Added value `keep` in version 3.13.0.
@@ -55,6 +55,10 @@ This implementation does exactly follow this approach, as it uses double precisi
smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally
and be serialized to `null`.
During deserialization (from JSON text or any of the binary formats), a finite number that does not fit into
`number_float_t` is rejected with [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406), for
example a double-precision number in a binary format when `number_float_t` is `#!cpp float`.
#### Storage
Floating-point number values are stored directly inside a `basic_json` type.
@@ -47,8 +47,9 @@ With the default values for `NumberIntegerType` (`std::int64_t`), the default va
When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and
the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of
range will yield over/underflow when used in a constructor. During deserialization, too large or small integer numbers
will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md).
range will yield over/underflow when used in a constructor. During deserialization (from JSON text or any of the binary
formats), too large or small integer numbers will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md)
or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
@@ -48,8 +48,9 @@ With the default values for `NumberUnsignedType` (`std::uint64_t`), the default
When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and
the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow
when used in a constructor. During deserialization, too large or small integer numbers will automatically be stored
as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md).
when used in a constructor. During deserialization (from JSON text or any of the binary formats), too large or small
integer numbers will automatically be stored as [`number_integer_t`](number_integer_t.md) or
[`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
@@ -1,4 +1,3 @@
#include <iomanip>
#include <iostream>
#include <nlohmann/json.hpp>
@@ -22,12 +21,4 @@ int main()
<< "\nstring with ignored invalid characters: "
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
<< '\n';
// the invalid byte is kept; print the result byte-wise to make it visible
std::cout << "string with kept invalid characters:";
for (const unsigned char c : j_invalid.dump(-1, ' ', false, json::error_handler_t::keep))
{
std::cout << ' ' << std::hex << std::setw(2) << std::setfill('0') << static_cast<int>(c);
}
std::cout << '\n';
}
@@ -1,4 +1,3 @@
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
string with replaced invalid characters: "ä�ü"
string with ignored invalid characters: "äü"
string with kept invalid characters: 22 c3 a4 a9 c3 bc 22
@@ -168,9 +168,9 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Negative integer overflow"
CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the
result to fit into `number_integer_t` (`std::int64_t` by default), parsing fails with a
[`parse_error.112`](../../home/exceptions.md#jsonexceptionparse_error112) exception rather than overflowing
silently.
result to fit into `number_integer_t` (`std::int64_t` by default), the result is stored as `number_float_t`, like
a too small integer in JSON text. For example, `-18446744073709551616` (`0x3B` followed by eight `0xFF` bytes) is
stored as `-1.8446744073709552e+19`.
!!! warning "Object keys"
@@ -64,7 +64,6 @@ serialization fails by default. The fourth argument of `dump` selects an
- `strict` (default) — throw a [`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) exception.
- `replace` — replace invalid bytes with the Unicode replacement character U+FFFD (`�`).
- `ignore` — silently drop invalid bytes.
- `keep` — copy invalid bytes to the output unchanged; the result is not valid UTF-8.
??? example
+7 -6
View File
@@ -331,9 +331,6 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1
```
```
[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow
```
```
[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)
```
@@ -755,7 +752,6 @@ The `dump()` function only works with UTF-8 encoded strings; that is, if you ass
- Pass an error handler as last parameter to the `dump()` function to avoid this exception:
- `json::error_handler_t::replace` will replace invalid bytes sequences with `U+FFFD`
- `json::error_handler_t::ignore` will silently ignore invalid byte sequences
- `json::error_handler_t::keep` will copy invalid byte sequences to the output unchanged
### json.exception.type_error.317
@@ -848,13 +844,18 @@ The JSON Patch operations 'remove' and 'add' cannot be applied to the root eleme
### json.exception.out_of_range.406
A parsed number could not be stored as without changing it to NaN or INF.
A parsed number could not be stored without changing it to NaN or INF. For the binary formats, this happens when a
finite floating-point number does not fit into [`number_float_t`](../api/basic_json/number_float_t.md), for example a
double-precision number when `number_float_t` is `#!cpp float`.
!!! failure "Example message"
!!! failure "Example messages"
```
number overflow parsing '10E1000'
```
```
[json.exception.out_of_range.406] syntax error while parsing CBOR value: number overflow
```
### json.exception.out_of_range.407
+1 -1
View File
@@ -85,7 +85,7 @@ The library supports **Unicode input** as follows:
- The library will not replace [Unicode noncharacters](http://www.unicode.org/faq/private_use.html#nonchar1).
- Invalid surrogates (e.g., incomplete pairs such as `\uDEAD`) will yield parse errors.
- The strings stored in the library are UTF-8 encoded. When using the default string type (`std::string`), note that its length/size functions return the number of stored bytes rather than the number of characters or glyphs.
- When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace`, `json::error_handler_t::ignore`, or `json::error_handler_t::keep` are used as error handlers.
- When you store strings with different encodings in the library, calling [`dump()`](https://nlohmann.github.io/json/classnlohmann_1_1basic__json_a50ec80b02d0f3f51130d4abb5d1cfdc5.html#a50ec80b02d0f3f51130d4abb5d1cfdc5) may throw an exception unless `json::error_handler_t::replace` or `json::error_handler_t::ignore` are used as error handlers.
In most cases, the parser is right to complain, because the input is not UTF-8 encoded. This is especially true for Microsoft Windows, where Latin-1 or ISO 8859-1 is often the standard encoding.
+122 -45
View File
@@ -559,7 +559,7 @@ class binary_reader
case 0x01: // double
{
double number{};
return get_number<double, true>(input_format_t::bson, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number<double, true>(input_format_t::bson, number) && emit_float(input_format_t::bson, number);
}
case 0x02: // string
@@ -600,19 +600,19 @@ class binary_reader
case 0x10: // int32
{
std::int32_t value{};
return get_number<std::int32_t, true>(input_format_t::bson, value) && sax->number_integer(value);
return get_number<std::int32_t, true>(input_format_t::bson, value) && emit_signed(value);
}
case 0x12: // int64
{
std::int64_t value{};
return get_number<std::int64_t, true>(input_format_t::bson, value) && sax->number_integer(value);
return get_number<std::int64_t, true>(input_format_t::bson, value) && emit_signed(value);
}
case 0x11: // uint64
{
std::uint64_t value{};
return get_number<std::uint64_t, true>(input_format_t::bson, value) && sax->number_unsigned(value);
return get_number<std::uint64_t, true>(input_format_t::bson, value) && emit_unsigned(value);
}
default: // anything else is not supported (yet)
@@ -638,14 +638,17 @@ class binary_reader
{
return false;
}
const auto max_val = static_cast<NumberType>((std::numeric_limits<number_integer_t>::max)());
if (number > max_val)
// the value is -1 - number, which fits into number_integer_t
// whenever number does
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(112, chars_read,
exception_message(input_format_t::cbor, "negative integer overflow", "value"), nullptr));
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
}
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
// like the lexer does for JSON text, store a value too small for
// number_integer_t as number_float_t
return sax->number_float(static_cast<number_float_t>(-1) - static_cast<number_float_t>(number), "");
}
/*!
@@ -702,25 +705,25 @@ class binary_reader
case 0x18: // Unsigned integer (one-byte uint8_t follows)
{
std::uint8_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
case 0x19: // Unsigned integer (two-byte uint16_t follows)
{
std::uint16_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
case 0x1A: // Unsigned integer (four-byte uint32_t follows)
{
std::uint32_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
case 0x1B: // Unsigned integer (eight-byte uint64_t follows)
{
std::uint64_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
// Negative integer -1-0x00..-1-0x17 (-1..-24)
@@ -1165,13 +1168,13 @@ class binary_reader
case 0xFA: // Single-Precision Float (four-byte IEEE 754)
{
float number{};
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
}
case 0xFB: // Double-Precision Float (eight-byte IEEE 754)
{
double number{};
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
}
default: // anything else (0xFF is handled inside the other types)
@@ -1861,61 +1864,61 @@ class binary_reader
case 0xCA: // float 32
{
float number{};
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
}
case 0xCB: // float 64
{
double number{};
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
}
case 0xCC: // uint 8
{
std::uint8_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xCD: // uint 16
{
std::uint16_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xCE: // uint 32
{
std::uint32_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xCF: // uint 64
{
std::uint64_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xD0: // int 8
{
std::int8_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xD1: // int 16
{
std::int16_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xD2: // int 32
{
std::int32_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xD3: // int 64
{
std::int64_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xDC: // array 16
@@ -2756,7 +2759,7 @@ class binary_reader
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(static_cast<number_unsigned_t>(i))))
if (JSON_HEDLEY_UNLIKELY(!emit_unsigned(i)))
{
return false;
}
@@ -2888,37 +2891,37 @@ class binary_reader
break;
}
std::uint8_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'U':
{
std::uint8_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'i':
{
std::int8_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'I':
{
std::int16_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'l':
{
std::int32_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'L':
{
std::int64_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'u':
@@ -2928,7 +2931,7 @@ class binary_reader
break;
}
std::uint16_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'm':
@@ -2938,7 +2941,7 @@ class binary_reader
break;
}
std::uint32_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'M':
@@ -2948,7 +2951,7 @@ class binary_reader
break;
}
std::uint64_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'h':
@@ -3006,13 +3009,13 @@ class binary_reader
case 'd':
{
float number{};
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format, number) && emit_float(input_format, number);
}
case 'D':
{
double number{};
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format, number) && emit_float(input_format, number);
}
case 'H':
@@ -3479,13 +3482,13 @@ class binary_reader
case 0x8E: // binary32
{
float number{};
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
}
case 0x8F: // binary64
{
double number{};
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
}
case 0xF8:
@@ -3551,7 +3554,9 @@ class binary_reader
@brief pass an integer to the SAX parser
Non-negative integers are passed as unsigned, negative integers as signed
numbers, like the other binary formats do.
numbers, like the other binary formats do. A value that does not fit the
number type is passed as described for @ref emit_unsigned and
@ref emit_signed.
@param[in] number the integer
@return whether the SAX parser accepted the value
@@ -3560,9 +3565,9 @@ class binary_reader
{
if (number >= 0)
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
return emit_unsigned(static_cast<std::uint64_t>(number));
}
return sax->number_integer(static_cast<number_integer_t>(number));
return emit_signed(number);
}
/*!
@@ -3619,8 +3624,7 @@ class binary_reader
value = (value << 8) | static_cast<std::int64_t>(current);
}
return negative ? sax->number_integer(static_cast<number_integer_t>(-(value + offset)))
: sax->number_unsigned(static_cast<number_unsigned_t>(value + offset));
return emit_bon8_integer(negative ? -(value + offset) : value + offset);
}
/*!
@@ -3917,6 +3921,79 @@ class binary_reader
return true;
}
/*!
@brief pass a signed integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_integer_t is passed as number_unsigned_t if it is non-negative and
fits there, and as number_float_t otherwise. With the default number
types, every integer the binary formats can encode fits, so this only
matters for narrower custom number types.
@tparam NumberType a signed integer type
@param[in] number the integer
@return whether the SAX parser accepted the value
*/
template<typename NumberType>
bool emit_signed(const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->number_integer(static_cast<number_integer_t>(number));
}
if (value_in_range_of<number_unsigned_t>(number))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return sax->number_float(static_cast<number_float_t>(number), "");
}
/*!
@brief pass an unsigned integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_unsigned_t is passed as number_float_t.
@tparam NumberType an unsigned integer type
@param[in] number the integer
@return whether the SAX parser accepted the value
*/
template<typename NumberType>
bool emit_unsigned(const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_unsigned_t>(number)))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return sax->number_float(static_cast<number_float_t>(number), "");
}
/*!
@brief pass a floating-point number read from the input to the SAX parser
Like the lexer does for JSON text, a finite value that overflows
number_float_t is rejected instead of silently becoming infinity. Infinity
and NaN in the input are passed on unchanged.
@tparam NumberType a floating-point type
@param[in] format the current format (for diagnostics)
@param[in] number the number
@return whether the SAX parser accepted the value
@throw out_of_range.406 if a finite @a number overflows number_float_t
*/
template<typename NumberType>
bool emit_float(const input_format_t format, const NumberType number)
{
const auto result = static_cast<number_float_t>(number);
if (JSON_HEDLEY_UNLIKELY(std::isfinite(number) && !std::isfinite(result)))
{
return sax->parse_error(chars_read, get_token_string(),
out_of_range::create(406, exception_message(format, "number overflow", "value"), nullptr));
}
return sax->number_float(result, "");
}
/*!
@brief create a string by reading characters from the input
+1 -52
View File
@@ -48,8 +48,7 @@ enum class error_handler_t
{
strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
ignore ///< ignore invalid UTF-8 sequences
};
template<typename BasicJsonType>
@@ -1020,47 +1019,6 @@ class serializer
break;
}
case error_handler_t::keep:
{
// drop whatever the incomplete sequence left in
// the buffer (only copied if !EnsureAscii) and copy
// the ill-formed bytes from the input instead
bytes = bytes_after_last_accept;
if (undumped_chars > 0)
{
// the pending bytes of the incomplete sequence
// are ill-formed; the current byte may be OK for
// itself, so we would like to read it again
for (std::size_t j = i - undumped_chars; j < i; ++j)
{
string_buffer[bytes++] = s[j];
}
--i;
}
else
{
// the current byte cannot start any sequence
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -1106,15 +1064,6 @@ class serializer
break;
}
case error_handler_t::keep:
{
// write all accepted bytes
put_buffer(string_buffer, bytes_after_last_accept);
// copy the bytes of the incomplete sequence unchanged
put_string(s, s.size() - undumped_chars, s.size());
break;
}
case error_handler_t::replace:
{
// write all accepted bytes
+123 -97
View File
@@ -13294,7 +13294,7 @@ class binary_reader
case 0x01: // double
{
double number{};
return get_number<double, true>(input_format_t::bson, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number<double, true>(input_format_t::bson, number) && emit_float(input_format_t::bson, number);
}
case 0x02: // string
@@ -13335,19 +13335,19 @@ class binary_reader
case 0x10: // int32
{
std::int32_t value{};
return get_number<std::int32_t, true>(input_format_t::bson, value) && sax->number_integer(value);
return get_number<std::int32_t, true>(input_format_t::bson, value) && emit_signed(value);
}
case 0x12: // int64
{
std::int64_t value{};
return get_number<std::int64_t, true>(input_format_t::bson, value) && sax->number_integer(value);
return get_number<std::int64_t, true>(input_format_t::bson, value) && emit_signed(value);
}
case 0x11: // uint64
{
std::uint64_t value{};
return get_number<std::uint64_t, true>(input_format_t::bson, value) && sax->number_unsigned(value);
return get_number<std::uint64_t, true>(input_format_t::bson, value) && emit_unsigned(value);
}
default: // anything else is not supported (yet)
@@ -13373,14 +13373,17 @@ class binary_reader
{
return false;
}
const auto max_val = static_cast<NumberType>((std::numeric_limits<number_integer_t>::max)());
if (number > max_val)
// the value is -1 - number, which fits into number_integer_t
// whenever number does
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(112, chars_read,
exception_message(input_format_t::cbor, "negative integer overflow", "value"), nullptr));
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
}
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
// like the lexer does for JSON text, store a value too small for
// number_integer_t as number_float_t
return sax->number_float(static_cast<number_float_t>(-1) - static_cast<number_float_t>(number), "");
}
/*!
@@ -13437,25 +13440,25 @@ class binary_reader
case 0x18: // Unsigned integer (one-byte uint8_t follows)
{
std::uint8_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
case 0x19: // Unsigned integer (two-byte uint16_t follows)
{
std::uint16_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
case 0x1A: // Unsigned integer (four-byte uint32_t follows)
{
std::uint32_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
case 0x1B: // Unsigned integer (eight-byte uint64_t follows)
{
std::uint64_t number{};
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
return get_number(input_format_t::cbor, number) && emit_unsigned(number);
}
// Negative integer -1-0x00..-1-0x17 (-1..-24)
@@ -13900,13 +13903,13 @@ class binary_reader
case 0xFA: // Single-Precision Float (four-byte IEEE 754)
{
float number{};
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
}
case 0xFB: // Double-Precision Float (eight-byte IEEE 754)
{
double number{};
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
}
default: // anything else (0xFF is handled inside the other types)
@@ -14596,61 +14599,61 @@ class binary_reader
case 0xCA: // float 32
{
float number{};
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
}
case 0xCB: // float 64
{
double number{};
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
}
case 0xCC: // uint 8
{
std::uint8_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xCD: // uint 16
{
std::uint16_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xCE: // uint 32
{
std::uint32_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xCF: // uint 64
{
std::uint64_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
return get_number(input_format_t::msgpack, number) && emit_unsigned(number);
}
case 0xD0: // int 8
{
std::int8_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xD1: // int 16
{
std::int16_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xD2: // int 32
{
std::int32_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xD3: // int 64
{
std::int64_t number{};
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
return get_number(input_format_t::msgpack, number) && emit_signed(number);
}
case 0xDC: // array 16
@@ -15491,7 +15494,7 @@ class binary_reader
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(static_cast<number_unsigned_t>(i))))
if (JSON_HEDLEY_UNLIKELY(!emit_unsigned(i)))
{
return false;
}
@@ -15623,37 +15626,37 @@ class binary_reader
break;
}
std::uint8_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'U':
{
std::uint8_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'i':
{
std::int8_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'I':
{
std::int16_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'l':
{
std::int32_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'L':
{
std::int64_t number{};
return get_number(input_format, number) && sax->number_integer(number);
return get_number(input_format, number) && emit_signed(number);
}
case 'u':
@@ -15663,7 +15666,7 @@ class binary_reader
break;
}
std::uint16_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'm':
@@ -15673,7 +15676,7 @@ class binary_reader
break;
}
std::uint32_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'M':
@@ -15683,7 +15686,7 @@ class binary_reader
break;
}
std::uint64_t number{};
return get_number(input_format, number) && sax->number_unsigned(number);
return get_number(input_format, number) && emit_unsigned(number);
}
case 'h':
@@ -15741,13 +15744,13 @@ class binary_reader
case 'd':
{
float number{};
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format, number) && emit_float(input_format, number);
}
case 'D':
{
double number{};
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format, number) && emit_float(input_format, number);
}
case 'H':
@@ -16214,13 +16217,13 @@ class binary_reader
case 0x8E: // binary32
{
float number{};
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
}
case 0x8F: // binary64
{
double number{};
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
}
case 0xF8:
@@ -16286,7 +16289,9 @@ class binary_reader
@brief pass an integer to the SAX parser
Non-negative integers are passed as unsigned, negative integers as signed
numbers, like the other binary formats do.
numbers, like the other binary formats do. A value that does not fit the
number type is passed as described for @ref emit_unsigned and
@ref emit_signed.
@param[in] number the integer
@return whether the SAX parser accepted the value
@@ -16295,9 +16300,9 @@ class binary_reader
{
if (number >= 0)
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
return emit_unsigned(static_cast<std::uint64_t>(number));
}
return sax->number_integer(static_cast<number_integer_t>(number));
return emit_signed(number);
}
/*!
@@ -16354,8 +16359,7 @@ class binary_reader
value = (value << 8) | static_cast<std::int64_t>(current);
}
return negative ? sax->number_integer(static_cast<number_integer_t>(-(value + offset)))
: sax->number_unsigned(static_cast<number_unsigned_t>(value + offset));
return emit_bon8_integer(negative ? -(value + offset) : value + offset);
}
/*!
@@ -16652,6 +16656,79 @@ class binary_reader
return true;
}
/*!
@brief pass a signed integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_integer_t is passed as number_unsigned_t if it is non-negative and
fits there, and as number_float_t otherwise. With the default number
types, every integer the binary formats can encode fits, so this only
matters for narrower custom number types.
@tparam NumberType a signed integer type
@param[in] number the integer
@return whether the SAX parser accepted the value
*/
template<typename NumberType>
bool emit_signed(const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->number_integer(static_cast<number_integer_t>(number));
}
if (value_in_range_of<number_unsigned_t>(number))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return sax->number_float(static_cast<number_float_t>(number), "");
}
/*!
@brief pass an unsigned integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_unsigned_t is passed as number_float_t.
@tparam NumberType an unsigned integer type
@param[in] number the integer
@return whether the SAX parser accepted the value
*/
template<typename NumberType>
bool emit_unsigned(const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_unsigned_t>(number)))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return sax->number_float(static_cast<number_float_t>(number), "");
}
/*!
@brief pass a floating-point number read from the input to the SAX parser
Like the lexer does for JSON text, a finite value that overflows
number_float_t is rejected instead of silently becoming infinity. Infinity
and NaN in the input are passed on unchanged.
@tparam NumberType a floating-point type
@param[in] format the current format (for diagnostics)
@param[in] number the number
@return whether the SAX parser accepted the value
@throw out_of_range.406 if a finite @a number overflows number_float_t
*/
template<typename NumberType>
bool emit_float(const input_format_t format, const NumberType number)
{
const auto result = static_cast<number_float_t>(number);
if (JSON_HEDLEY_UNLIKELY(std::isfinite(number) && !std::isfinite(result)))
{
return sax->parse_error(chars_read, get_token_string(),
out_of_range::create(406, exception_message(format, "number overflow", "value"), nullptr));
}
return sax->number_float(result, "");
}
/*!
@brief create a string by reading characters from the input
@@ -23931,8 +24008,7 @@ enum class error_handler_t
{
strict, ///< throw a type_error exception in case of invalid UTF-8
replace, ///< replace invalid UTF-8 sequences with U+FFFD
ignore, ///< ignore invalid UTF-8 sequences
keep ///< keep invalid UTF-8 sequences; their bytes are copied unchanged
ignore ///< ignore invalid UTF-8 sequences
};
template<typename BasicJsonType>
@@ -24903,47 +24979,6 @@ class serializer
break;
}
case error_handler_t::keep:
{
// drop whatever the incomplete sequence left in
// the buffer (only copied if !EnsureAscii) and copy
// the ill-formed bytes from the input instead
bytes = bytes_after_last_accept;
if (undumped_chars > 0)
{
// the pending bytes of the incomplete sequence
// are ill-formed; the current byte may be OK for
// itself, so we would like to read it again
for (std::size_t j = i - undumped_chars; j < i; ++j)
{
string_buffer[bytes++] = s[j];
}
--i;
}
else
{
// the current byte cannot start any sequence
string_buffer[bytes++] = s[i];
}
// write buffer and reset index; there must be 13 bytes
// left, as this is the maximal number of bytes to be
// written ("\uxxxx\uxxxx\0") for one code point
if (string_buffer.size() - bytes < 13)
{
put_buffer(string_buffer, bytes);
bytes = 0;
}
bytes_after_last_accept = bytes;
undumped_chars = 0;
// continue processing the string
state = UTF8_ACCEPT;
break;
}
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
@@ -24989,15 +25024,6 @@ class serializer
break;
}
case error_handler_t::keep:
{
// write all accepted bytes
put_buffer(string_buffer, bytes_after_last_accept);
// copy the bytes of the incomplete sequence unchanged
put_string(s, s.size() - undumped_chars, s.size());
break;
}
case error_handler_t::replace:
{
// write all accepted bytes
+116
View File
@@ -11,7 +11,12 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <cmath>
#include <fstream>
#include <limits>
#include <map>
#include <string>
#include <vector>
#include "make_test_data_available.hpp"
TEST_CASE("Binary Formats" * doctest::skip())
@@ -224,3 +229,114 @@ TEST_CASE("Binary Formats" * doctest::skip())
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450));
}
}
TEST_CASE("Binary formats with narrow number types")
{
// Numbers that do not fit the number types are handled like the lexer
// handles them in JSON text: an integer that fits neither integer type is
// stored as a floating-point number, and a finite floating-point number
// that overflows number_float_t is rejected with out_of_range.406.
using narrow_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int32_t, std::uint32_t, float>;
using bytes = std::vector<std::uint8_t>;
struct binary_format
{
const char* name;
bytes (*encode)(const json&);
narrow_json (*decode)(const bytes&, bool);
};
const std::vector<binary_format> formats =
{
{
"CBOR", [](const json & j) { return json::to_cbor(j); },
[](const bytes & v, bool allow_exceptions)
{
return narrow_json::from_cbor(v, true, allow_exceptions);
}
},
{
"MessagePack", [](const json & j) { return json::to_msgpack(j); },
[](const bytes & v, bool allow_exceptions)
{
return narrow_json::from_msgpack(v, true, allow_exceptions);
}
},
{
"UBJSON", [](const json & j) { return json::to_ubjson(j); },
[](const bytes & v, bool allow_exceptions)
{
return narrow_json::from_ubjson(v, true, allow_exceptions);
}
},
{
"BJData", [](const json & j) { return json::to_bjdata(j); },
[](const bytes & v, bool allow_exceptions)
{
return narrow_json::from_bjdata(v, true, allow_exceptions);
}
},
{
// BSON can only store numbers as object members
"BSON", [](const json & j) { return json::to_bson(json{{"a", j}}); },
[](const bytes & v, bool allow_exceptions)
{
const auto result = narrow_json::from_bson(v, true, allow_exceptions);
return result.is_discarded() ? result : result.at("a");
}
},
{
"BON8", [](const json & j) { return json::to_bon8(j); },
[](const bytes & v, bool allow_exceptions)
{
return narrow_json::from_bon8(v, true, allow_exceptions);
}
},
};
for (const auto& format : formats)
{
const std::string name = format.name;
INFO("format := ", name);
const auto roundtrip = [&format](const json & j)
{
return format.decode(format.encode(j), true);
};
// integers that fit keep their type
CHECK(roundtrip(json(-5)).is_number_integer());
CHECK(roundtrip(json(-5)).get<std::int32_t>() == -5);
CHECK(roundtrip(json(3000000000u)).is_number_unsigned());
CHECK(roundtrip(json(3000000000u)).get<std::uint32_t>() == 3000000000u);
// integers that fit neither integer type are stored as float
CHECK(roundtrip(json(5000000000u)).is_number_float());
CHECK(roundtrip(json(5000000000u)).get<float>() == 5000000000.0f);
if (name != "BON8") // BON8 cannot encode integers above INT64_MAX
{
CHECK(roundtrip(json(10000000000000000000u)).is_number_float());
CHECK(roundtrip(json(10000000000000000000u)).get<float>() == 10000000000000000000.0f);
}
CHECK(roundtrip(json(-3000000000)).is_number_float());
CHECK(roundtrip(json(-3000000000)).get<float>() == -3000000000.0f);
CHECK(roundtrip(json(-5000000000)).is_number_float());
CHECK(roundtrip(json(-5000000000)).get<float>() == -5000000000.0f);
// floating-point numbers that fit
CHECK(roundtrip(json(1.5)).get<float>() == 1.5f);
const auto just_above_max = std::nextafter(static_cast<double>((std::numeric_limits<float>::max)()),
std::numeric_limits<double>::infinity());
CHECK(roundtrip(json(just_above_max)).get<float>() == (std::numeric_limits<float>::max)());
// infinity and NaN are passed on
CHECK(std::isinf(roundtrip(json(std::numeric_limits<double>::infinity())).get<float>()));
CHECK(std::isnan(roundtrip(json(std::numeric_limits<double>::quiet_NaN())).get<float>()));
// finite floating-point numbers that overflow number_float_t are rejected
const std::string message = "[json.exception.out_of_range.406] syntax error while parsing " + name
+ " value: number overflow";
CHECK_THROWS_WITH_AS(roundtrip(json(1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK_THROWS_WITH_AS(roundtrip(json(-1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK(format.decode(format.encode(json(1e300)), false).is_discarded());
}
}
+16 -14
View File
@@ -3146,7 +3146,8 @@ TEST_CASE("Tagged values")
// CBOR encodes negative integers as: result = -1 - n
// For type 0x3B, n is an 8-byte uint64_t. Valid range for n with
// the default int64_t is [0, INT64_MAX], producing results in [INT64_MIN, -1].
// When n > INT64_MAX, the result exceeds int64_t range and is rejected.
// When n > INT64_MAX, the result exceeds int64_t range and is stored
// as a floating-point number, as the lexer does for JSON text.
SECTION("n = 0 is valid (result = -1)")
{
@@ -3167,33 +3168,34 @@ TEST_CASE("Tagged values")
CHECK(result.get<int64_t>() == (std::numeric_limits<int64_t>::min)());
}
SECTION("n = INT64_MAX + 1 is rejected (overflow)")
SECTION("n = INT64_MAX + 1 is stored as float")
{
// n = INT64_MAX + 1 (0x8000000000000000)
// result = -1 - n = -9223372036854775809, which exceeds int64_t range
// result = -1 - n = -9223372036854775809, which exceeds int64_t range;
// the nearest double is -9223372036854775808.0
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
json::parse_error);
const auto result = json::from_cbor(input);
CHECK(result.is_number_float());
CHECK(result.get<double>() == -9223372036854775808.0);
CHECK(result == json::parse("-9223372036854775809"));
}
SECTION("n = UINT64_MAX is rejected (overflow)")
SECTION("n = UINT64_MAX is stored as float")
{
// n = UINT64_MAX (0xFFFFFFFFFFFFFFFF)
// result = -1 - n = -18446744073709551616, which exceeds int64_t range
const std::vector<uint8_t> input = {0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
json::parse_error);
const auto result = json::from_cbor(input);
CHECK(result.is_number_float());
CHECK(result.get<double>() == -18446744073709551616.0);
CHECK(result == json::parse("-18446744073709551616"));
}
SECTION("overflow with allow_exceptions=false returns discarded")
SECTION("overflow with allow_exceptions=false is not an error")
{
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
const auto result = json::from_cbor(input, true, false);
CHECK(result.is_discarded());
CHECK(result.is_number_float());
}
}
-9
View File
@@ -766,15 +766,6 @@ TEST_CASE("regression tests 2")
CHECK(j == k);
}
SECTION("issue #4552 - UTF-8 invalid characters are not always ignored when dumping with error_handler_t::ignore")
{
json node;
node["test"] = "test\334\005";
CHECK(node.dump(-1, ' ', false, json::error_handler_t::ignore) == "{\"test\":\"test\\u0005\"}");
CHECK(node.dump(-1, ' ', false, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
CHECK(node.dump(-1, ' ', true, json::error_handler_t::keep) == "{\"test\":\"test\334\\u0005\"}");
}
}
TEST_CASE("regression test - parser callback must not lose a duplicate key's prior value")
-37
View File
@@ -92,8 +92,6 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"äü\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"ä\xEF\xBF\xBDü\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"\\u00e4\\ufffd\\u00fc\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"ä\xA9ü\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\\u00e4\xA9\\u00fc\"");
}
SECTION("invalid character (regression guard for shared UTF-8 decoder, see #5529)")
@@ -116,8 +114,6 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xC2\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xC2\"");
}
SECTION("unexpected character")
@@ -130,39 +126,6 @@ TEST_CASE("serialization")
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"123456\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"123\xEF\xBF\xBD\x34\x35\x36\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"123\\ufffd456\"");
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"123\xF1\xB0\x34\x35\x36\"");
}
SECTION("keep: valid characters are still escaped")
{
// an invalid byte followed by characters that must be escaped
const json j = "\xC2\"\\\n\xFF\x05";
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
CHECK(j.dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xC2\\\"\\\\\\n\xFF\\u0005\"");
}
SECTION("keep: truncated multibyte sequences")
{
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
CHECK(json("\xF0\x9F\x98").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98\"");
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', false, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
CHECK(json("\xF0\x9F\x98" "a").dump(-1, ' ', true, json::error_handler_t::keep) == "\"\xF0\x9F\x98" "a\"");
}
SECTION("keep: long string with many invalid bytes")
{
// exceeds the internal string buffer several times
std::string input;
std::string expected = "\"";
for (int i = 0; i < 2000; ++i)
{
input += "\xFF\xE2\x82\n\xC3\xA4";
expected += "\xFF\xE2\x82\\n\xC3\xA4";
}
expected += "\"";
const json j = input;
CHECK(j.dump(-1, ' ', false, json::error_handler_t::keep) == expected);
}
SECTION("U+FFFD Substitution of Maximal Subparts")
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);
+1 -25
View File
@@ -14,7 +14,6 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <fstream>
#include <sstream>
#include <iostream>
@@ -76,11 +75,8 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
static std::string s_replaced2;
static std::string s_replaced_ascii;
static std::string s_replaced2_ascii;
static std::string s_kept;
static std::string s_kept2;
static std::string s_kept_ascii;
// dumping with ignore/replace/keep must not throw in any case
// dumping with ignore/replace must not throw in any case
s_ignored = j.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored2 = j2.dump(-1, ' ', false, json::error_handler_t::ignore);
s_ignored_ascii = j.dump(-1, ' ', true, json::error_handler_t::ignore);
@@ -89,9 +85,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
s_replaced2 = j2.dump(-1, ' ', false, json::error_handler_t::replace);
s_replaced_ascii = j.dump(-1, ' ', true, json::error_handler_t::replace);
s_replaced2_ascii = j2.dump(-1, ' ', true, json::error_handler_t::replace);
s_kept = j.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept2 = j2.dump(-1, ' ', false, json::error_handler_t::keep);
s_kept_ascii = j.dump(-1, ' ', true, json::error_handler_t::keep);
if (success_expected)
{
@@ -101,7 +94,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// all dumps should agree on the string
CHECK(s_strict == s_ignored);
CHECK(s_strict == s_replaced);
CHECK(s_strict == s_kept);
}
else
{
@@ -113,20 +105,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
// check that replace string contains a replacement character
CHECK(s_replaced.find("\xEF\xBF\xBD") != std::string::npos);
// ignore drops the invalid bytes, keep copies them
CHECK(s_ignored != s_kept);
CHECK(s_ignored_ascii != s_kept_ascii);
// unless a byte needs escaping, keep copies the input unchanged
const bool needs_escaping = std::any_of(json_string.begin(), json_string.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x20 || c == '"' || c == '\\';
});
if (!needs_escaping)
{
CHECK(s_kept == "\"" + json_string + "\"");
}
}
// check that prefix and suffix are preserved
@@ -138,8 +116,6 @@ void check_utf8dump(bool success_expected, int byte1, int byte2 = -1, int byte3
CHECK(s_replaced2.substr(s_replaced2.size() - 4, 3) == "xyz");
CHECK(s_replaced2_ascii.substr(1, 3) == "abc");
CHECK(s_replaced2_ascii.substr(s_replaced2_ascii.size() - 4, 3) == "xyz");
CHECK(s_kept2.substr(1, 3) == "abc");
CHECK(s_kept2.substr(s_kept2.size() - 4, 3) == "xyz");
}
void check_utf8string(bool success_expected, int byte1, int byte2, int byte3, int byte4);