Handle numbers that do not fit narrow number types in the binary readers (#5607)

* Handle numbers that do not fit narrow number types in the binary readers

With custom number types narrower than the values in a binary document,
for example basic_json<..., std::int32_t, std::uint32_t, float>, every
binary reader (CBOR, MessagePack, UBJSON, BJData, BSON, BON8) passed the
decoded number to the SAX interface with an implicit conversion: the
integer 5000000000 silently became 705032704, and a finite double such as
1e300 became infinity. The lexer handles the same values in JSON text: an
integer that fits neither integer type is stored as number_float_t, and a
finite number that overflows number_float_t is rejected with
out_of_range.406.

Pass every number read from binary input through three helpers that
apply the lexer's rules:
- emit_signed(): number_integer_t, else number_unsigned_t for a
  non-negative value, else number_float_t
- emit_unsigned(): number_unsigned_t, else number_float_t
- emit_float(): out_of_range.406 if a finite value overflows
  number_float_t; infinity and NaN are passed on

For consistency, a CBOR negative integer below the range of
number_integer_t is now stored as number_float_t, like a too small
integer in JSON text, instead of being rejected with parse_error.112.
With the default number types, this is the only change in behavior.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix MSVC and clang 3.5 in the narrow number type test

MSVC types 3000000000 and 5000000000 as unsigned long, so
json(-3000000000) triggered C4146 (unary minus on an unsigned type),
which /WX turns into an error. Use LL literals, as elsewhere in the
tests.

clang 3.5 cannot convert the lambdas in the braced initializer of the
format table to function pointers. Use named functions instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Check integer-to-float fallbacks for overflow in the binary readers

emit_signed, emit_unsigned, and the CBOR negative integer fallback now
pass their number_float_t fallback through emit_float, so a value that
overflows number_float_t is rejected with out_of_range.406 like a
floating-point value, instead of silently becoming infinity. This only
matters for a number_float_t that cannot represent 2^64, such as a
half-precision type. The CBOR value -1 - n is computed as long double so
that emit_float sees a finite value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Use the input_format member instead of passing the format to binary_reader helpers

The helpers (get_number, get_to, get_string, get_binary, get_bytes,
emit_signed, emit_unsigned, emit_float, unexpect_eof, exception_message)
are members of binary_reader, which already stores the format it was
constructed with, so the parameter was redundant.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann authored and GitHub committed 2026-10-04 17:53:29 +02:00
1 parent d8d47be4a5
commit 6b4b825af2
9 files changed
+743 -448

No files matched your search

@@ -55,6 +55,10 @@ This implementation does exactly follow this approach, as it uses double precisi
smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally
and be serialized to `null`.
During deserialization (from JSON text or any of the binary formats), a finite number that does not fit into
`number_float_t` is rejected with [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406), for
example a double-precision number in a binary format when `number_float_t` is `#!cpp float`.
### Storage
Floating-point number values are stored directly inside a `basic_json` type.
@@ -47,8 +47,9 @@ With the default values for `NumberIntegerType` (`std::int64_t`), the default va
When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and
the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of
range will yield over/underflow when used in a constructor. During deserialization, too large or small integer numbers
will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md).
range will yield over/underflow when used in a constructor. During deserialization (from JSON text or any of the binary
formats), too large or small integer numbers will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md)
or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are
@@ -48,8 +48,9 @@ With the default values for `NumberUnsignedType` (`std::uint64_t`), the default
When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and
the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow
when used in a constructor. During deserialization, too large or small integer numbers will automatically be stored
as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md).
when used in a constructor. During deserialization (from JSON text or any of the binary formats), too large or small
integer numbers will automatically be stored as [`number_integer_t`](number_integer_t.md) or
[`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are