Convert floats independently of the C locale's decimal point

Under a locale whose decimal point is longer than one byte (e.g. U+066B
in fa_IR.UTF-8 or ar_EG.UTF-8), every float that reached the strtod
fallback was truncated at the decimal point: "3.14159265358979323846"
became 3.0, and "1.5e400" became 1.0 instead of throwing. With libc++
and in C++11/14, that fallback was taken for most floats.

The lexer now converts floats in this order:

1. std::from_chars, now also for float and double with libc++ 20 or
   later, which does not define __cpp_lib_to_chars (on Apple platforms
   only if the deployment target provides it);
2. Clinger's fast path (double only);
3. strtof_l/strtod_l/strtold_l with a "C" locale created once, on
   glibc, Apple platforms, and MSVC;
4. strtof/strtod/strtold with the decimal point of the current locale,
   which now puts a multi-byte decimal point into a copy of the token.

If std::from_chars reports a value out of range, the result is derived
from the token (+-infinity or +-0) instead of calling strtod, because
implementations disagree on the stored value (P4168). Values that may
be subnormal are left to the next step, because libstdc++ before
GCC 13 reports some of them as out of range.

The conversion helpers moved from the lexer to number_parse.hpp, so the
last-resort path can be tested directly.

Fixes #5660.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-30 09:51:50 +02:00
parent 633de8e44b
commit 436bfb1358
8 changed files with 1003 additions and 241 deletions
+2
View File
@@ -254,6 +254,8 @@ outside of a string, invalid) byte; see the [FAQ entry](../../home/faq.md#nul-by
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- `JSON_STRICT_NUL_HANDLING` added in version 3.13.0 to optionally reject a NUL byte in the input instead of treating
it as end of input; planned to become the default in version 4.0.0.
- The result of converting floating-point numbers no longer depends on the C locale in version 3.13.0; before, a
locale whose decimal point is longer than one byte (e.g., `fa_IR.UTF-8`) truncated them at the decimal point.
!!! warning "Deprecation"
@@ -75,6 +75,13 @@ otherwise, it uses unsigned integer storage.
[`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul),
[`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof), respectively.
- The result of converting floating-point numbers does not depend on the C locale (`LC_NUMERIC`). They are
converted with [`std::from_chars`](https://en.cppreference.com/w/cpp/utility/from_chars) where the standard
library implements it for the number type (including libc++ 20 or later for `#!c float` and `#!c double`),
otherwise with `strtod_l` and the "C" locale where the C library provides it (glibc, macOS, MSVC), and otherwise
with `std::strtod` and the decimal point of the current locale. Before version 3.13.0, the last way was used much
more often, and a locale whose decimal point is longer than one byte (e.g., `fa_IR.UTF-8`) truncated numbers at
the decimal point.
!!! example "Examples"
@@ -85,10 +92,11 @@ otherwise, it uses unsigned integer storage.
### Number limits
- Any 64-bit signed or unsigned integer can be stored without loss of precision.
- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion via
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) are not satisfying
- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion are not satisfying
[`std::isfinite`](https://en.cppreference.com/w/cpp/numeric/math/isfinite) such as `#!c 1E400`) will throw exception
[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing.
- Numbers too close to zero to be represented as `#!c double`, not even as subnormal number (such as `#!c 1E-400`), are
stored as `#!c 0.0`, or as `#!c -0.0` if they are negative.
- Floating-point numbers are rounded to the next number representable as `double`. For instance
`#!c 3.141592653589793238462643383279` is stored as [`0x400921fb54442d18`](https://float.exposed/0x400921fb54442d18).
This is the same behavior as the code `#!c double x = 3.141592653589793238462643383279;`.