mirror of
https://github.com/nlohmann/json.git
synced 2026-10-07 06:57:14 +00:00
Speed up the lexer: own float parser, string scan, and \u table
Give the library its own correctly rounded float converter for binary32 and binary64 (IEEE 754), and speed up the lexer's string and escape scanning. The converter splits a number token into sign, significand, and decimal exponent, then tries Clinger's fast path, then a templated Eisel-Lemire step, and falls back to an exact big-integer digit comparison for tokens with more than 19 significant digits whose two candidate values round differently. This replaces std::from_chars and strtod/strtof for both formats, so parsed values no longer depend on the C/C++ library or the current locale. The strtold fallback kept for other long double formats (x87, binary128) now also copies a multi-byte decimal point correctly, fixing #5660. eisel_lemire() and decimal_to_float() are always inlined so callers keep the whole conversion in their hot loop. The string-scanning kernels in string_scan.hpp find a stop byte with the trailing-zero count of the SWAR mask instead of a byte loop, and scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one after another instead of re-searching after each one. get_codepoint() decodes a contiguous \uXXXX escape with one table lookup per byte instead of four range-checked get() calls; the streaming path and all error positions are unchanged. Adds 508 generated hard float-parsing cases with expected binary32 and binary64 bits, and kernel-comparison tests for the string scans and the escape table against byte-by-byte references. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
1 parent
5ecb704f6b
commit
953d74ddcd
13 files changed
+2925
-980
No files matched your search
@@ -23,9 +23,10 @@ type to use.
|
||||
## Template parameters
|
||||
|
||||
`NumberFloatType`
|
||||
: the type to store floating-point numbers. Parsing and serialization are implemented in terms of
|
||||
`#!cpp std::strtof`/`#!cpp std::strtod`/`#!cpp std::strtold` and `#!cpp std::snprintf`, so the type must be
|
||||
`#!cpp float`, `#!cpp double`, or `#!cpp long double`. The
|
||||
: the type to store floating-point numbers. The parser converts `#!cpp float`, `#!cpp double`, and a
|
||||
`#!cpp long double` that is IEEE 754 binary64 itself and other `#!cpp long double` formats with
|
||||
`#!cpp std::from_chars` or `#!cpp std::strtold`, and serialization falls back to `#!cpp std::snprintf`, so the
|
||||
type must be `#!cpp float`, `#!cpp double`, or `#!cpp long double`. The
|
||||
[binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`,
|
||||
because they have no encoding for `#!cpp long double`. See
|
||||
[Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype).
|
||||
|
||||
@@ -82,12 +82,13 @@ flowchart TD
|
||||
|
||||
- Numbers with a decimal digit or scientific notation are always stored as `#!c double`.
|
||||
- The number types can be changed, see [Template number types](#template-number-types).
|
||||
- Integers are converted by the library's own digit parser. Floating-point numbers are converted with
|
||||
[`std::from_chars`](https://en.cppreference.com/w/cpp/utility/from_chars) if the library is compiled with C++17
|
||||
and the standard library supports it, then with an exact fast path for `#!c double` values with few significant
|
||||
digits, and otherwise with the locale-aware
|
||||
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) (`std::strtof`/`std::strtold` for the
|
||||
other floating-point types). Before version 3.13.0, the conversion was realized by
|
||||
- The library converts integers and floating-point numbers itself, independent of the locale. Floating-point
|
||||
numbers are correctly rounded (to nearest, ties to even). Only a `#!c long double` that is not IEEE 754 binary64
|
||||
(e.g., the 80-bit x87 format) is converted with `#!cpp std::from_chars` where available, or else with
|
||||
[`std::strtold`](https://en.cppreference.com/w/cpp/string/byte/strtof). For that call, the library temporarily
|
||||
replaces the `.` with the decimal point of the current locale (which may be longer than one byte, e.g., in
|
||||
`fa_IR.UTF-8`), so the result does not depend on the locale either. Changing the locale in another thread during
|
||||
parsing is undefined behavior of the C library, though. Before version 3.13.0, the conversion was realized by
|
||||
[`std::strtoull`](https://en.cppreference.com/w/cpp/string/byte/strtoul),
|
||||
[`std::strtoll`](https://en.cppreference.com/w/cpp/string/byte/strtol), and `std::strtod`, respectively.
|
||||
|
||||
@@ -100,10 +101,10 @@ flowchart TD
|
||||
### Number limits
|
||||
|
||||
- Any 64-bit signed or unsigned integer can be stored without loss of precision.
|
||||
- Numbers exceeding the limits of `#!c double` (i.e., numbers that after conversion via
|
||||
[`std::strtod`](https://en.cppreference.com/w/cpp/string/byte/strtof) are not satisfying
|
||||
- Numbers exceeding the limits of `#!c double` (i.e., numbers whose rounded value is not satisfying
|
||||
[`std::isfinite`](https://en.cppreference.com/w/cpp/numeric/math/isfinite) such as `#!c 1E400`) will throw exception
|
||||
[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing.
|
||||
[`json.exception.out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406) during parsing. Numbers too
|
||||
small for `#!c double` (such as `#!c 1E-400`) become zero, with the sign of the number.
|
||||
- Floating-point numbers are rounded to the next number representable as `double`. For instance
|
||||
`#!c 3.141592653589793238462643383279` is stored as [`0x400921fb54442d18`](https://float.exposed/0x400921fb54442d18).
|
||||
This is the same behavior as the code `#!c double x = 3.141592653589793238462643383279;`.
|
||||
|
||||
@@ -26,9 +26,9 @@ Requirements are split into two groups:
|
||||
diagnosed with dedicated error messages, and violating most of them results in a compiler error somewhere inside
|
||||
the library. Four violations are not caught at compile time at all:
|
||||
|
||||
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and can silently misparse
|
||||
floating-point numbers, because the lexer may hand the buffer to `#!cpp std::strtod`, which reads up to the
|
||||
terminating null character.
|
||||
- A [`StringType`](#stringtype) whose `data()` is not null-terminated compiles and silently misparses numbers
|
||||
stored as a `#!cpp long double` that is not IEEE 754 binary64 (e.g., the 80-bit x87 format), because the lexer
|
||||
hands the buffer to `#!cpp std::strtold`.
|
||||
- A stateful [`AllocatorType`](#allocatortype) compiles and silently ignores its state: allocation, deallocation,
|
||||
and [`get_allocator()`](../../api/basic_json/get_allocator.md) each use a different default-constructed instance.
|
||||
- The two [cross-specialization conversions](#cross-specialization-conversions) below. These abort on an assertion
|
||||
@@ -537,9 +537,10 @@ therefore silently changes parse results rather than raising an error. See
|
||||
|
||||
`NumberFloatType` must be one of `#!cpp float`, `#!cpp double`, or `#!cpp long double`:
|
||||
|
||||
- The [parser](../parsing/index.md) converts number literals with `#!cpp std::from_chars` or, as a fallback, with
|
||||
`#!cpp std::strtof`, `#!cpp std::strtod`, or `#!cpp std::strtold`; the library provides overloads for exactly these
|
||||
three types.
|
||||
- The [parser](../parsing/index.md) converts number literals to `#!cpp float`, `#!cpp double`, and a
|
||||
`#!cpp long double` that is IEEE 754 binary64 itself; other `#!cpp long double` formats are converted with
|
||||
`#!cpp std::from_chars` where available, or with `#!cpp std::strtold`. The library provides overloads for exactly
|
||||
these three types.
|
||||
- [`dump`](../../api/basic_json/dump.md) falls back to `#!cpp std::snprintf` with the `%g` and `%Lg` conversion
|
||||
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
|
||||
(`#!cpp float` is promoted to `#!cpp double`).
|
||||
|
||||
@@ -20,4 +20,4 @@ The class contains a slightly modified version of the Grisu2 algorithm from Flor
|
||||
|
||||
The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
|
||||
|
||||
The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors
|
||||
The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors
|
||||
Reference in new issue
Block a user