mirror of
https://github.com/nlohmann/json.git
synced 2026-10-07 15:07:13 +00:00
Write doubles with the shortest digits (Żmij), digits in registers
Write doubles with the conversion of Zmij by Victor Zverovich (MIT), ported to C++11 (detail/conversions/zmij.hpp). It finds the shortest decimal that reads back as the same double, and the closest one if there are several. Grisu2, used until now, is fast but not always shortest: it sometimes writes a 17th digit where 16 suffice, or a last digit that is not the closest. The layout is unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2. Digits are converted eight at a time with the BCD conversion of Xiang JunBo, as in Zmij, and written with one byte swap per eight digits and fixed-size moves instead of per-digit loops. Leading and trailing zeros are counted from those bytes. to_chars() uses a local buffer when the caller's is shorter than the 41 bytes this may write. The powers of ten come from the number-parsing table, adjusted where it holds values rounded up, and extended with Zmij's compressed tables beyond 10^308. write_shortest() converts its 16 digits in one vector register (SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the decimal point inside the register, avoiding a store-forwarding stall that cost about 25% of the time to write a double. dump() writes floats and integers straight into the serializer's write buffer instead of copying them from a member buffer, and small integers eight digits at a time. read_eight_bytes() and parse_eight_digits() are marked always-inline, which GCC had been calling out of line in the number-parsing loops. Of one million random doubles, about 0.14% are now written with different digits, always to a value that still reads back as the same double. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
15 files changed
+1823
-107
No files matched your search
@@ -62,6 +62,9 @@ Linear.
|
||||
|
||||
## Notes
|
||||
|
||||
Floating-point numbers are written with the fewest digits that read back as the same value (for `#!cpp double`; see
|
||||
[number handling](../../features/types/number_handling.md#number-serialization)).
|
||||
|
||||
Binary values are serialized as an object containing two keys:
|
||||
|
||||
- "bytes": an array of bytes as integers
|
||||
@@ -97,3 +100,5 @@ Binary values are serialized as an object containing two keys:
|
||||
- Error handlers added in version 3.4.0.
|
||||
- Serialization of binary values added in version 3.8.0.
|
||||
- Error handler `keep` added in version 3.13.0.
|
||||
- Doubles are written with the shortest digits (Żmij instead of Grisu2) since version 3.13.0; about 0.1% of doubles are
|
||||
written differently, most of them with fewer digits.
|
||||
@@ -134,9 +134,10 @@ That is, `-0` is stored as a signed integer, but the serialization does not repr
|
||||
### Number serialization
|
||||
|
||||
- Integer numbers are serialized as is; that is, no scientific notation is used.
|
||||
- Floating-point numbers are serialized as specified by the `#!c %g` printf modifier with
|
||||
[`std::numeric_limits<double>::max_digits10`](https://en.cppreference.com/w/cpp/types/numeric_limits/max_digits10)
|
||||
significant digits. The rationale is to use the shortest representation while still allowing round-tripping.
|
||||
- Floating-point numbers are serialized with the fewest digits that read back as the same value (the closest such
|
||||
digits if there are several), in the layout of the `#!c %g` printf modifier: `#!c 1.5`, `#!c 100.0`, `#!c 1e+100`.
|
||||
Doubles are converted with the algorithm of [Żmij](https://github.com/vitaut/zmij), floats with Grisu2, which
|
||||
can write more digits than necessary.
|
||||
|
||||
!!! hint "Notes regarding precision of floating-point numbers"
|
||||
|
||||
|
||||
@@ -545,9 +545,10 @@ therefore silently changes parse results rather than raising an error. See
|
||||
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
|
||||
(`#!cpp float` is promoted to `#!cpp double`).
|
||||
|
||||
If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary32 or binary64 number, `dump` uses the
|
||||
Grisu2 algorithm, which produces the shortest representation that round-trips. Otherwise the `snprintf` fallback with
|
||||
`max_digits10` digits is used.
|
||||
If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary64 number, `dump` uses the algorithm of
|
||||
Żmij, which produces the shortest representation that round-trips. For IEEE 754 binary32 numbers, it uses Grisu2,
|
||||
which produces a short representation that round-trips. Otherwise the `snprintf` fallback with `max_digits10` digits is
|
||||
used.
|
||||
|
||||
### Required for the binary formats
|
||||
|
||||
@@ -559,7 +560,7 @@ binary32 or binary64 field and have no encoding for `#!cpp long double`.
|
||||
|
||||
| Type | Support |
|
||||
|--------------------------|-----------------------------------------------------------------------------------------------------------------------|
|
||||
| `#!cpp double` (default) | full; short round-trip output through Grisu2 |
|
||||
| `#!cpp double` (default) | full; shortest round-trip output through Żmij |
|
||||
| `#!cpp float` | full; short round-trip output through Grisu2 |
|
||||
| `#!cpp long double` | `dump` and `parse` only; the binary format writers do not compile, as they only handle IEEE 754 binary32 and binary64 |
|
||||
| any other type | not usable |
|
||||
|
||||
@@ -18,6 +18,8 @@ The class contains the UTF-8 Decoder from Bjoern Hoehrmann which is licensed und
|
||||
|
||||
The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2009 [Florian Loitsch](https://florian.loitsch.com/)
|
||||
|
||||
The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2025 [Victor Zverovich](https://github.com/vitaut)
|
||||
|
||||
The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
|
||||
|
||||
The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors
|
||||
Reference in new issue
Block a user