Write doubles with the shortest digits (Żmij), digits in registers

Write doubles with the conversion of Zmij by Victor Zverovich (MIT),
ported to C++11 (detail/conversions/zmij.hpp). It finds the
shortest decimal that reads back as the same double, and the
closest one if there are several. Grisu2, used until now, is fast
but not always shortest: it sometimes writes a 17th digit where 16
suffice, or a last digit that is not the closest. The layout is
unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2.

Digits are converted eight at a time with the BCD conversion of
Xiang JunBo, as in Zmij, and written with one byte swap per eight
digits and fixed-size moves instead of per-digit loops. Leading and
trailing zeros are counted from those bytes. to_chars() uses a
local buffer when the caller's is shorter than the 41 bytes this
may write. The powers of ten come from the number-parsing table,
adjusted where it holds values rounded up, and extended with Zmij's
compressed tables beyond 10^308.

write_shortest() converts its 16 digits in one vector register
(SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the
decimal point inside the register, avoiding a store-forwarding
stall that cost about 25% of the time to write a double. dump()
writes floats and integers straight into the serializer's write
buffer instead of copying them from a member buffer, and small
integers eight digits at a time. read_eight_bytes() and
parse_eight_digits() are marked always-inline, which GCC had been
calling out of line in the number-parsing loops.

Of one million random doubles, about 0.14% are now written with
different digits, always to a value that still reads back as the
same double.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann committed 2026-10-07 16:41:47 +02:00
1 parent 68d61b6aef
commit bfa4886f0f
15 files changed
+1823 -107

No files matched your search

+5
View File
@@ -62,6 +62,9 @@ Linear.
## Notes
Floating-point numbers are written with the fewest digits that read back as the same value (for `#!cpp double`; see
[number handling](../../features/types/number_handling.md#number-serialization)).
Binary values are serialized as an object containing two keys:
- "bytes": an array of bytes as integers
@@ -97,3 +100,5 @@ Binary values are serialized as an object containing two keys:
- Error handlers added in version 3.4.0.
- Serialization of binary values added in version 3.8.0.
- Error handler `keep` added in version 3.13.0.
- Doubles are written with the shortest digits (Żmij instead of Grisu2) since version 3.13.0; about 0.1% of doubles are
written differently, most of them with fewer digits.
@@ -134,9 +134,10 @@ That is, `-0` is stored as a signed integer, but the serialization does not repr
### Number serialization
- Integer numbers are serialized as is; that is, no scientific notation is used.
- Floating-point numbers are serialized as specified by the `#!c %g` printf modifier with
[`std::numeric_limits<double>::max_digits10`](https://en.cppreference.com/w/cpp/types/numeric_limits/max_digits10)
significant digits. The rationale is to use the shortest representation while still allowing round-tripping.
- Floating-point numbers are serialized with the fewest digits that read back as the same value (the closest such
digits if there are several), in the layout of the `#!c %g` printf modifier: `#!c 1.5`, `#!c 100.0`, `#!c 1e+100`.
Doubles are converted with the algorithm of [Żmij](https://github.com/vitaut/zmij), floats with Grisu2, which
can write more digits than necessary.
!!! hint "Notes regarding precision of floating-point numbers"
@@ -545,9 +545,10 @@ therefore silently changes parse results rather than raising an error. See
specifiers, for which the library likewise provides only `#!cpp double` and `#!cpp long double` overloads
(`#!cpp float` is promoted to `#!cpp double`).
If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary32 or binary64 number, `dump` uses the
Grisu2 algorithm, which produces the shortest representation that round-trips. Otherwise the `snprintf` fallback with
`max_digits10` digits is used.
If `#!cpp std::numeric_limits<NumberFloatType>` describes an IEEE 754 binary64 number, `dump` uses the algorithm of
Żmij, which produces the shortest representation that round-trips. For IEEE 754 binary32 numbers, it uses Grisu2,
which produces a short representation that round-trips. Otherwise the `snprintf` fallback with `max_digits10` digits is
used.
### Required for the binary formats
@@ -559,7 +560,7 @@ binary32 or binary64 field and have no encoding for `#!cpp long double`.
| Type | Support |
|--------------------------|-----------------------------------------------------------------------------------------------------------------------|
| `#!cpp double` (default) | full; short round-trip output through Grisu2 |
| `#!cpp double` (default) | full; shortest round-trip output through Żmij |
| `#!cpp float` | full; short round-trip output through Grisu2 |
| `#!cpp long double` | `dump` and `parse` only; the binary format writers do not compile, as they only handle IEEE 754 binary32 and binary64 |
| any other type | not usable |
+2
View File
@@ -18,6 +18,8 @@ The class contains the UTF-8 Decoder from Bjoern Hoehrmann which is licensed und
The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2009 [Florian Loitsch](https://florian.loitsch.com/)
The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright &copy; 2025 [Victor Zverovich](https://github.com/vitaut)
The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors