diff --git a/README.md b/README.md index cffb1cc41..81a4cae2c 100644 --- a/README.md +++ b/README.md @@ -1402,7 +1402,7 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I - The class contains the UTF-8 Decoder from Bjoern Hoehrmann which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2008-2009 [Björn Hoehrmann](https://bjoern.hoehrmann.de/) - The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2009 [Florian Loitsch](https://florian.loitsch.com/) -- The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2025 [Victor Zverovich](https://github.com/vitaut) +- The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, including the conversion of the digits to text by Xiang JunBo and the SIMD instruction sequence of Dougall Johnson, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2025 [Victor Zverovich](https://github.com/vitaut) - The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). - The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0). - The class contains an adapted version of the Eisel-Lemire algorithm, its table of powers of five, and its digit comparison for long numbers from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright © 2021 The fast_float authors diff --git a/docs/mkdocs/docs/api/basic_json/number_float_t.md b/docs/mkdocs/docs/api/basic_json/number_float_t.md index aa41f33f0..1cde41960 100644 --- a/docs/mkdocs/docs/api/basic_json/number_float_t.md +++ b/docs/mkdocs/docs/api/basic_json/number_float_t.md @@ -23,10 +23,10 @@ type to use. ## Template parameters `NumberFloatType` -: the type to store floating-point numbers. The parser converts `#!cpp float`, `#!cpp double`, and a - `#!cpp long double` that is IEEE 754 binary64 itself and other `#!cpp long double` formats with - `#!cpp std::from_chars` or `#!cpp std::strtold`, and serialization falls back to `#!cpp std::snprintf`, so the - type must be `#!cpp float`, `#!cpp double`, or `#!cpp long double`. The +: the type to store floating-point numbers. The type must be `#!cpp float`, `#!cpp double`, or + `#!cpp long double`. The parser converts `#!cpp float`, `#!cpp double`, and a `#!cpp long double` that is IEEE 754 + binary64 itself. It converts other `#!cpp long double` formats with `#!cpp std::from_chars` where available, or + with `#!cpp std::strtold` otherwise. Serialization falls back to `#!cpp std::snprintf`. The [binary formats](../../features/binary_formats/index.md) additionally require `#!cpp float` or `#!cpp double`, because they have no encoding for `#!cpp long double`. See [Template Parameter Requirements](../../features/types/template_parameters.md#numberfloattype). diff --git a/docs/mkdocs/docs/features/types/template_parameters.md b/docs/mkdocs/docs/features/types/template_parameters.md index b2560f372..162cfbad5 100644 --- a/docs/mkdocs/docs/features/types/template_parameters.md +++ b/docs/mkdocs/docs/features/types/template_parameters.md @@ -353,16 +353,21 @@ using array_t = ArrayType>; ### Always required - A member type `value_type` that is one byte wide and `char`-compatible. The library stores and processes UTF-8 - encoded `char` data and passes `data()` to functions that take a `#!cpp const char*`, such as `#!cpp std::strtod`. + encoded `char` data and passes `data()` to functions that take a `#!cpp const char*`, such as `#!cpp std::strtold` + (only used to parse a `#!cpp long double` that is not IEEE 754 binary64, see + [`NumberFloatType`](#numberfloattype)). `#!cpp std::wstring`, `#!cpp std::u16string`, and `#!cpp std::u32string` are **not** valid choices; see the FAQ on [wide string handling](../../home/faq.md#wide-string-handling). - Constructors: default, copy, move, from `#!cpp const char*` (which must not be `#!cpp explicit`), from `#!cpp (const char*, size_type)`, and from `#!cpp (size_type, char)`; and copy or move assignment. - Member functions `size()`, `clear()`, `resize(n, c)`, `data()`, `push_back(char)`, and `operator[]` (const and non-const, returning references). `c_str()` and `back()` are **not** required. -- `data()` must return a pointer to a contiguous, **null-terminated** buffer -- the parser may hand it to - `#!cpp std::strtod`, which reads up to the null character. A type whose `data()` is not null-terminated does not - fail to compile; it can silently misparse floating-point numbers. +- `data()` must return a pointer to a contiguous, **null-terminated** buffer. `#!cpp float`, `#!cpp double`, and a + `#!cpp long double` that is IEEE 754 binary64 are converted by the library itself and do not depend on this. For any + other `NumberFloatType` (a `#!cpp long double` of another format), the parser falls back to `#!cpp std::strtold` when + `#!cpp std::from_chars` is not available or declines the token, and `std::strtold` reads up to the null character. A type whose `data()` + is not null-terminated does not fail to compile; with such a `NumberFloatType` it can silently misparse + floating-point numbers. - `append(const char*, size_type)`, used by [`dump`](../../api/basic_json/dump.md), and `append(const StringType&)`, used by the CBOR reader for indefinite-length strings. The library's internal string concatenation additionally has to append a `#!cpp char` and a `#!cpp const char*`; for each it selects between `append(arg)`, `#!cpp operator+=`, diff --git a/docs/mkdocs/docs/home/license.md b/docs/mkdocs/docs/home/license.md index 3327ad791..ef877fb1b 100644 --- a/docs/mkdocs/docs/home/license.md +++ b/docs/mkdocs/docs/home/license.md @@ -18,7 +18,7 @@ The class contains the UTF-8 Decoder from Bjoern Hoehrmann which is licensed und The class contains a slightly modified version of the Grisu2 algorithm from Florian Loitsch which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2009 [Florian Loitsch](https://florian.loitsch.com/) -The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2025 [Victor Zverovich](https://github.com/vitaut) +The class contains a port of the shortest double-to-decimal conversion of [Żmij](https://github.com/vitaut/zmij) by Victor Zverovich, including the conversion of the digits to text by Xiang JunBo and the SIMD instruction sequence of Dougall Johnson, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above). Copyright © 2025 [Victor Zverovich](https://github.com/vitaut) The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/). diff --git a/include/nlohmann/detail/bit_ops.hpp b/include/nlohmann/detail/bit_ops.hpp index 9da853faf..42a09b68e 100644 --- a/include/nlohmann/detail/bit_ops.hpp +++ b/include/nlohmann/detail/bit_ops.hpp @@ -9,8 +9,9 @@ #pragma once #include // uint64_t -#if !defined(__SIZEOF_INT128__) && defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) - #include // __umulh, _umul128 +#include // memcpy +#if defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) && (!defined(__SIZEOF_INT128__) || (!defined(__GNUC__) && !defined(__clang__))) + #include // __umulh, _umul128, _BitScanForward64, _BitScanReverse64 #endif #include // JSON_HEDLEY_ALWAYS_INLINE, NLOHMANN_JSON_NAMESPACE_BEGIN @@ -28,6 +29,10 @@ inline int count_leading_zeros(std::uint64_t x) noexcept { #if defined(__GNUC__) || defined(__clang__) return __builtin_clzll(x); +#elif defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) + unsigned long index = 0; + _BitScanReverse64(&index, x); + return 63 - static_cast(index); #else int n = 0; for (int shift = 32; shift != 0; shift >>= 1) @@ -47,6 +52,10 @@ inline int count_trailing_zeros(std::uint64_t x) noexcept { #if defined(__GNUC__) || defined(__clang__) return __builtin_ctzll(x); +#elif defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) + unsigned long index = 0; + _BitScanForward64(&index, x); + return static_cast(index); #else int n = 0; for (int shift = 32; shift != 0; shift >>= 1) @@ -94,15 +103,21 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc #endif } -/// eight bytes as a little-endian word (compilers fold this into one load on -/// little-endian targets; always inlined, as GCC otherwise calls it in the -/// number loops) +/// eight bytes as a little-endian word (a single load on little-endian +/// targets; always inlined, as GCC otherwise calls it in the number loops) JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const unsigned char* b) noexcept { +#if defined(_MSC_VER) || defined(__x86_64__) || defined(__i386__) || (defined(__BYTE_ORDER__) && defined(__ORDER_LITTLE_ENDIAN__) && __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__) + // the byte order already matches (all MSVC targets are little-endian) + std::uint64_t result = 0; + std::memcpy(&result, b, sizeof(result)); + return result; +#else return static_cast(b[0]) | (static_cast(b[1]) << 8u) | (static_cast(b[2]) << 16u) | (static_cast(b[3]) << 24u) | (static_cast(b[4]) << 32u) | (static_cast(b[5]) << 40u) | (static_cast(b[6]) << 48u) | (static_cast(b[7]) << 56u); +#endif } /// eight bytes as a little-endian word diff --git a/include/nlohmann/detail/conversions/to_chars.hpp b/include/nlohmann/detail/conversions/to_chars.hpp index 0e16152c3..db50795b5 100644 --- a/include/nlohmann/detail/conversions/to_chars.hpp +++ b/include/nlohmann/detail/conversions/to_chars.hpp @@ -4,6 +4,7 @@ // |_____|_____|_____|_|___| https://github.com/nlohmann/json // // SPDX-FileCopyrightText: 2009 Florian Loitsch +// SPDX-FileCopyrightText: 2025 Victor Zverovich // SPDX-FileCopyrightText: 2013-2026 Niels Lohmann // SPDX-License-Identifier: MIT @@ -939,88 +940,6 @@ void grisu2(char* buf, int& len, int& decimal_exponent, FloatType value) grisu2(buf, len, decimal_exponent, w.minus, w.w, w.plus); } -/*! -@brief the shortest digits of a positive finite float (other than double): Grisu2 -*/ -template -JSON_HEDLEY_NON_NULL(1) -void shortest_digits(char* buf, int& len, int& decimal_exponent, FloatType value) -{ - grisu2(buf, len, decimal_exponent, value); -} - -/*! -@brief the shortest digits of a positive finite double: the conversion of -Zmij (see zmij.hpp), which always finds the shortest digits that read back as -the same value (Grisu2 does not for about one double in a thousand), and the -closest of them if there are several - -v = buf * 10^decimal_exponent, as for grisu2() -*/ -JSON_HEDLEY_NON_NULL(1) -inline void shortest_digits(char* buf, int& len, int& decimal_exponent, double value) -{ - static_assert(std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53, - "internal error: the conversion of Zmij needs IEEE 754 binary64 doubles"); - JSON_ASSERT(std::isfinite(value)); - JSON_ASSERT(value > 0); - - std::uint64_t bits = 0; - std::memcpy(&bits, &value, sizeof(bits)); - zmij::decimal d = zmij::to_decimal(bits); - // without trailing zeros (up to 16): 8, 4, 2, 1 at a time - while (d.significand % 100000000 == 0) - { - d.significand /= 100000000; - d.exponent += 8; - } - if (d.significand % 10000 == 0) - { - d.significand /= 10000; - d.exponent += 4; - } - if (d.significand % 100 == 0) - { - d.significand /= 100; - d.exponent += 2; - } - if (d.significand % 10 == 0) - { - d.significand /= 10; - d.exponent += 1; - } - // at most 17 digits, written from the back two at a time - static constexpr const char* pairs = - "00010203040506070809101112131415161718192021222324252627282930313233343536373839" - "40414243444546474849505152535455565758596061626364656667686970717273747576777879" - "8081828384858687888990919293949596979899"; - std::array digits{}; - std::size_t n = digits.size(); - while (d.significand >= 100) - { - const std::uint64_t two_digits = d.significand % 100; // a variable: GCC calls a cast of the remainder useless where std::uint64_t is std::size_t - const auto i = static_cast(two_digits) * 2; - d.significand /= 100; - n -= 2; - digits[n] = pairs[i]; - digits[n + 1] = pairs[i + 1]; - } - if (d.significand >= 10) - { - const auto i = static_cast(d.significand) * 2; - n -= 2; - digits[n] = pairs[i]; - digits[n + 1] = pairs[i + 1]; - } - else - { - digits[--n] = static_cast('0' + d.significand); - } - len = static_cast(digits.size() - n); - std::memcpy(buf, digits.data() + n, static_cast(len)); - decimal_exponent = d.exponent; -} - /*! @brief appends a decimal representation of e to buf @return a pointer to the element following the exponent. @@ -1423,53 +1342,30 @@ inline char* write_shortest(char* first, const zmij::shortest_decimal d) noexcep return end + (three ? 5 : 4); } -/// the powers of ten up to 10^16 -inline const std::array& powers_of_ten_16() noexcept -{ - static const std::array powers = - { - { - 1u, 10u, 100u, 1000u, 10000u, 100000u, 1000000u, 10000000u, 100000000u, 1000000000u, 10000000000u, - 100000000000u, 1000000000000u, 10000000000000u, 100000000000000u, 1000000000000000u, 10000000000000000u - } - }; - return powers; -} - /*! -@brief digits * 10^exp, as write_decimal() writes it, for the digits of a -double that need no conversion (count digits, at most 15, the first not 0; -trailing zeros allowed): extended to 16 digits and written by write_shortest() +@brief whether FloatType is an IEEE 754 binary64 type (a double, or a long double +that has the same format, as with MSVC and on Apple's Arm CPUs) -@return a pointer past the text; up to 41 bytes at @a first are written - (some beyond the returned end) +These are the types the conversion of Zmij (see zmij.hpp) is used for; all +others (binary32, or a format the library does not know) use Grisu2. */ -JSON_HEDLEY_NON_NULL(1) -JSON_HEDLEY_RETURNS_NON_NULL -inline char* write_short_decimal(char* first, std::uint64_t digits, int count, int exp) noexcept +template +constexpr bool has_binary64_format() noexcept { - JSON_ASSERT(digits >= powers_of_ten_16()[static_cast(count - 1)] && count <= 15); - const int scale = 16 - count; - return write_shortest(first, zmij::shortest_decimal{digits * powers_of_ten_16()[static_cast(scale)], exp - scale - 1, 0, false}); + return std::numeric_limits::is_iec559 + && std::numeric_limits::digits == 53 + && std::numeric_limits::max_exponent == 1024 + && sizeof(FloatType) == sizeof(std::uint64_t); } -/// as write_short_decimal(), counting the digits (not 0, less than 10^15) -JSON_HEDLEY_NON_NULL(1) -JSON_HEDLEY_RETURNS_NON_NULL -inline char* write_short_decimal(char* first, std::uint64_t digits, int exp) noexcept -{ - JSON_ASSERT(digits != 0 && digits < 1000000000000000u); - // floor(log10(2^bits)) + 1 digits, or one less - const int log2_bound = ((64 - count_leading_zeros(digits)) * 1233) >> 12; - const int count = log2_bound + (digits >= powers_of_ten_16()[static_cast(log2_bound)] ? 1 : 0); - return write_short_decimal(first, digits, count, exp); -} +template +struct is_binary64 : std::integral_constant()> {}; -/// a positive finite float (other than double): Grisu2 and format_buffer() +/// a positive finite float (other than binary64): Grisu2 and format_buffer() template JSON_HEDLEY_NON_NULL(1, 2) JSON_HEDLEY_RETURNS_NON_NULL -char* write_positive(char* first, const char* last, FloatType value) +char* write_positive_grisu2(char* first, const char* last, FloatType value) { JSON_ASSERT(last - first >= std::numeric_limits::max_digits10); static_cast(last); // (only used in the assertion) @@ -1480,7 +1376,7 @@ char* write_positive(char* first, const char* last, FloatType value) // len is the length of the buffer, i.e., the number of decimal digits. int len = 0; int decimal_exponent = 0; - shortest_digits(first, len, decimal_exponent, value); + grisu2(first, len, decimal_exponent, value); JSON_ASSERT(len <= std::numeric_limits::max_digits10); @@ -1496,15 +1392,16 @@ char* write_positive(char* first, const char* last, FloatType value) return format_buffer(first, len, decimal_exponent, kMinExp, kMaxExp); } -/// a positive finite double: the shortest digits (Zmij), laid out by +/// a positive finite binary64 number: the shortest digits (Zmij), laid out by /// write_shortest() (through a local buffer if [first, last) is shorter than /// the 41 bytes it may write) +template JSON_HEDLEY_NON_NULL(1, 2) JSON_HEDLEY_RETURNS_NON_NULL -inline char* write_positive(char* first, const char* last, double value) +char* write_positive_zmij(char* first, const char* last, FloatType value) { - static_assert(std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53, - "internal error: the conversion of Zmij needs IEEE 754 binary64 doubles"); + static_assert(is_binary64::value, + "internal error: the conversion of Zmij needs IEEE 754 binary64 numbers"); std::uint64_t bits = 0; std::memcpy(&bits, &value, sizeof(bits)); const zmij::shortest_decimal d = zmij::to_shortest(bits); @@ -1519,6 +1416,34 @@ inline char* write_positive(char* first, const char* last, double value) return first + len; } +/// a positive finite binary64 number: Zmij (as a long double has the format of +/// a double here, its bits are those of the double of the same value) +template +JSON_HEDLEY_NON_NULL(1, 2) +JSON_HEDLEY_RETURNS_NON_NULL +char* write_positive(char* first, const char* last, FloatType value, std::true_type /*is_binary64*/) +{ + return write_positive_zmij(first, last, value); +} + +/// a positive finite float of any other format: Grisu2 +template +JSON_HEDLEY_NON_NULL(1, 2) +JSON_HEDLEY_RETURNS_NON_NULL +char* write_positive(char* first, const char* last, FloatType value, std::false_type /*is_binary64*/) +{ + return write_positive_grisu2(first, last, value); +} + +/// a positive finite float: Zmij for binary64 numbers, Grisu2 otherwise +template +JSON_HEDLEY_NON_NULL(1, 2) +JSON_HEDLEY_RETURNS_NON_NULL +char* write_positive(char* first, const char* last, FloatType value) +{ + return write_positive(first, last, value, is_binary64 {}); +} + } // namespace dtoa_impl /*! diff --git a/include/nlohmann/detail/conversions/zmij.hpp b/include/nlohmann/detail/conversions/zmij.hpp index ffff8380e..40e477f7b 100644 --- a/include/nlohmann/detail/conversions/zmij.hpp +++ b/include/nlohmann/detail/conversions/zmij.hpp @@ -36,13 +36,6 @@ computed from the compressed tables of Zmij beyond it. namespace zmij { -/// significand * 10^exponent -struct decimal -{ - std::uint64_t significand; - int exponent; -}; - /// the compressed powers of ten of Zmij inline const std::array& pow10_minor() noexcept { @@ -221,18 +214,6 @@ JSON_HEDLEY_ALWAYS_INLINE shortest_decimal to_shortest(std::uint64_t bits) noexc return shortest_decimal{integral, dec_exp, static_cast(digit), !round_up && !round_down}; } -/// The shortest decimal in the rounding interval of a positive finite double -/// given by its bits, as one number. The significand can end in zeros. -inline decimal to_decimal(std::uint64_t bits) noexcept -{ - const shortest_decimal d = to_shortest(bits); - if (d.has_digit) - { - return decimal{(d.integral * 10) + d.digit, d.exponent}; - } - return decimal{d.integral, d.exponent + 1}; -} - } // namespace zmij } // namespace detail NLOHMANN_JSON_NAMESPACE_END diff --git a/single_include/nlohmann/json.hpp b/single_include/nlohmann/json.hpp index e4fd9eeb8..839d1a2a7 100644 --- a/single_include/nlohmann/json.hpp +++ b/single_include/nlohmann/json.hpp @@ -8829,8 +8829,9 @@ NLOHMANN_JSON_NAMESPACE_END #include // uint64_t -#if !defined(__SIZEOF_INT128__) && defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) - #include // __umulh, _umul128 +#include // memcpy +#if defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) && (!defined(__SIZEOF_INT128__) || (!defined(__GNUC__) && !defined(__clang__))) + #include // __umulh, _umul128, _BitScanForward64, _BitScanReverse64 #endif // #include @@ -8849,6 +8850,10 @@ inline int count_leading_zeros(std::uint64_t x) noexcept { #if defined(__GNUC__) || defined(__clang__) return __builtin_clzll(x); +#elif defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) + unsigned long index = 0; + _BitScanReverse64(&index, x); + return 63 - static_cast(index); #else int n = 0; for (int shift = 32; shift != 0; shift >>= 1) @@ -8868,6 +8873,10 @@ inline int count_trailing_zeros(std::uint64_t x) noexcept { #if defined(__GNUC__) || defined(__clang__) return __builtin_ctzll(x); +#elif defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64)) + unsigned long index = 0; + _BitScanForward64(&index, x); + return static_cast(index); #else int n = 0; for (int shift = 32; shift != 0; shift >>= 1) @@ -8915,15 +8924,21 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc #endif } -/// eight bytes as a little-endian word (compilers fold this into one load on -/// little-endian targets; always inlined, as GCC otherwise calls it in the -/// number loops) +/// eight bytes as a little-endian word (a single load on little-endian +/// targets; always inlined, as GCC otherwise calls it in the number loops) JSON_HEDLEY_ALWAYS_INLINE std::uint64_t read_eight_bytes(const unsigned char* b) noexcept { +#if defined(_MSC_VER) || defined(__x86_64__) || defined(__i386__) || (defined(__BYTE_ORDER__) && defined(__ORDER_LITTLE_ENDIAN__) && __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__) + // the byte order already matches (all MSVC targets are little-endian) + std::uint64_t result = 0; + std::memcpy(&result, b, sizeof(result)); + return result; +#else return static_cast(b[0]) | (static_cast(b[1]) << 8u) | (static_cast(b[2]) << 16u) | (static_cast(b[3]) << 24u) | (static_cast(b[4]) << 32u) | (static_cast(b[5]) << 40u) | (static_cast(b[6]) << 48u) | (static_cast(b[7]) << 56u); +#endif } /// eight bytes as a little-endian word @@ -25117,6 +25132,7 @@ NLOHMANN_JSON_NAMESPACE_END // |_____|_____|_____|_|___| https://github.com/nlohmann/json // // SPDX-FileCopyrightText: 2009 Florian Loitsch +// SPDX-FileCopyrightText: 2025 Victor Zverovich // SPDX-FileCopyrightText: 2013-2026 Niels Lohmann // SPDX-License-Identifier: MIT @@ -25192,13 +25208,6 @@ computed from the compressed tables of Zmij beyond it. namespace zmij { -/// significand * 10^exponent -struct decimal -{ - std::uint64_t significand; - int exponent; -}; - /// the compressed powers of ten of Zmij inline const std::array& pow10_minor() noexcept { @@ -25377,18 +25386,6 @@ JSON_HEDLEY_ALWAYS_INLINE shortest_decimal to_shortest(std::uint64_t bits) noexc return shortest_decimal{integral, dec_exp, static_cast(digit), !round_up && !round_down}; } -/// The shortest decimal in the rounding interval of a positive finite double -/// given by its bits, as one number. The significand can end in zeros. -inline decimal to_decimal(std::uint64_t bits) noexcept -{ - const shortest_decimal d = to_shortest(bits); - if (d.has_digit) - { - return decimal{(d.integral * 10) + d.digit, d.exponent}; - } - return decimal{d.integral, d.exponent + 1}; -} - } // namespace zmij } // namespace detail NLOHMANN_JSON_NAMESPACE_END @@ -26296,88 +26293,6 @@ void grisu2(char* buf, int& len, int& decimal_exponent, FloatType value) grisu2(buf, len, decimal_exponent, w.minus, w.w, w.plus); } -/*! -@brief the shortest digits of a positive finite float (other than double): Grisu2 -*/ -template -JSON_HEDLEY_NON_NULL(1) -void shortest_digits(char* buf, int& len, int& decimal_exponent, FloatType value) -{ - grisu2(buf, len, decimal_exponent, value); -} - -/*! -@brief the shortest digits of a positive finite double: the conversion of -Zmij (see zmij.hpp), which always finds the shortest digits that read back as -the same value (Grisu2 does not for about one double in a thousand), and the -closest of them if there are several - -v = buf * 10^decimal_exponent, as for grisu2() -*/ -JSON_HEDLEY_NON_NULL(1) -inline void shortest_digits(char* buf, int& len, int& decimal_exponent, double value) -{ - static_assert(std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53, - "internal error: the conversion of Zmij needs IEEE 754 binary64 doubles"); - JSON_ASSERT(std::isfinite(value)); - JSON_ASSERT(value > 0); - - std::uint64_t bits = 0; - std::memcpy(&bits, &value, sizeof(bits)); - zmij::decimal d = zmij::to_decimal(bits); - // without trailing zeros (up to 16): 8, 4, 2, 1 at a time - while (d.significand % 100000000 == 0) - { - d.significand /= 100000000; - d.exponent += 8; - } - if (d.significand % 10000 == 0) - { - d.significand /= 10000; - d.exponent += 4; - } - if (d.significand % 100 == 0) - { - d.significand /= 100; - d.exponent += 2; - } - if (d.significand % 10 == 0) - { - d.significand /= 10; - d.exponent += 1; - } - // at most 17 digits, written from the back two at a time - static constexpr const char* pairs = - "00010203040506070809101112131415161718192021222324252627282930313233343536373839" - "40414243444546474849505152535455565758596061626364656667686970717273747576777879" - "8081828384858687888990919293949596979899"; - std::array digits{}; - std::size_t n = digits.size(); - while (d.significand >= 100) - { - const std::uint64_t two_digits = d.significand % 100; // a variable: GCC calls a cast of the remainder useless where std::uint64_t is std::size_t - const auto i = static_cast(two_digits) * 2; - d.significand /= 100; - n -= 2; - digits[n] = pairs[i]; - digits[n + 1] = pairs[i + 1]; - } - if (d.significand >= 10) - { - const auto i = static_cast(d.significand) * 2; - n -= 2; - digits[n] = pairs[i]; - digits[n + 1] = pairs[i + 1]; - } - else - { - digits[--n] = static_cast('0' + d.significand); - } - len = static_cast(digits.size() - n); - std::memcpy(buf, digits.data() + n, static_cast(len)); - decimal_exponent = d.exponent; -} - /*! @brief appends a decimal representation of e to buf @return a pointer to the element following the exponent. @@ -26780,53 +26695,30 @@ inline char* write_shortest(char* first, const zmij::shortest_decimal d) noexcep return end + (three ? 5 : 4); } -/// the powers of ten up to 10^16 -inline const std::array& powers_of_ten_16() noexcept -{ - static const std::array powers = - { - { - 1u, 10u, 100u, 1000u, 10000u, 100000u, 1000000u, 10000000u, 100000000u, 1000000000u, 10000000000u, - 100000000000u, 1000000000000u, 10000000000000u, 100000000000000u, 1000000000000000u, 10000000000000000u - } - }; - return powers; -} - /*! -@brief digits * 10^exp, as write_decimal() writes it, for the digits of a -double that need no conversion (count digits, at most 15, the first not 0; -trailing zeros allowed): extended to 16 digits and written by write_shortest() +@brief whether FloatType is an IEEE 754 binary64 type (a double, or a long double +that has the same format, as with MSVC and on Apple's Arm CPUs) -@return a pointer past the text; up to 41 bytes at @a first are written - (some beyond the returned end) +These are the types the conversion of Zmij (see zmij.hpp) is used for; all +others (binary32, or a format the library does not know) use Grisu2. */ -JSON_HEDLEY_NON_NULL(1) -JSON_HEDLEY_RETURNS_NON_NULL -inline char* write_short_decimal(char* first, std::uint64_t digits, int count, int exp) noexcept +template +constexpr bool has_binary64_format() noexcept { - JSON_ASSERT(digits >= powers_of_ten_16()[static_cast(count - 1)] && count <= 15); - const int scale = 16 - count; - return write_shortest(first, zmij::shortest_decimal{digits * powers_of_ten_16()[static_cast(scale)], exp - scale - 1, 0, false}); + return std::numeric_limits::is_iec559 + && std::numeric_limits::digits == 53 + && std::numeric_limits::max_exponent == 1024 + && sizeof(FloatType) == sizeof(std::uint64_t); } -/// as write_short_decimal(), counting the digits (not 0, less than 10^15) -JSON_HEDLEY_NON_NULL(1) -JSON_HEDLEY_RETURNS_NON_NULL -inline char* write_short_decimal(char* first, std::uint64_t digits, int exp) noexcept -{ - JSON_ASSERT(digits != 0 && digits < 1000000000000000u); - // floor(log10(2^bits)) + 1 digits, or one less - const int log2_bound = ((64 - count_leading_zeros(digits)) * 1233) >> 12; - const int count = log2_bound + (digits >= powers_of_ten_16()[static_cast(log2_bound)] ? 1 : 0); - return write_short_decimal(first, digits, count, exp); -} +template +struct is_binary64 : std::integral_constant()> {}; -/// a positive finite float (other than double): Grisu2 and format_buffer() +/// a positive finite float (other than binary64): Grisu2 and format_buffer() template JSON_HEDLEY_NON_NULL(1, 2) JSON_HEDLEY_RETURNS_NON_NULL -char* write_positive(char* first, const char* last, FloatType value) +char* write_positive_grisu2(char* first, const char* last, FloatType value) { JSON_ASSERT(last - first >= std::numeric_limits::max_digits10); static_cast(last); // (only used in the assertion) @@ -26837,7 +26729,7 @@ char* write_positive(char* first, const char* last, FloatType value) // len is the length of the buffer, i.e., the number of decimal digits. int len = 0; int decimal_exponent = 0; - shortest_digits(first, len, decimal_exponent, value); + grisu2(first, len, decimal_exponent, value); JSON_ASSERT(len <= std::numeric_limits::max_digits10); @@ -26853,15 +26745,16 @@ char* write_positive(char* first, const char* last, FloatType value) return format_buffer(first, len, decimal_exponent, kMinExp, kMaxExp); } -/// a positive finite double: the shortest digits (Zmij), laid out by +/// a positive finite binary64 number: the shortest digits (Zmij), laid out by /// write_shortest() (through a local buffer if [first, last) is shorter than /// the 41 bytes it may write) +template JSON_HEDLEY_NON_NULL(1, 2) JSON_HEDLEY_RETURNS_NON_NULL -inline char* write_positive(char* first, const char* last, double value) +char* write_positive_zmij(char* first, const char* last, FloatType value) { - static_assert(std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53, - "internal error: the conversion of Zmij needs IEEE 754 binary64 doubles"); + static_assert(is_binary64::value, + "internal error: the conversion of Zmij needs IEEE 754 binary64 numbers"); std::uint64_t bits = 0; std::memcpy(&bits, &value, sizeof(bits)); const zmij::shortest_decimal d = zmij::to_shortest(bits); @@ -26876,6 +26769,34 @@ inline char* write_positive(char* first, const char* last, double value) return first + len; } +/// a positive finite binary64 number: Zmij (as a long double has the format of +/// a double here, its bits are those of the double of the same value) +template +JSON_HEDLEY_NON_NULL(1, 2) +JSON_HEDLEY_RETURNS_NON_NULL +char* write_positive(char* first, const char* last, FloatType value, std::true_type /*is_binary64*/) +{ + return write_positive_zmij(first, last, value); +} + +/// a positive finite float of any other format: Grisu2 +template +JSON_HEDLEY_NON_NULL(1, 2) +JSON_HEDLEY_RETURNS_NON_NULL +char* write_positive(char* first, const char* last, FloatType value, std::false_type /*is_binary64*/) +{ + return write_positive_grisu2(first, last, value); +} + +/// a positive finite float: Zmij for binary64 numbers, Grisu2 otherwise +template +JSON_HEDLEY_NON_NULL(1, 2) +JSON_HEDLEY_RETURNS_NON_NULL +char* write_positive(char* first, const char* last, FloatType value) +{ + return write_positive(first, last, value, is_binary64 {}); +} + } // namespace dtoa_impl /*! diff --git a/tests/src/unit-class_lexer.cpp b/tests/src/unit-class_lexer.cpp index 82690a669..e5f82b53c 100644 --- a/tests/src/unit-class_lexer.cpp +++ b/tests/src/unit-class_lexer.cpp @@ -1792,4 +1792,11 @@ TEST_CASE("string scanning kernels") CHECK(nlohmann::detail::count_trailing_zeros(bit) == k); CHECK(nlohmann::detail::count_trailing_zeros(bit | (bit << 1u) | 0x8000000000000000u) == k); } + + // eight bytes as a little-endian word, at any alignment + const unsigned char bytes[16] = {0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08, 0x09, 0x0A, 0x0B, 0x0C, 0x0D, 0x0E, 0x0F, 0xFF}; + CHECK(nlohmann::detail::read_eight_bytes(bytes) == 0x0807060504030201u); + CHECK(nlohmann::detail::read_eight_bytes(bytes + 1) == 0x0908070605040302u); + CHECK(nlohmann::detail::read_eight_bytes(bytes + 8) == 0xFF0F0E0D0C0B0A09u); + CHECK(nlohmann::detail::read_eight_bytes(reinterpret_cast(bytes) + 3) == 0x0B0A090807060504u); } diff --git a/tests/src/unit-to_chars.cpp b/tests/src/unit-to_chars.cpp index 8ced43ad0..d8eb44434 100644 --- a/tests/src/unit-to_chars.cpp +++ b/tests/src/unit-to_chars.cpp @@ -15,6 +15,7 @@ #include using nlohmann::detail::dtoa_impl::reinterpret_bits; +#include #include #include #include @@ -666,13 +667,24 @@ void check_shortest(double v) const std::string text(buf.data(), end); CAPTURE(text) CHECK(parse_double(text) == v); - // the layout is that of format_buffer() for the same digits + // the layout is that of format_buffer() for the digits of Zmij + const auto sd = nlohmann::detail::zmij::to_shortest(reinterpret_bits(v)); + const std::uint64_t significand = sd.has_digit ? (sd.integral * 10) + sd.digit : sd.integral; + int exponent = sd.has_digit ? sd.exponent : sd.exponent + 1; + std::string significand_digits = std::to_string(significand); + while (significand_digits.size() > 1 && significand_digits.back() == '0') + { + significand_digits.pop_back(); + ++exponent; + } std::array reference{}; - int len = 0; - int exponent = 0; - nlohmann::detail::dtoa_impl::shortest_digits(reference.data(), len, exponent, v); - const char* const reference_end = nlohmann::detail::dtoa_impl::format_buffer(reference.data(), len, exponent, -4, 15); + std::copy(significand_digits.begin(), significand_digits.end(), reference.begin()); + const char* const reference_end = nlohmann::detail::dtoa_impl::format_buffer(reference.data(), static_cast(significand_digits.size()), exponent, -4, 15); CHECK(text == std::string(reference.data(), static_cast(reference_end - reference.data()))); + // and write_positive() is what to_chars() calls + std::array positive{}; + const char* const positive_end = nlohmann::detail::dtoa_impl::write_positive(positive.data(), positive.data() + positive.size(), v); + CHECK(text == std::string(positive.data(), static_cast(positive_end - positive.data()))); const auto de = digits_and_exponent(text); const std::string& digits = de.first; if (digits.size() > 1) @@ -785,3 +797,58 @@ TEST_CASE("shortest digits of doubles") } } } + +TEST_CASE("choice of the conversion") +{ + using nlohmann::detail::dtoa_impl::is_binary64; + + SECTION("by the format of the type") + { + // Zmij needs binary64 numbers; everything else uses Grisu2 + static_assert(!is_binary64::value, "float is not binary64"); + static_assert(is_binary64::value == (std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53), + "double is binary64 where it is IEEE 754 with 53 digits"); + static_assert(!is_binary64::value, "integers are not binary64"); + static_assert(is_binary64::value == (std::numeric_limits::is_iec559 && std::numeric_limits::digits == 53 && sizeof(long double) == 8), + "long double is binary64 where it has the format of a double"); + CHECK(!is_binary64::value); + CHECK(is_binary64::value); + } + + SECTION("float: Grisu2, double: Zmij") + { + // 5.3165205877497296e+16 is one of the doubles for which Grisu2 does not find the shortest digits + constexpr double value = 5.3165205877497296e+16; + std::array buf{}; + const char* const last = buf.data() + buf.size(); + + char* end = nlohmann::detail::dtoa_impl::write_positive(buf.data(), last, value); + CHECK(std::string(buf.data(), end) == "5.31652058774973e+16"); + end = nlohmann::detail::dtoa_impl::write_positive_grisu2(buf.data(), last, value); + CHECK(std::string(buf.data(), end) == "5.3165205877497296e+16"); + + constexpr float f = 1.1754944e-38f; + end = nlohmann::detail::dtoa_impl::write_positive(buf.data(), last, f); + const std::string dispatched(buf.data(), end); + end = nlohmann::detail::dtoa_impl::write_positive_grisu2(buf.data(), last, f); + CHECK(dispatched == std::string(buf.data(), end)); + } + + SECTION("long double with the format of a double: Zmij") + { + // (on platforms where long double is wider, Grisu2 does not apply either: the snprintf fallback does) + if (std::numeric_limits::digits == 53 && std::numeric_limits::is_iec559) + { + using long_double_json = nlohmann::json::with_float_t; + for (const double d : + { + 5.3165205877497296e+16, 1.0, 0.1, 123456.789, 2.2250738585072014e-308, 1.7976931348623157e+308, -5.3165205877497296e+16 + }) + { + CAPTURE(d) + CHECK(long_double_json(static_cast(d)).dump() == nlohmann::json(d).dump()); + } + CHECK(long_double_json(5.3165205877497296e+16L).dump() == "5.31652058774973e+16"); + } + } +}