Compare commits

..
Author SHA1 Message Date
Suyog Verma a269794db7 Use MSVC intrinsics for full multiplication (#5782)
* Use MSVC intrinsics for full multiplication

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

* Fix formatting in unit-class_lexer

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

* Address review feedback

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>

---------

Signed-off-by: Suyog Verma <suyogverma0057@gmail.com>
2026-10-08 08:49:32 +02:00
4 changed files with 65 additions and 30 deletions

No files matched your search

+9 -23
View File
@@ -36,26 +36,13 @@ work items are tracked in the [GitHub milestones](https://github.com/nlohmann/js
## API stability
Releases follow [semantic versioning](https://semver.org): a minor or patch release of version 3.x does not break code
that uses the public API, unless that code opts in to a change with a macro as described [below](#version-40). In
particular, a 3.x release does not:
that uses the public API. In particular, a 3.x release does not:
- make breaking changes to the signature of a function: the types or order of its existing parameters, its return type,
its `noexcept` or `constexpr` specifier, or the const-ness of a member function. New parameters may be added if they
have a default value;
- remove or rename a function or class, or change the template parameters of a public class template;
- change the signature of a function (its parameter types, return type, number of parameters, or the const-ness of a
member function);
- remove or rename a function or class;
- change which exceptions a function throws, or the [exception ids](../home/exceptions.md);
- change access specifiers, or change or remove existing default arguments. New default arguments may be added;
- change the JSON type that a valid input parses to, or the text that `dump()` produces for a valid value;
- accept input that was rejected before, or reject input that was accepted before;
- change the order in which the keys of an object are iterated. The default type sorts keys, and
[`ordered_json`](../api/ordered_json.md) keeps insertion order;
- change when iterators, pointers, or references are invalidated, or the state of a moved-from `basic_json`;
- add or remove implicit conversions from `basic_json`;
- change how `to_json` and `from_json` functions are found, or the behavior of
[`adl_serializer`](../api/adl_serializer/index.md);
- add pure virtual functions to the [`json_sax`](../api/json_sax/index.md) interface;
- remove, rename, renumber, or add enumerators of `value_t`;
- remove or rename a documented macro, CMake option, CMake target, or header, or change what a documented macro does.
- change access specifiers or default arguments.
Exceptions to these rules, for instance when fixing a bug requires changing the exception a function throws, are
documented in the [release notes](../home/releases.md).
@@ -64,14 +51,13 @@ The following are **not** part of the public API and may change in any release,
- The text of exception messages returned by `what()`. Use the [exception id](../home/exceptions.md) to tell errors
apart.
- The ABI, including `sizeof(basic_json)` and the memory layout of its values. The
[versioned inline namespace](../features/namespace.md) turns mixing versions into a link error.
- The hash values returned by `std::hash` for `basic_json`. Numbers that compare equal still hash equally.
- The ABI, including `sizeof(basic_json)` and the memory layout of its values. Recompile your code when you upgrade the
library. The [versioned inline namespace](../features/namespace.md) turns mixing versions into a link error.
- Everything in namespace `nlohmann::detail`, and macros and type traits that are not documented in the
[API reference](../api/basic_json/index.md).
Breaking changes are only added behind a macro whose default keeps the 3.x behavior. See [Version 4.0](#version-40) and
the [macro overview](../features/macros.md).
Changes that would break the public API are only added behind a macro whose default keeps the 3.x behavior, see
[Version 4.0](#version-40).
## Version 4.0
+11 -2
View File
@@ -9,12 +9,15 @@
#pragma once
#include <cstdint> // uint64_t
#if !defined(__SIZEOF_INT128__) && defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64))
#include <intrin0.h> // __umulh, _umul128
#endif
#include <nlohmann/detail/abi_macros.hpp>
// Portable bit-level helpers for the number and string scanners. They use
// compiler builtins where available and plain C++ otherwise, so they need no
// platform headers and work regardless of byte order.
// compiler builtins or platform-specific intrinsics where available and plain
// C++ otherwise, so they work regardless of byte order.
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -52,6 +55,12 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc
__extension__ using uint128 = unsigned __int128;
const uint128 r = static_cast<uint128>(a) * b;
return {static_cast<std::uint64_t>(r), static_cast<std::uint64_t>(r >> 64u)};
#elif defined(_MSC_VER) && defined(_M_X64)
std::uint64_t high = 0;
const std::uint64_t low = _umul128(a, b, &high);
return {low, high};
#elif defined(_MSC_VER) && defined(_M_ARM64)
return {a * b, __umulh(a, b)};
#else
const std::uint64_t a_lo = a & 0xFFFFFFFFu;
const std::uint64_t a_hi = a >> 32u;
+11 -2
View File
@@ -8791,13 +8791,16 @@ NLOHMANN_JSON_NAMESPACE_END
#include <cstdint> // uint64_t
#if !defined(__SIZEOF_INT128__) && defined(_MSC_VER) && (defined(_M_X64) || defined(_M_ARM64))
#include <intrin0.h> // __umulh, _umul128
#endif
// #include <nlohmann/detail/abi_macros.hpp>
// Portable bit-level helpers for the number and string scanners. They use
// compiler builtins where available and plain C++ otherwise, so they need no
// platform headers and work regardless of byte order.
// compiler builtins or platform-specific intrinsics where available and plain
// C++ otherwise, so they work regardless of byte order.
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -8835,6 +8838,12 @@ inline uint128_parts full_multiplication(std::uint64_t a, std::uint64_t b) noexc
__extension__ using uint128 = unsigned __int128;
const uint128 r = static_cast<uint128>(a) * b;
return {static_cast<std::uint64_t>(r), static_cast<std::uint64_t>(r >> 64u)};
#elif defined(_MSC_VER) && defined(_M_X64)
std::uint64_t high = 0;
const std::uint64_t low = _umul128(a, b, &high);
return {low, high};
#elif defined(_MSC_VER) && defined(_M_ARM64)
return {a * b, __umulh(a, b)};
#else
const std::uint64_t a_lo = a & 0xFFFFFFFFu;
const std::uint64_t a_hi = a >> 32u;
+34 -3
View File
@@ -17,6 +17,7 @@ using nlohmann::json;
#include <cstdint> // uint32_t, uint64_t
#include <cstdlib> // strtod
#include <cstring> // memcpy
#include <limits> // numeric_limits
#include <sstream> // stringstream
#include <string> // string
#include <utility> // pair
@@ -891,8 +892,39 @@ TEST_CASE("Eisel-Lemire float conversion")
SECTION("128-bit products and leading zeros")
{
const auto check_product = [](std::uint64_t a, std::uint64_t b)
{
const auto product = nlohmann::detail::full_multiplication(a, b);
CHECK(big_from(product.high, product.low) == big_mul(big_from(0, a), big_from(0, b)));
};
const std::uint64_t max = (std::numeric_limits<std::uint64_t>::max)();
const std::array<std::pair<std::uint64_t, std::uint64_t>, 13> edge_cases =
{
{
{0, 0},
{0, 1},
{1, 1},
{1, max},
{0xFFFFFFFFu, 0x100000000u},
{0x100000000u, 0x100000000u},
{0x100000001u, 0x100000001u},
{max, max},
{max, 2},
{0xFFFFFFFF00000000u, 0x100000001u},
{0x100000001u, 0xFFFFFFFF00000000u},
{max, 1},
{2, max},
}
};
for (const auto& test : edge_cases)
{
check_product(test.first, test.second);
}
// whichever implementation the compiler gets (with or without a
// 128-bit integer type or a builtin)
// 128-bit integer type or a builtin / intrinsic)
std::uint64_t state = 42;
for (int i = 0; i < 10000; ++i)
{
@@ -901,8 +933,7 @@ TEST_CASE("Eisel-Lemire float conversion")
state ^= state << 17u;
const std::uint64_t a = state;
const std::uint64_t b = (state * 0x9E3779B97F4A7C15u) >> (i % 64);
const auto product = nlohmann::detail::full_multiplication(a, b);
CHECK(big_from(product.high, product.low) == big_mul(big_from(0, a), big_from(0, b)));
check_product(a, b);
const int k = i % 64;
const std::uint64_t x = (std::uint64_t{1} << k) | (a & ((std::uint64_t{1} << k) - 1));