Compare commits

...
Author SHA1 Message Date
Niels Lohmann f1027d8f38 Clarify hash documentation after review
- Say "may hash differently" for null, false, and numbers, since a
  collision across types is possible.
- Name the storage types (signed integer, unsigned integer,
  floating-point number) instead of example literals.
- Explain that the hash survives converting an integer to
  number_float_t but not the lossy conversion back, and that unequal
  numbers may share a hash.
- State that the example hash values are illustrative only and vary by
  platform, compiler, compiler version, and library version.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:17 +02:00
Niels Lohmann cf67f8c57a Normalize -0.0 in number hashes and test range ends
operator== compares numbers exactly since #5459, so equal numbers
share one value and convert to the same number_float_t. Update the
comment accordingly, map -0.0 to 0.0 before hashing (std::hash need
not do that), and test -0.0 and the ends of the integer ranges. Show
hash(0.0) in the docs example and note the change in the version
history.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:15 +02:00
Afonso Januário deec1d0a67 Mark the unordered_set in the hash regression test const
clang-tidy's misc-const-correctness check flagged it: the set is
never mutated after construction, only read via size().

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:13 +02:00
Afonso Januário d223021995 Remove now-unused number_integer_t/number_unsigned_t typedefs in hash()
Merging the three numeric branches into one that only reads
number_float_t left these two aliases unused, which several CI
configurations treat as a build error under -Wunused-local-typedefs.

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:11 +02:00
Afonso Januário c66be708eb Make std::hash<basic_json> consistent with operator== for numbers
operator== converts between number_integer, number_unsigned, and
number_float before comparing, so json(0), json(0U), and json(0.0)
all compare equal. hash() folded the specific value_t into the
result for each of the three numeric cases, giving each a distinct
hash and breaking the standard Hash requirement that a == b implies
hash(a) == hash(b). A std::unordered_set could therefore hold all
three as separate elements even though they compare equal.

hash() now treats all three numeric variants the same way: it
converts the value to number_float_t and combines it with a single
shared type tag, so any two numbers operator== considers equal hash
identically regardless of which internal type actually holds them.

Updated the accompanying test to check this consistency directly
(including via an actual unordered_set) instead of asserting that 0,
0U, and 0.0 hash differently, since that assumption was the bug.
Also corrected the function's own doc comment and the std::hash API
docs, which described the old behavior as intended.

Fixes #5400

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:08 +02:00
6 changed files with 95 additions and 47 deletions

No files matched your search

+12 -3
View File
@@ -7,8 +7,14 @@ namespace std {
``` ```
Return a hash value for a JSON object. The hash function tries to rely on `std::hash` where possible. Furthermore, the Return a hash value for a JSON object. The hash function tries to rely on `std::hash` where possible. Furthermore, the
type of the JSON value is taken into account to have different hash values for `#!json null`, `#!cpp 0`, `#!cpp 0U`, and type of the JSON value is taken into account, so `#!json null`, `#!cpp false`, and numbers may hash differently from
`#!cpp false`, etc. each other. Numbers that compare equal under [`operator==`](operator_eq.md) always hash equally, regardless of
whether they are stored as signed integer, unsigned integer, or floating-point number.
Numbers are hashed by their value converted to `number_float_t`. Converting an integer to `number_float_t` therefore
keeps its hash, but converting a floating-point number to an integer type is lossy and can change it: `#!cpp 0.5`
converts to `#!cpp 0`, which need not have the same hash. Unequal numbers can also share a hash value, for example two
large integers that convert to the same `number_float_t`.
## Examples ## Examples
@@ -26,7 +32,8 @@ type of the JSON value is taken into account to have different hash values for `
--8<-- "examples/std_hash.output" --8<-- "examples/std_hash.output"
``` ```
Note the output is platform-dependent. The hash values shown are examples only. They depend on the platform, the compiler, and the compiler version, and
they can change between versions of this library. Do not persist them or rely on specific values.
## See also ## See also
@@ -36,3 +43,5 @@ type of the JSON value is taken into account to have different hash values for `
- Added in version 1.0.0. - Added in version 1.0.0.
- Extended for arbitrary basic_json types in version 3.10.5. - Extended for arbitrary basic_json types in version 3.10.5.
- Numbers that compare equal hash equally since version 3.13.0; before, `#!cpp 0`, `#!cpp 0U`, and `#!cpp 0.0` had
different hash values.
+1
View File
@@ -11,6 +11,7 @@ int main()
<< "hash(false) = " << std::hash<json> {}(json(false)) << '\n' << "hash(false) = " << std::hash<json> {}(json(false)) << '\n'
<< "hash(0) = " << std::hash<json> {}(json(0)) << '\n' << "hash(0) = " << std::hash<json> {}(json(0)) << '\n'
<< "hash(0U) = " << std::hash<json> {}(json(0U)) << '\n' << "hash(0U) = " << std::hash<json> {}(json(0U)) << '\n'
<< "hash(0.0) = " << std::hash<json> {}(json(0.0)) << '\n'
<< "hash(\"\") = " << std::hash<json> {}(json("")) << '\n' << "hash(\"\") = " << std::hash<json> {}(json("")) << '\n'
<< "hash({}) = " << std::hash<json> {}(json::object()) << '\n' << "hash({}) = " << std::hash<json> {}(json::object()) << '\n'
<< "hash([]) = " << std::hash<json> {}(json::array()) << '\n' << "hash([]) = " << std::hash<json> {}(json::array()) << '\n'
+5 -4
View File
@@ -1,8 +1,9 @@
hash(null) = 2654435769 hash(null) = 2654435769
hash(false) = 2654436030 hash(false) = 2654436030
hash(0) = 2654436095 hash(0) = 2654436221
hash(0U) = 2654436156 hash(0U) = 2654436221
hash("") = 6142509191626859748 hash(0.0) = 2654436221
hash("") = 11160318156688833227
hash({}) = 2654435832 hash({}) = 2654435832
hash([]) = 2654435899 hash([]) = 2654435899
hash({"hello": "world"}) = 4469488738203676328 hash({"hello": "world"}) = 3701319991624763853
+19 -16
View File
@@ -35,8 +35,10 @@ std::size_t hash_iteratively(const BasicJsonType& j);
@brief hash a JSON value @brief hash a JSON value
The hash function tries to rely on std::hash where possible. Furthermore, the The hash function tries to rely on std::hash where possible. Furthermore, the
type of the JSON value is taken into account to have different hash values for type of the JSON value is taken into account, so null, false, and numbers may
null, 0, 0U, and false, etc. hash differently from each other, but any two numbers that compare equal
under operator== hash equally regardless of which of number_integer,
number_unsigned, or number_float actually holds the value.
Hashing an array or an object hashes its elements, which used to call this Hashing an array or an object hashes its elements, which used to call this
function again once per nesting level, so a value nested deeply enough function again once per nesting level, so a value nested deeply enough
@@ -55,8 +57,6 @@ template<typename BasicJsonType>
std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0) std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
{ {
using string_t = typename BasicJsonType::string_t; using string_t = typename BasicJsonType::string_t;
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t; using number_float_t = typename BasicJsonType::number_float_t;
const auto type = static_cast<std::size_t>(j.type()); const auto type = static_cast<std::size_t>(j.type());
@@ -113,21 +113,24 @@ std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
} }
case BasicJsonType::value_t::number_integer: case BasicJsonType::value_t::number_integer:
{
const auto h = std::hash<number_integer_t> {}(j.template get<number_integer_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_unsigned: case BasicJsonType::value_t::number_unsigned:
{
const auto h = std::hash<number_unsigned_t> {}(j.template get<number_unsigned_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_float: case BasicJsonType::value_t::number_float:
{ {
const auto h = std::hash<number_float_t> {}(j.template get<number_float_t>()); // operator== compares numbers by their mathematical value across
return combine(type, h); // number_integer, number_unsigned, and number_float, so equal
// numbers of different internal types (0, 0U, 0.0) must hash the
// same. Two equal numbers have the same value, which converts to
// the same number_float_t, so all numbers share one type tag and
// hash that converted value. Adding zero turns -0.0 (equal to 0)
// into 0.0, as std::hash need not map both to the same hash.
// The converse does not hold: converting a number_float_t value
// to an integer type is lossy, so the result can hash
// differently, and unequal numbers that convert to the same
// number_float_t (e.g., 2^53 and 2^53 + 1) share a hash.
const auto number_type = static_cast<std::size_t>(BasicJsonType::value_t::number_float);
const auto value = j.template get<number_float_t>() + static_cast<number_float_t>(0);
const auto h = std::hash<number_float_t> {}(value);
return combine(number_type, h);
} }
case BasicJsonType::value_t::binary: case BasicJsonType::value_t::binary:
+19 -16
View File
@@ -7583,8 +7583,10 @@ std::size_t hash_iteratively(const BasicJsonType& j);
@brief hash a JSON value @brief hash a JSON value
The hash function tries to rely on std::hash where possible. Furthermore, the The hash function tries to rely on std::hash where possible. Furthermore, the
type of the JSON value is taken into account to have different hash values for type of the JSON value is taken into account, so null, false, and numbers may
null, 0, 0U, and false, etc. hash differently from each other, but any two numbers that compare equal
under operator== hash equally regardless of which of number_integer,
number_unsigned, or number_float actually holds the value.
Hashing an array or an object hashes its elements, which used to call this Hashing an array or an object hashes its elements, which used to call this
function again once per nesting level, so a value nested deeply enough function again once per nesting level, so a value nested deeply enough
@@ -7603,8 +7605,6 @@ template<typename BasicJsonType>
std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0) std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
{ {
using string_t = typename BasicJsonType::string_t; using string_t = typename BasicJsonType::string_t;
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t; using number_float_t = typename BasicJsonType::number_float_t;
const auto type = static_cast<std::size_t>(j.type()); const auto type = static_cast<std::size_t>(j.type());
@@ -7661,21 +7661,24 @@ std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
} }
case BasicJsonType::value_t::number_integer: case BasicJsonType::value_t::number_integer:
{
const auto h = std::hash<number_integer_t> {}(j.template get<number_integer_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_unsigned: case BasicJsonType::value_t::number_unsigned:
{
const auto h = std::hash<number_unsigned_t> {}(j.template get<number_unsigned_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_float: case BasicJsonType::value_t::number_float:
{ {
const auto h = std::hash<number_float_t> {}(j.template get<number_float_t>()); // operator== compares numbers by their mathematical value across
return combine(type, h); // number_integer, number_unsigned, and number_float, so equal
// numbers of different internal types (0, 0U, 0.0) must hash the
// same. Two equal numbers have the same value, which converts to
// the same number_float_t, so all numbers share one type tag and
// hash that converted value. Adding zero turns -0.0 (equal to 0)
// into 0.0, as std::hash need not map both to the same hash.
// The converse does not hold: converting a number_float_t value
// to an integer type is lossy, so the result can hash
// differently, and unequal numbers that convert to the same
// number_float_t (e.g., 2^53 and 2^53 + 1) share a hash.
const auto number_type = static_cast<std::size_t>(BasicJsonType::value_t::number_float);
const auto value = j.template get<number_float_t>() + static_cast<number_float_t>(0);
const auto h = std::hash<number_float_t> {}(value);
return combine(number_type, h);
} }
case BasicJsonType::value_t::binary: case BasicJsonType::value_t::binary:
+39 -8
View File
@@ -12,8 +12,10 @@
using json = nlohmann::json; using json = nlohmann::json;
using ordered_json = nlohmann::ordered_json; using ordered_json = nlohmann::ordered_json;
#include <limits>
#include <set> #include <set>
#include <string> #include <string>
#include <unordered_set>
namespace namespace
{ {
@@ -91,6 +93,9 @@ TEST_CASE("hash<nlohmann::json>")
// Collect hashes for different JSON values and make sure that they are distinct // Collect hashes for different JSON values and make sure that they are distinct
// We cannot compare against fixed values, because the implementation of // We cannot compare against fixed values, because the implementation of
// std::hash may differ between compilers. // std::hash may differ between compilers.
//
// numbers that compare equal under operator== (0 == 0U == 0.0) must hash
// equally, so they are only inserted once below and checked separately.
std::set<std::size_t> hashes; std::set<std::size_t> hashes;
@@ -107,10 +112,7 @@ TEST_CASE("hash<nlohmann::json>")
// number // number
hashes.insert(std::hash<json> {}(json(0))); hashes.insert(std::hash<json> {}(json(0)));
hashes.insert(std::hash<json> {}(json(static_cast<unsigned>(0))));
hashes.insert(std::hash<json> {}(json(-1))); hashes.insert(std::hash<json> {}(json(-1)));
hashes.insert(std::hash<json> {}(json(0.0)));
hashes.insert(std::hash<json> {}(json(42.23))); hashes.insert(std::hash<json> {}(json(42.23)));
// array // array
@@ -132,7 +134,36 @@ TEST_CASE("hash<nlohmann::json>")
// discarded // discarded
hashes.insert(std::hash<json> {}(json(json::value_t::discarded))); hashes.insert(std::hash<json> {}(json(json::value_t::discarded)));
CHECK(hashes.size() == 21); CHECK(hashes.size() == 19);
// numbers that compare equal under operator== must hash equally,
// regardless of which of number_integer, number_unsigned, or
// number_float actually holds the value
CHECK(json(0) == json(static_cast<unsigned>(0)));
CHECK(json(0) == json(0.0));
CHECK(std::hash<json> {}(json(0)) == std::hash<json> {}(json(static_cast<unsigned>(0))));
CHECK(std::hash<json> {}(json(0)) == std::hash<json> {}(json(0.0)));
CHECK(std::hash<json> {}(json(-1)) == std::hash<json> {}(json(-1.0)));
// a std::unordered_set relies on this same consistency between == and hash
const std::unordered_set<json> numbers {json(0), json(static_cast<unsigned>(0)), json(0.0)};
CHECK(numbers.size() == 1);
// -0.0 compares equal to 0 and 0.0
CHECK(json(-0.0) == json(0));
CHECK(std::hash<json> {}(json(-0.0)) == std::hash<json> {}(json(0)));
CHECK(std::hash<json> {}(json(-0.0)) == std::hash<json> {}(json(0.0)));
// the ends of the integer ranges, which equal floats exactly
const auto int_min = (std::numeric_limits<json::number_integer_t>::min)();
const auto int_max = (std::numeric_limits<json::number_integer_t>::max)();
const auto two_63 = json::number_unsigned_t(1) << 63U;
CHECK(json(int_min) == json(-9223372036854775808.0));
CHECK(std::hash<json> {}(json(int_min)) == std::hash<json> {}(json(-9223372036854775808.0)));
CHECK(json(two_63) == json(9223372036854775808.0));
CHECK(std::hash<json> {}(json(two_63)) == std::hash<json> {}(json(9223372036854775808.0)));
CHECK(json(json::number_unsigned_t(int_max)) == json(int_max));
CHECK(std::hash<json> {}(json(json::number_unsigned_t(int_max))) == std::hash<json> {}(json(int_max)));
} }
TEST_CASE("hash<nlohmann::ordered_json>") TEST_CASE("hash<nlohmann::ordered_json>")
@@ -156,10 +187,7 @@ TEST_CASE("hash<nlohmann::ordered_json>")
// number // number
hashes.insert(std::hash<ordered_json> {}(ordered_json(0))); hashes.insert(std::hash<ordered_json> {}(ordered_json(0)));
hashes.insert(std::hash<ordered_json> {}(ordered_json(static_cast<unsigned>(0))));
hashes.insert(std::hash<ordered_json> {}(ordered_json(-1))); hashes.insert(std::hash<ordered_json> {}(ordered_json(-1)));
hashes.insert(std::hash<ordered_json> {}(ordered_json(0.0)));
hashes.insert(std::hash<ordered_json> {}(ordered_json(42.23))); hashes.insert(std::hash<ordered_json> {}(ordered_json(42.23)));
// array // array
@@ -181,7 +209,10 @@ TEST_CASE("hash<nlohmann::ordered_json>")
// discarded // discarded
hashes.insert(std::hash<ordered_json> {}(ordered_json(ordered_json::value_t::discarded))); hashes.insert(std::hash<ordered_json> {}(ordered_json(ordered_json::value_t::discarded)));
CHECK(hashes.size() == 21); CHECK(hashes.size() == 19);
CHECK(std::hash<ordered_json> {}(ordered_json(0)) == std::hash<ordered_json> {}(ordered_json(static_cast<unsigned>(0))));
CHECK(std::hash<ordered_json> {}(ordered_json(0)) == std::hash<ordered_json> {}(ordered_json(0.0)));
} }
TEST_CASE("hash of deeply nested values") TEST_CASE("hash of deeply nested values")