Compare commits

..
Author SHA1 Message Date
Niels Lohmann f1027d8f38 Clarify hash documentation after review
- Say "may hash differently" for null, false, and numbers, since a
  collision across types is possible.
- Name the storage types (signed integer, unsigned integer,
  floating-point number) instead of example literals.
- Explain that the hash survives converting an integer to
  number_float_t but not the lossy conversion back, and that unequal
  numbers may share a hash.
- State that the example hash values are illustrative only and vary by
  platform, compiler, compiler version, and library version.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:17 +02:00
Niels Lohmann cf67f8c57a Normalize -0.0 in number hashes and test range ends
operator== compares numbers exactly since #5459, so equal numbers
share one value and convert to the same number_float_t. Update the
comment accordingly, map -0.0 to 0.0 before hashing (std::hash need
not do that), and test -0.0 and the ends of the integer ranges. Show
hash(0.0) in the docs example and note the change in the version
history.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:15 +02:00
Afonso Januário deec1d0a67 Mark the unordered_set in the hash regression test const
clang-tidy's misc-const-correctness check flagged it: the set is
never mutated after construction, only read via size().

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:13 +02:00
Afonso Januário d223021995 Remove now-unused number_integer_t/number_unsigned_t typedefs in hash()
Merging the three numeric branches into one that only reads
number_float_t left these two aliases unused, which several CI
configurations treat as a build error under -Wunused-local-typedefs.

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:11 +02:00
Afonso Januário c66be708eb Make std::hash<basic_json> consistent with operator== for numbers
operator== converts between number_integer, number_unsigned, and
number_float before comparing, so json(0), json(0U), and json(0.0)
all compare equal. hash() folded the specific value_t into the
result for each of the three numeric cases, giving each a distinct
hash and breaking the standard Hash requirement that a == b implies
hash(a) == hash(b). A std::unordered_set could therefore hold all
three as separate elements even though they compare equal.

hash() now treats all three numeric variants the same way: it
converts the value to number_float_t and combines it with a single
shared type tag, so any two numbers operator== considers equal hash
identically regardless of which internal type actually holds them.

Updated the accompanying test to check this consistency directly
(including via an actual unordered_set) instead of asserting that 0,
0U, and 0.0 hash differently, since that assumption was the bug.
Also corrected the function's own doc comment and the std::hash API
docs, which described the old behavior as intended.

Fixes #5400

Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:40:08 +02:00
8 changed files with 343 additions and 1381 deletions

No files matched your search

+12 -3
View File
@@ -7,8 +7,14 @@ namespace std {
``` ```
Return a hash value for a JSON object. The hash function tries to rely on `std::hash` where possible. Furthermore, the Return a hash value for a JSON object. The hash function tries to rely on `std::hash` where possible. Furthermore, the
type of the JSON value is taken into account to have different hash values for `#!json null`, `#!cpp 0`, `#!cpp 0U`, and type of the JSON value is taken into account, so `#!json null`, `#!cpp false`, and numbers may hash differently from
`#!cpp false`, etc. each other. Numbers that compare equal under [`operator==`](operator_eq.md) always hash equally, regardless of
whether they are stored as signed integer, unsigned integer, or floating-point number.
Numbers are hashed by their value converted to `number_float_t`. Converting an integer to `number_float_t` therefore
keeps its hash, but converting a floating-point number to an integer type is lossy and can change it: `#!cpp 0.5`
converts to `#!cpp 0`, which need not have the same hash. Unequal numbers can also share a hash value, for example two
large integers that convert to the same `number_float_t`.
## Examples ## Examples
@@ -26,7 +32,8 @@ type of the JSON value is taken into account to have different hash values for `
--8<-- "examples/std_hash.output" --8<-- "examples/std_hash.output"
``` ```
Note the output is platform-dependent. The hash values shown are examples only. They depend on the platform, the compiler, and the compiler version, and
they can change between versions of this library. Do not persist them or rely on specific values.
## See also ## See also
@@ -36,3 +43,5 @@ type of the JSON value is taken into account to have different hash values for `
- Added in version 1.0.0. - Added in version 1.0.0.
- Extended for arbitrary basic_json types in version 3.10.5. - Extended for arbitrary basic_json types in version 3.10.5.
- Numbers that compare equal hash equally since version 3.13.0; before, `#!cpp 0`, `#!cpp 0U`, and `#!cpp 0.0` had
different hash values.
+1
View File
@@ -11,6 +11,7 @@ int main()
<< "hash(false) = " << std::hash<json> {}(json(false)) << '\n' << "hash(false) = " << std::hash<json> {}(json(false)) << '\n'
<< "hash(0) = " << std::hash<json> {}(json(0)) << '\n' << "hash(0) = " << std::hash<json> {}(json(0)) << '\n'
<< "hash(0U) = " << std::hash<json> {}(json(0U)) << '\n' << "hash(0U) = " << std::hash<json> {}(json(0U)) << '\n'
<< "hash(0.0) = " << std::hash<json> {}(json(0.0)) << '\n'
<< "hash(\"\") = " << std::hash<json> {}(json("")) << '\n' << "hash(\"\") = " << std::hash<json> {}(json("")) << '\n'
<< "hash({}) = " << std::hash<json> {}(json::object()) << '\n' << "hash({}) = " << std::hash<json> {}(json::object()) << '\n'
<< "hash([]) = " << std::hash<json> {}(json::array()) << '\n' << "hash([]) = " << std::hash<json> {}(json::array()) << '\n'
+5 -4
View File
@@ -1,8 +1,9 @@
hash(null) = 2654435769 hash(null) = 2654435769
hash(false) = 2654436030 hash(false) = 2654436030
hash(0) = 2654436095 hash(0) = 2654436221
hash(0U) = 2654436156 hash(0U) = 2654436221
hash("") = 6142509191626859748 hash(0.0) = 2654436221
hash("") = 11160318156688833227
hash({}) = 2654435832 hash({}) = 2654435832
hash([]) = 2654435899 hash([]) = 2654435899
hash({"hello": "world"}) = 4469488738203676328 hash({"hello": "world"}) = 3701319991624763853
+19 -16
View File
@@ -35,8 +35,10 @@ std::size_t hash_iteratively(const BasicJsonType& j);
@brief hash a JSON value @brief hash a JSON value
The hash function tries to rely on std::hash where possible. Furthermore, the The hash function tries to rely on std::hash where possible. Furthermore, the
type of the JSON value is taken into account to have different hash values for type of the JSON value is taken into account, so null, false, and numbers may
null, 0, 0U, and false, etc. hash differently from each other, but any two numbers that compare equal
under operator== hash equally regardless of which of number_integer,
number_unsigned, or number_float actually holds the value.
Hashing an array or an object hashes its elements, which used to call this Hashing an array or an object hashes its elements, which used to call this
function again once per nesting level, so a value nested deeply enough function again once per nesting level, so a value nested deeply enough
@@ -55,8 +57,6 @@ template<typename BasicJsonType>
std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0) std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
{ {
using string_t = typename BasicJsonType::string_t; using string_t = typename BasicJsonType::string_t;
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t; using number_float_t = typename BasicJsonType::number_float_t;
const auto type = static_cast<std::size_t>(j.type()); const auto type = static_cast<std::size_t>(j.type());
@@ -113,21 +113,24 @@ std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
} }
case BasicJsonType::value_t::number_integer: case BasicJsonType::value_t::number_integer:
{
const auto h = std::hash<number_integer_t> {}(j.template get<number_integer_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_unsigned: case BasicJsonType::value_t::number_unsigned:
{
const auto h = std::hash<number_unsigned_t> {}(j.template get<number_unsigned_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_float: case BasicJsonType::value_t::number_float:
{ {
const auto h = std::hash<number_float_t> {}(j.template get<number_float_t>()); // operator== compares numbers by their mathematical value across
return combine(type, h); // number_integer, number_unsigned, and number_float, so equal
// numbers of different internal types (0, 0U, 0.0) must hash the
// same. Two equal numbers have the same value, which converts to
// the same number_float_t, so all numbers share one type tag and
// hash that converted value. Adding zero turns -0.0 (equal to 0)
// into 0.0, as std::hash need not map both to the same hash.
// The converse does not hold: converting a number_float_t value
// to an integer type is lossy, so the result can hash
// differently, and unequal numbers that convert to the same
// number_float_t (e.g., 2^53 and 2^53 + 1) share a hash.
const auto number_type = static_cast<std::size_t>(BasicJsonType::value_t::number_float);
const auto value = j.template get<number_float_t>() + static_cast<number_float_t>(0);
const auto h = std::hash<number_float_t> {}(value);
return combine(number_type, h);
} }
case BasicJsonType::value_t::binary: case BasicJsonType::value_t::binary:
+124 -568
View File
@@ -28,7 +28,6 @@
#include <nlohmann/detail/macro_scope.hpp> #include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/error_handler.hpp> #include <nlohmann/detail/output/error_handler.hpp>
#include <nlohmann/detail/output/output_adapters.hpp> #include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/recursion_depth_limit.hpp>
#include <nlohmann/detail/string_concat.hpp> #include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/string_utils.hpp> #include <nlohmann/detail/string_utils.hpp>
@@ -158,29 +157,12 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@param[in] depth nesting level of @a j, counted from the top-level value
passed to @ref basic_json::to_cbor
@throw type_error.316 if a string value or an object key is not valid @throw type_error.316 if a string value or an object key is not valid
UTF-8 UTF-8
@throw type_error.321 if @a j or a value nested in it is discarded @throw type_error.321 if @a j or a value nested in it is discarded
Serializing a container descends into its elements, so a value nested deeply
enough used to exhaust the call stack and terminate the process with no
exception to catch. The descent is bounded here: once @ref recursion_depth_limit
levels have been entered, @ref write_cbor_iterative writes out what is left
without the call stack. A value nested less deeply than that - all but a
vanishing minority - is written by exactly the code that always wrote it.
@sa https://github.com/nlohmann/json/issues/5392
*/ */
void write_cbor(const BasicJsonType& j, const std::size_t depth = 0) void write_cbor(const BasicJsonType& j)
{ {
if (JSON_HEDLEY_UNLIKELY(depth >= recursion_depth_limit()) && (j.is_array() || j.is_object()))
{
write_cbor_iterative(j);
return;
}
switch (j.type()) switch (j.type())
{ {
case value_t::null: case value_t::null:
@@ -262,9 +244,10 @@ class binary_writer
// step 1: write control byte and the array size // step 1: write control byte and the array size
write_cbor_head(0x80, j.m_data.m_value.array->size()); write_cbor_head(0x80, j.m_data.m_value.array->size());
// step 2: write each element
for (const auto& el : *j.m_data.m_value.array) for (const auto& el : *j.m_data.m_value.array)
{ {
write_cbor(el, depth + 1); write_cbor(el);
} }
break; break;
} }
@@ -319,6 +302,7 @@ class binary_writer
// step 1: write control byte and the object size // step 1: write control byte and the object size
write_cbor_head(0xA0, j.m_data.m_value.object->size()); write_cbor_head(0xA0, j.m_data.m_value.object->size());
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
// el.first is checked here, against the object as // el.first is checked here, against the object as
@@ -333,7 +317,7 @@ class binary_writer
check_utf8(el.first, j); check_utf8(el.first, j);
} }
write_cbor(el.first); write_cbor(el.first);
write_cbor(el.second, depth + 1); write_cbor(el.second);
} }
break; break;
} }
@@ -400,21 +384,10 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@param[in] depth nesting level of @a j, counted from the top-level value
passed to @ref basic_json::to_msgpack
@throw type_error.321 if @a j or a value nested in it is discarded @throw type_error.321 if @a j or a value nested in it is discarded
@sa @ref write_cbor
@sa https://github.com/nlohmann/json/issues/5392
*/ */
void write_msgpack(const BasicJsonType& j, const std::size_t depth = 0) void write_msgpack(const BasicJsonType& j)
{ {
if (JSON_HEDLEY_UNLIKELY(depth >= recursion_depth_limit()) && (j.is_array() || j.is_object()))
{
write_msgpack_iterative(j);
return;
}
switch (j.type()) switch (j.type())
{ {
case value_t::null: // nil case value_t::null: // nil
@@ -530,11 +503,29 @@ class binary_writer
case value_t::array: case value_t::array:
{ {
// step 1: write control byte and the array size // step 1: write control byte and the array size
write_msgpack_array_prefix(j.m_data.m_value.array->size(), j); const auto N = to_msgpack_length(j.m_data.m_value.array->size(), j);
if (N <= 15)
{
// fixarray
write_number(static_cast<std::uint8_t>(0x90 | N));
}
else if (N <= (std::numeric_limits<std::uint16_t>::max)())
{
// array 16
oa.write_character(to_char_type(0xDC));
write_number(static_cast<std::uint16_t>(N));
}
else
{
// array 32
oa.write_character(to_char_type(0xDD));
write_number(static_cast<std::uint32_t>(N));
}
// step 2: write each element
for (const auto& el : *j.m_data.m_value.array) for (const auto& el : *j.m_data.m_value.array)
{ {
write_msgpack(el, depth + 1); write_msgpack(el);
} }
break; break;
} }
@@ -630,8 +621,26 @@ class binary_writer
case value_t::object: case value_t::object:
{ {
// step 1: write control byte and the object size // step 1: write control byte and the object size
write_msgpack_object_prefix(j.m_data.m_value.object->size(), j); const auto N = to_msgpack_length(j.m_data.m_value.object->size(), j);
if (N <= 15)
{
// fixmap
write_number(static_cast<std::uint8_t>(0x80 | (N & 0xF)));
}
else if (N <= (std::numeric_limits<std::uint16_t>::max)())
{
// map 16
oa.write_character(to_char_type(0xDE));
write_number(static_cast<std::uint16_t>(N));
}
else
{
// map 32
oa.write_character(to_char_type(0xDF));
write_number(static_cast<std::uint32_t>(N));
}
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
// as in write_cbor, el.first is checked here against the // as in write_cbor, el.first is checked here against the
@@ -642,7 +651,7 @@ class binary_writer
check_utf8(el.first, j); check_utf8(el.first, j);
} }
write_msgpack(el.first); write_msgpack(el.first);
write_msgpack(el.second, depth + 1); write_msgpack(el.second);
} }
break; break;
} }
@@ -660,26 +669,14 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value @param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false @param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2 @param[in] bjdata_version which BJData version to use, default is draft2
@param[in] depth nesting level of @a j, counted from the top-level value
passed to @ref basic_json::to_ubjson or @ref basic_json::to_bjdata
@throw type_error.316 if a string value or an object key is not valid @throw type_error.316 if a string value or an object key is not valid
UTF-8 UTF-8
@throw type_error.321 if @a j or a value nested in it is discarded @throw type_error.321 if @a j or a value nested in it is discarded
@sa @ref write_cbor
@sa https://github.com/nlohmann/json/issues/5392
*/ */
void write_ubjson(const BasicJsonType& j, const bool use_count, void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true, const bool use_type, const bool add_prefix = true,
const bool use_bjdata = false, const bjdata_version_t bjdata_version = bjdata_version_t::draft2, const bool use_bjdata = false, const bjdata_version_t bjdata_version = bjdata_version_t::draft2)
const std::size_t depth = 0)
{ {
if (JSON_HEDLEY_UNLIKELY(depth >= recursion_depth_limit()) && (j.is_array() || j.is_object()))
{
write_ubjson_iterative(j, use_count, use_type, add_prefix, use_bjdata, bjdata_version);
return;
}
const bool bjdata_draft3 = use_bjdata && bjdata_version == bjdata_version_t::draft3; const bool bjdata_draft3 = use_bjdata && bjdata_version == bjdata_version_t::draft3;
switch (j.type()) switch (j.type())
@@ -740,15 +737,55 @@ class binary_writer
case value_t::array: case value_t::array:
{ {
if (add_prefix)
{
oa.write_character(to_char_type('['));
}
bool prefix_required = true; bool prefix_required = true;
const bool write_closer = write_ubjson_start_array(j, use_count, use_type, add_prefix, use_bjdata, prefix_required); if (use_type && !j.m_data.m_value.array->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin() + 1, j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
// an optimized array of a valueless type carries no payload, so a
// reader has nothing but the declared count to bound the allocation
// by and refuses an excessive one. Write the unoptimized form for
// those, at one byte per element, so the result can be read back.
// Objects are not affected: every element is preceded by its key.
const bool valueless_type = (first_prefix == 'Z' || first_prefix == 'T' || first_prefix == 'F');
const bool excessive_valueless = valueless_type
&& j.m_data.m_value.array->size() > detail::max_valueless_container_size;
if (same_prefix && !excessive_valueless
&& !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.array->size(), true, use_bjdata);
}
for (const auto& el : *j.m_data.m_value.array) for (const auto& el : *j.m_data.m_value.array)
{ {
write_ubjson(el, use_count, use_type, prefix_required, use_bjdata, bjdata_version, depth + 1); write_ubjson(el, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
} }
if (write_closer) if (!use_count)
{ {
oa.write_character(to_char_type(']')); oa.write_character(to_char_type(']'));
} }
@@ -806,7 +843,7 @@ class binary_writer
case value_t::object: case value_t::object:
{ {
if (use_bjdata && is_bjdata_ndarray(j)) if (use_bjdata && j.m_data.m_value.object->size() == 3 && j.m_data.m_value.object->find("_ArrayType_") != j.m_data.m_value.object->end() && j.m_data.m_value.object->find("_ArraySize_") != j.m_data.m_value.object->end() && j.m_data.m_value.object->find("_ArrayData_") != j.m_data.m_value.object->end())
{ {
if (!write_bjdata_ndarray(*j.m_data.m_value.object, use_count, use_type, bjdata_version)) // decode bjdata ndarray in the JData format (https://github.com/NeuroJSON/jdata) if (!write_bjdata_ndarray(*j.m_data.m_value.object, use_count, use_type, bjdata_version)) // decode bjdata ndarray in the JData format (https://github.com/NeuroJSON/jdata)
{ {
@@ -814,8 +851,38 @@ class binary_writer
} }
} }
if (add_prefix)
{
oa.write_character(to_char_type('{'));
}
bool prefix_required = true; bool prefix_required = true;
const bool write_closer = write_ubjson_start_object(j, use_count, use_type, add_prefix, use_bjdata, prefix_required); if (use_type && !j.m_data.m_value.object->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin(), j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
if (same_prefix && !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.object->size(), true, use_bjdata);
}
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
@@ -825,10 +892,10 @@ class binary_writer
oa.write_characters( oa.write_characters(
reinterpret_cast<const CharType*>(key.data()), reinterpret_cast<const CharType*>(key.data()),
key.size()); key.size());
write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version, depth + 1); write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
} }
if (write_closer) if (!use_count)
{ {
oa.write_character(to_char_type('}')); oa.write_character(to_char_type('}'));
} }
@@ -869,517 +936,6 @@ class binary_writer
JSON_THROW(type_error::create(321, concat("cannot serialize discarded value to ", format_name), &j)); JSON_THROW(type_error::create(321, concat("cannot serialize discarded value to ", format_name), &j));
} }
void write_msgpack_array_prefix(const std::size_t N, const BasicJsonType& j)
{
const auto n = to_msgpack_length(N, j);
if (n <= 15)
{
// fixarray
write_number(static_cast<std::uint8_t>(0x90 | n));
}
else if (n <= (std::numeric_limits<std::uint16_t>::max)())
{
// array 16
oa.write_character(to_char_type(0xDC));
write_number(static_cast<std::uint16_t>(n));
}
else
{
// array 32
oa.write_character(to_char_type(0xDD));
write_number(static_cast<std::uint32_t>(n));
}
}
void write_msgpack_object_prefix(const std::size_t N, const BasicJsonType& j)
{
const auto n = to_msgpack_length(N, j);
if (n <= 15)
{
// fixmap
write_number(static_cast<std::uint8_t>(0x80 | (n & 0xF)));
}
else if (n <= (std::numeric_limits<std::uint16_t>::max)())
{
// map 16
oa.write_character(to_char_type(0xDE));
write_number(static_cast<std::uint16_t>(n));
}
else
{
// map 32
oa.write_character(to_char_type(0xDF));
write_number(static_cast<std::uint32_t>(n));
}
}
/// @brief a CBOR or MessagePack array or object whose elements
/// @ref write_cbor_iterative or @ref write_msgpack_iterative is
/// still writing
struct binary_container_frame
{
explicit binary_container_frame(const BasicJsonType* value_) noexcept
: value(value_)
{
if (value->is_object())
{
object_it = value->m_data.m_value.object->cbegin();
}
else
{
array_it = value->m_data.m_value.array->cbegin();
}
}
// declared for GCC's -Weffc++, which asks for them in a class with
// pointer members and a non-trivial destructor; the exception
// specifications are left implicit, as GCC 4.8 rejects explicit ones
// that differ from them
binary_container_frame(const binary_container_frame&) = default;
binary_container_frame(binary_container_frame&&) = default;
binary_container_frame& operator=(const binary_container_frame&) = default;
binary_container_frame& operator=(binary_container_frame&&) = default;
~binary_container_frame() = default;
/// the array or object being written
const BasicJsonType* value;
/// value's elements still to write; which of the two is live follows
/// from the type of value. They are kept side by side rather than in
/// a union, which would need its special members written out by
/// hand, see detail/iterators/internal_iterator.hpp
typename BasicJsonType::object_t::const_iterator object_it{};
typename BasicJsonType::array_t::const_iterator array_it{};
};
/*!
@brief write @a j with @ref write_cbor, or write its header and push a
frame for @ref write_cbor_iterative to continue with its elements
A scalar, and an empty array or object, are written out in full: there is
nothing below them for @ref write_cbor_iterative to come back to, so
nothing is pushed for them.
*/
void write_cbor_value_or_push(const BasicJsonType& j, std::vector<binary_container_frame>& stack)
{
if (j.is_array())
{
write_cbor_head(0x80, j.m_data.m_value.array->size());
if (!j.m_data.m_value.array->empty())
{
stack.emplace_back(&j);
}
return;
}
if (j.is_object())
{
write_cbor_head(0xA0, j.m_data.m_value.object->size());
if (!j.m_data.m_value.object->empty())
{
stack.emplace_back(&j);
}
return;
}
write_cbor(j);
}
/*!
@brief write out @a root and everything below it without the call stack
Emits the same bytes as @ref write_cbor, keeping the containers it has
entered on an explicit stack instead of descending into them. Only reached
for values nested deeper than @ref recursion_depth_limit, which is why it
is not written for speed.
*/
void write_cbor_iterative(const BasicJsonType& root)
{
// only a container with elements is ever pushed; see write_cbor_value_or_push
std::vector<binary_container_frame> stack;
write_cbor_value_or_push(root, stack);
while (!stack.empty())
{
const binary_container_frame current = stack.back();
if (current.value->is_array())
{
const auto& array = *current.value->m_data.m_value.array;
if (current.array_it == array.cend())
{
stack.pop_back();
continue;
}
// read the child before pushing: entering it can move every frame
const BasicJsonType* child = &(*current.array_it);
++stack.back().array_it;
write_cbor_value_or_push(*child, stack);
}
else
{
const auto& object = *current.value->m_data.m_value.object;
if (current.object_it == object.cend())
{
stack.pop_back();
continue;
}
// el.first is checked here, against the object as diagnostics
// context, like the matching check in write_cbor's object case
if (error_handler == error_handler_t::strict)
{
check_utf8(current.object_it->first, *current.value);
}
write_cbor(current.object_it->first);
const BasicJsonType* child = &(current.object_it->second);
++stack.back().object_it;
write_cbor_value_or_push(*child, stack);
}
}
}
/*!
@brief write @a j with @ref write_msgpack, or write its header and push a
frame for @ref write_msgpack_iterative to continue with its elements
@sa @ref write_cbor_value_or_push
*/
void write_msgpack_value_or_push(const BasicJsonType& j, std::vector<binary_container_frame>& stack)
{
if (j.is_array())
{
write_msgpack_array_prefix(j.m_data.m_value.array->size(), j);
if (!j.m_data.m_value.array->empty())
{
stack.emplace_back(&j);
}
return;
}
if (j.is_object())
{
write_msgpack_object_prefix(j.m_data.m_value.object->size(), j);
if (!j.m_data.m_value.object->empty())
{
stack.emplace_back(&j);
}
return;
}
write_msgpack(j);
}
/*!
@brief write out @a root and everything below it without the call stack
@sa @ref write_cbor_iterative
*/
void write_msgpack_iterative(const BasicJsonType& root)
{
std::vector<binary_container_frame> stack;
write_msgpack_value_or_push(root, stack);
while (!stack.empty())
{
const binary_container_frame current = stack.back();
if (current.value->is_array())
{
const auto& array = *current.value->m_data.m_value.array;
if (current.array_it == array.cend())
{
stack.pop_back();
continue;
}
const BasicJsonType* child = &(*current.array_it);
++stack.back().array_it;
write_msgpack_value_or_push(*child, stack);
}
else
{
const auto& object = *current.value->m_data.m_value.object;
if (current.object_it == object.cend())
{
stack.pop_back();
continue;
}
if (error_handler == error_handler_t::strict)
{
check_utf8(current.object_it->first, *current.value);
}
write_msgpack(current.object_it->first);
const BasicJsonType* child = &(current.object_it->second);
++stack.back().object_it;
write_msgpack_value_or_push(*child, stack);
}
}
}
/// @return true when a closing ']' still has to be written after the elements
bool write_ubjson_start_array(const BasicJsonType& j, const bool use_count, const bool use_type,
const bool add_prefix, const bool use_bjdata, bool& prefix_required)
{
prefix_required = true;
if (add_prefix)
{
oa.write_character(to_char_type('['));
}
if (use_type && !j.m_data.m_value.array->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin() + 1, j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
// an optimized array of a valueless type carries no payload, so a
// reader has nothing but the declared count to bound the allocation
// by and refuses an excessive one. Write the unoptimized form for
// those, at one byte per element, so the result can be read back.
// Objects are not affected: every element is preceded by its key.
const bool valueless_type = (first_prefix == 'Z' || first_prefix == 'T' || first_prefix == 'F');
const bool excessive_valueless = valueless_type
&& j.m_data.m_value.array->size() > detail::max_valueless_container_size;
if (same_prefix && !excessive_valueless
&& !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.array->size(), true, use_bjdata);
}
return !use_count;
}
/// @return true when a closing '}' still has to be written after the elements
bool write_ubjson_start_object(const BasicJsonType& j, const bool use_count, const bool use_type,
const bool add_prefix, const bool use_bjdata, bool& prefix_required)
{
prefix_required = true;
if (add_prefix)
{
oa.write_character(to_char_type('{'));
}
if (use_type && !j.m_data.m_value.object->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin(), j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
if (same_prefix && !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.object->size(), true, use_bjdata);
}
return !use_count;
}
/*!
@brief whether @a j is a BJData ND-array annotation object
(https://github.com/NeuroJSON/jdata)
Used by both the recursive object case of @ref write_ubjson and
@ref write_ubjson_value_or_push, which must agree on what counts as an
ND-array: @a j is only actually written as one once @ref
write_bjdata_ndarray has also accepted its contents.
@pre @a j.is_object()
*/
static bool is_bjdata_ndarray(const BasicJsonType& j)
{
const auto& object = *j.m_data.m_value.object;
return object.size() == 3
&& object.find("_ArrayType_") != object.end()
&& object.find("_ArraySize_") != object.end()
&& object.find("_ArrayData_") != object.end();
}
/// @brief an object or array @ref write_ubjson_iterative is still writing
/// the elements of
struct ubjson_frame
{
ubjson_frame(const BasicJsonType* value_, const bool prefix_required_) noexcept
: value(value_)
, prefix_required(prefix_required_)
{
if (value->is_object())
{
object_it = value->m_data.m_value.object->cbegin();
}
else
{
array_it = value->m_data.m_value.array->cbegin();
}
}
// declared for GCC's -Weffc++, which asks for them in a class with
// pointer members and a non-trivial destructor; the exception
// specifications are left implicit, as GCC 4.8 rejects explicit ones
// that differ from them
ubjson_frame(const ubjson_frame&) = default;
ubjson_frame(ubjson_frame&&) = default;
ubjson_frame& operator=(const ubjson_frame&) = default;
ubjson_frame& operator=(ubjson_frame&&) = default;
~ubjson_frame() = default;
/// the array or object being written
const BasicJsonType* value;
/// whether value's elements each carry their own type marker; an
/// optimized ($type) container writes it once for all of them instead
bool prefix_required;
typename BasicJsonType::object_t::const_iterator object_it{};
typename BasicJsonType::array_t::const_iterator array_it{};
};
/*!
@brief write @a j with @ref write_ubjson, or write its header and push a
frame for @ref write_ubjson_iterative to continue with its elements
@param[in] add_prefix whether @a j's own type marker is written now (the
elements of an optimized container, and everything below the
top level, never repeat it)
@sa @ref write_cbor_value_or_push
*/
void write_ubjson_value_or_push(const BasicJsonType& j, const bool add_prefix, const bool use_count,
const bool use_type, const bool use_bjdata, const bjdata_version_t bjdata_version,
std::vector<ubjson_frame>& stack)
{
if (!j.is_array() && !j.is_object())
{
write_ubjson(j, use_count, use_type, add_prefix, use_bjdata, bjdata_version);
return;
}
if (use_bjdata && j.is_object() && is_bjdata_ndarray(j)
&& !write_bjdata_ndarray(*j.m_data.m_value.object, use_count, use_type, bjdata_version))
{
// fully written as an ND-array: nothing below it to come back to
return;
}
const bool is_array = j.is_array();
bool prefix_required = true;
if (is_array)
{
write_ubjson_start_array(j, use_count, use_type, add_prefix, use_bjdata, prefix_required);
}
else
{
write_ubjson_start_object(j, use_count, use_type, add_prefix, use_bjdata, prefix_required);
}
const bool empty = is_array ? j.m_data.m_value.array->empty() : j.m_data.m_value.object->empty();
if (!empty)
{
stack.emplace_back(&j, prefix_required);
return;
}
// write_ubjson_start_array/_object return !use_count, i.e. whether a
// closer still has to be written; use_count is constant for the whole
// document, so that is recomputed here instead of being carried along
if (!use_count)
{
oa.write_character(to_char_type(is_array ? ']' : '}'));
}
}
/*!
@brief write out @a root and everything below it without the call stack
@sa @ref write_cbor_iterative
*/
void write_ubjson_iterative(const BasicJsonType& root, const bool use_count, const bool use_type,
const bool add_prefix, const bool use_bjdata, const bjdata_version_t bjdata_version)
{
std::vector<ubjson_frame> stack;
write_ubjson_value_or_push(root, add_prefix, use_count, use_type, use_bjdata, bjdata_version, stack);
while (!stack.empty())
{
const ubjson_frame current = stack.back();
const BasicJsonType& j = *current.value;
if (j.is_array())
{
const auto& array = *j.m_data.m_value.array;
if (current.array_it == array.cend())
{
if (!use_count)
{
oa.write_character(to_char_type(']'));
}
stack.pop_back();
continue;
}
const BasicJsonType* child = &(*current.array_it);
const bool child_prefix = current.prefix_required;
++stack.back().array_it;
write_ubjson_value_or_push(*child, child_prefix, use_count, use_type, use_bjdata, bjdata_version, stack);
}
else
{
const auto& object = *j.m_data.m_value.object;
if (current.object_it == object.cend())
{
if (!use_count)
{
oa.write_character(to_char_type('}'));
}
stack.pop_back();
continue;
}
string_t storage;
const string_t& key = sanitize_utf8_for_write(current.object_it->first, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(key.data()),
key.size());
const BasicJsonType* child = &(current.object_it->second);
const bool child_prefix = current.prefix_required;
++stack.back().object_it;
write_ubjson_value_or_push(*child, child_prefix, use_count, use_type, use_bjdata, bjdata_version, stack);
}
}
}
////////// //////////
// BSON // // BSON //
////////// //////////
+143 -585
View File
@@ -7583,8 +7583,10 @@ std::size_t hash_iteratively(const BasicJsonType& j);
@brief hash a JSON value @brief hash a JSON value
The hash function tries to rely on std::hash where possible. Furthermore, the The hash function tries to rely on std::hash where possible. Furthermore, the
type of the JSON value is taken into account to have different hash values for type of the JSON value is taken into account, so null, false, and numbers may
null, 0, 0U, and false, etc. hash differently from each other, but any two numbers that compare equal
under operator== hash equally regardless of which of number_integer,
number_unsigned, or number_float actually holds the value.
Hashing an array or an object hashes its elements, which used to call this Hashing an array or an object hashes its elements, which used to call this
function again once per nesting level, so a value nested deeply enough function again once per nesting level, so a value nested deeply enough
@@ -7603,8 +7605,6 @@ template<typename BasicJsonType>
std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0) std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
{ {
using string_t = typename BasicJsonType::string_t; using string_t = typename BasicJsonType::string_t;
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t; using number_float_t = typename BasicJsonType::number_float_t;
const auto type = static_cast<std::size_t>(j.type()); const auto type = static_cast<std::size_t>(j.type());
@@ -7661,21 +7661,24 @@ std::size_t hash(const BasicJsonType& j, const std::size_t depth = 0)
} }
case BasicJsonType::value_t::number_integer: case BasicJsonType::value_t::number_integer:
{
const auto h = std::hash<number_integer_t> {}(j.template get<number_integer_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_unsigned: case BasicJsonType::value_t::number_unsigned:
{
const auto h = std::hash<number_unsigned_t> {}(j.template get<number_unsigned_t>());
return combine(type, h);
}
case BasicJsonType::value_t::number_float: case BasicJsonType::value_t::number_float:
{ {
const auto h = std::hash<number_float_t> {}(j.template get<number_float_t>()); // operator== compares numbers by their mathematical value across
return combine(type, h); // number_integer, number_unsigned, and number_float, so equal
// numbers of different internal types (0, 0U, 0.0) must hash the
// same. Two equal numbers have the same value, which converts to
// the same number_float_t, so all numbers share one type tag and
// hash that converted value. Adding zero turns -0.0 (equal to 0)
// into 0.0, as std::hash need not map both to the same hash.
// The converse does not hold: converting a number_float_t value
// to an integer type is lossy, so the result can hash
// differently, and unequal numbers that convert to the same
// number_float_t (e.g., 2^53 and 2^53 + 1) share a hash.
const auto number_type = static_cast<std::size_t>(BasicJsonType::value_t::number_float);
const auto value = j.template get<number_float_t>() + static_cast<number_float_t>(0);
const auto h = std::hash<number_float_t> {}(value);
return combine(number_type, h);
} }
case BasicJsonType::value_t::binary: case BasicJsonType::value_t::binary:
@@ -21497,8 +21500,6 @@ class output_adapter
} // namespace detail } // namespace detail
NLOHMANN_JSON_NAMESPACE_END NLOHMANN_JSON_NAMESPACE_END
// #include <nlohmann/detail/recursion_depth_limit.hpp>
// #include <nlohmann/detail/string_concat.hpp> // #include <nlohmann/detail/string_concat.hpp>
// #include <nlohmann/detail/string_utils.hpp> // #include <nlohmann/detail/string_utils.hpp>
@@ -21630,29 +21631,12 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@param[in] depth nesting level of @a j, counted from the top-level value
passed to @ref basic_json::to_cbor
@throw type_error.316 if a string value or an object key is not valid @throw type_error.316 if a string value or an object key is not valid
UTF-8 UTF-8
@throw type_error.321 if @a j or a value nested in it is discarded @throw type_error.321 if @a j or a value nested in it is discarded
Serializing a container descends into its elements, so a value nested deeply
enough used to exhaust the call stack and terminate the process with no
exception to catch. The descent is bounded here: once @ref recursion_depth_limit
levels have been entered, @ref write_cbor_iterative writes out what is left
without the call stack. A value nested less deeply than that - all but a
vanishing minority - is written by exactly the code that always wrote it.
@sa https://github.com/nlohmann/json/issues/5392
*/ */
void write_cbor(const BasicJsonType& j, const std::size_t depth = 0) void write_cbor(const BasicJsonType& j)
{ {
if (JSON_HEDLEY_UNLIKELY(depth >= recursion_depth_limit()) && (j.is_array() || j.is_object()))
{
write_cbor_iterative(j);
return;
}
switch (j.type()) switch (j.type())
{ {
case value_t::null: case value_t::null:
@@ -21734,9 +21718,10 @@ class binary_writer
// step 1: write control byte and the array size // step 1: write control byte and the array size
write_cbor_head(0x80, j.m_data.m_value.array->size()); write_cbor_head(0x80, j.m_data.m_value.array->size());
// step 2: write each element
for (const auto& el : *j.m_data.m_value.array) for (const auto& el : *j.m_data.m_value.array)
{ {
write_cbor(el, depth + 1); write_cbor(el);
} }
break; break;
} }
@@ -21791,6 +21776,7 @@ class binary_writer
// step 1: write control byte and the object size // step 1: write control byte and the object size
write_cbor_head(0xA0, j.m_data.m_value.object->size()); write_cbor_head(0xA0, j.m_data.m_value.object->size());
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
// el.first is checked here, against the object as // el.first is checked here, against the object as
@@ -21805,7 +21791,7 @@ class binary_writer
check_utf8(el.first, j); check_utf8(el.first, j);
} }
write_cbor(el.first); write_cbor(el.first);
write_cbor(el.second, depth + 1); write_cbor(el.second);
} }
break; break;
} }
@@ -21872,21 +21858,10 @@ class binary_writer
/*! /*!
@param[in] j JSON value to serialize @param[in] j JSON value to serialize
@param[in] depth nesting level of @a j, counted from the top-level value
passed to @ref basic_json::to_msgpack
@throw type_error.321 if @a j or a value nested in it is discarded @throw type_error.321 if @a j or a value nested in it is discarded
@sa @ref write_cbor
@sa https://github.com/nlohmann/json/issues/5392
*/ */
void write_msgpack(const BasicJsonType& j, const std::size_t depth = 0) void write_msgpack(const BasicJsonType& j)
{ {
if (JSON_HEDLEY_UNLIKELY(depth >= recursion_depth_limit()) && (j.is_array() || j.is_object()))
{
write_msgpack_iterative(j);
return;
}
switch (j.type()) switch (j.type())
{ {
case value_t::null: // nil case value_t::null: // nil
@@ -22002,11 +21977,29 @@ class binary_writer
case value_t::array: case value_t::array:
{ {
// step 1: write control byte and the array size // step 1: write control byte and the array size
write_msgpack_array_prefix(j.m_data.m_value.array->size(), j); const auto N = to_msgpack_length(j.m_data.m_value.array->size(), j);
if (N <= 15)
{
// fixarray
write_number(static_cast<std::uint8_t>(0x90 | N));
}
else if (N <= (std::numeric_limits<std::uint16_t>::max)())
{
// array 16
oa.write_character(to_char_type(0xDC));
write_number(static_cast<std::uint16_t>(N));
}
else
{
// array 32
oa.write_character(to_char_type(0xDD));
write_number(static_cast<std::uint32_t>(N));
}
// step 2: write each element
for (const auto& el : *j.m_data.m_value.array) for (const auto& el : *j.m_data.m_value.array)
{ {
write_msgpack(el, depth + 1); write_msgpack(el);
} }
break; break;
} }
@@ -22102,8 +22095,26 @@ class binary_writer
case value_t::object: case value_t::object:
{ {
// step 1: write control byte and the object size // step 1: write control byte and the object size
write_msgpack_object_prefix(j.m_data.m_value.object->size(), j); const auto N = to_msgpack_length(j.m_data.m_value.object->size(), j);
if (N <= 15)
{
// fixmap
write_number(static_cast<std::uint8_t>(0x80 | (N & 0xF)));
}
else if (N <= (std::numeric_limits<std::uint16_t>::max)())
{
// map 16
oa.write_character(to_char_type(0xDE));
write_number(static_cast<std::uint16_t>(N));
}
else
{
// map 32
oa.write_character(to_char_type(0xDF));
write_number(static_cast<std::uint32_t>(N));
}
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
// as in write_cbor, el.first is checked here against the // as in write_cbor, el.first is checked here against the
@@ -22114,7 +22125,7 @@ class binary_writer
check_utf8(el.first, j); check_utf8(el.first, j);
} }
write_msgpack(el.first); write_msgpack(el.first);
write_msgpack(el.second, depth + 1); write_msgpack(el.second);
} }
break; break;
} }
@@ -22132,26 +22143,14 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value @param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false @param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2 @param[in] bjdata_version which BJData version to use, default is draft2
@param[in] depth nesting level of @a j, counted from the top-level value
passed to @ref basic_json::to_ubjson or @ref basic_json::to_bjdata
@throw type_error.316 if a string value or an object key is not valid @throw type_error.316 if a string value or an object key is not valid
UTF-8 UTF-8
@throw type_error.321 if @a j or a value nested in it is discarded @throw type_error.321 if @a j or a value nested in it is discarded
@sa @ref write_cbor
@sa https://github.com/nlohmann/json/issues/5392
*/ */
void write_ubjson(const BasicJsonType& j, const bool use_count, void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true, const bool use_type, const bool add_prefix = true,
const bool use_bjdata = false, const bjdata_version_t bjdata_version = bjdata_version_t::draft2, const bool use_bjdata = false, const bjdata_version_t bjdata_version = bjdata_version_t::draft2)
const std::size_t depth = 0)
{ {
if (JSON_HEDLEY_UNLIKELY(depth >= recursion_depth_limit()) && (j.is_array() || j.is_object()))
{
write_ubjson_iterative(j, use_count, use_type, add_prefix, use_bjdata, bjdata_version);
return;
}
const bool bjdata_draft3 = use_bjdata && bjdata_version == bjdata_version_t::draft3; const bool bjdata_draft3 = use_bjdata && bjdata_version == bjdata_version_t::draft3;
switch (j.type()) switch (j.type())
@@ -22212,15 +22211,55 @@ class binary_writer
case value_t::array: case value_t::array:
{ {
if (add_prefix)
{
oa.write_character(to_char_type('['));
}
bool prefix_required = true; bool prefix_required = true;
const bool write_closer = write_ubjson_start_array(j, use_count, use_type, add_prefix, use_bjdata, prefix_required); if (use_type && !j.m_data.m_value.array->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin() + 1, j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
// an optimized array of a valueless type carries no payload, so a
// reader has nothing but the declared count to bound the allocation
// by and refuses an excessive one. Write the unoptimized form for
// those, at one byte per element, so the result can be read back.
// Objects are not affected: every element is preceded by its key.
const bool valueless_type = (first_prefix == 'Z' || first_prefix == 'T' || first_prefix == 'F');
const bool excessive_valueless = valueless_type
&& j.m_data.m_value.array->size() > detail::max_valueless_container_size;
if (same_prefix && !excessive_valueless
&& !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.array->size(), true, use_bjdata);
}
for (const auto& el : *j.m_data.m_value.array) for (const auto& el : *j.m_data.m_value.array)
{ {
write_ubjson(el, use_count, use_type, prefix_required, use_bjdata, bjdata_version, depth + 1); write_ubjson(el, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
} }
if (write_closer) if (!use_count)
{ {
oa.write_character(to_char_type(']')); oa.write_character(to_char_type(']'));
} }
@@ -22278,7 +22317,7 @@ class binary_writer
case value_t::object: case value_t::object:
{ {
if (use_bjdata && is_bjdata_ndarray(j)) if (use_bjdata && j.m_data.m_value.object->size() == 3 && j.m_data.m_value.object->find("_ArrayType_") != j.m_data.m_value.object->end() && j.m_data.m_value.object->find("_ArraySize_") != j.m_data.m_value.object->end() && j.m_data.m_value.object->find("_ArrayData_") != j.m_data.m_value.object->end())
{ {
if (!write_bjdata_ndarray(*j.m_data.m_value.object, use_count, use_type, bjdata_version)) // decode bjdata ndarray in the JData format (https://github.com/NeuroJSON/jdata) if (!write_bjdata_ndarray(*j.m_data.m_value.object, use_count, use_type, bjdata_version)) // decode bjdata ndarray in the JData format (https://github.com/NeuroJSON/jdata)
{ {
@@ -22286,8 +22325,38 @@ class binary_writer
} }
} }
if (add_prefix)
{
oa.write_character(to_char_type('{'));
}
bool prefix_required = true; bool prefix_required = true;
const bool write_closer = write_ubjson_start_object(j, use_count, use_type, add_prefix, use_bjdata, prefix_required); if (use_type && !j.m_data.m_value.object->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin(), j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
if (same_prefix && !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.object->size(), true, use_bjdata);
}
for (const auto& el : *j.m_data.m_value.object) for (const auto& el : *j.m_data.m_value.object)
{ {
@@ -22297,10 +22366,10 @@ class binary_writer
oa.write_characters( oa.write_characters(
reinterpret_cast<const CharType*>(key.data()), reinterpret_cast<const CharType*>(key.data()),
key.size()); key.size());
write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version, depth + 1); write_ubjson(el.second, use_count, use_type, prefix_required, use_bjdata, bjdata_version);
} }
if (write_closer) if (!use_count)
{ {
oa.write_character(to_char_type('}')); oa.write_character(to_char_type('}'));
} }
@@ -22341,517 +22410,6 @@ class binary_writer
JSON_THROW(type_error::create(321, concat("cannot serialize discarded value to ", format_name), &j)); JSON_THROW(type_error::create(321, concat("cannot serialize discarded value to ", format_name), &j));
} }
void write_msgpack_array_prefix(const std::size_t N, const BasicJsonType& j)
{
const auto n = to_msgpack_length(N, j);
if (n <= 15)
{
// fixarray
write_number(static_cast<std::uint8_t>(0x90 | n));
}
else if (n <= (std::numeric_limits<std::uint16_t>::max)())
{
// array 16
oa.write_character(to_char_type(0xDC));
write_number(static_cast<std::uint16_t>(n));
}
else
{
// array 32
oa.write_character(to_char_type(0xDD));
write_number(static_cast<std::uint32_t>(n));
}
}
void write_msgpack_object_prefix(const std::size_t N, const BasicJsonType& j)
{
const auto n = to_msgpack_length(N, j);
if (n <= 15)
{
// fixmap
write_number(static_cast<std::uint8_t>(0x80 | (n & 0xF)));
}
else if (n <= (std::numeric_limits<std::uint16_t>::max)())
{
// map 16
oa.write_character(to_char_type(0xDE));
write_number(static_cast<std::uint16_t>(n));
}
else
{
// map 32
oa.write_character(to_char_type(0xDF));
write_number(static_cast<std::uint32_t>(n));
}
}
/// @brief a CBOR or MessagePack array or object whose elements
/// @ref write_cbor_iterative or @ref write_msgpack_iterative is
/// still writing
struct binary_container_frame
{
explicit binary_container_frame(const BasicJsonType* value_) noexcept
: value(value_)
{
if (value->is_object())
{
object_it = value->m_data.m_value.object->cbegin();
}
else
{
array_it = value->m_data.m_value.array->cbegin();
}
}
// declared for GCC's -Weffc++, which asks for them in a class with
// pointer members and a non-trivial destructor; the exception
// specifications are left implicit, as GCC 4.8 rejects explicit ones
// that differ from them
binary_container_frame(const binary_container_frame&) = default;
binary_container_frame(binary_container_frame&&) = default;
binary_container_frame& operator=(const binary_container_frame&) = default;
binary_container_frame& operator=(binary_container_frame&&) = default;
~binary_container_frame() = default;
/// the array or object being written
const BasicJsonType* value;
/// value's elements still to write; which of the two is live follows
/// from the type of value. They are kept side by side rather than in
/// a union, which would need its special members written out by
/// hand, see detail/iterators/internal_iterator.hpp
typename BasicJsonType::object_t::const_iterator object_it{};
typename BasicJsonType::array_t::const_iterator array_it{};
};
/*!
@brief write @a j with @ref write_cbor, or write its header and push a
frame for @ref write_cbor_iterative to continue with its elements
A scalar, and an empty array or object, are written out in full: there is
nothing below them for @ref write_cbor_iterative to come back to, so
nothing is pushed for them.
*/
void write_cbor_value_or_push(const BasicJsonType& j, std::vector<binary_container_frame>& stack)
{
if (j.is_array())
{
write_cbor_head(0x80, j.m_data.m_value.array->size());
if (!j.m_data.m_value.array->empty())
{
stack.emplace_back(&j);
}
return;
}
if (j.is_object())
{
write_cbor_head(0xA0, j.m_data.m_value.object->size());
if (!j.m_data.m_value.object->empty())
{
stack.emplace_back(&j);
}
return;
}
write_cbor(j);
}
/*!
@brief write out @a root and everything below it without the call stack
Emits the same bytes as @ref write_cbor, keeping the containers it has
entered on an explicit stack instead of descending into them. Only reached
for values nested deeper than @ref recursion_depth_limit, which is why it
is not written for speed.
*/
void write_cbor_iterative(const BasicJsonType& root)
{
// only a container with elements is ever pushed; see write_cbor_value_or_push
std::vector<binary_container_frame> stack;
write_cbor_value_or_push(root, stack);
while (!stack.empty())
{
const binary_container_frame current = stack.back();
if (current.value->is_array())
{
const auto& array = *current.value->m_data.m_value.array;
if (current.array_it == array.cend())
{
stack.pop_back();
continue;
}
// read the child before pushing: entering it can move every frame
const BasicJsonType* child = &(*current.array_it);
++stack.back().array_it;
write_cbor_value_or_push(*child, stack);
}
else
{
const auto& object = *current.value->m_data.m_value.object;
if (current.object_it == object.cend())
{
stack.pop_back();
continue;
}
// el.first is checked here, against the object as diagnostics
// context, like the matching check in write_cbor's object case
if (error_handler == error_handler_t::strict)
{
check_utf8(current.object_it->first, *current.value);
}
write_cbor(current.object_it->first);
const BasicJsonType* child = &(current.object_it->second);
++stack.back().object_it;
write_cbor_value_or_push(*child, stack);
}
}
}
/*!
@brief write @a j with @ref write_msgpack, or write its header and push a
frame for @ref write_msgpack_iterative to continue with its elements
@sa @ref write_cbor_value_or_push
*/
void write_msgpack_value_or_push(const BasicJsonType& j, std::vector<binary_container_frame>& stack)
{
if (j.is_array())
{
write_msgpack_array_prefix(j.m_data.m_value.array->size(), j);
if (!j.m_data.m_value.array->empty())
{
stack.emplace_back(&j);
}
return;
}
if (j.is_object())
{
write_msgpack_object_prefix(j.m_data.m_value.object->size(), j);
if (!j.m_data.m_value.object->empty())
{
stack.emplace_back(&j);
}
return;
}
write_msgpack(j);
}
/*!
@brief write out @a root and everything below it without the call stack
@sa @ref write_cbor_iterative
*/
void write_msgpack_iterative(const BasicJsonType& root)
{
std::vector<binary_container_frame> stack;
write_msgpack_value_or_push(root, stack);
while (!stack.empty())
{
const binary_container_frame current = stack.back();
if (current.value->is_array())
{
const auto& array = *current.value->m_data.m_value.array;
if (current.array_it == array.cend())
{
stack.pop_back();
continue;
}
const BasicJsonType* child = &(*current.array_it);
++stack.back().array_it;
write_msgpack_value_or_push(*child, stack);
}
else
{
const auto& object = *current.value->m_data.m_value.object;
if (current.object_it == object.cend())
{
stack.pop_back();
continue;
}
if (error_handler == error_handler_t::strict)
{
check_utf8(current.object_it->first, *current.value);
}
write_msgpack(current.object_it->first);
const BasicJsonType* child = &(current.object_it->second);
++stack.back().object_it;
write_msgpack_value_or_push(*child, stack);
}
}
}
/// @return true when a closing ']' still has to be written after the elements
bool write_ubjson_start_array(const BasicJsonType& j, const bool use_count, const bool use_type,
const bool add_prefix, const bool use_bjdata, bool& prefix_required)
{
prefix_required = true;
if (add_prefix)
{
oa.write_character(to_char_type('['));
}
if (use_type && !j.m_data.m_value.array->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin() + 1, j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
// an optimized array of a valueless type carries no payload, so a
// reader has nothing but the declared count to bound the allocation
// by and refuses an excessive one. Write the unoptimized form for
// those, at one byte per element, so the result can be read back.
// Objects are not affected: every element is preceded by its key.
const bool valueless_type = (first_prefix == 'Z' || first_prefix == 'T' || first_prefix == 'F');
const bool excessive_valueless = valueless_type
&& j.m_data.m_value.array->size() > detail::max_valueless_container_size;
if (same_prefix && !excessive_valueless
&& !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.array->size(), true, use_bjdata);
}
return !use_count;
}
/// @return true when a closing '}' still has to be written after the elements
bool write_ubjson_start_object(const BasicJsonType& j, const bool use_count, const bool use_type,
const bool add_prefix, const bool use_bjdata, bool& prefix_required)
{
prefix_required = true;
if (add_prefix)
{
oa.write_character(to_char_type('{'));
}
if (use_type && !j.m_data.m_value.object->empty())
{
if (!use_count)
{
JSON_THROW(other_error::create(502, "use_type requires use_size = true", &j));
}
const CharType first_prefix = ubjson_prefix(j.front(), use_bjdata);
const bool same_prefix = std::all_of(j.begin(), j.end(),
[this, first_prefix, use_bjdata](const BasicJsonType & v)
{
return ubjson_prefix(v, use_bjdata) == first_prefix;
});
if (same_prefix && !(use_bjdata && is_bjdata_excluded_type_marker(first_prefix)))
{
prefix_required = false;
oa.write_character(to_char_type('$'));
oa.write_character(first_prefix);
}
}
if (use_count)
{
oa.write_character(to_char_type('#'));
write_number_with_ubjson_prefix(j.m_data.m_value.object->size(), true, use_bjdata);
}
return !use_count;
}
/*!
@brief whether @a j is a BJData ND-array annotation object
(https://github.com/NeuroJSON/jdata)
Used by both the recursive object case of @ref write_ubjson and
@ref write_ubjson_value_or_push, which must agree on what counts as an
ND-array: @a j is only actually written as one once @ref
write_bjdata_ndarray has also accepted its contents.
@pre @a j.is_object()
*/
static bool is_bjdata_ndarray(const BasicJsonType& j)
{
const auto& object = *j.m_data.m_value.object;
return object.size() == 3
&& object.find("_ArrayType_") != object.end()
&& object.find("_ArraySize_") != object.end()
&& object.find("_ArrayData_") != object.end();
}
/// @brief an object or array @ref write_ubjson_iterative is still writing
/// the elements of
struct ubjson_frame
{
ubjson_frame(const BasicJsonType* value_, const bool prefix_required_) noexcept
: value(value_)
, prefix_required(prefix_required_)
{
if (value->is_object())
{
object_it = value->m_data.m_value.object->cbegin();
}
else
{
array_it = value->m_data.m_value.array->cbegin();
}
}
// declared for GCC's -Weffc++, which asks for them in a class with
// pointer members and a non-trivial destructor; the exception
// specifications are left implicit, as GCC 4.8 rejects explicit ones
// that differ from them
ubjson_frame(const ubjson_frame&) = default;
ubjson_frame(ubjson_frame&&) = default;
ubjson_frame& operator=(const ubjson_frame&) = default;
ubjson_frame& operator=(ubjson_frame&&) = default;
~ubjson_frame() = default;
/// the array or object being written
const BasicJsonType* value;
/// whether value's elements each carry their own type marker; an
/// optimized ($type) container writes it once for all of them instead
bool prefix_required;
typename BasicJsonType::object_t::const_iterator object_it{};
typename BasicJsonType::array_t::const_iterator array_it{};
};
/*!
@brief write @a j with @ref write_ubjson, or write its header and push a
frame for @ref write_ubjson_iterative to continue with its elements
@param[in] add_prefix whether @a j's own type marker is written now (the
elements of an optimized container, and everything below the
top level, never repeat it)
@sa @ref write_cbor_value_or_push
*/
void write_ubjson_value_or_push(const BasicJsonType& j, const bool add_prefix, const bool use_count,
const bool use_type, const bool use_bjdata, const bjdata_version_t bjdata_version,
std::vector<ubjson_frame>& stack)
{
if (!j.is_array() && !j.is_object())
{
write_ubjson(j, use_count, use_type, add_prefix, use_bjdata, bjdata_version);
return;
}
if (use_bjdata && j.is_object() && is_bjdata_ndarray(j)
&& !write_bjdata_ndarray(*j.m_data.m_value.object, use_count, use_type, bjdata_version))
{
// fully written as an ND-array: nothing below it to come back to
return;
}
const bool is_array = j.is_array();
bool prefix_required = true;
if (is_array)
{
write_ubjson_start_array(j, use_count, use_type, add_prefix, use_bjdata, prefix_required);
}
else
{
write_ubjson_start_object(j, use_count, use_type, add_prefix, use_bjdata, prefix_required);
}
const bool empty = is_array ? j.m_data.m_value.array->empty() : j.m_data.m_value.object->empty();
if (!empty)
{
stack.emplace_back(&j, prefix_required);
return;
}
// write_ubjson_start_array/_object return !use_count, i.e. whether a
// closer still has to be written; use_count is constant for the whole
// document, so that is recomputed here instead of being carried along
if (!use_count)
{
oa.write_character(to_char_type(is_array ? ']' : '}'));
}
}
/*!
@brief write out @a root and everything below it without the call stack
@sa @ref write_cbor_iterative
*/
void write_ubjson_iterative(const BasicJsonType& root, const bool use_count, const bool use_type,
const bool add_prefix, const bool use_bjdata, const bjdata_version_t bjdata_version)
{
std::vector<ubjson_frame> stack;
write_ubjson_value_or_push(root, add_prefix, use_count, use_type, use_bjdata, bjdata_version, stack);
while (!stack.empty())
{
const ubjson_frame current = stack.back();
const BasicJsonType& j = *current.value;
if (j.is_array())
{
const auto& array = *j.m_data.m_value.array;
if (current.array_it == array.cend())
{
if (!use_count)
{
oa.write_character(to_char_type(']'));
}
stack.pop_back();
continue;
}
const BasicJsonType* child = &(*current.array_it);
const bool child_prefix = current.prefix_required;
++stack.back().array_it;
write_ubjson_value_or_push(*child, child_prefix, use_count, use_type, use_bjdata, bjdata_version, stack);
}
else
{
const auto& object = *j.m_data.m_value.object;
if (current.object_it == object.cend())
{
if (!use_count)
{
oa.write_character(to_char_type('}'));
}
stack.pop_back();
continue;
}
string_t storage;
const string_t& key = sanitize_utf8_for_write(current.object_it->first, j, storage);
write_number_with_ubjson_prefix(key.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(key.data()),
key.size());
const BasicJsonType* child = &(current.object_it->second);
const bool child_prefix = current.prefix_required;
++stack.back().object_it;
write_ubjson_value_or_push(*child, child_prefix, use_count, use_type, use_bjdata, bjdata_version, stack);
}
}
}
////////// //////////
// BSON // // BSON //
////////// //////////
+39 -8
View File
@@ -12,8 +12,10 @@
using json = nlohmann::json; using json = nlohmann::json;
using ordered_json = nlohmann::ordered_json; using ordered_json = nlohmann::ordered_json;
#include <limits>
#include <set> #include <set>
#include <string> #include <string>
#include <unordered_set>
namespace namespace
{ {
@@ -91,6 +93,9 @@ TEST_CASE("hash<nlohmann::json>")
// Collect hashes for different JSON values and make sure that they are distinct // Collect hashes for different JSON values and make sure that they are distinct
// We cannot compare against fixed values, because the implementation of // We cannot compare against fixed values, because the implementation of
// std::hash may differ between compilers. // std::hash may differ between compilers.
//
// numbers that compare equal under operator== (0 == 0U == 0.0) must hash
// equally, so they are only inserted once below and checked separately.
std::set<std::size_t> hashes; std::set<std::size_t> hashes;
@@ -107,10 +112,7 @@ TEST_CASE("hash<nlohmann::json>")
// number // number
hashes.insert(std::hash<json> {}(json(0))); hashes.insert(std::hash<json> {}(json(0)));
hashes.insert(std::hash<json> {}(json(static_cast<unsigned>(0))));
hashes.insert(std::hash<json> {}(json(-1))); hashes.insert(std::hash<json> {}(json(-1)));
hashes.insert(std::hash<json> {}(json(0.0)));
hashes.insert(std::hash<json> {}(json(42.23))); hashes.insert(std::hash<json> {}(json(42.23)));
// array // array
@@ -132,7 +134,36 @@ TEST_CASE("hash<nlohmann::json>")
// discarded // discarded
hashes.insert(std::hash<json> {}(json(json::value_t::discarded))); hashes.insert(std::hash<json> {}(json(json::value_t::discarded)));
CHECK(hashes.size() == 21); CHECK(hashes.size() == 19);
// numbers that compare equal under operator== must hash equally,
// regardless of which of number_integer, number_unsigned, or
// number_float actually holds the value
CHECK(json(0) == json(static_cast<unsigned>(0)));
CHECK(json(0) == json(0.0));
CHECK(std::hash<json> {}(json(0)) == std::hash<json> {}(json(static_cast<unsigned>(0))));
CHECK(std::hash<json> {}(json(0)) == std::hash<json> {}(json(0.0)));
CHECK(std::hash<json> {}(json(-1)) == std::hash<json> {}(json(-1.0)));
// a std::unordered_set relies on this same consistency between == and hash
const std::unordered_set<json> numbers {json(0), json(static_cast<unsigned>(0)), json(0.0)};
CHECK(numbers.size() == 1);
// -0.0 compares equal to 0 and 0.0
CHECK(json(-0.0) == json(0));
CHECK(std::hash<json> {}(json(-0.0)) == std::hash<json> {}(json(0)));
CHECK(std::hash<json> {}(json(-0.0)) == std::hash<json> {}(json(0.0)));
// the ends of the integer ranges, which equal floats exactly
const auto int_min = (std::numeric_limits<json::number_integer_t>::min)();
const auto int_max = (std::numeric_limits<json::number_integer_t>::max)();
const auto two_63 = json::number_unsigned_t(1) << 63U;
CHECK(json(int_min) == json(-9223372036854775808.0));
CHECK(std::hash<json> {}(json(int_min)) == std::hash<json> {}(json(-9223372036854775808.0)));
CHECK(json(two_63) == json(9223372036854775808.0));
CHECK(std::hash<json> {}(json(two_63)) == std::hash<json> {}(json(9223372036854775808.0)));
CHECK(json(json::number_unsigned_t(int_max)) == json(int_max));
CHECK(std::hash<json> {}(json(json::number_unsigned_t(int_max))) == std::hash<json> {}(json(int_max)));
} }
TEST_CASE("hash<nlohmann::ordered_json>") TEST_CASE("hash<nlohmann::ordered_json>")
@@ -156,10 +187,7 @@ TEST_CASE("hash<nlohmann::ordered_json>")
// number // number
hashes.insert(std::hash<ordered_json> {}(ordered_json(0))); hashes.insert(std::hash<ordered_json> {}(ordered_json(0)));
hashes.insert(std::hash<ordered_json> {}(ordered_json(static_cast<unsigned>(0))));
hashes.insert(std::hash<ordered_json> {}(ordered_json(-1))); hashes.insert(std::hash<ordered_json> {}(ordered_json(-1)));
hashes.insert(std::hash<ordered_json> {}(ordered_json(0.0)));
hashes.insert(std::hash<ordered_json> {}(ordered_json(42.23))); hashes.insert(std::hash<ordered_json> {}(ordered_json(42.23)));
// array // array
@@ -181,7 +209,10 @@ TEST_CASE("hash<nlohmann::ordered_json>")
// discarded // discarded
hashes.insert(std::hash<ordered_json> {}(ordered_json(ordered_json::value_t::discarded))); hashes.insert(std::hash<ordered_json> {}(ordered_json(ordered_json::value_t::discarded)));
CHECK(hashes.size() == 21); CHECK(hashes.size() == 19);
CHECK(std::hash<ordered_json> {}(ordered_json(0)) == std::hash<ordered_json> {}(ordered_json(static_cast<unsigned>(0))));
CHECK(std::hash<ordered_json> {}(ordered_json(0)) == std::hash<ordered_json> {}(ordered_json(0.0)));
} }
TEST_CASE("hash of deeply nested values") TEST_CASE("hash of deeply nested values")
-197
View File
@@ -13,7 +13,6 @@ using nlohmann::json;
#include <algorithm> #include <algorithm>
#include <string> #include <string>
#include <utility>
#include <vector> #include <vector>
TEST_CASE("tests on very large JSONs") TEST_CASE("tests on very large JSONs")
@@ -355,199 +354,3 @@ TEST_CASE("tests on deeply nested JSONs")
} }
} }
namespace
{
json nested_array(const std::size_t depth, json leaf)
{
json j = std::move(leaf);
for (std::size_t i = 0; i < depth; ++i)
{
json a = json::array();
a.push_back(std::move(j));
j = std::move(a);
}
return j;
}
json nested_object(const std::size_t depth, json leaf)
{
json j = std::move(leaf);
for (std::size_t i = 0; i < depth; ++i)
{
json o = json::object();
o["k"] = std::move(j);
j = std::move(o);
}
return j;
}
} // namespace
TEST_CASE("issue #5392 - binary writers on deeply nested values")
{
// 200 is past the point where the writers stop recursing, and still
// shallow enough that from_* and operator== (which still recurse) are fine.
const json deep_array = nested_array(200, json(0));
const json deep_object = nested_object(200, json("x"));
const json empty_array = nested_array(200, json::array());
const json empty_object = nested_object(200, json::object());
const json mixed = nested_object(80, nested_array(80, json(true)));
SECTION("roundtrip past the recursion bound")
{
CHECK(json::from_cbor(json::to_cbor(deep_array)) == deep_array);
CHECK(json::from_msgpack(json::to_msgpack(deep_array)) == deep_array);
CHECK(json::from_ubjson(json::to_ubjson(deep_array)) == deep_array);
CHECK(json::from_ubjson(json::to_ubjson(deep_array, true, false)) == deep_array);
CHECK(json::from_ubjson(json::to_ubjson(deep_array, true, true)) == deep_array);
CHECK(json::from_bjdata(json::to_bjdata(deep_array)) == deep_array);
CHECK(json::from_cbor(json::to_cbor(deep_object)) == deep_object);
CHECK(json::from_msgpack(json::to_msgpack(deep_object)) == deep_object);
CHECK(json::from_ubjson(json::to_ubjson(deep_object)) == deep_object);
CHECK(json::from_ubjson(json::to_ubjson(deep_object, true, true)) == deep_object);
CHECK(json::from_bjdata(json::to_bjdata(deep_object)) == deep_object);
CHECK(json::from_cbor(json::to_cbor(empty_array)) == empty_array);
CHECK(json::from_msgpack(json::to_msgpack(empty_array)) == empty_array);
CHECK(json::from_ubjson(json::to_ubjson(empty_array)) == empty_array);
CHECK(json::from_ubjson(json::to_ubjson(empty_array, true, true)) == empty_array);
CHECK(json::from_cbor(json::to_cbor(empty_object)) == empty_object);
CHECK(json::from_msgpack(json::to_msgpack(empty_object)) == empty_object);
CHECK(json::from_ubjson(json::to_ubjson(empty_object)) == empty_object);
CHECK(json::from_cbor(json::to_cbor(mixed)) == mixed);
CHECK(json::from_msgpack(json::to_msgpack(mixed)) == mixed);
CHECK(json::from_ubjson(json::to_ubjson(mixed)) == mixed);
CHECK(json::from_bjdata(json::to_bjdata(mixed)) == mixed);
}
SECTION("the two ways of writing a value meet at the bound")
{
for (std::size_t depth = 120; depth <= 140; ++depth)
{
CAPTURE(depth);
const json array = nested_array(depth, json(7));
CHECK(json::from_cbor(json::to_cbor(array)) == array);
CHECK(json::from_msgpack(json::to_msgpack(array)) == array);
CHECK(json::from_ubjson(json::to_ubjson(array, true, true)) == array);
const json object = nested_object(depth, json(7));
CHECK(json::from_cbor(json::to_cbor(object)) == object);
CHECK(json::from_msgpack(json::to_msgpack(object)) == object);
CHECK(json::from_bjdata(json::to_bjdata(object)) == object);
}
}
SECTION("a BJData ndarray below the bound is still an ndarray")
{
const json ndarray = json({{"_ArrayType_", "uint8"}, {"_ArraySize_", {2, 3}}, {"_ArrayData_", {1, 2, 3, 4, 5, 6}}});
const json invalid = json({{"_ArrayType_", "nope"}, {"_ArraySize_", {1}}, {"_ArrayData_", {1}}});
const json deep_ndarray = nested_array(140, ndarray);
const json deep_invalid = nested_array(140, invalid);
CHECK(json::from_bjdata(json::to_bjdata(deep_ndarray)) == deep_ndarray);
CHECK(json::from_bjdata(json::to_bjdata(deep_invalid)) == deep_invalid);
CHECK(json::from_bjdata(json::to_bjdata(ndarray)) == ndarray);
}
SECTION("byte-exact across the switch-over")
{
// nested one-element arrays around the recursion bound: the exact
// bytes a writer produces do not depend on whether it stayed on the
// call stack or moved to the heap one partway through
for (const std::size_t depth :
{
nlohmann::detail::recursion_depth_limit() - 1, nlohmann::detail::recursion_depth_limit(),
nlohmann::detail::recursion_depth_limit() + 1, nlohmann::detail::recursion_depth_limit() + 2
})
{
CAPTURE(depth);
const json array = nested_array(depth, json(0));
std::vector<std::uint8_t> expected_cbor(depth, 0x81);
expected_cbor.push_back(0x00);
CHECK(json::to_cbor(array) == expected_cbor);
std::vector<std::uint8_t> expected_msgpack(depth, 0x91);
expected_msgpack.push_back(0x00);
CHECK(json::to_msgpack(array) == expected_msgpack);
std::string expected_ubjson(depth, '[');
expected_ubjson += "i";
expected_ubjson += '\0';
expected_ubjson.append(depth, ']');
const auto packed_ubjson = json::to_ubjson(array);
CHECK(std::string(packed_ubjson.begin(), packed_ubjson.end()) == expected_ubjson);
}
}
SECTION("a deep object, and a BJData ndarray, past the recursion bound")
{
const std::size_t depth = nlohmann::detail::recursion_depth_limit() + 50;
const json object = nested_object(depth, json(42));
CHECK(json::from_cbor(json::to_cbor(object)) == object);
CHECK(json::from_msgpack(json::to_msgpack(object)) == object);
CHECK(json::from_ubjson(json::to_ubjson(object, true, true)) == object);
CHECK(json::from_bjdata(json::to_bjdata(object)) == object);
const json ndarray = json({{"_ArrayType_", "uint8"}, {"_ArraySize_", {2, 3}}, {"_ArrayData_", {1, 2, 3, 4, 5, 6}}});
const json deep_ndarray = nested_array(depth, ndarray);
CHECK(json::from_bjdata(json::to_bjdata(deep_ndarray)) == deep_ndarray);
}
SECTION("a discarded value past the recursion bound still throws type_error.321")
{
const std::size_t depth = nlohmann::detail::recursion_depth_limit() + 50;
const json discarded_leaf(json::value_t::discarded);
const json deep_discarded = nested_array(depth, discarded_leaf);
CHECK_THROWS_WITH_AS(json::to_cbor(deep_discarded), "[json.exception.type_error.321] cannot serialize discarded value to CBOR", json::type_error);
CHECK_THROWS_WITH_AS(json::to_msgpack(deep_discarded), "[json.exception.type_error.321] cannot serialize discarded value to MessagePack", json::type_error);
CHECK_THROWS_WITH_AS(json::to_ubjson(deep_discarded), "[json.exception.type_error.321] cannot serialize discarded value to UBJSON", json::type_error);
CHECK_THROWS_WITH_AS(json::to_bjdata(deep_discarded), "[json.exception.type_error.321] cannot serialize discarded value to BJData", json::type_error);
}
SECTION("does not overflow the C++ stack")
{
const std::size_t depth = 100000;
const json j = json::parse(std::string(depth, '[') + "0" + std::string(depth, ']'));
std::vector<std::uint8_t> packed;
CHECK_NOTHROW(packed = json::to_cbor(j));
CHECK(json::from_cbor(packed) == j);
CHECK_NOTHROW(packed = json::to_msgpack(j));
CHECK(json::from_msgpack(packed) == j);
CHECK_NOTHROW(packed = json::to_ubjson(j));
CHECK(json::from_ubjson(packed) == j);
CHECK_NOTHROW(packed = json::to_ubjson(j, true, false));
CHECK(json::from_ubjson(packed) == j);
CHECK_NOTHROW(packed = json::to_bjdata(j));
CHECK(json::from_bjdata(packed) == j);
}
SECTION("regression test for https://issues.oss-fuzz.com/issues/566583014")
{
// 200000 nested one-element CBOR arrays, the innermost holding null;
// round-tripping this used to recurse once per level on the way back
// out through to_cbor(), deep enough to overflow the stack
std::vector<std::uint8_t> v(200000, 0x81);
v.push_back(0xf6);
const json j = json::from_cbor(v);
CHECK(json::to_cbor(j) == v);
// the MessagePack analogue: fixarray of 1 nesting down to nil
std::vector<std::uint8_t> v_msgpack(200000, 0x91);
v_msgpack.push_back(0xc0);
const json j_msgpack = json::from_msgpack(v_msgpack);
CHECK(json::to_msgpack(j_msgpack) == v_msgpack);
}
}