Compare commits

..
Author SHA1 Message Date
Niels Lohmann af394fca9c Share one nesting depth limit between all bounded descents
Copying and comparing stopped their descent at
basic_json::nesting_depth_limit(), while serializing, hashing and
merging used detail::recursion_depth_limit(). Both were 128, but
nothing kept them equal. The thread-local count now tests against
detail::recursion_depth_limit() as well, and a static_assert keeps
the limit small enough for the byte that holds the count.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 16:40:54 +02:00
5 changed files with 40 additions and 98 deletions
@@ -445,10 +445,8 @@ struct wide_string_input_helper<BaseInputAdapter, 4>
}
else
{
// get the current character; converted to an unsigned type so that
// a negative unit (wint_t is signed on some platforms) is not
// mistaken for an ASCII character or for EOF
const auto wc = static_cast<std::uint32_t>(input.get_character());
// get the current character
const auto wc = input.get_character();
// UTF-32 to UTF-8 encoding
if (wc < 0x80)
@@ -543,11 +541,9 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{
// only consume the next unit if it completes the pair
const auto wc2 = static_cast<unsigned int>(*input.current);
const auto wc2 = static_cast<unsigned int>(input.get_character());
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
{
input.get_character();
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
@@ -560,8 +556,7 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
if (!valid_pair)
{
// emit a byte that is never valid UTF-8 (see the UTF-32 case)
utf8_bytes[0] = 0xFF;
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
}
}
@@ -19,11 +19,17 @@ namespace detail
/*!
@brief the number of nesting levels an operation recurses into
Operations that walk a value (serializing, hashing, merging, ...) recurse once
Operations that walk a value (copying, comparing, serializing, hashing, merging,
...) recurse once
per nesting level, which is fastest, but a value nested deeply enough would
exhaust the call stack. So they recurse only this many levels deep and finish
whatever lies below with an explicit stack. All of them share this limit.
Most of them pass the depth down as an argument. The copy constructor and the
comparison operators cannot, as their signatures are fixed, so they count it
in basic_json::nesting_depth() instead, a byte per thread; the limit must
therefore stay below 255.
@sa https://github.com/nlohmann/json/issues/5387
*/
constexpr std::size_t recursion_depth_limit() noexcept
+7 -11
View File
@@ -898,12 +898,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
}
#ifndef JSON_NO_THREAD_LOCAL
/// the number of levels an operation descends into before it finishes the
/// value below it without the call stack
static constexpr std::uint8_t nesting_depth_limit()
{
return 128;
}
// nesting_depth() is a byte and may exceed the limit by one level
static_assert(detail::recursion_depth_limit() < 255, "the nesting depth count must fit in a byte");
/*!
@brief how many levels the operation going on in this thread has descended into
@@ -945,7 +941,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static_cast<void>(may_descend);
return true;
#else
return !may_descend || nesting_depth() >= nesting_depth_limit();
return !may_descend || nesting_depth() >= detail::recursion_depth_limit();
#endif
}
@@ -969,7 +965,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
#ifdef JSON_NO_THREAD_LOCAL
: m_okay(false)
#else
: m_okay(nesting_depth() < nesting_depth_limit())
: m_okay(nesting_depth() < detail::recursion_depth_limit())
#endif
{
#ifndef JSON_NO_THREAD_LOCAL
@@ -1173,7 +1169,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
The values whose copy has not been created yet are kept on an explicit
worklist rather than on the call stack. This is only reached for values
nested deeper than @ref nesting_depth_limit levels, which is why it copies
nested deeper than @ref detail::recursion_depth_limit levels, which is why it copies
every container by hand instead of letting the container do it: the fast
ways of doing so would descend into the elements and defeat the purpose.
*/
@@ -1239,7 +1235,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
Copying a container copies its elements, so a value nested deeply enough
used to exhaust the call stack. The descent is bounded here: the first
@ref nesting_depth_limit levels are copied by the containers themselves, just
@ref detail::recursion_depth_limit levels are copied by the containers themselves, just
as they always were, and anything below that is copied without the call
stack by @ref copy_iteratively. Copying a value can therefore no longer
exhaust the stack, however deeply it is nested, just like destroying one
@@ -1377,7 +1373,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/*!
@brief compare @a lhs and @a rhs without descending into them
Reached once a comparison has descended @ref nesting_depth_limit levels, so
Reached once a comparison has descended @ref detail::recursion_depth_limit levels, so
that comparing values cannot exhaust the call stack however deeply they are
nested. The two values are walked in lockstep on an explicit stack and
compared lexicographically, element by element in the order the containers
+18 -21
View File
@@ -7281,11 +7281,17 @@ namespace detail
/*!
@brief the number of nesting levels an operation recurses into
Operations that walk a value (serializing, hashing, merging, ...) recurse once
Operations that walk a value (copying, comparing, serializing, hashing, merging,
...) recurse once
per nesting level, which is fastest, but a value nested deeply enough would
exhaust the call stack. So they recurse only this many levels deep and finish
whatever lies below with an explicit stack. All of them share this limit.
Most of them pass the depth down as an argument. The copy constructor and the
comparison operators cannot, as their signatures are fixed, so they count it
in basic_json::nesting_depth() instead, a byte per thread; the limit must
therefore stay below 255.
@sa https://github.com/nlohmann/json/issues/5387
*/
constexpr std::size_t recursion_depth_limit() noexcept
@@ -7989,10 +7995,8 @@ struct wide_string_input_helper<BaseInputAdapter, 4>
}
else
{
// get the current character; converted to an unsigned type so that
// a negative unit (wint_t is signed on some platforms) is not
// mistaken for an ASCII character or for EOF
const auto wc = static_cast<std::uint32_t>(input.get_character());
// get the current character
const auto wc = input.get_character();
// UTF-32 to UTF-8 encoding
if (wc < 0x80)
@@ -8087,11 +8091,9 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{
// only consume the next unit if it completes the pair
const auto wc2 = static_cast<unsigned int>(*input.current);
const auto wc2 = static_cast<unsigned int>(input.get_character());
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
{
input.get_character();
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
@@ -8104,8 +8106,7 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
if (!valid_pair)
{
// emit a byte that is never valid UTF-8 (see the UTF-32 case)
utf8_bytes[0] = 0xFF;
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
}
}
@@ -26984,12 +26985,8 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
}
#ifndef JSON_NO_THREAD_LOCAL
/// the number of levels an operation descends into before it finishes the
/// value below it without the call stack
static constexpr std::uint8_t nesting_depth_limit()
{
return 128;
}
// nesting_depth() is a byte and may exceed the limit by one level
static_assert(detail::recursion_depth_limit() < 255, "the nesting depth count must fit in a byte");
/*!
@brief how many levels the operation going on in this thread has descended into
@@ -27031,7 +27028,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static_cast<void>(may_descend);
return true;
#else
return !may_descend || nesting_depth() >= nesting_depth_limit();
return !may_descend || nesting_depth() >= detail::recursion_depth_limit();
#endif
}
@@ -27055,7 +27052,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
#ifdef JSON_NO_THREAD_LOCAL
: m_okay(false)
#else
: m_okay(nesting_depth() < nesting_depth_limit())
: m_okay(nesting_depth() < detail::recursion_depth_limit())
#endif
{
#ifndef JSON_NO_THREAD_LOCAL
@@ -27259,7 +27256,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
The values whose copy has not been created yet are kept on an explicit
worklist rather than on the call stack. This is only reached for values
nested deeper than @ref nesting_depth_limit levels, which is why it copies
nested deeper than @ref detail::recursion_depth_limit levels, which is why it copies
every container by hand instead of letting the container do it: the fast
ways of doing so would descend into the elements and defeat the purpose.
*/
@@ -27325,7 +27322,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
Copying a container copies its elements, so a value nested deeply enough
used to exhaust the call stack. The descent is bounded here: the first
@ref nesting_depth_limit levels are copied by the containers themselves, just
@ref detail::recursion_depth_limit levels are copied by the containers themselves, just
as they always were, and anything below that is copied without the call
stack by @ref copy_iteratively. Copying a value can therefore no longer
exhaust the stack, however deeply it is nested, just like destroying one
@@ -27463,7 +27460,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/*!
@brief compare @a lhs and @a rhs without descending into them
Reached once a comparison has descended @ref nesting_depth_limit levels, so
Reached once a comparison has descended @ref detail::recursion_depth_limit levels, so
that comparing values cannot exhaust the call stack however deeply they are
nested. The two values are walked in lockstep on an explicit stack and
compared lexicographically, element by element in the order the containers
+4 -56
View File
@@ -8,7 +8,6 @@
#include "doctest_compatibility.h"
#include <cwchar>
#include <nlohmann/json.hpp>
using nlohmann::json;
@@ -99,15 +98,15 @@ TEST_CASE("wide strings")
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// a lone low surrogate cannot start a pair
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xFF'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a high surrogate followed by a non-low-surrogate unit is invalid
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xD800, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xFF'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xD800, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// ... also when the unit is above the low surrogates
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xD800, 0xE000, u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xFF'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xD800, 0xE000, u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a lone low surrogate must not swallow the following unit: pairing
// it with any second unit would produce valid UTF-8, so the error
// has to report an ill-formed byte at the surrogate's own position
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xFF'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a valid surrogate pair is still decoded (U+1F600)
CHECK(json::parse(std::u16string{u'"', 0xD83D, 0xDE00, u'"'}).get<std::string>() == "\xF0\x9F\x98\x80");
}
@@ -142,56 +141,5 @@ TEST_CASE("wide strings")
CHECK_THROWS_WITH_AS(_ = json::parse(std::u32string{U'"', static_cast<char32_t>(0xFFFFFFFF), U'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xFF'", json::parse_error&);
}
}
SECTION("malformed wide-string input outside strings (#5645)")
{
json _;
// a lone low surrogate inside a literal must not be truncated to its
// low byte and mistaken for the letter the literal expects next
// (0xDC72 truncates to 'r', which is what "true" expects after 't')
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u't', static_cast<char16_t>(0xDC72), u'u', u'e'}),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid literal; last read: 't\xFF'", json::parse_error&);
// ... also when the lone surrogate is the last unit of the input
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'f', u'a', u'l', u's', static_cast<char16_t>(0xDD65)}),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: 'fals\xFF'", json::parse_error&);
// a high surrogate followed by a unit that is not its low surrogate
// must not silently swallow that unit
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u't', static_cast<char16_t>(0xD872), u'X', u'u', u'e'}),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid literal; last read: 't\xFF'", json::parse_error&);
// ... in particular, if the swallowed unit is the newline that ends a
// // comment, the comment must not extend over the following line
CHECK(json::parse(std::u16string{u'[', u'1', u' ', u'/', u'/', static_cast<char16_t>(0xD800), u'\n',
u',', u'2', u' ', u'/', u'/', u'\n', u']'},
nullptr, true, /*ignore_comments*/true) == json::parse("[1,2]"));
CHECK(json::accept(std::u16string{u'[', u'1', u' ', u'/', u'/', static_cast<char16_t>(0xD800), u'\n',
u',', u'2', u' ', u'/', u'/', u'\n', u']'}, /*ignore_comments*/true));
// cases 5 and 6 use a 32-bit wchar_t (Linux, macOS, the BSDs) to reach
// the UTF-32 helper tested above via u32string; the 16-bit wchar_t of
// Windows goes through the UTF-16 helper instead, already covered by
// the u16string cases above
#if WCHAR_MAX > 0xFFFFu
// a negative wchar_t must not be mistaken for
// char_traits<char>::eof() and silently end the input, letting
// trailing garbage pass the strict end-of-input check (only observable
// where wint_t is signed, e.g. macOS/the BSDs; on Linux wint_t is
// unsigned and this was already handled by #5348)
std::wstring w = L"[1]";
w.push_back(static_cast<wchar_t>(-1));
w += L"garbage";
CHECK(!json::accept(w));
CHECK_THROWS_WITH_AS(_ = json::parse(w),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: '1]\xFF'; expected end of input", json::parse_error&);
// other negative wchar_t units must not be truncated to their low
// byte (0xFFFFFF72 truncates to 'r', as in the u16string case above)
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L't', static_cast<wchar_t>(0xFFFFFF72), L'u', L'e'}),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid literal; last read: 't\xFF'", json::parse_error&);
#endif
}
}
#endif