Bound UBJSON optimized arrays of a valueless type

An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized UBJSON array of one of those has no
payload: reading an element consumes no input at all. Its declared count is
therefore the only thing that decides how much is allocated, and nothing
bounded it. "[$Z#l" and a four-byte count is nine bytes of input describing
two billion values; #2793 reports 35 GB and 150 seconds from ten bytes, and
OSS-Fuzz has an out-of-memory and a timeout report for the same shape.

Every other type costs at least one byte per element, so the end of the input
bounds it. 'N' (no-op) is already skipped rather than stored. Objects are not
affected either: each element is preceded by its key, which costs bytes. And
BJData already refuses these markers as an optimized type, so this is a plain
UBJSON matter.

Reject a count above 1,048,576 elements for those three types with
out_of_range.408, the code this reader already uses for a declared size it
will not honour. The check runs before the SAX start event, so no container
is opened and then abandoned.

Rejecting on the read side alone would break the guarantee that anything
to_ubjson() writes can be read back, and would trip the round-trip assertion
in fuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized
encoding, one byte per element, for arrays of these types above the same
limit. Its decision depends only on the array's size, which is identical for
a value and for anything parsed back from it, so the round trip is stable.

No existing test changes: the largest such count in the test suite is 65,793.
The excessive-size test that already used this shape still passes, now
rejected a little earlier than by the max_size() check it used to reach.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-06 18:12:41 +02:00
parent bcd5af62c7
commit d89acce09a
6 changed files with 154 additions and 2 deletions
@@ -58,6 +58,26 @@ inline bool little_endianness(int num = 1) noexcept
return *reinterpret_cast<char*>(&num) == 1;
}
/*!
@brief largest element count accepted for a UBJSON container of a valueless type
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized container of one of those types has no
payload at all and its declared count is the only thing that decides how much
is allocated: `[$Z#L` followed by a large count turns some ten bytes of input
into that many values (see #2793, which reports 35 GB and 150 seconds). Every
other type costs at least one byte per element and is bounded by the end of
the input.
This is a sanity bound rather than a security boundary, and it is far above
any container met in practice. @ref binary_writer falls back to the
unoptimized encoding for longer containers, so that a value serialized by
this library can always be read back.
@sa https://github.com/nlohmann/json/issues/2793
*/
JSON_INLINE_VARIABLE constexpr std::size_t max_valueless_container_size = 1 << 20;
///////////////////
// binary reader //
///////////////////
@@ -2799,6 +2819,17 @@ class binary_reader
if (size_and_type.first != npos)
{
// reading an element of a valueless type consumes no input, so the
// declared count alone decides how much is allocated; the check is
// made before the start event so that no container is opened that
// is then abandoned. See @ref max_valueless_container_size.
if (JSON_HEDLEY_UNLIKELY((size_and_type.second == 'Z' || size_and_type.second == 'T' || size_and_type.second == 'F')
&& size_and_type.first > max_valueless_container_size))
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408,
exception_message(input_format, "excessive array size", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(size_and_type.first)))
{
return false;