Compare commits

..
Author SHA1 Message Date
Niels Lohmann 5b26411166 Read MessagePack containers without recursing per nesting level
get_msgpack_array() and get_msgpack_object() read their elements by calling
back into parse_msgpack_internal(), which calls them again for a nested
container. The native call stack therefore grew with the nesting depth of the
input, and each level costs only one byte to encode: 0x91 is a one-element
array, so a few hundred thousand of them crash the process before any of the
input is rejected (#5104).

Keep the open containers on a heap stack instead, the way
parser::sax_parse_internal() has always done for JSON text. A frame records
how many elements are left and whether to close with end_object() or
end_array(); parse_msgpack_value() reads a single value and, for a container,
only opens it; and parse_msgpack_internal() loops, resuming the innermost
container after each element and closing it when its count runs out. Whether
the value that was begun is complete is answered by the stack being empty, so
no separate bookkeeping is needed.

The switch that decodes a value is untouched apart from the six container
cases, which now call enter_container() rather than a reader that loops. That
keeps this diff to the control flow and leaves the decoding of every other
type byte-identical.

enter_container() is the only place a binary reader emits start_object() or
start_array(), so a check that rejects a container can be added there once and
is guaranteed to run before the start event. The frame type and the stack are
shared, ready for the other three formats.

Verified against develop over empty, nested, counted (array 16/32, map 16/32)
and truncated inputs: identical values, error codes, messages and byte
offsets. 300,000 levels now report parse_error.110 instead of crashing, and a
well-formed 300,000-level value is read to completion through the SAX
interface, where develop crashes.

Reading such a value into a basic_json needs the return-by-move change as
well, without which the recursive copy constructor overflows on the way out;
that is the parent commit, and the test for the value path covers the two
together. Timing is unchanged: parsing 60,000 small objects and one array of
a million integers is within run-to-run noise of develop either way.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-06 18:12:42 +02:00
Niels Lohmann d89acce09a Bound UBJSON optimized arrays of a valueless type
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized UBJSON array of one of those has no
payload: reading an element consumes no input at all. Its declared count is
therefore the only thing that decides how much is allocated, and nothing
bounded it. "[$Z#l" and a four-byte count is nine bytes of input describing
two billion values; #2793 reports 35 GB and 150 seconds from ten bytes, and
OSS-Fuzz has an out-of-memory and a timeout report for the same shape.

Every other type costs at least one byte per element, so the end of the input
bounds it. 'N' (no-op) is already skipped rather than stored. Objects are not
affected either: each element is preceded by its key, which costs bytes. And
BJData already refuses these markers as an optimized type, so this is a plain
UBJSON matter.

Reject a count above 1,048,576 elements for those three types with
out_of_range.408, the code this reader already uses for a declared size it
will not honour. The check runs before the SAX start event, so no container
is opened and then abandoned.

Rejecting on the read side alone would break the guarantee that anything
to_ubjson() writes can be read back, and would trip the round-trip assertion
in fuzzer-parse_ubjson.cpp. So the writer falls back to the unoptimized
encoding, one byte per element, for arrays of these types above the same
limit. Its decision depends only on the array's size, which is identical for
a value and for anything parsed back from it, so the round trip is stable.

No existing test changes: the largest such count in the test suite is 65,793.
The excessive-size test that already used this shape still passes, now
rejected a little earlier than by the max_size() check it used to reach.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-06 18:12:41 +02:00
Niels Lohmann bcd5af62c7 Reject a nested BJData ndarray dimension vector where it is read
get_ubjson_size_type() takes an inside_ndarray parameter saying whether it is
being called for an ndarray's dimension vector, where another ndarray is not
allowed. It then seeded the flag it passes down to get_ubjson_size_value()
with `false` rather than with that parameter, and only consulted
inside_ndarray afterwards, on the '$' branch.

So on the '#' branch nothing stopped the descent: every "#[" pair of an input
like "[" followed by "#[#[#[..." opened another dimension vector, several
native stack frames deeper each time, and the recursion was only reported on
the way back out. 100,000 pairs crash the process. This is #5104 again, in a
path that has nothing to do with containers.

Seed the flag with inside_ndarray, which is what get_ubjson_size_value()
documents it wants: "for input, `true` means already inside an ndarray vector
or ndarray dimension is not allowed". The nested '[' is then refused where it
is read, so the length of the chain no longer matters.

Both post-checks gain `&& !inside_ndarray`, because an ndarray was found
*here* only if the flag flipped -- get_ubjson_size_value() only ever returns
`true` when its initial value was `false`, as its documentation says. With
that, the "ndarray can not be recursive" branch is unreachable: a recursive
ndarray is now caught one level earlier, and reported as "ndarray dimensional
vector is not allowed" like every other nested dimension vector.

Three existing expectations move accordingly (vR2, vR4, vR6). All three now
fail earlier, and all three now report the same error that vR1, vR5 and vH
already reported for the same shape, which is the more consistent outcome.
Everything else is unchanged: valid 1D and 2D ndarrays, optimized containers
and plain arrays produce identical results, and unit-ubjson is untouched.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-06 18:12:40 +02:00
Niels Lohmann 25a4333a31 Stop CBOR indefinite-length strings from recursing per chunk
get_cbor_string() and get_cbor_binary() handled the indefinite-length forms
(0x7F and 0x5F) by calling themselves once per chunk. Each chunk therefore
cost a native stack frame, and since a chunk may itself be an indefinite-
length string, an input of repeated 0x7F bytes reached one frame per input
byte: 200,000 of them crash the process with SIGSEGV before a single byte is
rejected. This is the same defect as #5104, in a path the container-level
work does not touch.

Count the open levels instead of recursing through them. That is enough here
because every chunk is appended to the same result -- get_bytes() writes at
result.size() -- so there is no per-level state to keep. The temporary chunk
string and its copy into the result go away with the recursion.

The definite-length cases move to get_cbor_string_chunk() and
get_cbor_binary_chunk() unchanged, including their error messages, which
still name 0x7F and 0x5F because those are handled one level up.

Behaviour is unchanged. Comparing against develop over the interesting byte
sequences -- empty, single-chunk, nested, over-closed and truncated forms,
both strings and byte arrays, and an indefinite-length map key -- produces
identical values, error codes, messages and byte offsets. The 200,000-level
input now reports parse_error.110 at byte 200001 instead of crashing.

Note that nesting these is not valid CBOR: RFC 8949, Section 3.2.3 forbids
it. This does not change that either way -- it has always been accepted, and
rejecting it is a separate decision (#5317, #5325). Should it be rejected
later, that is now one condition on the level counter rather than a change to
the control flow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-06 18:12:39 +02:00
Niels Lohmann c3219fdc30 Return the parsed value by move from from_cbor() and friends
The binary entry points end with

    return res ? result : basic_json(value_t::discarded);

The condition operator's second operand is an lvalue, so this is not a case
where the return value can be elided or implicitly moved from: every
successful from_cbor(), from_msgpack(), from_ubjson(), from_bjdata() and
from_bson() call deep-copies the value it just parsed, and then destroys the
original.

The copy is not cheap, and it is not incidental: basic_json's copy
constructor walks the whole value. Parsing a 2 MB CBOR document with 60,000
objects, median of 25 runs, clang 17 -O3:

    from_cbor      26.99 ms  ->  14.65 ms
    from_msgpack   26.82 ms  ->  14.82 ms

Moving instead of copying is the entire change; the parsed value is not used
again after the return expression is evaluated.

There is a second reason to prefer the move. The copy constructor recurses
once per nesting level, so the copy is also a stack-overflow path on the
return side, on a value the reader has already accepted. That is currently
masked because the readers themselves recurse and overflow first (#5104), but
it has to be fixed for making them iterative to have any effect.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-06 16:41:35 +02:00
10 changed files with 857 additions and 241 deletions
@@ -69,6 +69,12 @@ The library uses the following mapping from JSON values types to UBJSON types ac
Note that `use_size = true` alone may result in larger representations - the benefit of this parameter is that the
receiving side is immediately informed on the number of elements of the container.
An array whose type marker is `Z` (null), `T` (true) or `F` (false) stores no payload at all, because the marker
already is the value. Its declared count is therefore the only thing that decides how much memory the receiving side
allocates, and a handful of bytes can describe billions of elements. `from_ubjson` rejects such an array with
[`out_of_range.408`](../../home/exceptions.md#jsonexceptionout_of_range408) when the count exceeds 1,048,576, and
`to_ubjson` writes longer arrays of these types without the annotation, so any value it produces can be read back.
!!! info "Binary values"
If the JSON data contains the binary type, the value stored is a list of integers, as suggested by the UBJSON
+9
View File
@@ -868,6 +868,12 @@ The size of an array or object in a [binary format](../features/binary_formats/i
the size following `#` for [UBJSON](../features/binary_formats/ubjson.md)/[BJData](../features/binary_formats/bjdata.md),
or the encoded length for [CBOR](../features/binary_formats/cbor.md).
The exception is also thrown for a [UBJSON](../features/binary_formats/ubjson.md) array of a type that is encoded by its
marker alone (`Z`, `T` or `F`) whose declared count exceeds 1,048,576. Such an array has no payload, so its count alone
decides how much memory is allocated, and a handful of bytes would otherwise describe billions of values.
[`to_ubjson`](../api/basic_json/to_ubjson.md) writes longer arrays of these types without the size and type annotation,
so any value it produces can still be read back.
!!! failure "Example messages"
```
@@ -879,6 +885,9 @@ or the encoded length for [CBOR](../features/binary_formats/cbor.md).
```
[json.exception.out_of_range.408] syntax error while parsing CBOR size: excessive map size
```
```
[json.exception.out_of_range.408] syntax error while parsing UBJSON size: excessive array size
```
### json.exception.out_of_range.409
+303 -104
View File
@@ -58,6 +58,26 @@ inline bool little_endianness(int num = 1) noexcept
return *reinterpret_cast<char*>(&num) == 1;
}
/*!
@brief largest element count accepted for a UBJSON container of a valueless type
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized container of one of those types has no
payload at all and its declared count is the only thing that decides how much
is allocated: `[$Z#L` followed by a large count turns some ten bytes of input
into that many values (see #2793, which reports 35 GB and 150 seconds). Every
other type costs at least one byte per element and is bounded by the end of
the input.
This is a sanity bound rather than a security boundary, and it is far above
any container met in practice. @ref binary_writer falls back to the
unoptimized encoding for longer containers, so that a value serialized by
this library can always be read back.
@sa https://github.com/nlohmann/json/issues/2793
*/
JSON_INLINE_VARIABLE constexpr std::size_t max_valueless_container_size = 1 << 20;
///////////////////
// binary reader //
///////////////////
@@ -110,6 +130,7 @@ class binary_reader
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
{
sax = sax_;
container_stack.clear();
bool result = false;
switch (format)
@@ -159,6 +180,57 @@ class binary_reader
}
private:
////////////////////////
// nested containers //
////////////////////////
/*!
@brief a container that has been opened and not closed yet
The binary readers do not call themselves once per nesting level. Like
@ref parser::sax_parse_internal, which does the same for JSON text, they
keep the containers they are inside of on a heap-allocated stack, so that
the native call stack does not grow with the nesting depth of the input
and a deeply nested value is bounded by memory rather than by the stack
(see #5104).
The members are ordered widest first: frames are stored in a vector, and
declaring the `bool` first would pad the struct out for no reason.
*/
struct container_frame
{
/// number of elements that have not been read yet
std::size_t remaining = 0;
/// whether to close this container with end_object() or end_array()
bool is_object = false;
};
/*!
@brief open a nested array or object
Emits the SAX start event and records the container. This is the only
place the binary readers start a container, so a check that rejects one
can be made here and is then guaranteed to run before the start event.
@param[in] is_object whether an object (true) or an array (false) begins
@param[in] len number of elements the container declares
@return whether the SAX parser accepted the start event
*/
bool enter_container(const bool is_object, const std::size_t len)
{
if (JSON_HEDLEY_UNLIKELY(is_object ? !sax->start_object(len) : !sax->start_array(len)))
{
return false;
}
container_frame frame;
frame.remaining = len;
frame.is_object = is_object;
container_stack.push_back(frame);
return true;
}
//////////
// BSON //
//////////
@@ -996,23 +1068,21 @@ class binary_reader
}
/*!
@brief reads a CBOR string
@brief reads a definite-length CBOR string
This function first reads starting bytes to determine the expected
string length and then copies this number of bytes into a string.
Additionally, CBOR's strings with indefinite lengths are supported.
Reads everything @ref get_cbor_string accepts except the indefinite-length
form, which that function handles itself. The bytes are appended to @a
result, so consecutive chunks of an indefinite-length string can be read
into the same string.
@param[out] result created string
@param[out] result string the bytes are appended to
@return whether string creation completed
*/
bool get_cbor_string(string_t& result)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "string")))
{
return false;
}
@pre @a current is not EOF
*/
bool get_cbor_string_chunk(string_t& result)
{
switch (current)
{
// UTF-8 string (0x00..0x17 bytes follow)
@@ -1068,20 +1138,6 @@ class binary_reader
return get_number(input_format_t::cbor, len) && get_string(input_format_t::cbor, len, result);
}
case 0x7F: // UTF-8 string (indefinite length)
{
while (get() != 0xFF)
{
string_t chunk;
if (!get_cbor_string(chunk))
{
return false;
}
result.append(chunk);
}
return true;
}
default:
{
auto last_token = get_token_string();
@@ -1092,23 +1148,82 @@ class binary_reader
}
/*!
@brief reads a CBOR byte array
@brief reads a CBOR string
This function first reads starting bytes to determine the expected
byte array length and then copies this number of bytes into the byte array.
Additionally, CBOR's byte arrays with indefinite lengths are supported.
string length and then copies this number of bytes into a string.
Additionally, CBOR's strings with indefinite lengths are supported.
@param[out] result created byte array
@param[out] result created string
@return whether string creation completed
*/
bool get_cbor_string(string_t& result)
{
// number of indefinite-length strings that have been opened and not
// closed yet. RFC 8949, Section 3.2.3 does not permit nesting them,
// but this reader has always accepted it, so the open levels are
// counted instead of recursed through, which overflowed the stack for
// an input of repeated 0x7F bytes (see #5104). Every chunk is appended
// to the same result, so no per-level state is needed.
std::size_t open = 0;
while (true)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "string")))
{
return false;
}
if (current == 0x7F) // UTF-8 string (indefinite length)
{
++open;
get();
continue;
}
// a break marker closes the innermost indefinite-length string;
// outside of one it is not a string and falls through to the error
if (open != 0 && current == 0xFF)
{
if (--open == 0)
{
return true;
}
get();
continue;
}
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string_chunk(result)))
{
return false;
}
if (open == 0)
{
return true;
}
get();
}
}
/*!
@brief reads a definite-length CBOR byte array
Reads everything @ref get_cbor_binary accepts except the indefinite-length
form, which that function handles itself. The bytes are appended to @a
result, so consecutive chunks of an indefinite-length byte array can be
read into the same byte array.
@param[out] result byte array the bytes are appended to
@return whether byte array creation completed
*/
bool get_cbor_binary(binary_t& result)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "binary")))
{
return false;
}
@pre @a current is not EOF
*/
bool get_cbor_binary_chunk(binary_t& result)
{
switch (current)
{
// Binary data (0x00..0x17 bytes follow)
@@ -1168,20 +1283,6 @@ class binary_reader
get_binary(input_format_t::cbor, len, result);
}
case 0x5F: // Binary data (indefinite length)
{
while (get() != 0xFF)
{
binary_t chunk;
if (!get_cbor_binary(chunk))
{
return false;
}
result.insert(result.end(), chunk.begin(), chunk.end());
}
return true;
}
default:
{
auto last_token = get_token_string();
@@ -1191,6 +1292,63 @@ class binary_reader
}
}
/*!
@brief reads a CBOR byte array
This function first reads starting bytes to determine the expected
byte array length and then copies this number of bytes into the byte array.
Additionally, CBOR's byte arrays with indefinite lengths are supported.
@param[out] result created byte array
@return whether byte array creation completed
*/
bool get_cbor_binary(binary_t& result)
{
// the open indefinite-length byte arrays are counted rather than
// recursed through, for the reason given in @ref get_cbor_string
std::size_t open = 0;
while (true)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "binary")))
{
return false;
}
if (current == 0x5F) // Binary data (indefinite length)
{
++open;
get();
continue;
}
// a break marker closes the innermost indefinite-length byte
// array; outside of one it falls through to the error below
if (open != 0 && current == 0xFF)
{
if (--open == 0)
{
return true;
}
get();
continue;
}
if (JSON_HEDLEY_UNLIKELY(!get_cbor_binary_chunk(result)))
{
return false;
}
if (open == 0)
{
return true;
}
get();
}
}
/*!
@brief narrow a definite CBOR array/map length to std::size_t
@@ -1316,7 +1474,17 @@ class binary_reader
/*!
@return whether a valid MessagePack value was passed to the SAX parser
*/
bool parse_msgpack_internal()
/*!
@brief read one MessagePack value
Reads a single value and passes it to the SAX parser. A value that begins
a container is not read to its end: the container is opened with
@ref enter_container and its elements are read by
@ref parse_msgpack_internal, so that nesting does not consume native stack.
@return whether reading the value succeeded
*/
bool parse_msgpack_value()
{
switch (get())
{
@@ -1472,7 +1640,7 @@ class binary_reader
case 0x8D:
case 0x8E:
case 0x8F:
return get_msgpack_object(conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
return enter_container(/*is_object*/true, conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
// fixarray
case 0x90:
@@ -1491,7 +1659,7 @@ class binary_reader
case 0x9D:
case 0x9E:
case 0x9F:
return get_msgpack_array(conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
return enter_container(/*is_object*/false, conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
// fixstr
case 0xA0:
@@ -1622,25 +1790,25 @@ class binary_reader
case 0xDC: // array 16
{
std::uint16_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_array(static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/false, static_cast<std::size_t>(len));
}
case 0xDD: // array 32
{
std::uint32_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_array(conditional_static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/false, conditional_static_cast<std::size_t>(len));
}
case 0xDE: // map 16
{
std::uint16_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_object(static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/true, static_cast<std::size_t>(len));
}
case 0xDF: // map 32
{
std::uint32_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_object(conditional_static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/true, conditional_static_cast<std::size_t>(len));
}
// negative fixint
@@ -1888,55 +2056,69 @@ class binary_reader
}
/*!
@param[in] len the length of the array
@return whether array creation completed
@brief read a MessagePack value and everything nested inside it
Reads values until the one that was begun here is complete, resuming the
enclosing container each time an element ends, so that the nesting depth
of the input costs heap rather than native stack (see #5104).
@return whether reading the value succeeded
*/
bool get_msgpack_array(const std::size_t len)
bool parse_msgpack_internal()
{
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(len)))
{
return false;
}
for (std::size_t i = 0; i < len; ++i)
{
if (JSON_HEDLEY_UNLIKELY(!parse_msgpack_internal()))
{
return false;
}
}
return sax->end_array();
}
/*!
@param[in] len the length of the object
@return whether object creation completed
*/
bool get_msgpack_object(const std::size_t len)
{
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(len)))
{
return false;
}
// the key currently being read; hoisted out of the loop so that its
// capacity is reused across elements and across nesting levels
string_t key;
for (std::size_t i = 0; i < len; ++i)
while (true)
{
get();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
if (!container_stack.empty())
{
// copied out before anything can push onto the stack and
// invalidate a reference into it
const bool is_object = container_stack.back().is_object;
if (container_stack.back().remaining == 0)
{
container_stack.pop_back();
if (JSON_HEDLEY_UNLIKELY(is_object ? !sax->end_object() : !sax->end_array()))
{
return false;
}
// the value begun here is complete once its container is
if (container_stack.empty())
{
return true;
}
continue;
}
// claim the element about to be read
--container_stack.back().remaining;
if (is_object)
{
get();
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
{
return false;
}
}
}
if (JSON_HEDLEY_UNLIKELY(!parse_msgpack_value()))
{
return false;
}
if (JSON_HEDLEY_UNLIKELY(!parse_msgpack_internal()))
// a value that opened a container left it on the stack; one that
// did not, and that was not inside a container, was the whole value
if (container_stack.empty())
{
return false;
return true;
}
key.clear();
}
return sax->end_object();
}
////////////
@@ -2391,7 +2573,12 @@ class binary_reader
{
result.first = npos; // size
result.second = 0; // type
bool is_ndarray = false;
// seed the flag with the caller's context: inside an ndarray dimension
// vector another ndarray is not allowed, and get_ubjson_size_value()
// rejects it up front instead of reading it and reporting afterwards.
// Seeding it with `false` made every '#' of a "[#[#[..." chain descend
// another level, which overflowed the stack (see #5104).
bool is_ndarray = inside_ndarray;
get_ignore_noop();
@@ -2424,13 +2611,11 @@ class binary_reader
}
const bool is_error = get_ubjson_size_value(result.first, is_ndarray);
if (input_format == input_format_t::bjdata && is_ndarray)
// an ndarray was read here only if the flag flipped; when it was
// seeded true, get_ubjson_size_value() already rejected the nested
// dimension vector
if (input_format == input_format_t::bjdata && is_ndarray && !inside_ndarray)
{
if (inside_ndarray)
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format, "ndarray can not be recursive", "size"), nullptr));
}
result.second |= (1 << 8); // use bit 8 to indicate ndarray, all UBJSON and BJData markers should be ASCII letters
}
return is_error;
@@ -2439,7 +2624,7 @@ class binary_reader
if (current == '#')
{
const bool is_error = get_ubjson_size_value(result.first, is_ndarray);
if (input_format == input_format_t::bjdata && is_ndarray)
if (input_format == input_format_t::bjdata && is_ndarray && !inside_ndarray)
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format, "ndarray requires both type and size", "size"), nullptr));
@@ -2710,6 +2895,17 @@ class binary_reader
if (size_and_type.first != npos)
{
// reading an element of a valueless type consumes no input, so the
// declared count alone decides how much is allocated; the check is
// made before the start event so that no container is opened that
// is then abandoned. See @ref max_valueless_container_size.
if (JSON_HEDLEY_UNLIKELY((size_and_type.second == 'Z' || size_and_type.second == 'T' || size_and_type.second == 'F')
&& size_and_type.first > max_valueless_container_size))
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408,
exception_message(input_format, "excessive array size", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(size_and_type.first)))
{
return false;
@@ -3227,6 +3423,9 @@ class binary_reader
/// the SAX parser
json_sax_t* sax = nullptr;
/// the containers that have been opened and not closed yet; see @ref container_frame
std::vector<container_frame> container_stack{};
// excluded markers in bjdata optimized type
#define JSON_BINARY_READER_MAKE_BJD_OPTIMIZED_TYPE_MARKERS_ \
make_array<char_int_type>('F', 'H', 'N', 'S', 'T', 'Z', '[', '{')
@@ -826,7 +826,17 @@ class binary_writer
std::vector<CharType> bjdx = {'[', '{', 'S', 'H', 'T', 'F', 'N', 'Z'}; // excluded markers in bjdata optimized type
if (same_prefix && !(use_bjdata && std::find(bjdx.begin(), bjdx.end(), first_prefix) != bjdx.end()))
// an optimized array of a valueless type carries no payload, so a
// reader has nothing but the declared count to bound the allocation
// by and refuses an excessive one. Write the unoptimized form for
// those, at one byte per element, so the result can be read back.
// Objects are not affected: every element is preceded by its key.
const bool valueless_type = (first_prefix == 'Z' || first_prefix == 'T' || first_prefix == 'F');
const bool excessive_valueless = valueless_type
&& j.m_data.m_value.array->size() > detail::max_valueless_container_size;
if (same_prefix && !excessive_valueless
&& !(use_bjdata && std::find(bjdx.begin(), bjdx.end(), first_prefix) != bjdx.end()))
{
prefix_required = false;
oa->write_character(to_char_type('$'));
+14 -14
View File
@@ -4474,7 +4474,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in CBOR format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4491,7 +4491,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -4517,7 +4517,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in MessagePack format
@@ -4532,7 +4532,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in MessagePack format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4548,7 +4548,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -4572,7 +4572,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in UBJSON format
@@ -4587,7 +4587,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in UBJSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4603,7 +4603,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -4627,7 +4627,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BJData format
@@ -4642,7 +4642,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BJData format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4658,7 +4658,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BSON format
@@ -4673,7 +4673,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -4689,7 +4689,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -4713,7 +4713,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @}
+328 -119
View File
@@ -10745,6 +10745,26 @@ inline bool little_endianness(int num = 1) noexcept
return *reinterpret_cast<char*>(&num) == 1;
}
/*!
@brief largest element count accepted for a UBJSON container of a valueless type
An element of type 'Z' (null), 'T' (true) or 'F' (false) is encoded by its
type marker alone, so an optimized container of one of those types has no
payload at all and its declared count is the only thing that decides how much
is allocated: `[$Z#L` followed by a large count turns some ten bytes of input
into that many values (see #2793, which reports 35 GB and 150 seconds). Every
other type costs at least one byte per element and is bounded by the end of
the input.
This is a sanity bound rather than a security boundary, and it is far above
any container met in practice. @ref binary_writer falls back to the
unoptimized encoding for longer containers, so that a value serialized by
this library can always be read back.
@sa https://github.com/nlohmann/json/issues/2793
*/
JSON_INLINE_VARIABLE constexpr std::size_t max_valueless_container_size = 1 << 20;
///////////////////
// binary reader //
///////////////////
@@ -10797,6 +10817,7 @@ class binary_reader
const cbor_tag_handler_t tag_handler = cbor_tag_handler_t::error)
{
sax = sax_;
container_stack.clear();
bool result = false;
switch (format)
@@ -10846,6 +10867,57 @@ class binary_reader
}
private:
////////////////////////
// nested containers //
////////////////////////
/*!
@brief a container that has been opened and not closed yet
The binary readers do not call themselves once per nesting level. Like
@ref parser::sax_parse_internal, which does the same for JSON text, they
keep the containers they are inside of on a heap-allocated stack, so that
the native call stack does not grow with the nesting depth of the input
and a deeply nested value is bounded by memory rather than by the stack
(see #5104).
The members are ordered widest first: frames are stored in a vector, and
declaring the `bool` first would pad the struct out for no reason.
*/
struct container_frame
{
/// number of elements that have not been read yet
std::size_t remaining = 0;
/// whether to close this container with end_object() or end_array()
bool is_object = false;
};
/*!
@brief open a nested array or object
Emits the SAX start event and records the container. This is the only
place the binary readers start a container, so a check that rejects one
can be made here and is then guaranteed to run before the start event.
@param[in] is_object whether an object (true) or an array (false) begins
@param[in] len number of elements the container declares
@return whether the SAX parser accepted the start event
*/
bool enter_container(const bool is_object, const std::size_t len)
{
if (JSON_HEDLEY_UNLIKELY(is_object ? !sax->start_object(len) : !sax->start_array(len)))
{
return false;
}
container_frame frame;
frame.remaining = len;
frame.is_object = is_object;
container_stack.push_back(frame);
return true;
}
//////////
// BSON //
//////////
@@ -11683,23 +11755,21 @@ class binary_reader
}
/*!
@brief reads a CBOR string
@brief reads a definite-length CBOR string
This function first reads starting bytes to determine the expected
string length and then copies this number of bytes into a string.
Additionally, CBOR's strings with indefinite lengths are supported.
Reads everything @ref get_cbor_string accepts except the indefinite-length
form, which that function handles itself. The bytes are appended to @a
result, so consecutive chunks of an indefinite-length string can be read
into the same string.
@param[out] result created string
@param[out] result string the bytes are appended to
@return whether string creation completed
*/
bool get_cbor_string(string_t& result)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "string")))
{
return false;
}
@pre @a current is not EOF
*/
bool get_cbor_string_chunk(string_t& result)
{
switch (current)
{
// UTF-8 string (0x00..0x17 bytes follow)
@@ -11755,20 +11825,6 @@ class binary_reader
return get_number(input_format_t::cbor, len) && get_string(input_format_t::cbor, len, result);
}
case 0x7F: // UTF-8 string (indefinite length)
{
while (get() != 0xFF)
{
string_t chunk;
if (!get_cbor_string(chunk))
{
return false;
}
result.append(chunk);
}
return true;
}
default:
{
auto last_token = get_token_string();
@@ -11779,23 +11835,82 @@ class binary_reader
}
/*!
@brief reads a CBOR byte array
@brief reads a CBOR string
This function first reads starting bytes to determine the expected
byte array length and then copies this number of bytes into the byte array.
Additionally, CBOR's byte arrays with indefinite lengths are supported.
string length and then copies this number of bytes into a string.
Additionally, CBOR's strings with indefinite lengths are supported.
@param[out] result created byte array
@param[out] result created string
@return whether string creation completed
*/
bool get_cbor_string(string_t& result)
{
// number of indefinite-length strings that have been opened and not
// closed yet. RFC 8949, Section 3.2.3 does not permit nesting them,
// but this reader has always accepted it, so the open levels are
// counted instead of recursed through, which overflowed the stack for
// an input of repeated 0x7F bytes (see #5104). Every chunk is appended
// to the same result, so no per-level state is needed.
std::size_t open = 0;
while (true)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "string")))
{
return false;
}
if (current == 0x7F) // UTF-8 string (indefinite length)
{
++open;
get();
continue;
}
// a break marker closes the innermost indefinite-length string;
// outside of one it is not a string and falls through to the error
if (open != 0 && current == 0xFF)
{
if (--open == 0)
{
return true;
}
get();
continue;
}
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string_chunk(result)))
{
return false;
}
if (open == 0)
{
return true;
}
get();
}
}
/*!
@brief reads a definite-length CBOR byte array
Reads everything @ref get_cbor_binary accepts except the indefinite-length
form, which that function handles itself. The bytes are appended to @a
result, so consecutive chunks of an indefinite-length byte array can be
read into the same byte array.
@param[out] result byte array the bytes are appended to
@return whether byte array creation completed
*/
bool get_cbor_binary(binary_t& result)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "binary")))
{
return false;
}
@pre @a current is not EOF
*/
bool get_cbor_binary_chunk(binary_t& result)
{
switch (current)
{
// Binary data (0x00..0x17 bytes follow)
@@ -11855,20 +11970,6 @@ class binary_reader
get_binary(input_format_t::cbor, len, result);
}
case 0x5F: // Binary data (indefinite length)
{
while (get() != 0xFF)
{
binary_t chunk;
if (!get_cbor_binary(chunk))
{
return false;
}
result.insert(result.end(), chunk.begin(), chunk.end());
}
return true;
}
default:
{
auto last_token = get_token_string();
@@ -11878,6 +11979,63 @@ class binary_reader
}
}
/*!
@brief reads a CBOR byte array
This function first reads starting bytes to determine the expected
byte array length and then copies this number of bytes into the byte array.
Additionally, CBOR's byte arrays with indefinite lengths are supported.
@param[out] result created byte array
@return whether byte array creation completed
*/
bool get_cbor_binary(binary_t& result)
{
// the open indefinite-length byte arrays are counted rather than
// recursed through, for the reason given in @ref get_cbor_string
std::size_t open = 0;
while (true)
{
if (JSON_HEDLEY_UNLIKELY(!unexpect_eof(input_format_t::cbor, "binary")))
{
return false;
}
if (current == 0x5F) // Binary data (indefinite length)
{
++open;
get();
continue;
}
// a break marker closes the innermost indefinite-length byte
// array; outside of one it falls through to the error below
if (open != 0 && current == 0xFF)
{
if (--open == 0)
{
return true;
}
get();
continue;
}
if (JSON_HEDLEY_UNLIKELY(!get_cbor_binary_chunk(result)))
{
return false;
}
if (open == 0)
{
return true;
}
get();
}
}
/*!
@brief narrow a definite CBOR array/map length to std::size_t
@@ -12003,7 +12161,17 @@ class binary_reader
/*!
@return whether a valid MessagePack value was passed to the SAX parser
*/
bool parse_msgpack_internal()
/*!
@brief read one MessagePack value
Reads a single value and passes it to the SAX parser. A value that begins
a container is not read to its end: the container is opened with
@ref enter_container and its elements are read by
@ref parse_msgpack_internal, so that nesting does not consume native stack.
@return whether reading the value succeeded
*/
bool parse_msgpack_value()
{
switch (get())
{
@@ -12159,7 +12327,7 @@ class binary_reader
case 0x8D:
case 0x8E:
case 0x8F:
return get_msgpack_object(conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
return enter_container(/*is_object*/true, conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
// fixarray
case 0x90:
@@ -12178,7 +12346,7 @@ class binary_reader
case 0x9D:
case 0x9E:
case 0x9F:
return get_msgpack_array(conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
return enter_container(/*is_object*/false, conditional_static_cast<std::size_t>(static_cast<unsigned int>(current) & 0x0Fu));
// fixstr
case 0xA0:
@@ -12309,25 +12477,25 @@ class binary_reader
case 0xDC: // array 16
{
std::uint16_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_array(static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/false, static_cast<std::size_t>(len));
}
case 0xDD: // array 32
{
std::uint32_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_array(conditional_static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/false, conditional_static_cast<std::size_t>(len));
}
case 0xDE: // map 16
{
std::uint16_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_object(static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/true, static_cast<std::size_t>(len));
}
case 0xDF: // map 32
{
std::uint32_t len{};
return get_number(input_format_t::msgpack, len) && get_msgpack_object(conditional_static_cast<std::size_t>(len));
return get_number(input_format_t::msgpack, len) && enter_container(/*is_object*/true, conditional_static_cast<std::size_t>(len));
}
// negative fixint
@@ -12575,55 +12743,69 @@ class binary_reader
}
/*!
@param[in] len the length of the array
@return whether array creation completed
@brief read a MessagePack value and everything nested inside it
Reads values until the one that was begun here is complete, resuming the
enclosing container each time an element ends, so that the nesting depth
of the input costs heap rather than native stack (see #5104).
@return whether reading the value succeeded
*/
bool get_msgpack_array(const std::size_t len)
bool parse_msgpack_internal()
{
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(len)))
{
return false;
}
for (std::size_t i = 0; i < len; ++i)
{
if (JSON_HEDLEY_UNLIKELY(!parse_msgpack_internal()))
{
return false;
}
}
return sax->end_array();
}
/*!
@param[in] len the length of the object
@return whether object creation completed
*/
bool get_msgpack_object(const std::size_t len)
{
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(len)))
{
return false;
}
// the key currently being read; hoisted out of the loop so that its
// capacity is reused across elements and across nesting levels
string_t key;
for (std::size_t i = 0; i < len; ++i)
while (true)
{
get();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
if (!container_stack.empty())
{
// copied out before anything can push onto the stack and
// invalidate a reference into it
const bool is_object = container_stack.back().is_object;
if (container_stack.back().remaining == 0)
{
container_stack.pop_back();
if (JSON_HEDLEY_UNLIKELY(is_object ? !sax->end_object() : !sax->end_array()))
{
return false;
}
// the value begun here is complete once its container is
if (container_stack.empty())
{
return true;
}
continue;
}
// claim the element about to be read
--container_stack.back().remaining;
if (is_object)
{
get();
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
{
return false;
}
}
}
if (JSON_HEDLEY_UNLIKELY(!parse_msgpack_value()))
{
return false;
}
if (JSON_HEDLEY_UNLIKELY(!parse_msgpack_internal()))
// a value that opened a container left it on the stack; one that
// did not, and that was not inside a container, was the whole value
if (container_stack.empty())
{
return false;
return true;
}
key.clear();
}
return sax->end_object();
}
////////////
@@ -13078,7 +13260,12 @@ class binary_reader
{
result.first = npos; // size
result.second = 0; // type
bool is_ndarray = false;
// seed the flag with the caller's context: inside an ndarray dimension
// vector another ndarray is not allowed, and get_ubjson_size_value()
// rejects it up front instead of reading it and reporting afterwards.
// Seeding it with `false` made every '#' of a "[#[#[..." chain descend
// another level, which overflowed the stack (see #5104).
bool is_ndarray = inside_ndarray;
get_ignore_noop();
@@ -13111,13 +13298,11 @@ class binary_reader
}
const bool is_error = get_ubjson_size_value(result.first, is_ndarray);
if (input_format == input_format_t::bjdata && is_ndarray)
// an ndarray was read here only if the flag flipped; when it was
// seeded true, get_ubjson_size_value() already rejected the nested
// dimension vector
if (input_format == input_format_t::bjdata && is_ndarray && !inside_ndarray)
{
if (inside_ndarray)
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format, "ndarray can not be recursive", "size"), nullptr));
}
result.second |= (1 << 8); // use bit 8 to indicate ndarray, all UBJSON and BJData markers should be ASCII letters
}
return is_error;
@@ -13126,7 +13311,7 @@ class binary_reader
if (current == '#')
{
const bool is_error = get_ubjson_size_value(result.first, is_ndarray);
if (input_format == input_format_t::bjdata && is_ndarray)
if (input_format == input_format_t::bjdata && is_ndarray && !inside_ndarray)
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(112, chars_read,
exception_message(input_format, "ndarray requires both type and size", "size"), nullptr));
@@ -13397,6 +13582,17 @@ class binary_reader
if (size_and_type.first != npos)
{
// reading an element of a valueless type consumes no input, so the
// declared count alone decides how much is allocated; the check is
// made before the start event so that no container is opened that
// is then abandoned. See @ref max_valueless_container_size.
if (JSON_HEDLEY_UNLIKELY((size_and_type.second == 'Z' || size_and_type.second == 'T' || size_and_type.second == 'F')
&& size_and_type.first > max_valueless_container_size))
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408,
exception_message(input_format, "excessive array size", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(size_and_type.first)))
{
return false;
@@ -13914,6 +14110,9 @@ class binary_reader
/// the SAX parser
json_sax_t* sax = nullptr;
/// the containers that have been opened and not closed yet; see @ref container_frame
std::vector<container_frame> container_stack{};
// excluded markers in bjdata optimized type
#define JSON_BINARY_READER_MAKE_BJD_OPTIMIZED_TYPE_MARKERS_ \
make_array<char_int_type>('F', 'H', 'N', 'S', 'T', 'Z', '[', '{')
@@ -17834,7 +18033,17 @@ class binary_writer
std::vector<CharType> bjdx = {'[', '{', 'S', 'H', 'T', 'F', 'N', 'Z'}; // excluded markers in bjdata optimized type
if (same_prefix && !(use_bjdata && std::find(bjdx.begin(), bjdx.end(), first_prefix) != bjdx.end()))
// an optimized array of a valueless type carries no payload, so a
// reader has nothing but the declared count to bound the allocation
// by and refuses an excessive one. Write the unoptimized form for
// those, at one byte per element, so the result can be read back.
// Objects are not affected: every element is preceded by its key.
const bool valueless_type = (first_prefix == 'Z' || first_prefix == 'T' || first_prefix == 'F');
const bool excessive_valueless = valueless_type
&& j.m_data.m_value.array->size() > detail::max_valueless_container_size;
if (same_prefix && !excessive_valueless
&& !(use_bjdata && std::find(bjdx.begin(), bjdx.end(), first_prefix) != bjdx.end()))
{
prefix_required = false;
oa->write_character(to_char_type('$'));
@@ -25902,7 +26111,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in CBOR format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -25919,7 +26128,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -25945,7 +26154,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in MessagePack format
@@ -25960,7 +26169,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in MessagePack format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -25976,7 +26185,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -26000,7 +26209,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in UBJSON format
@@ -26015,7 +26224,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in UBJSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -26031,7 +26240,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -26055,7 +26264,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BJData format
@@ -26070,7 +26279,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BJData format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -26086,7 +26295,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BSON format
@@ -26101,7 +26310,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @brief create a JSON value from an input in BSON format (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -26117,7 +26326,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
template<typename T>
@@ -26141,7 +26350,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
const bool res = binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict); // cppcheck-suppress[accessMoved]
return res ? result : basic_json(value_t::discarded);
return res ? std::move(result) : basic_json(value_t::discarded);
}
/// @}
+18 -3
View File
@@ -3288,8 +3288,10 @@ TEST_CASE("BJData")
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR1), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR1, true, false).is_discarded());
// a dimension vector that opens another one is rejected where the
// nested '[' is read, rather than after it has been descended into
std::vector<uint8_t> const vR2 = {'[', '$', 'i', '#', '[', '#', '[', 'i', 1, ']', ']', 1};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR2), "[json.exception.parse_error.113] parse error at byte 11: syntax error while parsing BJData size: expected length type specification (U, i, u, I, m, l, M, L) after '#'; last byte: 0x5D", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR2), "[json.exception.parse_error.113] parse error at byte 7: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR2, true, false).is_discarded());
std::vector<uint8_t> const vR3 = {'[', '#', '[', 'i', '2', 'i', 2, ']'};
@@ -3297,7 +3299,7 @@ TEST_CASE("BJData")
CHECK(json::from_bjdata(vR3, true, false).is_discarded());
std::vector<uint8_t> const vR4 = {'[', '$', 'i', '#', '[', '$', 'i', '#', '[', 'i', 1, ']', 1};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR4), "[json.exception.parse_error.110] parse error at byte 14: syntax error while parsing BJData number: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR4), "[json.exception.parse_error.113] parse error at byte 9: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR4, true, false).is_discarded());
std::vector<uint8_t> const vR5 = {'[', '$', 'i', '#', '[', '[', '[', ']', ']', ']'};
@@ -3305,12 +3307,25 @@ TEST_CASE("BJData")
CHECK(json::from_bjdata(vR5, true, false).is_discarded());
std::vector<uint8_t> const vR6 = {'[', '$', 'i', '#', '[', '$', 'i', '#', '[', 'i', '2', 'i', 2, ']'};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR6), "[json.exception.parse_error.112] parse error at byte 14: syntax error while parsing BJData size: ndarray can not be recursive", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vR6), "[json.exception.parse_error.113] parse error at byte 9: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vR6, true, false).is_discarded());
std::vector<uint8_t> const vH = {'[', 'H', '[', '#', '[', '$', 'i', '#', '[', 'i', '2', 'i', 2, ']'};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vH), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vH, true, false).is_discarded());
// Every "#[" of this chain used to open another dimension vector
// and cost several stack frames before anything was rejected, so a
// long enough chain crashed the process (see #5104). The nested
// vector is refused where it is read, so the length is irrelevant.
std::vector<uint8_t> vRdeep = {'['};
for (std::size_t i = 0; i < 100000; ++i)
{
vRdeep.push_back('#');
vRdeep.push_back('[');
}
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vRdeep), "[json.exception.parse_error.113] parse error at byte 5: syntax error while parsing BJData size: ndarray dimensional vector is not allowed", json::parse_error&);
CHECK(json::from_bjdata(vRdeep, true, false).is_discarded());
}
SECTION("objects")
+52
View File
@@ -2035,6 +2035,58 @@ TEST_CASE("CBOR definite length equal to the indefinite-length sentinel")
}
}
TEST_CASE("CBOR indefinite-length strings do not recurse per chunk")
{
// Reading an indefinite-length string or byte array used to call itself
// once per chunk, so a payload of repeated 0x7F (or 0x5F) bytes exhausted
// the call stack before any of the input was rejected. The open levels are
// counted now, and the levels below prove the reader still reads the same
// values and reports the same errors at the same byte offsets.
json _;
SECTION("many open levels are reported, not crashed on")
{
const std::vector<uint8_t> input(200000, 0x7F);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), "[json.exception.parse_error.110] parse error at byte 200001: syntax error while parsing CBOR string: unexpected end of input", json::parse_error&);
CHECK(json::from_cbor(input, true, false).is_discarded());
}
SECTION("many open levels are reported, not crashed on (binary)")
{
const std::vector<uint8_t> input(200000, 0x5F);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input), "[json.exception.parse_error.110] parse error at byte 200001: syntax error while parsing CBOR binary: unexpected end of input", json::parse_error&);
CHECK(json::from_cbor(input, true, false).is_discarded());
}
SECTION("chunks are still concatenated")
{
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0xFF})) == json(""));
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0x61, 0x61, 0xFF})) == json("a"));
// nested indefinite-length strings are concatenated across levels
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0x7F, 0x61, 0x61, 0xFF, 0x61, 0x62, 0xFF})) == json("ab"));
CHECK(json::from_cbor(std::vector<uint8_t>({0x7F, 0x7F, 0x7F, 0x61, 0x7A, 0xFF, 0xFF, 0xFF})) == json("z"));
CHECK(json::from_cbor(std::vector<uint8_t>({0xA1, 0x7F, 0x61, 0x61, 0xFF, 0x01})) == json({{"a", 1}}));
}
SECTION("chunks are still concatenated (binary)")
{
CHECK(json::from_cbor(std::vector<uint8_t>({0x5F, 0x41, 0x61, 0xFF})) == json::binary({0x61}));
CHECK(json::from_cbor(std::vector<uint8_t>({0x5F, 0x5F, 0x41, 0x61, 0xFF, 0x41, 0x62, 0xFF})) == json::binary({0x61, 0x62}));
}
SECTION("a chunk that is not a string is still rejected")
{
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7F, 0x7F, 0x00})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x00", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x5F, 0x5F, 0x00})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR binary: expected length specification (0x40-0x5B) or indefinite binary array type (0x5F); last byte: 0x00", json::parse_error&);
}
SECTION("a break marker outside an indefinite-length string is not a string")
{
// 0xFF only closes a string that was opened; on its own it is not one
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&);
}
}
TEST_CASE("CBOR roundtrips" * doctest::skip())
{
SECTION("input from flynn")
+61
View File
@@ -1598,6 +1598,67 @@ TEST_CASE("MessagePack")
}
// use this testcase outside [hide] to run it with Valgrind
TEST_CASE("MessagePack nesting does not consume the call stack")
{
// Reading a container used to call back into the value reader once per
// element, so the native call stack grew with the nesting depth of the
// input: one frame per byte for repeated 0x91 (a one-element array), which
// crashes the process long before the input is exhausted (#5104). The
// containers are kept on a heap stack now.
//
// Note that deeply nested values must not be compared, copied or dumped
// here: those operations are still recursive, and would reintroduce the
// very crash this checks for. Depth is measured by descending instead.
SECTION("an unterminated chain is reported, not crashed on")
{
json _;
const std::vector<uint8_t> input(300000, 0x91);
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(input), "[json.exception.parse_error.110] parse error at byte 300001: syntax error while parsing MessagePack value: unexpected end of input", json::parse_error&);
CHECK(json::from_msgpack(input, true, false).is_discarded());
}
SECTION("a well-formed deep value is read through the SAX interface")
{
std::vector<uint8_t> input(300000, 0x91);
input.push_back(0x01); // innermost value
SaxCountdown accept_all(600001);
CHECK(json::sax_parse(input, &accept_all, json::input_format_t::msgpack));
}
SECTION("a well-formed deep value is read into a value")
{
const std::size_t depth = 10000;
std::vector<uint8_t> input(depth, 0x91);
input.push_back(0x01);
json j = json::from_msgpack(input);
std::size_t measured = 0;
const json* p = &j;
while (p->is_array() && !p->empty())
{
p = &p->front();
++measured;
}
CHECK(measured == depth);
CHECK(p->is_number());
}
SECTION("containers are still read the same way")
{
CHECK(json::from_msgpack(std::vector<uint8_t>({0x90})) == json::array());
CHECK(json::from_msgpack(std::vector<uint8_t>({0x80})) == json::object());
CHECK(json::from_msgpack(std::vector<uint8_t>({0x92, 0x90, 0x80})) == json({json::array(), json::object()}));
CHECK(json::from_msgpack(std::vector<uint8_t>({0x91, 0x91, 0x91, 0x90})) == json({{{json::array()}}}));
CHECK(json::from_msgpack(std::vector<uint8_t>({0x81, 0xA1, 'a', 0x81, 0xA1, 'b', 0x92, 0x01, 0x02})) == json({{"a", {{"b", {1, 2}}}}}));
// array 16 and map 32, i.e. the counted forms
CHECK(json::from_msgpack(std::vector<uint8_t>({0xDC, 0x00, 0x02, 0x01, 0x02})) == json({1, 2}));
CHECK(json::from_msgpack(std::vector<uint8_t>({0xDF, 0x00, 0x00, 0x00, 0x01, 0xA1, 'k', 0xC3})) == json({{"k", true}}));
}
}
TEST_CASE("single MessagePack roundtrip")
{
SECTION("sample.json")
+55
View File
@@ -2149,6 +2149,61 @@ TEST_CASE("UBJSON")
}
}
TEST_CASE("UBJSON optimized arrays of a valueless type are bounded")
{
// An element of type 'Z', 'T' or 'F' is encoded by its marker alone, so an
// optimized array of one of those has no payload and the declared count is
// the only thing deciding how much is allocated. Ten bytes used to produce
// billions of values (#2793); every other type costs at least one byte per
// element and is bounded by the end of the input.
json _;
SECTION("an excessive count is rejected")
{
// 'l' is a big-endian int32: 0x7FFFFFFF elements, about 34 GB of value
for (const auto marker :
{'Z', 'T', 'F'
})
{
const std::vector<uint8_t> input = {'[', '$', static_cast<uint8_t>(marker), '#', 'l', 0x7F, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(input), "[json.exception.out_of_range.408] syntax error while parsing UBJSON size: excessive array size", json::out_of_range&);
CHECK(json::from_ubjson(input, true, false).is_discarded());
}
}
SECTION("ordinary counts are unaffected")
{
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'Z', '#', 'i', 3})) == json({nullptr, nullptr, nullptr}));
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'T', '#', 'i', 2})) == json({true, true}));
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'F', '#', 'i', 2})) == json({false, false}));
// 'N' is a no-op rather than a value, and still yields an empty array
CHECK(json::from_ubjson(std::vector<uint8_t>({'[', '$', 'N', '#', 'i', 2})) == json::array());
}
SECTION("a type with a payload is unaffected")
{
// the same count for 'U' is bounded by the end of the input instead
const std::vector<uint8_t> input = {'[', '$', 'U', '#', 'l', 0x7F, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_AS(_ = json::from_ubjson(input), json::parse_error&);
}
SECTION("the writer stays within what the reader accepts")
{
// below the limit the optimized form is used and is tiny; above it the
// writer falls back so that the result can still be read back
json const at_limit(1048576, nullptr);
const auto v_at_limit = json::to_ubjson(at_limit, true, true);
CHECK(v_at_limit.size() == 9);
CHECK(v_at_limit.at(1) == '$');
CHECK(json::from_ubjson(v_at_limit) == at_limit);
json const above_limit(1048577, nullptr);
const auto v_above_limit = json::to_ubjson(above_limit, true, true);
CHECK(v_above_limit.at(1) != '$');
CHECK(json::from_ubjson(v_above_limit) == above_limit);
}
}
TEST_CASE("Universal Binary JSON Specification Examples 1")
{
SECTION("Null Value")