Repair complete items in binary formats when parse_error() returns true (#3989)

When the SAX parser asks to recover, the binary readers now repair an
item whose end is known and read on after it, as RFC 8949, Section 5.3
describes for CBOR:

- CBOR: tags are ignored, and simple values other than false, true, and
  null become null (RFC 8949, Section 6.1); a negative integer below the
  range of number_integer_t becomes the nearest floating-point number.
- Strings that are not valid UTF-8 get U+FFFD for each ill-formed
  sequence, as in JSON text; so does a UBJSON/BJData char above 0x7F.
- UBJSON/BJData high-precision numbers keep their longest valid
  beginning (via the lexer's recover_token()), or become infinity.
- Members whose key is not a string are skipped (CBOR, MessagePack,
  BON8), like members without a key in JSON text.
- BSON elements of types the library does not read (ObjectId, datetime,
  decimal128, ...) become null; a string without its terminator and a
  document whose size does not match are kept.

Where the end of an item is unknown, reading stops as before, except
that BSON skips to the end of the document, whose size it knows.

The value read before such an error is now completed by the reader from
its container stack, as the JSON parser does, instead of by a proxy SAX
parser, which is removed. Like the parser, binary_reader gets an
AllowRecovery template parameter, so that from_*() compile without the
new code.

Tests: a table of repairs, numbers out of range, errors that stop, and
a sweep over changed and removed bytes of eight encodings that checks
balanced events and that the first error is the one from_*() reports.
All fuzzers now run a recovering checker; the binary ones also check
that it reports an error exactly when from_*() fails.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-27 21:40:59 +02:00
parent c8735246d0
commit 437a95cfdb
16 changed files with 2925 additions and 673 deletions
+3 -3
View File
@@ -24,9 +24,9 @@ A parse error occurred.
Whether to recover from the error:
- `#!cpp false` stops parsing.
- `#!cpp true` recovers from the error: JSON text is repaired and parsing continues; for the binary formats, the value
read so far is completed and parsing stops. See [error recovery](../../features/parsing/error_recovery.md) for how
errors are repaired.
- `#!cpp true` recovers from the error: the error is repaired and parsing continues. If that is not possible, which
happens in the binary formats when the end of the item with the error is unknown, the value read so far is completed
and parsing stops. See [error recovery](../../features/parsing/error_recovery.md) for how errors are repaired.
Either way, [`sax_parse`](../basic_json/sax_parse.md) returns `#!cpp false`.
@@ -65,11 +65,36 @@ The input after the top-level value is not repaired: as without recovery, it is
The binary formats ([BJData](../binary_formats/bjdata.md), [BON8](../binary_formats/bon8.md),
[BSON](../binary_formats/bson.md), [CBOR](../binary_formats/cbor.md), [MessagePack](../binary_formats/messagepack.md),
and [UBJSON](../binary_formats/ubjson.md)) cannot be repaired: a value's size is stored before its content, and every
byte is a valid type marker, so after an error there is no way to tell where the next value begins. Parsing therefore
always stops at the first error. If `parse_error` returns `#!cpp true`, the value read so far is completed before
parsing stops: a key that waits for its value gets `#!json null`, and all open arrays and objects are closed. This keeps
everything before the error of an input that was cut off.
and [UBJSON](../binary_formats/ubjson.md)) have no delimiters to find the next value by. So what can be repaired depends
on whether the end of the item with the error is known, a distinction that
[RFC 8949, Section 5.3](https://www.rfc-editor.org/rfc/rfc8949.html#section-5.3) makes for CBOR, too.
If the item is complete, but cannot be passed on as it is, it is replaced, and parsing continues after it:
| Mistake | Formats | Repair |
|---------------------------------------------------------------------|-----------------------------------------|-------------------------------------------------------------------------|
| tag | CBOR | ignored |
| simple value other than `false`, `true`, and `null`, like undefined | CBOR | `#!json null` |
| negative integer below the range of `number_integer_t` | CBOR | the nearest floating-point number |
| string that is not valid UTF-8 | BJData, BSON, CBOR, MessagePack, UBJSON | each ill-formed sequence becomes U+FFFD |
| character (`C`) that is not ASCII | BJData, UBJSON | U+FFFD |
| invalid high-precision number (`H`) | BJData, UBJSON | the longest valid beginning is kept, as for JSON text, or `#!json null` |
| high-precision number too large | BJData, UBJSON | passed as infinity, together with its text |
| object key that is not a string | BON8, CBOR, MessagePack | the member is skipped |
| element of a type the library does not read, like ObjectId or date | BSON | `#!json null` |
| string without its terminator | BSON | kept |
| document whose size does not match its content | BSON | kept |
CBOR tags and simple values are repaired as [RFC 8949, Section 6.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-6.1)
suggests for converting CBOR to JSON. Note that [`sax_parse`](../../api/basic_json/sax_parse.md) has no parameter for
CBOR tags, so every tag is an error there; when recovering, tags are ignored like with
[`cbor_tag_handler_t::ignore`](../../api/basic_json/cbor_tag_handler_t.md).
After any other error, the end of the item is unknown: the input ended, a byte is not a valid type marker, or a size
cannot be right. Parsing then stops, and the value read so far is completed: a key that waits for its value gets
`#!json null`, and all open arrays and objects are closed. This keeps everything before the error of an input that was
cut off. The exception is BSON, which stores the size of every document: an element whose end is unknown gets
`#!json null`, the rest of its document is skipped, and parsing continues after the document.
## Limitations
@@ -80,6 +105,8 @@ everything before the error of an input that was cut off.
repair differs from the intention: `#!json {"a": {"b": [1, 2}, "c": 3}` is repaired to
`#!json {"a": {"b": [1, 2], "c": 3}}`, although `#!json {"a": {"b": [1, 2]}, "c": 3}` may have been meant.
- Keys without quotes, and strings in single quotes, are not supported; such members are skipped.
- In the binary formats, a member that is skipped because its key is not a string is lost, and so are the elements of a
BSON document after one whose end is unknown.
- A number that is too large for `number_float_t` is passed as positive or negative infinity. The SAX parser's
`number_float` also gets the number's text, but a JSON value cannot store it, and
[`dump`](../../api/basic_json/dump.md) serializes infinity as `#!json null`.
File diff suppressed because it is too large Load Diff
+2 -173
View File
@@ -132,8 +132,8 @@ struct json_sax
@param[in] last_token the last read token
@param[in] ex an exception object describing the error
@return whether to recover from the error: false stops parsing; true
repairs JSON text and continues, or, for the binary formats, stops
after closing the containers read so far
repairs the error and continues, or, if that is not possible,
stops after completing the value read so far
*/
virtual bool parse_error(std::size_t position,
const std::string& last_token,
@@ -1212,176 +1212,5 @@ class json_sax_acceptor
}
};
/*!
@brief SAX proxy that lets the binary readers keep what was read before an error
The binary formats cannot continue after an error: a value's size is given
before its payload, and every byte value is a valid type marker, so there is no
way to find where the next value begins. When the SAX parser's parse_error()
returns true to ask for error recovery, the best the binary readers can offer is
the value read up to the error.
This proxy forwards every event to the SAX parser and records which containers
are open and whether a key still waits for its value. After an error the SAX
parser asked to recover from, @ref close_open_containers then completes the
value with null for a pending key and the missing end events, so the SAX parser
sees balanced events (see #3989).
@tparam BasicJsonType the JSON type
@tparam SAX the SAX parser to forward the events to
*/
template<typename BasicJsonType, typename SAX>
class json_sax_salvager
{
public:
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t;
using string_t = typename BasicJsonType::string_t;
using binary_t = typename BasicJsonType::binary_t;
explicit json_sax_salvager(SAX* sax_) noexcept
: sax(sax_)
{}
bool null()
{
key_pending = false;
return sax->null();
}
bool boolean(bool val)
{
key_pending = false;
return sax->boolean(val);
}
bool number_integer(number_integer_t val)
{
key_pending = false;
return sax->number_integer(val);
}
bool number_unsigned(number_unsigned_t val)
{
key_pending = false;
return sax->number_unsigned(val);
}
bool number_float(number_float_t val, const string_t& s)
{
key_pending = false;
return sax->number_float(val, s);
}
bool string(string_t& val)
{
key_pending = false;
return sax->string(val);
}
bool binary(binary_t& val)
{
key_pending = false;
return sax->binary(val);
}
bool start_object(std::size_t len)
{
key_pending = false;
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(len)))
{
return false;
}
open_containers.push_back(true);
return true;
}
bool key(string_t& val)
{
key_pending = true;
return sax->key(val);
}
bool end_object()
{
JSON_ASSERT(!open_containers.empty() && open_containers.back());
open_containers.pop_back();
return sax->end_object();
}
bool start_array(std::size_t len)
{
key_pending = false;
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(len)))
{
return false;
}
open_containers.push_back(false);
return true;
}
bool end_array()
{
JSON_ASSERT(!open_containers.empty() && !open_containers.back());
open_containers.pop_back();
return sax->end_array();
}
template<class Exception>
bool parse_error(std::size_t position, const std::string& last_token,
const Exception& ex)
{
recovery_requested = sax->parse_error(position, last_token, ex);
// the binary readers stop after an error anyway
return false;
}
/*!
@brief complete the value read before an error
Does nothing unless the SAX parser's parse_error() returned true. Otherwise
passes null for a key that waits for its value and closes the containers
that are still open, innermost first, until an event returns false.
*/
void close_open_containers()
{
if (!recovery_requested)
{
return;
}
recovery_requested = false;
if (key_pending)
{
key_pending = false;
if (JSON_HEDLEY_UNLIKELY(!sax->null()))
{
return;
}
}
while (!open_containers.empty())
{
const bool is_object = open_containers.back();
open_containers.pop_back();
if (JSON_HEDLEY_UNLIKELY(is_object ? !sax->end_object() : !sax->end_array()))
{
return;
}
}
}
private:
/// the SAX parser the events are forwarded to
SAX* sax = nullptr;
/// the containers that are open, innermost last; true for an object
std::vector<bool> open_containers {}; // NOLINT(readability-redundant-member-init)
/// whether a key was passed whose value has not been passed yet
bool key_pending = false;
/// whether the SAX parser's parse_error() asked to recover from the error
bool recovery_requested = false;
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+74
View File
@@ -12,6 +12,7 @@
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t
#include <string> // string, to_string
#include <utility> // move
#include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/macro_scope.hpp>
@@ -133,5 +134,78 @@ inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noex
return state == UTF8_ACCEPT;
}
/*!
@brief append U+FFFD REPLACEMENT CHARACTER, encoded in UTF-8
@param[in,out] s the string to append to
*/
template<typename StringType>
inline void append_replacement_character(StringType& s)
{
s.push_back(static_cast<typename StringType::value_type>(0xEFu));
s.push_back(static_cast<typename StringType::value_type>(0xBFu));
s.push_back(static_cast<typename StringType::value_type>(0xBDu));
}
/*!
@brief replace ill-formed UTF-8 with U+FFFD REPLACEMENT CHARACTER
Each maximal subpart of an ill-formed sequence becomes one U+FFFD, as the
Unicode Standard recommends (Section 3.9, "U+FFFD Substitution of Maximal
Subparts"), and as the parser for JSON text does when it recovers from errors.
@param[in,out] s the string to repair
@param[in] first index of the first byte to repair; the bytes before it are
assumed to be valid UTF-8 that ends on a code point boundary
*/
template<typename StringType>
inline void replace_invalid_utf8(StringType& s, const std::size_t first = 0)
{
StringType result = s;
result.resize(first);
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
// the first byte of the sequence being decoded
std::size_t sequence_start = first;
std::size_t i = first;
while (i < s.size())
{
switch (decode(state, codepoint, static_cast<std::uint8_t>(s[i])))
{
case UTF8_ACCEPT:
for (++i; sequence_start < i; ++sequence_start)
{
result.push_back(s[sequence_start]);
}
break;
case UTF8_REJECT:
append_replacement_character(result);
// the byte that made the sequence ill-formed begins the next
// one, unless it began this one
if (i == sequence_start)
{
++i;
}
state = UTF8_ACCEPT;
sequence_start = i;
break;
default: // in the middle of a sequence
++i;
break;
}
}
// a sequence that the string ends in the middle of
if (state != UTF8_ACCEPT)
{
append_replacement_character(result);
}
s = std::move(result);
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+4 -24
View File
@@ -143,7 +143,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
friend class ::nlohmann::detail::iter_impl;
template<typename BasicJsonType, typename CharType, typename OutputSinkType>
friend class ::nlohmann::detail::binary_writer;
template<typename BasicJsonType, typename InputType, typename SAX>
template<typename BasicJsonType, typename InputType, typename SAX, bool AllowRecovery>
friend class ::nlohmann::detail::binary_reader;
template<typename BasicJsonType, typename InputAdapterType>
friend class ::nlohmann::detail::json_sax_dom_parser;
@@ -4979,26 +4979,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
return parser(i.get(), nullptr, false, ignore_comments, ignore_trailing_commas, true).accept(true);
}
private:
/// read a binary format and pass it to a SAX parser; if the SAX parser
/// asks to recover from an error, the value read so far is completed
/// (see detail::json_sax_salvager and #3989)
template<typename InputAdapterType, typename SAX>
static bool sax_parse_binary(InputAdapterType ia, SAX* sax,
const input_format_t format, const bool strict)
{
(void)detail::is_sax_static_asserts<SAX, basic_json> {};
using salvager_t = detail::json_sax_salvager<basic_json, SAX>;
salvager_t salvager(sax);
const bool result = detail::binary_reader<basic_json, InputAdapterType, salvager_t>(std::move(ia), format).sax_parse(format, &salvager, strict);
if (!result)
{
salvager.close_open_containers();
}
return result;
}
public:
/// @brief generate SAX events
/// @sa https://json.nlohmann.me/api/basic_json/sax_parse/
template <typename InputType, typename SAX>
@@ -5012,7 +4992,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: sax_parse_binary(std::move(ia), sax, format, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(format, sax, strict);
}
/// @brief generate SAX events (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -5029,7 +5009,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: sax_parse_binary(std::move(ia), sax, format, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(format, sax, strict);
}
/// @brief generate SAX events
@@ -5051,7 +5031,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
: sax_parse_binary(std::move(ia), sax, format, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(format, sax, strict);
}
#ifndef JSON_NO_IO
/// @brief deserialize from stream
File diff suppressed because it is too large Load Diff
+12
View File
@@ -45,6 +45,10 @@ dumps is stable under exactly the same values that break operator==.
The unit tests run the same checks on a fixed corpus (see the "BJData round-trip
invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bjdata() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -59,6 +63,8 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// value-stable comparison for the round-trip checks below; see the note
@@ -71,11 +77,15 @@ static bool is_value_stable(const json& lhs, const json& rhs)
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bjdata).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bjdata(vec1);
assert(recovered_without_errors);
try
{
@@ -109,6 +119,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -117,6 +128,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -19,6 +19,10 @@ It also checks that reading the data from a stream, which reads strings byte by
byte, gives the same value or error as reading it from contiguous memory, which
copies strings in bulk.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bon8() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -33,6 +37,8 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
namespace
@@ -56,6 +62,9 @@ std::string read_bon8(InputType&& input)
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bon8).errors == 0;
// contiguous and stream input must be read alike
{
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
@@ -67,6 +76,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bon8(vec1);
assert(recovered_without_errors);
try
{
@@ -88,6 +98,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -96,6 +107,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_bson(vec)
- assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -29,16 +33,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bson).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bson(vec1);
assert(recovered_without_errors);
if (j1.is_discarded())
{
@@ -65,6 +75,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -73,6 +84,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors can occur during parsing, too
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_cbor(vec)
- assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_cbor() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -29,16 +33,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::cbor).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_cbor(vec1);
assert(recovered_without_errors);
try
{
@@ -60,6 +70,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -68,6 +79,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors can occur during parsing, too
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+4 -130
View File
@@ -28,7 +28,6 @@ drivers.
#include <iostream>
#include <sstream>
#include <string>
#include <vector>
#include <nlohmann/json.hpp>
// the round-trip checks below are assertions; NDEBUG would compile them away
@@ -36,143 +35,18 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
namespace
{
// a SAX parser that recovers from every error and checks that the events are
// balanced and that every key is followed by exactly one value
class recovering_checker : public nlohmann::json_sax<json>
{
public:
bool null() override
{
return value();
}
bool boolean(bool /*val*/) override
{
return value();
}
bool number_integer(number_integer_t /*val*/) override
{
return value();
}
bool number_unsigned(number_unsigned_t /*val*/) override
{
return value();
}
bool number_float(number_float_t /*val*/, const string_t& /*s*/) override
{
return value();
}
bool string(string_t& /*val*/) override
{
return value();
}
bool binary(binary_t& /*val*/) override
{
return value();
}
bool start_object(std::size_t /*elements*/) override
{
value();
stack.push_back('o');
return true;
}
bool key(string_t& /*val*/) override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.back() = 'v';
return true;
}
bool end_object() override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.pop_back();
return true;
}
bool start_array(std::size_t /*elements*/) override
{
value();
stack.push_back('a');
return true;
}
bool end_array() override
{
++events;
assert(!stack.empty() && stack.back() == 'a');
stack.pop_back();
return true;
}
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const nlohmann::detail::exception& /*ex*/) override
{
++errors;
return true;
}
bool complete() const
{
return stack.empty();
}
std::size_t events = 0;
std::size_t errors = 0;
private:
bool value()
{
++events;
if (!stack.empty())
{
// an array element, or the value of a key
assert(stack.back() != 'o');
if (stack.back() == 'v')
{
stack.back() = 'o';
}
}
return true;
}
// 'a' for an array, 'o' for an object that expects a key, 'v' for an
// object that expects the value of a key
std::vector<char> stack;
};
} // namespace
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
{
recovering_checker checker;
const bool ok = json::sax_parse(data, data + size, &checker);
assert(checker.complete());
assert(checker.errors <= size + 1);
const auto checker = check_recovering_parse(data, size, json::input_format_t::json);
assert(checker.events <= (4 * size) + 4);
assert(ok == json::accept(data, data + size));
assert(ok == (checker.errors == 0));
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
recovering_checker stream_checker;
assert(json::sax_parse(stream, &stream_checker) == ok);
assert(stream_checker.complete());
assert(stream_checker.events == checker.events);
assert(stream_checker.errors == checker.errors);
assert((checker.errors == 0) == json::accept(data, data + size));
}
try
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_msgpack(vec)
- assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_msgpack() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -29,16 +33,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::msgpack).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_msgpack(vec1);
assert(recovered_without_errors);
try
{
@@ -60,6 +70,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -68,6 +79,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+12
View File
@@ -24,6 +24,10 @@ array data, it performs the following steps:
The unit tests run the same checks on a fixed corpus (see the "UBJSON round-trip
invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_ubjson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -38,16 +42,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::ubjson).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_ubjson(vec1);
assert(recovered_without_errors);
try
{
@@ -79,6 +89,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -87,6 +98,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
+154
View File
@@ -0,0 +1,154 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cassert>
#include <cstddef>
#include <cstdint>
#include <sstream>
#include <string>
#include <vector>
#include <nlohmann/json.hpp>
namespace
{
// a SAX parser that recovers from every error and checks that the events are
// balanced and that every key is followed by exactly one value
class recovering_checker : public nlohmann::json_sax<nlohmann::json>
{
public:
bool null() override
{
return value();
}
bool boolean(bool /*val*/) override
{
return value();
}
bool number_integer(number_integer_t /*val*/) override
{
return value();
}
bool number_unsigned(number_unsigned_t /*val*/) override
{
return value();
}
bool number_float(number_float_t /*val*/, const string_t& /*s*/) override
{
return value();
}
bool string(string_t& /*val*/) override
{
return value();
}
bool binary(binary_t& /*val*/) override
{
return value();
}
bool start_object(std::size_t /*elements*/) override
{
value();
stack.push_back('o');
return true;
}
bool key(string_t& /*val*/) override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.back() = 'v';
return true;
}
bool end_object() override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.pop_back();
return true;
}
bool start_array(std::size_t /*elements*/) override
{
value();
stack.push_back('a');
return true;
}
bool end_array() override
{
++events;
assert(!stack.empty() && stack.back() == 'a');
stack.pop_back();
return true;
}
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const nlohmann::detail::exception& /*ex*/) override
{
++errors;
return true;
}
bool complete() const
{
return stack.empty();
}
std::size_t events = 0;
std::size_t errors = 0;
private:
bool value()
{
++events;
if (!stack.empty())
{
// an array element, or the value of a key
assert(stack.back() != 'o');
if (stack.back() == 'v')
{
stack.back() = 'o';
}
}
return true;
}
// 'a' for an array, 'o' for an object that expects a key, 'v' for an
// object that expects the value of a key
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
};
/// parses @a data with a recovering_checker from memory and from a stream,
/// checks that both see the same, that the events are balanced, and that the
/// number of errors is bounded, and returns the checker (see #3989)
inline recovering_checker check_recovering_parse(const std::uint8_t* data, const std::size_t size, const nlohmann::json::input_format_t format)
{
recovering_checker checker;
const bool ok = nlohmann::json::sax_parse(data, data + size, &checker, format);
assert(checker.complete());
assert(checker.errors <= size + 1);
assert(ok == (checker.errors == 0));
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
recovering_checker stream_checker;
assert(nlohmann::json::sax_parse(stream, &stream_checker, format) == ok);
assert(stream_checker.complete());
assert(stream_checker.events == checker.events);
assert(stream_checker.errors == checker.errors);
return checker;
}
} // namespace
+314 -3
View File
@@ -972,8 +972,9 @@ class RecoveringParser : public nlohmann::detail::json_sax_dom_parser<json>
return base::end_array();
}
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& /*unused*/)
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& ex)
{
messages.emplace_back(ex.what());
// a limit, so that a reader that does not stop fails the test
// instead of making it hang
return ++errors < 100;
@@ -986,6 +987,7 @@ class RecoveringParser : public nlohmann::detail::json_sax_dom_parser<json>
}
std::size_t errors = 0;
std::vector<std::string> messages {}; // NOLINT(readability-redundant-member-init)
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
bool well_formed = true;
@@ -1011,6 +1013,7 @@ struct BinaryParseResult
{
json value;
std::size_t errors;
std::vector<std::string> messages;
bool ok;
bool balanced;
};
@@ -1020,13 +1023,117 @@ BinaryParseResult parse_binary_recovering(const std::vector<std::uint8_t>& input
json j;
RecoveringParser sax(j);
const bool ok = json::sax_parse(input, &sax, format);
return {j, sax.errors, ok, sax.balanced()};
return {j, sax.errors, sax.messages, ok, sax.balanced()};
}
/// the message of the exception that reading @a input into a JSON value
/// throws, or an empty string if reading succeeds
std::string binary_error_message(const std::vector<std::uint8_t>& input, const json::input_format_t format)
{
try
{
json _;
switch (format)
{
case json::input_format_t::cbor:
_ = json::from_cbor(input);
break;
case json::input_format_t::msgpack:
_ = json::from_msgpack(input);
break;
case json::input_format_t::ubjson:
_ = json::from_ubjson(input);
break;
case json::input_format_t::bjdata:
_ = json::from_bjdata(input);
break;
case json::input_format_t::bson:
_ = json::from_bson(input);
break;
case json::input_format_t::bon8:
_ = json::from_bon8(input);
break;
case json::input_format_t::json:
default:
break;
}
}
catch (const json::exception& e)
{
return e.what();
}
return "";
}
/// a BSON element: its type, its name, and its value
std::vector<std::uint8_t> bson_element(const std::uint8_t type, const std::string& name, const std::vector<std::uint8_t>& value)
{
std::vector<std::uint8_t> result = {type};
result.insert(result.end(), name.begin(), name.end());
result.push_back(0x00);
result.insert(result.end(), value.begin(), value.end());
return result;
}
/// a BSON document of the given elements; @a size_offset is added to the
/// size it declares
std::vector<std::uint8_t> bson_document(const std::vector<std::vector<std::uint8_t>>& elements, const int size_offset = 0)
{
std::vector<std::uint8_t> body;
for (const auto& element : elements)
{
body.insert(body.end(), element.begin(), element.end());
}
const auto size = static_cast<std::uint32_t>(static_cast<int>(body.size()) + 5 + size_offset);
std::vector<std::uint8_t> result = {static_cast<std::uint8_t>(size & 0xFFu), static_cast<std::uint8_t>((size >> 8u) & 0xFFu),
static_cast<std::uint8_t>((size >> 16u) & 0xFFu), static_cast<std::uint8_t>((size >> 24u) & 0xFFu)
};
result.insert(result.end(), body.begin(), body.end());
result.push_back(0x00);
return result;
}
/// a BSON int32 value
std::vector<std::uint8_t> bson_int32(const std::int32_t value)
{
const auto u = static_cast<std::uint32_t>(value);
return {static_cast<std::uint8_t>(u & 0xFFu), static_cast<std::uint8_t>((u >> 8u) & 0xFFu),
static_cast<std::uint8_t>((u >> 16u) & 0xFFu), static_cast<std::uint8_t>((u >> 24u) & 0xFFu)};
}
/// a BSON string value, whose length is @a length_offset off
std::vector<std::uint8_t> bson_string(const std::string& value, const std::int32_t length_offset = 0)
{
auto result = bson_int32(static_cast<std::int32_t>(value.size() + 1) + length_offset);
result.insert(result.end(), value.begin(), value.end());
result.push_back(0x00);
return result;
}
/// @a count bytes of value 0xAB
std::vector<std::uint8_t> bytes(const std::size_t count)
{
return std::vector<std::uint8_t>(count, 0xAB);
}
template<typename... Parts>
std::vector<std::uint8_t> concatenated(const std::vector<std::uint8_t>& first, const Parts& ... rest)
{
std::vector<std::uint8_t> result = first;
for (const auto& part : std::initializer_list<std::vector<std::uint8_t>> {rest...})
{
result.insert(result.end(), part.begin(), part.end());
}
return result;
}
/// U+FFFD REPLACEMENT CHARACTER
const std::string replacement_character = "\xEF\xBF\xBD";
} // namespace
TEST_CASE("regression test - #3989 SAX parse_error() returning true")
{
SECTION("binary formats stop after an error and complete what was read")
SECTION("binary formats complete what was read before the input ends")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}};
@@ -1108,6 +1215,210 @@ TEST_CASE("regression test - #3989 SAX parse_error() returning true")
CHECK(result.value == json({{"_ArrayType_", "int8"}, {"_ArraySize_", {2, 3}}, {"_ArrayData_", {1, 2}}}));
}
SECTION("binary formats repair items whose end is known")
{
struct Repair
{
json::input_format_t format;
std::vector<std::uint8_t> input;
json expected;
std::size_t errors;
};
const std::vector<Repair> repairs =
{
// CBOR: tags are ignored (here tag 1 and the self-describe tag 55799)
{json::input_format_t::cbor, {0x82, 0xC1, 0x05, 0xD9, 0xD9, 0xF7, 0x06}, {5, 6}, 2},
// CBOR: undefined and other simple values become null
{json::input_format_t::cbor, {0x84, 0xF7, 0xE0, 0xF8, 0x20, 0x01}, {nullptr, nullptr, nullptr, 1}, 3},
// CBOR: ill-formed UTF-8 becomes U+FFFD, also in keys
{json::input_format_t::cbor, {0xA1, 0x61, 0xFF, 0x62, 0xC3, 0x28}, {{replacement_character, replacement_character + "("}}, 2},
// CBOR: members whose key is not a string are skipped, whatever their key and value
{json::input_format_t::cbor, {0xA4, 0x01, 0x02, 0x82, 0x01, 0x02, 0xA1, 0x61, 'x', 0x9F, 0xFF, 0xC1, 0x01, 0x5F, 0x41, 0x00, 0xFF, 0x61, 'a', 0x03}, {{"a", 3}}, 3},
{json::input_format_t::cbor, {0xBF, 0xF5, 0xBF, 0x61, 'x', 0x7F, 0x61, 'y', 0xFF, 0xFF, 0x61, 'a', 0x03, 0xFF}, {{"a", 3}}, 1},
// MessagePack: members whose key is not a string are skipped
{json::input_format_t::msgpack, {0x84, 0x01, 0x02, 0x81, 0xA1, 'x', 0x01, 0x92, 0x01, 0x02, 0xD4, 0x01, 0x02, 0xC0, 0xA1, 'a', 0x04}, {{"a", 4}}, 3},
// MessagePack: ill-formed UTF-8 becomes U+FFFD
{json::input_format_t::msgpack, {0x92, 0xA2, 0xC3, 0x28, 0xA3, 0xE2, 0x82, 'x'}, {replacement_character + "(", replacement_character + "x"}, 2},
// UBJSON: a char that is not ASCII becomes U+FFFD
{json::input_format_t::ubjson, {'[', 'C', 0x80, 'C', 'A', ']'}, {replacement_character, "A"}, 1},
// UBJSON: the longest beginning of a high-precision number is kept
{json::input_format_t::ubjson, {'[', 'H', 'i', 5, '1', '2', 'a', 'b', 'c', 'H', 'i', 2, '1', '.', 'H', 'i', 3, 'a', 'b', 'c', 'H', 'i', 3, '4', '.', '5', ']'}, {12, 1, nullptr, 4.5}, 3},
// BJData, too
{json::input_format_t::bjdata, {'[', 'C', 0xFF, 'H', 'i', 2, '-', '1', 'H', 'i', 2, '-', 'x', ']'}, {replacement_character, -1, nullptr}, 2},
// BON8: members whose key is not a string are skipped
{json::input_format_t::bon8, {0x89, 0x91, 0x92, 0xC9, 0x40, 0x82, 0x91, 0x92, 0x61, 0x93}, {{"a", 3}}, 2},
{json::input_format_t::bon8, {0x8B, 0x91, 0x85, 0x91, 0xFE, 0xFA, 0x8B, 'x', 0x91, 0xFE, 0x61, 0x93, 0xFE}, {{"a", 3}}, 2},
// BSON: elements of types the library does not read become null
{
json::input_format_t::bson, bson_document(
{
bson_element(0x07, "_id", bytes(12)), // ObjectId
bson_element(0x09, "date", bytes(8)), // UTC datetime
bson_element(0x13, "decimal", bytes(16)), // 128-bit decimal
bson_element(0x0B, "regex", {'a', '+', 0, 'i', 0}), // regular expression
bson_element(0x0D, "code", bson_string("f()")), // JavaScript code
bson_element(0x0E, "symbol", bson_string("s")), // symbol
bson_element(0x0C, "pointer", concatenated(bson_string("c"), bytes(12))), // DBPointer
bson_element(0x0F, "scope", concatenated(bson_int32(15), bson_string("g"), bson_document({}))), // code with scope
bson_element(0x06, "undefined", {}), // undefined
bson_element(0xFF, "min", {}), // min key
bson_element(0x7F, "max", {}), // max key
bson_element(0x10, "z", bson_int32(7)),
}),
{{"_id", nullptr}, {"date", nullptr}, {"decimal", nullptr}, {"regex", nullptr}, {"code", nullptr}, {"symbol", nullptr}, {"pointer", nullptr}, {"scope", nullptr}, {"undefined", nullptr}, {"min", nullptr}, {"max", nullptr}, {"z", 7}},
11
},
// BSON: an element of an unknown type becomes null, and the rest of its document is skipped
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x04, "array", bson_document({bson_element(0x10, "0", bson_int32(1)), bson_element(0x42, "1", bytes(3))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"a", 1}, {"x", nullptr}}}, {"array", {1, nullptr}}, {"after", 3}},
2
},
// BSON: so does a string or byte array whose length cannot be right
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x02, "s", bson_string("abc", -10)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x03, "bin", bson_document({bson_element(0x05, "b", concatenated(bson_int32(-1), bytes(1))), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"s", nullptr}}}, {"bin", {{"b", nullptr}}}, {"after", 3}},
2
},
// BSON: a string without its terminator, and a document whose size does not match, are kept
{
json::input_format_t::bson, bson_document(
{
bson_element(0x02, "s", {2, 0, 0, 0, 'a', 'X'}),
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1))}, 1)),
}),
{{"s", "a"}, {"inner", {{"a", 1}}}},
2
},
};
for (const auto& repair : repairs)
{
CAPTURE(repair.format);
CAPTURE(repair.input);
const auto result = parse_binary_recovering(repair.input, repair.format);
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.errors == repair.errors);
CHECK(result.value == repair.expected);
// the first error is the one reported without recovering
REQUIRE(!result.messages.empty());
CHECK(result.messages.front() == binary_error_message(repair.input, repair.format));
}
}
SECTION("binary formats repair numbers that are out of range")
{
// CBOR: a negative integer below the range of number_integer_t
const auto cbor = parse_binary_recovering({0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value.is_number_float());
CHECK(cbor.value.get<double>() == -18446744073709551616.0);
// UBJSON: a high-precision number too large for number_float_t
const auto ubjson = parse_binary_recovering({'H', 'i', 5, '1', 'e', '9', '9', '9'}, json::input_format_t::ubjson);
CHECK(ubjson.errors == 1);
CHECK(ubjson.value.is_number_float());
CHECK(std::isinf(ubjson.value.get<double>()));
}
SECTION("binary formats stop where the end of an item is not known")
{
// a byte that begins no item
const auto cbor = parse_binary_recovering({0x82, 0x01, 0x1C, 0x02}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value == json({1}));
// a key that is no item: the unused MessagePack byte, a CBOR break
// in a map of known size, and the end of a BON8 container
const auto msgpack = parse_binary_recovering({0x82, 0xA1, 'a', 0x01, 0xC1, 0x02}, json::input_format_t::msgpack);
CHECK(msgpack.errors == 1);
CHECK(msgpack.value == json({{"a", 1}}));
const auto cbor_break = parse_binary_recovering({0xA2, 0x61, 'a', 0x01, 0xFF, 0x02}, json::input_format_t::cbor);
CHECK(cbor_break.errors == 1);
CHECK(cbor_break.value == json({{"a", 1}}));
const auto bon8 = parse_binary_recovering({0x88, 0x61, 0x91, 0xFE}, json::input_format_t::bon8);
CHECK(bon8.errors == 1);
CHECK(bon8.value == json({{"a", 1}}));
// a skipped member that the input ends in
const auto truncated = parse_binary_recovering({0xA2, 0x01, 0x82, 0x01}, json::input_format_t::cbor);
CHECK(truncated.errors == 2);
CHECK(truncated.balanced);
CHECK(truncated.value == json::object());
// a BSON element of an unknown type in a document whose size cannot be right
const auto bson = parse_binary_recovering(bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3))}, -10), json::input_format_t::bson);
CHECK(bson.errors == 1);
CHECK(bson.value == json({{"a", 1}, {"x", nullptr}}));
}
SECTION("changed bytes in binary input")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}, {"i", "\xC3\xA4"}};
const std::vector<std::pair<json::input_format_t, std::vector<std::uint8_t>>> encodings =
{
{json::input_format_t::cbor, json::to_cbor(j)},
{json::input_format_t::msgpack, json::to_msgpack(j)},
{json::input_format_t::ubjson, json::to_ubjson(j)},
{json::input_format_t::ubjson, json::to_ubjson(j, true, true)},
{json::input_format_t::bjdata, json::to_bjdata(j)},
{json::input_format_t::bjdata, json::to_bjdata(j, true, true)},
{json::input_format_t::bson, json::to_bson(j)},
{json::input_format_t::bon8, json::to_bon8(j)},
};
const std::vector<std::uint8_t> replacements = {0x00, 0x01, 0x7F, 0x80, 0xC1, 0xD9, 0xE0, 0xF7, 0xFE, 0xFF};
for (const auto& encoding : encodings)
{
const auto format = encoding.first;
const auto& original = encoding.second;
CAPTURE(format);
std::vector<std::vector<std::uint8_t>> inputs;
for (std::size_t position = 0; position < original.size(); ++position)
{
for (const auto replacement : replacements)
{
auto changed = original;
changed[position] = replacement;
inputs.push_back(changed);
}
auto removed = original;
removed.erase(removed.begin() + static_cast<std::ptrdiff_t>(position));
inputs.push_back(removed);
}
for (const auto& input : inputs)
{
CAPTURE(input);
const auto result = parse_binary_recovering(input, format);
CHECK(result.balanced);
CHECK(result.errors <= input.size() + 1);
// an error is reported exactly if reading into a JSON value
// fails, and the first one is the same
const auto message = binary_error_message(input, format);
CHECK(result.ok == message.empty());
if (!result.ok && result.errors < 100)
{
CHECK(result.messages.front() == message);
}
}
}
}
SECTION("JSON text")
{
// the parser stopped, but reported success