Repair complete items in binary formats when parse_error() returns true (#3989)

When the SAX parser asks to recover, the binary readers now repair an
item whose end is known and read on after it, as RFC 8949, Section 5.3
describes for CBOR:

- CBOR: tags are ignored, and simple values other than false, true, and
  null become null (RFC 8949, Section 6.1); a negative integer below the
  range of number_integer_t becomes the nearest floating-point number.
- Strings that are not valid UTF-8 get U+FFFD for each ill-formed
  sequence, as in JSON text; so does a UBJSON/BJData char above 0x7F.
- UBJSON/BJData high-precision numbers keep their longest valid
  beginning (via the lexer's recover_token()), or become infinity.
- Members whose key is not a string are skipped (CBOR, MessagePack,
  BON8), like members without a key in JSON text.
- BSON elements of types the library does not read (ObjectId, datetime,
  decimal128, ...) become null; a string without its terminator and a
  document whose size does not match are kept.

Where the end of an item is unknown, reading stops as before, except
that BSON skips to the end of the document, whose size it knows.

The value read before such an error is now completed by the reader from
its container stack, as the JSON parser does, instead of by a proxy SAX
parser, which is removed. Like the parser, binary_reader gets an
AllowRecovery template parameter, so that from_*() compile without the
new code.

Tests: a table of repairs, numbers out of range, errors that stop, and
a sweep over changed and removed bytes of eight encodings that checks
balanced events and that the first error is the one from_*() reports.
All fuzzers now run a recovering checker; the binary ones also check
that it reports an error exactly when from_*() fails.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann committed 2026-09-27 21:40:59 +02:00
1 parent c8735246d0
commit 437a95cfdb
16 files changed
+2925 -673

No files matched your search

+12
View File
@@ -19,6 +19,10 @@ It also checks that reading the data from a stream, which reads strings byte by
byte, gives the same value or error as reading it from contiguous memory, which
copies strings in bulk.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bon8() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -33,6 +37,8 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
namespace
@@ -56,6 +62,9 @@ std::string read_bon8(InputType&& input)
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bon8).errors == 0;
// contiguous and stream input must be read alike
{
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
@@ -67,6 +76,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bon8(vec1);
assert(recovered_without_errors);
try
{
@@ -88,6 +98,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -96,6 +107,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use