An image is a document stored so that loading it needs no
parsing: save() writes the node index, the text and the decoded
strings; the static load() reads an image written by save().
load() takes a pointer and size, a borrowed vector, or an owned
rvalue vector; the nodes are copied so they are aligned and can
be edited, while the text and decoded strings stay in the image.
image_check controls how much load() trusts the input: full
checks structure, bounds, strings and numbers, the parser's own
guarantees; bounds checks structure and bounds only; none skips
all checks, for images from a trusted source.
Layout is little-endian only ("NJVI" header, nodes, text, decoded
strings), following the idea of zero-copy formats such as
FlatBuffers and YaFF; the check follows FlatBuffers' Verifier.
New errors: parse_error.116 for a malformed image or a failed
check, type_error.320 for a discarded document or a big-endian
target.
A dedicated fuzzer and 6,000 seeded corruptions, checked under
ASan/UBSan, found and fixed two gaps: unchecked reserved header
fields, and unbounded null/boolean offsets that could make
dump() throw std::length_error.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
7.0 KiB
Fuzz testing
Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently, libFuzzer and afl++ are supported.
Additionally, parse_json_view_fuzzer (tests/src/fuzzer-parse_json_view.cpp) cross-checks json_document/json_view
(the zero-copy, read-only view declared in json_view.hpp) against basic_json on the same JSON text: it asserts that
json_document::accept agrees with json::accept, that an accepted input materializes to the same value json::parse
produces, and that a rejected input makes both parsers throw with an identical what(). It takes plain JSON text, so it
reuses the corpus_json corpus rather than a format of its own.
json_view_image_fuzzer (tests/src/fuzzer-json_view_image.cpp) tests the images of json_document (save() and
load()). It uses each input twice: as an image, which load() must either reject with parse_error.116 or read
safely (with image_check::full, the document must also serialize to the JSON it reads as), and as a JSON text, whose
image must load and serialize to the same text. A corpus of images can be made from JSON files with a small program
that calls json_document::parse(text).save(); plain JSON files work as well.
What the fuzzers check
Each fuzzer driver (tests/src/fuzzer-parse_*.cpp) parses its input twice: once with allow_exceptions = false and
once with exceptions. Both calls must agree. Where parsing with exceptions fails, the call without exceptions must
return a discarded value (or throw the same kind of non-parse error), and it must never throw a parse_error. Where
parsing succeeds, both calls must return the same value. The drivers then serialize the value, parse the result back,
and check that nothing was lost. The drivers check all of this with assert, so they refuse to build with NDEBUG.
Corpus creation
For most effective fuzzing, a corpus should be provided. A corpus is a directory with some simple input files that cover several features of the parser and is hence a good starting point for mutations.
TEST_DATA_VERSION=3.2.0
wget https://github.com/nlohmann/json_test_data/archive/refs/tags/v$TEST_DATA_VERSION.zip
unzip v$TEST_DATA_VERSION.zip
rm v$TEST_DATA_VERSION.zip
for FORMAT in json bjdata bon8 bson cbor msgpack ubjson
do
rm -fr corpus_$FORMAT
mkdir corpus_$FORMAT
find json_test_data-$TEST_DATA_VERSION -size -5k -name "*.$FORMAT" -exec cp "{}" "corpus_$FORMAT" \;
done
rm -fr json_test_data-$TEST_DATA_VERSION
The generated corpus can be used with both libFuzzer and afl++. The remainder of this documentation assumes the corpus
directories have been created in the tests directory.
libFuzzer
To use libFuzzer, you need to pass -fsanitize=fuzzer as FUZZER_ENGINE. In the tests directory, call
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
This creates a fuzz tester binary for each parser that supports these command line options.
In case your default compiler is not a Clang compiler that includes libFuzzer (Clang 6.0 or later), you need to set the
CXX variable accordingly. Note the compiler provided by Xcode (AppleClang) does not contain libFuzzer. Please install
Clang via Homebrew calling brew install llvm and add CXX=$(brew --prefix llvm)/bin/clang to the make call:
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" CXX=$(brew --prefix llvm)/bin/clang
Then pass the corpus directory as command-line argument (assuming it is located in tests):
./parse_cbor_fuzzer corpus_cbor
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into
a file starting with crash-.
To also detect memory leaks, build with AddressSanitizer (FUZZER_ENGINE="-fsanitize=fuzzer,address"): libFuzzer then
runs LeakSanitizer by default (-detect_leaks=1). LeakSanitizer is not available with Apple Clang on macOS.
afl++
To use afl++, you need to pass -fsanitize=fuzzer as FUZZER_ENGINE. It will be replaced by a libAFLDriver.a to
re-use the same code written for libFuzzer with afl++. Furthermore, set afl-clang-fast++ as compiler.
CXX=afl-clang-fast++ make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
Then the fuzzer is called like this in the tests directory:
afl-fuzz -i corpus_cbor -o out -- ./parse_cbor_fuzzer
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is written to the
directory out.
OSS-Fuzz
The library is further fuzz-tested 24/7 by Google's OSS-Fuzz project. It uses
the same fuzzers target as above and also relies on the FUZZER_ENGINE variable. See the used
build script for more information. Its default
address sanitizer includes LeakSanitizer, so OSS-Fuzz and the CIFuzz workflow (.github/workflows/cifuzz.yml) report
memory leaks, too.
In case the build at OSS-Fuzz fails, an issue will be created automatically.
Handling OSS-Fuzz reports
OSS-Fuzz files the crashes it finds in its own issue tracker, not on GitHub. So that each report can be traced to the change that fixed it, and each fix to the report it answers, fixes follow these conventions:
- Reference the OSS-Fuzz issue in the pull request, next to any GitHub issue it closes, as
OSS-Fuzz: <id>(for example,OSS-Fuzz: 563659413), and in the commit message. The ID alone does not disclose the crash. If the report was triaged into a GitHub issue, link the OSS-Fuzz issue there too. - Turn the reproducer into a unit test. Download the testcase from the OSS-Fuzz report, reduce it if possible, and
add it as a regression test to the unit test of the affected format (e.g.,
tests/src/unit-bjdata.cpp), with a comment naming the OSS-Fuzz issue. This way the input is checked by every CI run rather than only by OSS-Fuzz, and it stays covered even if OSS-Fuzz later closes the report as not reproducible. - Keep the fuzzer drivers and the unit tests in sync. The round-trip checks of the BJData, BON8, BSON, CBOR,
MessagePack and UBJSON drivers are also run on a fixed corpus in the unit tests (see
tests/src/round_trip_corpus.hppand the "round-trip invariants" test cases), so a regression shows up in CI first. When a driver's checks change, change the unit tests with them. - Record in the report whether the bug shipped. OSS-Fuzz asks whether a crash was a short-lived regression or affects a released version; answer it when the fix is merged, as it decides whether the fix needs a release note or a security advisory (see the security policy).
After the fix is merged, OSS-Fuzz re-runs the reproducer on its next build and marks the report as verified and closed. If it does not, the fix is incomplete.