Files
json/tests/fuzzing.md
T
Niels Lohmann 759dd2e1b1 Add json_document and json_view: node index, parser, and document
Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).

A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.

The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.

basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.

detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.

A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-07 16:42:22 +02:00

125 lines
6.4 KiB
Markdown

# Fuzz testing
Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently,
[libFuzzer](https://llvm.org/docs/LibFuzzer.html) and [afl++](https://github.com/AFLplusplus/AFLplusplus) are supported.
Additionally, `parse_json_view_fuzzer` (`tests/src/fuzzer-parse_json_view.cpp`) cross-checks `json_document`/`json_view`
(the zero-copy, read-only view declared in `json_view.hpp`) against `basic_json` on the same JSON text: it asserts that
`json_document::accept` agrees with `json::accept`, that an accepted input materializes to the same value `json::parse`
produces, and that a rejected input makes both parsers throw with an identical `what()`. It takes plain JSON text, so it
reuses the `corpus_json` corpus rather than a format of its own.
## What the fuzzers check
Each fuzzer driver (`tests/src/fuzzer-parse_*.cpp`) parses its input twice: once with `allow_exceptions = false` and
once with exceptions. Both calls must agree. Where parsing with exceptions fails, the call without exceptions must
return a discarded value (or throw the same kind of non-parse error), and it must never throw a `parse_error`. Where
parsing succeeds, both calls must return the same value. The drivers then serialize the value, parse the result back,
and check that nothing was lost. The drivers check all of this with `assert`, so they refuse to build with `NDEBUG`.
## Corpus creation
For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a
directory with some simple input files that cover several features of the parser and is hence a good starting point
for mutations.
```shell
TEST_DATA_VERSION=3.2.0
wget https://github.com/nlohmann/json_test_data/archive/refs/tags/v$TEST_DATA_VERSION.zip
unzip v$TEST_DATA_VERSION.zip
rm v$TEST_DATA_VERSION.zip
for FORMAT in json bjdata bon8 bson cbor msgpack ubjson
do
rm -fr corpus_$FORMAT
mkdir corpus_$FORMAT
find json_test_data-$TEST_DATA_VERSION -size -5k -name "*.$FORMAT" -exec cp "{}" "corpus_$FORMAT" \;
done
rm -fr json_test_data-$TEST_DATA_VERSION
```
The generated corpus can be used with both libFuzzer and afl++. The remainder of this documentation assumes the corpus
directories have been created in the `tests` directory.
## libFuzzer
To use libFuzzer, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. In the `tests` directory, call
```shell
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
```
This creates a fuzz tester binary for each parser that supports these
[command line options](https://llvm.org/docs/LibFuzzer.html#options).
In case your default compiler is not a Clang compiler that includes libFuzzer (Clang 6.0 or later), you need to set the
`CXX` variable accordingly. Note the compiler provided by Xcode (AppleClang) does not contain libFuzzer. Please install
Clang via Homebrew calling `brew install llvm` and add `CXX=$(brew --prefix llvm)/bin/clang` to the `make` call:
```shell
make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer" CXX=$(brew --prefix llvm)/bin/clang
```
Then pass the corpus directory as command-line argument (assuming it is located in `tests`):
```shell
./parse_cbor_fuzzer corpus_cbor
```
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into
a file starting with `crash-`.
To also detect memory leaks, build with AddressSanitizer (`FUZZER_ENGINE="-fsanitize=fuzzer,address"`): libFuzzer then
runs LeakSanitizer by default (`-detect_leaks=1`). LeakSanitizer is not available with Apple Clang on macOS.
## afl++
To use afl++, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. It will be replaced by a `libAFLDriver.a` to
re-use the same code written for libFuzzer with afl++. Furthermore, set `afl-clang-fast++` as compiler.
```shell
CXX=afl-clang-fast++ make fuzzers FUZZER_ENGINE="-fsanitize=fuzzer"
```
Then the fuzzer is called like this in the `tests` directory:
```shell
afl-fuzz -i corpus_cbor -o out -- ./parse_cbor_fuzzer
```
The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is written to the
directory `out`.
## OSS-Fuzz
The library is further fuzz-tested 24/7 by Google's [OSS-Fuzz project](https://github.com/google/oss-fuzz). It uses
the same `fuzzers` target as above and also relies on the `FUZZER_ENGINE` variable. See the used
[build script](https://github.com/google/oss-fuzz/blob/master/projects/json/build.sh) for more information. Its default
`address` sanitizer includes LeakSanitizer, so OSS-Fuzz and the CIFuzz workflow (`.github/workflows/cifuzz.yml`) report
memory leaks, too.
In case the build at OSS-Fuzz fails, an issue will be created automatically.
### Handling OSS-Fuzz reports
OSS-Fuzz files the crashes it finds in its own [issue tracker](https://issues.oss-fuzz.com), not on GitHub. So that
each report can be traced to the change that fixed it, and each fix to the report it answers, fixes follow these
conventions:
- **Reference the OSS-Fuzz issue in the pull request**, next to any GitHub issue it closes, as `OSS-Fuzz: <id>` (for
example, `OSS-Fuzz: 563659413`), and in the commit message. The ID alone does not disclose the crash. If the report
was triaged into a GitHub issue, link the OSS-Fuzz issue there too.
- **Turn the reproducer into a unit test.** Download the testcase from the OSS-Fuzz report, reduce it if possible, and
add it as a regression test to the unit test of the affected format (e.g., `tests/src/unit-bjdata.cpp`), with a
comment naming the OSS-Fuzz issue. This way the input is checked by every CI run rather than only by OSS-Fuzz, and
it stays covered even if OSS-Fuzz later closes the report as not reproducible.
- **Keep the fuzzer drivers and the unit tests in sync.** The round-trip checks of the BJData, BON8, BSON, CBOR,
MessagePack and UBJSON drivers are also run on a fixed corpus in the unit tests (see
`tests/src/round_trip_corpus.hpp` and the "round-trip invariants" test cases), so a regression shows up in CI
first. When a driver's checks change, change the unit tests with them.
- **Record in the report whether the bug shipped.** OSS-Fuzz asks whether a crash was a short-lived regression or
affects a released version; answer it when the fix is merged, as it decides whether the fix needs a release note or
a security advisory (see the [security policy](../.github/SECURITY.md)).
After the fix is merged, OSS-Fuzz re-runs the reproducer on its next build and marks the report as verified and
closed. If it does not, the fix is incomplete.