diff --git a/.github/CONTRIBUTING.md b/.github/CONTRIBUTING.md index 6f3d0bef3..9c3bb2c04 100644 --- a/.github/CONTRIBUTING.md +++ b/.github/CONTRIBUTING.md @@ -205,6 +205,9 @@ API of the 3.x.y version is broken. This includes: - Changing access specifiers. - Changing default arguments. +What is and is not covered by this guarantee is described in the +[roadmap](https://json.nlohmann.me/community/roadmap/#api-stability). + Although these guidelines may seem restrictive, they are essential for maintaining the library’s utility. Breaking changes may be introduced when they are guarded with a feature macro such as diff --git a/README.md b/README.md index 354d7a871..967ba7f6c 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,7 @@ [![Coverage Status](https://coveralls.io/repos/github/nlohmann/json/badge.svg?branch=develop)](https://coveralls.io/github/nlohmann/json?branch=develop) [![Coverity Scan Build Status](https://scan.coverity.com/projects/5550/badge.svg)](https://scan.coverity.com/projects/nlohmann-json) [![Codacy Badge](https://app.codacy.com/project/badge/Grade/e0d1a9d5d6fd46fcb655c4cb930bb3e8)](https://app.codacy.com/gh/nlohmann/json/dashboard?utm_source=gh&utm_medium=referral&utm_content=&utm_campaign=Badge_grade) -[![Fuzzing Status](https://oss-fuzz-build-logs.storage.googleapis.com/badges/json.svg)](https://bugs.chromium.org/p/oss-fuzz/issues/list?sort=-opened&can=1&q=proj:json) +[![Fuzzing Status](https://oss-fuzz-build-logs.storage.googleapis.com/badges/json.svg)](https://issues.oss-fuzz.com/issues?q=project:json) [![Try online](https://img.shields.io/badge/try-online-blue.svg)](https://wandbox.org/permlink/1mp10JbaANo6FUc7) [![Documentation](https://img.shields.io/badge/docs-mkdocs-blue.svg)](https://json.nlohmann.me) [![GitHub license](https://img.shields.io/badge/license-MIT-blue.svg)](https://raw.githubusercontent.com/nlohmann/json/develop/LICENSE.MIT) diff --git a/docs/mkdocs/docs/community/roadmap.md b/docs/mkdocs/docs/community/roadmap.md index 36d5d7e7f..27afeb4c5 100644 --- a/docs/mkdocs/docs/community/roadmap.md +++ b/docs/mkdocs/docs/community/roadmap.md @@ -25,9 +25,7 @@ work items are tracked in the [GitHub milestones](https://github.com/nlohmann/js ## What the project will not do -- **Break the public API of version 3.x.** See the - [contribution guidelines](https://github.com/nlohmann/json/blob/develop/.github/CONTRIBUTING.md#break-the-public-api) - for what counts as a breaking change. +- **Break the public API of version 3.x.** See [API stability](#api-stability) for what this covers. - **Require a newer C++ standard than C++11.** - **Break JSON conformance** or enable non-standard extensions by default. - **Add dependencies** or require a build step. The library remains header-only, and the single header @@ -35,6 +33,32 @@ work items are tracked in the [GitHub milestones](https://github.com/nlohmann/js - **Trade simplicity for speed or memory efficiency.** Performance improvements are welcome, but the library is not meant to compete with the fastest JSON libraries, see [Design goals](../home/design_goals.md). +## API stability + +Releases follow [semantic versioning](https://semver.org): a minor or patch release of version 3.x does not break code +that uses the public API. In particular, a 3.x release does not: + +- change the signature of a function (its parameter types, return type, number of parameters, or the const-ness of a + member function); +- remove or rename a function or class; +- change which exceptions a function throws, or the [exception ids](../home/exceptions.md); +- change access specifiers or default arguments. + +Exceptions to these rules, for instance when fixing a bug requires changing the exception a function throws, are +documented in the [release notes](../home/releases.md). + +The following are **not** part of the public API and may change in any release, including patch releases: + +- The text of exception messages returned by `what()`. Use the [exception id](../home/exceptions.md) to tell errors + apart. +- The ABI, including `sizeof(basic_json)` and the memory layout of its values. Recompile your code when you upgrade the + library. The [versioned inline namespace](../features/namespace.md) turns mixing versions into a link error. +- Everything in namespace `nlohmann::detail`, and macros and type traits that are not documented in the + [API reference](../api/basic_json/index.md). + +Changes that would break the public API are only added behind a macro whose default keeps the 3.x behavior, see +[Version 4.0](#version-40). + ## Version 4.0 There is no release date for version 4.0 yet. Proposals that need a major version, for instance stricter type diff --git a/tests/fuzzing.md b/tests/fuzzing.md index b7c2da27f..329cdbf58 100644 --- a/tests/fuzzing.md +++ b/tests/fuzzing.md @@ -3,6 +3,14 @@ Each parser of the library (JSON, BJData, BON8, BSON, CBOR, MessagePack, and UBJSON) can be fuzz tested. Currently, [libFuzzer](https://llvm.org/docs/LibFuzzer.html) and [afl++](https://github.com/AFLplusplus/AFLplusplus) are supported. +## What the fuzzers check + +Each fuzzer driver (`tests/src/fuzzer-parse_*.cpp`) parses its input twice: once with `allow_exceptions = false` and +once with exceptions. Both calls must agree. Where parsing with exceptions fails, the call without exceptions must +return a discarded value (or throw the same kind of non-parse error), and it must never throw a `parse_error`. Where +parsing succeeds, both calls must return the same value. The drivers then serialize the value, parse the result back, +and check that nothing was lost. The drivers check all of this with `assert`, so they refuse to build with `NDEBUG`. + ## Corpus creation For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a @@ -54,6 +62,9 @@ Then pass the corpus directory as command-line argument (assuming it is located The fuzzer should be able to run indefinitely without crashing. In case of a crash, the tested input is dumped into a file starting with `crash-`. +To also detect memory leaks, build with AddressSanitizer (`FUZZER_ENGINE="-fsanitize=fuzzer,address"`): libFuzzer then +runs LeakSanitizer by default (`-detect_leaks=1`). LeakSanitizer is not available with Apple Clang on macOS. + ## afl++ To use afl++, you need to pass `-fsanitize=fuzzer` as `FUZZER_ENGINE`. It will be replaced by a `libAFLDriver.a` to @@ -76,7 +87,9 @@ directory `out`. The library is further fuzz-tested 24/7 by Google's [OSS-Fuzz project](https://github.com/google/oss-fuzz). It uses the same `fuzzers` target as above and also relies on the `FUZZER_ENGINE` variable. See the used -[build script](https://github.com/google/oss-fuzz/blob/master/projects/json/build.sh) for more information. +[build script](https://github.com/google/oss-fuzz/blob/master/projects/json/build.sh) for more information. Its default +`address` sanitizer includes LeakSanitizer, so OSS-Fuzz and the CIFuzz workflow (`.github/workflows/cifuzz.yml`) report +memory leaks, too. In case the build at OSS-Fuzz fails, an issue will be created automatically. diff --git a/tests/src/fuzzer-parse_bjdata.cpp b/tests/src/fuzzer-parse_bjdata.cpp index 5db76caae..470765583 100644 --- a/tests/src/fuzzer-parse_bjdata.cpp +++ b/tests/src/fuzzer-parse_bjdata.cpp @@ -10,7 +10,9 @@ This file implements a parser test suitable for fuzz testing. Given a byte array data, it performs the following steps: +- j0 = from_bjdata(data, allow_exceptions = false) - j1 = from_bjdata(data) +- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise) - vec2 = to_bjdata(j1, use_size = false, use_type = false) - vec3 = to_bjdata(j1, use_size = true, use_type = false) - vec4 = to_bjdata(j1, use_size = true, use_type = true) @@ -65,6 +67,13 @@ drivers. using json = nlohmann::json; +// compares dumps rather than values, because NaN != NaN; keep writes strings +// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw +static bool same_value(const json& lhs, const json& rhs) +{ + return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep); +} + // value-stable comparison for the round-trip checks below; see the note // above on why this compares dump()s rather than the json values directly static bool is_value_stable(const json& lhs, const json& rhs) @@ -75,14 +84,42 @@ static bool is_value_stable(const json& lhs, const json& rhs) // see http://llvm.org/docs/LibFuzzer.html extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) { - // step 0: recover from all errors, reading from memory and from a stream + // recover from all errors, reading from memory and from a stream const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bjdata).errors == 0; + std::vector const vec1(data, data + size); + + // step 0: parse input without exceptions; a parse error must then be + // reported as a discarded value, never thrown + json j_noexcept; + bool noexcept_threw = false; + try + { + j_noexcept = json::from_bjdata(vec1, true, false); + } + catch (const json::parse_error&) + { + assert(false); + } + catch (const json::exception&) + { + // type and out-of-range errors are not parse errors and still throw + noexcept_threw = true; + } + // whether step 1 succeeded; if not, the catch blocks below check that + // step 0 failed, too + bool parsed = false; + try { // step 1: parse input - std::vector const vec1(data, data + size); json const j1 = json::from_bjdata(vec1); + parsed = true; + + // without exceptions, the same input must give the same value + assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1)); + + // the recovering parser must not have reported an error either assert(recovered_without_errors); try @@ -117,16 +154,19 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) catch (const json::parse_error&) { // parse errors are ok, because input may be random bytes - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } catch (const json::type_error&) { // type errors can occur during parsing, too + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } catch (const json::out_of_range&) { // out of range errors may happen if provided sizes are excessive - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } // return 0 - non-zero return values are reserved for future use diff --git a/tests/src/fuzzer-parse_bon8.cpp b/tests/src/fuzzer-parse_bon8.cpp index 210ed2519..a4c6700b0 100644 --- a/tests/src/fuzzer-parse_bon8.cpp +++ b/tests/src/fuzzer-parse_bon8.cpp @@ -10,7 +10,9 @@ This file implements a parser test suitable for fuzz testing. Given a byte array data, it performs the following steps: +- j0 = from_bon8(data, allow_exceptions = false) - j1 = from_bon8(data) +- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise) - vec = to_bon8(j1) - j2 = from_bon8(vec) - assert(to_bon8(j2) == vec) @@ -40,6 +42,13 @@ drivers. using json = nlohmann::json; +// compares dumps rather than values, because NaN != NaN; keep writes strings +// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw +static bool same_value(const json& lhs, const json& rhs) +{ + return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep); +} + namespace { // the serialization of the value read from @a input, or the error message @@ -61,7 +70,7 @@ std::string read_bon8(InputType&& input) // see http://llvm.org/docs/LibFuzzer.html extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) { - // step 0: recover from all errors, reading from memory and from a stream + // recover from all errors, reading from memory and from a stream const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bon8).errors == 0; // contiguous and stream input must be read alike @@ -70,11 +79,39 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) assert(read_bon8(std::vector(data, data + size)) == read_bon8(stream)); } + std::vector const vec1(data, data + size); + + // step 0: parse input without exceptions; a parse error must then be + // reported as a discarded value, never thrown + json j_noexcept; + bool noexcept_threw = false; + try + { + j_noexcept = json::from_bon8(vec1, true, false); + } + catch (const json::parse_error&) + { + assert(false); + } + catch (const json::exception&) + { + // type and out-of-range errors are not parse errors and still throw + noexcept_threw = true; + } + // whether step 1 succeeded; if not, the catch blocks below check that + // step 0 failed, too + bool parsed = false; + try { // step 1: parse input - std::vector const vec1(data, data + size); json const j1 = json::from_bon8(vec1); + parsed = true; + + // without exceptions, the same input must give the same value + assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1)); + + // the recovering parser must not have reported an error either assert(recovered_without_errors); try @@ -97,16 +134,19 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) catch (const json::parse_error&) { // parse errors are ok, because input may be random bytes - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } catch (const json::type_error&) { // type errors can occur during parsing, too + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } catch (const json::out_of_range&) { // out of range errors may happen if provided sizes are excessive - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } // return 0 - non-zero return values are reserved for future use diff --git a/tests/src/fuzzer-parse_bson.cpp b/tests/src/fuzzer-parse_bson.cpp index e35b9ce7d..e0415160a 100644 --- a/tests/src/fuzzer-parse_bson.cpp +++ b/tests/src/fuzzer-parse_bson.cpp @@ -10,7 +10,9 @@ This file implements a parser test suitable for fuzz testing. Given a byte array data, it performs the following steps: +- j0 = from_bson(data, allow_exceptions = false) - j1 = from_bson(data) +- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise) - vec = to_bson(j1) - j2 = from_bson(vec) - assert(to_bson(j2) == vec) @@ -35,17 +37,52 @@ drivers. using json = nlohmann::json; +// compares dumps rather than values, because NaN != NaN; keep writes strings +// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw +static bool same_value(const json& lhs, const json& rhs) +{ + return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep); +} + // see http://llvm.org/docs/LibFuzzer.html extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) { - // step 0: recover from all errors, reading from memory and from a stream + // recover from all errors, reading from memory and from a stream const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bson).errors == 0; + std::vector const vec1(data, data + size); + + // step 0: parse input without exceptions; a parse error must then be + // reported as a discarded value, never thrown + json j_noexcept; + bool noexcept_threw = false; + try + { + j_noexcept = json::from_bson(vec1, true, false); + } + catch (const json::parse_error&) + { + assert(false); + } + catch (const json::exception&) + { + // type and out-of-range errors are not parse errors and still throw + noexcept_threw = true; + } + // whether step 1 succeeded; if not, the catch blocks below check that + // step 0 failed, too + bool parsed = false; + try { // step 1: parse input - std::vector const vec1(data, data + size); json const j1 = json::from_bson(vec1); + parsed = true; + + // without exceptions, the same input must give the same value + assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1)); + + // the recovering parser must not have reported an error either assert(recovered_without_errors); try @@ -68,16 +105,19 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) catch (const json::parse_error&) { // parse errors are ok, because input may be random bytes - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } catch (const json::type_error&) { // type errors can occur during parsing, too + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } catch (const json::out_of_range&) { // out of range errors can occur during parsing, too - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } // return 0 - non-zero return values are reserved for future use diff --git a/tests/src/fuzzer-parse_cbor.cpp b/tests/src/fuzzer-parse_cbor.cpp index fe5344f32..7207e17de 100644 --- a/tests/src/fuzzer-parse_cbor.cpp +++ b/tests/src/fuzzer-parse_cbor.cpp @@ -10,7 +10,9 @@ This file implements a parser test suitable for fuzz testing. Given a byte array data, it performs the following steps: +- j0 = from_cbor(data, allow_exceptions = false) - j1 = from_cbor(data) +- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise) - vec = to_cbor(j1) - j2 = from_cbor(vec) - assert(to_cbor(j2) == vec) @@ -35,17 +37,52 @@ drivers. using json = nlohmann::json; +// compares dumps rather than values, because NaN != NaN; keep writes strings +// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw +static bool same_value(const json& lhs, const json& rhs) +{ + return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep); +} + // see http://llvm.org/docs/LibFuzzer.html extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) { - // step 0: recover from all errors, reading from memory and from a stream + // recover from all errors, reading from memory and from a stream const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::cbor).errors == 0; + std::vector const vec1(data, data + size); + + // step 0: parse input without exceptions; a parse error must then be + // reported as a discarded value, never thrown + json j_noexcept; + bool noexcept_threw = false; + try + { + j_noexcept = json::from_cbor(vec1, true, false); + } + catch (const json::parse_error&) + { + assert(false); + } + catch (const json::exception&) + { + // type and out-of-range errors are not parse errors and still throw + noexcept_threw = true; + } + // whether step 1 succeeded; if not, the catch blocks below check that + // step 0 failed, too + bool parsed = false; + try { // step 1: parse input - std::vector const vec1(data, data + size); json const j1 = json::from_cbor(vec1); + parsed = true; + + // without exceptions, the same input must give the same value + assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1)); + + // the recovering parser must not have reported an error either assert(recovered_without_errors); try @@ -68,16 +105,19 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) catch (const json::parse_error&) { // parse errors are ok, because input may be random bytes - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } catch (const json::type_error&) { // type errors can occur during parsing, too + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } catch (const json::out_of_range&) { // out of range errors can occur during parsing, too - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } // return 0 - non-zero return values are reserved for future use diff --git a/tests/src/fuzzer-parse_json.cpp b/tests/src/fuzzer-parse_json.cpp index c084ba015..1a5f9044b 100644 --- a/tests/src/fuzzer-parse_json.cpp +++ b/tests/src/fuzzer-parse_json.cpp @@ -10,7 +10,9 @@ This file implements a parser test suitable for fuzz testing. Given a byte array data, it performs the following steps: +- j0 = parse(data, allow_exceptions = false) - j1 = parse(data) +- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise) - s1 = serialize(j1) - j2 = parse(s1) - s2 = serialize(j2) @@ -36,20 +38,52 @@ drivers. using json = nlohmann::json; +// compares dumps rather than values, because NaN != NaN; keep writes strings +// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw +static bool same_value(const json& lhs, const json& rhs) +{ + return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep); +} + // see http://llvm.org/docs/LibFuzzer.html extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) { - // step 0: recover from all errors, reading from memory and from a stream + // recover from all errors, reading from memory and from a stream { const auto checker = check_recovering_parse(data, size, json::input_format_t::json); assert(checker.events <= (4 * size) + 4); assert((checker.errors == 0) == json::accept(data, data + size)); } + // step 0: parse input without exceptions; a parse error must then be + // reported as a discarded value, never thrown + json j_noexcept; + bool noexcept_threw = false; + try + { + j_noexcept = json::parse(data, data + size, nullptr, false); + } + catch (const json::parse_error&) + { + assert(false); + } + catch (const json::exception&) + { + // type and out-of-range errors are not parse errors and still throw + noexcept_threw = true; + } + // whether step 1 succeeded; if not, the catch blocks below check that + // step 0 failed, too + bool parsed = false; + try { // step 1: parse input json const j1 = json::parse(data, data + size); + parsed = true; + + // without exceptions, the same input must give the same value + assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1)); try { @@ -76,10 +110,12 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) catch (const json::parse_error&) { // parse errors are ok, because input may be random bytes + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } catch (const json::out_of_range&) { // out of range errors may happen if provided sizes are excessive + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } // return 0 - non-zero return values are reserved for future use diff --git a/tests/src/fuzzer-parse_msgpack.cpp b/tests/src/fuzzer-parse_msgpack.cpp index 4283da263..26b45ca8c 100644 --- a/tests/src/fuzzer-parse_msgpack.cpp +++ b/tests/src/fuzzer-parse_msgpack.cpp @@ -10,7 +10,9 @@ This file implements a parser test suitable for fuzz testing. Given a byte array data, it performs the following steps: +- j0 = from_msgpack(data, allow_exceptions = false) - j1 = from_msgpack(data) +- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise) - vec = to_msgpack(j1) - j2 = from_msgpack(vec) - assert(to_msgpack(j2) == vec) @@ -35,17 +37,52 @@ drivers. using json = nlohmann::json; +// compares dumps rather than values, because NaN != NaN; keep writes strings +// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw +static bool same_value(const json& lhs, const json& rhs) +{ + return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep); +} + // see http://llvm.org/docs/LibFuzzer.html extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) { - // step 0: recover from all errors, reading from memory and from a stream + // recover from all errors, reading from memory and from a stream const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::msgpack).errors == 0; + std::vector const vec1(data, data + size); + + // step 0: parse input without exceptions; a parse error must then be + // reported as a discarded value, never thrown + json j_noexcept; + bool noexcept_threw = false; + try + { + j_noexcept = json::from_msgpack(vec1, true, false); + } + catch (const json::parse_error&) + { + assert(false); + } + catch (const json::exception&) + { + // type and out-of-range errors are not parse errors and still throw + noexcept_threw = true; + } + // whether step 1 succeeded; if not, the catch blocks below check that + // step 0 failed, too + bool parsed = false; + try { // step 1: parse input - std::vector const vec1(data, data + size); json const j1 = json::from_msgpack(vec1); + parsed = true; + + // without exceptions, the same input must give the same value + assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1)); + + // the recovering parser must not have reported an error either assert(recovered_without_errors); try @@ -68,16 +105,19 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) catch (const json::parse_error&) { // parse errors are ok, because input may be random bytes - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } catch (const json::type_error&) { // type errors can occur during parsing, too + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } catch (const json::out_of_range&) { // out of range errors may happen if provided sizes are excessive - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } // return 0 - non-zero return values are reserved for future use diff --git a/tests/src/fuzzer-parse_ubjson.cpp b/tests/src/fuzzer-parse_ubjson.cpp index 94acfb9b7..98cea4167 100644 --- a/tests/src/fuzzer-parse_ubjson.cpp +++ b/tests/src/fuzzer-parse_ubjson.cpp @@ -10,7 +10,9 @@ This file implements a parser test suitable for fuzz testing. Given a byte array data, it performs the following steps: +- j0 = from_ubjson(data, allow_exceptions = false) - j1 = from_ubjson(data) +- assert(j0 is discarded if parsing j1 fails, and j0 == j1 otherwise) - vec2 = to_ubjson(j1, use_size = false, use_type = false) - vec3 = to_ubjson(j1, use_size = true, use_type = false) - vec4 = to_ubjson(j1, use_size = true, use_type = true) @@ -44,17 +46,52 @@ drivers. using json = nlohmann::json; +// compares dumps rather than values, because NaN != NaN; keep writes strings +// byte for byte, so ill-formed UTF-8 that a binary reader accepts cannot throw +static bool same_value(const json& lhs, const json& rhs) +{ + return lhs.dump(-1, ' ', false, json::error_handler_t::keep) == rhs.dump(-1, ' ', false, json::error_handler_t::keep); +} + // see http://llvm.org/docs/LibFuzzer.html extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) { - // step 0: recover from all errors, reading from memory and from a stream + // recover from all errors, reading from memory and from a stream const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::ubjson).errors == 0; + std::vector const vec1(data, data + size); + + // step 0: parse input without exceptions; a parse error must then be + // reported as a discarded value, never thrown + json j_noexcept; + bool noexcept_threw = false; + try + { + j_noexcept = json::from_ubjson(vec1, true, false); + } + catch (const json::parse_error&) + { + assert(false); + } + catch (const json::exception&) + { + // type and out-of-range errors are not parse errors and still throw + noexcept_threw = true; + } + // whether step 1 succeeded; if not, the catch blocks below check that + // step 0 failed, too + bool parsed = false; + try { // step 1: parse input - std::vector const vec1(data, data + size); json const j1 = json::from_ubjson(vec1); + parsed = true; + + // without exceptions, the same input must give the same value + assert(!noexcept_threw && !j_noexcept.is_discarded() && same_value(j_noexcept, j1)); + + // the recovering parser must not have reported an error either assert(recovered_without_errors); try @@ -87,16 +124,19 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) catch (const json::parse_error&) { // parse errors are ok, because input may be random bytes - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } catch (const json::type_error&) { // type errors can occur during parsing, too + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); } catch (const json::out_of_range&) { // out of range errors may happen if provided sizes are excessive - assert(!recovered_without_errors); + assert(parsed || noexcept_threw || j_noexcept.is_discarded()); + assert(parsed || !recovered_without_errors); } // return 0 - non-zero return values are reserved for future use diff --git a/tests/src/unit-bjdata.cpp b/tests/src/unit-bjdata.cpp index af1883827..3927ab2e3 100644 --- a/tests/src/unit-bjdata.cpp +++ b/tests/src/unit-bjdata.cpp @@ -4603,3 +4603,49 @@ TEST_CASE("issue #5648 - from_bjdata(ptr, len) must read len bytes, not treat pt CHECK(json::from_bjdata(packed.data(), packed.size(), false) == j); #endif } + +TEST_CASE("BJData large strings and binaries (chunked reader)") +{ + // Strings share get_ubjson_string() -> get_string() -> get_bytes() with + // plain UBJSON. Binary values are different: only a Draft 3 optimized + // array (type marker 'B') is read back as a binary value, through + // get_binary() -> get_bytes() (see parse_ubjson_internal()'s "If BJData + // type marker is 'B'" branch); Draft 2 (the default) writes a binary + // value as a plain array of uint8_t numbers instead (see the "round trip + // of a binary value is value-stable, not byte-stable" test above), which + // never reaches get_bytes(). Both reads happen in bounded chunks + // (binary_reader.hpp, chunk_size == 4096); check lengths around and + // beyond that size, for both vector (iterator) and pointer inputs. + for (const std::size_t len : + { + std::size_t{0}, std::size_t{1}, std::size_t{4095}, std::size_t{4096}, + std::size_t{4097}, std::size_t{8192}, std::size_t{100000} + }) + { + CAPTURE(len) + + // string + const json j_string = std::string(len, 'x'); + const std::vector v_string = json::to_bjdata(j_string); + CHECK(json::from_bjdata(v_string) == j_string); + // pointer input exercises the std::memcpy fast path + CHECK(json::from_bjdata(reinterpret_cast(v_string.data()), + reinterpret_cast(v_string.data()) + v_string.size()) == j_string); + + // binary, forced into the Draft 3 optimized ('B' marker) encoding + const json j_binary = json::binary(std::vector(len, 0xCD)); + const std::vector v_binary = json::to_bjdata(j_binary, true, true, json::bjdata_version_t::draft3); + CHECK(json::from_bjdata(v_binary) == j_binary); + CHECK(json::from_bjdata(reinterpret_cast(v_binary.data()), + reinterpret_cast(v_binary.data()) + v_binary.size()) == j_binary); + + // a truncated payload must still be reported as an error + if (len > 16) + { + std::vector truncated = v_string; + truncated.resize(truncated.size() - 8); + json _; + CHECK_THROWS_AS(_ = json::from_bjdata(truncated), json::parse_error); + } + } +} diff --git a/tests/src/unit-bson.cpp b/tests/src/unit-bson.cpp index 352ad6f1a..ea5ca3130 100644 --- a/tests/src/unit-bson.cpp +++ b/tests/src/unit-bson.cpp @@ -1993,3 +1993,47 @@ TEST_CASE("Invalid document size handling") CHECK(json::from_bson(v, true, false).is_discarded()); } } + +TEST_CASE("BSON large strings and binaries (chunked reader)") +{ + // get_bson_string()/get_bson_binary() both read through get_string()/ + // get_binary(), which read in bounded chunks (binary_reader.hpp, + // chunk_size == 4096); make sure roundtripping is correct for lengths + // around and beyond that chunk size, for both vector (iterator) and + // pointer inputs. BSON only accepts an object at the top level, so the + // string/binary value is wrapped in one. + for (const std::size_t len : + { + std::size_t{0}, std::size_t{1}, std::size_t{4095}, std::size_t{4096}, + std::size_t{4097}, std::size_t{8192}, std::size_t{100000} + }) + { + CAPTURE(len) + + // string + const json j_string = {{"k", std::string(len, 'x')}}; + const std::vector v_string = json::to_bson(j_string); + CHECK(json::from_bson(v_string) == j_string); + // pointer input exercises the std::memcpy fast path + CHECK(json::from_bson(reinterpret_cast(v_string.data()), + reinterpret_cast(v_string.data()) + v_string.size()) == j_string); + + // binary (BSON binary values always carry a subtype, so give one + // explicitly; otherwise from_bson() would round-trip to subtype 0 + // rather than back to the original "no subtype" value) + const json j_binary = {{"k", json::binary(std::vector(len, 0xCD), std::uint8_t{0})}}; + const std::vector v_binary = json::to_bson(j_binary); + CHECK(json::from_bson(v_binary) == j_binary); + CHECK(json::from_bson(reinterpret_cast(v_binary.data()), + reinterpret_cast(v_binary.data()) + v_binary.size()) == j_binary); + + // a truncated payload must still be reported as an error + if (len > 16) + { + std::vector truncated = v_string; + truncated.resize(truncated.size() - 8); + json _; + CHECK_THROWS_AS(_ = json::from_bson(truncated), json::parse_error); + } + } +} diff --git a/tests/src/unit-msgpack.cpp b/tests/src/unit-msgpack.cpp index 47cccbf13..b705f94b7 100644 --- a/tests/src/unit-msgpack.cpp +++ b/tests/src/unit-msgpack.cpp @@ -2482,3 +2482,43 @@ TEST_CASE("MessagePack numbers use the active union member (see #5644)") CHECK(json::from_msgpack(result) == j); } } + +TEST_CASE("MessagePack large strings and binaries (chunked reader)") +{ + // get_msgpack_string()/get_msgpack_binary() both read through get_binary(), + // which reads in bounded chunks (binary_reader.hpp, chunk_size == 4096); + // make sure roundtripping is correct for lengths around and beyond that + // chunk size, for both vector (iterator) and pointer inputs. + for (const std::size_t len : + { + std::size_t{0}, std::size_t{1}, std::size_t{4095}, std::size_t{4096}, + std::size_t{4097}, std::size_t{8192}, std::size_t{100000} + }) + { + CAPTURE(len) + + // string + const json j_string = std::string(len, 'x'); + const std::vector v_string = json::to_msgpack(j_string); + CHECK(json::from_msgpack(v_string) == j_string); + // pointer input exercises the std::memcpy fast path + CHECK(json::from_msgpack(reinterpret_cast(v_string.data()), + reinterpret_cast(v_string.data()) + v_string.size()) == j_string); + + // binary + const json j_binary = json::binary(std::vector(len, 0xCD)); + const std::vector v_binary = json::to_msgpack(j_binary); + CHECK(json::from_msgpack(v_binary) == j_binary); + CHECK(json::from_msgpack(reinterpret_cast(v_binary.data()), + reinterpret_cast(v_binary.data()) + v_binary.size()) == j_binary); + + // a truncated payload must still be reported as an error + if (len > 16) + { + std::vector truncated = v_string; + truncated.resize(truncated.size() - 8); + json _; + CHECK_THROWS_AS(_ = json::from_msgpack(truncated), json::parse_error); + } + } +} diff --git a/tests/src/unit-regression3.cpp b/tests/src/unit-regression3.cpp index d4f75a36e..712be6b1a 100644 --- a/tests/src/unit-regression3.cpp +++ b/tests/src/unit-regression3.cpp @@ -533,7 +533,7 @@ TEST_CASE("regression tests 3") } #endif -#if JSON_HAS_RANGES && !defined(__MINGW32__) +#if JSON_HAS_RANGE_VIEW_CONVERSION SECTION("issue #4916 - constructing array from C++20 ranges view does not work") { std::vector nums{1, 2, 37, 42, 21}; @@ -548,7 +548,7 @@ TEST_CASE("regression tests 3") #endif // owning_view is not available in libstdc++ < 12 -#if JSON_HAS_RANGES && !defined(__MINGW32__) && !(defined(__GLIBCXX__) && _GLIBCXX_RELEASE < 12) +#if JSON_HAS_RANGE_VIEW_CONVERSION && !(defined(__GLIBCXX__) && _GLIBCXX_RELEASE < 12) SECTION("issue #4916 - constructing array from prvalue C++20 ranges view (owning_view)") { json const j(std::vector {1, 2, 37, 42, 21} | std::views::filter([](int i) @@ -560,7 +560,7 @@ TEST_CASE("regression tests 3") } #endif -#if JSON_HAS_RANGES && !defined(__MINGW32__) +#if JSON_HAS_RANGE_VIEW_CONVERSION SECTION("issue #4916 - constructing array from C++20 transform view (prvalue elements)") { std::vector nums{1, 2, 3}; diff --git a/tests/src/unit-serialization.cpp b/tests/src/unit-serialization.cpp index 4b7a53b10..9d7c41b33 100644 --- a/tests/src/unit-serialization.cpp +++ b/tests/src/unit-serialization.cpp @@ -809,3 +809,178 @@ TEST_CASE("serializer buffers are flushed mid-string and mid-binary") CHECK(j.dump(2) == "{\n \"bytes\": [" + expected_pretty_bytes + "],\n \"subtype\": null\n}"); } } + +TEST_CASE("serialization boundary values for the write buffer") +{ + // write_buffer is a std::array (write_buffer_size). put_string() + // guards it with two checks, and each must be exercised exactly on and one + // past its own boundary: a heap overflow in a different manual buffer path + // (the dump(1100) indent buffer) once survived 100% line coverage because + // every test that touched it only ever grew the buffer by a single step, + // never landing on the exact edge of the comparison that protects it. + // + // - straight-through: put_string() bypasses write_buffer entirely and + // writes directly to the output adapter once `length >= write_buffer.size()`. + // - flush-then-copy: otherwise, if `write_buffer_pos + length > write_buffer.size()`, + // put_string() flushes what is pending and then memcpy's the new run into + // the freshly emptied buffer. + + SECTION("top-level string exercises the straight-through guard (length >= 1024)") + { + // dump() of a bare string writes the opening quote with put_char() + // (write_buffer_pos: 0 -> 1), then the body with put_string(). With + // write_buffer_pos == 1, `1 + length > 1024` and `length >= 1024` flip + // together at length 1024, so 1023/1024/1025 cover "just under", + // "exactly at" and "just over" the guard in one move: 1023 is copied + // into the buffer (filling it exactly), 1024 and 1025 bypass it. + for (const std::size_t len : + { + std::size_t{1023}, std::size_t{1024}, std::size_t{1025} + }) + { + CAPTURE(len) + const std::string body(len, 'a'); + const json j = body; + const std::string expected = '"' + body + '"'; + + CHECK(j.dump() == expected); + + std::ostringstream o; + o << j; + CHECK(o.str() == expected); + } + } + + SECTION("string nested in an array exercises the flush-then-copy guard") + { + // json::array({body}) writes '[' then '"' before the body, so + // write_buffer_pos == 2 when put_string() is entered for it. The + // body's last byte then lands at logical offset 2 + len: len == 1022 + // lands exactly on offset 1024 (2 + 1022 == write_buffer.size(), so the + // strict "> " guard does not fire and the body fits snugly), while + // len == 1023 lands one past it at offset 1025 (2 + 1023 > 1024), + // which must flush what's pending before copying the body in. + for (const std::size_t len : + { + std::size_t{1022}, std::size_t{1023} + }) + { + CAPTURE(len) + const std::string body(len, 'a'); + const json j = json::array({body}); + const std::string expected = "[\"" + body + "\"]"; + + CHECK(j.dump() == expected); + + std::ostringstream o; + o << j; + CHECK(o.str() == expected); + + CHECK(json::parse(j.dump()) == j); + } + } +} + +TEST_CASE("serialization boundary values for the string buffer") +{ + // string_buffer is a std::array. dump_escaped_impl() flushes it + // mid-string once fewer than 13 bytes remain (`string_buffer.size() - bytes + // < 13`), 13 being one more than the most a single code point can ever + // write at once (a surrogate pair: two back-to-back "\uXXXX" escapes, 12 + // bytes). Every write into string_buffer that this check protects happens + // in steps of 2 (a simple "\\x" escape) or 6 (one "\uXXXX" unit), so + // `bytes` only ever takes even values at the point the check runs - the + // tightest values actually reachable are therefore 498 (512 - 498 == 14, + // one simple escape away from the threshold) and 500 (512 - 500 == 12, + // where the flush fires immediately and resets bytes to 0). + + SECTION("a run of 2-byte escapes lands bytes on, and one step past, the flush threshold") + { + for (const int count : + { + 249, 250, 251 + }) + { + CAPTURE(count) + const json j = std::string(static_cast(count), '\n'); + std::string expected = "\""; + for (int i = 0; i < count; ++i) + { + expected += "\\n"; + } + expected += '"'; + CHECK(j.dump() == expected); + } + } + + SECTION("an ASCII prefix leaves the tightest reachable margin before a 12-byte surrogate pair") + { + // U+1F600 (the "\xF0\x9F\x98\x80" UTF-8 bytes) is dumped under + // ensure_ascii as the 12-byte surrogate pair "\ud83d\ude00"; that + // write happens in a single step with no intermediate flush check, so + // it is the write most exposed by an off-by-one in the "< 13" guard. + // A prefix of 249 newlines leaves exactly 14 bytes of headroom + // (512 - 498), the smallest margin the guard ever actually allows + // into a new code point; 250 newlines instead trigger the guard's own + // flush first, so the emoji starts from a freshly emptied (512-byte) + // buffer, and 251 repeats that with one more escape already past the + // reset. Together they cover the margin the guard allows landing on, + // one step before, and one step after - all must still produce the + // identical, correct escapes. + for (const int prefix_count : + { + 249, 250, 251 + }) + { + CAPTURE(prefix_count) + const std::string prefix(static_cast(prefix_count), '\n'); + const std::string emoji = "\xF0\x9F\x98\x80"; + const json j = prefix + emoji; + + std::string expected_prefix; + for (int i = 0; i < prefix_count; ++i) + { + expected_prefix += "\\n"; + } + + // newline escaping does not depend on ensure_ascii: only the + // emoji differs (raw UTF-8 bytes vs. a \u-escaped surrogate pair) + CHECK(j.dump(-1, ' ', false) == '"' + expected_prefix + emoji + '"'); + CHECK(j.dump(-1, ' ', true) == '"' + expected_prefix + "\\ud83d\\ude00\""); + CHECK(json::parse(j.dump(-1, ' ', true)) == j); + CHECK(json::parse(j.dump(-1, ' ', false)) == j); + } + } + + SECTION("SWAR bulk-copy stride: k plain bytes followed by a byte handled individually") + { + // string_bulk_run()/find_ascii_copyable_run() (string_scan.hpp) scan 8 + // bytes at a time and fall back to a byte-at-a-time tail scan for + // what is left over. k from 0 to 17 spans zero, one and two full + // 8-byte strides plus a 1-byte tail, so every possible stopping point + // within and right after the SIMD stride is covered. + for (std::size_t k = 0; k <= 17; ++k) + { + CAPTURE(k) + const std::string prefix(k, 'a'); + + // (a) the run is stopped by a quote that must itself be escaped + { + const json j = prefix + "\""; + CHECK(j.dump() == '"' + prefix + "\\\"" + '"'); + } + + // (b) the run is stopped by a control character + { + const json j = prefix + "\x01"; + CHECK(j.dump() == '"' + prefix + "\\u0001" + '"'); + } + + // (c) the run is stopped by a non-ASCII byte under ensure_ascii + { + const json j = prefix + "\xC3\xA9"; // prefix + 'é' + CHECK(j.dump(-1, ' ', true) == '"' + prefix + "\\u00e9" + '"'); + } + } + } +} diff --git a/tests/src/unit-ubjson.cpp b/tests/src/unit-ubjson.cpp index 62abd50fc..4cce74dff 100644 --- a/tests/src/unit-ubjson.cpp +++ b/tests/src/unit-ubjson.cpp @@ -3308,3 +3308,43 @@ TEST_CASE("UBJSON and BJData integer markers at every range edge") } } } + +TEST_CASE("UBJSON large strings (chunked reader)") +{ + // get_ubjson_string() reads through get_string(), which reads in bounded + // chunks (binary_reader.hpp, chunk_size == 4096); make sure roundtripping + // is correct for lengths around and beyond that chunk size, for both + // vector (iterator) and pointer inputs. + // + // A binary value is not included here: plain UBJSON (unlike BJData, see + // the "BJData large strings and binaries" test) has no reader-side binary + // type, so even the optimized uint8_t-array encoding of a binary value is + // read back element-by-element as a JSON array of numbers rather than + // through get_binary() - it never reaches the chunked path this test is + // about (see the "roundtrip only works to an array of numbers" case + // above). + for (const std::size_t len : + { + std::size_t{0}, std::size_t{1}, std::size_t{4095}, std::size_t{4096}, + std::size_t{4097}, std::size_t{8192}, std::size_t{100000} + }) + { + CAPTURE(len) + + const json j_string = std::string(len, 'x'); + const std::vector v_string = json::to_ubjson(j_string); + CHECK(json::from_ubjson(v_string) == j_string); + // pointer input exercises the std::memcpy fast path + CHECK(json::from_ubjson(reinterpret_cast(v_string.data()), + reinterpret_cast(v_string.data()) + v_string.size()) == j_string); + + // a truncated payload must still be reported as an error + if (len > 16) + { + std::vector truncated = v_string; + truncated.resize(truncated.size() - 8); + json _; + CHECK_THROWS_AS(_ = json::from_ubjson(truncated), json::parse_error); + } + } +}