mirror of
https://github.com/nlohmann/json.git
synced 2026-10-07 15:07:13 +00:00
c6ee5a64bed45a3ff8641e71de079b6b6108f484
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
759dd2e1b1 |
Add json_document and json_view: node index, parser, and document
Add json_document and json_view, a read-only, zero-copy index of a JSON text, as the first public slice of the zero-copy view (#5295). A parse produces a flat array of 16-byte nodes in document order, one per value and one per object key. Strings stay in the source text; escaped strings are decoded into an arena. Integers are converted while their digits are in the cache; floats keep only their digit layout and are converted on read. Containers store the size of their subtree, so a reader can step over one in constant time. A document makes a handful of allocations, however many values it has. The parser accepts exactly what json::parse accepts, with every combination of ignore_comments and ignore_trailing_commas, with and without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is portable C++11 and does not depend on byte order. basic_json_document adds parse, parse_copy, accept, read (reuses a document's memory), root, is_discarded, source, owns_source, node_count, memory_usage, and shrink_to_fit. basic_json_view adds type, the is_* queries, operator bool, size, empty, materialize, and source_offset. A parse error throws the same exception basic_json::parse would throw for the same input, message and position included. detail::abi_config keeps JSON_STRICT_NUL_HANDLING and JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp undefines them, in the ABI namespace so they always match the basic_json in use. A NUL byte that ends a // comment is the end of the input, as in parse() since #5696. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e6f32bd28a |
Stricter fuzzer checks and boundary-value tests for buffers (#5774)
* Check in the fuzzers that parsing without exceptions agrees Each fuzzer driver now also parses its input with allow_exceptions = false. That call must never throw a parse_error, must return a discarded value where parsing with exceptions fails, and must return the same value where it succeeds. Values are compared by their dump(), because NaN is not equal to itself. A plain !is_discarded() assertion, as suggested in #3642, would never fail: the drivers parse with exceptions, so a result can never be discarded. tests/fuzzing.md describes the checks and notes that OSS-Fuzz and CIFuzz already run LeakSanitizer, because their default address sanitizer includes it. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Use JSON_HAS_RANGE_VIEW_CONVERSION in the range view regression tests #5728 combined the JSON_HAS_RANGES and MinGW conditions into JSON_HAS_RANGE_VIEW_CONVERSION, but three test guards still spelled them out. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Test the serializer's buffers at their boundaries The dump() indent overflow survived full line coverage because the tests grew its buffer by only one step. This adds tests that land exactly on, and one past, the limits of the other two serializer buffers: - write_buffer (1024 bytes): strings of 1023, 1024 and 1025 bytes at the top level, and of 1022 and 1023 bytes inside an array, so that both guards in put_string() are hit at their boundary. Each is checked for dump() and for stream output. - string_buffer (512 bytes, flushed when fewer than 13 bytes remain): runs of two-byte escapes, and a surrogate pair written with 14 bytes of room, right after a flush, and one escape later. - The 8-byte bulk scan from the serializer side: 0 to 17 plain bytes followed by a quote, a control character, or a non-ASCII character. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Test the chunked string and binary reads of all binary formats The binary readers read strings and binary values in chunks of 4096 bytes. Only CBOR tested lengths around that size. MessagePack, UBJSON, BJData and BSON now round-trip lengths 0, 1, 4095, 4096, 4097, 8192 and 100000 from vector and pointer input, and must report a truncated payload as a parse error. UBJSON reads binary values as arrays of numbers, so it is tested with strings only. BJData binary values reach the chunked read only in Draft 3. BON8 decodes strings byte by byte and does not use this path. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
b54ed188e6 |
Remove test debt: dead guards, discarded results, and unreferenced files (#5732)
* Run the README test case in JSON_FastTests jobs The "README" test case was marked doctest::skip() when the tests moved from Catch to doctest in 2019, where it replaced Catch's hidden tag. It is not slow (17 assertions, about 0.00 s), but cmake/test.cmake only passes --no-skip when JSON_FastTests is off, so the per-compiler ci_test_*_cxxNN matrix, macOS, Windows Release/ARM, icpc, icpx and nvhpc compiled the README examples without running them. Drop the skip decorator so every job runs the case. Test-only change. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove stale clang ranges guards in unit-iterators2.cpp The "algorithms" and "views" sections were guarded by clang/libstdc++ checks written for a clang 15 (04/2022) bug. The first guard's condition contradicts its own comment: it skips clang+libc++ and keeps clang+libstdc++. Both sections already sit inside `#if JSON_HAS_RANGES`, which macro_scope.hpp excludes for the toolchains these guards targeted, so the inner guards never let the sections run on the platforms they meant to protect and are redundant on the rest. Verified locally with Apple clang 21/libc++ and clang 16.0.6/libstdc++ 12 (Docker): both pass all 1355 assertions with the guards removed. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix copy-pasted CBOR half-float checks; enable stale encode checks In the RFC 8949 Appendix A test case, the decode checks for 5.960464477539063e-8 (0xf9 0x00 0x01) and 0.00006103515625 (0xf9 0x04 0x00) were copy-pasted from the neighboring -4.0 example, so those two half-float byte sequences were never actually decoded and checked, and -4.0 was checked three times instead. The two float32 encode checks for 100000.0 and 3.4028234663852886e+38 were commented out before the writer supported emitting float32 and are now verified to match byte for byte, so they are enabled. The remaining commented-out half-precision to_cbor checks are collapsed into a single explanatory comment, since the writer never emits half-precision floats. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Assert on the result of STL container conversions in tests The "object-like STL containers" and "array-like STL containers" sections converted json values into std::map, unordered_map, multimap, unordered_multimap, list, forward_list, array, valarray, vector, deque, set and unordered_set and discarded the result, so these ~60 conversions only proved that the code compiles and does not throw; a conversion that dropped or reordered elements would still pass. Bind each result and compare it against the expected container. Also fix a copy-paste slip in the deque section (`j2.get<std::deque<double>>()` instead of j3, so j3's doubles were never converted to a deque), and remove the dead `// CHECK(m5["one"] == "eins")` comments that referred to a variable that did not exist by asserting the equivalent through the bound result. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate SaxCountdown and other test helpers across formats SaxCountdown was copied byte-for-byte into six binary-format test files (unit-cbor.cpp, unit-msgpack.cpp, unit-ubjson.cpp, unit-bjdata.cpp, unit-bon8.cpp, unit-bson.cpp), about 370 redundant lines. Move it into tests/src/sax_countdown.hpp (namespace utils, alongside test_utils.hpp and round_trip_corpus.hpp) and include it from all six. trait_test_arg and the "value_in_range_of trait" TEST_CASE_TEMPLATE_DEFINE were duplicated between unit-32bit.cpp and unit-bjdata.cpp; the trait is a detail/meta trait, not specific to either file. Move it into tests/src/value_in_range_of_test.hpp; unit-32bit.cpp keeps its own include, since JSON_32bitTest=ONLY builds only that file. Each file keeps its own TEST_CASE_TEMPLATE_INVOKE list. sax_no_exception and the "issue #2824" section were duplicated in unit-regression2.cpp and unit-disabled_exceptions.cpp. Drop the copy from unit-regression2.cpp; unit-disabled_exceptions.cpp already covers the no-exceptions case that #2824 was about, and ci_test_noexceptions reruns it. No behavior change. Verified by building and running unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bon8, unit-bson, unit-32bit, unit-regression2 and unit-disabled_exceptions against include/ (clang++ -std=c++11, ASan/UBSan where applicable); assertion counts are unchanged from before the refactor. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the unreferenced vendored libFuzzer tests/thirdparty/Fuzzer (155 files, ~776 KB of vendored Apache-2.0 LLVM code from the 2016 OSS-Fuzz import) is not referenced by any CMakeLists, Makefile or workflow: the fuzz drivers link against -fsanitize=fuzzer or the repo's own tests/src/fuzzer-driver_afl.cpp. Its vendored README only points at llvm.org's own libFuzzer docs. Being dead code, it also adds noise to the flawfinder code-scanning workflow, which scans the whole tree. Remove the directory and its .reuse/dep5 entry. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove unreferenced 2016 benchmark and fuzz reports tests/reports (1.6 MB) holds AFL status pages and plots from 2016-08-29 and 2016-10-02, and a nativejson-benchmark snapshot from 2016 with links to rawgit.com, which shut down in 2019. Nothing references this directory: no doc, README section, script or workflow points at it, and it describes a ten-years-old, pre-2.0 snapshot of the library. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Run the CBOR, MessagePack, BSON and BON8 round-trip invariants in CI tests/src/round_trip_corpus.hpp exists so that the byte-stability invariant the fuzzer drivers check also runs on a fixed corpus in CI, instead of only at OSS-Fuzz. So far only the UBJSON and BJData drivers had a matching unit test; the CBOR, MessagePack, BSON and BON8 drivers assert the same invariant (assert(to_X(j2) == vec)) but nothing ran it outside OSS-Fuzz. Add "<FORMAT> round-trip invariants" test cases to unit-cbor.cpp, unit-msgpack.cpp, unit-bson.cpp and unit-bon8.cpp, modeled on the UBJSON case: seed j1 from the corpus (skipping values that do not survive the format's own round trip, as the fuzzer drivers only ever see values from_X() actually produced), then require from_X(to_X(j1)) not to throw and check to_X(j2) == to_X(j1). BSON only serializes objects, so non-object corpus values are skipped. Update the comments in round_trip_corpus.hpp and tests/fuzzing.md to name all six formats. The stream-versus-contiguous check in the BON8 driver is left out, as #5601 reworks it. A local probe confirms no violations on the current corpus (CBOR 3849 checked, MessagePack 3909, BSON 2958, BON8 3841 - matching the counts already recorded for this probe in the issue). Closes #5714 item 1. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix fuzzer driver step lists to match the checks the code performs The header comment of six of the seven binary-format fuzzer drivers listed an invariant the code does not check: CBOR, MessagePack, BSON and BON8 said "assert(j1 == j2)", but the code checks byte stability, assert(to_X(j2) == vec). UBJSON and BJData still described the old "assert(j1 == j2/j3/j4)" byte-exact check from before PR #5494 replaced it with a use_size/use_type-aware round trip (UBJSON) and a value-stability check (BJData); BJData's added paragraph already explained the new check, but the step list above it did not. Also remove a dead branch in fuzzer-parse_bson.cpp: from_bson() is called with allow_exceptions = true, so it throws instead of returning a discarded value, and the "if (j1.is_discarded()) return 0;" guard could never trigger. Drop the unused <iostream> include from all seven drivers and <sstream> from all but fuzzer-parse_bon8.cpp, which is the only one that uses std::istringstream. Overlaps #5601, which edits all seven drivers in the same hunks. Closes #5714 item 4. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Silence the CMP0169 deprecation in cmake_fetch_content, fix stale guards tests/cmake_fetch_content/project calls the single-argument FetchContent_Populate(json) after FetchContent_Declare(), which CMake 3.30 deprecated as CMP0169. Since the project declares cmake_minimum_required(VERSION 3.11...3.14), the policy stays unset, so every configure with a current CMake prints the deprecation warning. The test is kept on purpose: it is the only coverage of the FetchContent_Populate + add_subdirectory pattern for CMake 3.11-3.13 users, which the docs still describe as supported. Explicitly set CMP0169 to OLD, with a comment explaining why. Also fix two stale version guards: - tests/cmake_fetch_content/CMakeLists.txt guarded the test with VERSION_GREATER "3.11.0", which is dead now that tests/CMakeLists.txt requires CMake 3.13. - tests/cmake_fetch_content2/CMakeLists.txt guarded with VERSION_GREATER "3.14.0", which skips exactly 3.14.0, the first version with FetchContent_MakeAvailable. Change it to VERSION_GREATER_EQUAL "3.14". Verified locally: `ctest -R cmake_fetch_content` passes with CMake 4.1, and the CMP0169 deprecation warning that appeared before this change is gone. Closes #5714 item 5. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Make the CMake integration-test wrappers consistent The six tests/cmake_* integration-test wrappers had drifted: - Only cmake_import and cmake_import_minver forwarded -A "${CMAKE_GENERATOR_PLATFORM}" to the inner configure, and none forwarded -T "${CMAKE_GENERATOR_TOOLSET}". The Windows workflow configures the outer build with -A Win32 -T ClangCL, so without forwarding, the inner projects of cmake_add_subdirectory, cmake_fetch_content, cmake_fetch_content2 and cmake_target_include_directories built with the generator defaults instead of matching the outer build's platform and toolset. Forward both consistently from all six wrappers. - cmake_fetch_content and cmake_fetch_content2 passed -Dnlohmann_json_source to their inner projects, which never read it (CMake warns "manually-specified variables were not used"); the inner projects fetch their own copy of the library instead. Drop it. - tests/CMakeLists.txt set JSON_FORCED_GLOBAL_COMPILE_OPTIONS from the matching environment variable but never read the cache variable again; the lines right below it read $ENV{JSON_FORCED_GLOBAL_COMPILE_OPTIONS} directly, like the LINK_OPTIONS counterpart already does. Remove the dead set(). This changes which platform and toolset the Win32 and ClangCL CI jobs build the four newly-forwarding wrappers' inner projects with, which may surface new failures there; CI has to confirm those jobs. Verified locally with Ninja (empty -A ""/-T "" is accepted): all 12 cmake_* tests still pass, and the inner fetch_content configures no longer warn about the unused variable. Closes #5714 item 6. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Turn the #972 fifo_map regression test into a real test The #972 regression test in unit-regression1.cpp only built a my_json array from a string literal (the original crash) and had no CHECK, so the fifo_map object type it exists to demonstrate was never exercised. Meanwhile the docs recommend fifo_map for keeping object keys in insertion order (object_order.md, template_parameters.md), and nothing tested that recommendation. Extend the section: after the original array assignment, parse an object with my_json::parse() (not via the "..."_json UDL, which returns a plain nlohmann::json and would exercise the cross-basic_json conversion constructor instead of the parser's own key insertion - and, as tried locally, does not keep fifo order for this stateful comparator) and check that dump() keeps insertion order, and that it survives erase() and inserting a new key. Also narrow thirdparty/fifo_map off the include path of every other test-* target: it was a PUBLIC include directory of test_main, even though unit-regression1.cpp is its only user. Add a small fifo_map_include INTERFACE library with that include directory and attach it to test-regression1 only via json_test_set_test_options(). Verified locally (test-regression1_cpp11, default build and -fsanitize=address,undefined): the new checks pass; `git grep fifo_map tests` still only finds unit-regression1.cpp and the vendored header. Closes #5714 item 7. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Document the vendored doctest.h patch; fix stale doctest_compatibility.h comments tests/thirdparty/doctest/doctest.h is doctest 2.4.12, imported in #4771. Two weeks later, #4801 hand-edited translateActiveException() to declare "String res;" inside the translator loop instead of before it, so a translator that does not match does not leave a previous translator's result in "res" for the next iteration to see. Nothing recorded this, so re-vendoring doctest.h from upstream would silently drop the fix. Add a comment at the patched site naming the version, the PR and the reason, so a future re-vendor knows to re-apply it. Also fix two stale comments in doctest_compatibility.h: - The DOCTEST_THREAD_LOCAL comment referenced Xcode 6/7, which is no longer supported; reword it to explain why the define must stay regardless (it keeps doctest's own thread_local usage out of the way of the same Clang/MinGW crash that JSON_NO_THREAD_LOCAL works around in the library, see ci_test_no_thread_local). - The <iosfwd> include's comment justified it with tests that define "private" as "public"; no test under tests/src does that any more (removed by #2352). Reword the comment instead of dropping the include, since confirming it is safe to drop needs the full CI matrix including MSVC 2015+. Verified locally that tests/src/unit-readme.cpp still builds and passes 17/17 with these headers. Closes #5714 item 9. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop compiling unit-wstring.cpp out entirely on classic ICC tests/src/unit-wstring.cpp wrapped the whole file in #ifndef __INTEL_COMPILER, with the comment "ICPC errors out on multibyte character sequences in source files". The ci_icpc job (intel/oneapi-hpckit:2023.2.1) still exists, so that job ran none of the wstring/u16string/u32string input adapter tests, including the malformed-input checks #5704 (open) extends. Only 9 lines contained non-ASCII bytes: the three *_is_utf16()/ *_is_utf32() probe functions, and three std::wstring/u16string/ u32string literals plus their narrow-string dump() expectations. Rewrite all of them with \u/\U escapes in the wide/u16/u32 literals and \x escapes (split into separate string-literal tokens so a following byte is never read as part of the same hex escape, e.g. "\xE1\x83\x85" "a") in the narrow ones. Remove the #ifndef __INTEL_COMPILER/#endif guard along with it. The *_is_utf16()/*_is_utf32() probes compared a raw multibyte literal against an escape-based one to detect a compiler that misreads the source file's encoding; with no raw literals left to misread, the comparison is now tautological, so drop the probes and the "if" guards around each SECTION's body instead of leaving them in as dead checks. The same non-ASCII-in-source-and-in-a-narrow-comparison pattern existed once more in unit-deserialization.cpp's "Using _json with char8_t literals #4945" test: a raw emoji character in a u8R"(...)" literal, guarded by a check_utf8() that returned false for ICC (same reason) and for Windows without the active UTF-8 code page. Rewrite the literal with a \U escape and compare it against a \x-escaped expectation instead of a second raw literal, and drop check_utf8() and the now-unused <windows.h> include along with the guard. Verified locally (clang, -std=c++11 and -std=c++20, -fsanitize=address,undefined, and a plain build): test-wstring keeps 18/18 assertions and unit-deserialization keeps 466/466 (c++11) and 477/477 (c++20) assertions, matching this branch before the change exactly - no coverage was gained or lost, only the source-encoding dependency was removed. ci_icpc has to confirm classic ICC actually builds and passes test-wstring now; if it does not, that is a real finding, not a reason to restore the guard. Overlaps #5704 (open), which edits unit-wstring.cpp inside the previously-guarded region (an include near the top, checks in the invalid-string sections, and a new section at the end). Closes #5713 item 5. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
1e101ecac1 |
Add BON8 support (#2998)
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
cc472af13f |
Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests (#5569)
* Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests The strongest correctness checks for the UBJSON and BJData writers lived only in the OSS-Fuzz drivers: anything from_ubjson()/from_bjdata() returns must serialize with every option combination, parse back, and re-serialize stably. Those checks only run at OSS-Fuzz, so regressions surfaced days later as external reports - the same BJData assert pair was reported five times over three years, and #5494's harness change was followed by OSS-Fuzz 563659413 within a day. Add "UBJSON round-trip invariants" and "BJData round-trip invariants" test cases that run the drivers' checks on a fixed, deterministic corpus (tests/src/round_trip_corpus.hpp): integer and float boundaries, non-finite numbers, strings, binary values, optimized containers, deep nesting, the JData annotated-array matrix, and seeded random containers. They also check two properties the drivers do not: the first round trip preserves the value, and re-serializing reproduces the exact bytes. For BJData both exclude values containing a binary value, which is read back as an array of integers unless it was written as a Draft 3 optimized binary array; this carve-out is now documented in bjdata.md. Run against the headers before #5542, the BJData test fails, including on the shape from OSS-Fuzz 563659413. Also document how OSS-Fuzz reports are handled (reference them as "OSS-Fuzz: <id>", turn the reproducer into a unit test, keep drivers and unit tests in sync) in tests/fuzzing.md, and link it from the PR template and the quality assurance page. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Add the OSS-Fuzz reproducers for 474400817 and 474480402 as unit tests Following the convention added to tests/fuzzing.md, the reproducers of the two BJData fuzzer asserts tracked since January are now unit tests: - 474400817 (assert(false)): an empty object _ArraySize_ was written as the ND-array header length, which from_bjdata() could not read back. Fixed by #5455. - 474480402 (to_bjdata(j2, false, false) == vec2): a one-byte Draft 3 binary array is written in Draft 2 mode as a uint8 array and then re-serialized with the int8 marker. This is the documented exception to byte stability, not a library bug; OSS-Fuzz closed it after #5494 relaxed the harness to value stability. The test pins the exact bytes so the exception stays deliberate. The 563659413 reproducer is already a unit test (#5542). A comment also ties the existing UBJSON excessive-count test to the timeout OSS-Fuzz reported for that shape (testcase 6347769435193344). OSS-Fuzz: 474400817 OSS-Fuzz: 474480402 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix GCC -Weffc++ and -Wuseless-cast warnings in the round-trip corpus Initialize the atoms in the member initialization list, and drop the cast of the generator's result, which already is std::size_t on 64-bit Linux. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
d6efe672b5 |
Document fuzzer usage (#3478)
* 📝 document fuzzer usage * 📝 address review comments |