Commit Graph
12 Commits
Author SHA1 Message Date
Niels Lohmann cc472af13f Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests (#5569)
* Check the fuzzers' UBJSON/BJData round-trip invariants in the unit tests

The strongest correctness checks for the UBJSON and BJData writers lived
only in the OSS-Fuzz drivers: anything from_ubjson()/from_bjdata()
returns must serialize with every option combination, parse back, and
re-serialize stably. Those checks only run at OSS-Fuzz, so regressions
surfaced days later as external reports - the same BJData assert pair
was reported five times over three years, and #5494's harness change
was followed by OSS-Fuzz 563659413 within a day.

Add "UBJSON round-trip invariants" and "BJData round-trip invariants"
test cases that run the drivers' checks on a fixed, deterministic corpus
(tests/src/round_trip_corpus.hpp): integer and float boundaries,
non-finite numbers, strings, binary values, optimized containers, deep
nesting, the JData annotated-array matrix, and seeded random containers.
They also check two properties the drivers do not: the first round trip
preserves the value, and re-serializing reproduces the exact bytes. For
BJData both exclude values containing a binary value, which is read back
as an array of integers unless it was written as a Draft 3 optimized
binary array; this carve-out is now documented in bjdata.md. Run against
the headers before #5542, the BJData test fails, including on the shape
from OSS-Fuzz 563659413.

Also document how OSS-Fuzz reports are handled (reference them as
"OSS-Fuzz: <id>", turn the reproducer into a unit test, keep drivers and
unit tests in sync) in tests/fuzzing.md, and link it from the PR
template and the quality assurance page.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add the OSS-Fuzz reproducers for 474400817 and 474480402 as unit tests

Following the convention added to tests/fuzzing.md, the reproducers of
the two BJData fuzzer asserts tracked since January are now unit tests:

- 474400817 (assert(false)): an empty object _ArraySize_ was written as
  the ND-array header length, which from_bjdata() could not read back.
  Fixed by #5455.

- 474480402 (to_bjdata(j2, false, false) == vec2): a one-byte Draft 3
  binary array is written in Draft 2 mode as a uint8 array and then
  re-serialized with the int8 marker. This is the documented exception to
  byte stability, not a library bug; OSS-Fuzz closed it after #5494
  relaxed the harness to value stability. The test pins the exact bytes
  so the exception stays deliberate.

The 563659413 reproducer is already a unit test (#5542). A comment also
ties the existing UBJSON excessive-count test to the timeout OSS-Fuzz
reported for that shape (testcase 6347769435193344).

OSS-Fuzz: 474400817
OSS-Fuzz: 474480402

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix GCC -Weffc++ and -Wuseless-cast warnings in the round-trip corpus

Initialize the atoms in the member initialization list, and drop the cast of
the generator's result, which already is std::size_t on 64-bit Linux.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-25 08:29:02 +02:00
Niels Lohmann 918da64657 Keep BJData ndarray annotations that would not survive a round trip as objects (#5542)
write_bjdata_ndarray() encoded a JData-annotated object as a BJData
ND-array whenever its dimensions' product matched _ArrayData_.size(),
which lost information in two ways:

- _ArrayData_ was never required to be an array. null has size 0, any
  other scalar has size 1, and iterating an object visits its values, so
  e.g. {"_ArraySize_":[1],"_ArrayData_":5} was written as the array [5],
  and an object _ArrayData_ came back as an array.

- The reader only restores an annotated object from an ND-array with at
  least two non-zero dimensions that is not a 1xN row vector; an empty,
  1-D, row-vector, or zero-sized shape is read back as a plain array. The
  writer nonetheless emitted ND-array headers for these shapes, so the
  annotation was silently dropped.

OSS-Fuzz issue 563659413 hit this in parse_bjdata_fuzzer: an empty binary
_ArraySize_ is written as a plain object and read back as an empty array,
after which {"_ArrayType_":"int16","_ArraySize_":[],"_ArrayData_":null}
was encoded as the ND-array header "[$I#[]" and re-read as [], failing the
harness's value-stability check.

Such objects now fall back to a plain object encoding, which round-trips.
Genuine ND-arrays (two or more positive dimensions, not a 1xN row vector)
are encoded exactly as before. Existing fallback tests that used 1-D
shapes are moved to 2-D shapes so they keep exercising the check they
were written for, and the BJData documentation is updated.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-24 17:01:31 +02:00
Qatadaha Bin Matloob 09b6b6b5ba Fix to_bjdata() emitting unparsable output when _ArraySize_ is not an array (#5455)
* Fix to_bjdata() emitting unparsable output when _ArraySize_ is not an array

write_bjdata_ndarray() never checked that _ArraySize_ is an array. The shape
is written verbatim as the header length, so a null shape emitted 'Z' and an
object shape emitted '{' after the '#', neither of which from_bjdata()
accepts, and the round-trip guarantee in the BJData docs was broken.

Both slipped through the existing validation: for null, empty() is true so
the element count starts at 0 and the per-dimension loop never runs, and for
an object the loop walks its values, which can satisfy the non-negative
integer check. When _ArrayData_ then matched that count, the writer took the
ndarray path.

Require the shape to be an array, so anything else falls back to a plain
object encoding that round-trips, as the fallback rule in the docs already
specifies.

Signed-off-by: qatcod <79017227+qatcod@users.noreply.github.com>

* Document that _ArraySize_ must be an array in the ndarray requirements

The list at bjdata.md is the exhaustive set of conditions for the ndarray
encoding, but it only implied this one through 'every entry of'.

Signed-off-by: qatcod <79017227+qatcod@users.noreply.github.com>

---------

Signed-off-by: qatcod <79017227+qatcod@users.noreply.github.com>
2026-09-05 20:30:22 +02:00
Niels Lohmann 9a091d2b82 Do not write BJData ndarrays whose size overflows std::size_t (#5362)
* Do not write BJData ndarrays whose size overflows std::size_t

write_bjdata_ndarray() multiplied the _ArraySize_ dimensions into a
std::size_t without checking for overflow. A product that wraps around
to a value that happens to match the size of _ArrayData_ passed the
length check, and the writer emitted an ndarray header announcing an
element count that cannot be represented:

    {"_ArrayType_":"uint8","_ArraySize_":[9223372036854775808,2],"_ArrayData_":[]}

was encoded as 5b 24 55 23 5b 4d 00 00 00 00 00 00 00 80 69 02 5d, an
ndarray of 2^64 elements followed by no data. Reading that back throws
out_of_range.408 ("excessive ndarray size caused overflow"), so to_bjdata
produced output that from_bjdata rejects. This is reachable by parsing
untrusted JSON and re-encoding it as BJData.

Mirror the overflow check the binary reader already performs, and also
reject a single dimension that does not fit into std::size_t, which the
previous cast silently truncated where std::size_t is narrower than 64
bits. Such objects now fall back to a plain object encoding, which is
what the surrounding type and length validation already does for
annotations it cannot represent, and they round-trip unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document when to_bjdata converts a JData annotation to an ND-array

The BJData page described the 1-D vector case as the only situation in
which an object carrying _ArrayType_/_ArraySize_/_ArrayData_ is not
written as a compact ND-array. The writer has always had several other
fallbacks -- an unknown _ArrayType_, a dimension that is not a
non-negative integer, an _ArrayData_ whose length does not match the
product of the dimensions, and elements that are not numbers of the
annotated kind -- all of which cause the value to be serialized as a
regular JSON object instead.

Spell out the conditions, including the size-overflow check added in the
preceding commit, so the documented behavior matches the implementation.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-08-05 13:44:15 +02:00
Niels Lohmann c034480c22 📝 add more docs (#5231)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-04 11:13:25 +02:00
co63oc 44bee1b138 chore: fix typos (#4921)
Signed-off-by: co63oc <co63oc@users.noreply.github.com>
2025-09-16 08:07:00 +02:00
Niels Lohmann 4424a0fcc1 📝 update documentation (#4723)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2025-04-05 18:54:35 +02:00
Niels Lohmann 0f9e6ae098 Fix broken links (#4605)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2025-01-18 23:20:45 +01:00
Nebojša Cvetković 2e50d5b2f3 BJData optimized binary array type (#4513) 2025-01-07 18:09:19 +01:00
Niels LohmannandFlorian Albrechtskirchinger 7b6cf5918b Documentation change (#3672)
Co-authored-by: Florian Albrechtskirchinger <falbrechtskirchinger@gmail.com>
2022-08-05 19:51:39 +02:00
Florian Albrechtskirchinger d3e347bd2d More documentation updates for 3.11.0 (#3553)
* mkdocs: add string_view examples

* mkdocs: reference underlying operators

* mkdocs: add operator<=> examples

* mkdocs: fix style check issues

* mkdocs: tweak BJData page

* mkdocs: add CMake option hints to macros

* mkdocs: fix JSON_DISABLE_ENUM_SERIALIZATION definition

* mkdocs: fix link to unit-udt.cpp

* mkdocs: fix "Arbitrary Type Conversions" title

* mkdocs: link to api/macros/*.md instead of features/macros.md

* mkdocs: document JSON_DisableEnumSerialization CMake option

* mkdocs: encode required C++ standard in example files

* docset: detect gsed/sed

* docset: update index

* docset: fix CSS patching

* docset: add list_missing_pages make target

* docset: add list_removed_paths make target

* docset: replace page titles with name from index

* docset: add install target for Zeal docset browser

* Use GCC_TOOL in ci_test_documentation target
2022-07-31 14:05:58 +02:00
6a7392058e Complete documentation for 3.11.0 (#3464)
* 👥 update contributor and sponsor list

* 🚧 document BJData format

* 🚧 document BJData format

* 📝 clarified documentation of [json.exception.parse_error.112]

* ✏️ adjust titles

* 📝 add more examples

* 🚨 adjust warnings for index.md files

* 📝 add more examples

* 🔥 remove example for deprecated code

* 📝 add missing enum entry

* 📝 overwork table for binary formats

* ✅ add test to create table for binary formats

* 📝 fix wording in example

* 📝 add more examples

* Update iterators.md (#3481)

* ✨ add check for overloads to linter #3455

* 👥 update contributor list

* 📝 add more examples

* 📝 fix documentation

* 📝 add more examples

* 🎨 fix indentation

* 🔥 remove example for destructor

* 📝 overwork documentation

* Updated BJData documentation, #3464 (#3493)

* update bjdata.md for #3464

* Minor edit

* Fix URL typo

* Add info on demoting ND array to a 1-D optimized array when singleton dimension

Co-authored-by: Chaoqi Zhang <prncoprs@163.com>
Co-authored-by: Qianqian Fang <fangqq@gmail.com>
2022-05-17 13:08:56 +02:00