Keep BJData ndarray annotations that would not survive a round trip as objects (#5542)

write_bjdata_ndarray() encoded a JData-annotated object as a BJData
ND-array whenever its dimensions' product matched _ArrayData_.size(),
which lost information in two ways:

- _ArrayData_ was never required to be an array. null has size 0, any
  other scalar has size 1, and iterating an object visits its values, so
  e.g. {"_ArraySize_":[1],"_ArrayData_":5} was written as the array [5],
  and an object _ArrayData_ came back as an array.

- The reader only restores an annotated object from an ND-array with at
  least two non-zero dimensions that is not a 1xN row vector; an empty,
  1-D, row-vector, or zero-sized shape is read back as a plain array. The
  writer nonetheless emitted ND-array headers for these shapes, so the
  annotation was silently dropped.

OSS-Fuzz issue 563659413 hit this in parse_bjdata_fuzzer: an empty binary
_ArraySize_ is written as a plain object and read back as an empty array,
after which {"_ArrayType_":"int16","_ArraySize_":[],"_ArrayData_":null}
was encoded as the ND-array header "[$I#[]" and re-read as [], failing the
harness's value-stability check.

Such objects now fall back to a plain object encoding, which round-trips.
Genuine ND-arrays (two or more positive dimensions, not a 1xN row vector)
are encoded exactly as before. Existing fallback tests that used 1-D
shapes are moved to 2-D shapes so they keep exercising the check they
were written for, and the BJData documentation is updated.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-24 17:01:31 +02:00
committed by GitHub
parent f3768d6868
commit 918da64657
4 changed files with 156 additions and 34 deletions
+26 -4
View File
@@ -20664,8 +20664,19 @@ class binary_writer
return true;
}
std::size_t len = (value.at(key).empty() ? 0 : 1);
for (const auto& el : value.at(key))
// the reader only restores an annotated object from an ND-array header
// with at least two dimensions: an empty dimension vector, a single
// dimension, or a 1xN row vector is read back as a plain array, which
// would silently drop the annotation, so such an object falls back to
// a plain object encoding instead
const auto& dims = value.at(key);
if (dims.size() < 2 || (dims.size() == 2 && dims.at(0).is_number_integer() && dims.at(0).template get<std::int64_t>() == 1))
{
return true;
}
std::size_t len = 1;
for (const auto& el : dims)
{
// a dimension is read as an unsigned value below, so anything that
// is not a non-negative integer is rejected: a non-integer entry
@@ -20687,15 +20698,26 @@ class binary_writer
return true;
}
const auto dim_size = static_cast<std::size_t>(dim);
if (dim_size != 0 && len > (std::numeric_limits<std::size_t>::max)() / dim_size)
// the reader turns an ND-array with any zero dimension into an
// empty plain array, dropping the annotation, so keep the object
if (dim_size == 0)
{
return true;
}
if (len > (std::numeric_limits<std::size_t>::max)() / dim_size)
{
return true;
}
len *= dim_size;
}
// the elements are written from _ArrayData_ as a flat list, so it has
// to be an array: size() is 0 for null and 1 for any other scalar, and
// iterating an object visits its values, so any of these could match
// the dimensions by accident and be encoded as an unrelated ND-array
key = "_ArrayData_";
if (value.at(key).size() != len)
if (!value.at(key).is_array() || value.at(key).size() != len)
{
return true;
}