Add BON8 support (#2998)

* Add BON8 support

Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format
that uses the byte values that cannot begin a UTF-8 character as type
markers, so strings need no length prefix. It is the most compact of the
supported binary formats on the benchmark files.

The reader is non-recursive like the other binary readers. A string ends
at the first byte that cannot continue it, so the reader hands the one or
two bytes it reads past a string back to the value that follows. The
writer produces the canonical representation of the specification, except
for NFC normalization; its output is identical to that of the reference
implementation (HikoGUI) on all files of the test data.

The round-trip tests need the .bon8 files of json_test_data 3.2.0.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Address review comments

- Reuse detail::validate_one_utf8 to check strings in to_bon8; the error
  now names the first byte of the invalid sequence.
- Document that to_bon8 leaves bytes in the output adapter on an
  exception, and that string_open is only an output of write_bon8_marker.
- Explain why the pushback buffer of the BON8 reader cannot overflow.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Select the BON8 float prefix by type

get_bon8_float_prefix only depends on the type of its argument, so make
the type a template parameter instead of passing an unused value.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Rename a test variable that Flawfinder mistakes for read()

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures

- compare the float in write_bon8_float with number_float_t constants,
  so GCC does not warn about a float-to-double conversion
- mark check_bon8_utf8's context as used when exceptions are disabled
- choose the compact float prefix in a helper rather than with nested
  conditional operators (clang-tidy)
- use auto for the cast in the BON8 integer reader (clang-tidy)
- write the int32 minimum test values as long long literals (MSVC C4146)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BON8 strings in bulk from contiguous input

- copy the valid UTF-8 of a string in one step when the input is
  contiguous (twitter.json is read in 1.68 instead of 2.52 ms,
  jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack)
- share the new valid_utf8_prefix() with the writer's UTF-8 check, which
  now skips ASCII 8 bytes at a time
- let the fuzzer check that contiguous and stream input give the same
  value or error, and test both paths in the unit tests
- clarify that a second 0xFF after a string is an empty string

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Link the BON8 functions from the other binary format pages

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Name the bulk scan flag after the input, not BON8

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read BSON keys in bulk from contiguous input

BSON keys (and array indices) are C-style strings, which were read byte
by byte. For contiguous input they are now read up to their \x00-byte in
one step, using the same bulk_scan flag as BON8 strings: twitter.json is
read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of
3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys
are almost all one-digit array indices, takes 2 % longer.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix the BON8 CI failures of the bulk-read tests

- skip the contiguous-versus-stream tests of BON8 strings and BSON keys
  when exceptions are disabled: they catch the parse errors of invalid
  input, and without exceptions the library aborts instead
- use static_cast for the int64 test value (google-readability-casting)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Move the explicit basic_json instantiation into its own test file

Linking test-regression3_cpp20 with clang and MinGW failed with
"relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'",
as test-regression2 did before #5511. The explicit instantiation of
basic_json<> for #4825 compiles every member function, including the
BON8 reader and writer, into that object, and it was already close to
the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1,
C++20).

Give the instantiation a file of its own: unit-regression3 is now
1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file
mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built
for the C++17 standard the regression was about.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert the bytes of the BON8 test strings explicitly

The str() helper constructed a std::string from a byte range, which
converts each unsigned char implicitly; -fsanitize=integer reports that
for bytes of 0x80 and above (ci_test_clang_sanitizer).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-27 16:56:21 +02:00
committed by GitHub
parent f682cd2ef1
commit 1e101ecac1
58 changed files with 3950 additions and 53 deletions
@@ -25,6 +25,7 @@
#endif
#include <nlohmann/detail/input/binary_reader.hpp>
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/output/output_adapters.hpp>
#include <nlohmann/detail/string_concat.hpp>
@@ -76,7 +77,7 @@ std::size_t binary_reserve_hint(const BasicJsonType& j)
}
/*!
@brief serialization to CBOR and MessagePack values
@brief serialization to BJData, BON8, BSON, CBOR, MessagePack, and UBJSON values
*/
template<typename BasicJsonType, typename CharType, typename OutputSinkType = output_adapter_sink<CharType>>
class binary_writer
@@ -873,6 +874,21 @@ class binary_writer
}
}
/*!
@param[in] j JSON value to serialize
*/
void write_bon8(const BasicJsonType& j)
{
bool string_open = false;
write_bon8_value(j, string_open);
// the last string of a message must be terminated
if (string_open)
{
oa.write_character(to_char_type(0xFF));
}
}
private:
//////////
// BSON //
@@ -1431,6 +1447,28 @@ class binary_writer
return to_char_type(0xCB); // float 64
}
/// @return the BON8 type marker for binary32 (float) or binary64 (double)
template<typename FloatType>
static constexpr CharType get_bon8_float_prefix()
{
return to_char_type(std::is_same<FloatType, float>::value ? 0x8E : 0x8F);
}
/// @return the type marker for a FloatType value in @a format (CBOR, MessagePack, or BON8)
template<typename FloatType>
static CharType get_compact_float_prefix(const detail::input_format_t format)
{
if (format == detail::input_format_t::cbor)
{
return get_cbor_float_prefix(FloatType{});
}
if (format == detail::input_format_t::bon8)
{
return get_bon8_float_prefix<FloatType>();
}
return get_msgpack_float_prefix(FloatType{});
}
////////////
// UBJSON //
////////////
@@ -2050,6 +2088,322 @@ class binary_writer
return false;
}
//////////
// BON8 //
//////////
/*!
@brief write a BON8 value
A string is written without length or terminator: it ends at the first
byte that cannot continue it, which is the first byte of any non-string
value and of the end-of-container marker 0xFE. It only needs an explicit
end-of-string marker (0xFF) when it is empty, when another string follows,
or when it is the last thing in the message.
@param[in] j JSON value to serialize
@param[in,out] string_open whether the output ends with a non-empty
string that has not been terminated with 0xFF
*/
void write_bon8_value(const BasicJsonType& j, bool& string_open)
{
switch (j.type())
{
case value_t::null:
{
write_bon8_marker(0xFA, string_open);
break;
}
case value_t::boolean:
{
write_bon8_marker(j.m_data.m_value.boolean ? 0xF9 : 0xF8, string_open);
break;
}
case value_t::number_unsigned:
{
if (j.m_data.m_value.number_unsigned > static_cast<typename BasicJsonType::number_unsigned_t>((std::numeric_limits<std::int64_t>::max)()))
{
JSON_THROW(out_of_range::create(407, concat("integer number ", std::to_string(j.m_data.m_value.number_unsigned), " cannot be represented by BON8 as it does not fit int64"), &j));
}
write_bon8_integer(static_cast<std::int64_t>(j.m_data.m_value.number_unsigned));
string_open = false;
break;
}
case value_t::number_integer:
{
write_bon8_integer(static_cast<std::int64_t>(j.m_data.m_value.number_integer));
string_open = false;
break;
}
case value_t::number_float:
{
write_bon8_float(j.m_data.m_value.number_float);
string_open = false;
break;
}
case value_t::string:
{
write_bon8_string(*j.m_data.m_value.string, string_open, j);
break;
}
case value_t::array:
{
const auto N = j.m_data.m_value.array->size();
// 0x80..0x84: array with 0..4 elements; 0x85: array ended by 0xFE
write_bon8_marker(static_cast<std::uint8_t>(N <= 4 ? 0x80 + N : 0x85), string_open);
for (const auto& el : *j.m_data.m_value.array)
{
write_bon8_value(el, string_open);
}
if (N > 4)
{
write_bon8_marker(0xFE, string_open);
}
break;
}
case value_t::object:
{
const auto N = j.m_data.m_value.object->size();
// 0x86..0x8A: object with 0..4 members; 0x8B: object ended by 0xFE
write_bon8_marker(static_cast<std::uint8_t>(N <= 4 ? 0x86 + N : 0x8B), string_open);
for (const auto& el : *j.m_data.m_value.object)
{
write_bon8_string(el.first, string_open, j);
write_bon8_value(el.second, string_open);
}
if (N > 4)
{
write_bon8_marker(0xFE, string_open);
}
break;
}
case value_t::binary:
{
// BON8 has no binary type: write the bytes as an array of
// integers, like UBJSON and BJData do
const auto N = j.m_data.m_value.binary->size();
write_bon8_marker(static_cast<std::uint8_t>(N <= 4 ? 0x80 + N : 0x85), string_open);
for (std::size_t i = 0; i < N; ++i)
{
// the cast is needed for binary types whose value type
// is not an integer (e.g., std::byte)
write_bon8_integer(static_cast<std::uint8_t>(j.m_data.m_value.binary->data()[i]));
}
if (N > 4)
{
write_bon8_marker(0xFE, string_open);
}
break;
}
case value_t::discarded:
default:
break;
}
}
/*!
@brief write a single byte that is not part of a string
@param[in] marker the byte to write
@param[out] string_open set to false, because the output no longer ends
with a string; see @ref write_bon8_value
*/
void write_bon8_marker(const std::uint8_t marker, bool& string_open)
{
oa.write_character(to_char_type(marker));
string_open = false;
}
/*!
@brief write a string
@param[in] s the string to write
@param[in,out] string_open see @ref write_bon8_value
@param[in] context the value the string belongs to (for diagnostics)
@throw type_error.316 if @a s is not valid UTF-8, because the end of a
string is determined from its encoding
*/
void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context)
{
check_bon8_utf8(s, context);
// a string that follows another string terminates it
if (string_open)
{
oa.write_character(to_char_type(0xFF));
}
if (s.empty())
{
// the empty string is just the end-of-string marker
oa.write_character(to_char_type(0xFF));
string_open = false;
}
else
{
oa.write_characters(reinterpret_cast<const CharType*>(s.data()), s.size());
string_open = true;
}
}
/*!
@brief check that a string is valid UTF-8 (RFC 3629)
@param[in] s the string to check
@param[in] context the value the string belongs to (for diagnostics)
@throw type_error.316 if @a s is not valid UTF-8; the message names the
first byte of the first invalid or incomplete sequence
*/
static void check_bon8_utf8(const string_t& s, const BasicJsonType& context)
{
static_cast<void>(context); // only used when exceptions are enabled
const auto* data = reinterpret_cast<const unsigned char*>(s.data());
const std::size_t valid = valid_utf8_prefix(data, s.size());
if (JSON_HEDLEY_UNLIKELY(valid != s.size()))
{
JSON_THROW(type_error::create(316, concat("invalid UTF-8 byte at index ", std::to_string(valid), ": 0x", hex_byte(data[valid])), &context));
}
}
/// @return a byte as two uppercase hexadecimal digits
static std::string hex_byte(const std::uint8_t byte)
{
std::string result = "00";
constexpr const char* nibble_to_hex = "0123456789ABCDEF";
result[0] = nibble_to_hex[byte / 16];
result[1] = nibble_to_hex[byte % 16];
return result;
}
/*!
@brief write an integer in the shortest encoding
Integers from -10 to 39 take one byte. Up to -33818506 and 67637031, an
integer takes 2 to 4 bytes that begin with a UTF-8 lead byte (0xC2..0xF7)
followed by a byte that is not a continuation byte: 0x00..0x7F for
positive and 0xC0..0xFF for negative integers. Each range starts where the
shorter one ends. Larger integers are written as int32 (0x8C) or int64
(0x8D) in big-endian byte order.
@param[in] value the integer to write
*/
void write_bon8_integer(std::int64_t value)
{
if (value < (std::numeric_limits<std::int32_t>::min)() || value > (std::numeric_limits<std::int32_t>::max)())
{
oa.write_character(to_char_type(0x8D));
write_number(value);
}
else if (value < -33818506 || value > 67637031)
{
oa.write_character(to_char_type(0x8C));
write_number(static_cast<std::int32_t>(value));
}
else if (value <= -264075)
{
value = -(value + 264075);
write_bon8_bytes(0xF0 + ((value >> 22) & 0x07), 0xC0 + ((value >> 16) & 0x3F), value >> 8, value);
}
else if (value <= -1931)
{
value = -(value + 1931);
write_bon8_bytes(0xE0 + ((value >> 14) & 0x0F), 0xC0 + ((value >> 8) & 0x3F), value);
}
else if (value <= -11)
{
value = -(value + 11);
write_bon8_bytes(0xC2 + ((value >> 6) & 0x1F), 0xC0 + (value & 0x3F));
}
else if (value <= -1)
{
write_bon8_bytes(0xB8 - (value + 1));
}
else if (value <= 39)
{
write_bon8_bytes(0x90 + value);
}
else if (value <= 3879)
{
value -= 40;
write_bon8_bytes(0xC2 + ((value >> 7) & 0x1F), value & 0x7F);
}
else if (value <= 528167)
{
value -= 3880;
write_bon8_bytes(0xE0 + ((value >> 15) & 0x0F), (value >> 8) & 0x7F, value);
}
else
{
value -= 528168;
write_bon8_bytes(0xF0 + ((value >> 23) & 0x07), (value >> 16) & 0x7F, value >> 8, value);
}
}
/// write the low byte of each argument
template<typename... Bytes>
void write_bon8_bytes(const Bytes... bytes)
{
const std::array<CharType, sizeof...(Bytes)> buffer{{to_char_type(static_cast<std::uint8_t>(bytes & 0xFF))...}};
oa.write_characters(buffer.data(), buffer.size());
}
/*!
@brief write a floating-point number
-1.0, +0.0, and 1.0 take one byte. Other numbers are written as binary32
(0x8E) if that loses no precision, and as binary64 (0x8F) otherwise; -0.0,
infinities, and NaN are always written as binary32, NaN as 0x7F800001.
@param[in] n the number to write
*/
void write_bon8_float(const number_float_t n)
{
#ifdef __GNUC__
JSON_HEDLEY_DIAGNOSTIC_PUSH
JSON_HEDLEY_PRAGMA(GCC diagnostic ignored "-Wfloat-equal")
#endif
if (n == static_cast<number_float_t>(-1))
{
oa.write_character(to_char_type(0xFB));
}
else if (n == static_cast<number_float_t>(0) && !std::signbit(n))
{
oa.write_character(to_char_type(0xFC));
}
else if (n == static_cast<number_float_t>(1))
{
oa.write_character(to_char_type(0xFD));
}
else if (std::isnan(n))
{
write_bon8_bytes(0x8E, 0x7F, 0x80, 0x00, 0x01);
}
else
{
write_compact_float(n, detail::input_format_t::bon8);
}
#ifdef __GNUC__
JSON_HEDLEY_DIAGNOSTIC_POP
#endif
}
///////////////////////
// Utility functions //
///////////////////////
@@ -2184,16 +2538,12 @@ class binary_writer
static_cast<double>(n) <= static_cast<double>((std::numeric_limits<float>::max)()) &&
static_cast<double>(static_cast<float>(n)) == static_cast<double>(n))))
{
oa.write_character(format == detail::input_format_t::cbor
? get_cbor_float_prefix(static_cast<float>(n))
: get_msgpack_float_prefix(static_cast<float>(n)));
oa.write_character(get_compact_float_prefix<float>(format));
write_number(static_cast<float>(n));
}
else
{
oa.write_character(format == detail::input_format_t::cbor
? get_cbor_float_prefix(n)
: get_msgpack_float_prefix(n));
oa.write_character(get_compact_float_prefix<number_float_t>(format));
write_number(n);
}
#ifdef __GNUC__