Reject ill-formed UTF-8 in CBOR/UBJSON/BJData/BSON writers, accept it in MessagePack reader

to_cbor(), to_ubjson(), to_bjdata(), and to_bson() wrote a string or object
key with ill-formed UTF-8 byte for byte, but from_cbor()/from_ubjson()/
from_bjdata()/from_bson() reject such strings with parse_error.113 since
d19f7f5dc (#5185, not yet released): a value that serialized without error
could not be read back by the same library (#5651).

CBOR (RFC 8949 Section 5), UBJSON, BJData, and BSON all require text strings
and object keys/element names to be valid UTF-8, so their writers now
validate and throw type_error.316, like to_bon8() already does. For BSON,
the check runs in the size-computation pass, before any byte is written, the
same way the binary subtype check works. check_bon8_utf8() is renamed to
check_utf8() since it is now shared by all of these writers.

MessagePack's specification explicitly allows a str object to contain an
invalid byte sequence and expects a deserializer to hand the bytes back
unchanged, so its writer is unaffected and from_msgpack() (object keys
included) no longer validates UTF-8, restoring its pre-#5185 behavior.
dump() still rejects ill-formed UTF-8 with type_error.316 unless an error
handler is passed.

Docs: add the type_error.316 exception to to_cbor/to_ubjson/to_bjdata/to_bson,
document the writer-side check on the cbor/bson/ubjson/bjdata format pages,
and rewrite the MessagePack UTF-8 warning to describe the round-trip and
dump() behavior instead of a validation requirement the spec does not have.

Tests: add ill-formed value/key cases (invalid byte, truncated sequence,
encoded surrogate, overlong encoding) expecting type_error.316 to
unit-cbor.cpp, unit-ubjson.cpp, unit-bjdata.cpp, and unit-bson.cpp (which
also checks the output vector stays empty), and turn unit-msgpack.cpp's
former parse_error.113 ill-formed-UTF-8 tests into byte-for-byte round-trip
tests for both values and keys.

This text was written by Claude Code.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-30 22:19:11 +02:00
parent 7d7055ec50
commit 73d116547a
17 changed files with 297 additions and 41 deletions
+24
View File
@@ -4025,6 +4025,30 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_bjdata(j) == v);
CHECK(json::from_bjdata(v) == j);
}
SECTION("ill-formed UTF-8 (see #5651)")
{
// a string value whose bytes are not valid UTF-8 (0xC0 0xAE is an
// overlong encoding of '.') is rejected at decode time, matching
// every other kind of malformed binary input, and to_bjdata()
// rejects it as well, so a value it accepts can always be read
// back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(v), "[json.exception.parse_error.113] parse error at byte 5: syntax error while parsing BJData string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_bjdata(v, true, false).is_discarded());
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_bjdata(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")
+38
View File
@@ -151,6 +151,44 @@ TEST_CASE("BSON")
#endif
}
SECTION("ill-formed UTF-8 (see #5651)")
{
// a BSON document {"s": "\xC0\xAE"} (0xC0 0xAE is an overlong
// encoding of '.'); the reader rejects an ill-formed string value at
// decode time
const std::vector<uint8_t> v =
{
0x0F, 0x00, 0x00, 0x00, // document length
0x02, 's', 0x00, // type 0x02 (string), key "s"
0x03, 0x00, 0x00, 0x00, // string length (including null)
0xc0, 0xae, 0x00, // string content and its null terminator
0x00 // document terminator
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_bson(v), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing BSON string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_bson(v, true, false).is_discarded());
// to_bson() rejects the same kind of ill-formed string value, before
// any bytes reach the output adapter (the BSON document length
// prefix must be known up front, so nothing is written incrementally)
std::vector<std::uint8_t> out{0x42}; // a sentinel byte the writer must not touch
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}, nlohmann::detail::output_adapter<std::uint8_t>(out)), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
CHECK(out == std::vector<std::uint8_t> {0x42});
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC3"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xED\xA0\x80"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC0\xAF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected as well; unlike
// the reader (which never validates element names), the writer
// checks both string values and object keys
CHECK_THROWS_WITH_AS(json::to_bson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON")
{
// out_of_range.412 is thrown from a single shared helper
+19
View File
@@ -1896,6 +1896,25 @@ TEST_CASE("CBOR")
CHECK(json::from_cbor(json::to_cbor(j)) == j);
}
SECTION("to_cbor rejects ill-formed UTF-8 (see #5651)")
{
// to_cbor() must reject the same ill-formed strings from_cbor()
// rejects, so a value it accepts can always be read back
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_cbor(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// binary values are not text and are unaffected
CHECK_NOTHROW(json::to_cbor(json::binary(std::vector<std::uint8_t>({0xFF}))));
}
SECTION("invalid UTF-8 in indefinite-length string")
{
json _;
+28 -8
View File
@@ -1614,19 +1614,39 @@ TEST_CASE("MessagePack")
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
SECTION("ill-formed UTF-8 in string (see #5529, #5651)")
{
// the MessagePack specification explicitly allows a str object to
// contain a byte sequence that is not valid UTF-8 and expects a
// deserializer to hand the original bytes back unchanged; this
// library follows that, unlike CBOR/UBJSON/BJData/BSON, whose
// specifications require text strings to be valid UTF-8
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
// (0xC0 0xAE is an overlong encoding of '.') must be rejected at
// decode time, matching every other kind of malformed binary
// input, rather than only failing later when the resulting
// value is dumped
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing MessagePack string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae}), true, false).is_discarded());
// (0xC0 0xAE is an overlong encoding of '.') round-trips byte for
// byte as a string value
const std::vector<uint8_t> ill_formed_value = {0xa2, 0xc0, 0xae};
json j_value;
CHECK_NOTHROW(j_value = json::from_msgpack(ill_formed_value));
REQUIRE(j_value.is_string());
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK(json::from_msgpack(json::to_msgpack(j_value)) == j_value);
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0x81, 0xa2, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_msgpack(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_msgpack(json::to_msgpack(j_key)) == j_key);
// a MessagePack bin8 blob with the very same bytes is NOT text
// and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_msgpack(std::vector<uint8_t>({0xc4, 0x02, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
+24
View File
@@ -2580,6 +2580,30 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_ubjson(j) == v);
CHECK(json::from_ubjson(v) == j);
}
SECTION("ill-formed UTF-8 (see #5651)")
{
// a string value whose bytes are not valid UTF-8 (0xC0 0xAE is an
// overlong encoding of '.') is rejected at decode time, matching
// every other kind of malformed binary input, and to_ubjson()
// rejects it as well, so a value it accepts can always be read
// back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(v), "[json.exception.parse_error.113] parse error at byte 5: syntax error while parsing UBJSON string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_ubjson(v, true, false).is_discarded());
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_ubjson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")