Compare commits

..
Author SHA1 Message Date
Niels Lohmann 92c729f657 Accept ill-formed UTF-8 in all binary readers again
RFC 8949 and the MessagePack/BSON/UBJSON/BJData specs leave UTF-8
well-formedness checking up to the decoder, so following #5529 the
binary readers are lenient by default again, as in release 3.12.0
(the reader-side check was added by #5185/#5531, not in any release);
reader-side validation becomes opt-in in a follow-up PR. The writers
stay strict and throw type_error.316 for ill-formed UTF-8. BON8 is
unchanged, since UTF-8 lead bytes are structural there.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:59:14 +02:00
Niels Lohmann f3e30fa5b7 Merge branch 'develop' into claude/binary-utf8-roundtrip-5651
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:44:29 +02:00
Niels Lohmann 73d116547a Reject ill-formed UTF-8 in CBOR/UBJSON/BJData/BSON writers, accept it in MessagePack reader
to_cbor(), to_ubjson(), to_bjdata(), and to_bson() wrote a string or object
key with ill-formed UTF-8 byte for byte, but from_cbor()/from_ubjson()/
from_bjdata()/from_bson() reject such strings with parse_error.113 since
d19f7f5dc (#5185, not yet released): a value that serialized without error
could not be read back by the same library (#5651).

CBOR (RFC 8949 Section 5), UBJSON, BJData, and BSON all require text strings
and object keys/element names to be valid UTF-8, so their writers now
validate and throw type_error.316, like to_bon8() already does. For BSON,
the check runs in the size-computation pass, before any byte is written, the
same way the binary subtype check works. check_bon8_utf8() is renamed to
check_utf8() since it is now shared by all of these writers.

MessagePack's specification explicitly allows a str object to contain an
invalid byte sequence and expects a deserializer to hand the bytes back
unchanged, so its writer is unaffected and from_msgpack() (object keys
included) no longer validates UTF-8, restoring its pre-#5185 behavior.
dump() still rejects ill-formed UTF-8 with type_error.316 unless an error
handler is passed.

Docs: add the type_error.316 exception to to_cbor/to_ubjson/to_bjdata/to_bson,
document the writer-side check on the cbor/bson/ubjson/bjdata format pages,
and rewrite the MessagePack UTF-8 warning to describe the round-trip and
dump() behavior instead of a validation requirement the spec does not have.

Tests: add ill-formed value/key cases (invalid byte, truncated sequence,
encoded surrogate, overlong encoding) expecting type_error.316 to
unit-cbor.cpp, unit-ubjson.cpp, unit-bjdata.cpp, and unit-bson.cpp (which
also checks the output vector stays empty), and turn unit-msgpack.cpp's
former parse_error.113 ill-formed-UTF-8 tests into byte-for-byte round-trip
tests for both values and keys.

This text was written by Claude Code.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:19:11 +02:00
22 changed files with 394 additions and 288 deletions
+4 -1
View File
@@ -56,6 +56,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8
## Complexity
@@ -89,4 +91,5 @@ Linear in the size of the JSON value `j`.
## Version history
- Added in version 3.11.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
- BJData version parameter (for draft3 binary encoding) added in version 3.12.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
@@ -46,6 +46,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`out_of_range.415`](../../home/exceptions.md#jsonexceptionout_of_range415) if the subtype of a binary value
exceeds 255, the maximum of the BSON binary subtype; example:
`"subtype 70000 is too large for the BSON binary subtype (max 255)"`
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key is
not valid UTF-8
## Complexity
@@ -82,3 +84,5 @@ pass before anything is written.
- Added in version 3.4.0.
- Linear in the size of `j`, and no longer limited by the call stack for deeply nested values, since version 3.13.0.
- `out_of_range.415` is now detected before anything is written, like the other exceptions above, since version 3.13.0.
- Throwing `type_error.316` for a string value or object key that is not valid UTF-8, detected before anything is
written, added in version 3.13.0.
@@ -35,6 +35,11 @@ The exact mapping and its limitations are described on a [dedicated page](../../
Strong guarantee: if an exception is thrown, there are no changes in the JSON value.
## Exceptions
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8
## Complexity
Linear in the size of the JSON value `j`.
@@ -68,3 +73,4 @@ Linear in the size of the JSON value `j`.
- Added in version 2.0.9.
- Compact representation of floating-point numbers added in version 3.8.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
@@ -49,6 +49,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
- Throws [`other_error.502`](../../home/exceptions.md#jsonexceptionother_error502) if `use_type` is true and `use_size`
is false.
- Throws [type_error.316](../../home/exceptions.md#jsonexceptiontype_error316) if a string or object key in `j` is
not valid UTF-8
## Complexity
@@ -82,3 +84,4 @@ Linear in the size of the JSON value `j`.
## Version history
- Added in version 3.1.0.
- Throwing `type_error.316` for a string or object key that is not valid UTF-8 added in version 3.13.0.
@@ -63,6 +63,12 @@ The library uses the following mapping from JSON values types to BJData types ac
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
BJData strings must use UTF-8 encoding. `to_bjdata()` validates the bytes of every string value and object key
and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8, so a
value with such a string cannot be serialized in the first place.
!!! info "Unused BJData markers"
The following markers are not used in the conversion:
@@ -208,6 +214,15 @@ The library maps BJData types to JSON value types as follows:
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
BJData strings must use UTF-8 encoding, but this is not enforced on read: `from_bjdata()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()` is strict as well (see above), so
a value read this way cannot be written back to BJData.
!!! info "Round trips"
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
@@ -109,14 +109,17 @@ The library maps BSON record types to JSON value types as follows:
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
`from_bson()`.
!!! warning "UTF-8 validation of string values"
!!! warning "Ill-formed UTF-8 in string values"
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
bytes of every such string at decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
(key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
byte-by-byte as a C string, or are not required to hold text, respectively.
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
decoder. `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged.
However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_bson()` is strict as well and throws the
same exception for a string value or element (key) name that is not valid UTF-8, so an object with such a key
or value cannot be produced in the first place, even though `from_bson()` would accept it from another source.
Element (key) names are never validated on read, since they are read byte-by-byte as a C string. `binary`
values (type `0x05`) are unaffected, since they are not required to hold text.
??? example
@@ -189,15 +189,16 @@ The library maps CBOR types to JSON value types as follows:
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "UTF-8 validation of text strings"
!!! warning "Ill-formed UTF-8 in text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
text.
(major type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this. This library does
not: `from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them
back unchanged. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_cbor()` is strict as well and throws the
same exception for a string value or object key that is not valid UTF-8, so such a value cannot be written back
to CBOR. Byte strings (major type 2) are unaffected, since they are not required to hold text.
!!! warning "Tagged items"
@@ -153,14 +153,15 @@ The library maps MessagePack types to JSON value types as follows:
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values"
!!! warning "Ill-formed UTF-8 in string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
This library validates the bytes of every such string (object keys included) at decode time and rejects
ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
to hold text.
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
This library follows that: `from_msgpack()` reads `str` bytes (object keys included) as-is, without validating
them, and `to_msgpack()` writes them back as-is, so such a value round-trips through `from_msgpack(to_msgpack(j))`
byte for byte. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way, unless an
error handler is passed that replaces or ignores the ill-formed bytes.
??? example
@@ -47,6 +47,12 @@ The library uses the following mapping from JSON values types to UBJSON types ac
- strings with more than 9223372036854775807 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
UBJSON's required string encoding is UTF-8. `to_ubjson()` validates the bytes of every string value and object
key and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8, so
a value with such a string cannot be serialized in the first place.
!!! info "Unused UBJSON markers"
The following markers are not used in the conversion:
@@ -120,6 +126,15 @@ The library maps UBJSON types to JSON value types as follows:
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
UBJSON's required string encoding is UTF-8, but this is not enforced on read: `from_ubjson()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()` is strict as well (see above), so
a value read this way cannot be written back to UBJSON.
??? example
```cpp
+3 -5
View File
@@ -340,8 +340,9 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
### json.exception.parse_error.113
A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8.
string was read where one was required (for instance as a map key), or the string's length specification is invalid.
The bytes of a string itself are not checked for valid UTF-8 on read; see the ill-formed UTF-8 notes on the
individual [binary format](../features/binary_formats/index.md) pages for how such a string is handled afterward.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
@@ -364,9 +365,6 @@ type (for instance integers or `null`) are therefore not supported; see the note
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative
```
```
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte
```
### json.exception.parse_error.114
@@ -4031,28 +4031,13 @@ class binary_reader
const NumberType len,
string_t& result)
{
// get_bytes() appends to result, and CBOR indefinite-length strings
// collect all their chunks in the same result; validating only the
// newly read bytes keeps the check linear in the input size
const std::size_t old_size = result.size();
if (JSON_HEDLEY_UNLIKELY(!get_bytes(format, len, "string", result)))
{
return false;
}
// RFC 8949 (CBOR) §3.1 and the MessagePack/BSON/UBJSON specifications
// all require text strings to be valid UTF-8; reject anything else
// right here so malformed input is caught at decode time instead of
// only surfacing later as a type_error.316 when the value is dumped
// (which would defeat allow_exceptions=false / strict discarding).
if (JSON_HEDLEY_UNLIKELY(!is_valid_utf8(result, old_size)))
{
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(113, chars_read,
exception_message(format, "invalid string: ill-formed UTF-8 byte", "string"), nullptr));
}
return true;
// Strings are taken as is: none of CBOR (RFC 8949 §3.1 leaves the
// choice to the decoder), MessagePack (whose spec explicitly allows
// a str object to contain an invalid byte sequence), UBJSON, BJData,
// or BSON requires a decoder to reject ill-formed UTF-8. The bytes
// are kept unchanged; dump() and the binary writers are the ones
// that check them and report type_error.316 if they are not valid.
return get_bytes(format, len, "string", result);
}
/*!
@@ -115,6 +115,8 @@ class binary_writer
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
@throw type_error.317 if @a j is not an object
*/
void write_bson(const BasicJsonType& j)
@@ -145,6 +147,8 @@ class binary_writer
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_cbor(const BasicJsonType& j)
{
@@ -211,6 +215,8 @@ class binary_writer
case value_t::string:
{
check_utf8(*j.m_data.m_value.string, j);
// step 1: write control byte and the string length
write_cbor_head(0x60, j.m_data.m_value.string->size());
@@ -287,6 +293,11 @@ class binary_writer
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object)
{
// el.first is checked here, against the object as
// diagnostics context, because write_cbor(el.first)
// converts it to a temporary basic_json that would be
// used as the context instead
check_utf8(el.first, j);
write_cbor(el.first);
write_cbor(el.second);
}
@@ -629,6 +640,8 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true,
@@ -678,6 +691,8 @@ class binary_writer
case value_t::string:
{
check_utf8(*j.m_data.m_value.string, j);
if (add_prefix)
{
oa.write_character(to_char_type('S'));
@@ -840,6 +855,7 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object)
{
check_utf8(el.first, j);
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(el.first.data()),
@@ -884,6 +900,10 @@ class binary_writer
/*!
@return The size of a BSON document entry header, including the id marker
and the entry name size (and its null-terminator).
@throw out_of_range.409 if @a name contains U+0000, before anything is
written
@throw type_error.316 if @a name is not valid UTF-8, before anything is
written
*/
static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
{
@@ -893,7 +913,8 @@ class binary_writer
JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j));
}
static_cast<void>(j);
check_utf8(name, j);
return /*id*/ 1ul + name.size() + /*zero-terminator*/1u;
}
@@ -949,9 +970,21 @@ class binary_writer
/*!
@return The size of the BSON-encoded string in @a value
@throw type_error.316 if @a value is not valid UTF-8, before anything is
written
@note The UTF-8 check is skipped if @a value is already too long for the
32-bit BSON length field (@ref to_bson_length rejects it later, once
the size of the whole document is known); this also keeps the check
from reading past a StringType that reports a size larger than what
it actually holds.
*/
static std::size_t calc_bson_string_size(const string_t& value)
static std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<std::int32_t>(value.size())))
{
check_utf8(value, j);
}
return sizeof(std::int32_t) + value.size() + 1ul;
}
@@ -1080,6 +1113,8 @@ class binary_writer
is neither an object nor an array
@throw out_of_range.415 if @a j is binary with a subtype that does not fit
into a byte, before anything is written
@throw type_error.316 if @a j is a string that is not valid UTF-8, before
anything is written
*/
static std::size_t calc_bson_value_size(const BasicJsonType& j)
{
@@ -1101,7 +1136,7 @@ class binary_writer
return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned);
case value_t::string:
return calc_bson_string_size(*j.m_data.m_value.string);
return calc_bson_string_size(*j.m_data.m_value.string, j);
case value_t::null:
return 0ul;
@@ -1214,6 +1249,8 @@ class binary_writer
written
@throw out_of_range.415 if a binary value's subtype does not fit into a
byte, before anything is written
@throw type_error.316 if a string value or a key is not valid UTF-8,
before anything is written
*/
static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
{
@@ -2092,7 +2129,7 @@ class binary_writer
*/
void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context)
{
check_bon8_utf8(s, context);
check_utf8(s, context);
// a string that follows another string terminates it
if (string_open)
@@ -2122,7 +2159,7 @@ class binary_writer
@throw type_error.316 if @a s is not valid UTF-8; the message names the
first byte of the first invalid or incomplete sequence
*/
static void check_bon8_utf8(const string_t& s, const BasicJsonType& context)
static void check_utf8(const string_t& s, const BasicJsonType& context)
{
static_cast<void>(context); // only used when exceptions are enabled
const auto* data = reinterpret_cast<const unsigned char*>(s.data());
+6 -38
View File
@@ -117,13 +117,14 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
The library checks UTF-8 well-formedness (RFC 3629, section 4) in three
places, which differ in speed, diagnostics, and how they read the input:
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
text strings at decode time).
- decode() below: the serializer, to escape and, in strict mode, reject
ill-formed UTF-8 when dumping a string. The CBOR, MessagePack, BSON,
UBJSON and BJData readers do not use it: none of those specs requires a
decoder to reject ill-formed UTF-8 in text strings, so the readers keep
the bytes as is and leave the check to dump() and the binary writers.
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
@@ -178,38 +179,5 @@ inline std::uint8_t decode(std::uint8_t& state, std::uint32_t& codep, const std:
return state;
}
/*!
@brief check whether a string consists solely of valid UTF-8
Used by the CBOR/MessagePack/BSON/UBJSON binary readers to reject text
strings that are not valid UTF-8 at decode time (RFC 8949 §3.1 and the
MessagePack/BSON specifications all require text strings to be UTF-8), so
that malformed input is caught immediately instead of only surfacing later
as a type_error.316 when the resulting value is dumped.
@param[in] s the string to check
@param[in] first index of the first byte to check; the bytes before it are
assumed to have been validated already and to end on a
code point boundary
@return whether @a s (from index @a first on) is valid UTF-8
*/
template<typename StringType>
inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noexcept
{
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
for (std::size_t i = first; i < s.size(); ++i)
{
decode(state, codepoint, static_cast<std::uint8_t>(s[i]));
if (state == UTF8_REJECT)
{
return false;
}
}
return state == UTF8_ACCEPT;
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
-12
View File
@@ -763,15 +763,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static_cast<void>(check_parents);
}
// GCC 13 to at least 15 report a false -Warray-bounds error when set_parents()
// is inlined at -O3 right after a non-container value was created: the analysis
// does not use m_type to rule out the object/array branches and checks
// the std::map access against the allocation of, e.g., a string.
// See https://github.com/nlohmann/json/issues/5742 and #4819.
#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC diagnostic push
#pragma GCC diagnostic ignored "-Warray-bounds"
#endif
void set_parents()
{
#if JSON_DIAGNOSTICS
@@ -808,9 +799,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
}
#endif
}
#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC diagnostic pop
#endif
iterator set_parents(iterator it, std::ptrdiff_t count_set_parents)
{
+55 -77
View File
@@ -6321,13 +6321,14 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
The library checks UTF-8 well-formedness (RFC 3629, section 4) in three
places, which differ in speed, diagnostics, and how they read the input:
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
text strings at decode time).
- decode() below: the serializer, to escape and, in strict mode, reject
ill-formed UTF-8 when dumping a string. The CBOR, MessagePack, BSON,
UBJSON and BJData readers do not use it: none of those specs requires a
decoder to reject ill-formed UTF-8 in text strings, so the readers keep
the bytes as is and leave the check to dump() and the binary writers.
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
@@ -6382,39 +6383,6 @@ inline std::uint8_t decode(std::uint8_t& state, std::uint32_t& codep, const std:
return state;
}
/*!
@brief check whether a string consists solely of valid UTF-8
Used by the CBOR/MessagePack/BSON/UBJSON binary readers to reject text
strings that are not valid UTF-8 at decode time (RFC 8949 §3.1 and the
MessagePack/BSON specifications all require text strings to be UTF-8), so
that malformed input is caught immediately instead of only surfacing later
as a type_error.316 when the resulting value is dumped.
@param[in] s the string to check
@param[in] first index of the first byte to check; the bytes before it are
assumed to have been validated already and to end on a
code point boundary
@return whether @a s (from index @a first on) is valid UTF-8
*/
template<typename StringType>
inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noexcept
{
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
for (std::size_t i = first; i < s.size(); ++i)
{
decode(state, codepoint, static_cast<std::uint8_t>(s[i]));
if (state == UTF8_REJECT)
{
return false;
}
}
return state == UTF8_ACCEPT;
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -17544,28 +17512,13 @@ class binary_reader
const NumberType len,
string_t& result)
{
// get_bytes() appends to result, and CBOR indefinite-length strings
// collect all their chunks in the same result; validating only the
// newly read bytes keeps the check linear in the input size
const std::size_t old_size = result.size();
if (JSON_HEDLEY_UNLIKELY(!get_bytes(format, len, "string", result)))
{
return false;
}
// RFC 8949 (CBOR) §3.1 and the MessagePack/BSON/UBJSON specifications
// all require text strings to be valid UTF-8; reject anything else
// right here so malformed input is caught at decode time instead of
// only surfacing later as a type_error.316 when the value is dumped
// (which would defeat allow_exceptions=false / strict discarding).
if (JSON_HEDLEY_UNLIKELY(!is_valid_utf8(result, old_size)))
{
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(113, chars_read,
exception_message(format, "invalid string: ill-formed UTF-8 byte", "string"), nullptr));
}
return true;
// Strings are taken as is: none of CBOR (RFC 8949 §3.1 leaves the
// choice to the decoder), MessagePack (whose spec explicitly allows
// a str object to contain an invalid byte sequence), UBJSON, BJData,
// or BSON requires a decoder to reject ill-formed UTF-8. The bytes
// are kept unchanged; dump() and the binary writers are the ones
// that check them and report type_error.316 if they are not valid.
return get_bytes(format, len, "string", result);
}
/*!
@@ -21240,6 +21193,8 @@ class binary_writer
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
@throw type_error.317 if @a j is not an object
*/
void write_bson(const BasicJsonType& j)
@@ -21270,6 +21225,8 @@ class binary_writer
/*!
@param[in] j JSON value to serialize
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_cbor(const BasicJsonType& j)
{
@@ -21336,6 +21293,8 @@ class binary_writer
case value_t::string:
{
check_utf8(*j.m_data.m_value.string, j);
// step 1: write control byte and the string length
write_cbor_head(0x60, j.m_data.m_value.string->size());
@@ -21412,6 +21371,11 @@ class binary_writer
// step 2: write each element
for (const auto& el : *j.m_data.m_value.object)
{
// el.first is checked here, against the object as
// diagnostics context, because write_cbor(el.first)
// converts it to a temporary basic_json that would be
// used as the context instead
check_utf8(el.first, j);
write_cbor(el.first);
write_cbor(el.second);
}
@@ -21754,6 +21718,8 @@ class binary_writer
@param[in] add_prefix whether prefixes need to be used for this value
@param[in] use_bjdata whether write in BJData format, default is false
@param[in] bjdata_version which BJData version to use, default is draft2
@throw type_error.316 if a string value or an object key is not valid
UTF-8
*/
void write_ubjson(const BasicJsonType& j, const bool use_count,
const bool use_type, const bool add_prefix = true,
@@ -21803,6 +21769,8 @@ class binary_writer
case value_t::string:
{
check_utf8(*j.m_data.m_value.string, j);
if (add_prefix)
{
oa.write_character(to_char_type('S'));
@@ -21965,6 +21933,7 @@ class binary_writer
for (const auto& el : *j.m_data.m_value.object)
{
check_utf8(el.first, j);
write_number_with_ubjson_prefix(el.first.size(), true, use_bjdata);
oa.write_characters(
reinterpret_cast<const CharType*>(el.first.data()),
@@ -22009,6 +21978,10 @@ class binary_writer
/*!
@return The size of a BSON document entry header, including the id marker
and the entry name size (and its null-terminator).
@throw out_of_range.409 if @a name contains U+0000, before anything is
written
@throw type_error.316 if @a name is not valid UTF-8, before anything is
written
*/
static std::size_t calc_bson_entry_header_size(const string_t& name, const BasicJsonType& j)
{
@@ -22018,7 +21991,8 @@ class binary_writer
JSON_THROW(out_of_range::create(409, concat("BSON key cannot contain code point U+0000 (at byte ", std::to_string(it), ")"), &j));
}
static_cast<void>(j);
check_utf8(name, j);
return /*id*/ 1ul + name.size() + /*zero-terminator*/1u;
}
@@ -22074,9 +22048,21 @@ class binary_writer
/*!
@return The size of the BSON-encoded string in @a value
@throw type_error.316 if @a value is not valid UTF-8, before anything is
written
@note The UTF-8 check is skipped if @a value is already too long for the
32-bit BSON length field (@ref to_bson_length rejects it later, once
the size of the whole document is known); this also keeps the check
from reading past a StringType that reports a size larger than what
it actually holds.
*/
static std::size_t calc_bson_string_size(const string_t& value)
static std::size_t calc_bson_string_size(const string_t& value, const BasicJsonType& j)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<std::int32_t>(value.size())))
{
check_utf8(value, j);
}
return sizeof(std::int32_t) + value.size() + 1ul;
}
@@ -22205,6 +22191,8 @@ class binary_writer
is neither an object nor an array
@throw out_of_range.415 if @a j is binary with a subtype that does not fit
into a byte, before anything is written
@throw type_error.316 if @a j is a string that is not valid UTF-8, before
anything is written
*/
static std::size_t calc_bson_value_size(const BasicJsonType& j)
{
@@ -22226,7 +22214,7 @@ class binary_writer
return calc_bson_unsigned_size(j.m_data.m_value.number_unsigned);
case value_t::string:
return calc_bson_string_size(*j.m_data.m_value.string);
return calc_bson_string_size(*j.m_data.m_value.string, j);
case value_t::null:
return 0ul;
@@ -22339,6 +22327,8 @@ class binary_writer
written
@throw out_of_range.415 if a binary value's subtype does not fit into a
byte, before anything is written
@throw type_error.316 if a string value or a key is not valid UTF-8,
before anything is written
*/
static std::size_t calc_bson_sizes(const BasicJsonType& document, std::vector<std::size_t>& nested_sizes)
{
@@ -23217,7 +23207,7 @@ class binary_writer
*/
void write_bon8_string(const string_t& s, bool& string_open, const BasicJsonType& context)
{
check_bon8_utf8(s, context);
check_utf8(s, context);
// a string that follows another string terminates it
if (string_open)
@@ -23247,7 +23237,7 @@ class binary_writer
@throw type_error.316 if @a s is not valid UTF-8; the message names the
first byte of the first invalid or incomplete sequence
*/
static void check_bon8_utf8(const string_t& s, const BasicJsonType& context)
static void check_utf8(const string_t& s, const BasicJsonType& context)
{
static_cast<void>(context); // only used when exceptions are enabled
const auto* data = reinterpret_cast<const unsigned char*>(s.data());
@@ -27441,15 +27431,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static_cast<void>(check_parents);
}
// GCC 13 to at least 15 report a false -Warray-bounds error when set_parents()
// is inlined at -O3 right after a non-container value was created: the analysis
// does not use m_type to rule out the object/array branches and checks
// the std::map access against the allocation of, e.g., a string.
// See https://github.com/nlohmann/json/issues/5742 and #4819.
#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC diagnostic push
#pragma GCC diagnostic ignored "-Warray-bounds"
#endif
void set_parents()
{
#if JSON_DIAGNOSTICS
@@ -27486,9 +27467,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
}
#endif
}
#if defined(__GNUC__) && !defined(__clang__)
#pragma GCC diagnostic pop
#endif
iterator set_parents(iterator it, std::ptrdiff_t count_set_parents)
{
-5
View File
@@ -139,11 +139,6 @@ json_test_set_test_options(test-unicode4 TEST_PROPERTIES TIMEOUT 3000)
# only the #972 regression test needs thirdparty/fifo_map on its include path
json_test_set_test_options(test-regression1 LINK_LIBRARIES fifo_map_include)
# GCC's false -Warray-bounds error with JSON_DIAGNOSTICS only shows up when optimizing (#5742)
json_test_set_test_options(test-diagnostics-optimized
COMPILE_OPTIONS $<$<CXX_COMPILER_ID:GNU>:-O3 -Werror=array-bounds>
)
#############################################################################
# add unit tests
#############################################################################
+35
View File
@@ -3907,6 +3907,41 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_bjdata(j) == v);
CHECK(json::from_bjdata(v) == j);
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_bjdata() is
// strict, so such a value cannot be written back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_bjdata(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_bjdata(j), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_bjdata(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_bjdata(j_key), json::type_error&);
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bjdata(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_bjdata(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")
+45
View File
@@ -154,6 +154,51 @@ TEST_CASE("BSON")
#endif
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// a BSON document {"s": "\xC0\xAE"} (0xC0 0xAE is an overlong
// encoding of '.'); the BSON spec does not require a decoder to
// reject ill-formed UTF-8 in a string value, so the reader hands the
// bytes back unchanged
const std::vector<uint8_t> v =
{
0x0F, 0x00, 0x00, 0x00, // document length
0x02, 's', 0x00, // type 0x02 (string), key "s"
0x03, 0x00, 0x00, 0x00, // string length (including null)
0xc0, 0xae, 0x00, // string content and its null terminator
0x00 // document terminator
};
json j;
CHECK_NOTHROW(j = json::from_bson(v));
REQUIRE(j.is_object());
REQUIRE(j.contains("s"));
CHECK(j["s"].get_ref<const json::string_t&>() == std::string("\xc0\xae"));
// dump() still requires valid UTF-8 and throws for such a value
CHECK_THROWS_AS(j.dump(), json::type_error&);
// to_bson() is strict as well, so the value cannot be written back
CHECK_THROWS_AS(json::to_bson(j), json::type_error&);
// to_bson() rejects the same kind of ill-formed string value, before
// any bytes reach the output adapter (the BSON document length
// prefix must be known up front, so nothing is written incrementally)
std::vector<std::uint8_t> out{0x42}; // a sentinel byte the writer must not touch
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}, nlohmann::detail::output_adapter<std::uint8_t>(out)), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
CHECK(out == std::vector<std::uint8_t> {0x42});
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xFF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC3"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xED\xA0\x80"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_bson(json{{"s", "\xC0\xAF"}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected as well; unlike
// the reader (which never validates element names), the writer
// checks both string values and object keys
CHECK_THROWS_WITH_AS(json::to_bson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
SECTION("lengths exceeding INT32_MAX cannot be serialized to BSON")
{
// out_of_range.412 is thrown from a single shared helper
+65 -18
View File
@@ -1801,19 +1801,40 @@ TEST_CASE("CBOR")
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
SECTION("ill-formed UTF-8 in string (see #5529, #5651)")
{
// RFC 8949 §3.1 leaves it up to the decoder whether to reject
// ill-formed UTF-8 in a text string; this library does not, and
// hands the original bytes back unchanged, matching the
// MessagePack reader and the behavior before #5185/#5531 (not in
// any release)
// a two-character text string (major type 3) whose bytes are not
// valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') must be
// rejected at decode time, matching every other kind of
// malformed binary input, rather than only failing later when
// the resulting value is dumped
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0x62, 0xc0, 0xae}), true, false).is_discarded());
// valid UTF-8 (0xC0 0xAE is an overlong encoding of '.') round-trips
// byte for byte as a string value
const std::vector<uint8_t> ill_formed_value = {0x62, 0xc0, 0xae};
json j_value;
CHECK_NOTHROW(j_value = json::from_cbor(ill_formed_value));
REQUIRE(j_value.is_string());
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// to_cbor() is strict as well, so the value cannot be written back
CHECK_THROWS_AS(json::to_cbor(j_value), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0xa1, 0x62, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_cbor(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_cbor(j_key), json::type_error&);
// a CBOR byte string (major type 2) with the very same bytes is
// NOT text and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x42, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
@@ -1822,17 +1843,46 @@ TEST_CASE("CBOR")
CHECK(json::from_cbor(json::to_cbor(j)) == j);
}
SECTION("invalid UTF-8 in indefinite-length string")
SECTION("to_cbor rejects ill-formed UTF-8 (see #5651)")
{
// to_cbor() must reject the same ill-formed strings from_cbor()
// rejects, so a value it accepts can always be read back
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_cbor(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_cbor(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// binary values are not text and are unaffected
CHECK_NOTHROW(json::to_cbor(json::binary(std::vector<std::uint8_t>({0xFF}))));
}
SECTION("ill-formed UTF-8 in indefinite-length string")
{
json _;
// every chunk must be valid UTF-8 on its own (RFC 8949, Section
// 3.2.3), so a code point split across two chunks is rejected
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff}), true, false).is_discarded());
// the chunks are concatenated as is, without checking that each
// chunk is valid UTF-8 on its own (RFC 8949, Section 3.2.3), so
// a code point split across two chunks yields a valid string
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0x61, 0xa9, 0xff})));
CHECK(_ == "\xc3\xa9");
CHECK(_.dump() == "\"\xc3\xa9\"");
// an ill-formed later chunk is rejected after valid ones
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})), "[json.exception.parse_error.113] parse error at byte 7: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
// a truncated code point is kept as is
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x61, 0xc3, 0xff})));
CHECK(_ == "\xc3");
CHECK_THROWS_AS(_.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_cbor(_), json::type_error&);
// an ill-formed later chunk is kept after valid ones
CHECK_NOTHROW(_ = json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc0, 0xae, 0xff})));
CHECK(_ == "\xc3\xa9\xc0\xae");
CHECK_THROWS_AS(_.dump(), json::type_error&);
// valid multi-byte chunks are accepted
CHECK(json::from_cbor(std::vector<uint8_t>({0x7f, 0x62, 0xc3, 0xa9, 0x62, 0xc3, 0xb6, 0xff})) == "\xc3\xa9\xc3\xb6");
@@ -1840,9 +1890,6 @@ TEST_CASE("CBOR")
SECTION("many chunks in indefinite-length string")
{
// only the newly read chunk is validated, not the whole string
// collected so far; validating the latter made this input take
// quadratic time (about ten seconds for 100000 chunks)
constexpr std::size_t chunks = 100000;
std::vector<uint8_t> v{0x7f};
for (std::size_t i = 0; i < chunks; ++i)
-76
View File
@@ -1,76 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Regression test for https://github.com/nlohmann/json/issues/5742: with
// JSON_DIAGNOSTICS, GCC (13 to at least 15) reports a false -Warray-bounds
// error in the inlined set_parents() at -O3. The warning depends on GCC's
// inlining decisions, so the sections cover patterns that trigger it on
// different GCC versions (#4819: GCC 13 and 14; #5742: GCC 14 and 15).
// On GCC, this file is compiled with -O3 -Werror=array-bounds (see
// tests/CMakeLists.txt), so the test fails to build if the warning returns.
#include "doctest_compatibility.h"
#ifdef JSON_DIAGNOSTICS
#undef JSON_DIAGNOSTICS
#endif
#define JSON_DIAGNOSTICS 1
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <algorithm>
#include <iterator>
#include <utility>
#include <vector>
namespace
{
enum class diag_color
{
red,
green,
blue
};
void to_json(json& j, const diag_color& c)
{
static const std::pair<diag_color, json> m[] = // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
{
{diag_color::red, "r"},
{diag_color::green, "g"},
{diag_color::blue, "b"},
};
const auto* it = std::find_if(std::begin(m), std::end(m), [c](const std::pair<diag_color, json>& p)
{
return p.first == c;
});
j = it->second;
}
} // namespace
TEST_CASE("diagnostics with optimization")
{
SECTION("issue #4819 - object in vector")
{
std::vector<json> jsons{};
jsons.emplace_back(json({{"key", "value"}}));
CHECK(jsons.back()["key"] == "value");
}
SECTION("issue #5742 - string values from a static table")
{
json j = json::array();
j.push_back(diag_color::red);
j.push_back(diag_color::green);
j.push_back(diag_color::blue);
CHECK(j.dump() == R"(["r","g","b"])");
CHECK_THROWS_WITH_AS(j[1].get<int>(), "[json.exception.type_error.302] (/1) type must be number, but is string", json::type_error);
}
}
+28 -8
View File
@@ -1540,19 +1540,39 @@ TEST_CASE("MessagePack")
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
SECTION("ill-formed UTF-8 in string (see #5529, #5651)")
{
// the MessagePack specification explicitly allows a str object to
// contain a byte sequence that is not valid UTF-8 and expects a
// deserializer to hand the original bytes back unchanged; this
// library follows that, unlike CBOR/UBJSON/BJData/BSON, whose
// specifications require text strings to be valid UTF-8
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
// (0xC0 0xAE is an overlong encoding of '.') must be rejected at
// decode time, matching every other kind of malformed binary
// input, rather than only failing later when the resulting
// value is dumped
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae})), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing MessagePack string: invalid string: ill-formed UTF-8 byte", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0xa2, 0xc0, 0xae}), true, false).is_discarded());
// (0xC0 0xAE is an overlong encoding of '.') round-trips byte for
// byte as a string value
const std::vector<uint8_t> ill_formed_value = {0xa2, 0xc0, 0xae};
json j_value;
CHECK_NOTHROW(j_value = json::from_msgpack(ill_formed_value));
REQUIRE(j_value.is_string());
CHECK(j_value.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK(json::from_msgpack(json::to_msgpack(j_value)) == j_value);
// dump() still requires valid UTF-8 and throws for such a value,
// unless an error handler that replaces or ignores the bytes is
// passed
CHECK_THROWS_AS(j_value.dump(), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> ill_formed_key = {0x81, 0xa2, 0xc0, 0xae, 0x01};
json j_key;
CHECK_NOTHROW(j_key = json::from_msgpack(ill_formed_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK(json::from_msgpack(json::to_msgpack(j_key)) == j_key);
// a MessagePack bin8 blob with the very same bytes is NOT text
// and must still be accepted as-is
json _;
CHECK_NOTHROW(_ = json::from_msgpack(std::vector<uint8_t>({0xc4, 0x02, 0xc0, 0xae})));
CHECK(_ == json::binary(std::vector<std::uint8_t>({0xc0, 0xae})));
+35
View File
@@ -2505,6 +2505,41 @@ TEST_CASE("Universal Binary JSON Specification Examples 1")
CHECK(json::to_ubjson(j) == v);
CHECK(json::from_ubjson(v) == j);
}
SECTION("ill-formed UTF-8 (see #5529, #5651)")
{
// none of the binary format specs requires a decoder to reject
// ill-formed UTF-8 in a text string, so a value whose bytes are
// not valid UTF-8 (0xC0 0xAE is an overlong encoding of '.')
// round-trips byte for byte as a string value; to_ubjson() is
// strict, so such a value cannot be written back
const std::vector<uint8_t> v = {'S', 'i', 2, 0xc0, 0xae};
json j;
CHECK_NOTHROW(j = json::from_ubjson(v));
REQUIRE(j.is_string());
CHECK(j.get_ref<const json::string_t&>() == std::string("\xc0\xae"));
CHECK_THROWS_AS(j.dump(), json::type_error&);
CHECK_THROWS_AS(json::to_ubjson(j), json::type_error&);
// the same bytes as an object key round-trip as well
const std::vector<uint8_t> v_key = {'{', 'i', 2, 0xc0, 0xae, 'i', 1, '}'};
json j_key;
CHECK_NOTHROW(j_key = json::from_ubjson(v_key));
REQUIRE(j_key.is_object());
CHECK(j_key.contains(std::string("\xc0\xae")));
CHECK_THROWS_AS(json::to_ubjson(j_key), json::type_error&);
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xFF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
// a truncated multi-byte sequence
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC3")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC3", json::type_error&);
// an encoded surrogate half (U+D800)
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xED\xA0\x80")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xED", json::type_error&);
// an overlong encoding of '.'
CHECK_THROWS_WITH_AS(json::to_ubjson(json("\xC0\xAF")), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xC0", json::type_error&);
// an object key with ill-formed UTF-8 is rejected the same way
CHECK_THROWS_WITH_AS(json::to_ubjson(json{{"\xFF", 1}}), "[json.exception.type_error.316] invalid UTF-8 byte at index 0: 0xFF", json::type_error&);
}
}
SECTION("Array Type")