Follow each binary format's UTF-8 rule: strict writers (CBOR/UBJSON/BJData/BSON), lenient readers (#5741)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann authored and GitHub committed 2026-10-04 12:13:48 +02:00
1 parent 40021f38fb
commit f56b418c56
32 files changed
+651 -128

No files matched your search

@@ -63,6 +63,13 @@ The library uses the following mapping from JSON values types to BJData types ac
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
BJData strings must use UTF-8 encoding. By default, `to_bjdata()` writes the bytes of string values and object keys
unchanged, even if they are not valid UTF-8. If
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
!!! info "Unused BJData markers"
The following markers are not used in the conversion:
@@ -208,6 +215,15 @@ The library maps BJData types to JSON value types as follows:
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
BJData strings must use UTF-8 encoding, but this is not enforced on read: `from_bjdata()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_bjdata()` writes such a value
back unchanged (see above).
!!! info "Round trips"
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
@@ -109,14 +109,16 @@ The library maps BSON record types to JSON value types as follows:
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
`from_bson()`.
!!! warning "UTF-8 validation of string values"
!!! warning "Ill-formed UTF-8 in string values"
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
bytes of every such string at decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
(key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
byte-by-byte as a C string, or are not required to hold text, respectively.
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
decoder. `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged.
However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is
passed that replaces or ignores the ill-formed bytes. By default, `to_bson()` writes such a string value or element
(key) name unchanged; if [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it
throws the same exception instead. Element (key) names are never validated on read, since they are read byte-by-byte
as a C string. `binary` values (type `0x05`) are unaffected, since they are not required to hold text.
??? example "Example: deserialize a JSON value from BSON"
@@ -189,15 +189,16 @@ The library maps CBOR types to JSON value types as follows:
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
!!! warning "UTF-8 validation of text strings"
!!! warning "Ill-formed UTF-8 in text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
decode time and rejects ill-formed UTF-8 with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
text.
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings (major
type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this. This library does not:
`from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back
unchanged. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is
passed that replaces or ignores the ill-formed bytes. By default, `to_cbor()` writes such a value back unchanged; if
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws the same exception
instead. Byte strings (major type 2) are unaffected, since they are not required to hold text.
!!! warning "Tagged items"
@@ -153,14 +153,15 @@ The library maps MessagePack types to JSON value types as follows:
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values"
!!! warning "Ill-formed UTF-8 in string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
This library validates the bytes of every such string (object keys included) at decode time and rejects
ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
to hold text.
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
This library follows that: `from_msgpack()` reads `str` bytes (object keys included) as-is, without validating
them, and `to_msgpack()` writes them back as-is, so such a value round-trips through `from_msgpack(to_msgpack(j))`
byte for byte. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way, unless an
error handler is passed that replaces or ignores the ill-formed bytes.
??? example "Example: deserialize a JSON value from MessagePack"
@@ -47,6 +47,13 @@ The library uses the following mapping from JSON values types to UBJSON types ac
- strings with more than 9223372036854775807 bytes (theoretical)
!!! warning "UTF-8 validation of string values and object keys"
UBJSON's required string encoding is UTF-8. By default, `to_ubjson()` writes the bytes of string values and object
keys unchanged, even if they are not valid UTF-8. If
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
!!! info "Unused UBJSON markers"
The following markers are not used in the conversion:
@@ -120,6 +127,15 @@ The library maps UBJSON types to JSON value types as follows:
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
!!! warning "Ill-formed UTF-8 in string values and object keys"
UBJSON's required string encoding is UTF-8, but this is not enforced on read: `from_ubjson()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_ubjson()` writes such a value
back unchanged (see above).
??? example "Example: deserialize a JSON value from UBJSON"
```cpp
+14
View File
@@ -138,6 +138,20 @@ using the library with compilers that do not fully support C++11 and may only wo
See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](../api/macros/json_skip_unsupported_compiler_check.md).
## `JSON_STRICT_BINARY_UTF8`
When defined to `1`, [`to_cbor`](../api/basic_json/to_cbor.md), [`to_ubjson`](../api/basic_json/to_ubjson.md),
[`to_bjdata`](../api/basic_json/to_bjdata.md), and [`to_bson`](../api/basic_json/to_bson.md) throw
[`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) for a string value or object key that is not
valid UTF-8. The default value is `0`, which writes the bytes unchanged as before version 3.13.0; this is planned to
become the default in version 4.0.0.
The check can also be enabled with the CMake option
[`JSON_StrictBinaryUTF8`](../integration/cmake.md#json_strictbinaryutf8) (`OFF` by default) which sets
`JSON_STRICT_BINARY_UTF8` accordingly.
See [full documentation of `JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md).
## `JSON_STRICT_NUL_HANDLING`
When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other
+1
View File
@@ -20,6 +20,7 @@ The complete default namespace name is derived as follows:
`_bics`.
- [`JSON_PRECISE_STREAM_POSITION`](../api/macros/json_precise_stream_position.md) defined non-zero appends `_psp`.
- [`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md) defined non-zero appends `_snul`.
- [`JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md) defined non-zero appends `_sbu8`.
- The inline namespace ends with the suffix `_v` followed by the 3 components of the version number separated by
underscores. To omit the version component, see [Disabling the version component](#disabling-the-version-component)
below.