Add an error_handler parameter for UTF-8 to the binary readers and writers (#5746)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann authored and GitHub committed 2026-10-04 12:13:49 +02:00
1 parent f56b418c56
commit 73e9eae3c1
30 files changed
+1794 -422

No files matched your search

@@ -65,10 +65,12 @@ The library uses the following mapping from JSON values types to BJData types ac
!!! warning "UTF-8 validation of string values and object keys"
BJData strings must use UTF-8 encoding. By default, `to_bjdata()` writes the bytes of string values and object keys
unchanged, even if they are not valid UTF-8. If
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
BJData strings must use UTF-8 encoding. By default (the [`error_handler`](../../api/basic_json/to_bjdata.md)
parameter left at `keep`), `to_bjdata()` writes the bytes of string values and object keys unchanged, even if they
are not valid UTF-8. With `error_handler_t::strict`, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead;
`replace`/`ignore` sanitize the string. [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md)
makes `strict` the default.
!!! info "Unused BJData markers"
@@ -217,12 +219,16 @@ The library maps BJData types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in string values and object keys"
BJData strings must use UTF-8 encoding, but this is not enforced on read: `from_bjdata()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
BJData strings must use UTF-8 encoding, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_bjdata.md) parameter left at `keep` (the default), `from_bjdata()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_bjdata()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_bjdata()` writes such a value
back unchanged (see above).
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()`'s
own `error_handler` parameter defaults to `keep` (see above), so such a value is written back unchanged.
!!! info "Round trips"
@@ -112,13 +112,18 @@ The library maps BSON record types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in string values"
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
decoder. `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged.
However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is
passed that replaces or ignores the ill-formed bytes. By default, `to_bson()` writes such a string value or element
(key) name unchanged; if [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it
throws the same exception instead. Element (key) names are never validated on read, since they are read byte-by-byte
as a C string. `binary` values (type `0x05`) are unaffected, since they are not required to hold text.
decoder, so checking is opt-in: with the [`error_handler`](../../api/basic_json/from_bson.md) parameter left at
`keep` (the default), `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back
unchanged. Passing `error_handler_t::strict` makes `from_bson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a
value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the
ill-formed bytes. `to_bson()`'s own `error_handler` parameter defaults to `keep`, so such a string value or element
(key) name is written unchanged; with `strict` (the default if
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled), it throws the same exception
instead. Element (key) names are never validated on read, since they are read byte-by-byte as a C string. `binary`
values (type `0x05`) are unaffected, since they are not required to hold text.
??? example "Example: deserialize a JSON value from BSON"
@@ -192,12 +192,17 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in text strings"
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings (major
type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this. This library does not:
`from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back
unchanged. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is
passed that replaces or ignores the ill-formed bytes. By default, `to_cbor()` writes such a value back unchanged; if
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws the same exception
type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this, so checking is opt-in: with the
[`error_handler`](../../api/basic_json/from_cbor.md) parameter left at `keep` (the default), `from_cbor()` accepts a
text string (object keys included) whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_cbor()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a
value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the
ill-formed bytes. `to_cbor()`'s own [`error_handler`](../../api/basic_json/to_cbor.md) parameter defaults to `keep`,
so such a value is written back unchanged; with `strict` (the default if
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled), it throws the same exception
instead. Byte strings (major type 2) are unaffected, since they are not required to hold text.
!!! warning "Tagged items"
@@ -157,11 +157,19 @@ The library maps MessagePack types to JSON value types as follows:
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
This library follows that: `from_msgpack()` reads `str` bytes (object keys included) as-is, without validating
them, and `to_msgpack()` writes them back as-is, so such a value round-trips through `from_msgpack(to_msgpack(j))`
byte for byte. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way, unless an
error handler is passed that replaces or ignores the ill-formed bytes.
This library follows that by default: with its
[`error_handler`](../../api/basic_json/from_msgpack.md) parameter left at `keep` (the default),
`from_msgpack()` reads `str` bytes (object keys included) as-is, without validating them, so such a value
round-trips through `from_msgpack(to_msgpack(j))` byte for byte. Passing `error_handler_t::strict` makes
`from_msgpack()` check anyway and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. `to_msgpack()` also writes `str` bytes as-is by
default, since the specification permits it; its [`error_handler`](../../api/basic_json/to_msgpack.md) parameter
can be set to `strict` to throw [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) instead, or
to `replace`/`ignore` to sanitize the string, for instance for a decoder that rejects ill-formed UTF-8. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way with the
default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes.
??? example "Example: deserialize a JSON value from MessagePack"
@@ -49,10 +49,12 @@ The library uses the following mapping from JSON values types to UBJSON types ac
!!! warning "UTF-8 validation of string values and object keys"
UBJSON's required string encoding is UTF-8. By default, `to_ubjson()` writes the bytes of string values and object
keys unchanged, even if they are not valid UTF-8. If
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
UBJSON's required string encoding is UTF-8. By default (the [`error_handler`](../../api/basic_json/to_ubjson.md)
parameter left at `keep`), `to_ubjson()` writes the bytes of string values and object keys unchanged, even if they
are not valid UTF-8. With `error_handler_t::strict`, it throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead;
`replace`/`ignore` sanitize the string. [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md)
makes `strict` the default.
!!! info "Unused UBJSON markers"
@@ -129,12 +131,16 @@ The library maps UBJSON types to JSON value types as follows:
!!! warning "Ill-formed UTF-8 in string values and object keys"
UBJSON's required string encoding is UTF-8, but this is not enforced on read: `from_ubjson()` accepts a string
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
UBJSON's required string encoding is UTF-8, but checking it on read is opt-in: with the
[`error_handler`](../../api/basic_json/from_ubjson.md) parameter left at `keep` (the default), `from_ubjson()`
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
`error_handler_t::strict` makes `from_ubjson()` check and throw
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
`replace`/`ignore` sanitize the string instead of keeping it. However,
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_ubjson()` writes such a value
back unchanged (see above).
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()`'s
own `error_handler` parameter defaults to `keep` (see above), so such a value is written back unchanged.
??? example "Example: deserialize a JSON value from UBJSON"