This commit is contained in:
nlohmann
2026-10-04 15:46:44 +00:00
parent 19f538472d
commit c51ca275d0
374 changed files with 1998 additions and 551 deletions
+28 -4
View File
@@ -4,15 +4,31 @@
enum class error_handler_t {
strict,
replace,
ignore
ignore,
keep
};
```
This enumeration is used in the [`dump`](dump.md) function to choose how to treat decoding errors while serializing a
`basic_json` value. Three values are differentiated:
This enumeration is used to choose how to treat ill-formed UTF-8 in a string value or object key:
- [`dump`](dump.md) uses it while serializing a `basic_json` value to text.
- [`to_cbor`](to_cbor.md), [`to_msgpack`](to_msgpack.md), [`to_ubjson`](to_ubjson.md), [`to_bjdata`](to_bjdata.md),
and [`to_bson`](to_bson.md) use it while serializing a `basic_json` value to that binary format. Their default is
`keep`, as no binary writer checked before this parameter was added. CBOR, UBJSON, BJData, and BSON require valid
UTF-8, so for these four the default is `strict` if [`JSON_STRICT_BINARY_UTF8`](../macros/json_strict_binary_utf8.md)
is enabled; MessagePack's specification explicitly allows a string to contain ill-formed UTF-8, so `to_msgpack`
stays at `keep`. `to_bon8` does not take this parameter: BON8 always validates, since UTF-8 lead bytes are
structural to that format.
- [`from_cbor`](from_cbor.md), [`from_msgpack`](from_msgpack.md), [`from_ubjson`](from_ubjson.md),
[`from_bjdata`](from_bjdata.md), and [`from_bson`](from_bson.md) use it while parsing that binary format, to decide
whether to check a string value or object key for well-formed UTF-8 at all; by default (`keep`) they do not, as no
binary reader did before this parameter was added. `from_bon8` does not take this parameter, for the same reason
`to_bon8` does not.
Four values are differentiated:
strict
: throw a `type_error` exception in case of invalid UTF-8
: throw a `type_error`/`parse_error` exception in case of invalid UTF-8
replace
: replace invalid UTF-8 sequences with U+FFFD (� REPLACEMENT CHARACTER)
@@ -20,6 +36,12 @@ replace
ignore
: ignore invalid UTF-8 sequences; all valid bytes are copied to the output unchanged, and invalid bytes are dropped
keep
: keep invalid UTF-8 sequences unchanged; only meaningful for the binary formats mentioned above, since [`dump`]
(dump.md) itself must produce text, and `keep` there writes the ill-formed bytes to the output as is, so the
result is then not valid UTF-8 (but still equals the input bytes exactly, including around any well-formed
characters, which are still escaped as usual)
## Examples
??? example
@@ -45,3 +67,5 @@ ignore
## Version history
- Added in version 3.4.0.
- Added `keep`, and made this enumeration apply to the binary readers and writers in addition to `dump`, in version
3.13.0.