mirror of
https://github.com/nlohmann/json.git
synced 2026-10-10 16:37:14 +00:00
Follow each binary format's UTF-8 rule: strict writers (CBOR/UBJSON/BJData/BSON), lenient readers (#5741)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
1 parent
40021f38fb
commit
f56b418c56
32 files changed
+651
-128
No files matched your search
@@ -63,6 +63,13 @@ The library uses the following mapping from JSON values types to BJData types ac
|
||||
|
||||
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
|
||||
|
||||
!!! warning "UTF-8 validation of string values and object keys"
|
||||
|
||||
BJData strings must use UTF-8 encoding. By default, `to_bjdata()` writes the bytes of string values and object keys
|
||||
unchanged, even if they are not valid UTF-8. If
|
||||
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
|
||||
|
||||
!!! info "Unused BJData markers"
|
||||
|
||||
The following markers are not used in the conversion:
|
||||
@@ -208,6 +215,15 @@ The library maps BJData types to JSON value types as follows:
|
||||
|
||||
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
|
||||
|
||||
!!! warning "Ill-formed UTF-8 in string values and object keys"
|
||||
|
||||
BJData strings must use UTF-8 encoding, but this is not enforced on read: `from_bjdata()` accepts a string
|
||||
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
|
||||
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
|
||||
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_bjdata()` writes such a value
|
||||
back unchanged (see above).
|
||||
|
||||
!!! info "Round trips"
|
||||
|
||||
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
|
||||
|
||||
@@ -109,14 +109,16 @@ The library maps BSON record types to JSON value types as follows:
|
||||
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
|
||||
`from_bson()`.
|
||||
|
||||
!!! warning "UTF-8 validation of string values"
|
||||
!!! warning "Ill-formed UTF-8 in string values"
|
||||
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
|
||||
bytes of every such string at decode time and rejects ill-formed UTF-8 with a
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
|
||||
set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
|
||||
(key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
|
||||
byte-by-byte as a C string, or are not required to hold text, respectively.
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
|
||||
decoder. `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged.
|
||||
However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is
|
||||
passed that replaces or ignores the ill-formed bytes. By default, `to_bson()` writes such a string value or element
|
||||
(key) name unchanged; if [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it
|
||||
throws the same exception instead. Element (key) names are never validated on read, since they are read byte-by-byte
|
||||
as a C string. `binary` values (type `0x05`) are unaffected, since they are not required to hold text.
|
||||
|
||||
??? example "Example: deserialize a JSON value from BSON"
|
||||
|
||||
|
||||
@@ -189,15 +189,16 @@ The library maps CBOR types to JSON value types as follows:
|
||||
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
|
||||
general-purpose CBOR library instead.
|
||||
|
||||
!!! warning "UTF-8 validation of text strings"
|
||||
!!! warning "Ill-formed UTF-8 in text strings"
|
||||
|
||||
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
|
||||
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
|
||||
decode time and rejects ill-formed UTF-8 with a
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
|
||||
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
|
||||
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
|
||||
text.
|
||||
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings (major
|
||||
type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this. This library does not:
|
||||
`from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back
|
||||
unchanged. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error handler is
|
||||
passed that replaces or ignores the ill-formed bytes. By default, `to_cbor()` writes such a value back unchanged; if
|
||||
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws the same exception
|
||||
instead. Byte strings (major type 2) are unaffected, since they are not required to hold text.
|
||||
|
||||
!!! warning "Tagged items"
|
||||
|
||||
|
||||
@@ -153,14 +153,15 @@ The library maps MessagePack types to JSON value types as follows:
|
||||
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
|
||||
on. Such input needs a general-purpose MessagePack library instead.
|
||||
|
||||
!!! warning "UTF-8 validation of string values"
|
||||
!!! warning "Ill-formed UTF-8 in string values"
|
||||
|
||||
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
|
||||
This library validates the bytes of every such string (object keys included) at decode time and rejects
|
||||
ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
|
||||
with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
|
||||
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
|
||||
to hold text.
|
||||
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
|
||||
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
|
||||
This library follows that: `from_msgpack()` reads `str` bytes (object keys included) as-is, without validating
|
||||
them, and `to_msgpack()` writes them back as-is, so such a value round-trips through `from_msgpack(to_msgpack(j))`
|
||||
byte for byte. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way, unless an
|
||||
error handler is passed that replaces or ignores the ill-formed bytes.
|
||||
|
||||
??? example "Example: deserialize a JSON value from MessagePack"
|
||||
|
||||
|
||||
@@ -47,6 +47,13 @@ The library uses the following mapping from JSON values types to UBJSON types ac
|
||||
|
||||
- strings with more than 9223372036854775807 bytes (theoretical)
|
||||
|
||||
!!! warning "UTF-8 validation of string values and object keys"
|
||||
|
||||
UBJSON's required string encoding is UTF-8. By default, `to_ubjson()` writes the bytes of string values and object
|
||||
keys unchanged, even if they are not valid UTF-8. If
|
||||
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled, it throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead.
|
||||
|
||||
!!! info "Unused UBJSON markers"
|
||||
|
||||
The following markers are not used in the conversion:
|
||||
@@ -120,6 +127,15 @@ The library maps UBJSON types to JSON value types as follows:
|
||||
|
||||
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
|
||||
|
||||
!!! warning "Ill-formed UTF-8 in string values and object keys"
|
||||
|
||||
UBJSON's required string encoding is UTF-8, but this is not enforced on read: `from_ubjson()` accepts a string
|
||||
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
|
||||
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
|
||||
handler is passed that replaces or ignores the ill-formed bytes. By default, `to_ubjson()` writes such a value
|
||||
back unchanged (see above).
|
||||
|
||||
??? example "Example: deserialize a JSON value from UBJSON"
|
||||
|
||||
```cpp
|
||||
|
||||
@@ -138,6 +138,20 @@ using the library with compilers that do not fully support C++11 and may only wo
|
||||
|
||||
See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](../api/macros/json_skip_unsupported_compiler_check.md).
|
||||
|
||||
## `JSON_STRICT_BINARY_UTF8`
|
||||
|
||||
When defined to `1`, [`to_cbor`](../api/basic_json/to_cbor.md), [`to_ubjson`](../api/basic_json/to_ubjson.md),
|
||||
[`to_bjdata`](../api/basic_json/to_bjdata.md), and [`to_bson`](../api/basic_json/to_bson.md) throw
|
||||
[`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) for a string value or object key that is not
|
||||
valid UTF-8. The default value is `0`, which writes the bytes unchanged as before version 3.13.0; this is planned to
|
||||
become the default in version 4.0.0.
|
||||
|
||||
The check can also be enabled with the CMake option
|
||||
[`JSON_StrictBinaryUTF8`](../integration/cmake.md#json_strictbinaryutf8) (`OFF` by default) which sets
|
||||
`JSON_STRICT_BINARY_UTF8` accordingly.
|
||||
|
||||
See [full documentation of `JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md).
|
||||
|
||||
## `JSON_STRICT_NUL_HANDLING`
|
||||
|
||||
When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other
|
||||
|
||||
@@ -20,6 +20,7 @@ The complete default namespace name is derived as follows:
|
||||
`_bics`.
|
||||
- [`JSON_PRECISE_STREAM_POSITION`](../api/macros/json_precise_stream_position.md) defined non-zero appends `_psp`.
|
||||
- [`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md) defined non-zero appends `_snul`.
|
||||
- [`JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md) defined non-zero appends `_sbu8`.
|
||||
- The inline namespace ends with the suffix `_v` followed by the 3 components of the version number separated by
|
||||
underscores. To omit the version component, see [Disabling the version component](#disabling-the-version-component)
|
||||
below.
|
||||
|
||||
Reference in new issue
Block a user