mirror of
https://github.com/nlohmann/json.git
synced 2026-10-05 14:10:31 +00:00
deploy: d8d47be4a5
This commit is contained in:
File diff suppressed because one or more lines are too long
+31
-7
@@ -16,14 +16,15 @@ before including the `json.hpp` header.
|
||||
|
||||
## Function with runtime assertions
|
||||
|
||||
### Unchecked object access to a const value
|
||||
### Unchecked access to a const value
|
||||
|
||||
Function [`operator[]`](../api/basic_json/operator%5B%5D.md) implements unchecked access for objects. Whereas a missing
|
||||
key is added in the case of non-const objects, accessing a const object with a missing key is undefined behavior (think
|
||||
of a dereferenced null pointer) and yields a runtime assertion.
|
||||
Function [`operator[]`](../api/basic_json/operator%5B%5D.md) implements unchecked access for arrays and objects. Whereas
|
||||
a missing element is added in the case of non-const values, accessing a const value with a missing object key or an
|
||||
invalid array index is undefined behavior (think of a dereferenced null pointer) and yields a runtime assertion. This
|
||||
also applies to a [JSON pointer](json_pointer.md) that refers to a missing key or an invalid index.
|
||||
|
||||
If you are not sure whether an element in an object exists, use checked access with the
|
||||
[`at` function](../api/basic_json/at.md) or call the [`contains` function](../api/basic_json/contains.md) before.
|
||||
If you are not sure whether an element exists, use checked access with the [`at` function](../api/basic_json/at.md)
|
||||
or call the [`contains` function](../api/basic_json/contains.md) before.
|
||||
|
||||
See also the documentation on [element access](element_access/index.md).
|
||||
|
||||
@@ -46,7 +47,30 @@ See also the documentation on [element access](element_access/index.md).
|
||||
Output:
|
||||
|
||||
```
|
||||
Assertion failed: (m_value.object->find(key) != m_value.object->end()), function operator[], file json.hpp, line 2144.
|
||||
Assertion failed: (it != m_data.m_value.object->end()), function operator[], file json.hpp, line 28795.
|
||||
```
|
||||
|
||||
??? example "Example 2: Invalid array index in a JSON pointer"
|
||||
|
||||
The following code will trigger an assertion at runtime:
|
||||
|
||||
```cpp
|
||||
#include <nlohmann/json.hpp>
|
||||
|
||||
using json = nlohmann::json;
|
||||
using namespace nlohmann::literals;
|
||||
|
||||
int main()
|
||||
{
|
||||
const json j = {{"array", {1, 2, 3}}};
|
||||
auto v = j["/array/5"_json_pointer];
|
||||
}
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
```
|
||||
Assertion failed: (idx < m_data.m_value.array->size()), function operator[], file json.hpp, line 28758.
|
||||
```
|
||||
|
||||
### Constructing from an uninitialized iterator range
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -12,11 +12,11 @@ The behavior of runtime assertions can be changed by defining macro [`JSON_ASSER
|
||||
|
||||
## Function with runtime assertions
|
||||
|
||||
### Unchecked object access to a const value
|
||||
### Unchecked access to a const value
|
||||
|
||||
Function [`operator[]`](https://json.nlohmann.me/api/basic_json/operator%5B%5D/index.md) implements unchecked access for objects. Whereas a missing key is added in the case of non-const objects, accessing a const object with a missing key is undefined behavior (think of a dereferenced null pointer) and yields a runtime assertion.
|
||||
Function [`operator[]`](https://json.nlohmann.me/api/basic_json/operator%5B%5D/index.md) implements unchecked access for arrays and objects. Whereas a missing element is added in the case of non-const values, accessing a const value with a missing object key or an invalid array index is undefined behavior (think of a dereferenced null pointer) and yields a runtime assertion. This also applies to a [JSON pointer](https://json.nlohmann.me/features/json_pointer/index.md) that refers to a missing key or an invalid index.
|
||||
|
||||
If you are not sure whether an element in an object exists, use checked access with the [`at` function](https://json.nlohmann.me/api/basic_json/at/index.md) or call the [`contains` function](https://json.nlohmann.me/api/basic_json/contains/index.md) before.
|
||||
If you are not sure whether an element exists, use checked access with the [`at` function](https://json.nlohmann.me/api/basic_json/at/index.md) or call the [`contains` function](https://json.nlohmann.me/api/basic_json/contains/index.md) before.
|
||||
|
||||
See also the documentation on [element access](https://json.nlohmann.me/features/element_access/index.md).
|
||||
|
||||
@@ -39,7 +39,30 @@ int main()
|
||||
Output:
|
||||
|
||||
```
|
||||
Assertion failed: (m_value.object->find(key) != m_value.object->end()), function operator[], file json.hpp, line 2144.
|
||||
Assertion failed: (it != m_data.m_value.object->end()), function operator[], file json.hpp, line 28795.
|
||||
```
|
||||
|
||||
Example 2: Invalid array index in a JSON pointer
|
||||
|
||||
The following code will trigger an assertion at runtime:
|
||||
|
||||
```
|
||||
#include <nlohmann/json.hpp>
|
||||
|
||||
using json = nlohmann::json;
|
||||
using namespace nlohmann::literals;
|
||||
|
||||
int main()
|
||||
{
|
||||
const json j = {{"array", {1, 2, 3}}};
|
||||
auto v = j["/array/5"_json_pointer];
|
||||
}
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
```
|
||||
Assertion failed: (idx < m_data.m_value.array->size()), function operator[], file json.hpp, line 28758.
|
||||
```
|
||||
|
||||
### Constructing from an uninitialized iterator range
|
||||
|
||||
@@ -63,6 +63,15 @@ The library uses the following mapping from JSON values types to BJData types ac
|
||||
|
||||
- strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
|
||||
|
||||
!!! warning "UTF-8 validation of string values and object keys"
|
||||
|
||||
BJData strings must use UTF-8 encoding. By default (the [`error_handler`](../../api/basic_json/to_bjdata.md)
|
||||
parameter left at `keep`), `to_bjdata()` writes the bytes of string values and object keys unchanged, even if they
|
||||
are not valid UTF-8. With `error_handler_t::strict`, it throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead;
|
||||
`replace`/`ignore` sanitize the string. [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md)
|
||||
makes `strict` the default.
|
||||
|
||||
!!! info "Unused BJData markers"
|
||||
|
||||
The following markers are not used in the conversion:
|
||||
@@ -208,6 +217,19 @@ The library maps BJData types to JSON value types as follows:
|
||||
|
||||
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
|
||||
|
||||
!!! warning "Ill-formed UTF-8 in string values and object keys"
|
||||
|
||||
BJData strings must use UTF-8 encoding, but checking it on read is opt-in: with the
|
||||
[`error_handler`](../../api/basic_json/from_bjdata.md) parameter left at `keep` (the default), `from_bjdata()`
|
||||
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
|
||||
`error_handler_t::strict` makes `from_bjdata()` check and throw
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
|
||||
`replace`/`ignore` sanitize the string instead of keeping it. However,
|
||||
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
|
||||
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()`'s
|
||||
own `error_handler` parameter defaults to `keep` (see above), so such a value is written back unchanged.
|
||||
|
||||
!!! info "Round trips"
|
||||
|
||||
A value returned by [`from_bjdata`](../../api/basic_json/from_bjdata.md) can be serialized with
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -54,6 +54,10 @@ The following values can **not** be converted to a BJData value:
|
||||
|
||||
- strings with more than 18446744073709551615 bytes, i.e., 264-1 bytes (theoretical)
|
||||
|
||||
UTF-8 validation of string values and object keys
|
||||
|
||||
BJData strings must use UTF-8 encoding. By default (the [`error_handler`](https://json.nlohmann.me/api/basic_json/to_bjdata/index.md) parameter left at `keep`), `to_bjdata()` writes the bytes of string values and object keys unchanged, even if they are not valid UTF-8. With `error_handler_t::strict`, it throws [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for ill-formed UTF-8 instead; `replace`/`ignore` sanitize the string. [`JSON_STRICT_BINARY_UTF8`](https://json.nlohmann.me/api/macros/json_strict_binary_utf8/index.md) makes `strict` the default.
|
||||
|
||||
Unused BJData markers
|
||||
|
||||
The following markers are not used in the conversion:
|
||||
@@ -230,6 +234,10 @@ Complete mapping
|
||||
|
||||
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
|
||||
|
||||
Ill-formed UTF-8 in string values and object keys
|
||||
|
||||
BJData strings must use UTF-8 encoding, but checking it on read is opt-in: with the [`error_handler`](https://json.nlohmann.me/api/basic_json/from_bjdata/index.md) parameter left at `keep` (the default), `from_bjdata()` accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing `error_handler_t::strict` makes `from_bjdata()` check and throw [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) for ill-formed UTF-8, and `replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](https://json.nlohmann.me/api/basic_json/dump/index.md) still requires valid UTF-8 and throws [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()`'s own `error_handler` parameter defaults to `keep` (see above), so such a value is written back unchanged.
|
||||
|
||||
Round trips
|
||||
|
||||
A value returned by [`from_bjdata`](https://json.nlohmann.me/api/basic_json/from_bjdata/index.md) can be serialized with [`to_bjdata`](https://json.nlohmann.me/api/basic_json/to_bjdata/index.md) using any combination of options and parsed back into an equal value, and serializing that value again with the same options produces the same bytes. The exception is binary values: they are only written as an optimized binary array (`[$B`) if Draft 3 is enabled and both `use_size` and `use_type` are set. Otherwise, they are written as arrays of integers and parsed back as such (see the notes on binary values above), and serializing such an array again may choose different, but equally valid, type markers. The bytes can then differ, but parsing them again yields the same value.
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -109,14 +109,21 @@ The library maps BSON record types to JSON value types as follows:
|
||||
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
|
||||
`from_bson()`.
|
||||
|
||||
!!! warning "UTF-8 validation of string values"
|
||||
!!! warning "Ill-formed UTF-8 in string values"
|
||||
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
|
||||
bytes of every such string at decode time and rejects ill-formed UTF-8 with a
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
|
||||
set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
|
||||
(key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read
|
||||
byte-by-byte as a C string, or are not required to hold text, respectively.
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
|
||||
decoder, so checking is opt-in: with the [`error_handler`](../../api/basic_json/from_bson.md) parameter left at
|
||||
`keep` (the default), `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back
|
||||
unchanged. Passing `error_handler_t::strict` makes `from_bson()` check and throw
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
|
||||
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
|
||||
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a
|
||||
value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the
|
||||
ill-formed bytes. `to_bson()`'s own `error_handler` parameter defaults to `keep`, so such a string value or element
|
||||
(key) name is written unchanged; with `strict` (the default if
|
||||
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled), it throws the same exception
|
||||
instead. Element (key) names are never validated on read, since they are read byte-by-byte as a C string. `binary`
|
||||
values (type `0x05`) are unaffected, since they are not required to hold text.
|
||||
|
||||
??? example "Example: deserialize a JSON value from BSON"
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -118,9 +118,9 @@ The BSON reader is lenient in a few areas where the BSON specification is more r
|
||||
|
||||
If BSON input must be validated for strict specification compliance, validate it separately before passing it to `from_bson()`.
|
||||
|
||||
UTF-8 validation of string values
|
||||
Ill-formed UTF-8 in string values
|
||||
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the bytes of every such string at decode time and rejects ill-formed UTF-8 with a [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element (key) names and `binary` values (type `0x05`) are unaffected and are never validated, since they are read byte-by-byte as a C string, or are not required to hold text, respectively.
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a decoder, so checking is opt-in: with the [`error_handler`](https://json.nlohmann.me/api/basic_json/from_bson/index.md) parameter left at `keep` (the default), `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged. Passing `error_handler_t::strict` makes `from_bson()` check and throw [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) for ill-formed UTF-8, and `replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](https://json.nlohmann.me/api/basic_json/dump/index.md) still requires valid UTF-8 and throws [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_bson()`'s own `error_handler` parameter defaults to `keep`, so such a string value or element (key) name is written unchanged; with `strict` (the default if [`JSON_STRICT_BINARY_UTF8`](https://json.nlohmann.me/api/macros/json_strict_binary_utf8/index.md) is enabled), it throws the same exception instead. Element (key) names are never validated on read, since they are read byte-by-byte as a C string. `binary` values (type `0x05`) are unaffected, since they are not required to hold text.
|
||||
|
||||
Example: deserialize a JSON value from BSON
|
||||
|
||||
|
||||
@@ -189,15 +189,21 @@ The library maps CBOR types to JSON value types as follows:
|
||||
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
|
||||
general-purpose CBOR library instead.
|
||||
|
||||
!!! warning "UTF-8 validation of text strings"
|
||||
!!! warning "Ill-formed UTF-8 in text strings"
|
||||
|
||||
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
|
||||
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
|
||||
decode time and rejects ill-formed UTF-8 with a
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
|
||||
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
|
||||
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
|
||||
text.
|
||||
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings (major
|
||||
type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this, so checking is opt-in: with the
|
||||
[`error_handler`](../../api/basic_json/from_cbor.md) parameter left at `keep` (the default), `from_cbor()` accepts a
|
||||
text string (object keys included) whose bytes are not valid UTF-8 and hands them back unchanged. Passing
|
||||
`error_handler_t::strict` makes `from_cbor()` check and throw
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
|
||||
`replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](../../api/basic_json/dump.md)
|
||||
still requires valid UTF-8 and throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a
|
||||
value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the
|
||||
ill-formed bytes. `to_cbor()`'s own [`error_handler`](../../api/basic_json/to_cbor.md) parameter defaults to `keep`,
|
||||
so such a value is written back unchanged; with `strict` (the default if
|
||||
[`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md) is enabled), it throws the same exception
|
||||
instead. Byte strings (major type 2) are unaffected, since they are not required to hold text.
|
||||
|
||||
!!! warning "Tagged items"
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -193,9 +193,9 @@ CBOR allows map keys of any type, whereas JSON only allows strings as keys in ob
|
||||
|
||||
This applies to the [SAX interface](https://json.nlohmann.me/features/parsing/sax_interface/index.md) as well, as the key is read before it is passed on. This is a deliberate restriction of the library's JSON value model, not an oversight: formats built on CBOR maps with integer keys, such as COSE ([RFC 9052](https://www.rfc-editor.org/rfc/rfc9052.html)) or CWT ([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a general-purpose CBOR library instead.
|
||||
|
||||
UTF-8 validation of text strings
|
||||
Ill-formed UTF-8 in text strings
|
||||
|
||||
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings (major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at decode time and rejects ill-formed UTF-8 with a [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold text.
|
||||
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings (major type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this, so checking is opt-in: with the [`error_handler`](https://json.nlohmann.me/api/basic_json/from_cbor/index.md) parameter left at `keep` (the default), `from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them back unchanged. Passing `error_handler_t::strict` makes `from_cbor()` check and throw [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) for ill-formed UTF-8, and `replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](https://json.nlohmann.me/api/basic_json/dump/index.md) still requires valid UTF-8 and throws [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_cbor()`'s own [`error_handler`](https://json.nlohmann.me/api/basic_json/to_cbor/index.md) parameter defaults to `keep`, so such a value is written back unchanged; with `strict` (the default if [`JSON_STRICT_BINARY_UTF8`](https://json.nlohmann.me/api/macros/json_strict_binary_utf8/index.md) is enabled), it throws the same exception instead. Byte strings (major type 2) are unaffected, since they are not required to hold text.
|
||||
|
||||
Tagged items
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -153,14 +153,23 @@ The library maps MessagePack types to JSON value types as follows:
|
||||
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
|
||||
on. Such input needs a general-purpose MessagePack library instead.
|
||||
|
||||
!!! warning "UTF-8 validation of string values"
|
||||
!!! warning "Ill-formed UTF-8 in string values"
|
||||
|
||||
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
|
||||
This library validates the bytes of every such string (object keys included) at decode time and rejects
|
||||
ill-formed UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or,
|
||||
with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting
|
||||
value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required
|
||||
to hold text.
|
||||
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain
|
||||
a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged.
|
||||
This library follows that by default: with its
|
||||
[`error_handler`](../../api/basic_json/from_msgpack.md) parameter left at `keep` (the default),
|
||||
`from_msgpack()` reads `str` bytes (object keys included) as-is, without validating them, so such a value
|
||||
round-trips through `from_msgpack(to_msgpack(j))` byte for byte. Passing `error_handler_t::strict` makes
|
||||
`from_msgpack()` check anyway and throw
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
|
||||
`replace`/`ignore` sanitize the string instead of keeping it. `to_msgpack()` also writes `str` bytes as-is by
|
||||
default, since the specification permits it; its [`error_handler`](../../api/basic_json/to_msgpack.md) parameter
|
||||
can be set to `strict` to throw [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) instead, or
|
||||
to `replace`/`ignore` to sanitize the string, for instance for a decoder that rejects ill-formed UTF-8. However,
|
||||
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read this way with the
|
||||
default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes.
|
||||
|
||||
??? example "Example: deserialize a JSON value from MessagePack"
|
||||
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -162,9 +162,9 @@ MessagePack allows map keys of any type, whereas JSON only allows strings as key
|
||||
|
||||
This applies to the [SAX interface](https://json.nlohmann.me/features/parsing/sax_interface/index.md) as well, as the key is read before it is passed on. Such input needs a general-purpose MessagePack library instead.
|
||||
|
||||
UTF-8 validation of string values
|
||||
Ill-formed UTF-8 in string values
|
||||
|
||||
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8. This library validates the bytes of every such string (object keys included) at decode time and rejects ill-formed UTF-8 with a [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. `bin`/`ext`/`fixext` values are unaffected and are never validated, since they are not required to hold text.
|
||||
The MessagePack specification explicitly allows a `str` value (`fixstr`, `str 8`, `str 16`, `str 32`) to contain a byte sequence that is not valid UTF-8, and expects a deserializer to hand the original bytes back unchanged. This library follows that by default: with its [`error_handler`](https://json.nlohmann.me/api/basic_json/from_msgpack/index.md) parameter left at `keep` (the default), `from_msgpack()` reads `str` bytes (object keys included) as-is, without validating them, so such a value round-trips through `from_msgpack(to_msgpack(j))` byte for byte. Passing `error_handler_t::strict` makes `from_msgpack()` check anyway and throw [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) for ill-formed UTF-8, and `replace`/`ignore` sanitize the string instead of keeping it. `to_msgpack()` also writes `str` bytes as-is by default, since the specification permits it; its [`error_handler`](https://json.nlohmann.me/api/basic_json/to_msgpack/index.md) parameter can be set to `strict` to throw [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) instead, or to `replace`/`ignore` to sanitize the string, for instance for a decoder that rejects ill-formed UTF-8. However, [`dump()`](https://json.nlohmann.me/api/basic_json/dump/index.md) still requires valid UTF-8 and throws [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for a value read this way with the default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes.
|
||||
|
||||
Example: deserialize a JSON value from MessagePack
|
||||
|
||||
|
||||
@@ -47,6 +47,15 @@ The library uses the following mapping from JSON values types to UBJSON types ac
|
||||
|
||||
- strings with more than 9223372036854775807 bytes (theoretical)
|
||||
|
||||
!!! warning "UTF-8 validation of string values and object keys"
|
||||
|
||||
UBJSON's required string encoding is UTF-8. By default (the [`error_handler`](../../api/basic_json/to_ubjson.md)
|
||||
parameter left at `keep`), `to_ubjson()` writes the bytes of string values and object keys unchanged, even if they
|
||||
are not valid UTF-8. With `error_handler_t::strict`, it throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 instead;
|
||||
`replace`/`ignore` sanitize the string. [`JSON_STRICT_BINARY_UTF8`](../../api/macros/json_strict_binary_utf8.md)
|
||||
makes `strict` the default.
|
||||
|
||||
!!! info "Unused UBJSON markers"
|
||||
|
||||
The following markers are not used in the conversion:
|
||||
@@ -120,6 +129,19 @@ The library maps UBJSON types to JSON value types as follows:
|
||||
|
||||
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
|
||||
|
||||
!!! warning "Ill-formed UTF-8 in string values and object keys"
|
||||
|
||||
UBJSON's required string encoding is UTF-8, but checking it on read is opt-in: with the
|
||||
[`error_handler`](../../api/basic_json/from_ubjson.md) parameter left at `keep` (the default), `from_ubjson()`
|
||||
accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing
|
||||
`error_handler_t::strict` makes `from_ubjson()` check and throw
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) for ill-formed UTF-8, and
|
||||
`replace`/`ignore` sanitize the string instead of keeping it. However,
|
||||
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for a value read with the default
|
||||
`keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()`'s
|
||||
own `error_handler` parameter defaults to `keep` (see above), so such a value is written back unchanged.
|
||||
|
||||
??? example "Example: deserialize a JSON value from UBJSON"
|
||||
|
||||
```cpp
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -46,6 +46,10 @@ The following values can **not** be converted to a UBJSON value:
|
||||
|
||||
- strings with more than 9223372036854775807 bytes (theoretical)
|
||||
|
||||
UTF-8 validation of string values and object keys
|
||||
|
||||
UBJSON's required string encoding is UTF-8. By default (the [`error_handler`](https://json.nlohmann.me/api/basic_json/to_ubjson/index.md) parameter left at `keep`), `to_ubjson()` writes the bytes of string values and object keys unchanged, even if they are not valid UTF-8. With `error_handler_t::strict`, it throws [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for ill-formed UTF-8 instead; `replace`/`ignore` sanitize the string. [`JSON_STRICT_BINARY_UTF8`](https://json.nlohmann.me/api/macros/json_strict_binary_utf8/index.md) makes `strict` the default.
|
||||
|
||||
Unused UBJSON markers
|
||||
|
||||
The following markers are not used in the conversion:
|
||||
@@ -173,6 +177,10 @@ Complete mapping
|
||||
|
||||
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
|
||||
|
||||
Ill-formed UTF-8 in string values and object keys
|
||||
|
||||
UBJSON's required string encoding is UTF-8, but checking it on read is opt-in: with the [`error_handler`](https://json.nlohmann.me/api/basic_json/from_ubjson/index.md) parameter left at `keep` (the default), `from_ubjson()` accepts a string value or object key whose bytes are not valid UTF-8 and hands them back unchanged. Passing `error_handler_t::strict` makes `from_ubjson()` check and throw [`parse_error.113`](https://json.nlohmann.me/home/exceptions/#jsonexceptionparse_error113) for ill-formed UTF-8, and `replace`/`ignore` sanitize the string instead of keeping it. However, [`dump()`](https://json.nlohmann.me/api/basic_json/dump/index.md) still requires valid UTF-8 and throws [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for a value read with the default `keep` handler, unless an error handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()`'s own `error_handler` parameter defaults to `keep` (see above), so such a value is written back unchanged.
|
||||
|
||||
Example: deserialize a JSON value from UBJSON
|
||||
|
||||
```
|
||||
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+1
-1
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -138,6 +138,20 @@ using the library with compilers that do not fully support C++11 and may only wo
|
||||
|
||||
See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](../api/macros/json_skip_unsupported_compiler_check.md).
|
||||
|
||||
## `JSON_STRICT_BINARY_UTF8`
|
||||
|
||||
When defined to `1`, [`to_cbor`](../api/basic_json/to_cbor.md), [`to_ubjson`](../api/basic_json/to_ubjson.md),
|
||||
[`to_bjdata`](../api/basic_json/to_bjdata.md), and [`to_bson`](../api/basic_json/to_bson.md) throw
|
||||
[`type_error.316`](../home/exceptions.md#jsonexceptiontype_error316) for a string value or object key that is not
|
||||
valid UTF-8. The default value is `0`, which writes the bytes unchanged as before version 3.13.0; this is planned to
|
||||
become the default in version 4.0.0.
|
||||
|
||||
The check can also be enabled with the CMake option
|
||||
[`JSON_StrictBinaryUTF8`](../integration/cmake.md#json_strictbinaryutf8) (`OFF` by default) which sets
|
||||
`JSON_STRICT_BINARY_UTF8` accordingly.
|
||||
|
||||
See [full documentation of `JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md).
|
||||
|
||||
## `JSON_STRICT_NUL_HANDLING`
|
||||
|
||||
When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -104,6 +104,14 @@ When defined, the library will not create a compile error when a known unsupport
|
||||
|
||||
See [full documentation of `JSON_SKIP_UNSUPPORTED_COMPILER_CHECK`](https://json.nlohmann.me/api/macros/json_skip_unsupported_compiler_check/index.md).
|
||||
|
||||
## `JSON_STRICT_BINARY_UTF8`
|
||||
|
||||
When defined to `1`, [`to_cbor`](https://json.nlohmann.me/api/basic_json/to_cbor/index.md), [`to_ubjson`](https://json.nlohmann.me/api/basic_json/to_ubjson/index.md), [`to_bjdata`](https://json.nlohmann.me/api/basic_json/to_bjdata/index.md), and [`to_bson`](https://json.nlohmann.me/api/basic_json/to_bson/index.md) throw [`type_error.316`](https://json.nlohmann.me/home/exceptions/#jsonexceptiontype_error316) for a string value or object key that is not valid UTF-8. The default value is `0`, which writes the bytes unchanged as before version 3.13.0 unreleased; this is planned to become the default in version 4.0.0.
|
||||
|
||||
The check can also be enabled with the CMake option [`JSON_StrictBinaryUTF8`](https://json.nlohmann.me/integration/cmake/#json_strictbinaryutf8) (`OFF` by default) which sets `JSON_STRICT_BINARY_UTF8` accordingly.
|
||||
|
||||
See [full documentation of `JSON_STRICT_BINARY_UTF8`](https://json.nlohmann.me/api/macros/json_strict_binary_utf8/index.md).
|
||||
|
||||
## `JSON_STRICT_NUL_HANDLING`
|
||||
|
||||
When defined to `1`, a `'\0'` (NUL) byte anywhere in the input is rejected with `parse_error.101`, like any other unexpected byte, instead of being silently treated as end of input (see the [FAQ entry](https://json.nlohmann.me/home/faq/#nul-bytes-in-the-input) for background). The default value is `0`, which preserves the existing behavior; this is planned to become the default in version 4.0.0.
|
||||
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -20,6 +20,7 @@ The complete default namespace name is derived as follows:
|
||||
`_bics`.
|
||||
- [`JSON_PRECISE_STREAM_POSITION`](../api/macros/json_precise_stream_position.md) defined non-zero appends `_psp`.
|
||||
- [`JSON_STRICT_NUL_HANDLING`](../api/macros/json_strict_nul_handling.md) defined non-zero appends `_snul`.
|
||||
- [`JSON_STRICT_BINARY_UTF8`](../api/macros/json_strict_binary_utf8.md) defined non-zero appends `_sbu8`.
|
||||
- The inline namespace ends with the suffix `_v` followed by the 3 components of the version number separated by
|
||||
underscores. To omit the version component, see [Disabling the version component](#disabling-the-version-component)
|
||||
below.
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -14,6 +14,7 @@ The complete default namespace name is derived as follows:
|
||||
- [`JSON_BRACE_INIT_COPY_SEMANTICS`](https://json.nlohmann.me/api/macros/json_brace_init_copy_semantics/index.md) defined non-zero appends `_bics`.
|
||||
- [`JSON_PRECISE_STREAM_POSITION`](https://json.nlohmann.me/api/macros/json_precise_stream_position/index.md) defined non-zero appends `_psp`.
|
||||
- [`JSON_STRICT_NUL_HANDLING`](https://json.nlohmann.me/api/macros/json_strict_nul_handling/index.md) defined non-zero appends `_snul`.
|
||||
- [`JSON_STRICT_BINARY_UTF8`](https://json.nlohmann.me/api/macros/json_strict_binary_utf8/index.md) defined non-zero appends `_sbu8`.
|
||||
- The inline namespace ends with the suffix `_v` followed by the 3 components of the version number separated by underscores. To omit the version component, see [Disabling the version component](#disabling-the-version-component) below.
|
||||
|
||||
For example, the namespace name for version 3.11.2 with `JSON_DIAGNOSTICS` defined to `1` is:
|
||||
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -180,7 +180,8 @@ int main()
|
||||
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::replace)
|
||||
<< "\nstring with ignored invalid characters: "
|
||||
<< j_invalid.dump(-1, ' ', false, json::error_handler_t::ignore)
|
||||
<< '\n';
|
||||
<< "\nstring with the invalid byte kept as is (" << j_invalid.dump(-1, ' ', false, json::error_handler_t::keep).size()
|
||||
<< " bytes, not valid UTF-8 itself)\n";
|
||||
}
|
||||
```
|
||||
|
||||
@@ -190,6 +191,7 @@ Output:
|
||||
[json.exception.type_error.316] invalid UTF-8 byte at index 2: 0xA9
|
||||
string with replaced invalid characters: "ä�ü"
|
||||
string with ignored invalid characters: "äü"
|
||||
string with the invalid byte kept as is (7 bytes, not valid UTF-8 itself)
|
||||
```
|
||||
|
||||
Avoiding invalid UTF-8
|
||||
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user