mirror of
https://github.com/nlohmann/json.git
synced 2026-10-03 05:00:30 +00:00
Accept ill-formed UTF-8 in all binary readers again
RFC 8949 and the MessagePack/BSON/UBJSON/BJData specs leave UTF-8 well-formedness checking up to the decoder, so following #5529 the binary readers are lenient by default again, as in release 3.12.0 (the reader-side check was added by #5185/#5531, not in any release); reader-side validation becomes opt-in in a follow-up PR. The writers stay strict and throw type_error.316 for ill-formed UTF-8. BON8 is unchanged, since UTF-8 lead bytes are structural there. Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
@@ -214,12 +214,14 @@ The library maps BJData types to JSON value types as follows:
|
||||
|
||||
The mapping is **complete** in the sense that any BJData value can be converted to a JSON value.
|
||||
|
||||
!!! warning "UTF-8 validation of string values and object keys"
|
||||
!!! warning "Ill-formed UTF-8 in string values and object keys"
|
||||
|
||||
This library validates the bytes of every string value and object key at decode time and rejects ill-formed
|
||||
UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
|
||||
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value
|
||||
is dumped.
|
||||
BJData strings must use UTF-8 encoding, but this is not enforced on read: `from_bjdata()` accepts a string
|
||||
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
|
||||
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
|
||||
handler is passed that replaces or ignores the ill-formed bytes. `to_bjdata()` is strict as well (see above), so
|
||||
a value read this way cannot be written back to BJData.
|
||||
|
||||
!!! info "Round trips"
|
||||
|
||||
|
||||
@@ -109,18 +109,17 @@ The library maps BSON record types to JSON value types as follows:
|
||||
If BSON input must be validated for strict specification compliance, validate it separately before passing it to
|
||||
`from_bson()`.
|
||||
|
||||
!!! warning "UTF-8 validation of string values"
|
||||
!!! warning "Ill-formed UTF-8 in string values"
|
||||
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8. This library validates the
|
||||
bytes of every such string at decode time and rejects ill-formed UTF-8 with a
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions`
|
||||
set to `false`, a discarded value), rather than only failing later when the resulting value is dumped. Element
|
||||
(key) names and `binary` values (type `0x05`) are unaffected and are never validated on read, since they are read
|
||||
byte-by-byte as a C string, or are not required to hold text, respectively. `to_bson()` validates both string
|
||||
values and element names and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8 in either, so an
|
||||
object with such a key or value cannot be produced in the first place, even though `from_bson()` would accept it
|
||||
from another source.
|
||||
The BSON specification requires `string` values (type `0x02`) to be valid UTF-8, but this is not required of a
|
||||
decoder. `from_bson()` accepts a `string` value whose bytes are not valid UTF-8 and hands them back unchanged.
|
||||
However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
|
||||
handler is passed that replaces or ignores the ill-formed bytes. `to_bson()` is strict as well and throws the
|
||||
same exception for a string value or element (key) name that is not valid UTF-8, so an object with such a key
|
||||
or value cannot be produced in the first place, even though `from_bson()` would accept it from another source.
|
||||
Element (key) names are never validated on read, since they are read byte-by-byte as a C string. `binary`
|
||||
values (type `0x05`) are unaffected, since they are not required to hold text.
|
||||
|
||||
??? example
|
||||
|
||||
|
||||
@@ -189,17 +189,16 @@ The library maps CBOR types to JSON value types as follows:
|
||||
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
|
||||
general-purpose CBOR library instead.
|
||||
|
||||
!!! warning "UTF-8 validation of text strings"
|
||||
!!! warning "Ill-formed UTF-8 in text strings"
|
||||
|
||||
[RFC 8949, Section 3.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-3.1) requires CBOR text strings
|
||||
(major type 3) to be valid UTF-8. This library validates the bytes of every text string (object keys included) at
|
||||
decode time and rejects ill-formed UTF-8 with a
|
||||
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
|
||||
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value is
|
||||
dumped. Byte strings (major type 2) are unaffected and are never validated, since they are not required to hold
|
||||
text. `to_cbor()` validates string values and object keys the same way and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for ill-formed UTF-8, so a value with
|
||||
such a string cannot be serialized in the first place.
|
||||
(major type 3) to be valid UTF-8, but leaves it up to the decoder whether to enforce this. This library does
|
||||
not: `from_cbor()` accepts a text string (object keys included) whose bytes are not valid UTF-8 and hands them
|
||||
back unchanged. However, [`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
|
||||
handler is passed that replaces or ignores the ill-formed bytes. `to_cbor()` is strict as well and throws the
|
||||
same exception for a string value or object key that is not valid UTF-8, so such a value cannot be written back
|
||||
to CBOR. Byte strings (major type 2) are unaffected, since they are not required to hold text.
|
||||
|
||||
!!! warning "Tagged items"
|
||||
|
||||
|
||||
@@ -126,12 +126,14 @@ The library maps UBJSON types to JSON value types as follows:
|
||||
|
||||
The mapping is **complete** in the sense that any UBJSON value can be converted to a JSON value.
|
||||
|
||||
!!! warning "UTF-8 validation of string values and object keys"
|
||||
!!! warning "Ill-formed UTF-8 in string values and object keys"
|
||||
|
||||
This library validates the bytes of every string value and object key at decode time and rejects ill-formed
|
||||
UTF-8 with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
|
||||
`allow_exceptions` set to `false`, a discarded value), rather than only failing later when the resulting value
|
||||
is dumped.
|
||||
UBJSON's required string encoding is UTF-8, but this is not enforced on read: `from_ubjson()` accepts a string
|
||||
value or object key whose bytes are not valid UTF-8 and hands them back unchanged. However,
|
||||
[`dump()`](../../api/basic_json/dump.md) still requires valid UTF-8 and throws
|
||||
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) for such a value, unless an error
|
||||
handler is passed that replaces or ignores the ill-formed bytes. `to_ubjson()` is strict as well (see above), so
|
||||
a value read this way cannot be written back to UBJSON.
|
||||
|
||||
??? example
|
||||
|
||||
|
||||
@@ -340,8 +340,9 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
|
||||
### json.exception.parse_error.113
|
||||
|
||||
A string could not be read from a [binary format](../features/binary_formats/index.md): either a value that is not a
|
||||
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
|
||||
the string's bytes are not valid UTF-8.
|
||||
string was read where one was required (for instance as a map key), or the string's length specification is invalid.
|
||||
The bytes of a string itself are not checked for valid UTF-8 on read; see the ill-formed UTF-8 notes on the
|
||||
individual [binary format](../features/binary_formats/index.md) pages for how such a string is handled afterward.
|
||||
|
||||
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
|
||||
type (for instance integers or `null`) are therefore not supported; see the notes on
|
||||
@@ -364,9 +365,6 @@ type (for instance integers or `null`) are therefore not supported; see the note
|
||||
```
|
||||
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative
|
||||
```
|
||||
```
|
||||
[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing CBOR string: invalid string: ill-formed UTF-8 byte
|
||||
```
|
||||
|
||||
### json.exception.parse_error.114
|
||||
|
||||
|
||||
Reference in New Issue
Block a user