mirror of
https://github.com/nlohmann/json.git
synced 2026-10-01 04:00:31 +00:00
List the UTF-8 validators instead of calling the DFA the only one
The documentation of decode() called the Hoehrmann DFA the single source of truth for UTF-8 validation. It is used only by the serializer and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text strings). The lexer's scan_string() switch, validate_one_utf8() / valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer) and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges on their own. Replace the sentence with a list of the four validators, what each is used for, and a note that they must accept the same sequences. Sharing code between them was considered and dropped: it would save a few lines in a validator that is entangled with BON8 pushback, and #5677 is editing the BON8 byte path. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
@@ -51,11 +51,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
|
|||||||
written by Björn Hoehrmann. See
|
written by Björn Hoehrmann. See
|
||||||
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
||||||
|
|
||||||
This decoder is the single source of truth for UTF-8 validation in this
|
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
|
||||||
library: it is used both by the serializer (to escape and, in strict mode,
|
places, which differ in speed, diagnostics, and how they read the input:
|
||||||
reject ill-formed UTF-8 when dumping a string) and by the binary readers
|
|
||||||
(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at
|
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
|
||||||
decode time; see @ref is_valid_utf8 below).
|
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
|
||||||
|
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
|
||||||
|
text strings at decode time).
|
||||||
|
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
|
||||||
|
diagnostic for each kind of error.
|
||||||
|
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
|
||||||
|
bulk string scan, the bulk path of the BON8 reader, and the BON8 writer.
|
||||||
|
They must accept exactly what the lexer's switch accepts.
|
||||||
|
- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk
|
||||||
|
access, and the bytes the bulk path leaves to it.
|
||||||
|
|
||||||
|
All four must accept the same set of sequences, so a change to one needs a
|
||||||
|
matching change to the others.
|
||||||
|
|
||||||
@param[in,out] state the current decoder state
|
@param[in,out] state the current decoder state
|
||||||
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
||||||
|
|||||||
@@ -6249,11 +6249,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
|
|||||||
written by Björn Hoehrmann. See
|
written by Björn Hoehrmann. See
|
||||||
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
||||||
|
|
||||||
This decoder is the single source of truth for UTF-8 validation in this
|
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
|
||||||
library: it is used both by the serializer (to escape and, in strict mode,
|
places, which differ in speed, diagnostics, and how they read the input:
|
||||||
reject ill-formed UTF-8 when dumping a string) and by the binary readers
|
|
||||||
(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at
|
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
|
||||||
decode time; see @ref is_valid_utf8 below).
|
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
|
||||||
|
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
|
||||||
|
text strings at decode time).
|
||||||
|
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
|
||||||
|
diagnostic for each kind of error.
|
||||||
|
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
|
||||||
|
bulk string scan, the bulk path of the BON8 reader, and the BON8 writer.
|
||||||
|
They must accept exactly what the lexer's switch accepts.
|
||||||
|
- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk
|
||||||
|
access, and the bytes the bulk path leaves to it.
|
||||||
|
|
||||||
|
All four must accept the same set of sequences, so a change to one needs a
|
||||||
|
matching change to the others.
|
||||||
|
|
||||||
@param[in,out] state the current decoder state
|
@param[in,out] state the current decoder state
|
||||||
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
||||||
|
|||||||
Reference in New Issue
Block a user