List the UTF-8 validators instead of calling the DFA the only one

The documentation of decode() called the Hoehrmann DFA the single
source of truth for UTF-8 validation. It is used only by the serializer
and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text
strings). The lexer's scan_string() switch, validate_one_utf8() /
valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer)
and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges
on their own.

Replace the sentence with a list of the four validators, what each is
used for, and a note that they must accept the same sequences. Sharing
code between them was considered and dropped: it would save a few lines
in a validator that is entangled with BON8 pushback, and #5677 is
editing the BON8 byte path.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-30 09:48:09 +02:00
parent 5e41bf284a
commit b3753fbb1a
2 changed files with 34 additions and 10 deletions
+17 -5
View File
@@ -51,11 +51,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details. http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
This decoder is the single source of truth for UTF-8 validation in this The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
library: it is used both by the serializer (to escape and, in strict mode, places, which differ in speed, diagnostics, and how they read the input:
reject ill-formed UTF-8 when dumping a string) and by the binary readers
(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at - decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
decode time; see @ref is_valid_utf8 below). strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
text strings at decode time).
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
bulk string scan, the bulk path of the BON8 reader, and the BON8 writer.
They must accept exactly what the lexer's switch accepts.
- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk
access, and the bytes the bulk path leaves to it.
All four must accept the same set of sequences, so a change to one needs a
matching change to the others.
@param[in,out] state the current decoder state @param[in,out] state the current decoder state
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT) @param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
+17 -5
View File
@@ -6249,11 +6249,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details. http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
This decoder is the single source of truth for UTF-8 validation in this The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
library: it is used both by the serializer (to escape and, in strict mode, places, which differ in speed, diagnostics, and how they read the input:
reject ill-formed UTF-8 when dumping a string) and by the binary readers
(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at - decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
decode time; see @ref is_valid_utf8 below). strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
text strings at decode time).
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
bulk string scan, the bulk path of the BON8 reader, and the BON8 writer.
They must accept exactly what the lexer's switch accepts.
- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk
access, and the bytes the bulk path leaves to it.
All four must accept the same set of sequences, so a change to one needs a
matching change to the others.
@param[in,out] state the current decoder state @param[in,out] state the current decoder state
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT) @param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)