From b3753fbb1aa652a40c4e8d25ec65de57ff984258 Mon Sep 17 00:00:00 2001 From: Niels Lohmann Date: Wed, 30 Sep 2026 09:48:09 +0200 Subject: [PATCH] List the UTF-8 validators instead of calling the DFA the only one The documentation of decode() called the Hoehrmann DFA the single source of truth for UTF-8 validation. It is used only by the serializer and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text strings). The lexer's scan_string() switch, validate_one_utf8() / valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer) and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges on their own. Replace the sentence with a list of the four validators, what each is used for, and a note that they must accept the same sequences. Sharing code between them was considered and dropped: it would save a few lines in a validator that is entangled with BON8 pushback, and #5677 is editing the BON8 byte path. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann --- include/nlohmann/detail/string_utils.hpp | 22 +++++++++++++++++----- single_include/nlohmann/json.hpp | 22 +++++++++++++++++----- 2 files changed, 34 insertions(+), 10 deletions(-) diff --git a/include/nlohmann/detail/string_utils.hpp b/include/nlohmann/detail/string_utils.hpp index 142943cd6..8d9956716 100644 --- a/include/nlohmann/detail/string_utils.hpp +++ b/include/nlohmann/detail/string_utils.hpp @@ -51,11 +51,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally written by Björn Hoehrmann. See http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details. -This decoder is the single source of truth for UTF-8 validation in this -library: it is used both by the serializer (to escape and, in strict mode, -reject ill-formed UTF-8 when dumping a string) and by the binary readers -(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at -decode time; see @ref is_valid_utf8 below). +The library checks UTF-8 well-formedness (RFC 3629, section 4) in four +places, which differ in speed, diagnostics, and how they read the input: + +- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in + strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR, + MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in + text strings at decode time). +- the per-lead-byte switch in lexer::scan_string(): JSON text, with a + diagnostic for each kind of error. +- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's + bulk string scan, the bulk path of the BON8 reader, and the BON8 writer. + They must accept exactly what the lexer's switch accepts. +- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk + access, and the bytes the bulk path leaves to it. + +All four must accept the same set of sequences, so a change to one needs a +matching change to the others. @param[in,out] state the current decoder state @param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT) diff --git a/single_include/nlohmann/json.hpp b/single_include/nlohmann/json.hpp index 6061f3215..fcb8af1ec 100644 --- a/single_include/nlohmann/json.hpp +++ b/single_include/nlohmann/json.hpp @@ -6249,11 +6249,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally written by Björn Hoehrmann. See http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details. -This decoder is the single source of truth for UTF-8 validation in this -library: it is used both by the serializer (to escape and, in strict mode, -reject ill-formed UTF-8 when dumping a string) and by the binary readers -(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at -decode time; see @ref is_valid_utf8 below). +The library checks UTF-8 well-formedness (RFC 3629, section 4) in four +places, which differ in speed, diagnostics, and how they read the input: + +- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in + strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR, + MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in + text strings at decode time). +- the per-lead-byte switch in lexer::scan_string(): JSON text, with a + diagnostic for each kind of error. +- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's + bulk string scan, the bulk path of the BON8 reader, and the BON8 writer. + They must accept exactly what the lexer's switch accepts. +- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk + access, and the bytes the bulk path leaves to it. + +All four must accept the same set of sequences, so a change to one needs a +matching change to the others. @param[in,out] state the current decoder state @param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)