Share the code point to UTF-8 encoding between the wide-string helpers and the lexer

The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three
times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32
code point, in the UTF-16 helper for both a BMP code unit and a valid
surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The
copies had drifted: the UTF-32 helper masked the leading bits of each
byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift
alone, even though both give the same result for a code point that is
already known to be in range.

Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder
that invokes a callable once per output byte, most significant byte
first. Use it in the three valid-code-point branches (UTF-32 code
points up to U+10FFFF, UTF-16 code units outside the surrogate range,
and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where
out forwards to add(). The UTF-16 helper's deliberate pass-through of
malformed surrogate units and the UTF-32 helper's 0xFF sentinel for
code points above U+10FFFF are untouched, since neither reaches the new
helper.

Behavior-preserving: same bytes in the same order for every valid code
point, verified with unit-class_lexer, unit-class_parser,
unit-deserialization, unit-wstring and the non-test-data parts of
unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse
microbenchmark that shows no change (about 73 ms either way, median of
3, 1M escape sequences). single_include/ regenerated with make
amalgamate; make check-amalgamation leaves a clean tree.

Overlaps #5704, which rewrites the wide_string_input_helper
specializations touched here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5712 item 6
This commit is contained in:
Niels Lohmann
2026-09-30 18:15:03 +02:00
parent 32d5540d37
commit 4356d0fe50
4 changed files with 166 additions and 144 deletions
+55
View File
@@ -36,6 +36,61 @@ StringType to_string(std::size_t value)
return result;
}
///////////////////
// UTF-8 encoding //
///////////////////
/*!
@brief encode a Unicode code point as UTF-8
Used to turn a decoded code point back into bytes: by the wide-string input
adapters in input_adapters.hpp (one code point per UTF-32 unit, per UTF-16
unit outside the surrogate range, and per valid UTF-16 surrogate pair), and
by the lexer's `\uXXXX`/`\uXXXX\uYYYY` handling in lexer.hpp. Passing a
code point above U+10FFFF, or one in the surrogate range U+D800..U+DFFF, is
undefined behavior; callers are expected to have rejected those already
(the wide-string adapters pass malformed units through unencoded instead of
calling this function, and the lexer rejects unpaired surrogates before
reaching it).
@tparam Out a callable invoked with one byte (as std::uint32_t, 0x00..0xFF)
at a time, most significant byte first
@param[in] cp the code point to encode (at most U+10FFFF)
@param[in] out called once for each byte of the UTF-8 encoding of @a cp
*/
template<typename Out>
void encode_utf8(std::uint32_t cp, Out&& out)
{
JSON_ASSERT(cp <= 0x10FFFF);
if (cp < 0x80)
{
// 1-byte characters: 0xxxxxxx (ASCII)
out(cp);
}
else if (cp <= 0x7FF)
{
// 2-byte characters: 110xxxxx 10xxxxxx
out(0xC0u | (cp >> 6u));
out(0x80u | (cp & 0x3Fu));
}
else if (cp <= 0xFFFF)
{
// 3-byte characters: 1110xxxx 10xxxxxx 10xxxxxx
out(0xE0u | (cp >> 12u));
out(0x80u | ((cp >> 6u) & 0x3Fu));
out(0x80u | (cp & 0x3Fu));
}
else
{
// 4-byte characters: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
out(0xF0u | (cp >> 18u));
out(0x80u | ((cp >> 12u) & 0x3Fu));
out(0x80u | ((cp >> 6u) & 0x3Fu));
out(0x80u | (cp & 0x3Fu));
}
}
///////////////////
// UTF-8 decoding //
///////////////////