Do not scan for a copyable run that cannot exist

Under ensure_ascii, dump_escaped() calls find_ascii_copyable_run() at every
character boundary. When the text is dense non-ASCII - CJK, where every byte
is >= 0x80 - the scanner stops on its first byte and returns zero, so its SWAR
block runs once per character and buys nothing, on top of the escaping that
still has to happen afterwards.

A run can only be non-empty when the first byte is one the scanner may copy,
so test that single byte before calling it. Runs that do exist are found
exactly as before, so the bulk-copy win is unchanged; only the calls that were
always going to return zero are skipped.

Output is unchanged: the dump digest over canada/citm/twitter, in compact,
pretty and ensure_ascii form, matches develop byte for byte.

  dump(ensure_ascii=true)   develop    before     after
  CJK text                   3.54ms    4.25ms    3.36ms
  CJK, no ASCII at all       3.09ms    4.02ms    3.02ms
  Latin-1-ish text           4.39ms    3.04ms    2.93ms
  plain ASCII                3.92ms    0.80ms    0.79ms

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann
2026-09-02 22:28:50 +02:00
parent e73bfd0f25
commit 3426a41391
2 changed files with 16 additions and 2 deletions
@@ -844,8 +844,15 @@ class serializer
if (state == UTF8_ACCEPT)
{
const auto* const data = reinterpret_cast<const unsigned char*>(s.data());
// A run can only be non-empty when the very first byte is one
// the scanner may copy, so test that single byte before paying
// for the scan. Without it, text whose characters all have to be
// escaped - CJK under ensure_ascii, where every byte is >= 0x80 -
// runs the scanner once per character only to be told zero.
const std::size_t run = EnsureAscii
? find_ascii_copyable_run(data + i, s.size() - i)
? (is_ascii_copyable(data[i])
? find_ascii_copyable_run(data + i, s.size() - i)
: 0)
: string_bulk_run(data + i, s.size() - i);
if (run != 0)
{