The vector UTF-8 check of json_view needed SSSE3 at compile time
(JSON_VIEW_USE_SSSE3 with -mssse3), so default x86-64 builds validated
non-ASCII text one sequence at a time. The check is now compiled for
SSSE3 with a function attribute (GCC 4.9 and later, Clang; MSVC compiles
the intrinsics anyway) and used where CPUID reports SSSE3. The answer is
kept in an atomic that is initialized at compile time, so neither a
guard nor a global constructor is needed. The definitions do not depend
on compiler flags, so there is no ODR issue. JSON_VIEW_USE_SSSE3 now only
skips the CPU check.
On x86-64 Linux (Haswell), twitter.json parses 23% faster with Clang 18
and 25-34% faster with GCC 13, now ahead of yyjson.
Also always inline read_eight_bytes() and parse_eight_digits(): GCC
called both in the number loops of the lexer and of json_view (52 call
sites), which cost about 10% on citm_catalog.json at -O2.
Document that reusing a document with read() avoids the page faults of
a fresh node index (about 40% of a 55 MB parse on x86-64 Linux).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.
Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.
json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.
Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>