mirror of
https://github.com/nlohmann/json.git
synced 2026-10-06 06:30:31 +00:00
Validate non-ASCII strings with SSSE3 on every x86-64 CPU that has it
The vector UTF-8 check of json_view needed SSSE3 at compile time (JSON_VIEW_USE_SSSE3 with -mssse3), so default x86-64 builds validated non-ASCII text one sequence at a time. The check is now compiled for SSSE3 with a function attribute (GCC 4.9 and later, Clang; MSVC compiles the intrinsics anyway) and used where CPUID reports SSSE3. The answer is kept in an atomic that is initialized at compile time, so neither a guard nor a global constructor is needed. The definitions do not depend on compiler flags, so there is no ODR issue. JSON_VIEW_USE_SSSE3 now only skips the CPU check. On x86-64 Linux (Haswell), twitter.json parses 23% faster with Clang 18 and 25-34% faster with GCC 13, now ahead of yyjson. Also always inline read_eight_bytes() and parse_eight_digits(): GCC called both in the number loops of the lexer and of json_view (52 call sites), which cost about 10% on citm_catalog.json at -O2. Document that reusing a document with read() avoids the page faults of a fresh node index (about 40% of a 55 MB parse on x86-64 Linux). Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
@@ -49,6 +49,11 @@ whether or not the new parse succeeds; take fresh views from [`root()`](root.md)
|
||||
`input` is borrowed or owned by the same rules as [`parse()`](parse.md#notes); a document can borrow on one call and
|
||||
own on the next, since ownership is decided freshly each time.
|
||||
|
||||
Reusing a document matters most for large inputs: the operating system provides the memory of a fresh node index one
|
||||
page at a time, and every page costs a page fault the first time it is written. On x86-64 Linux (4 KiB pages), parsing
|
||||
a 55 MB document into a reused document took about 40 % less time than parsing it into a fresh one. Programs that parse
|
||||
many documents of similar size should therefore keep one document and call `read()`.
|
||||
|
||||
## Examples
|
||||
|
||||
??? example
|
||||
|
||||
@@ -5,13 +5,17 @@
|
||||
```
|
||||
|
||||
When defined on x86-64, the parser of [`basic_json_document`](../basic_json_document/index.md)
|
||||
(`<nlohmann/json_view.hpp>`) validates non-ASCII text in strings with SSSE3, 16 bytes at a time, using the "lookup4"
|
||||
algorithm of [simdjson](https://github.com/simdjson/simdjson). Without it, non-ASCII text is validated one UTF-8
|
||||
sequence at a time on x86-64; on AArch64, the vector check uses NEON and is always on.
|
||||
(`<nlohmann/json_view.hpp>`) validates non-ASCII text in strings with SSSE3 without asking the CPU first.
|
||||
|
||||
SSSE3 is not part of the x86-64 baseline, so the code must be compiled for it: define the macro only together with a
|
||||
compiler option that enables SSSE3 (e.g. `-mssse3`, or `-march=` with a CPU that has it), and only for programs that
|
||||
run on such CPUs. The same input is accepted or rejected either way; only the speed of non-ASCII text differs.
|
||||
By default, the parser checks once at run time whether the CPU has SSSE3 (all x86-64 CPUs since about 2011 have it)
|
||||
and then validates non-ASCII text 16 bytes at a time, using the "lookup4" algorithm of
|
||||
[simdjson](https://github.com/simdjson/simdjson); on CPUs without SSSE3, it validates one UTF-8 sequence at a time.
|
||||
The vector check is compiled for SSSE3 with a function attribute (GCC 4.9 and later, Clang), so this needs no compiler
|
||||
option. With MSVC, the check uses `__cpuid`. On AArch64, the vector check uses NEON and is always on.
|
||||
|
||||
Define the macro only together with a compiler option that enables SSSE3 (e.g. `-mssse3`, or `-march=` with a CPU that
|
||||
has it), and only for programs that run on such CPUs. It saves the check of the CPU, which costs little. The same
|
||||
input is accepted or rejected either way; only the speed of non-ASCII text differs.
|
||||
|
||||
!!! warning "Define consistently"
|
||||
|
||||
|
||||
Reference in New Issue
Block a user