- The output string reserves the size estimate (the source extent) and grows
in steps of 64 KiB within it, instead of being resized to the estimate at
once: resize() zero-fills, and for a pretty-printed source the estimate is
far larger than the compact output. citm_catalog dump: 342 -> 283 us on
x86-64, 144 -> 134 us on Apple M1.
- bench_corpus compares the dump with source numbers with yyjson writing
numbers read as raw text (YYJSON_READ_NUMBER_AS_RAW); both write the same
bytes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- compare.py: make --data, --corpus, and --build-dir absolute, since the
benchmarks run in the build directory; download into a .part file and
remove an archive whose SHA-256 does not match, so that an interrupted
download is not kept
- bench_view/bench_corpus/bench_edit: report files that cannot be opened instead of
aborting; run each engine once untimed before its timed call, so that
no engine pays for the allocator cleaning up after the previous one
(with glibc, json_view after json::parse looked 1.7x slower on
citm_catalog traverse); add "simdjson DOM (fresh)" and time
"json_view (reused)" for traverse and select too
- README: explain fresh vs. reused documents and page faults on Linux
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The commands are argument lists the script builds itself, and the
downloads are pinned https URLs whose SHA-256 is checked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
bench_edit.cpp joins the comparison: parse, apply the same logical edits
with each library's own API (a handful at fixed places, or one in every
record), and serialize; all outputs must describe the same value.
compare.py builds and runs it with the other two programs.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A workflow runs compare.py with pinned downloads on GitHub-hosted
Ubuntu runners and shows the results as the job summary and as an
artifact: started by hand (workflow_dispatch: x86-64 or AArch64, GCC or
Clang), or when a pull request gets the label "benchmark" (both
architectures, GCC). The label trigger gives numbers before the
workflow is on the default branch, which workflow_dispatch needs.
Shared runners are noisy, so the numbers show where json_view stands on
another architecture; published numbers still need a quiet machine.
compare.py takes the CPU name from lscpu where /proc/cpuinfo has none
(AArch64 Linux), and falls back to the architecture. Checked in Linux
containers (AArch64, Clang 15 and GCC 9, offline with the pinned
archives).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
compare.py --download now checks the SHA-256 of the yyjson 0.13.0,
simdjson 4.6.11, and Boost 1.92.0 archives. It unpacks each archive
once (Boost's directory is boost_1_92_0) and, where Python supports it,
with the 'data' filter.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
tests/benchmarks/json_view/ holds the comparison with other libraries,
which is not built by CMake or run by CI:
- bench_view.cpp: parse, traverse, select, and dump of twitter,
citm_catalog, canada, jeopardy, a single tweet, and a JSON-RPC request,
with json_view, yyjson, simdjson (DOM and On-Demand), Boost.JSON, and
json::parse; all engines must agree on every document before anything
is timed, and run interleaved in every round
- bench_corpus.cpp: parse, traverse, and dump of any list of files
- compare.py: builds both against include/ with the libraries of the
system (or pinned downloads), runs them, and writes the results with
what is needed to reproduce them (date, commit, CPU, OS, compiler,
flags, library versions) to results/<date>-<host>.md and .csv; only the
Python 3 standard library is used
- README.md: how to run it, what is measured, and which features the
engines have, so the numbers can be read correctly
Boost.JSON is optional (JSON_VIEW_BENCH_BOOST).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>