Add json_document and json_view: node index, parser, and document

Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).

A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.

The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.

basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.

detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.

A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
Niels Lohmann committed 2026-10-06 10:48:06 +02:00
1 parent 17842694f6
commit 4303ba674d
146 files changed
+9611 -13

No files matched your search

+7
View File
@@ -38,3 +38,10 @@ target_compile_features(json_benchmarks PRIVATE cxx_std_11)
target_link_libraries(json_benchmarks benchmark ${CMAKE_THREAD_LIBS_INIT})
add_dependencies(json_benchmarks download_test_data)
target_include_directories(json_benchmarks PRIVATE ${JSON_BENCHMARK_INCLUDE_DIR} ${CMAKE_BINARY_DIR}/include)
# the json_document benchmarks, unless the header to benchmark predates json_view.hpp
if(EXISTS "${JSON_BENCHMARK_INCLUDE_DIR}/nlohmann/json_view.hpp")
target_sources(json_benchmarks PRIVATE src/benchmarks_view.cpp)
else()
message(STATUS "${JSON_BENCHMARK_INCLUDE_DIR} has no nlohmann/json_view.hpp: skipping the json_document benchmarks")
endif()
+8 -1
View File
@@ -9,6 +9,7 @@ Micro-benchmarks for parsing, serialization and the binary formats, written with
| benchmark | what it does |
|---|---|
| `ParseFile`, `ParseString` | parse JSON from a file stream or a string |
| `Accept` | validate JSON from a string (`json::accept`) |
| `ParseIndented` | parse the large files re-indented by 4 spaces, for the lexer's whitespace handling |
| `Dump` | serialize, compact (`-`) and indented (`4`) |
| `ToCbor`, `BinaryToCbor` | write CBOR; `BinaryToCbor` writes binary values of growing size |
@@ -16,6 +17,10 @@ Micro-benchmarks for parsing, serialization and the binary formats, written with
| `FromBinaryBuffer`, `FromBinaryFile` | read CBOR, MessagePack, UBJSON, BJData and BSON from a buffer or a `FILE*` |
| `FromBinaryShape` | read deeply nested, container-heavy and scalar-heavy documents in every binary format |
| `FromCborChunkedString` | read CBOR strings split into indefinite-length chunks |
| `ViewParse`, `ViewRead` | parse with [`json_document`](https://json.nlohmann.me/features/json_view/) into a new document, or into one that is reused; compare with `ParseString` |
| `ViewParseIndented` | as `ParseIndented`, with a reused `json_document` |
| `ViewAccept` | validate with `json_document::accept`; compare with `Accept` |
| `ViewMaterialize` | convert a parsed `json_document` into a `json` value |
The input files are those of [nativejson-benchmark](https://github.com/miloyip/nativejson-benchmark) (`canada`,
`citm_catalog`, `twitter`), a large `jeopardy` file, and number-heavy files (`floats`, `signed_ints`, ...).
@@ -110,7 +115,9 @@ Python, unpinned versions work as well. In its output:
- `OVERALL_GEOMEAN` summarizes all benchmarks;
- `-a` shows only the aggregates, not every repetition.
The header you compare with must support everything the benchmarks use. The current benchmarks build against 3.12.0.
The header you compare with must support everything the benchmarks use. The current benchmarks build against 3.12.0;
the `View*` benchmarks (in `src/benchmarks_view.cpp`) are only built if the directory also holds
`nlohmann/json_view.hpp`.
Only benchmarks present in both result files are compared, so for older releases, either filter the benchmarks or
build that release's own `tests/benchmarks` against its own header.
+25
View File
@@ -81,6 +81,31 @@ BENCHMARK_CAPTURE(ParseString, signed_ints, TEST_DATA_DIRECTORY "/regressi
BENCHMARK_CAPTURE(ParseString, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json");
BENCHMARK_CAPTURE(ParseString, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json");
//////////////////////////////////////////////////////////////////////////////
// validate JSON from string
//////////////////////////////////////////////////////////////////////////////
static void Accept(benchmark::State& state, const char* filename)
{
std::ifstream f(filename);
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
while (state.KeepRunning())
{
benchmark::DoNotOptimize(json::accept(str));
}
state.SetBytesProcessed(state.iterations() * str.size());
}
BENCHMARK_CAPTURE(Accept, jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json");
BENCHMARK_CAPTURE(Accept, canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json");
BENCHMARK_CAPTURE(Accept, citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json");
BENCHMARK_CAPTURE(Accept, twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json");
BENCHMARK_CAPTURE(Accept, floats, TEST_DATA_DIRECTORY "/regression/floats.json");
BENCHMARK_CAPTURE(Accept, signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json");
BENCHMARK_CAPTURE(Accept, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json");
BENCHMARK_CAPTURE(Accept, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json");
//////////////////////////////////////////////////////////////////////////////
// parse pretty-printed JSON from string
//
+146
View File
@@ -0,0 +1,146 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Benchmarks of json_document (nlohmann/json_view.hpp). They use the inputs of
// ParseString and ParseIndented in benchmarks.cpp, so that each ViewParse row
// can be read against the json::parse row of the same file.
#include <benchmark/benchmark.h>
#include <nlohmann/json_view.hpp>
#include <fstream>
#include <iterator>
#include <string>
#include <test_data.hpp>
using json = nlohmann::json;
using json_document = nlohmann::json_document;
static std::string read_file(const char* filename)
{
std::ifstream f(filename, std::ios::binary);
return std::string((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
}
#define JSON_VIEW_BENCHMARK_FILES(fn) \
BENCHMARK_CAPTURE(fn, jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json"); \
BENCHMARK_CAPTURE(fn, canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json"); \
BENCHMARK_CAPTURE(fn, citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json"); \
BENCHMARK_CAPTURE(fn, twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json"); \
BENCHMARK_CAPTURE(fn, floats, TEST_DATA_DIRECTORY "/regression/floats.json"); \
BENCHMARK_CAPTURE(fn, signed_ints, TEST_DATA_DIRECTORY "/regression/signed_ints.json"); \
BENCHMARK_CAPTURE(fn, unsigned_ints, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json"); \
BENCHMARK_CAPTURE(fn, small_signed_ints, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json")
//////////////////////////////////////////////////////////////////////////////
// parse into a new document (compare with ParseString)
//////////////////////////////////////////////////////////////////////////////
static void ViewParse(benchmark::State& state, const char* filename)
{
const std::string str = read_file(filename);
while (state.KeepRunning())
{
state.PauseTiming();
auto* d = new json_document();
state.ResumeTiming();
*d = json_document::parse(str);
state.PauseTiming();
delete d;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * str.size());
}
JSON_VIEW_BENCHMARK_FILES(ViewParse);
//////////////////////////////////////////////////////////////////////////////
// parse into a document that is reused (its memory stays allocated)
//////////////////////////////////////////////////////////////////////////////
static void ViewRead(benchmark::State& state, const char* filename)
{
const std::string str = read_file(filename);
json_document d;
while (state.KeepRunning())
{
d.read(str);
benchmark::DoNotOptimize(d);
}
state.SetBytesProcessed(state.iterations() * str.size());
}
JSON_VIEW_BENCHMARK_FILES(ViewRead);
//////////////////////////////////////////////////////////////////////////////
// parse pretty-printed JSON (compare with ParseIndented)
//////////////////////////////////////////////////////////////////////////////
static void ViewParseIndented(benchmark::State& state, const char* filename, int indent)
{
const std::string indented = json::parse(read_file(filename)).dump(indent);
json_document d;
while (state.KeepRunning())
{
d.read(indented);
benchmark::DoNotOptimize(d);
}
state.SetBytesProcessed(state.iterations() * indented.size());
}
BENCHMARK_CAPTURE(ViewParseIndented, jeopardy / 4, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", 4);
BENCHMARK_CAPTURE(ViewParseIndented, canada / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", 4);
BENCHMARK_CAPTURE(ViewParseIndented, citm_catalog / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", 4);
BENCHMARK_CAPTURE(ViewParseIndented, twitter / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", 4);
//////////////////////////////////////////////////////////////////////////////
// validate only (compare with Accept)
//////////////////////////////////////////////////////////////////////////////
static void ViewAccept(benchmark::State& state, const char* filename)
{
const std::string str = read_file(filename);
while (state.KeepRunning())
{
benchmark::DoNotOptimize(json_document::accept(str));
}
state.SetBytesProcessed(state.iterations() * str.size());
}
JSON_VIEW_BENCHMARK_FILES(ViewAccept);
//////////////////////////////////////////////////////////////////////////////
// convert a parsed document into a json value
//////////////////////////////////////////////////////////////////////////////
static void ViewMaterialize(benchmark::State& state, const char* filename)
{
const std::string str = read_file(filename);
const json_document d = json_document::parse(str);
while (state.KeepRunning())
{
state.PauseTiming();
auto* j = new json();
state.ResumeTiming();
*j = d.root().materialize();
state.PauseTiming();
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * str.size());
}
JSON_VIEW_BENCHMARK_FILES(ViewMaterialize);