Compare commits

...
Author SHA1 Message Date
Niels Lohmann cbdc502fbf Write compact dumps of json_view without a library call per token
The default dump() (no indentation, no ensure_ascii) gets its own
writer: the same walk and output, with the write position in a local
variable (stores through char pointers would otherwise force a reload
of the buffer's members after each one), strings and number tokens of
the source copied by fixed-size moves of 32 bytes where the source has
that many bytes left (the buffer keeps 64 bytes of slack), and decoded
strings copied in runs up to the next quote, backslash, or control
character. Documents that are not edited are walked through the node
array in order, so that a frame only needs the end of its container,
and integer tokens are read from the source directly. The innermost
open container is kept in local variables, and the stack holds only
the ones around it; the stack starts in a local array of 32 and moves
to the heap only for deeper nesting (its address does not escape, so
its pointers stay in registers). Dumps of shallow documents thus
allocate only the output, whose first size includes the slack, so it
does not grow just before the end.

The long copies are out of line: otherwise, the compiler merges the
fixed-size moves into the same library call.

These techniques come from the prototype; the writer lost them when the
view was split into pull requests, which made dump() 2 to 3 times
slower.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:24 +02:00
Niels Lohmann 300ab011b0 Document how images store the node index of json_view
The node index section on the architecture page says how save() writes
the nodes and that a change of their layout must raise the image
version, and save's format note links to it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:23 +02:00
Niels Lohmann 8d9088ab05 Address the cpplint findings of json_document images
The exponent of the overflow check is an std::int64_t instead of a long
(runtime/int).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:22 +02:00
Niels Lohmann 5a936a2ef3 Document the images of json_documents
Add API pages for basic_json_document::save() and load(), including the
image_check enumeration and its three levels. Add examples that show
caching a parsed document as an image, loading it without parsing, the
difference between a borrowed and an owned image, and a full check
rejecting a damaged image that a bounds check still reads safely.

Add an "Images" section to the json_view feature page, register the new
pages in mkdocs.yml and docSet.sql, group basic_json_document's member
list by parsing/access/images/edits, and document parse_error.116 and
type_error.320 on the exceptions page, extending out_of_range.416 for
images.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:21 +02:00
Niels Lohmann 5876a22712 Add images of json_documents: save() and load()
An image is a document stored so that loading it needs no parsing: a
64-byte header, the nodes, the text, and the decoded strings
(little-endian; version 1).

- save() writes an edited document in its current state, in document
  order (floats that are not finite become null, as in dump()); the
  same document always gives the same bytes
- load(pointer, size) and load(const vector&) borrow the image;
  load(vector&&) keeps it without a copy. The nodes are copied (aligned,
  and editable); the hash indexes of large objects are rebuilt.
- image_check::full checks everything the parser guarantees (structure,
  bounds, UTF-8, strings of the source, number tokens and their values);
  bounds checks structure and bounds, so that reading and serializing
  stay safe; none trusts the image.

A malformed image or a failed check throws the new parse_error.116;
saving a discarded document (or images on a big-endian target) throws
the new type_error.320; images of 4 GiB or more out_of_range.416.

As images checked for bounds only can hold any bytes, the general float
conversion now checks the token's grammar (and locates the point and
the exponent itself), the exponent loop of the layout conversion takes
digits as unsigned, and the serializer validates each non-ASCII sequence
it decodes, throwing what basic_json::dump() throws for invalid UTF-8.
Parsed and edited documents are not affected.

The idea of images comes from zero-copy formats such as FlatBuffers and
YaFF, the check from FlatBuffers' Verifier; no code is taken from them.

Tests: round trips with every check (small documents, test files, large
objects, edited documents with every kind of edit), ownership, all
errors, one corruption per rejection branch of the check, and 12,000
seeded random corruptions, which must be rejected or read safely. The
fuzzer json_view_image_fuzzer uses each input as an image and as a JSON
text.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:20 +02:00
Niels Lohmann cc36e26254 Document insert() and erase() of editable json_documents
Add API reference pages for basic_json_document::insert and
basic_json_document::erase, matching the style of set.md and
push_back.md: signatures, parameters, return values, exception safety,
exceptions with their exact ids and messages, complexity, notes on
duplicate keys and view/iterator validity, and an example.

Add example programs basic_json_document__insert.cpp and
basic_json_document__erase.cpp with their expected output, each
comparing an edit on an editable document with the same edit on a
plain json value to show what is preserved: member order, the
spelling of untouched numbers, and, for insert, that a view taken
before the insert keeps referring to the same element after its index
shifts.

Register both new pages in mkdocs.yml, docSet.sql, and the member list
of basic_json_document/index.md, add cross-references to them from
set.md and push_back.md, and mention insert/erase in the "Editing a
document" section of features/json_view.md.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:19 +02:00
Niels Lohmann 3c5c000408 Address the clang-tidy findings of the structural edit tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:18 +02:00
Niels Lohmann 1e212bc50c Benchmark editing documents with json_view, yyjson, Boost.JSON, and json
bench_edit.cpp joins the comparison: parse, apply the same logical edits
with each library's own API (a handful at fixed places, or one in every
record), and serialize; all outputs must describe the same value.
compare.py builds and runs it with the other two programs.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:17 +02:00
Niels Lohmann 5d25f863c0 Add insert() and erase() to editable documents
- insert(array, index, value): insert before an element (index <= size)
- erase(object, key): remove all members with the key; returns their
  number
- erase(array, index): remove an element
- erase(json_pointer): remove the member or element a pointer names

The errors are those of basic_json (type_error.307/309,
out_of_range.401/403/405). A view of an erased value keeps its last value,
and views of other values keep referring to them when elements move.

Tests: the differential test now also inserts and erases members and
elements, directly and through JSON pointers; plus the errors, views
across inserts and erasures, duplicate keys, and large objects.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:16 +02:00
Niels Lohmann ada8f17230 Document how editable json_documents use the node index
The node index section on the architecture page says how edits are kept
(the edit buffer, moved elements, links, and new nodes), and the feature
page links to it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:15 +02:00
Niels Lohmann 266c2249fe Cover the edit storage in tests
Test edits of an empty document and assignments through a view of a
value that is no longer part of the document; copy the entries of a
block through deref() (one path for links and values); mark the
4 GiB limit and the returns of find_parent() that no document reaches.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:15 +02:00
Niels Lohmann 2e5419eebc Label the pages and examples of the editable aliases
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:14 +02:00
Niels Lohmann 7e577fb2d4 Document editable json_documents
- API pages for set and push_back of basic_json_document and for the
  editable aliases; the class overviews name the Editable parameter
- type_error.319 on the exceptions page
- the feature page explains editing, and the examples show when it
  helps: a configuration file changed without reformatting it

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:13 +02:00
Niels Lohmann baf4ccfda8 Address the clang-tidy findings of the editable documents
Pick the overloads of encode() with a first_true trait instead of
nested conditionals, name the pointer type in the copies of links,
mark the owning pointers of the edit storage, and compare doubles
by their bits in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:11 +02:00
Niels Lohmann 76897a9b9b Add editable documents with set() and push_back()
basic_json_document<BasicJsonType, true> (json_editable_document,
ordered_json_editable_document) can be edited:

- set(view, value): replace a value
- set(object, key, value): assign a member, or add it (a null becomes an
  object); with duplicate keys, the first is assigned and the others go
- set(array, index, value): assign an element
- set(json_pointer, value): the member, element, or ("-", or the size of
  the array) the end of an array a pointer names
- push_back(array, value): append (a null becomes an array)

Values are views (of any document, copied), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written, and the parsed index never moves: new values and element
sequences go to storage owned by the document (edit_storage.hpp), so views
stay valid, and a view keeps referring to its value (after an assignment,
it sees the new one). Read-only documents are unchanged; editing one does
not compile.

Errors are those of basic_json where the operation corresponds
(type_error.305/308, out_of_range.401/403/405, parse_error.106/109); a view
of another document is invalid_iterator.202. Strings are checked for UTF-8
when they enter the document, with the type_error.316 that
basic_json::dump() throws for the same string, so that a document only
holds valid UTF-8. Binary values cannot be stored (the new
type_error.319), and edits of 4 GiB or more end with out_of_range.416.
Views of editable and read-only documents compare with each other.

Tests (unit-json_view_edit.cpp): random assignments, member and element
changes, copies within and between documents, and pushes, applied to an
ordered_json_editable_document and to the ordered_json value; after every
edit both must serialize (also indented and with ensure_ascii),
materialize, compare, and read back the same. Further: the errors, strings
that stay valid while the edit arena grows, numbers (NaN, infinities,
extremes; number_format::source), nulls that become containers, the root
replaced, duplicate keys, values of other documents, large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:11 +02:00
Niels Lohmann 2b1a720fc0 Walk the index of json_view through a navigation policy
Preparation for editable documents, without a change in behavior: views and
documents get a template parameter Editable (false by default), and every
walk over the index (iterators, lookups, dump(), materialize()) goes
through detail::view::navigation<Editable>. For read-only documents it is
the plain node array, as before, so they compile without any of the edit
handling. For editable documents it also follows the representation of
edits, which this commit defines:

- node flags `edited` (a string or number token in the edit arena),
  `moved` (the elements of an array/object live in a separate sequence),
  and `is_new` (no source position), and link nodes (kind_link) that
  stand for a value stored elsewhere
- document_data::edit_state: the moved sequences, the storage of new
  values, and the edit arena

materialize() now keeps a frame per open container instead of returning
to the end of a closed one, as the serializer does, so that it can
follow moved sequences. Floats whose token lives in the edit arena (also
"nan", "inf", "-inf") are converted out of line. dump() copies only
strings of the source without escaping, and shrink_to_fit() leaves the
node array in place once there are edits, as they link into it.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:10 +02:00
Niels Lohmann 95c2fe08a6 Document the hash index of json_view on the architecture page
The node index section says what extra holds for objects and how large
objects are indexed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:09 +02:00
Niels Lohmann 0d59ae5d49 Address the clang-tidy findings of the object index tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:08 +02:00
Niels Lohmann 61b104606d Index the large objects of json_view
Lookups in objects are linear, as for ordered_json. Objects with 128
members or more now get a hash table after parsing (open addressing; the
first of duplicate keys is kept, as for the linear search), so that
operator[], at(), find(), contains(), count(), value(), and JSON pointers
take constant time on average in them; the idea of switching to a hash
table for large objects is Boost.JSON's. The parser notes such objects when
it closes them (out of line, so that the parse loop only has a call for
it), and the object node keeps the number of its table.

Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms.
Parsing (json_document::parse, best of 7, separate processes): most files
within 1%; canada +5%, mesh.pretty +3%, citm +3%.

Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty,
and duplicate keys, missing keys, comparisons), nested large objects, and
documents reused with read().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:07 +02:00
Niels Lohmann 57b5d56522 Address the clang-tidy findings of the SIMD scan
Hold the UTF-8 lookup tables in std::array, compute the length of a
sequence without nested conditionals, and use std::array in the tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:06 +02:00
Niels Lohmann ebfbbf29cb Scan the strings of json_view with NEON and SSE2
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with
GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction
sets. A signed compare with 0x20 finds control characters and non-ASCII
bytes at once. Keys keep 16 table checks before the vector loop (their
lengths repeat from record to record, so the branches predict well);
string values have 8, as their lengths vary more.

Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of
simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One
Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if
JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code
must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD
selects the portable code. The vector code sits in
detail/view/simd.hpp; the same input is accepted either way.

json_document::parse, best of 7 runs in separate processes (M1 Max):
poet.json (CJK text) -72%, random.json -25%, twitter.json -22%,
gsoc-2018.json -20%, semanticscholar -19%, github_events -11%,
apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%.

Tests: every two-byte sequence and three- and four-byte sequences with
continuation bytes at the edges of their ranges, at every offset around
the vector blocks of keys and values, cut short, and long runs of text
with a damaged byte, against json::accept and json::parse. CMake builds
the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with
JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:06 +02:00
Niels Lohmann a491cc7663 Run the json_view comparison on GitHub runners on demand
A workflow runs compare.py with pinned downloads on GitHub-hosted
Ubuntu runners and shows the results as the job summary and as an
artifact: started by hand (workflow_dispatch: x86-64 or AArch64, GCC or
Clang), or when a pull request gets the label "benchmark" (both
architectures, GCC). The label trigger gives numbers before the
workflow is on the default branch, which workflow_dispatch needs.
Shared runners are noisy, so the numbers show where json_view stands on
another architecture; published numbers still need a quiet machine.

compare.py takes the CPU name from lscpu where /proc/cpuinfo has none
(AArch64 Linux), and falls back to the architecture. Checked in Linux
containers (AArch64, Clang 15 and GCC 9, offline with the pinned
archives).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:05 +02:00
Niels Lohmann 6536c1b869 Pin the downloads of the json_view comparison
compare.py --download now checks the SHA-256 of the yyjson 0.13.0,
simdjson 4.6.11, and Boost 1.92.0 archives. It unpacks each archive
once (Boost's directory is boost_1_92_0) and, where Python supports it,
with the 'data' filter.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:04 +02:00
Niels Lohmann 038f448dec Compare json_view with yyjson, simdjson, and Boost.JSON
tests/benchmarks/json_view/ holds the comparison with other libraries,
which is not built by CMake or run by CI:

- bench_view.cpp: parse, traverse, select, and dump of twitter,
  citm_catalog, canada, jeopardy, a single tweet, and a JSON-RPC request,
  with json_view, yyjson, simdjson (DOM and On-Demand), Boost.JSON, and
  json::parse; all engines must agree on every document before anything
  is timed, and run interleaved in every round
- bench_corpus.cpp: parse, traverse, and dump of any list of files
- compare.py: builds both against include/ with the libraries of the
  system (or pinned downloads), runs them, and writes the results with
  what is needed to reproduce them (date, commit, CPU, OS, compiler,
  flags, library versions) to results/<date>-<host>.md and .csv; only the
  Python 3 standard library is used
- README.md: how to run it, what is measured, and which features the
  engines have, so the numbers can be read correctly

Boost.JSON is optional (JSON_VIEW_BENCH_BOOST).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:03 +02:00
Niels Lohmann bf6b5b4719 Convert the floats of json_view from the digit layout
The parser records where the integer digits, the fraction digits, and the
exponent of a float token are. For doubles with at most 19 digits, the
value is now read from that layout: the digits eight at a time, without
scanning the token, and rounded with Clinger's fast path where both
operands are exact, else with the Eisel-Lemire algorithm (which needs no
fallback for up to 19 digits). Both round correctly, so the values are
those of parse(); other tokens and types keep the library's conversion.

get<double>(), materialize(), dump(), and comparisons use it. Traversing
canada.json (111,000 floats, every number converted): 0.53 -> 0.86 GB/s.

Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22) to the
bit-for-bit comparison with parse().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:02 +02:00
Niels Lohmann 1c7041908d Address the cpplint findings of json_view's comparisons
compare.hpp includes <string> (build/include_what_you_use).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:01 +02:00
Niels Lohmann 20fd4c6a8b Address the clang-tidy findings of the comparisons
Separate the comparison of discarded values from the other types, so
that the conditional chain has no repeated branch bodies, and mark
the deliberate comparisons of views with empty containers in the
tests.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:01 +02:00
Niels Lohmann f57de1aa0a Document the comparisons of json_view
- API pages for operator== and operator!= of basic_json_view, linked
  both ways with the basic_json pages
- the feature page and the class overview list the comparisons
- the examples show when the view helps: detecting a changed document
  without building json values

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:03:00 +02:00
Niels Lohmann 401af52511 Add comparisons to json_view
basic_json_view gains operator== and operator!= with other views and with
basic_json values. Two views are equal if the values parse() would
produce for them are equal by basic_json's operator==: numbers compare by
value across their types, and objects by their members, with duplicate
keys resolved as parse() resolves them (the last value, at the position of
the first key). Objects are compared in member order if the object type
keeps an order (ordered_json), by key otherwise, as basic_json does.
Discarded views compare as discarded basic_json values do, which follows
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except
single numbers, and the walk is iterative.

Tests compare the results for pairs of 1,200 generated documents (also
written differently: sorted keys, canonical numbers) with those of
basic_json, for json and ordered_json, plus numbers, duplicate keys,
member order, discarded values, and 100,000 levels of nesting.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:02:59 +02:00
Niels Lohmann ecce914701 Mark the cases of the view's serializer that tests cannot reach
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:02:58 +02:00
Niels Lohmann 1318023103 Address the clang-tidy findings of dump()
The output buffer initializes its members in the initializer list, and the
escaping has no nested conditional operators; the test marks a fixed seed.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:02:57 +02:00
Niels Lohmann 0192171c37 Document dump() of json_view
- API pages for dump, number_format, and operator<< of basic_json_view,
  linked both ways with the basic_json pages
- the feature page describes document order and number_format::source
- the examples show when the view helps: forwarding part of a message
  and writing numbers exactly as they were read

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:02:56 +02:00
Niels Lohmann 1ef0582ed7 Add dump() to json_view
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.

The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.

Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 14:02:56 +02:00
105 changed files with 14661 additions and 418 deletions
+4 -3
View File
@@ -51,12 +51,13 @@ labels:
- "include/nlohmann/detail/view/.*"
- "single_include/nlohmann/json_view\\.hpp"
- "tests/src/unit-json_view.*"
- "tests/src/fuzzer-parse_json_view\\.cpp"
- "tests/src/fuzzer-(parse_json_view|json_view_image)\\.cpp"
- "tests/benchmarks/json_view/.*"
- "tools/amalgamate/config_json_view\\.json"
- "docs/mkdocs/docs/features/json_view\\.md"
- "docs/mkdocs/docs/api/basic_json_(document|view)/.*"
- "docs/mkdocs/docs/api/(ordered_)?json_(document|view)\\.md"
- "docs/mkdocs/docs/examples/(basic_json_(document|view)__|(ordered_)?json_(document|view)).*"
- "docs/mkdocs/docs/api/(ordered_)?json_(editable_)?(document|view)\\.md"
- "docs/mkdocs/docs/examples/(basic_json_(document|view)__|(ordered_)?json_(editable_)?(document|view)).*"
- "tests/benchmarks/src/benchmarks_view\\.cpp"
- label: "aspect: json_view"
@@ -0,0 +1,78 @@
name: "json_view benchmarks"
# On demand only: runs the comparison of json_view with yyjson, simdjson, and
# Boost.JSON (tests/benchmarks/json_view/compare.py) on GitHub-hosted runners,
# for numbers from x86-64 and AArch64 Linux. It runs when started by hand, or
# when a pull request gets the label "benchmark" (on both architectures, with
# GCC and the default settings). Shared runners are noisy: the results show
# where json_view stands, but published numbers need a quiet machine (see
# tests/benchmarks/json_view/README.md).
on:
pull_request:
types: [labeled]
workflow_dispatch:
inputs:
runner:
description: "Runner image"
type: choice
options:
- ubuntu-24.04
- ubuntu-24.04-arm
default: ubuntu-24.04
compiler:
description: "Compiler"
type: choice
options:
- g++
- clang++
default: g++
native:
description: "Compile for the runner's CPU (-march=native)"
type: boolean
default: false
rounds:
description: "Rounds of bench_view"
type: number
default: 30
permissions:
contents: read
jobs:
compare:
if: github.event_name == 'workflow_dispatch' || github.event.label.name == 'benchmark'
strategy:
matrix:
runner: ${{ fromJSON(github.event_name == 'workflow_dispatch' && format('["{0}"]', inputs.runner) || '["ubuntu-24.04", "ubuntu-24.04-arm"]') }}
runs-on: ${{ matrix.runner }}
steps:
- name: Harden Runner
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
with:
egress-policy: audit
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Download test data
run: |
cmake -S . -B build -DJSON_BuildTests=On
cmake --build build --target download_test_data
- name: Run the comparison
env:
CXX: ${{ inputs.compiler || 'g++' }}
CC: ${{ inputs.compiler == 'clang++' && 'clang' || 'gcc' }}
ROUNDS: ${{ inputs.rounds || 30 }}
NATIVE: ${{ inputs.native && '--native' || '' }}
run: python3 tests/benchmarks/json_view/compare.py --data build/test_files --download --rounds "$ROUNDS" $NATIVE
- name: Summary
run: cat tests/benchmarks/json_view/results/*.md >> "$GITHUB_STEP_SUMMARY"
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: json_view-benchmarks-${{ matrix.runner }}-${{ inputs.compiler || 'g++' }}
path: tests/benchmarks/json_view/results/
+7
View File
@@ -67,8 +67,12 @@ cc_library(
"include/nlohmann/detail/string_utils.hpp",
"include/nlohmann/detail/value_t.hpp",
"include/nlohmann/detail/view/builder.hpp",
"include/nlohmann/detail/view/compare.hpp",
"include/nlohmann/detail/view/document_data.hpp",
"include/nlohmann/detail/view/edit.hpp",
"include/nlohmann/detail/view/edit_storage.hpp",
"include/nlohmann/detail/view/errors.hpp",
"include/nlohmann/detail/view/image.hpp",
"include/nlohmann/detail/view/input.hpp",
"include/nlohmann/detail/view/iterator.hpp",
"include/nlohmann/detail/view/lookup.hpp",
@@ -77,8 +81,11 @@ cc_library(
"include/nlohmann/detail/view/materialize.hpp",
"include/nlohmann/detail/view/node.hpp",
"include/nlohmann/detail/view/number.hpp",
"include/nlohmann/detail/view/object_index.hpp",
"include/nlohmann/detail/view/pointer.hpp",
"include/nlohmann/detail/view/scan.hpp",
"include/nlohmann/detail/view/serializer.hpp",
"include/nlohmann/detail/view/simd.hpp",
"include/nlohmann/detail/view/string_ref.hpp",
"include/nlohmann/detail/view/value.hpp",
"include/nlohmann/json.hpp",
+1
View File
@@ -1405,6 +1405,7 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I
- The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0).
- The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
- The view's parser (`<nlohmann/json_view.hpp>`) contains techniques and code adapted from [yyjson](https://github.com/ibireme/yyjson) by YaoYuan, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above): table-driven decoding of `\u` escapes and fixed-offset unrolled checks.
- The view's parser (`<nlohmann/json_view.hpp>`) validates non-ASCII strings with the vector UTF-8 check of [simdjson](https://github.com/simdjson/simdjson) by Daniel Lemire, Geoff Langdale, John Keiser, and contributors (its "lookup4" algorithm and tables, after J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One Instruction Per Byte", 2021), which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here) and the Apache 2.0 License. Copyright &copy; 2018-2025 The simdjson authors
<img align="right" src="https://git.fsfe.org/reuse/reuse-ci/raw/branch/master/reuse-horizontal.png" alt="REUSE Software">
+17
View File
@@ -131,14 +131,20 @@ INSERT INTO searchIndex(name, type, path) VALUES ('basic_json::~basic_json', 'Me
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document', 'Class', 'api/basic_json_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::basic_json_document', 'Constructor', 'api/basic_json_document/basic_json_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::accept', 'Function', 'api/basic_json_document/accept/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::erase', 'Method', 'api/basic_json_document/erase/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::insert', 'Method', 'api/basic_json_document/insert/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::is_discarded', 'Method', 'api/basic_json_document/is_discarded/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::load', 'Function', 'api/basic_json_document/load/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::memory_usage', 'Method', 'api/basic_json_document/memory_usage/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::node_count', 'Method', 'api/basic_json_document/node_count/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::owns_source', 'Method', 'api/basic_json_document/owns_source/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::parse', 'Function', 'api/basic_json_document/parse/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::parse_copy', 'Function', 'api/basic_json_document/parse_copy/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::push_back', 'Method', 'api/basic_json_document/push_back/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::read', 'Method', 'api/basic_json_document/read/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::root', 'Method', 'api/basic_json_document/root/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::save', 'Method', 'api/basic_json_document/save/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::set', 'Method', 'api/basic_json_document/set/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::shrink_to_fit', 'Method', 'api/basic_json_document/shrink_to_fit/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_document::source', 'Method', 'api/basic_json_document/source/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view', 'Class', 'api/basic_json_view/index.html');
@@ -150,6 +156,7 @@ INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::cbegin', 'Me
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::cend', 'Method', 'api/basic_json_view/cend/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::contains', 'Method', 'api/basic_json_view/contains/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::count', 'Method', 'api/basic_json_view/count/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::dump', 'Method', 'api/basic_json_view/dump/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::empty', 'Method', 'api/basic_json_view/empty/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::end', 'Method', 'api/basic_json_view/end/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::find', 'Method', 'api/basic_json_view/find/index.html');
@@ -172,9 +179,13 @@ INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::is_string',
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::is_structured', 'Method', 'api/basic_json_view/is_structured/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::items', 'Method', 'api/basic_json_view/items/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::materialize', 'Method', 'api/basic_json_view/materialize/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::number_format', 'Enum', 'api/basic_json_view/number_format/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::number_token', 'Method', 'api/basic_json_view/number_token/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::operator bool', 'Method', 'api/basic_json_view/operator_bool/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::operator<<', 'Operator', 'api/basic_json_view/operator_ltlt/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::operator[]', 'Operator', 'api/basic_json_view/operator[]/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::operator==', 'Operator', 'api/basic_json_view/operator_eq/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::operator!=', 'Operator', 'api/basic_json_view/operator_ne/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::size', 'Method', 'api/basic_json_view/size/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::source_offset', 'Method', 'api/basic_json_view/source_offset/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::type', 'Method', 'api/basic_json_view/type/index.html');
@@ -182,6 +193,8 @@ INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::type_name',
INSERT INTO searchIndex(name, type, path) VALUES ('basic_json_view::value', 'Method', 'api/basic_json_view/value/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('json', 'Class', 'api/json/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('json_document', 'Class', 'api/json_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('json_editable_document', 'Class', 'api/json_editable_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('json_editable_view', 'Class', 'api/json_editable_view/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('json_view', 'Class', 'api/json_view/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('json_pointer', 'Class', 'api/json_pointer/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('json_pointer::back', 'Method', 'api/json_pointer/back/index.html');
@@ -217,6 +230,8 @@ INSERT INTO searchIndex(name, type, path) VALUES ('operator<<', 'Operator', 'api
INSERT INTO searchIndex(name, type, path) VALUES ('operator>>', 'Operator', 'api/operator_gtgt/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('ordered_json', 'Class', 'api/ordered_json/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('ordered_json_document', 'Class', 'api/ordered_json_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('ordered_json_editable_document', 'Class', 'api/ordered_json_editable_document/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('ordered_json_editable_view', 'Class', 'api/ordered_json_editable_view/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('ordered_json_view', 'Class', 'api/ordered_json_view/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('ordered_map', 'Class', 'api/ordered_map/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('std::hash<basic_json>', 'Class', 'api/basic_json/std_hash/index.html');
@@ -302,3 +317,5 @@ INSERT INTO searchIndex(name, type, path) VALUES ('NLOHMANN_JSON_SERIALIZE_ENUM'
INSERT INTO searchIndex(name, type, path) VALUES ('NLOHMANN_JSON_VERSION_MAJOR', 'Macro', 'api/macros/nlohmann_json_version_major/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('NLOHMANN_JSON_VERSION_MINOR', 'Macro', 'api/macros/nlohmann_json_version_major/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('NLOHMANN_JSON_VERSION_PATCH', 'Macro', 'api/macros/nlohmann_json_version_major/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_VIEW_NO_SIMD', 'Macro', 'api/macros/json_view_no_simd/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_VIEW_USE_SSSE3', 'Macro', 'api/macros/json_view_use_ssse3/index.html');
+2
View File
@@ -86,6 +86,8 @@ Binary values are serialized as an object containing two keys:
- [to_string](to_string.md) returns a string representation of a JSON value
- [operator<<](../operator_ltlt.md) serialize to stream
- [`basic_json_view::dump`](../basic_json_view/dump.md) the corresponding function of `basic_json_view`, serializing
directly from a flat index without building a `basic_json` value
- [Serialization](../../features/serialization.md) - the serialization article
## Version history
@@ -166,6 +166,8 @@ Linear.
- [operator!=](operator_ne.md) compare for inequality
- [operator<=>](operator_spaceship.md) comparison: 3-way (C++20)
- [basic_json_view::operator==](../basic_json_view/operator_eq.md) - the same comparison on a zero-copy view, without
building a `basic_json` value for it
## Version history
@@ -89,6 +89,12 @@ Linear.
--8<-- "examples/operator__notequal__nullptr_t.output"
```
## See also
- [operator==](operator_eq.md) compare for equality
- [basic_json_view::operator!=](../basic_json_view/operator_ne.md) - the same comparison on a zero-copy view, without
building a `basic_json` value for it
## Version history
1. Added in version 1.0.0. Added C++20 member functions in version 3.11.0. Changed in version 3.13.0 to remove
@@ -0,0 +1,128 @@
# <small>nlohmann::basic_json_document::</small>erase
```cpp
// (1)
std::size_t erase(view_type object, string_view_t key);
// (2)
template<typename I>
void erase(view_type array, I idx);
// (3)
std::size_t erase(const json_pointer& ptr);
```
Only an **editable** document (`#!cpp Editable == true`, e.g. [`json_editable_document`](../json_editable_document.md))
has `erase`; calling it on a read-only `basic_json_document` fails to compile (`#!cpp static_assert`).
1. Removes every member of `object` whose key is `key` (see [Notes](#notes) on duplicate keys) and returns how many
were removed; `#!cpp 0` if `object` has no member with this key.
2. Removes the element at index `idx` of `array`, which must already exist (`#!cpp idx < array.size()`).
3. Removes the value the JSON pointer `ptr` refers to, relative to [`root()`](root.md), and returns how many values
were removed: the *parent* of the target must already exist, and the target itself is removed as in 1. (an object
member; `#!cpp 0` or more) or 2. (an array element; always `#!cpp 1`). `ptr` must not be empty -- [`root()`](root.md)
itself cannot be erased.
## Template parameters
`I`
: an integral type other than `#!cpp bool`, deduced (overloads taking a `#!cpp bool` or a non-integral type for
`idx` do not participate in overload resolution).
## Parameters
`object` (in)
: the object to remove a member of
`array` (in)
: the array to remove an element of
`key` (in)
: the key of the member(s) to remove
`idx` (in)
: the index of the element to remove; a negative value throws (see [Exceptions](#exceptions))
`ptr` (in)
: a JSON pointer to the value to remove, relative to `root()`
## Return value
1. the number of removed members (`#!cpp 0` if `object` had none with this `key`)
2. (nothing)
3. the number of removed values (`#!cpp 0` or more for an object member, always `#!cpp 1` for an array element)
## Exceptions
1. Throws [`type_error.307`](../../home/exceptions.md#jsonexceptiontype_error307) if `object` is not an object -- the
same message [`BasicJsonType::erase`](../basic_json/erase.md) throws for the same type.
2. Throws `type_error.307` if `array` is not an array. Throws
[`out_of_range.401`](../../home/exceptions.md#jsonexceptionout_of_range401) if `idx` is negative, or if
`#!cpp idx >= array.size()`.
3. Throws [`out_of_range.405`](../../home/exceptions.md#jsonexceptionout_of_range405) ("JSON pointer has no parent")
if `ptr` is empty. Throws what [`at`](../basic_json_view/at.md) throws (overload 3) for resolving `ptr`'s parent.
For the last reference token itself: if the parent is an array, throws what 2. throws for an index that is out of
range, or, for a token that is not a valid array index,
[`parse_error.106`](../../home/exceptions.md#jsonexceptionparse_error106) (a leading `#!cpp '0'`),
[`parse_error.109`](../../home/exceptions.md#jsonexceptionparse_error109) (not a number),
[`out_of_range.410`](../../home/exceptions.md#jsonexceptionout_of_range410) (too large for `size_type`), or
[`out_of_range.404`](../../home/exceptions.md#jsonexceptionout_of_range404) (an empty token); otherwise (an
object, or a primitive value the pointer's parent resolves to) throws what 1. throws.
Every overload also throws [`invalid_iterator.202`](../../home/exceptions.md#jsonexceptioninvalid_iterator202) ("view
does not belong to this document") if `object`/`array` is a [discarded](../basic_json_view/is_discarded.md) view or a
view of a *different* document (overloads 1-2 only; overload 3 always starts from this document's own
[`root()`](root.md)).
## Complexity
1. Linear in the number of members of `object`.
2. Linear in the number of elements of `array` at or after `idx` (they move one slot over).
3. Linear in the number of reference tokens of `ptr` and, for each token, in the number of members of the object at
that level or the index into the array (as [`at`](../basic_json_view/at.md)), plus the complexity of 1. or 2. for
the last token.
## Notes
!!! info "Duplicate keys"
Overload 1. removes *every* member with `key`, not just the first -- unlike [`set`](set.md), which assigns the
first occurrence and drops the rest. This is why it returns a count rather than a single view: there may be
more than one member removed, or none.
Like [`set`](set.md) and [`push_back`](push_back.md), `erase` never moves an element's *value*: a view still
referring to a removed member or element keeps showing what it last held (see [Edits](index.md#edits)) -- it just no
longer appears when `array`/`object` is read, dumped, or iterated. Removing an element of `array` (2.) does shift the
*links* to the elements after it, the same way `insert`, `set`, or `push_back` on the same array would; any iterator
already taken over `array`/`object` is invalidated by an erase, since it was walking the old layout.
## Examples
??? example
The example below drops a deprecated field and a decommissioned entry from a configuration document -- using all
three overloads -- and shows what stays intact that would not with a plain `json`/`ordered_json` value: the order
of the fields around the ones removed, and the exact spelling of a number that was never touched.
```cpp
--8<-- "examples/basic_json_document__erase.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__erase.output"
```
## See also
- [insert](insert.md) - insert an element into an array
- [set](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers to
- [push_back](push_back.md) - append to an array
- [root](root.md) - the view of the root value, the starting point of overload 3
- [`BasicJsonType::erase`](../basic_json/erase.md) - the corresponding function of `basic_json`
- [Edits](index.md#edits) - what an edit guarantees, for every overload
## Version history
- Added in version 3.13.0.
@@ -3,7 +3,7 @@
<small>Defined in header `<nlohmann/json_view.hpp>`</small>
```cpp
template<typename BasicJsonType>
template<typename BasicJsonType, bool Editable = false>
class basic_json_document;
```
@@ -19,6 +19,11 @@ it (a copy, or an rvalue `#!cpp std::string` that was moved in); see [`owns_sour
is move-only: copying a document would either duplicate a potentially large index and text, or leave two documents
claiming to borrow the same buffer, so it is disabled.
With `#!cpp Editable == true`, the document also offers [`set`](set.md), [`push_back`](push_back.md),
[`insert`](insert.md), and [`erase`](erase.md) to change values in place, see [Edits](#edits) below. The source text
itself is never written; a read-only document (`#!cpp Editable == false`, the default) does not carry any of the
bookkeeping edits need, and calling any of them on one fails to compile (`#!cpp static_assert`).
## Template parameters
`BasicJsonType`
@@ -26,23 +31,37 @@ claiming to borrow the same buffer, so it is disabled.
[`ordered_json`](../ordered_json.md). Only 64-bit `number_integer_t`/`number_unsigned_t` types are supported; this
is checked with a `static_assert`.
`Editable`
: whether the document supports [`set`](set.md), [`push_back`](push_back.md), [`insert`](insert.md), and
[`erase`](erase.md) (optional, `#!cpp false` by default). See [Edits](#edits) below.
## Specializations
- [**json_document**](../json_document.md) - documents of the default specialization [`json`](../json.md)
- [**ordered_json_document**](../ordered_json_document.md) - documents of [`ordered_json`](../ordered_json.md)
- [**json_document**](../json_document.md) - read-only documents of the default specialization [`json`](../json.md)
- [**ordered_json_document**](../ordered_json_document.md) - read-only documents of
[`ordered_json`](../ordered_json.md)
- [**json_editable_document**](../json_editable_document.md) - editable documents of [`json`](../json.md)
- [**ordered_json_editable_document**](../ordered_json_editable_document.md) - editable documents of
[`ordered_json`](../ordered_json.md)
## Member types
- **view_type** - the type of view returned by [`root()`](root.md) (`#!cpp basic_json_view<BasicJsonType>`)
- **view_type** - the type of view returned by [`root()`](root.md) (`#!cpp basic_json_view<BasicJsonType, Editable>`)
- **value_t** - the JSON type enumeration, see [`basic_json::value_t`](../basic_json/value_t.md)
## Member functions
- [(constructor)](basic_json_document.md)
### Parsing
- [**parse**](parse.md) (_static_) - deserialize from a compatible input, borrowing or owning it as appropriate
- [**parse_copy**](parse_copy.md) (_static_) - deserialize a copy of a compatible input
- [**accept**](accept.md) (_static_) - check whether the input is valid JSON
- [**read**](read.md) - (re-)parse into this document, reusing its memory
### Access
- [**root**](root.md) - the view of the root value
- [**is_discarded**](is_discarded.md) - return whether the last parse failed
- [**source**](source.md) - the parsed text
@@ -51,6 +70,49 @@ claiming to borrow the same buffer, so it is disabled.
- [**memory_usage**](memory_usage.md) - the number of bytes held by the document
- [**shrink_to_fit**](shrink_to_fit.md) - release unused index capacity
### Images
- [**save**](save.md) - the document as an image that `load()` reads without parsing
- [**load**](load.md) (_static_) - read an image written by `save()`
### Edits
- [**set**](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers
to (`#!cpp Editable` documents only)
- [**push_back**](push_back.md) - append to an array (`#!cpp Editable` documents only)
- [**insert**](insert.md) - insert an element into an array before a given position (`#!cpp Editable` documents only)
- [**erase**](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
(`#!cpp Editable` documents only)
## Edits
An editable document (`#!cpp Editable == true`) can be changed after parsing, with [`set`](set.md),
[`push_back`](push_back.md), [`insert`](insert.md), and [`erase`](erase.md);
[`json_editable_document`](../json_editable_document.md) and
[`ordered_json_editable_document`](../ordered_json_editable_document.md) are the corresponding specializations. A few
points apply to every edit:
- **The source text is never written**, and the parsed index never moves: every value keeps the node it was parsed
into, so [views](../basic_json_view/index.md) taken before an edit stay valid, including
[`root()`](root.md). New values (and the element sequences of an edited array/object) go to storage owned by the
document, allocated on demand.
- **A view keeps referring to the same value.** After [`set`](set.md) replaces the value a view refers to, that view
sees the new value; a view of a value that a later edit replaces or drops keeps showing what it last held. An edit
of an array or object, however, **invalidates the iterators taken over it** (its members may now live in a
different sequence), and a string obtained with [`get_string()`](../basic_json_view/get_string.md) stays valid even
as further edits happen (earlier buffers of edited text are kept alive, not overwritten).
- **Values are accepted three ways:** a [`basic_json_view`](../basic_json_view/index.md) of *any* document
(read-only or editable; it is copied, nothing is shared with the source document), a `BasicJsonType` value, or
anything `BasicJsonType` can be constructed from (numbers, strings, `#!cpp bool`, `#!cpp nullptr`, containers, ...).
- [`dump()`](../basic_json_view/dump.md) writes an edited document with members in document order, new members at
the end, and, with [`number_format::source`](../basic_json_view/number_format.md), keeps the spelling of every
number that was not itself edited -- see [Editing a document](../../features/json_view.md#editing-a-document) for
why this matters.
- [`read()`](read.md) discards all edits, [`shrink_to_fit()`](shrink_to_fit.md) does not move the node index once
there are edits, and [`memory_usage()`](memory_usage.md) includes the memory edits use.
[`source_offset()`](../basic_json_view/source_offset.md) of a value introduced by an edit is
`#!cpp static_cast<std::size_t>(-1)`, the same value it reports for a decoded string.
## Version history
- Added in version 3.13.0.
@@ -0,0 +1,114 @@
# <small>nlohmann::basic_json_document::</small>insert
```cpp
template<typename I, typename V>
view_type insert(view_type array, I idx, V&& value);
```
Only an **editable** document (`#!cpp Editable == true`, e.g. [`json_editable_document`](../json_editable_document.md))
has `insert`; calling it on a read-only `basic_json_document` fails to compile (`#!cpp static_assert`).
Inserts `value` into `array` as a new element before position `idx`, which must not be past the end
(`#!cpp idx <= array.size()`; `#!cpp idx == array.size()` appends, like [`push_back`](push_back.md)). Unlike
[`push_back`](push_back.md), a [null](../basic_json_view/is_null.md) `array` does *not* first become an empty array:
`array` must already be an array.
`value` is accepted three ways: a [`basic_json_view`](../basic_json_view/index.md) of *any* document -- read-only or
editable, and it does not have to be `array`'s own document -- which is copied so that nothing is shared with the
source document afterward; a `BasicJsonType` value; or anything `BasicJsonType` can be constructed from (numbers,
strings, `#!cpp bool`, `#!cpp nullptr`, containers, ...).
## Template parameters
`I`
: an integral type other than `#!cpp bool`, deduced (overloads taking a `#!cpp bool` or a non-integral type for
`idx` do not participate in overload resolution).
`V`
: the type of `value`, deduced; see above for what is accepted.
## Parameters
`array` (in)
: the array to insert into
`idx` (in)
: the position to insert `value` before; a negative value throws (see [Exceptions](#exceptions))
`value` (in)
: the value to insert
## Return value
a view of the new element, now holding `value`
## Exception safety
Basic exception safety: `value` is fully encoded -- including the checks below -- into storage owned by the document
before anything already reachable from [`root()`](root.md) is touched, so a failure while encoding `value` (an
invalid argument, or `#!cpp std::bad_alloc`) leaves the document completely unchanged, other than memory allocated
for the encoding that is not reclaimed. A failure of a later allocation -- while `array` switches from its parsed
layout to a growable block, or while that block grows, see [Notes](#notes) -- can still leave `array` already
switched to that layout even though `value` itself was not inserted.
## Exceptions
Throws [`type_error.309`](../../home/exceptions.md#jsonexceptiontype_error309) if `array` is not an array -- the same
message [`BasicJsonType::insert`](../basic_json/insert.md) throws for the same type; a null `array` throws this too
(see above). Throws [`out_of_range.401`](../../home/exceptions.md#jsonexceptionout_of_range401) if `idx` is negative,
or if `#!cpp idx > array.size()`. Throws
[`invalid_iterator.202`](../../home/exceptions.md#jsonexceptioninvalid_iterator202) ("view does not belong to this
document") if `array` is a [discarded](../basic_json_view/is_discarded.md) view or a view of a *different* document.
Throws [`type_error.302`](../../home/exceptions.md#jsonexceptiontype_error302) if `value` is a
[discarded](../basic_json_view/is_discarded.md) view or a [discarded](../basic_json/is_discarded.md) `BasicJsonType`
value, and [`type_error.319`](../../home/exceptions.md#jsonexceptiontype_error319) if `value` is (or contains) a
binary value -- `BasicJsonType` can hold one, but a `json_document` cannot. Throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) if `value` is (or contains) a string that is
not valid UTF-8, with the same message [`BasicJsonType::dump()`](../basic_json/dump.md) gives for that string.
## Complexity
Linear in the number of elements of `array` at or after `idx` (they move one slot over), plus time linear in the
size of `value` to encode it into the document's storage (constant for a scalar, linear in the number of nested
values for an array or object): like [`push_back`](push_back.md), the elements of `array` move to a growable block
of links the first time it is inserted into (or [`set`](set.md)/[`push_back`](push_back.md) on), and that block
grows in amortized constant time; inserting before the end within that block still shifts every later element.
## Notes
Like [`set`](set.md) on a member or an element, `insert` never moves an existing *element's value* -- only where
`array`'s *links* to its elements live -- so a view of an existing element of `array` stays valid across an
`insert`, and keeps referring to the same element even though its index shifts. Any iterator already taken over
`array` is invalidated, since it was walking the old layout. See [Edits](index.md#edits) for what stays valid across
an edit in general.
## Examples
??? example
The example below inserts a step into the middle of a deployment plan, without touching the steps that come
after it, and shows that a view taken before the insert keeps referring to the same element even though its
index shifts -- something a plain `json`/`ordered_json` array, or its `std::vector`-based storage, has no
equivalent for.
```cpp
--8<-- "examples/basic_json_document__insert.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__insert.output"
```
## See also
- [push_back](push_back.md) - append to an array
- [erase](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
- [set](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers to
- [`BasicJsonType::insert`](../basic_json/insert.md) - the corresponding function of `basic_json`
- [Edits](index.md#edits) - what an edit guarantees, for every overload
## Version history
- Added in version 3.13.0.
@@ -0,0 +1,169 @@
# <small>nlohmann::basic_json_document::</small>load
```cpp
// (1)
static basic_json_document load(const std::uint8_t* image, std::size_t size,
const image_check check = image_check::full);
// (2)
static basic_json_document load(const std::vector<std::uint8_t>& image,
const image_check check = image_check::full);
// (3)
static basic_json_document load(std::vector<std::uint8_t>&& image,
const image_check check = image_check::full);
```
1. Reads an image [`save()`](save.md) wrote, from a pointer and a byte count. The image is **borrowed**: `image`
must stay alive and unchanged for as long as the returned document, and any view taken from it, is used.
2. Reads an image from a `#!cpp std::vector`. Also **borrowed** -- equivalent to overload 1 called with
`#!cpp image.data()` and `#!cpp image.size()`.
3. Reads an image, keeping the vector instead of copying it: `image` is moved into the document (no copy), which
then owns it for as long as it needs the text and the decoded strings. [`owns_source()`](owns_source.md) is
`#!cpp true` afterward.
In every overload, the node index is copied into storage the document itself owns -- so that it is properly aligned,
and, for an [editable](index.md#edits) document, can be edited -- while the text and the decoded strings stay in
`image`. The hash indexes [large objects](../../features/json_view.md) use for lookup are rebuilt, exactly as after
parsing.
## Parameters
`image` (in)
: the image [`save()`](save.md) wrote (overloads 1 and 2), or one to take ownership of (overload 3)
`size` (in)
: the number of bytes at `image` (overload 1)
`check` (in)
: how thoroughly to validate `image` before trusting it; see [`image_check`](#image_check) below (optional,
`#!cpp image_check::full` by default)
## Return value
The document read from the image.
## Exception safety
Overloads 1 and 2 give the strong guarantee: `image` is only read, never written, so a thrown exception leaves the
caller's buffer untouched.
Overload 3 moves `image` into the document *before* validating it, so that a good image is kept without a copy. If
loading then fails, the partially built document -- and the vector now inside it -- is discarded along with the
exception, and `image` itself is left **empty**, not restored to what was passed in. Move a copy in instead, or
validate with overload 2 first, if the original vector must survive a failed load.
## Exceptions
On a big-endian target, throws [`type_error.320`](../../home/exceptions.md#jsonexceptiontype_error320) -- the same
exception [`save()`](save.md#exceptions) throws there, since the image format is little-endian only.
Otherwise throws [`parse_error.116`](../../home/exceptions.md#jsonexceptionparse_error116) if `image` is not one
`save()` could have written, or fails the requested `check`:
| message | when |
|------------------------|------------------------------------------------------------------------------------------------------|
| `too short` | `image` is `#!cpp nullptr`, or `size` is smaller than the 64-byte header |
| `unknown format` | the header's magic bytes or version do not match, or a reserved header field is not zero |
| `sizes out of range` | the node count, text size, or decoded-string size the header describes does not fit `size`, or the `#!cpp '\0'` after the text or after the decoded strings is missing |
| `the check failed` | `check` is not `#!cpp image_check::none`, and the image fails it -- see [`image_check`](#image_check) |
!!! failure "Example messages"
```
[json.exception.parse_error.116] parse error: invalid json_document image: too short
```
```
[json.exception.parse_error.116] parse error: invalid json_document image: unknown format
```
```
[json.exception.parse_error.116] parse error: invalid json_document image: sizes out of range
```
```
[json.exception.parse_error.116] parse error: invalid json_document image: the check failed
```
## Complexity
Linear in the number of nodes, which are always copied into the document. With `#!cpp check == image_check::full`,
additionally linear in the combined length of the text and the decoded strings; `#!cpp image_check::bounds` and
`#!cpp image_check::none` do not read them.
## `image_check`
```cpp
using image_check = detail::view::image_check;
enum class image_check
{
full,
bounds,
none
};
```
How thoroughly `load()` validates `image` before trusting it.
| value | checks | guarantees |
|----------|--------------------------------------------------------------------------------------------------------------|------------|
| `full` | everything the parser itself guarantees: structure and bounds; that every string is valid UTF-8 (and, for a string still in the source text, that it contains no quote, backslash, or control character); and that every number token is well-formed and matches the value stored for it | reading and serializing a checked image is safe and always produces valid JSON, exactly as for a parsed document |
| `bounds` | structure and bounds only -- that every offset and count in the node index stays inside the image | reading and serializing stay memory-safe, but a crafted image can hold strings that are not valid UTF-8 or that serialize to invalid JSON ([`dump()`](../basic_json_view/dump.md) writes them unchanged or throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316)), and numbers whose values differ from their text |
| `none` | nothing | images from a trusted source only -- reading a damaged image is undefined behavior |
`full` is the default and the right choice for an image from anything you do not fully control -- a file, a cache
shared with other processes, a peer on the network. `bounds` skips scanning the text and the decoded strings, so it
fits a cache your own process just wrote and reads straight back, where damage would mean a bug or a hardware fault
rather than adversarial input; it still cannot crash or read out of bounds. `none` skips validation entirely and
should only be used for an image you trust as much as your own memory.
## Notes
**Lifetime.** Overloads 1 and 2 borrow `image`: it must stay alive and byte-for-byte unchanged for as long as the
returned document, and any [view](../basic_json_view/index.md) taken from it, is used -- exactly like a document
[`parse()`](parse.md) borrowed its input for. Overload 3 avoids this by keeping the vector itself; see
[`owns_source`](owns_source.md).
!!! warning "Experimental"
The image format is versioned but not yet stable, and may change in an incompatible way before it is declared
stable; `load()` already rejects an image written by a different format version with `parse_error.116`
("unknown format"). Use images to cache a document within one build of the library, or to hand one to another
process running the *same* build on the *same* (little-endian) machine -- not as a long-term storage format.
**What `image_check::bounds` does not guarantee.** A bounds-checked image can never make `load()`,
[`root()`](root.md), element access, or [`materialize()`](../basic_json_view/materialize.md) read outside the image,
so those stay safe on a damaged one. It does *not* guarantee that the image describes valid JSON: a string
that a `full` check would have rejected can make [`dump()`](../basic_json_view/dump.md) write invalid UTF-8 or invalid
JSON, or throw `type_error.316`, and a number can read back with a value that does not match how it is spelled. Reserve `bounds` for images you already trust to be well-formed, and use it
only to skip the extra scan.
## Examples
??? example "Caching a document, ownership, and a rejected image"
The example below saves a parsed document as an image, checks that `load()` reproduces the original
[`dump()`](../basic_json_view/dump.md) without parsing, and shows the difference between
`load(std::move(image))` (owned) and `load(image)` (borrowed). It then damages one byte of the image and shows
`image_check::full` rejecting it with `parse_error.116`, while `image_check::bounds` -- meant for a cache the
process already trusts -- still reads it without going out of bounds.
```cpp
--8<-- "examples/basic_json_document__load.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__load.output"
```
## See also
- [save](save.md) - write the document as an image
- [owns_source](owns_source.md) - return whether the document holds its own copy of the text
- [parse](parse.md) - deserialize from JSON text instead of an image
- [Images](../../features/json_view.md#images) - why and when to use images
## Version history
- Added in version 3.13.0.
@@ -132,6 +132,7 @@ integer type becomes a floating-point value.
- [accept](accept.md) - check whether the input is valid JSON
- [read](read.md) - (re-)parse into this document, reusing its memory
- [owns_source](owns_source.md) - return whether the document holds its own copy of the text
- [load](load.md) - read a document from an image instead of parsing JSON text
- [`BasicJsonType::parse`](../basic_json/parse.md) - the corresponding function of `basic_json`
## Version history
@@ -0,0 +1,102 @@
# <small>nlohmann::basic_json_document::</small>push_back
```cpp
template<typename V>
view_type push_back(view_type array, V&& value);
```
Appends `value` as a new last element of `array`. A [null](../basic_json_view/is_null.md) `array` first becomes an
empty array, the same way [`set`](set.md) turns a null `object` into an empty object.
`value` is accepted three ways: a [`basic_json_view`](../basic_json_view/index.md) of *any* document -- read-only or
editable, and it does not have to be `array`'s own document -- which is copied so that nothing is shared with the
source document afterward; a `BasicJsonType` value; or anything `BasicJsonType` can be constructed from (numbers,
strings, `#!cpp bool`, `#!cpp nullptr`, containers, ...).
Only an **editable** document (`#!cpp Editable == true`, e.g. [`json_editable_document`](../json_editable_document.md))
has `push_back`; calling it on a read-only `basic_json_document` fails to compile (`#!cpp static_assert`).
## Template parameters
`V`
: the type of `value`, deduced; see above for what is accepted.
## Parameters
`array` (in)
: the array (or null value) to append to
`value` (in)
: the value to append
## Return value
a view of the new last element of `array`, now holding `value`
## Exception safety
Basic exception safety: `value` is fully encoded -- including the checks below -- into storage owned by the document
before anything already reachable from [`root()`](root.md) is touched, so a failure while encoding `value` (an
invalid argument, or `#!cpp std::bad_alloc`) leaves the document completely unchanged, other than memory allocated
for the encoding that is not reclaimed. A failure of a later allocation -- while `array` switches from its parsed
layout to a growable block, or while that block grows, see [Notes](#notes) -- can still leave a partial effect, such
as a null `array` argument already turned into an empty array even though `value` itself was not appended.
## Exceptions
Throws [`type_error.308`](../../home/exceptions.md#jsonexceptiontype_error308) if `array` is neither an array nor
null -- the same message [`BasicJsonType::push_back`](../basic_json/push_back.md) throws for the same type. Throws
[`invalid_iterator.202`](../../home/exceptions.md#jsonexceptioninvalid_iterator202) ("view does not belong to this
document") if `array` is a [discarded](../basic_json_view/is_discarded.md) view or a view of a *different* document.
Throws [`type_error.302`](../../home/exceptions.md#jsonexceptiontype_error302) if `value` is a
[discarded](../basic_json_view/is_discarded.md) view or a [discarded](../basic_json/is_discarded.md) `BasicJsonType`
value, and [`type_error.319`](../../home/exceptions.md#jsonexceptiontype_error319) if `value` is (or contains) a
binary value -- `BasicJsonType` can hold one, but a `json_document` cannot. Throws
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) if `value` is (or contains) a string that is
not valid UTF-8, with the same message [`BasicJsonType::dump()`](../basic_json/dump.md) gives for that string.
## Complexity
Amortized constant, plus time linear in the size of `value` to encode it into the document's storage (constant for
a scalar, linear in the number of nested values for an array or object): the elements of `array` move to a growable
block of links the first time it is appended to (or [`set`](set.md) on), and that block itself grows -- doubling its
capacity, so the cost of growing it amortizes to constant per element -- only once it runs out of room. See
[Notes](#notes).
## Notes
Like [`set`](set.md) on a member or an element, `push_back` never moves an existing element itself -- only where
`array`'s *links* to its elements live -- so a view of an existing element of `array` stays valid across a
`push_back`, but any iterator already taken over `array` is invalidated, since it was walking the old layout. See
[Edits](index.md#edits) for what stays valid across an edit in general.
## Examples
??? example
The example below appends records to an array one at a time, as they might arrive from a stream of events,
without ever building a `BasicJsonType` value for the array or for the records already in it, and shows that a
view taken from an earlier `push_back` still refers to the same element once later ones have run.
```cpp
--8<-- "examples/basic_json_document__push_back.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__push_back.output"
```
## See also
- [set](set.md) - replace a value, or set an object member, an array element, or the value a JSON pointer refers to
- [insert](insert.md) - insert an element into an array before a given position
- [erase](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
- [root](root.md) - the view of the root value
- [`BasicJsonType::push_back`](../basic_json/push_back.md) - the corresponding function of `basic_json`
- [Edits](index.md#edits) - what an edit guarantees, for every overload
## Version history
- Added in version 3.13.0.
@@ -70,6 +70,7 @@ own on the next, since ownership is decided freshly each time.
- [parse](parse.md) - deserialize from a compatible input
- [root](root.md) - the view of the root value
- [load](load.md) - read a document from an image instead of parsing JSON text
## Version history
@@ -0,0 +1,104 @@
# <small>nlohmann::basic_json_document::</small>save
```cpp
std::vector<std::uint8_t> save() const;
```
Writes the document as an *image*: a byte buffer that [`load`](load.md) reads back without parsing. The image holds
the node index, the source text (plus, for an edited document, the number tokens edits wrote), and the decoded
strings (plus the strings edits wrote) -- everything [`root()`](root.md) needs, with nothing left to parse.
An edited document is written in its *current* state, with its values in document order, the way the library's own
parser would have produced them for that JSON text: a member [`set`](set.md) added goes at the end, an
[`erase`](erase.md)d member leaves no trace, and a float that is not finite (NaN or positive/negative infinity)
becomes null, the same substitution [`dump()`](../basic_json_view/dump.md) makes. The same document always saves to
the same bytes -- also across `BasicJsonType` and `#!cpp Editable`, since the image reflects document order and
values only, not which specialization produced them.
## Return value
The image, as a `#!cpp std::vector<std::uint8_t>`. Pass it, or a pointer to its data together with its size, to
[`load`](load.md) to read the document back.
## Exception safety
Strong guarantee: `save()` does not modify `#!cpp *this` (it is `#!cpp const`), so if it throws, the document is left
exactly as it was, and the partially built image is discarded with the exception.
## Exceptions
Throws [`type_error.320`](../../home/exceptions.md#jsonexceptiontype_error320) if the document is
[discarded](is_discarded.md) -- a default-constructed document, or one a failed [`parse()`](parse.md)/
[`read()`](read.md) with `allow_exceptions == false` left discarded.
On a big-endian target, throws `type_error.320` with a different message instead: the image format is little-endian
only (see [Notes](#notes)).
Throws [`out_of_range.416`](../../home/exceptions.md#jsonexceptionout_of_range416) if the node count, the text, or
the decoded strings of the image would individually reach 4 GiB -- the same 32-bit offsets
[`parse()`](parse.md#exceptions) and, for edits, [`set`](set.md)/[`push_back`](push_back.md) are already limited to.
!!! failure "Example messages"
```
[json.exception.type_error.320] cannot save a discarded json_document
```
```
[json.exception.type_error.320] json_document images need a little-endian target
```
```
[json.exception.out_of_range.416] images of 4 GiB or more are not supported by json_document
```
## Complexity
Linear in the size of the document: the number of nodes, plus the length of the text and the decoded strings that end
up in the image.
## Notes
**Format.** The image begins with a 64-byte header (the magic bytes `#!cpp "NJVI"`, a version number, the node count,
and the sizes of the text and the decoded strings, all little-endian), followed by the nodes
([16 bytes each](../../home/architecture.md#node-index-of-json-views)), the text and a `#!cpp '\0'`, and the decoded
strings and a `#!cpp '\0'`. [`load`](load.md) checks the header, and the sizes it describes, before reading anything
else -- see [`load`'s Exceptions](load.md#exceptions).
!!! warning "Experimental"
The image format is versioned but not yet stable: it may change in an incompatible way before it is declared
stable. Use images to cache a document within one build of the library, or to hand one to another process running
the *same* build on the *same* (little-endian) machine -- not as a long-term storage format. Keep the original
JSON text if you need to read a saved document back with a future library version.
**Little-endian only.** The image is written as raw little-endian bytes, with no byte-swapping. `save()` (and
[`load`](load.md)) throw `type_error.320` on a big-endian target rather than silently produce bytes a big-endian
reader could not interpret correctly.
## Examples
??? example "Caching a parsed document as an image"
The example below saves a parsed configuration as an image -- the way a service might cache one to answer later
requests without parsing the text again -- and confirms that loading it back gives exactly the same result as
parsing did, and that saving is deterministic.
```cpp
--8<-- "examples/basic_json_document__save.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__save.output"
```
## See also
- [load](load.md) - read an image written by `save()`
- [owns_source](owns_source.md) - return whether the document holds its own copy of the text
- [`basic_json_view::dump`](../basic_json_view/dump.md) - serialize the document to JSON text instead of an image
- [Images](../../features/json_view.md#images) - why and when to use images
## Version history
- Added in version 3.13.0.
@@ -0,0 +1,194 @@
# <small>nlohmann::basic_json_document::</small>set
```cpp
// (1)
template<typename V>
view_type set(view_type target, V&& value);
// (2)
template<typename V>
view_type set(view_type object, string_view_t key, V&& value);
// (3)
template<typename I, typename V>
view_type set(view_type array, I idx, V&& value);
// (4)
template<typename V>
view_type set(const json_pointer& ptr, V&& value);
```
Only an **editable** document (`#!cpp Editable == true`, e.g. [`json_editable_document`](../json_editable_document.md))
has `set`; calling it on a read-only `basic_json_document` fails to compile (`#!cpp static_assert`).
1. Replaces the value `target` refers to with `value`.
2. Sets the member `key` of the object `object` to `value`: assigns it if `object` already has a member with this
key -- the first one, should the key occur more than once, and the later duplicates are then dropped (see the
[Notes](#notes) below) -- or appends a new member at the end otherwise. A [null](../basic_json_view/is_null.md)
`object` first becomes an empty object.
3. Assigns `value` to the element at index `idx` of the array `array`, which must already exist (`#!cpp idx <
array.size()`).
4. Sets the value the JSON pointer `ptr` refers to, relative to [`root()`](root.md), to `value`. The *parent* of the
target must already exist: an object member is set as in 2. (added if it does not exist yet), an array element is
assigned as in 3., and a last reference token of `#!cpp "-"`, or equal to the size of the array, appends `value`
instead, exactly as [`push_back`](push_back.md) would. An empty `ptr` sets [`root()`](root.md) itself, as in 1.
In every overload, `value` is accepted three ways: a [`basic_json_view`](../basic_json_view/index.md) of *any*
document -- read-only or editable, and it does not have to be `target`'s/`object`'s/`array`'s own document -- which
is copied so that nothing is shared with the source document afterward; a `BasicJsonType` value; or anything
`BasicJsonType` can be constructed from (numbers, strings, `#!cpp bool`, `#!cpp nullptr`, containers, ...).
## Template parameters
`V`
: the type of `value`, deduced; see above for what is accepted.
`I`
: an integral type other than `#!cpp bool`, deduced (overloads taking a `#!cpp bool` or a non-integral type for
`idx` do not participate in overload resolution).
## Parameters
`target` (in)
: the value to replace
`object` (in)
: the object (or null value) whose member to set
`array` (in)
: the array whose element to assign
`key` (in)
: the key of the member to set
`idx` (in)
: the index of the element to assign; a negative value throws (see [Exceptions](#exceptions))
`ptr` (in)
: a JSON pointer to the value to set, relative to `root()`
`value` (in)
: the new value
## Return value
1. a view of `target`, now holding `value`
2. a view of the member `key` of `object`, now holding `value`
3. a view of the element `idx` of `array`, now holding `value`
4. a view of the value `ptr` refers to, now holding `value`
## Exception safety
Basic exception safety: `value` is fully encoded -- including the checks below -- into storage owned by the document
before anything already reachable from [`root()`](root.md) is touched, so a failure while encoding `value` (an
invalid argument, or `#!cpp std::bad_alloc`) leaves the document completely unchanged, other than memory allocated
for the encoding that is not reclaimed. A failure of a later allocation -- while an edited array or object switches
from its parsed layout to a growable block, see [Notes](#notes) -- can still leave a partial effect, such as a
[null](../basic_json_view/is_null.md) `object`/`array` argument already turned into an empty object/array even
though `value` itself was not linked in.
## Exceptions
1. Throws [`type_error.302`](../../home/exceptions.md#jsonexceptiontype_error302) if `value` is a
[discarded](../basic_json_view/is_discarded.md) view, or a [discarded](../basic_json/is_discarded.md)
`BasicJsonType` value (e.g. `#!cpp BasicJsonType(value_t::discarded)`) -- an object or array `value`, of either
kind, is fine and is encoded as a whole subtree.
2. Throws [`type_error.305`](../../home/exceptions.md#jsonexceptiontype_error305) if `object` is neither an object
nor null -- the same message [`operator[]`](../basic_json_view/operator%5B%5D.md) throws for a string argument on
such a value. Throws [`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316) if `key` is not
valid UTF-8, with the same message [`BasicJsonType::dump()`](../basic_json/dump.md) gives for that string.
Also throws what 1. throws for `value`.
3. Throws `type_error.305` if `array` is not an array -- the same message `operator[]` throws for a numeric argument
on such a value. Throws [`out_of_range.401`](../../home/exceptions.md#jsonexceptionout_of_range401) if `idx` is
negative, or if `#!cpp idx >= array.size()`. Also throws what 1. throws for `value`.
4. Throws what [`at`](../basic_json_view/at.md) throws (overload 3) for resolving `ptr`'s parent, except that a
missing object member or an array index equal to the array's size at the very last reference token is not an
error there (it becomes a new member or an appended element) instead of
[`out_of_range.403`](../../home/exceptions.md#jsonexceptionout_of_range403)/[`out_of_range.402`](../../home/exceptions.md#jsonexceptionout_of_range402).
For the last reference token itself: if the parent is an object (or a primitive value, where it throws
`type_error.305`), throws what 2. throws; if the parent is an array, throws what 3. throws for an index that is
out of range, or, for a token that is not a valid array index,
[`parse_error.106`](../../home/exceptions.md#jsonexceptionparse_error106) (a leading `#!cpp '0'`),
[`parse_error.109`](../../home/exceptions.md#jsonexceptionparse_error109) (not a number),
[`out_of_range.410`](../../home/exceptions.md#jsonexceptionout_of_range410) (too large for `size_type`), or
[`out_of_range.404`](../../home/exceptions.md#jsonexceptionout_of_range404) (an empty token). Also throws what 1.
throws for `value`.
Every overload also throws [`type_error.319`](../../home/exceptions.md#jsonexceptiontype_error319) if `value` is (or
contains) a binary value -- `BasicJsonType` can hold one, but a `json_document` cannot -- and
[`invalid_iterator.202`](../../home/exceptions.md#jsonexceptioninvalid_iterator202) ("view does not belong to this
document") if `target`/`object`/`array` is a [discarded](../basic_json_view/is_discarded.md) view or a view of a
*different* document (overloads 1-3 only; overload 4 always starts from this document's own [`root()`](root.md)).
## Complexity
1. Linear in the size of `value` (encoding it into the document's storage): constant for a scalar, linear in the
number of nested values for an array or object. If `target` is itself an array or object that spans more than one
node in its parent's original, unedited layout, and `value` is a scalar, replacing it additionally costs time
linear in the number of elements of that parent, the *first* time -- see [Notes](#notes).
2. Linear in the number of members of `object`, to find an existing member with `key`, plus the complexity of 1. for
`value`.
3. Constant, plus the complexity of 1. for `value`.
4. Linear in the number of reference tokens of `ptr` and, for each token, in the number of members of the object at
that level or the index into the array (as [`at`](../basic_json_view/at.md)), plus the complexity of 2. or 3. for
the last token.
## Notes
!!! info "Duplicate keys"
If `object` already has more than one member with `key` (2.), the *first* one is assigned `value` and every
later member with the same key is removed -- so that a lookup, an iteration, and
[`materialize()`](../basic_json_view/materialize.md) of `object` afterward all agree on a single value for
`key`, the same way [`operator[]`](../basic_json_view/operator%5B%5D.md) already picks the first occurrence of a
duplicate key for reading. See the [Notes on duplicate keys](../basic_json_view/operator%5B%5D.md#notes) of
`operator[]`.
Setting a member (2.) or an element (3., through 4.) of an array or object whose elements have not been edited
before switches it from its parsed layout to a growable block holding links to its elements; a later
[`push_back`](push_back.md) or `set` on the same container reuses that block, growing it (amortized constant time)
only once it runs out of room. This never moves an element itself -- only where the container's *links* to its
elements live -- so a view of an element stays valid, but any iterator already taken over the container is
invalidated, since it was walking the old layout. See [Edits](index.md#edits) for what stays valid across an edit in
general.
The same switch happens, for the same reason, when overload 1. replaces a multi-node array/object value with a
scalar: the *parent's* element sequence is what has to switch to links, not `target` itself, because the parent
originally stepped over `target`'s whole subtree by its node count, which no longer applies once `target` is a
one-node scalar.
## Examples
??? example "Example: (1)/(2)/(3)/(4) replace a value, set a member, assign an element, set via a JSON pointer"
The example below edits a small configuration document -- replacing a value, adding an object member, assigning
an array element, and reaching a field through a JSON pointer -- and shows what
[`dump()`](../basic_json_view/dump.md) preserves that is lost once the same edits are made on a `BasicJsonType`
value instead: the order object members were written in, and the exact spelling of a number that was never
touched.
```cpp
--8<-- "examples/basic_json_document__set.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__set.output"
```
## See also
- [push_back](push_back.md) - append to an array
- [insert](insert.md) - insert an element into an array
- [erase](erase.md) - remove an object member, an array element, or the value a JSON pointer refers to
- [root](root.md) - the view of the root value, the starting point of overload 4
- [`basic_json_view::dump`](../basic_json_view/dump.md) - serialize the document, keeping an untouched number's
spelling with `#!cpp number_format::source`
- [Edits](index.md#edits) - what an edit guarantees, for every overload
- [Editing a document](../../features/json_view.md#editing-a-document) - why editable documents keep the source
text's order and number spelling
## Version history
- Added in version 3.13.0.
@@ -74,6 +74,8 @@ None of these exceptions carry a [`JSON_DIAGNOSTICS`](../macros/json_diagnostics
1. Linear in the number of members: as for [`ordered_json`](../ordered_json.md), members are compared one after
another, in document order, stopping at the first match. Each comparison first checks the key's length --
already known from the index, without reading the key bytes -- before comparing its content.
Objects with 128 or more members get a hash index while parsing, so that a lookup in them takes constant time
on average.
2. Linear in `idx`: elements are skipped one at a time from the first one, since they are not a fixed size in the
index (unlike `BasicJsonType`'s array, which is random-access).
3. Linear in the number of reference tokens of `ptr` and, for each token, in the number of members of the object at
@@ -35,6 +35,8 @@ No-throw guarantee: this function never throws exceptions.
1. Linear in the number of members: as for [`ordered_json`](../ordered_json.md), members are compared one after
another, in document order, stopping at the first match. Each comparison first checks the key's length -- already
known from the index, without reading the key bytes -- before comparing its content.
Objects with 128 or more members get a hash index while parsing, so that a lookup in them takes constant time
on average.
2. Linear in the number of reference tokens of `ptr` and, for each token, in the number of members of the object at
that level or the index into the array -- as for [`operator[]`](operator[].md#complexity) and
[`at`](at.md#complexity) with a JSON pointer.
@@ -26,6 +26,8 @@ No-throw guarantee: this function never throws exceptions.
Linear in the number of members: as for [`ordered_json`](../ordered_json.md), members are compared one after
another, in document order, stopping at the first match. Each comparison first checks the key's length -- already
known from the index, without reading the key bytes -- before comparing its content.
Objects with 128 or more members get a hash index while parsing, so that a lookup in them takes constant time on
average.
## Notes
@@ -0,0 +1,102 @@
# <small>nlohmann::basic_json_view::</small>dump
```cpp
string_t dump(const int indent = -1,
const char indent_char = ' ',
const bool ensure_ascii = false,
const number_format numbers = number_format::shortest) const;
```
Serializes this value (and its subtree) directly from the flat index, without ever building a `BasicJsonType` value
first. With the default `#!cpp numbers == number_format::shortest`, the result is the same string
[`BasicJsonType::dump`](../basic_json/dump.md) would produce for the value
[`BasicJsonType::parse()`](../basic_json/parse.md) builds from the same source text, called with the same `indent`,
`indent_char`, and `ensure_ascii` -- except that members of an object appear in document order rather than sorted by
key, and *every* occurrence of a repeated key is written rather than only the last one (see
[Notes on duplicate keys](operator[].md#notes)). For a `json_view` (whose `BasicJsonType` is not ordered), this means
`dump()` can print an object's members in a different order than [`materialize()`](materialize.md)`.dump()` of the
same subtree.
## Parameters
`indent` (in)
: If `indent` is nonnegative, array elements and object members are pretty-printed with that indent level. An
indent level of `0` only inserts newlines. `-1` (the default) selects the most compact representation.
`indent_char` (in)
: The character used for indentation if `indent` is greater than `0`. The default is ` ` (space).
`ensure_ascii` (in)
: If `ensure_ascii` is `#!cpp true`, all non-ASCII characters in the output are escaped with `\uXXXX` sequences, and
the result consists of ASCII characters only.
`numbers` (in)
: how to write numbers, see [`number_format`](number_format.md): `shortest` (the default) writes them the way
[`BasicJsonType::dump`](../basic_json/dump.md) would; `source` copies every number exactly as it appears in the
source text.
## Return value
string containing the serialization of this value, or `#!cpp "<discarded>"` if the view is
[discarded](is_discarded.md).
## Exception safety
Strong exception safety: if an exception is thrown, there are no changes to the view or the document it refers to.
## Exceptions
May throw `#!cpp std::bad_alloc` if allocating the output string fails. Unlike
[`BasicJsonType::dump`](../basic_json/dump.md), there is no `error_handler` parameter and no
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316): the view only ever holds text the parser
already validated as UTF-8, so there is nothing to replace or ignore.
## Complexity
Linear in the size of the output text.
## Notes
The walk over the subtree is iterative, so the nesting depth it can write is limited by available memory only, not by
the call stack -- as for [`materialize()`](materialize.md).
Strings are escaped by the same rules as [`BasicJsonType::dump`](../basic_json/dump.md). With
`#!cpp numbers == number_format::shortest`, floats are written with the library's shortest round-trip conversion,
exactly as [`BasicJsonType::dump`](../basic_json/dump.md) would (e.g. `#!cpp 1.5`, `#!cpp 100.0`, `#!cpp 1e+100`), and
integers are copied from the source text -- already canonical in JSON, so this matches their shortest form too --
except that `#!cpp -0` is written as `#!cpp 0`, the way [`BasicJsonType::parse()`](../basic_json/parse.md) reads it.
`#!cpp number_format::source` copies every number exactly as written in the source text instead, with no exception
for `#!cpp -0` -- `#!cpp 1.50`, `#!cpp 1E2`, `#!cpp -0.0`, `#!cpp -0`, or all digits of an integer literal with more
digits than any number type holds (such a literal is itself classified as a float, see
[What is different](../../features/json_view.md#what-is-different)) -- something `BasicJsonType` cannot do, since
parsing already reduces every number to its parsed value.
## Examples
??? example
The example below forwards a single record out of a larger batch, and re-serializes a configuration file, both
without ever building a `BasicJsonType` value for the surrounding array or for the parts of it that were not
needed. It also shows that [`materialize()`](materialize.md)`.dump()` of the configuration sorts its keys, where
`dump()` on the view keeps the order they appear in the source text.
```cpp
--8<-- "examples/basic_json_view__dump.cpp"
```
Output:
```json
--8<-- "examples/basic_json_view__dump.output"
```
## See also
- [`number_format`](number_format.md) - how `dump()` writes numbers
- [operator<<](operator_ltlt.md) - serialize this value to a stream
- [materialize](materialize.md) - build a `BasicJsonType` value, e.g. to use `BasicJsonType::dump`'s `error_handler`
- [`BasicJsonType::dump`](../basic_json/dump.md) - the corresponding function of `basic_json`
## Version history
- Added in version 3.13.0.
@@ -28,6 +28,8 @@ No-throw guarantee: this function never throws exceptions.
Linear in the number of members: as for [`ordered_json`](../ordered_json.md), members are compared one after
another, in document order, stopping at the first match. Each comparison first checks the key's length -- already
known from the index, without reading the key bytes -- before comparing its content.
Objects with 128 or more members get a hash index while parsing, so that a lookup in them takes constant time on
average.
## Notes
+30 -4
View File
@@ -3,7 +3,7 @@
<small>Defined in header `<nlohmann/json_view.hpp>`</small>
```cpp
template<typename BasicJsonType>
template<typename BasicJsonType, bool Editable = false>
class basic_json_view;
```
@@ -20,11 +20,19 @@ Moving the document itself does not invalidate its views: the index is heap-allo
`basic_json_document` object.
`basic_json_view` provides the read-only part of the `BasicJsonType` interface: the type-inspection functions, element
access, lookup, iteration, and conversion -- [`get<T>()`](get.md), [`get_string()`](get_string.md),
access, lookup, iteration, conversion, and comparison -- [`get<T>()`](get.md), [`get_string()`](get_string.md),
[`number_token()`](number_token.md), and [`materialize()`](materialize.md) to build the `BasicJsonType` value of a
subtree on demand. [`operator[]`](operator%5B%5D.md), [`at`](at.md), [`contains`](contains.md), and
[`value`](value.md) also accept a [`json_pointer`](../json_pointer/index.md). It does not (yet) provide `dump()` or
comparison.
[`value`](value.md) also accept a [`json_pointer`](../json_pointer/index.md). [`operator==`](operator_eq.md) and
[`operator!=`](operator_ne.md) compare two views, or a view and a `BasicJsonType` value, without ever building a
`BasicJsonType` value for a view; no ordering comparison (`#!cpp operator<`) is provided.
`basic_json_view` itself is always read-only -- it never has a `set` or `push_back` of its own. A view of an
**editable** document (`#!cpp Editable == true`) sees every edit made through
[`basic_json_document::set`](../basic_json_document/set.md) and
[`basic_json_document::push_back`](../basic_json_document/push_back.md): once a value is changed, every view that
still refers to it -- including ones taken before the change -- reads the new value. See
[Edits](../basic_json_document/index.md#edits).
## Template parameters
@@ -32,10 +40,17 @@ comparison.
: a specialization of [`basic_json`](../basic_json/index.md), matching the
[`basic_json_document`](../basic_json_document/index.md) the view was taken from.
`Editable`
: whether the view is of an editable document, matching the [`basic_json_document`](../basic_json_document/index.md)
it was taken from (optional, `#!cpp false` by default). See [Edits](../basic_json_document/index.md#edits).
## Specializations
- [**json_view**](../json_view.md) - views of a [`json_document`](../json_document.md)
- [**ordered_json_view**](../ordered_json_view.md) - views of an [`ordered_json_document`](../ordered_json_document.md)
- [**json_editable_view**](../json_editable_view.md) - views of a [`json_editable_document`](../json_editable_document.md)
- [**ordered_json_editable_view**](../ordered_json_editable_view.md) - views of an
[`ordered_json_editable_document`](../ordered_json_editable_document.md)
## Member types
@@ -47,6 +62,7 @@ comparison.
- **iterator**, **const_iterator** - a forward iterator over the elements of an array or the member values of an
object, in document order; both names refer to the same type, since a view is always read-only
- **item** - a (key, value) pair produced by [`items()`](items.md)
- [**number_format**](number_format.md) - how [`dump()`](dump.md) writes numbers
## Member functions
@@ -106,6 +122,16 @@ comparison.
- [**number_token**](number_token.md) - get a number's token text without a copy
- [**materialize**](materialize.md) - build the `BasicJsonType` value of this subtree
### Comparison
- [**operator==**](operator_eq.md) - comparison: equal
- [**operator!=**](operator_ne.md) - comparison: not equal
### Serialization
- [**dump**](dump.md) - serialize to a JSON-formatted string
- [**operator<<**](operator_ltlt.md) - serialize to stream
### Source access
- [**source_offset**](source_offset.md) - byte offset of this value in the document's source text
@@ -0,0 +1,51 @@
# <small>nlohmann::basic_json_view::</small>number_format
```cpp
enum class number_format {
shortest,
source
};
```
This enumeration is used in [`dump`](dump.md) to choose how numbers are written. Two values are differentiated:
shortest
: integers are copied from the source text -- already canonical in JSON -- except that `#!cpp -0` becomes
`#!cpp 0`, the way [`BasicJsonType::parse()`](../basic_json/parse.md) reads it; floats are written with the
library's shortest round-trip conversion, exactly as [`BasicJsonType::dump()`](../basic_json/dump.md) would (e.g.
`#!cpp 1.5`, `#!cpp 100.0`, `#!cpp 1e+100`)
source
: every number is copied exactly as it appears in the source text -- `#!cpp 1.50`, `#!cpp 1E2`, `#!cpp -0`, all
digits of an integer literal with more digits than any number type holds -- something `BasicJsonType` cannot do,
since parsing already reduces every number to its parsed value
## Examples
??? example
The example below writes back a price list received from a supplier: with `number_format::shortest` (the
default), a trailing zero and scientific notation are normalized away and a long account number that overflows
every number type is rounded, the same way `#!cpp materialize().dump()` (or `basic_json::dump()`) would;
`number_format::source` keeps every number exactly as it was written in the source text instead.
```cpp
--8<-- "examples/basic_json_view__number_format.cpp"
```
Output:
```json
--8<-- "examples/basic_json_view__number_format.output"
```
## See also
- [dump](dump.md) - serialize to a JSON-formatted string
- [number_token](number_token.md) - get a single number's token text without dumping the whole value
- [`BasicJsonType::error_handler_t`](../basic_json/error_handler_t.md) - the analogous enumeration for
`BasicJsonType::dump`'s decoding-error behavior
## Version history
- Added in version 3.13.0.
@@ -76,6 +76,8 @@ None of these exceptions carry a [`JSON_DIAGNOSTICS`](../macros/json_diagnostics
another, in document order, stopping at the first match. Each comparison first checks the key's length --
already known from the index, without reading the key bytes -- before comparing its content, so a key of a
different length than `key` is rejected without touching the source text.
Objects with 128 or more members get a hash index while parsing, so that a lookup in them takes constant time
on average.
2. Linear in `idx`: elements are skipped one at a time from the first one, since they are not a fixed size in the
index (unlike `BasicJsonType`'s array, which is random-access).
3. Linear in the number of reference tokens of `ptr` and, for each token, in the number of members of the object at
@@ -0,0 +1,107 @@
# <small>nlohmann::basic_json_view::</small>operator==
```cpp
// (1)
bool operator==(const basic_json_view& lhs, const basic_json_view& rhs);
// (2)
bool operator==(const basic_json_view& lhs, const BasicJsonType& rhs);
bool operator==(const BasicJsonType& lhs, const basic_json_view& rhs);
```
1. Compares two views for equality: whether the values [`BasicJsonType::parse()`](../basic_json/parse.md) would
produce for `lhs` and `rhs` are equal, according to `BasicJsonType`'s [`operator==`](../basic_json/operator_eq.md).
2. Compares a view and a `BasicJsonType` value for equality, in either order: whether the value `parse()` would
produce for the view and the other operand are equal, according to `BasicJsonType`'s
[`operator==`](../basic_json/operator_eq.md).
Neither overload builds a `BasicJsonType` value for a view to do the comparison (see [Notes](#notes) below). Numbers
compare by value across their types (`#!cpp 1 == 1.0`), and an object compares by its members, with duplicate keys
resolved exactly as `parse()` resolves them -- the last value, at the position of the first occurrence of the key.
## Parameters
`lhs` (in)
: first value to consider
`rhs` (in)
: second value to consider
## Return value
whether the values `lhs` and `rhs` are equal
## Exception safety
Strong exception safety: if an exception is thrown, there are no changes to either operand, or to the document(s) a
view refers to.
## Exceptions
May throw `#!cpp std::bad_alloc`. Unlike the other comparison and most other `basic_json_view` functions,
`operator==` is not `#!cpp noexcept`: resolving an object's members needs a temporary array to sort them by key (see
[Complexity](#complexity) below), and that allocation can fail.
## Complexity
Linear in the size of the compared values: every number, string, array element, and object member is visited at most
once, and the walk is iterative, so the nesting depth it can compare is limited by available memory only, not by the
call stack (as for [`materialize()`](materialize.md)). Resolving an object's members takes an additional O(n log n)
in the number of members at that level, since they are sorted by key to detect and resolve duplicates before being
compared. Two arrays of different [`size()`](size.md) are rejected without visiting either one's elements.
## Notes
Only a single number, boolean, or `#!cpp null` value is ever materialized into a `BasicJsonType`, to reuse its
`operator==` -- for numbers, so that values written differently in the source text but equal in value (e.g. an
integer and a floating-point literal) still compare equal, following the same rules `BasicJsonType` does for special
values such as `#!cpp NaN`. Constructing one of these scalars never allocates. Strings are compared directly, without
allocating, either from the source text on both sides or, for overload 2, against `BasicJsonType`'s own string.
Arrays and objects are never materialized at all; only their elements or members are visited, one pair at a time.
!!! info "How objects are compared"
For a [`json_view`](../json_view.md) (`BasicJsonType::object_t` is `#!cpp std::map`), members are compared by
key, regardless of the order they appear in the source text. For an
[`ordered_json_view`](../ordered_json_view.md) (`object_t` is `ordered_map`), they are compared in the order
they occur, so the very same two objects with their members reordered can compare equal as `json_view`s but not
as `ordered_json_view`s. This is exactly how [`json`](../json.md) and [`ordered_json`](../ordered_json.md)
compare, see ["Comparing different `basic_json` specializations"](../basic_json/operator_eq.md#notes).
!!! info "Discarded views"
A [discarded](is_discarded.md) view compares the same way a discarded `BasicJsonType` value does, which is
governed by
[`JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON`](../macros/json_use_legacy_discarded_value_comparison.md): by
default, a discarded view is never equal to anything, not even another discarded view.
No ordering comparison (`#!cpp operator<`) is provided for `basic_json_view`; [`materialize()`](materialize.md) is
the way to get a `BasicJsonType` value that supports it.
## Examples
??? example
The example below checks whether a newly received configuration differs from the previous one, and whether a
received document matches what a test expects -- directly on views, without ever materializing a `BasicJsonType`
value for either side.
```cpp
--8<-- "examples/basic_json_view__operator_eq.cpp"
```
Output:
```json
--8<-- "examples/basic_json_view__operator_eq.output"
```
## See also
- [operator!=](operator_ne.md) - compare for inequality
- [materialize](materialize.md) - build a `BasicJsonType` value, e.g. to keep comparing after the document is gone
- [`BasicJsonType::operator==`](../basic_json/operator_eq.md) - the corresponding function of `basic_json`
## Version history
- Added in version 3.13.0.
@@ -0,0 +1,74 @@
# <small>nlohmann::basic_json_view::</small>operator<<
```cpp
std::ostream& operator<<(std::ostream& o, const basic_json_view& v);
```
Not available when [`JSON_NO_IO`](../macros/json_no_io.md) is defined.
Serializes the given view `v` to the output stream `o`, using [`dump`](dump.md) -- exactly as
`#!cpp operator<<(std::ostream&, const basic_json&)` does for a `basic_json` value.
- The indentation of the output can be controlled with the member variable `width` of the output stream `o`. For
instance, using the manipulator `std::setw(4)` on `o` sets the indentation level to `4`, and the serialization
result is the same as calling `#!cpp v.dump(4)`. A `width` of `0` or less (the default) selects the most compact
representation, as `#!cpp v.dump(-1)` does.
- The indentation character can be controlled with the member variable `fill` of the output stream `o`. For instance,
the manipulator `std::setfill('\t')` sets indentation to use a tab character rather than the default space
character.
- As for `basic_json`, `o`'s `width` is reset to `0` after this call, whether or not it was greater than `0` before.
Numbers are always written as `#!cpp v.dump()` writes them by default, i.e. as with
[`number_format::shortest`](number_format.md); there is no way to select `#!cpp number_format::source` through the
stream.
## Parameters
`o` (in, out)
: stream to write to
`v` (in)
: view to serialize
## Return value
the stream `o`
## Exceptions
May throw `#!cpp std::bad_alloc`, propagated from [`dump`](dump.md#exceptions). Unlike
`#!cpp operator<<(std::ostream&, const basic_json&)`, there is no UTF-8 decoding step that could throw
[`type_error.316`](../../home/exceptions.md#jsonexceptiontype_error316), and no `error_handler` to choose between --
see the [Exceptions](dump.md#exceptions) of `dump`.
## Complexity
Linear, as [`dump`](dump.md#complexity).
## Examples
??? example
The example below writes one record out of a larger batch straight to a log stream -- compact for a one-line
entry, and pretty-printed with `std::setw`/`std::setfill` for a readable dump -- without ever building a
`BasicJsonType` value for the record, or for the rest of the batch.
```cpp
--8<-- "examples/basic_json_view__operator_ltlt.cpp"
```
Output:
```json
--8<-- "examples/basic_json_view__operator_ltlt.output"
```
## See also
- [dump](dump.md) - serialize to a JSON-formatted string
- [`operator<<(std::ostream&)`](../operator_ltlt.md) - the corresponding operator for `basic_json`
- [`JSON_NO_IO`](../macros/json_no_io.md) - switch off functions relying on certain C++ I/O headers
## Version history
- Added in version 3.13.0.
@@ -0,0 +1,82 @@
# <small>nlohmann::basic_json_view::</small>operator!=
```cpp
// (1)
bool operator!=(const basic_json_view& lhs, const basic_json_view& rhs);
// (2)
bool operator!=(const basic_json_view& lhs, const BasicJsonType& rhs);
bool operator!=(const BasicJsonType& lhs, const basic_json_view& rhs);
```
1. Compares two views for inequality. Returns `#!cpp !(lhs == rhs)`, see [operator==](operator_eq.md).
2. Compares a view and a `BasicJsonType` value for inequality, in either order. Returns `#!cpp !(lhs == rhs)` (or,
for the reversed order, `#!cpp !(rhs == lhs)`), see [operator==](operator_eq.md).
Since `operator!=` is defined as the negation of [`operator==`](operator_eq.md), it follows the same rules for
special cases: for instance, since a [discarded](is_discarded.md) view is never equal to anything by default (see
[operator=='s Notes](operator_eq.md#notes)), it is never *unequal* to anything either -- `#!cpp discarded != discarded`
is also `#!cpp false`, exactly as for a discarded `BasicJsonType` value.
## Parameters
`lhs` (in)
: first value to consider
`rhs` (in)
: second value to consider
## Return value
whether the values `lhs` and `rhs` are not equal
## Exception safety
Strong exception safety: if an exception is thrown, there are no changes to either operand, or to the document(s) a
view refers to.
## Exceptions
May throw `#!cpp std::bad_alloc`, propagated from [`operator==`](operator_eq.md#exceptions). Unlike most other
`basic_json_view` functions, `operator!=` is not `#!cpp noexcept`.
## Complexity
Linear, as [`operator==`](operator_eq.md#complexity).
## Notes
See the [Notes](operator_eq.md#notes) of `operator==` -- in particular for how an object's members are compared
(order matters for [`ordered_json_view`](../ordered_json_view.md) but not for [`json_view`](../json_view.md)) and
for how discarded views compare.
No ordering comparison (`#!cpp operator<`) is provided for `basic_json_view`; [`materialize()`](materialize.md) is
the way to get a `BasicJsonType` value that supports it.
## Examples
??? example
The example below asserts, as a test would, that a received document differs from an unwanted value, and shows
that -- as for [`json`](../json.md)/[`ordered_json`](../ordered_json.md) -- reordering an object's members is
detected as a difference for an `ordered_json_view` but not for a `json_view`.
```cpp
--8<-- "examples/basic_json_view__operator_ne.cpp"
```
Output:
```json
--8<-- "examples/basic_json_view__operator_ne.output"
```
## See also
- [operator==](operator_eq.md) - compare for equality
- [materialize](materialize.md) - build a `BasicJsonType` value, e.g. to keep comparing after the document is gone
- [`BasicJsonType::operator!=`](../basic_json/operator_ne.md) - the corresponding function of `basic_json`
## Version history
- Added in version 3.13.0.
@@ -70,6 +70,8 @@ None of these exceptions carry a [`JSON_DIAGNOSTICS`](../macros/json_diagnostics
1. Linear in the number of members: as for [`operator[]`](operator[].md#complexity), members are compared one after
another, in document order, stopping at the first match. Plus the complexity of converting the found member to
`T` (see [`get`](get.md)).
Objects with 128 or more members get a hash index while parsing, so that a lookup in them takes constant time
on average.
2. Linear in the number of reference tokens of `ptr` and, for each token, in the number of members of the object at
that level or the index into the array -- as for the [`operator[]`](operator[].md#complexity) and
[`at`](at.md#complexity) overloads that take a JSON pointer. Plus the complexity of converting the resolved value
@@ -0,0 +1,43 @@
# <small>nlohmann::</small>json_editable_document
<small>Defined in header `<nlohmann/json_view.hpp>`</small>
```cpp
using json_editable_document = basic_json_document<json, true>;
```
This type is an **editable** [`basic_json_document`](basic_json_document/index.md) of the default
[`json`](json.md) specialization: in addition to everything [`json_document`](json_document.md) offers,
[`set`](basic_json_document/set.md) and [`push_back`](basic_json_document/push_back.md) change values after
parsing, without ever rewriting the source text -- see [Edits](basic_json_document/index.md#edits) and
[Editing a document](../features/json_view.md#editing-a-document).
## Examples
??? example
The example below patches two fields of a small configuration document -- changing one and adding another --
and dumps it back out with the member order and the spelling of the untouched number preserved, something a
plain [`json`](json.md) value cannot do.
```cpp
--8<-- "examples/json_editable_document.cpp"
```
Output:
```json
--8<-- "examples/json_editable_document.output"
```
## See also
- [json_editable_view](json_editable_view.md) - a view of a value of a `json_editable_document`
- [json_document](json_document.md) - the read-only document this type adds edits to
- [ordered_json_editable_document](ordered_json_editable_document.md) - the corresponding editable document for
`ordered_json`
- [Edits](basic_json_document/index.md#edits) - what an edit guarantees
## Version history
Since version 3.13.0.
@@ -0,0 +1,41 @@
# <small>nlohmann::</small>json_editable_view
<small>Defined in header `<nlohmann/json_view.hpp>`</small>
```cpp
using json_editable_view = basic_json_view<json, true>;
```
This type is a [`basic_json_view`](basic_json_view/index.md) of a value of a
[`json_editable_document`](json_editable_document.md). It offers the same read-only interface as
[`json_view`](json_view.md); what is different is what it can be a view *of* -- a value that
[`set`](basic_json_document/set.md) and [`push_back`](basic_json_document/push_back.md) can change, with every view
still referring to it seeing the change, see [Edits](basic_json_document/index.md#edits).
## Examples
??? example
The example below is the same as [`json_editable_document`'s](json_editable_document.md): every view read back
out of the document -- `#!cpp doc.root()` and the views nested under it -- sees the edits made through `set`.
```cpp
--8<-- "examples/json_editable_document.cpp"
```
Output:
```json
--8<-- "examples/json_editable_document.output"
```
## See also
- [json_editable_document](json_editable_document.md) - the document type this view refers into
- [json_view](json_view.md) - the corresponding read-only view
- [ordered_json_editable_view](ordered_json_editable_view.md) - the corresponding view for
`ordered_json_editable_document`
## Version history
Since version 3.13.0.
+2
View File
@@ -33,6 +33,8 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
- [**JSON_USE_SIMDUTF**](json_use_simdutf.md) - use the simdutf library to accelerate UTF-8 validation
- [**JSON_VIEW_NO_SIMD**](json_view_no_simd.md) - use only portable code in the parser of `json_view.hpp`
- [**JSON_VIEW_USE_SSSE3**](json_view_use_ssse3.md) - validate non-ASCII strings with SSSE3 in the parser of `json_view.hpp`
## Library version
@@ -0,0 +1,50 @@
# JSON_VIEW_NO_SIMD
```cpp
#define JSON_VIEW_NO_SIMD
```
When defined, the parser of [`basic_json_document`](../basic_json_document/index.md) (`<nlohmann/json_view.hpp>`)
uses only portable C++ to scan strings. By default, it scans long runs of string bytes 16 at a time with NEON on
AArch64 (with GCC and Clang) and SSE2 on x86-64, which are part of the baseline instruction sets of these
architectures, and validates non-ASCII text with NEON (or SSSE3, see
[`JSON_VIEW_USE_SSSE3`](json_view_use_ssse3.md)).
The same input is accepted or rejected either way, with the same values, and errors are reported the same way; only
the speed differs. The macro exists for platforms whose compilers lack the intrinsics headers, and to test the portable
code.
!!! warning "Define consistently"
The macro selects between two definitions of the same inline functions. It must therefore be defined identically for
**every** translation unit that includes `<nlohmann/json_view.hpp>`; prefer a compile definition on the target.
## Default definition
By default, `#!cpp JSON_VIEW_NO_SIMD` is not defined, and the vector code is used where available.
```cpp
#undef JSON_VIEW_NO_SIMD
```
## Examples
??? example
The code below uses the portable string scanning of the view.
```cpp
#define JSON_VIEW_NO_SIMD
#include <nlohmann/json_view.hpp>
...
```
## See also
- [JSON_VIEW_USE_SSSE3](json_view_use_ssse3.md) - validate non-ASCII strings with SSSE3 on x86-64
- [json_view](../../features/json_view.md) - the zero-copy view
## Version history
- Added in version 3.13.0.
@@ -0,0 +1,49 @@
# JSON_VIEW_USE_SSSE3
```cpp
#define JSON_VIEW_USE_SSSE3
```
When defined on x86-64, the parser of [`basic_json_document`](../basic_json_document/index.md)
(`<nlohmann/json_view.hpp>`) validates non-ASCII text in strings with SSSE3, 16 bytes at a time, using the "lookup4"
algorithm of [simdjson](https://github.com/simdjson/simdjson). Without it, non-ASCII text is validated one UTF-8
sequence at a time on x86-64; on AArch64, the vector check uses NEON and is always on.
SSSE3 is not part of the x86-64 baseline, so the code must be compiled for it: define the macro only together with a
compiler option that enables SSSE3 (e.g. `-mssse3`, or `-march=` with a CPU that has it), and only for programs that
run on such CPUs. The same input is accepted or rejected either way; only the speed of non-ASCII text differs.
!!! warning "Define consistently"
The macro selects between two definitions of the same inline functions. It must therefore be defined identically,
with the same compiler options, for **every** translation unit that includes `<nlohmann/json_view.hpp>`; mixing
translation units that define it with ones that do not is an ODR violation. Prefer a compile definition on the
target.
## Default definition
By default, `#!cpp JSON_VIEW_USE_SSSE3` is not defined.
```cpp
#undef JSON_VIEW_USE_SSSE3
```
## Examples
??? example
With CMake, for a program that only runs on CPUs with SSSE3:
```cmake
target_compile_definitions(your_target PRIVATE JSON_VIEW_USE_SSSE3)
target_compile_options(your_target PRIVATE -mssse3)
```
## See also
- [JSON_VIEW_NO_SIMD](json_view_no_simd.md) - use only portable code in the view's parser
- [JSON_USE_SIMDUTF](json_use_simdutf.md) - validate UTF-8 with simdutf in `basic_json`'s parser
## Version history
- Added in version 3.13.0.
+2
View File
@@ -84,6 +84,8 @@ Linear.
## See also
- [dump](basic_json/dump.md) - serialize to a JSON-formatted string
- [`basic_json_view::operator<<`](basic_json_view/operator_ltlt.md) - the corresponding operator for
`basic_json_view`
- [Serialization](../features/serialization.md) - the serialization article
## Version history
@@ -0,0 +1,47 @@
# <small>nlohmann::</small>ordered_json_editable_document
<small>Defined in header `<nlohmann/json_view.hpp>`</small>
```cpp
using ordered_json_editable_document = basic_json_document<ordered_json, true>;
```
This type is an **editable** [`basic_json_document`](basic_json_document/index.md) of the
[`ordered_json`](ordered_json.md) specialization: [`set`](basic_json_document/set.md) and
[`push_back`](basic_json_document/push_back.md) change values after parsing, as for
[`json_editable_document`](json_editable_document.md), and
[`materialize()`](basic_json_view/materialize.md) preserves the document order of object members -- including
members [`set`](basic_json_document/set.md) added -- instead of sorting them like
[`json_editable_document`](json_editable_document.md) does.
## Examples
??? example
The example below edits a document with `set`, then shows that `materialize()` keeps the member order of the
source text (with the new member at the end) for `ordered_json_editable_document`, where it would sort the
members alphabetically for [`json_editable_document`](json_editable_document.md).
```cpp
--8<-- "examples/ordered_json_editable_document.cpp"
```
Output:
```json
--8<-- "examples/ordered_json_editable_document.output"
```
## See also
- [ordered_json_editable_view](ordered_json_editable_view.md) - a view of a value of an
`ordered_json_editable_document`
- [ordered_json_document](ordered_json_document.md) - the read-only document this type adds edits to
- [json_editable_document](json_editable_document.md) - the corresponding editable document for the default `json`
specialization
- [Object Order](../features/object_order.md)
- [Edits](basic_json_document/index.md#edits) - what an edit guarantees
## Version history
Since version 3.13.0.
@@ -0,0 +1,40 @@
# <small>nlohmann::</small>ordered_json_editable_view
<small>Defined in header `<nlohmann/json_view.hpp>`</small>
```cpp
using ordered_json_editable_view = basic_json_view<ordered_json, true>;
```
This type is a [`basic_json_view`](basic_json_view/index.md) of a value of an
[`ordered_json_editable_document`](ordered_json_editable_document.md), the corresponding view for
[`ordered_json_view`](ordered_json_view.md) the way [`json_editable_view`](json_editable_view.md) is for
[`json_view`](json_view.md).
## Examples
??? example
The example below is the same as [`ordered_json_editable_document`'s](ordered_json_editable_document.md): the
views `set` returns see the document's member order preserved on `materialize()`, unlike for a
[`json_editable_document`](json_editable_document.md).
```cpp
--8<-- "examples/ordered_json_editable_document.cpp"
```
Output:
```json
--8<-- "examples/ordered_json_editable_document.output"
```
## See also
- [ordered_json_editable_document](ordered_json_editable_document.md) - the document type this view refers into
- [ordered_json_view](ordered_json_view.md) - the corresponding read-only view
- [json_editable_view](json_editable_view.md) - the corresponding view for `json_editable_document`
## Version history
Since version 3.13.0.
@@ -0,0 +1,38 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_editable_document = nlohmann::json_editable_document;
using json_editable_view = nlohmann::json_editable_view;
int main()
{
// a deprecated field is dropped from a configuration file, and a
// decommissioned replica is removed from the list -- "price" keeps its
// trailing zero, and the fields around the removed ones keep their order
const std::string text = R"({
"name": "cache",
"legacy_host": "db0",
"host": "db1",
"price": 19.90,
"replicas": ["db2", "db3", "db4"]
})";
json_editable_document doc = json_editable_document::parse(text);
doc.erase(doc.root(), "legacy_host"); // (1) an object member
doc.erase(doc.root()["replicas"], 1); // (2) an array element ("db3")
const std::size_t removed = doc.erase(json::json_pointer("/replicas/0")); // (3) via a JSON pointer
std::cout << removed << '\n';
std::cout << doc.root().dump(2, ' ', false, json_editable_view::number_format::source) << "\n\n";
// the same edits on a plain json value: object_t is a std::map, so
// parsing already sorted the keys, and dump() rewrites every number to
// its shortest form, even "price", which was never touched
json plain = json::parse(text);
plain.erase("legacy_host");
plain["replicas"].erase(1);
plain["replicas"].erase(0);
std::cout << plain.dump(2) << '\n';
}
@@ -0,0 +1,18 @@
1
{
"name": "cache",
"host": "db1",
"price": 19.90,
"replicas": [
"db4"
]
}
{
"host": "db1",
"name": "cache",
"price": 19.9,
"replicas": [
"db4"
]
}
@@ -0,0 +1,38 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_editable_document = nlohmann::json_editable_document;
using json_editable_view = nlohmann::json_editable_view;
int main()
{
// a deployment plan -- "budget" is written with a trailing zero that has
// no effect on its value
const std::string text = R"({
"release": "2026.09",
"steps": ["build", "test", "deploy"],
"budget": 19.90
})";
json_editable_document doc = json_editable_document::parse(text);
const std::size_t deploy_index = 2;
const auto deploy = doc.root()["steps"][deploy_index]; // held across the insert
doc.insert(doc.root()["steps"], deploy_index, "smoke-test"); // insert before "deploy"
// the held view still refers to "deploy", even though its index moved
// from 2 to 3, and nothing else in the document was touched
std::cout << deploy.dump() << '\n';
std::cout << doc.root().dump(2, ' ', false, json_editable_view::number_format::source) << "\n\n";
// the same edit on a plain json value: an index held from before the
// insert now refers to whatever moved into that slot, and dump()
// rewrites "budget" to its shortest form even though it was never
// touched
json plain = json::parse(text);
plain["steps"].insert(plain["steps"].begin() + static_cast<std::ptrdiff_t>(deploy_index), "smoke-test");
std::cout << plain["steps"][deploy_index].dump() << '\n';
std::cout << plain.dump(2) << '\n';
}
@@ -0,0 +1,23 @@
"deploy"
{
"release": "2026.09",
"steps": [
"build",
"test",
"smoke-test",
"deploy"
],
"budget": 19.90
}
"smoke-test"
{
"budget": 19.9,
"release": "2026.09",
"steps": [
"build",
"test",
"smoke-test",
"deploy"
]
}
@@ -0,0 +1,52 @@
#include <cstdint>
#include <iostream>
#include <string>
#include <vector>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_document = nlohmann::json_document;
using image_check = json_document::image_check;
int main()
{
std::cout << std::boolalpha;
// the image of a parsed document -- as if read back from a cache file or
// received from another process running the same build of the library
const std::string text = R"({"name": "cache", "note": "caf\u00e9", "replicas": ["db2", "db3"]})";
const json_document parsed = json_document::parse(text);
const std::vector<std::uint8_t> image = parsed.save();
// (1)/(2) load() needs no parsing, yet dumps exactly what parsing did
const json_document borrowed = json_document::load(image);
std::cout << (borrowed.root().dump() == parsed.root().dump()) << '\n';
std::cout << borrowed.owns_source() << '\n'; // borrowed: still points into `image`
// (3) load(std::move(image)) keeps the vector instead of copying it
std::vector<std::uint8_t> to_move = image;
const json_document owned = json_document::load(std::move(to_move));
std::cout << owned.owns_source() << '\n';
// a damaged image -- the last byte of the decoded string "note" holds
// (an escape sequence, so it was unescaped into the document's own
// buffer), flipped, as storage or transport corruption might do
std::vector<std::uint8_t> damaged = image;
damaged[damaged.size() - 2] = 0xFF;
// image_check::full inspects strings and numbers, so it catches the damage
try
{
static_cast<void>(json_document::load(damaged, image_check::full));
}
catch (const json::parse_error& e)
{
std::cout << e.id << '\n';
}
// image_check::bounds only checks structure and bounds, so a cache the
// process already trusts loads without the extra scan -- reading a value
// the damage did not touch is still safe
const json_document trusted = json_document::load(damaged, image_check::bounds);
std::cout << trusted.root()["name"].get<std::string>() << '\n';
}
@@ -0,0 +1,5 @@
true
false
true
116
cache
@@ -0,0 +1,24 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_editable_document = nlohmann::json_editable_document;
int main()
{
// "events" starts out null -- the first push_back() turns it into an
// array, exactly like set() turns a null object member into an object
json_editable_document doc = json_editable_document::parse(R"({"source": "sensor-1", "events": null})");
const auto first = doc.push_back(doc.root()["events"], json{{"type", "start"}, {"t", 0}});
for (int t = 1; t <= 3; ++t)
{
doc.push_back(doc.root()["events"], json{{"type", "tick"}, {"t", t}});
}
// push_back() never moves an existing element: a view taken from an
// earlier call still refers to the same element after later ones
std::cout << first.dump() << '\n';
std::cout << doc.root()["events"].size() << '\n';
std::cout << doc.root().dump(2) << '\n';
}
@@ -0,0 +1,23 @@
{"t":0,"type":"start"}
4
{
"source": "sensor-1",
"events": [
{
"t": 0,
"type": "start"
},
{
"t": 1,
"type": "tick"
},
{
"t": 2,
"type": "tick"
},
{
"t": 3,
"type": "tick"
}
]
}
@@ -0,0 +1,30 @@
#include <cstdint>
#include <iostream>
#include <string>
#include <vector>
#include <nlohmann/json_view.hpp>
using json_document = nlohmann::json_document;
int main()
{
std::cout << std::boolalpha;
// a configuration a service parses once and then caches as an image, so
// that later requests can load() it instead of parsing the text again
const std::string text = R"({"name": "cache", "host": "db1", "port": 6379, "replicas": ["db2", "db3"]})";
const json_document config = json_document::parse(text);
// save() turns the parsed document into a byte buffer: a 64-byte header,
// the node index, the source text, and the decoded strings
const std::vector<std::uint8_t> image = config.save();
std::cout << image.size() << '\n';
// the same document always saves to the same bytes
std::cout << (image == json_document::parse(text).save()) << '\n';
// loading the image back needs no parsing, yet dumps exactly what
// parsing the text produced
const json_document reloaded = json_document::load(image);
std::cout << (reloaded.root().dump() == config.root().dump()) << '\n';
}
@@ -0,0 +1,3 @@
316
true
true
@@ -0,0 +1,41 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_editable_document = nlohmann::json_editable_document;
using json_editable_view = nlohmann::json_editable_view;
int main()
{
// a configuration file, as it might be read from disk -- "price" is
// written with a trailing zero that has no effect on its value
const std::string text = R"({
"name": "cache",
"host": "db1",
"port": 6379,
"price": 19.90,
"replicas": ["db2", "db3"],
"timeout": 30
})";
json_editable_document doc = json_editable_document::parse(text);
doc.set(doc.root()["port"], 6380); // (1) replace a value
doc.set(doc.root(), "region", "us-east"); // (2) add a member
doc.set(doc.root()["replicas"], 0, "db4"); // (3) assign an element
doc.set(json::json_pointer("/timeout"), 45); // (4) via a JSON pointer
// members stay in document order (the new one at the end), and a number
// that was not itself edited keeps its exact spelling
std::cout << doc.root().dump(2, ' ', false, json_editable_view::number_format::source) << "\n\n";
// the same edits on a plain json value: object_t is a std::map, so
// parsing already sorted the keys, and dump() rewrites every number to
// its shortest form, even "price", which was never touched
json plain = json::parse(text);
plain["port"] = 6380;
plain["region"] = "us-east";
plain["replicas"][0] = "db4";
plain[json::json_pointer("/timeout")] = 45;
std::cout << plain.dump(2) << '\n';
}
@@ -0,0 +1,25 @@
{
"name": "cache",
"host": "db1",
"port": 6380,
"price": 19.90,
"replicas": [
"db4",
"db3"
],
"timeout": 45,
"region": "us-east"
}
{
"host": "db1",
"name": "cache",
"port": 6380,
"price": 19.9,
"region": "us-east",
"replicas": [
"db4",
"db3"
],
"timeout": 45
}
@@ -0,0 +1,25 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json_document = nlohmann::json_document;
using json_view = nlohmann::json_view;
int main()
{
// a large batch of sensor readings -- forward just the one that changed,
// without ever building a basic_json value for the batch or for the
// readings that are not needed
const json_document batch = json_document::parse(R"(
[{"id": 1, "temp": 21.5}, {"id": 2, "temp": 87.3}, {"id": 3, "temp": 21.7}]
)");
const json_view readings = batch.root();
std::cout << readings[1].dump() << '\n';
// a configuration file -- dump() on the view keeps the member order of
// the source text; a json value's object_t is std::map, so
// materialize().dump() of the very same view sorts the keys instead
const json_document config = json_document::parse(
R"({"name": "cache", "host": "db1", "port": 6379, "timeout": 30})");
std::cout << config.root().dump(2) << "\n\n";
std::cout << config.root().materialize().dump(2) << '\n';
}
@@ -0,0 +1,14 @@
{"id":2,"temp":87.3}
{
"name": "cache",
"host": "db1",
"port": 6379,
"timeout": 30
}
{
"host": "db1",
"name": "cache",
"port": 6379,
"timeout": 30
}
@@ -0,0 +1,28 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json_document = nlohmann::json_document;
using json_view = nlohmann::json_view;
int main()
{
// a price list received from a supplier feed -- prices and account
// numbers must be forwarded exactly, e.g. into an invoice
const json_document doc = json_document::parse(R"(
[{"sku": "A1", "price": 19.90, "account_id": 12345678901234567890123456},
{"sku": "A2", "price": 1E2, "account_id": 98765432109876543210987654}]
)");
const json_view list = doc.root();
// number_format::shortest (the default) writes numbers the way
// basic_json::dump() would: "19.90" becomes "19.9", "1E2" becomes
// "100.0", and each account number -- far beyond any 64-bit integer --
// is rounded to the nearest double, exactly as materialize().dump()
// (or a plain nlohmann::json) would round it
std::cout << list.dump() << '\n';
// number_format::source copies every number exactly as it was written
// in the source text instead -- something basic_json cannot do at all,
// since parsing already reduces every number to its parsed value
std::cout << list.dump(-1, ' ', false, json_view::number_format::source) << '\n';
}
@@ -0,0 +1,2 @@
[{"sku":"A1","price":19.9,"account_id":1.2345678901234568e+25},{"sku":"A2","price":100.0,"account_id":9.876543210987655e+25}]
[{"sku":"A1","price":19.90,"account_id":12345678901234567890123456},{"sku":"A2","price":1E2,"account_id":98765432109876543210987654}]
@@ -0,0 +1,30 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json_document = nlohmann::json_document;
using json = nlohmann::json;
int main()
{
// two snapshots of a polled configuration endpoint -- compare them
// directly as views, without ever building a nlohmann::json value for
// either one
const json_document previous = json_document::parse(
R"({"name": "cache", "port": 6379, "timeout": 30})");
const json_document current = json_document::parse(
R"({"port": 6379.0, "timeout": 30, "name": "cache"})");
// same members, reordered, and 6379 written as a float -- operator==
// treats them the same way BasicJsonType::operator== would
std::cout << std::boolalpha << (previous.root() == current.root()) << '\n';
// an actually changed value is detected the same way
const json_document changed = json_document::parse(
R"({"name": "cache", "port": 6380, "timeout": 30})");
std::cout << (previous.root() == changed.root()) << '\n';
// comparing a view directly against an expected json value -- handy in a
// test, without materializing the received document at all
const json expected = {{"name", "cache"}, {"port", 6379}, {"timeout", 30}};
std::cout << (previous.root() == expected) << '\n';
}
@@ -0,0 +1,3 @@
true
false
true
@@ -0,0 +1,26 @@
#include <iostream>
#include <iomanip>
#include <nlohmann/json_view.hpp>
using json_document = nlohmann::json_document;
using json_view = nlohmann::json_view;
int main()
{
// one order out of a large incoming batch -- write it straight to a log
// stream without ever building a basic_json value for it, or for the
// rest of the batch
const json_document doc = json_document::parse(R"(
[{"id": 1, "item": "cable"}, {"id": 2, "item": "adapter"}]
)");
const json_view orders = doc.root();
// compact, for a one-line log entry
std::cout << orders[1] << '\n';
// std::setw sets the indentation level, exactly as for basic_json
std::cout << std::setw(2) << orders[1] << "\n\n";
// std::setfill changes the indentation character
std::cout << std::setw(1) << std::setfill('\t') << orders[1] << '\n';
}
@@ -0,0 +1,10 @@
{"id":2,"item":"adapter"}
{
"id": 2,
"item": "adapter"
}
{
"id": 2,
"item": "adapter"
}
@@ -0,0 +1,28 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json_document = nlohmann::json_document;
using ordered_json_document = nlohmann::ordered_json_document;
using json = nlohmann::json;
int main()
{
// assert, as a test would, that a received document differs from an
// unwanted shape -- without ever materializing it into a json value just
// to compare
const json_document received = json_document::parse(
R"({"status": "ok", "code": 200})");
const json unwanted = {{"status", "error"}, {"code", 500}};
std::cout << std::boolalpha << (received.root() != unwanted) << '\n';
// json (std::map) compares object members regardless of order ...
const json_document a = json_document::parse(R"({"a": 1, "b": 2})");
const json_document b = json_document::parse(R"({"b": 2, "a": 1})");
std::cout << (a.root() != b.root()) << '\n';
// ... but ordered_json (ordered_map) compares them in the order they
// appear, so the very same reordering is detected as a difference
const ordered_json_document oa = ordered_json_document::parse(R"({"a": 1, "b": 2})");
const ordered_json_document ob = ordered_json_document::parse(R"({"b": 2, "a": 1})");
std::cout << (oa.root() != ob.root()) << '\n';
}
@@ -0,0 +1,3 @@
true
false
true
@@ -0,0 +1,30 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using json = nlohmann::json;
using json_editable_document = nlohmann::json_editable_document;
using json_editable_view = nlohmann::json_editable_view;
int main()
{
// a configuration file, as it might be read from disk
const std::string text = R"({"name": "cache", "host": "db1", "port": 6379, "price": 19.90})";
std::cout << text << "\n\n";
// patch two fields -- "price" is never touched
json_editable_document doc = json_editable_document::parse(text);
doc.set(doc.root(), "host", "db2");
doc.set(doc.root(), "retries", 3);
// member order (the new member at the end) and the untouched number's
// exact spelling survive
std::cout << doc.root().dump(-1, ' ', false, json_editable_view::number_format::source) << '\n';
// the same patch on a plain json value: keys are sorted (object_t is a
// std::map), and "price" is rewritten even though the patch never
// touched it
json plain = json::parse(text);
plain["host"] = "db2";
plain["retries"] = 3;
std::cout << plain.dump() << '\n';
}
@@ -0,0 +1,4 @@
{"name": "cache", "host": "db1", "port": 6379, "price": 19.90}
{"name":"cache","host":"db2","port":6379,"price":19.90,"retries":3}
{"host":"db2","name":"cache","port":6379,"price":19.9,"retries":3}
@@ -0,0 +1,15 @@
#include <iostream>
#include <nlohmann/json_view.hpp>
using ordered_json_editable_document = nlohmann::ordered_json_editable_document;
int main()
{
// ordered_json_editable_document is basic_json_document<nlohmann::ordered_json, true>
ordered_json_editable_document doc = ordered_json_editable_document::parse(R"({"z": 1, "a": 2, "m": 3})");
doc.set(doc.root(), "b", 4); // set() always appends a new member at the end
// materialize() preserves the document order (with "b" at the end),
// instead of sorting the keys the way json_editable_document does
std::cout << doc.root().materialize().dump() << '\n';
}
@@ -0,0 +1 @@
{"z":1,"a":2,"m":3,"b":4}
+2
View File
@@ -13,6 +13,8 @@ C++ types, and finally serialize it again.
[SAX interface](parsing/sax_interface.md), and [error handling](parsing/parse_exceptions.md).
- [Zero-copy JSON views](json_view.md) — read a JSON text through a flat index instead of building a `json` tree;
strings and numbers stay in the input and are only decoded when needed.
[Editable documents](json_view.md#editing-a-document) can also be modified, and [images](json_view.md#images) load a
parsed document again without parsing it.
- [Comments](comments.md) and [trailing commas](trailing_commas.md) — opt-in relaxations of the JSON grammar.
## Accessing and modifying values
+144 -8
View File
@@ -139,8 +139,12 @@ whenever any of the other conditions above was not met.
element access and lookup functions never carry the JSON Pointer path `JSON_DIAGNOSTICS` would otherwise add: the
view has no `basic_json` value to point at, so the exception is created without one, regardless of how
`BasicJsonType` was built.
- **`dump()` and comparison are not (yet) provided** by `basic_json_view`. For now,
[`materialize()`](../api/basic_json_view/materialize.md) is the way to get a value you can do those things with.
- **Ordering comparisons are not provided** by `basic_json_view` -- there is no `#!cpp operator<`.
[`operator==`](../api/basic_json_view/operator_eq.md) and [`operator!=`](../api/basic_json_view/operator_ne.md) are
provided, though: two views, or a view and a `BasicJsonType` value, compare equal exactly when
[`materialize()`](../api/basic_json_view/materialize.md) or [`parse()`](../api/basic_json/parse.md) would produce
equal values for them, without ever building a tree to do it. For ordering, too,
[`materialize()`](../api/basic_json_view/materialize.md) is the way to get a value you can compare.
## Getting values out without copying
@@ -166,14 +170,146 @@ Two conversions never copy at all:
Both results are only valid as long as the view -- and, for a string with no escapes, the borrowed source text -- is.
## Writing a view back
[`dump()`](../api/basic_json_view/dump.md) serializes a view directly from the flat index, without ever building a
`basic_json` value. An object's members are written in document order, not sorted by key, and *every* occurrence of a
repeated key is written, not only the last one -- the same two ways [iteration](#what-is-different) already differs
from a [`materialize()`](../api/basic_json_view/materialize.md)d value, see above. `#!cpp materialize().dump()` gives
a different result in both respects for a `json_view`.
By default, numbers are written the way [`basic_json::dump()`](../api/basic_json/dump.md) would.
[`number_format::source`](../api/basic_json_view/number_format.md) instead copies every number exactly as it was
written in the source text -- a price like `#!cpp 19.90`, a long order or account ID with more digits than any number
type holds, or a high-precision coordinate -- something `basic_json` cannot do at all, since parsing already reduces
a number to its parsed `#!cpp double`/`#!cpp int64_t` value.
[`operator<<`](../api/basic_json_view/operator_ltlt.md) writes a view to a stream the way `basic_json`'s does, using
the stream's `width`/`fill` for indentation.
## Editing a document
Everything above is read-only: a `json_document`/`json_view` lets you look at a parsed text without copying it, but
not change it. [`basic_json_document<BasicJsonType, true>`](../api/basic_json_document/index.md) -- more conveniently
spelled [`json_editable_document`](../api/json_editable_document.md) or
[`ordered_json_editable_document`](../api/ordered_json_editable_document.md) -- also lets you
[`set`](../api/basic_json_document/set.md) a value, [`push_back`](../api/basic_json_document/push_back.md) onto or
[`insert`](../api/basic_json_document/insert.md) into an array, and [`erase`](../api/basic_json_document/erase.md)
an object member or an array element, still without ever building a `basic_json` tree for parts you do not touch.
`#!cpp Editable` defaults to `#!cpp false`, so `json_document`/`ordered_json_document` are unaffected -- they carry
none of the bookkeeping edits need, and calling `set`/`push_back`/`insert`/`erase` on one is a compile error, not a
runtime one.
### Why: editing without reformatting
The `#!cpp 19.90` price from [above](#writing-a-view-back) is exactly the kind of value that makes editing a `json`
or `ordered_json` value in place lossy. Say you parse a configuration file, patch one field, and write it back:
- **`json`** re-sorts every key on the way in (`object_t` is a `#!cpp std::map`) and rewrites every number to its
shortest round-trip form on the way out -- a one-field patch turns into a diff that reorders the whole file and
rewrites `#!cpp 19.90` to `#!cpp 19.9`.
- **`ordered_json`** keeps the key order, but still rewrites every number the same way: parsing has already reduced
it to a `#!cpp double`/`#!cpp int64_t`, and there is no way back to how it was spelled in the source text.
An editable document keeps both. [`dump()`](../api/basic_json_view/dump.md) of an edited document writes members in
document order -- a member [`set`](../api/basic_json_document/set.md) added goes at the end, exactly where it was
inserted, and an [`erase`](../api/basic_json_document/erase.md)d member simply leaves a gap: everything around it
keeps its place -- and [`number_format::source`](../api/basic_json_view/number_format.md) keeps the exact spelling
of every number an edit did not itself touch; a number an edit *did* touch is written the way
[`BasicJsonType::dump()`](../api/basic_json/dump.md) would write it, since there is no source spelling for a brand
new value.
??? example "Example: patch a configuration, keeping member order and number spellings"
```cpp
--8<-- "examples/json_editable_document.cpp"
```
Output:
```json
--8<-- "examples/json_editable_document.output"
```
### What stays valid, and what an edit costs
The source text itself is **never written**, and the parsed index never moves -- a value keeps the node it was
parsed into for as long as it is not itself replaced. So every [view](../api/basic_json_view/index.md) taken before
an edit, including a previously obtained [`root()`](../api/basic_json_document/root.md), stays valid and, if it
still refers to the edited value, sees the edit; a view of a value a later edit drops or replaces just keeps showing
what it last held. New values go to storage the document allocates and owns on demand. The one thing an edit does
invalidate is the **iterators** taken over an edited array or object: the first time one of its elements is set,
appended to, inserted into, or erased, its elements move from the parsed, fixed layout to a growable block of links
so that [`push_back`](../api/basic_json_document/push_back.md) can later grow it in amortized constant time --
existing elements are not touched, but an iterator that was walking the old layout no longer matches. A string
obtained with
[`get_string()`](../api/basic_json_view/get_string.md) is unaffected either way and stays valid across further
edits. See [`basic_json_document`'s Edits](../api/basic_json_document/index.md#edits) for the details, and
[`set`'s Exception safety](../api/basic_json_document/set.md#exception-safety) for what an edit guarantees if it
throws (the *basic* guarantee, not the strong one `dump()` and the read-only functions provide). How edits are kept in
the index is described in the [architecture overview](../home/architecture.md#node-index-of-json-views).
## Images
[`save()`](../api/basic_json_document/save.md) writes a document as an *image*: a byte buffer that
[`load()`](../api/basic_json_document/load.md) reads back into a document without parsing -- no lexing, no building
the node index, nothing but copying the nodes and pointing the text and the decoded strings at the image. Where
[`parse_copy()`](../api/basic_json_document/parse_copy.md) still has to scan the whole input,
[`load()`](../api/basic_json_document/load.md) turns that scan into a copy of the node index alone.
**Why.** A document that is parsed once and then read many times -- a configuration loaded at startup, a template
rendered on every request, a large reference dataset a worker process needs in memory -- pays for parsing once but
can amortize [`save()`](../api/basic_json_document/save.md)'s cost across every later load. That makes images useful
for a cache: save a document the first time it is parsed (to a file, a shared-memory segment, an in-process cache),
and [`load()`](../api/basic_json_document/load.md) it on every later use instead of parsing the source text again.
They are just as useful for handing a parsed document to another process (or a forked worker) running the same build
of the library, since [`load()`](../api/basic_json_document/load.md) turns the transfer into a copy of the node index
plus pointers into the received bytes, not a re-parse.
**Choosing a check.** [`load()`](../api/basic_json_document/load.md) takes an
[`image_check`](../api/basic_json_document/load.md#image_check) that trades validation against speed:
`image_check::full` (the default) checks everything the parser itself guarantees, so a checked image is exactly as
safe to read and serialize as a freshly parsed document -- the right choice whenever the image did not come straight
from this process's own [`save()`](../api/basic_json_document/save.md), such as a file or a network peer.
`image_check::bounds` only checks structure and bounds -- cheaper, since it skips scanning the text and the decoded
strings -- and fits a cache the process trusts, one it wrote and reads back itself. `image_check::none` skips
validation entirely, for an image trusted as much as the process's own memory. See
[`load()`'s Notes](../api/basic_json_document/load.md#notes) for exactly what each level does and does not guarantee.
??? example "Example: cache a parsed configuration as an image"
```cpp
--8<-- "examples/basic_json_document__save.cpp"
```
Output:
```json
--8<-- "examples/basic_json_document__save.output"
```
!!! warning "Experimental"
The image format is versioned but not yet stable, and may change in an incompatible way before it is declared
stable. It is little-endian only, and tied to the library build that wrote it -- use it to cache a document or to
hand one to another process running the *same* build, not as a long-term storage format; keep the original JSON
text if a saved document needs to be readable by a future library version.
The idea of a document you can read without parsing comes from zero-copy formats such as
[FlatBuffers](https://github.com/google/flatbuffers) and [YaFF](https://github.com/yandex/yaff); the check
[`load()`](../api/basic_json_document/load.md) runs follows the idea of FlatBuffers' Verifier. No code is taken from
either.
## Choosing between `json`, `ordered_json`, the SAX interface, and `json_view`
| | [`json`](../api/json.md) / [`ordered_json`](../api/ordered_json.md) | [SAX interface](parsing/sax_interface.md) | [`json_document`](../api/json_document.md) / [`json_view`](../api/json_view.md) |
|---|---|---|---|
| **Ownership** | owns every value | owns nothing; you decide what to keep, in your handler | borrows or owns the *text*; the index is always owned by the document |
| **Mutability** | freely mutable | not applicable (a one-shot event stream) | read-only |
| **What you get** | a full tree you can read, write, and keep as long as you like | a sequence of callbacks; whatever your handler builds from them | a flat index plus, on demand, [`materialize()`](../api/basic_json_view/materialize.md)d `json`/`ordered_json` values for the parts you actually use |
| **Typical use** | general-purpose JSON handling: config, request/response bodies you build or modify, anything you hold onto | validating or projecting a text into your own data structure without ever holding the whole thing as JSON | large or high-volume input where you only need part of it, or need it repeatedly, and can keep the source text (or a copy) alive for as long as the document lives |
| | [`json`](../api/json.md) / [`ordered_json`](../api/ordered_json.md) | [SAX interface](parsing/sax_interface.md) | [`json_document`](../api/json_document.md) / [`json_view`](../api/json_view.md) | [`json_editable_document`](../api/json_editable_document.md) / [`json_editable_view`](../api/json_editable_view.md) |
|---|---|---|---|---|
| **Ownership** | owns every value | owns nothing; you decide what to keep, in your handler | borrows or owns the *text*; the index is always owned by the document | same as `json_document`; edits go to storage the document owns |
| **Mutability** | freely mutable | not applicable (a one-shot event stream) | read-only | [`set`](../api/basic_json_document/set.md)/[`push_back`](../api/basic_json_document/push_back.md)/[`insert`](../api/basic_json_document/insert.md)/[`erase`](../api/basic_json_document/erase.md) edit in place; the source text is never rewritten |
| **What you get** | a full tree you can read, write, and keep as long as you like | a sequence of callbacks; whatever your handler builds from them | a flat index plus, on demand, [`materialize()`](../api/basic_json_view/materialize.md)d `json`/`ordered_json` values for the parts you actually use | the same, plus [`dump()`](../api/basic_json_view/dump.md) of an edited document that keeps the member order and, with [`number_format::source`](../api/basic_json_view/number_format.md), the spelling of every untouched number |
| **Typical use** | general-purpose JSON handling: config, request/response bodies you build or modify, anything you hold onto | validating or projecting a text into your own data structure without ever holding the whole thing as JSON | large or high-volume input where you only need part of it, or need it repeatedly, and can keep the source text (or a copy) alive for as long as the document lives | a document you read, patch a few fields of, and write back -- a configuration file, for instance -- where the rest of it should come back exactly as it was |
| **Caching/reload** | not applicable -- re-parse, or roll your own serialization | not applicable | [`save()`](../api/basic_json_document/save.md)/[`load()`](../api/basic_json_document/load.md): cache the parsed index as an image and reload it without parsing | same, saving the document's current -- possibly edited -- state |
## Version history
+31 -3
View File
@@ -194,9 +194,9 @@ packet-beta
| Bytes | Field | Type | Contents |
|-------|---------|------------|-------------------------------------------------------------------------------------------------------------------------------|
| 0 | `kind` | `uint8_t` | the type, numbered as [`value_t`](../api/basic_json/value_t.md): 0 null, 1 object, 2 array, 3 string, 4 boolean, 5 signed integer, 6 unsigned integer, 7 float |
| 1 | `flags` | `uint8_t` | bits 0-1: where a string's bytes are (0: the source text, 1: the buffer of decoded strings, for strings with escapes); bit 2: the value of a boolean |
| 2-3 | `extra` | `uint16_t` | numbers: the number of integer digits (low byte) and fraction digits (high byte), 255 for more; otherwise 0 |
| 0 | `kind` | `uint8_t` | the type, numbered as [`value_t`](../api/basic_json/value_t.md): 0 null, 1 object, 2 array, 3 string, 4 boolean, 5 signed integer, 6 unsigned integer, 7 float; 10 for a link (see below) |
| 1 | `flags` | `uint8_t` | bits 0-1: where a string's bytes (or a number's token) are (0: the source text, 1: the buffer of decoded strings, for strings with escapes, 2: the edit buffer); bit 2: the value of a boolean; bits 3 and 4: moved and new (see below) |
| 2-3 | `extra` | `uint16_t` | numbers: the number of integer digits (low byte) and fraction digits (high byte), 255 for more; objects: the number of their hash index (1-based), or 0; otherwise 0 |
| 4-7 | `off` | `uint32_t` | where the value starts: the first byte after a string's opening quote (or its position in the buffer of decoded strings), the first byte of a number or literal, the bracket of an array or object |
| 8-11 | `len` | `uint32_t` | strings: the length after decoding; floats and literals: the length of the token; arrays and objects: the number of elements |
| 12-15 | `next` | `uint32_t` | arrays and objects: the number of nodes of the subtree, including the node itself |
@@ -211,6 +211,11 @@ packet-beta
subtree is `next` nodes further for an array or object, and the next node otherwise (`document_data::after`). Views
step from element to element this way and skip whole subtrees in constant time.
- **Offsets** are 32 bits wide, so a document is limited to 4 GiB (`out_of_range.416`).
- **Large objects** (128 members or more) get a hash index after parsing
([`detail/view/object_index.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/view/object_index.hpp)):
an open-addressing table whose slots hold the distance from the object's node to a key's node, so that a lookup does
not compare every key. The object's `extra` holds the number of its table. Only 65,535 tables fit into `extra`;
objects beyond them are searched linearly.
For example, `#!json {"a": [1, 2.5]}` becomes five nodes. Each node's elements follow it, and `next` leads from an
array or object past its subtree:
@@ -239,6 +244,29 @@ flowchart LR
All `flags` are 0. The integer's bytes 8-15 hold its value, 1; its `extra` says it has one digit. The float's `extra`
says it has one integer and one fraction digit, and its `len` is that of the token `2.5`.
Editable documents ([`json_editable_document`](../api/json_editable_document.md),
[`detail/view/edit.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/view/edit.hpp) and
[`detail/view/edit_storage.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/view/edit_storage.hpp))
never write the source text and never move or resize the parsed index, so views stay valid while the document is
edited:
- A new scalar is written over its node. Its text (a string, or the token of a number as `dump()` writes it) goes to
the edit buffer, which `flags` bits 0-1 then name.
- An array or object whose elements change gets the flag *moved* (bit 3): its elements then live in a separate
sequence (a header node, then the entries), whose number is in `off`. The entries are links (`kind` 10), whose bytes
8-15 hold the address of the value's node, so values never move.
- A node written by an edit gets the flag *new* (bit 4): it has no position in the source text.
- Views of read-only documents compile without any of this: how views walk the index is a template parameter
(`navigation<Editable>`).
Images ([`save`](../api/basic_json_document/save.md) and [`load`](../api/basic_json_document/load.md),
[`detail/view/image.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/view/image.hpp)) store
the nodes as they are: a 64-byte header (the magic bytes `NJVI`, a format version, the sizes, and reserved bytes that
must be zero), the nodes, the text, and the decoded strings. An edited document is first written in document order, as
the parser would have written it (without links), and the numbers of the hash indexes are cleared, since `load`
rebuilds the indexes. So a change of the node layout is a change of the image format: it must raise `image_version`,
and `load` then rejects images of other versions (`parse_error.116`) instead of misreading them.
## Input adapters
Input is read via **input adapters** that abstract a source. Every input adapter provides this interface:
+68 -2
View File
@@ -388,6 +388,23 @@ A UBJSON high-precision number could not be parsed.
[json.exception.parse_error.115] parse error at byte 5: syntax error while parsing UBJSON high-precision number: invalid number text: 1A
```
### json.exception.parse_error.116
[`basic_json_document::load()`](../api/basic_json_document/load.md) rejected an
[image](../features/json_view.md#images): either the bytes are not one [`save()`](../api/basic_json_document/save.md)
could have written (too short, an unknown magic number or format version, or sizes that do not fit the buffer), or
they are, but fail the requested [`image_check`](../api/basic_json_document/load.md#image_check).
!!! failure "Example message"
```
[json.exception.parse_error.116] parse error: invalid json_document image: the check failed
```
!!! note
This exception was added in version 3.13.0, together with [images](../features/json_view.md#images).
## Iterator errors
This exception is thrown if iterators passed to a library function do not match
@@ -782,6 +799,44 @@ The dynamic type of the object cannot be represented in the requested serializat
Encapsulate the JSON value in an object. That is, instead of serializing `#!json true`, serialize `#!json {"value": true}`
### json.exception.type_error.319
[`basic_json_document::set`](../api/basic_json_document/set.md) and
[`basic_json_document::push_back`](../api/basic_json_document/push_back.md) can store any `basic_json` value except
a binary one: a `json_document` has no representation for [binary values](../features/binary_values.md), which only
ever arise from parsing a binary format or from an explicit [`json::binary`](../api/basic_json/binary.md) value, not
from JSON text.
!!! failure "Example message"
```
[json.exception.type_error.319] cannot store a binary value in a json_document
```
!!! note
This exception was added in version 3.13.0, together with editable [`json_document`s](../features/json_view.md).
### json.exception.type_error.320
[`basic_json_document::save()`](../api/basic_json_document/save.md) cannot write an
[image](../features/json_view.md#images) of a [discarded](../api/basic_json_document/is_discarded.md) document.
[`save()`](../api/basic_json_document/save.md) and [`load()`](../api/basic_json_document/load.md) also throw this
exception on a big-endian target, since the image format is little-endian only.
!!! failure "Example messages"
```
[json.exception.type_error.320] cannot save a discarded json_document
```
```
[json.exception.type_error.320] json_document images need a little-endian target
```
!!! note
This exception was added in version 3.13.0, together with [images](../features/json_view.md#images).
## Out of range
This exception is thrown in case a library function is called on an input parameter that exceeds the expected range, for instance, in the case of array indices or nonexisting object keys.
@@ -1006,13 +1061,24 @@ MessagePack's ext type and BSON's binary subtype are each stored in a single byt
[`basic_json_document::parse()`](../api/basic_json_document/parse.md) and the other parsing functions of
[`basic_json_document`](../api/basic_json_document/index.md) index a value's position in the source text in 32 bits,
so they do not support an input of 4 GiB or more.
so they do not support an input of 4 GiB or more. The same 32-bit limit applies to an **editable** document's own
storage: [`set`](../api/basic_json_document/set.md) and [`push_back`](../api/basic_json_document/push_back.md) throw
this exception once the strings and number tokens written by edits reach 4 GiB in total, or once more than
4294967295 arrays/objects have had an element set or appended to them. The same limit applies to an
[image](../features/json_view.md#images): [`save()`](../api/basic_json_document/save.md) throws it if the node
count, the text, or the decoded strings it would write would individually reach 4 GiB.
!!! failure "Example message"
!!! failure "Example messages"
```
[json.exception.out_of_range.416] input of 4 GiB or more is not supported by json_document
```
```
[json.exception.out_of_range.416] edits of 4 GiB or more are not supported by json_document
```
```
[json.exception.out_of_range.416] images of 4 GiB or more are not supported by json_document
```
!!! note
+2
View File
@@ -23,3 +23,5 @@ The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Eva
The class contains an adapted version of the Eisel-Lemire algorithm and its table of powers of five from [fast_float](https://github.com/fastfloat/fast_float) by Daniel Lemire and contributors, which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here), the Apache 2.0 License, and the Boost Software License. Copyright &copy; 2021 The fast_float authors
The view's parser (`<nlohmann/json_view.hpp>`) contains techniques and code adapted from [yyjson](https://github.com/ibireme/yyjson) by YaoYuan, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above): table-driven decoding of `\u` escapes and fixed-offset unrolled checks.
The view's parser (`<nlohmann/json_view.hpp>`) validates non-ASCII strings with the vector UTF-8 check of [simdjson](https://github.com/simdjson/simdjson) by Daniel Lemire, Geoff Langdale, John Keiser, and contributors (its "lookup4" algorithm and tables, after J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One Instruction Per Byte", 2021), which is available under the [MIT License](https://opensource.org/licenses/MIT) (used here) and the Apache 2.0 License. Copyright &copy; 2018-2025 The simdjson authors
+17
View File
@@ -233,14 +233,20 @@ nav:
- 'Overview': api/basic_json_document/index.md
- '(Constructor)': api/basic_json_document/basic_json_document.md
- 'accept': api/basic_json_document/accept.md
- 'erase': api/basic_json_document/erase.md
- 'insert': api/basic_json_document/insert.md
- 'is_discarded': api/basic_json_document/is_discarded.md
- 'load': api/basic_json_document/load.md
- 'memory_usage': api/basic_json_document/memory_usage.md
- 'node_count': api/basic_json_document/node_count.md
- 'owns_source': api/basic_json_document/owns_source.md
- 'parse': api/basic_json_document/parse.md
- 'parse_copy': api/basic_json_document/parse_copy.md
- 'push_back': api/basic_json_document/push_back.md
- 'read': api/basic_json_document/read.md
- 'root': api/basic_json_document/root.md
- 'save': api/basic_json_document/save.md
- 'set': api/basic_json_document/set.md
- 'shrink_to_fit': api/basic_json_document/shrink_to_fit.md
- 'source': api/basic_json_document/source.md
- basic_json_view:
@@ -253,6 +259,7 @@ nav:
- 'cend': api/basic_json_view/cend.md
- 'contains': api/basic_json_view/contains.md
- 'count': api/basic_json_view/count.md
- 'dump': api/basic_json_view/dump.md
- 'empty': api/basic_json_view/empty.md
- 'end': api/basic_json_view/end.md
- 'find': api/basic_json_view/find.md
@@ -275,9 +282,13 @@ nav:
- 'is_structured': api/basic_json_view/is_structured.md
- 'items': api/basic_json_view/items.md
- 'materialize': api/basic_json_view/materialize.md
- 'number_format': api/basic_json_view/number_format.md
- 'number_token': api/basic_json_view/number_token.md
- 'operator bool': api/basic_json_view/operator_bool.md
- 'operator<<': api/basic_json_view/operator_ltlt.md
- 'operator[]': api/basic_json_view/operator[].md
- 'operator==': api/basic_json_view/operator_eq.md
- 'operator!=': api/basic_json_view/operator_ne.md
- 'size': api/basic_json_view/size.md
- 'source_offset': api/basic_json_view/source_offset.md
- 'type': api/basic_json_view/type.md
@@ -296,6 +307,8 @@ nav:
- 'to_json': api/adl_serializer/to_json.md
- 'json': api/json.md
- 'json_document': api/json_document.md
- 'json_editable_document': api/json_editable_document.md
- 'json_editable_view': api/json_editable_view.md
- json_pointer:
- 'Overview': api/json_pointer/index.md
- '(Constructor)': api/json_pointer/json_pointer.md
@@ -336,6 +349,8 @@ nav:
- 'operator""_json_pointer': api/operator_literal_json_pointer.md
- 'ordered_json': api/ordered_json.md
- 'ordered_json_document': api/ordered_json_document.md
- 'ordered_json_editable_document': api/ordered_json_editable_document.md
- 'ordered_json_editable_view': api/ordered_json_editable_view.md
- 'ordered_json_view': api/ordered_json_view.md
- 'ordered_map': api/ordered_map.md
- macros:
@@ -363,6 +378,8 @@ nav:
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
- 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md
- 'JSON_USE_SIMDUTF': api/macros/json_use_simdutf.md
- 'JSON_VIEW_NO_SIMD': api/macros/json_view_no_simd.md
- 'JSON_VIEW_USE_SSSE3': api/macros/json_view_use_ssse3.md
- 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md
- 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md
- 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md
+23 -7
View File
@@ -115,6 +115,13 @@ class builder
frame shallow[64]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays): not initialized on purpose; filled as containers open
std::vector<frame> deep{};
/// remember an object to index after parsing (out of line, so that the
/// parse loop only has a call for it)
NLOHMANN_VIEW_NOINLINE void note_large_object(std::uint32_t idx)
{
doc.large_objects.push_back(idx);
}
NLOHMANN_VIEW_NOINLINE bool fail(error_code c, const unsigned char* at) noexcept
{
m_failure.code = c;
@@ -492,7 +499,7 @@ class builder
switch (cur()) \
{ \
case '"': \
if (NLOHMANN_VIEW_UNLIKELY(!string())) { return false; } \
if (NLOHMANN_VIEW_UNLIKELY(!string<true>())) { return false; } \
goto NEXT; \
case '{': \
open(value_t::object); \
@@ -576,7 +583,7 @@ obj_key:
{
return fail(error_code::expected_key);
}
if (NLOHMANN_VIEW_UNLIKELY(!string()))
if (NLOHMANN_VIEW_UNLIKELY(!string<false>()))
{
return false;
}
@@ -617,19 +624,27 @@ obj_next:
if (TrailingCommas && cur() == '}')
{
++p;
goto close_container;
goto close_object;
}
goto obj_key;
}
if (cur() == '}')
{
++p;
goto close_container;
goto close_object;
}
return fail(error_code::expected_object_end);
#undef NLOHMANN_VIEW_VALUE
close_object:
// a large object gets a hash index (objects only, so that closing
// an array pays nothing for this)
if (NLOHMANN_VIEW_UNLIKELY(cur_count >= document_data::index_min_members))
{
cold.note_large_object(cur_idx);
}
close_container:
close();
if (NLOHMANN_VIEW_UNLIKELY(depth == 0))
@@ -679,7 +694,7 @@ root_done:
switch (cur())
{
case '"':
return string();
return string<true>();
case 't':
return literal("true", 4, value_t::boolean, node_flags::is_true);
case 'f':
@@ -965,12 +980,13 @@ indent_done:
return true;
}
/// a string (value or key) at p
/// a string at p: a value (Value) or a key
template<bool Value>
NLOHMANN_VIEW_ALWAYS_INLINE bool string()
{
++p; // opening quote
const unsigned char* const s = p;
p = scan_string_run(p, e);
p = scan_string_run<Value>(p, e);
if (NLOHMANN_VIEW_LIKELY(p != e && *p == '"'))
{
emit(value_t::string, 0, 0, static_cast<std::size_t>(s - b), static_cast<std::uint64_t>(p - s));
+317
View File
@@ -0,0 +1,317 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <algorithm> // sort, stable_sort
#include <cstddef> // size_t
#include <string> // string
#include <utility> // move, pair
#include <vector> // vector
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
// Equality of views, and of views with basic_json values, with the semantics
// of basic_json's operator== applied to the values parse() would produce:
// numbers compare by value across their types, an object is compared by its
// members with duplicate keys resolved as parse() resolves them (the last
// value, at the position of the first occurrence), and in document order if
// the object type keeps an order (ordered_json), by key otherwise.
/// one side of a comparison: a view
template<typename BasicJsonType, typename View>
class view_side
{
public:
using string_view_t = typename View::string_view_t;
explicit view_side(const View& v) noexcept
: m_view(v)
{}
value_t type() const noexcept
{
return m_view.type();
}
std::size_t size() const noexcept
{
return m_view.size();
}
string_view_t string() const
{
return m_view.get_string();
}
/// a number, boolean, or null as a basic_json value (no allocation)
BasicJsonType scalar() const
{
switch (m_view.type())
{
case value_t::number_integer:
return BasicJsonType(m_view.template get<typename BasicJsonType::number_integer_t>());
case value_t::number_unsigned:
return BasicJsonType(m_view.template get<typename BasicJsonType::number_unsigned_t>());
case value_t::number_float:
return BasicJsonType(m_view.template get<typename BasicJsonType::number_float_t>());
case value_t::boolean:
return BasicJsonType(m_view.template get<bool>());
case value_t::null:
case value_t::object:
case value_t::array:
case value_t::string:
case value_t::binary:
case value_t::discarded:
default:
return BasicJsonType(nullptr);
}
}
void elements(std::vector<view_side>& out) const
{
out.reserve(m_view.size());
for (const View e : m_view)
{
out.emplace_back(e);
}
}
/// the members as parse() keeps them: one per key, the last value at the
/// position of the first occurrence; in that order, or sorted by key
void members(std::vector<std::pair<string_view_t, view_side>>& out, bool ordered) const
{
struct member
{
string_view_t key;
View value;
std::size_t position;
};
std::vector<member> all;
all.reserve(m_view.size());
std::size_t position = 0;
for (auto it = m_view.begin(); it != m_view.end(); ++it)
{
all.push_back(member{it.key(), it.value(), position++});
}
std::stable_sort(all.begin(), all.end(), [](const member & a, const member & b)
{
return a.key < b.key;
});
std::vector<member> unique;
unique.reserve(all.size());
for (std::size_t i = 0; i < all.size();)
{
std::size_t last = i;
while (last + 1 < all.size() && all[last + 1].key == all[i].key)
{
++last;
}
unique.push_back(member{all[i].key, all[last].value, all[i].position});
i = last + 1;
}
if (ordered)
{
std::sort(unique.begin(), unique.end(), [](const member & a, const member & b)
{
return a.position < b.position;
});
}
out.reserve(unique.size());
for (const member& m : unique)
{
out.emplace_back(m.key, view_side(m.value));
}
}
private:
View m_view;
};
/// the other side of a comparison: a basic_json value
template<typename BasicJsonType, typename StringView>
class json_side
{
public:
using string_view_t = StringView;
explicit json_side(const BasicJsonType& j) noexcept
: m_json(&j)
{}
value_t type() const noexcept
{
return m_json->type();
}
std::size_t size() const noexcept
{
return m_json->size();
}
string_view_t string() const
{
const auto& s = m_json->template get_ref<const typename BasicJsonType::string_t&>();
return string_view_t(s.data(), s.size());
}
BasicJsonType scalar() const
{
return *m_json;
}
void elements(std::vector<json_side>& out) const
{
out.reserve(m_json->size());
for (const auto& e : *m_json)
{
out.emplace_back(e);
}
}
void members(std::vector<std::pair<string_view_t, json_side>>& out, bool ordered) const
{
out.reserve(m_json->size());
for (auto it = m_json->cbegin(); it != m_json->cend(); ++it)
{
out.emplace_back(string_view_t(it.key().data(), it.key().size()), json_side(it.value()));
}
if (!ordered)
{
std::sort(out.begin(), out.end(), [](const std::pair<string_view_t, json_side>& a, const std::pair<string_view_t, json_side>& b)
{
return a.first < b.first;
});
}
}
private:
const BasicJsonType* m_json;
};
/// whether two sides are equal; iterative, so that the nesting depth is
/// limited by memory only
template<typename BasicJsonType, typename A, typename B>
bool equal(const A& a0, const B& b0)
{
using string_view_t = typename A::string_view_t;
const bool ordered = is_ordered_map<typename BasicJsonType::object_t>::value;
struct frame
{
std::vector<A> elements_a{};
std::vector<B> elements_b{};
std::vector<std::pair<string_view_t, A>> members_a{};
std::vector<std::pair<string_view_t, B>> members_b{};
bool object = false;
std::size_t next = 0;
};
std::vector<frame> stack;
A a = a0;
B b = b0;
for (;;)
{
const value_t ta = a.type();
const value_t tb = b.type();
const bool numbers = (ta == value_t::number_integer || ta == value_t::number_unsigned || ta == value_t::number_float)
&& (tb == value_t::number_integer || tb == value_t::number_unsigned || tb == value_t::number_float);
if (ta == value_t::discarded || tb == value_t::discarded)
{
// basic_json decides (JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON)
if (ta != tb || !(BasicJsonType(value_t::discarded) == BasicJsonType(value_t::discarded)))
{
return false;
}
}
else
{
if (!numbers && ta != tb)
{
return false;
}
if (ta == value_t::string)
{
if (!(a.string() == b.string()))
{
return false;
}
}
else if (ta == value_t::array || ta == value_t::object)
{
if (a.size() != b.size() && ta == value_t::array)
{
return false;
}
frame f;
f.object = ta == value_t::object;
if (f.object)
{
a.members(f.members_a, ordered);
b.members(f.members_b, ordered);
if (f.members_a.size() != f.members_b.size())
{
return false;
}
}
else
{
a.elements(f.elements_a);
b.elements(f.elements_b);
}
stack.push_back(std::move(f));
}
else if (!(a.scalar() == b.scalar())) // numbers (also of different types), null, boolean
{
return false;
}
}
// the next pair of values
for (;;)
{
if (stack.empty())
{
return true;
}
frame& f = stack.back();
const std::size_t count = f.object ? f.members_a.size() : f.elements_a.size();
if (f.next == count)
{
stack.pop_back();
continue;
}
if (f.object)
{
if (!(f.members_a[f.next].first == f.members_b[f.next].first))
{
return false;
}
a = f.members_a[f.next].second;
b = f.members_b[f.next].second;
}
else
{
a = f.elements_a[f.next];
b = f.elements_b[f.next];
}
++f.next;
break;
}
}
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+112 -1
View File
@@ -10,9 +10,14 @@
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t
#include <cstring> // memcpy
#include <functional> // less
#include <map> // map
#include <memory> // unique_ptr
#include <new> // operator new, placement new
#include <string> // string
#include <vector> // vector
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
@@ -36,10 +41,44 @@ struct document_data
node* inline_tape = nullptr; ///< node array allocated together with this header
std::size_t inline_cap = 0;
std::string arena{}; ///< decoded strings that contained escapes // NOLINT(readability-redundant-member-init)
std::size_t arena_size = 0; ///< bytes of decoded strings at base[1] (the arena, or those of a loaded image)
std::string owned{}; ///< owned copy of the input, if any // NOLINT(readability-redundant-member-init)
std::array<const char*, 4> base = {{nullptr, nullptr, nullptr, nullptr}}; ///< string bases: source, arena (indexed by flags & node_flags::storage)
std::vector<std::uint8_t> owned_image{}; ///< a loaded image the document owns (the text and the decoded strings point into it) // NOLINT(readability-redundant-member-init)
// hash indexes of large objects (see object_index.hpp)
static constexpr std::uint32_t index_min_members = 128;
struct object_index
{
std::size_t start; ///< first slot in index_slots
std::uint32_t mask; ///< slot count - 1 (a power of two minus one)
};
std::vector<object_index> indexes{}; // NOLINT(readability-redundant-member-init)
std::vector<std::uint32_t> index_slots{}; // NOLINT(readability-redundant-member-init)
std::vector<std::uint32_t> large_objects{}; ///< positions of the objects to index (noted while parsing) // NOLINT(readability-redundant-member-init)
std::array<const char*, 4> base = {{nullptr, nullptr, nullptr, nullptr}}; ///< string bases: source, arena, edit arena (indexed by flags & node_flags::storage)
bool discarded = true;
/// The storage of edits (editable documents only; see edit_storage.hpp).
/// Edits never move or resize the parsed index, so views stay valid: an
/// array/object whose elements change gets node_flags::moved, and its
/// elements then live in a separate sequence (a header node, then the
/// entries), whose entries link to the values.
struct edit_state
{
std::vector<node*> moved{}; ///< element sequences of moved arrays/objects (header node first) // NOLINT(readability-redundant-member-init)
std::vector<std::size_t> moved_cap{}; ///< capacity in nodes of a growable block; 0: a fixed sequence (a new value) // NOLINT(readability-redundant-member-init)
std::vector<std::unique_ptr<node[]>> chunks{}; ///< storage of new values and blocks; never moved // NOLINT(readability-redundant-member-init,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
std::map<const node*, node*, std::less<const node*>> regions{}; ///< new arrays/objects: root -> container that uses it as its element sequence (nullptr: linked from a block) // NOLINT(readability-redundant-member-init)
node* chunk_cur = nullptr;
node* chunk_end = nullptr;
std::size_t chunk_next = 64;
std::vector<std::unique_ptr<char[]>> texts{}; ///< edit arena, the current buffer last; earlier ones stay alive for string views // NOLINT(readability-redundant-member-init,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
std::size_t text_used = 0;
std::size_t text_cap = 0;
std::size_t bytes = 0; ///< memory held by edits
};
std::unique_ptr<edit_state> edits{}; ///< created by the first edit // NOLINT(readability-redundant-member-init)
/// one allocation for the header and room for `nodes` nodes; large
/// documents get a separate node array instead (so it can be trimmed)
static document_data* create(std::size_t nodes)
@@ -123,6 +162,78 @@ struct document_data
{
return n + n->next;
}
/// (editable documents) first element or key, also of a moved container
NLOHMANN_VIEW_ALWAYS_INLINE const node* first_child_edited(const node* n) const noexcept
{
return NLOHMANN_VIEW_LIKELY((n->flags & node_flags::moved) == 0) ? n + 1 : edits->moved[n->off] + 1;
}
/// (editable documents) end of the elements, also of a moved container
NLOHMANN_VIEW_ALWAYS_INLINE const node* child_end_edited(const node* n) const noexcept
{
if (NLOHMANN_VIEW_LIKELY((n->flags & node_flags::moved) == 0))
{
return n + n->next;
}
const node* const h = edits->moved[n->off];
return h + h->next;
}
/// (editable documents) the value at an element position: entries of
/// moved sequences are links. The link case is out of line, so that this
/// compiles to a predicted branch rather than a select that delays the
/// following loads.
static NLOHMANN_VIEW_ALWAYS_INLINE const node* deref(const node* n) noexcept
{
return NLOHMANN_VIEW_LIKELY(n->kind != kind_link) ? n : follow_link(n);
}
static NLOHMANN_VIEW_NOINLINE const node* follow_link(const node* n) noexcept
{
return link_target(*n);
}
};
/// How the index is walked: views of read-only documents follow the node
/// array alone and compile without any of the edit handling; views of
/// editable documents also follow moved element sequences and links.
template<bool Editable>
struct navigation
{
static NLOHMANN_VIEW_ALWAYS_INLINE const node* first(const document_data& /*d*/, const node* n) noexcept
{
return n + 1;
}
static NLOHMANN_VIEW_ALWAYS_INLINE const node* end(const document_data& /*d*/, const node* n) noexcept
{
return n + n->next;
}
static NLOHMANN_VIEW_ALWAYS_INLINE const node* value(const node* n) noexcept
{
return n;
}
};
template<>
struct navigation<true>
{
static NLOHMANN_VIEW_ALWAYS_INLINE const node* first(const document_data& d, const node* n) noexcept
{
return d.first_child_edited(n);
}
static NLOHMANN_VIEW_ALWAYS_INLINE const node* end(const document_data& d, const node* n) noexcept
{
return d.child_end_edited(n);
}
static NLOHMANN_VIEW_ALWAYS_INLINE const node* value(const node* n) noexcept
{
return document_data::deref(n);
}
};
} // namespace view
+767
View File
@@ -0,0 +1,767 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cmath> // isinf, isnan
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint8_t, uint32_t, uint64_t
#include <cstring> // memcmp, memmove
#include <limits> // numeric_limits
#include <string> // string, to_string
#include <type_traits> // decay, enable_if, integral_constant, is_arithmetic, is_convertible, is_floating_point, is_same, is_signed
#include <utility> // forward
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/edit_storage.hpp>
#include <nlohmann/detail/view/errors.hpp>
#include <nlohmann/detail/view/lookup.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
template<typename BasicJsonType, bool Editable>
class basic_json_view;
namespace detail
{
namespace view
{
/// the index of the first true condition (the number of conditions if none is)
template<bool... Conditions>
struct first_true : std::integral_constant<int, 0> {};
template<bool... Conditions>
struct first_true<false, Conditions...> : std::integral_constant < int, 1 + first_true<Conditions...>::value > {};
/// Checks a string the way basic_json's serializer does when it writes it
/// (type_error.316 with the same message), so that an editable document
/// only holds valid UTF-8: the error is at the first byte that no
/// well-formed sequence can continue with (Unicode, Table 3-7).
inline void check_utf8(const char* s, std::size_t n)
{
const auto* const p = reinterpret_cast<const unsigned char*>(s); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
const auto hex = [](unsigned char c)
{
constexpr const char* digits = "0123456789ABCDEF";
return std::string{digits[c >> 4u], digits[c & 0xFu]};
};
for (std::size_t i = 0; i < n;)
{
const unsigned char c = p[i];
if (c < 0x80)
{
++i;
continue;
}
std::size_t len = 0;
unsigned char lo = 0x80;
unsigned char hi = 0xBF;
if (c >= 0xC2 && c <= 0xDF)
{
len = 2;
}
else if (c >= 0xE0 && c <= 0xEF)
{
len = 3;
lo = c == 0xE0 ? 0xA0 : 0x80;
hi = c == 0xED ? 0x9F : 0xBF;
}
else if (c >= 0xF0 && c <= 0xF4)
{
len = 4;
lo = c == 0xF0 ? 0x90 : 0x80;
hi = c == 0xF4 ? 0x8F : 0xBF;
}
else
{
throw_type_error(316, concat("invalid UTF-8 byte at index ", std::to_string(i), ": 0x", hex(c)));
}
for (std::size_t k = 1; k < len; ++k)
{
if (i + k == n)
{
throw_type_error(316, concat("incomplete UTF-8 string; last byte: 0x", hex(p[n - 1])));
}
const unsigned char b = p[i + k];
if (b < (k == 1 ? lo : 0x80) || b > (k == 1 ? hi : 0xBF))
{
throw_type_error(316, concat("invalid UTF-8 byte at index ", std::to_string(i + k), ": 0x", hex(b)));
}
}
i += len;
}
}
/*!
@brief the edits of an editable basic_json_document
Values are accepted as views (of any document), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written: new values go to storage owned by the document (see
edit_storage.hpp).
*/
template<typename BasicJsonType, typename View>
class editor
{
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t;
using string_t = typename BasicJsonType::string_t;
using string_view_t = typename View::string_view_t;
using nav = navigation<true>;
public:
explicit editor(document_data& d) noexcept
: m_doc(d)
{}
/// replace a value; returns its view
template<typename V>
View set(const View& target, V&& value)
{
node* const slot = own(target);
const encoded e = encode(std::forward<V>(value));
assign(slot, e, nullptr, false);
return View(&m_doc, slot);
}
/// set a member (appended if missing; a null value becomes an object);
/// returns a view of the member value
template<typename V>
View set(const View& object, string_view_t key, V&& value)
{
node* const o = own(object);
if (o->kind != static_cast<std::uint8_t>(value_t::object) && o->kind != static_cast<std::uint8_t>(value_t::null))
{
throw_type_error(305, "cannot use operator[] with a string argument with ", object.type_name());
}
check_utf8(key.data(), key.size());
const encoded e = encode(std::forward<V>(value));
if (o->kind == static_cast<std::uint8_t>(value_t::null))
{
become_empty(o, value_t::object);
}
// an existing member: assign it (and drop later duplicates, so that
// lookups, iteration, and materialize() agree)
node* slot = nullptr;
bool duplicates = false;
for (const node* k = nav::first(m_doc, o), *end = nav::end(m_doc, o); k != end; k = document_data::after(k + 1))
{
if (key_equals(*k, key))
{
if (slot != nullptr)
{
duplicates = true;
break;
}
slot = const_cast<node*>(nav::value(k + 1)); // NOLINT(cppcoreguidelines-pro-type-const-cast): the nodes belong to this document
}
}
if (slot != nullptr)
{
if (duplicates)
{
erase_members(o, key, true);
}
assign(slot, e, o, true);
return View(&m_doc, slot);
}
const node k = string_node(key.data(), key.size());
slot = new_slot(e);
node* const h = block_of(m_doc, o, 2);
h[h->next] = k;
make_link(h[h->next + 1], slot);
h->next += 2;
++h->len;
++o->len;
return View(&m_doc, slot);
}
/// assign an existing array element; returns a view of it
template<typename V>
View set(const View& array, std::size_t idx, V&& value)
{
node* const a = own(array);
if (a->kind != static_cast<std::uint8_t>(value_t::array))
{
throw_type_error(305, "cannot use operator[] with a numeric argument with ", array.type_name());
}
check_index(idx, a->len);
const encoded e = encode(std::forward<V>(value));
node* const slot = const_cast<node*>(nav::value(element_at<true>(m_doc, a, idx))); // NOLINT(cppcoreguidelines-pro-type-const-cast)
assign(slot, e, a, true);
return View(&m_doc, slot);
}
/// append to an array (a null value becomes an array); returns a view of
/// the new element
template<typename V>
View push_back(const View& array, V&& value)
{
node* const a = own(array);
if (a->kind != static_cast<std::uint8_t>(value_t::array) && a->kind != static_cast<std::uint8_t>(value_t::null))
{
throw_type_error(308, "cannot use push_back() with ", array.type_name());
}
const encoded e = encode(std::forward<V>(value));
if (a->kind == static_cast<std::uint8_t>(value_t::null))
{
become_empty(a, value_t::array);
}
node* const slot = new_slot(e);
node* const h = block_of(m_doc, a, 1);
make_link(h[h->next], slot);
++h->next;
++h->len;
++a->len;
return View(&m_doc, slot);
}
/// insert into an array before position idx (idx <= size()); returns a
/// view of the new element
template<typename V>
View insert(const View& array, std::size_t idx, V&& value)
{
node* const a = own(array);
if (a->kind != static_cast<std::uint8_t>(value_t::array))
{
throw_type_error(309, "cannot use insert() with ", array.type_name());
}
check_index(idx, a->len + 1);
const encoded e = encode(std::forward<V>(value));
node* const slot = new_slot(e);
node* const h = block_of(m_doc, a, 1);
std::memmove(h + 2 + idx, h + 1 + idx, (h->next - 1 - idx) * sizeof(node));
make_link(h[1 + idx], slot);
++h->next;
++h->len;
++a->len;
return View(&m_doc, slot);
}
/// remove all members with this key; returns their number
std::size_t erase(const View& object, string_view_t key)
{
node* const o = own(object);
if (o->kind != static_cast<std::uint8_t>(value_t::object))
{
throw_type_error(307, "cannot use erase() with ", object.type_name());
}
for (const node* k = nav::first(m_doc, o), *end = nav::end(m_doc, o); k != end; k = document_data::after(k + 1))
{
if (key_equals(*k, key))
{
return erase_members(o, key, false);
}
}
return 0;
}
/// remove an array element
void erase(const View& array, std::size_t idx)
{
node* const a = own(array);
if (a->kind != static_cast<std::uint8_t>(value_t::array))
{
throw_type_error(307, "cannot use erase() with ", array.type_name());
}
check_index(idx, a->len);
node* const h = block_of(m_doc, a, 0);
std::memmove(h + 1 + idx, h + 2 + idx, (h->next - 2 - idx) * sizeof(node));
--h->next;
--h->len;
--a->len;
}
private:
/// an encoded value: a scalar node, or the root of a new array/object
struct encoded
{
node scalar{};
node* region = nullptr;
};
/// the node of a view of this document
node* own(const View& v)
{
if (NLOHMANN_VIEW_UNLIKELY(v.m_doc != &m_doc || v.m_node == nullptr))
{
throw_invalid_iterator(202, "view does not belong to this document");
}
edit_state_of(m_doc);
return const_cast<node*>(v.m_node); // NOLINT(cppcoreguidelines-pro-type-const-cast): the nodes belong to this document
}
static void check_index(std::size_t idx, std::size_t limit)
{
if (idx >= limit)
{
throw_out_of_range(401, concat("array index ", std::to_string(idx), " is out of range"));
}
}
bool key_equals(const node& k, string_view_t key) const noexcept
{
return k.len == key.size() && (key.size() == 0 || std::memcmp(m_doc.str(k), key.data(), key.size()) == 0);
}
/// remove the members with this key (all, or all but the first) from an object
std::size_t erase_members(node* o, string_view_t key, bool keep_first)
{
node* const h = block_of(m_doc, o, 0);
node* w = h + 1;
std::size_t erased = 0;
bool kept = false;
for (node* r = h + 1, *end = h + h->next; r != end; r += 2)
{
const bool match = key_equals(*r, key);
if (match && (kept || !keep_first))
{
++erased;
continue;
}
kept = kept || match;
if (w != r)
{
w[0] = r[0];
w[1] = r[1];
}
w += 2;
}
h->next = static_cast<std::uint32_t>(w - h);
h->len -= static_cast<std::uint32_t>(erased);
o->len -= static_cast<std::uint32_t>(erased);
return erased;
}
/// turn a null into an empty array/object in place
static void become_empty(node* n, value_t k) noexcept
{
*n = node{};
n->kind = static_cast<std::uint8_t>(k);
n->flags = node_flags::is_new;
n->next = 1;
}
/// Replace the value at slot; `parent` is the container whose elements
/// include slot (if known).
void assign(node* slot, const encoded& e, node* parent, bool parent_known)
{
if (e.region == nullptr)
{
if (is_container(*slot) && slot->next > 1 && slot != m_doc.tape)
{
// The slot spans its old elements in the enclosing sequence, but
// a scalar is one node: the enclosing container first switches to
// links (then the extent of the slot no longer matters).
node* const p = parent_known ? parent : find_parent(m_doc, slot);
if (p != nullptr && ((p->flags & node_flags::moved) == 0 || moved_capacity(m_doc, p) == 0))
{
block_of(m_doc, p, 0);
}
}
*slot = e.scalar;
return;
}
// an array/object: the slot keeps its extent (so that the enclosing
// sequence still steps over it), and the elements come from the new
// sequence
const node* const r = e.region;
const std::uint32_t extent = is_container(*slot) ? slot->next : 1;
const bool was_moved = (slot->flags & node_flags::moved) != 0;
slot->kind = r->kind;
slot->extra = 0;
slot->len = r->len;
slot->next = extent;
slot->flags = was_moved ? static_cast<std::uint8_t>(node_flags::moved | node_flags::is_new) : std::uint8_t{0};
set_moved(m_doc, slot, e.region, 0);
edit_state_of(m_doc).regions[e.region] = slot;
}
/// a node for a new element (links point to it; it never moves)
node* new_slot(const encoded& e)
{
if (e.region != nullptr)
{
return e.region;
}
node* const s = alloc_nodes(m_doc, 1);
*s = e.scalar;
return s;
}
//////////////
// encoding //
//////////////
template<int N>
using encode_tag = std::integral_constant<int, N>;
template<typename T>
struct is_view : std::false_type {};
template<typename J, bool E>
struct is_view<basic_json_view<J, E>> : std::true_type {};
template<typename V>
encoded encode(V&& v)
{
using D = typename std::decay<V>::type;
return encode_impl(std::forward<V>(v), encode_tag<first_true<is_view<D>::value,
std::is_same<D, BasicJsonType>::value,
std::is_same<D, std::nullptr_t>::value,
std::is_same<D, bool>::value,
std::is_arithmetic<D>::value,
std::is_convertible<const D&, string_view_t>::value>::value> {});
}
/// a view of any document (copied; nothing is shared with it)
template<typename J, bool E>
encoded encode_impl(const basic_json_view<J, E>& v, encode_tag<0> /*view*/)
{
if (NLOHMANN_VIEW_UNLIKELY(v.m_node == nullptr))
{
throw_type_error(302, "type must be a value, but is ", "discarded");
}
encoded r;
if (!is_container(*v.m_node))
{
r.scalar = copy_scalar(*v.m_doc, *v.m_node);
return r;
}
r.region = alloc_nodes(m_doc, count_nodes<E>(*v.m_doc, v.m_node));
fill_nodes<E>(*v.m_doc, v.m_node, r.region);
edit_state_of(m_doc).regions.emplace(r.region, nullptr);
return r;
}
encoded encode_impl(const BasicJsonType& j, encode_tag<1> /*json*/)
{
encoded r;
if (!j.is_structured())
{
r.scalar = json_scalar(j);
return r;
}
r.region = alloc_nodes(m_doc, count_nodes(j));
fill_nodes(j, r.region);
edit_state_of(m_doc).regions.emplace(r.region, nullptr);
return r;
}
encoded encode_impl(std::nullptr_t /*unused*/, encode_tag<2> /*null*/)
{
encoded r;
r.scalar = plain_node(value_t::null);
return r;
}
encoded encode_impl(bool b, encode_tag<3> /*boolean*/)
{
encoded r;
r.scalar = plain_node(value_t::boolean);
r.scalar.flags = static_cast<std::uint8_t>(r.scalar.flags | (b ? node_flags::is_true : 0));
return r;
}
template<typename T>
encoded encode_impl(T x, encode_tag<4> /*number*/)
{
encoded r;
r.scalar = number_node(x, std::integral_constant<int, first_true<std::is_floating_point<T>::value, std::is_signed<T>::value>::value> {});
return r;
}
template<typename T>
encoded encode_impl(const T& s, encode_tag<5> /*string*/)
{
const string_view_t sv(s);
check_utf8(sv.data(), sv.size());
encoded r;
r.scalar = string_node(sv.data(), sv.size());
return r;
}
template<typename T>
encoded encode_impl(T&& x, encode_tag<6> /*other*/)
{
return encode_impl(BasicJsonType(std::forward<T>(x)), encode_tag<1> {});
}
static node plain_node(value_t k) noexcept
{
node n{};
n.kind = static_cast<std::uint8_t>(k);
n.flags = node_flags::is_new;
return n;
}
template<typename T>
node number_node(T x, std::integral_constant<int, 0> /*floating-point*/)
{
return float_node(static_cast<number_float_t>(x));
}
template<typename T>
node number_node(T x, std::integral_constant<int, 1> /*signed*/)
{
return integer_node(static_cast<std::uint64_t>(static_cast<std::int64_t>(x)), value_t::number_integer);
}
template<typename T>
node number_node(T x, std::integral_constant<int, 2> /*unsigned*/)
{
return integer_node(static_cast<std::uint64_t>(x), value_t::number_unsigned);
}
/// an integer with its canonical token in the edit arena
node integer_node(std::uint64_t bits, value_t k)
{
const bool negative = k == value_t::number_integer && static_cast<std::int64_t>(bits) < 0;
std::uint64_t magnitude = negative ? 0 - bits : bits;
std::array<char, 24> buf{};
char* p = buf.data() + buf.size();
do
{
*--p = static_cast<char>('0' + (magnitude % 10));
magnitude /= 10;
}
while (magnitude != 0);
if (negative)
{
*--p = '-';
}
const auto len = static_cast<std::size_t>(buf.data() + buf.size() - p);
node n = plain_node(k);
n.flags = static_cast<std::uint8_t>(n.flags | node_flags::edited);
n.off = append_text(m_doc, p, len);
// number_length() adds one for the sign of number_integer nodes
n.extra = static_cast<std::uint16_t>(k == value_t::number_integer ? len - 1 : len);
set_integer_bits(n, bits);
return n;
}
/// a float with its shortest round-trip token (as basic_json::dump()
/// writes it), or nan, inf, -inf, in the edit arena
node float_node(number_float_t x)
{
string_t text;
if (std::isnan(x))
{
text = "nan";
}
else if (std::isinf(x))
{
text = x > 0 ? "inf" : "-inf";
}
else
{
text = BasicJsonType(x).dump();
}
node n = plain_node(value_t::number_float);
n.flags = static_cast<std::uint8_t>(n.flags | node_flags::edited);
n.extra = 0xFFFFu; // (the digit layout is not recorded)
n.off = append_text(m_doc, text.data(), text.size());
n.len = static_cast<std::uint32_t>(text.size());
return n;
}
/// a string (or key) in the edit arena
node string_node(const char* s, std::size_t len)
{
if (NLOHMANN_VIEW_UNLIKELY(len >= 0xFFFFFFFFu))
{
throw_out_of_range(416, "strings of 4 GiB or more are not supported by json_document"); // LCOV_EXCL_LINE
}
node n = plain_node(value_t::string);
n.flags = static_cast<std::uint8_t>(n.flags | node_flags::edited);
n.off = append_text(m_doc, s, len);
n.len = static_cast<std::uint32_t>(len);
return n;
}
/// a scalar of a view (of any document) as a node of this document
node copy_scalar(const document_data& from, const node& n)
{
if (&from == &m_doc)
{
return n; // the same storage
}
switch (static_cast<value_t>(n.kind))
{
case value_t::string:
return string_node(from.str(n), n.len);
case value_t::number_integer:
case value_t::number_unsigned:
return integer_node(integer_bits(n), static_cast<value_t>(n.kind));
case value_t::number_float:
{
node r = plain_node(value_t::number_float);
r.flags = static_cast<std::uint8_t>(r.flags | node_flags::edited);
r.off = append_text(m_doc, from.str(n), n.len);
r.len = n.len;
r.extra = n.extra;
return r;
}
case value_t::boolean:
{
node r = plain_node(value_t::boolean);
r.flags = static_cast<std::uint8_t>(r.flags | (n.flags & node_flags::is_true));
return r;
}
case value_t::null:
case value_t::object:
case value_t::array:
case value_t::binary:
case value_t::discarded:
default:
return plain_node(value_t::null);
}
}
node json_scalar(const BasicJsonType& j)
{
switch (j.type())
{
case value_t::null:
return plain_node(value_t::null);
case value_t::boolean:
{
node r = plain_node(value_t::boolean);
r.flags = static_cast<std::uint8_t>(r.flags | (j.template get<bool>() ? node_flags::is_true : 0));
return r;
}
case value_t::number_integer:
return integer_node(static_cast<std::uint64_t>(static_cast<std::int64_t>(j.template get<number_integer_t>())), value_t::number_integer);
case value_t::number_unsigned:
return integer_node(static_cast<std::uint64_t>(j.template get<number_unsigned_t>()), value_t::number_unsigned);
case value_t::number_float:
return float_node(j.template get<number_float_t>());
case value_t::string:
{
const auto& s = j.template get_ref<const string_t&>();
check_utf8(s.data(), s.size());
return string_node(s.data(), s.size());
}
case value_t::binary:
throw_type_error(319, "cannot store a binary value in a json_document", "");
case value_t::discarded:
case value_t::object:
case value_t::array:
default:
throw_type_error(302, "type must be a value, but is ", "discarded");
}
}
/// number of nodes of a subtree (containers, keys, scalars)
template<bool E>
static std::size_t count_nodes(const document_data& d, const node* n)
{
if (!is_container(*n))
{
return 1;
}
const bool object = n->kind == static_cast<std::uint8_t>(value_t::object);
std::size_t r = 1;
for (const node* c = navigation<E>::first(d, n), *end = navigation<E>::end(d, n); c != end;)
{
const node* const v = object ? c + 1 : c;
r += (object ? 1 : 0) + count_nodes<E>(d, navigation<E>::value(v));
c = document_data::after(v);
}
return r;
}
/// copy a subtree (of any document) as a contiguous sequence; returns its end
template<bool E>
node* fill_nodes(const document_data& d, const node* n, node* out)
{
if (!is_container(*n))
{
*out = copy_scalar(d, *n);
return out + 1;
}
node* const self = out++;
*self = plain_node(static_cast<value_t>(n->kind));
self->len = n->len;
const bool object = n->kind == static_cast<std::uint8_t>(value_t::object);
for (const node* c = navigation<E>::first(d, n), *end = navigation<E>::end(d, n); c != end;)
{
if (object)
{
*out++ = copy_scalar(d, *c);
++c;
}
out = fill_nodes<E>(d, navigation<E>::value(c), out);
c = document_data::after(c);
}
self->next = static_cast<std::uint32_t>(out - self);
return out;
}
static std::size_t count_nodes(const BasicJsonType& j)
{
std::size_t r = 1;
if (j.is_object())
{
for (const auto& member : j.items())
{
r += 1 + count_nodes(member.value());
}
}
else if (j.is_array())
{
for (const auto& e : j)
{
r += count_nodes(e);
}
}
return r;
}
node* fill_nodes(const BasicJsonType& j, node* out)
{
if (!j.is_structured())
{
*out = json_scalar(j);
return out + 1;
}
node* const self = out++;
*self = plain_node(j.type());
self->len = static_cast<std::uint32_t>(j.size());
if (j.is_object())
{
for (const auto& member : j.items())
{
check_utf8(member.key().data(), member.key().size());
*out++ = string_node(member.key().data(), member.key().size());
out = fill_nodes(member.value(), out);
}
}
else
{
for (const auto& e : j)
{
out = fill_nodes(e, out);
}
}
self->next = static_cast<std::uint32_t>(out - self);
return out;
}
document_data& m_doc;
};
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -0,0 +1,242 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <algorithm> // max, min
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t
#include <cstring> // memcpy
#include <functional> // less
#include <memory> // unique_ptr
#include <utility> // move
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/errors.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
// The storage of edits. Edits never move or resize the parsed index: every
// value keeps its node, so views stay valid. New values and element sequences
// live in chunks that never move; strings and number tokens written by edits
// live in the edit arena. An array/object whose elements change gets
// node_flags::moved: its elements then live in a separate sequence (a header
// node, then the entries), whose entries link to the values (kind_link).
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
inline document_data::edit_state& edit_state_of(document_data& d)
{
if (!d.edits)
{
d.edits.reset(new document_data::edit_state()); // NOLINT(cppcoreguidelines-owning-memory): owned by the unique_ptr
}
return *d.edits;
}
/// k consecutive nodes that never move (new values and blocks)
inline node* alloc_nodes(document_data& d, std::size_t k)
{
document_data::edit_state& e = edit_state_of(d);
if (NLOHMANN_VIEW_UNLIKELY(static_cast<std::size_t>(e.chunk_end - e.chunk_cur) < k))
{
const std::size_t count = (std::max)(k, e.chunk_next);
std::unique_ptr<node[]> fresh(new node[count]()); // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
e.chunks.push_back(std::move(fresh));
e.chunk_cur = e.chunks.back().get();
e.chunk_end = e.chunk_cur + count;
e.chunk_next = (std::min)(e.chunk_next * 2, std::size_t{65536});
e.bytes += count * sizeof(node);
}
node* const r = e.chunk_cur;
e.chunk_cur += k;
return r;
}
/// copy n bytes into the edit arena and return their offset; a new buffer
/// leaves the old one alive, so that string views into it remain valid
inline std::uint32_t append_text(document_data& d, const char* s, std::size_t n)
{
document_data::edit_state& e = edit_state_of(d);
if (NLOHMANN_VIEW_UNLIKELY(e.text_cap - e.text_used < n))
{
const std::size_t cap = (std::max)(e.text_cap * 2, e.text_used + n + 256);
if (cap > 0xFFFFFFFFu)
{
throw_out_of_range(416, "edits of 4 GiB or more are not supported by json_document"); // LCOV_EXCL_LINE (4 GiB)
}
std::unique_ptr<char[]> fresh(new char[cap]); // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
if (e.text_used != 0)
{
std::memcpy(fresh.get(), e.texts.back().get(), e.text_used);
}
e.texts.push_back(std::move(fresh));
e.text_cap = cap;
e.bytes += cap;
d.base[2] = e.texts.back().get();
}
const auto off = static_cast<std::uint32_t>(e.text_used);
if (n != 0)
{
std::memcpy(e.texts.back().get() + e.text_used, s, n);
}
e.text_used += n;
return off;
}
/// the capacity in nodes of the block of a moved container (0: a fixed
/// sequence, the elements of a new value)
inline std::size_t moved_capacity(const document_data& d, const node* n) noexcept
{
return d.edits->moved_cap[n->off];
}
/// let container n take its elements from `seq` (header node first)
inline void set_moved(document_data& d, node* n, node* seq, std::size_t cap)
{
document_data::edit_state& e = edit_state_of(d);
if ((n->flags & node_flags::moved) != 0)
{
e.moved[n->off] = seq;
e.moved_cap[n->off] = cap;
return;
}
if (e.moved.size() >= 0xFFFFFFFFu)
{
throw_out_of_range(416, "more than 4294967295 edited arrays and objects are not supported by json_document"); // LCOV_EXCL_LINE
}
if (e.moved.size() == e.moved.capacity() || e.moved_cap.size() == e.moved_cap.capacity())
{
// both grow before either changes, so that the push_backs cannot throw
e.moved.reserve((2 * e.moved.size()) + 16);
e.moved_cap.reserve((2 * e.moved.size()) + 16);
}
e.moved.push_back(seq);
e.moved_cap.push_back(cap);
n->off = static_cast<std::uint32_t>(e.moved.size() - 1);
n->flags = static_cast<std::uint8_t>(n->flags | node_flags::moved | node_flags::is_new);
}
/// Make the elements of container n a growable block with room for `extra`
/// more nodes, and return its header. The entries link to the existing
/// values, which stay where they are. A block that grows is copied (its old
/// space is not reused).
inline node* block_of(document_data& d, node* n, std::size_t extra)
{
if ((n->flags & node_flags::moved) != 0 && moved_capacity(d, n) != 0)
{
node* const h = d.edits->moved[n->off];
if (h->next + extra <= moved_capacity(d, n))
{
return h;
}
const std::size_t cap = (std::max)(2 * moved_capacity(d, n), h->next + extra);
node* const nh = alloc_nodes(d, cap);
std::memcpy(nh, h, h->next * sizeof(node));
set_moved(d, n, nh, cap);
return nh;
}
const bool object = n->kind == static_cast<std::uint8_t>(value_t::object);
const std::size_t used = 1 + (static_cast<std::size_t>(n->len) * (object ? 2 : 1));
const std::size_t cap = used + extra;
node* const h = alloc_nodes(d, cap);
*h = node{};
h->kind = n->kind;
h->len = n->len;
h->next = static_cast<std::uint32_t>(used);
node* o = h + 1;
for (const node* c = d.first_child_edited(n), *e = d.child_end_edited(n); c != e;)
{
if (object)
{
*o++ = *c++; // the key
}
make_link(*o, document_data::deref(c));
++o;
c = document_data::after(c);
}
set_moved(d, n, h, cap);
return h;
}
/// The container whose elements include `target`; nullptr for the root, for
/// a value that is no longer part of the document, and for a value that is
/// only reached through a link. Values never move between allocations, so
/// the path to `target` stays inside the allocation that holds it (the parsed
/// index, or one new value), where the extent of each container (`next`)
/// still covers its original subtree.
inline node* find_parent(const document_data& d, const node* target)
{
const std::less<const node*> lt;
const node* lo = d.tape;
const node* hi = d.tape + d.tape_size;
const node* c = d.tape;
if (lt(target, lo) || !lt(target, hi))
{
if (!d.edits)
{
return nullptr; // LCOV_EXCL_LINE (nodes outside the index exist only after edits)
}
auto it = d.edits->regions.upper_bound(target);
if (it == d.edits->regions.begin())
{
return nullptr; // LCOV_EXCL_LINE (an array/object with elements is in the index or a new value)
}
--it;
lo = it->first;
hi = lo + lo->next;
if (!lt(target, hi))
{
return nullptr; // LCOV_EXCL_LINE (a single-node value, reached through a link)
}
// the root of a new value is the element sequence of its owner, or a linked value
c = it->second != nullptr ? it->second : lo;
}
if (target == lo)
{
return nullptr;
}
for (;;)
{
if (!is_container(*c))
{
return nullptr; // LCOV_EXCL_LINE (the value is inside c)
}
const bool object = c->kind == static_cast<std::uint8_t>(value_t::object);
const node* down = nullptr;
for (const node* p = d.first_child_edited(c), *e = d.child_end_edited(c); p != e;)
{
const node* const at = object ? p + 1 : p;
const node* const v = document_data::deref(at);
if (v == target)
{
return const_cast<node*>(c); // NOLINT(cppcoreguidelines-pro-type-const-cast): the nodes belong to the document
}
if (is_container(*v) && !lt(v, lo) && lt(v, target) && lt(target, v + v->next))
{
down = v;
break;
}
p = document_data::after(at);
}
if (down == nullptr)
{
return nullptr;
}
c = down;
}
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+5
View File
@@ -30,6 +30,11 @@ namespace view
NLOHMANN_VIEW_THROW(type_error::create(id, concat(prefix, type), nullptr));
}
[[noreturn]] NLOHMANN_VIEW_NOINLINE inline void throw_type_error(int id, const std::string& msg)
{
NLOHMANN_VIEW_THROW(type_error::create(id, msg, nullptr));
}
[[noreturn]] NLOHMANN_VIEW_NOINLINE inline void throw_out_of_range(int id, const std::string& msg)
{
NLOHMANN_VIEW_THROW(out_of_range::create(id, msg, nullptr));
+602
View File
@@ -0,0 +1,602 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint8_t, uint16_t, uint32_t, uint64_t
#include <cstring> // memcmp, memcpy
#include <limits> // numeric_limits
#include <string> // string
#include <vector> // vector
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/errors.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
#include <nlohmann/detail/view/number.hpp>
#include <nlohmann/detail/view/object_index.hpp>
#include <nlohmann/detail/view/scan.hpp>
// Images: a document stored so that loading it needs no parsing.
//
// Layout (little-endian): a 64-byte header, the nodes, the text (the source,
// followed by the number tokens written by edits), a NUL, the decoded strings
// (followed by the strings written by edits), a NUL. The idea is that of
// zero-copy formats such as FlatBuffers (https://github.com/google/flatbuffers)
// and YaFF (https://github.com/yandex/yaff); no code is taken from them.
// check_image follows the idea of FlatBuffers' Verifier (bounds and
// structure) and also checks what the parser guarantees about strings and
// numbers, so that reading and serializing a checked image is safe and yields
// valid JSON.
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
/// how load() checks an image
enum class image_check
{
/// everything the parser guarantees: structure and bounds, strings (valid
/// UTF-8; source strings without quotes, backslashes, and control
/// characters), and numbers (well-formed, matching the stored values)
full,
/// structure and bounds only: reading and serializing are safe, but a
/// crafted image can yield invalid UTF-8, strings that serialize to
/// invalid JSON, or numbers that differ from their text
bounds,
/// none: for images from a trusted source only (a damaged image is
/// undefined behavior)
none,
};
struct image_header
{
std::array<char, 4> magic; ///< "NJVI"
std::uint32_t version; ///< 1
std::uint64_t node_count;
std::uint64_t text_size;
std::uint64_t arena_size;
std::array<std::uint64_t, 4> reserved; ///< zero (for later versions)
};
static_assert(sizeof(image_header) == 64, "the image header must be 64 bytes");
constexpr std::uint32_t image_version = 1;
/// the largest node count and text or string size of an image (as for parsed
/// documents, offsets and counts must fit 32 bits)
constexpr std::uint64_t image_limit = 0xFFFFFFF0u;
/// Copy the current structure of an edited document into nodes in document
/// order, as the parser would have written them. Text written by edits is
/// appended to text_tail (number tokens) and arena_tail (strings); floats that
/// are not finite become null, as dump() writes them.
inline void compact_nodes(const document_data& d, std::size_t arena_size, std::vector<node>& out, std::string& text_tail, std::string& arena_tail)
{
struct frame
{
const node* cur;
const node* end;
std::size_t index; ///< the container's node in out
std::uint32_t count;
bool object;
};
std::vector<frame> stack;
const auto string_node = [&](const node & s)
{
node r = s;
r.extra = 0;
r.flags = static_cast<std::uint8_t>(s.flags & node_flags::storage);
if (r.flags == node_flags::edited)
{
r.off = static_cast<std::uint32_t>(arena_size + arena_tail.size());
arena_tail.append(d.str(s), s.len);
r.flags = node_flags::escaped;
}
return r;
};
const auto emit = [&](const node * v)
{
node r = *v;
switch (static_cast<value_t>(v->kind))
{
case value_t::object:
case value_t::array:
r.flags = 0;
r.extra = 0;
r.off = (v->flags & (node_flags::moved | node_flags::is_new)) != 0 ? 0 : v->off;
r.len = 0; // counted below
r.next = 0; // set when the container is complete
stack.push_back(frame{d.first_child_edited(v), d.child_end_edited(v), out.size(), 0, v->kind == static_cast<std::uint8_t>(value_t::object)});
break;
case value_t::string:
r = string_node(*v);
break;
case value_t::number_integer:
case value_t::number_unsigned:
if ((v->flags & node_flags::storage) == node_flags::edited)
{
r.off = static_cast<std::uint32_t>(d.size + text_tail.size());
text_tail.append(d.str(*v), number_length(*v));
}
r.flags = 0;
break;
case value_t::number_float:
if ((v->flags & node_flags::storage) == node_flags::edited)
{
const char* const t = d.str(*v);
if (t[0] == 'n' || t[0] == 'i' || (v->len > 1 && t[1] == 'i'))
{
r = node{}; // nan and infinity: null, as dump() writes them
r.kind = static_cast<std::uint8_t>(value_t::null);
break;
}
r.off = static_cast<std::uint32_t>(d.size + text_tail.size());
text_tail.append(t, v->len);
r.extra = 0xFFFFu; // the digit layout is not recorded
}
r.flags = 0;
break;
case value_t::boolean:
r.flags = static_cast<std::uint8_t>(v->flags & node_flags::is_true);
break;
case value_t::null:
case value_t::binary:
case value_t::discarded:
default:
r.flags = 0;
break;
}
out.push_back(r);
};
emit(d.tape);
while (!stack.empty())
{
frame& top = stack.back();
if (top.cur == top.end)
{
node& c = out[top.index];
c.len = top.count;
c.next = static_cast<std::uint32_t>(out.size() - top.index);
stack.pop_back();
continue;
}
++top.count;
const node* v = nullptr;
if (top.object)
{
out.push_back(string_node(*top.cur));
v = document_data::deref(top.cur + 1);
top.cur = document_data::after(top.cur + 1);
}
else
{
v = document_data::deref(top.cur);
top.cur = document_data::after(top.cur);
}
emit(v); // may grow the stack (top is not used afterwards)
}
}
/// the document as an image
inline std::vector<std::uint8_t> save_image(const document_data& d)
{
#if !NLOHMANN_VIEW_LITTLE_ENDIAN
throw_type_error(320, "json_document images need a little-endian target"); // LCOV_EXCL_LINE
#endif
const std::size_t arena_size = d.arena_size;
const node* nodes = d.tape;
std::size_t count = d.tape_size;
std::vector<node> compacted;
std::string text_tail;
std::string arena_tail;
if (d.edits)
{
compact_nodes(d, arena_size, compacted, text_tail, arena_tail);
nodes = compacted.data();
count = compacted.size();
}
const std::size_t text_size = d.size + text_tail.size();
const std::size_t total_arena = arena_size + arena_tail.size();
if (NLOHMANN_VIEW_UNLIKELY(text_size >= image_limit || total_arena >= image_limit || count >= image_limit))
{
// LCOV_EXCL_START (4 GiB)
throw_out_of_range(416, "images of 4 GiB or more are not supported by json_document");
// LCOV_EXCL_STOP
}
image_header h{};
h.magic = {{'N', 'J', 'V', 'I'}};
h.version = image_version;
h.node_count = count;
h.text_size = text_size;
h.arena_size = total_arena;
std::vector<std::uint8_t> image(sizeof(h) + (count * sizeof(node)) + text_size + 1 + total_arena + 1);
std::uint8_t* o = image.data();
std::memcpy(o, &h, sizeof(h));
o += sizeof(h);
std::memcpy(o, nodes, count * sizeof(node));
// the hash indexes are rebuilt by load()
for (std::size_t i = 0; i < count; ++i)
{
if (nodes[i].kind == static_cast<std::uint8_t>(value_t::object) && nodes[i].extra != 0)
{
node n = nodes[i];
n.extra = 0;
std::memcpy(o + (i * sizeof(node)), &n, sizeof(node));
}
}
o += count * sizeof(node);
const auto append = [&o](const char* s, std::size_t n)
{
if (n != 0)
{
std::memcpy(o, s, n);
o += n;
}
};
append(d.src, d.size);
append(text_tail.data(), text_tail.size());
*o++ = 0;
append(d.base[1], arena_size);
append(arena_tail.data(), arena_tail.size());
*o = 0;
return image;
}
/// whether a number node matches its token the way the parser records it
/// (after the bounds check)
inline bool check_number(const node& n, const unsigned char* text)
{
const std::size_t len = number_length(n);
const unsigned char* const s = text + n.off;
const unsigned char* const e = s + len;
const unsigned char* p = s;
const bool negative = *p == '-';
p += negative ? 1 : 0;
const unsigned char* const int_start = p;
if (p == e)
{
return false;
}
if (*p == '0')
{
++p;
}
else if (*p >= '1' && *p <= '9')
{
while (p != e && is_digit(*p))
{
++p;
}
}
else
{
return false;
}
const auto int_digits = static_cast<std::size_t>(p - int_start);
std::size_t frac_digits = 0;
bool is_float = false;
if (p != e && *p == '.')
{
const unsigned char* const f0 = ++p;
while (p != e && is_digit(*p))
{
++p;
}
if (p == f0)
{
return false;
}
frac_digits = static_cast<std::size_t>(p - f0);
is_float = true;
}
std::int64_t exponent = 0;
if (p != e && (*p | 0x20u) == 'e')
{
++p;
const bool exp_negative = p != e && *p == '-';
p += (p != e && (*p == '+' || *p == '-')) ? 1 : 0;
if (p == e || !is_digit(*p))
{
return false;
}
while (p != e && is_digit(*p))
{
exponent = exponent < 100000 ? (exponent * 10) + (*p - '0') : exponent;
++p;
}
exponent = exp_negative ? -exponent : exponent;
is_float = true;
}
if (p != e)
{
return false;
}
if (n.kind == static_cast<std::uint8_t>(value_t::number_float))
{
// the digit layout the parser records (or "many", as compaction
// writes it), and a finite value
const auto layout = static_cast<std::uint16_t>((int_digits < 255 ? int_digits : 255) | ((frac_digits < 255 ? frac_digits : 255) << 8u));
if (n.extra != layout && n.extra != 0xFFFFu)
{
return false;
}
// parse() rejects floats that overflow; as there, only a number whose
// magnitude could reach 1e308 needs the conversion
if (static_cast<std::int64_t>(int_digits) + exponent > 300)
{
const auto v = float_value<double>(reinterpret_cast<const char*>(s), n); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
return v <= (std::numeric_limits<double>::max)() && v >= -(std::numeric_limits<double>::max)();
}
return true;
}
// integers: the token's value is the stored one; number_integer nodes of
// edits can be non-negative (as basic_json keeps the type of a value)
const bool integer = n.kind == static_cast<std::uint8_t>(value_t::number_integer);
if (is_float || int_digits > 20 || (negative && !integer))
{
return false;
}
// (at most 19 digits cannot overflow; 20 digits are compared with 2^64 - 1)
if (int_digits == 20 && std::memcmp(int_start, "18446744073709551615", 20) > 0)
{
return false;
}
std::uint64_t m = 0;
for (const unsigned char* d = int_start; d != int_start + int_digits; ++d)
{
m = (m * 10) + static_cast<std::uint64_t>(*d - '0');
}
if (integer && m > (negative ? std::uint64_t{1} << 63u : (std::uint64_t{1} << 63u) - 1))
{
return false;
}
return integer_bits(n) == (negative ? 0 - m : m);
}
/// Check the nodes of a loaded image against its text and decoded strings:
/// kinds, flags, and `extra`; extents and element counts of arrays and
/// objects; keys; bounds; string contents (source strings as the parser
/// leaves them: no quotes, backslashes, or control characters; all strings
/// valid UTF-8); and number tokens.
inline bool check_image(const node* nodes, std::size_t count, const unsigned char* text, std::size_t text_size,
const unsigned char* arena, std::size_t arena_size, bool full)
{
struct frame
{
std::size_t end;
std::uint32_t len;
std::uint32_t seen;
bool object;
bool expect_key;
};
std::vector<frame> stack;
const auto check_string = [&](const node & n) -> bool
{
if ((n.flags & ~node_flags::escaped) != 0 || n.extra != 0)
{
return false;
}
const bool decoded = (n.flags & node_flags::escaped) != 0;
const unsigned char* const base = decoded ? arena : text;
const std::size_t limit = decoded ? arena_size : text_size;
if (n.off > limit || n.len > limit - n.off)
{
return false;
}
if (!full)
{
return true;
}
const unsigned char* const b = base + n.off;
return decoded ? valid_utf8_prefix(b, n.len) == n.len : scan_string_run(b, b + n.len) == b + n.len;
};
// bounds of a number token; the recorded digit layout must lie within it
const auto number_in_bounds = [&](const node & n) -> bool
{
const std::size_t len = number_length(n);
if (len == 0 || n.off > text_size || len > text_size - n.off)
{
return false;
}
if (n.kind != static_cast<std::uint8_t>(value_t::number_float))
{
return (n.extra >> 8u) == 0;
}
// float_value() reads the sign, the integer digits, and the point and
// fraction digits the layout records (a layout of more than 19 digits
// means the general conversion, which stays within the token)
const std::size_t int_digits = n.extra & 0xFFu;
const std::size_t frac_digits = n.extra >> 8u;
const std::size_t need = (text[n.off] == '-' ? 1u : 0u) + int_digits + (frac_digits != 0 ? frac_digits + 1 : 0);
return int_digits + frac_digits > 19 || need <= len;
};
std::size_t i = 0;
for (;;)
{
// close finished arrays and objects
while (!stack.empty() && i == stack.back().end)
{
const frame f = stack.back();
if (f.seen != f.len || (f.object && !f.expect_key))
{
return false;
}
stack.pop_back();
if (!stack.empty())
{
++stack.back().seen;
stack.back().expect_key = true;
}
}
if (i == count)
{
return stack.empty();
}
if (i != 0 && stack.empty())
{
return false; // nodes after the root
}
const node& n = nodes[i];
if (!stack.empty() && stack.back().object && stack.back().expect_key)
{
if (n.kind != static_cast<std::uint8_t>(value_t::string) || !check_string(n))
{
return false;
}
stack.back().expect_key = false;
++i;
continue;
}
bool complete = true;
switch (static_cast<value_t>(n.kind))
{
case value_t::null:
// (the offset of a literal is read to size the output of dump())
if (n.flags != 0 || n.extra != 0 || n.off > text_size)
{
return false;
}
break;
case value_t::boolean:
if ((n.flags & ~node_flags::is_true) != 0 || n.extra != 0 || n.off > text_size)
{
return false;
}
break;
case value_t::string:
if (!check_string(n))
{
return false;
}
break;
case value_t::number_integer:
case value_t::number_unsigned:
case value_t::number_float:
if (n.flags != 0 || !number_in_bounds(n) || (full && !check_number(n, text)))
{
return false;
}
break;
case value_t::array:
case value_t::object:
{
const std::size_t limit = stack.empty() ? count : stack.back().end;
if (n.flags != 0 || n.extra != 0 || n.next == 0 || n.next > limit - i || n.off > text_size)
{
return false;
}
stack.push_back(frame{i + n.next, n.len, 0, n.kind == static_cast<std::uint8_t>(value_t::object), true});
complete = false;
break;
}
case value_t::binary:
case value_t::discarded:
default:
return false;
}
++i;
if (complete && !stack.empty())
{
++stack.back().seen;
stack.back().expect_key = true;
}
}
}
[[noreturn]] NLOHMANN_VIEW_NOINLINE inline void throw_invalid_image(const char* what)
{
throw_parse_error(116, concat("invalid json_document image: ", what));
}
/// Read an image into d. The text and the decoded strings stay in the image;
/// the nodes are copied (so that they are aligned, and edits can change them).
inline void load_image(document_data& d, const std::uint8_t* image, std::size_t size, image_check check)
{
#if !NLOHMANN_VIEW_LITTLE_ENDIAN
throw_type_error(320, "json_document images need a little-endian target"); // LCOV_EXCL_LINE
#endif
if (image == nullptr || size < sizeof(image_header))
{
throw_invalid_image("too short");
}
image_header h{};
std::memcpy(&h, image, sizeof(h));
// (the reserved fields are for later versions)
if (std::memcmp(h.magic.data(), "NJVI", 4) != 0 || h.version != image_version
|| (h.reserved[0] | h.reserved[1] | h.reserved[2] | h.reserved[3]) != 0)
{
throw_invalid_image("unknown format");
}
const std::size_t room = size - sizeof(h);
if (h.node_count == 0 || h.node_count > room / sizeof(node) || h.node_count >= image_limit || h.text_size >= image_limit || h.arena_size >= image_limit)
{
throw_invalid_image("sizes out of range");
}
const auto count = static_cast<std::size_t>(h.node_count);
const auto text_size = static_cast<std::size_t>(h.text_size);
const auto arena_size = static_cast<std::size_t>(h.arena_size);
const std::size_t text_at = sizeof(h) + (count * sizeof(node));
// the text, a NUL, the decoded strings, a NUL, and nothing after them
if (size - text_at < 2 || text_size > size - text_at - 2 || arena_size != size - text_at - text_size - 2
|| image[text_at + text_size] != 0 || image[size - 1] != 0)
{
throw_invalid_image("sizes out of range");
}
d.discarded = true;
d.edits.reset();
d.base[2] = nullptr;
d.owned.clear();
if (d.owned_image.empty() || image != d.owned_image.data())
{
d.owned_image.clear();
}
d.arena.clear();
d.indexes.clear();
d.index_slots.clear();
d.large_objects.clear();
d.tape_size = 0;
d.reserve(count);
std::memcpy(d.tape, image + sizeof(h), count * sizeof(node));
d.tape_size = count;
const std::uint8_t* const text = image + text_at;
const std::uint8_t* const arena = text + text_size + 1;
d.src = reinterpret_cast<const char*>(text); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
d.size = text_size;
d.base[0] = d.src;
d.base[1] = reinterpret_cast<const char*>(arena); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
d.arena_size = arena_size;
if (check != image_check::none && !check_image(d.tape, count, text, text_size, arena, arena_size, check == image_check::full))
{
throw_invalid_image("the check failed");
}
// the hash indexes of large objects, as after parsing
for (std::size_t i = 0; i < count; ++i)
{
node& n = d.tape[i];
if (n.kind == static_cast<std::uint8_t>(value_t::object))
{
n.extra = 0;
if (n.len >= document_data::index_min_members)
{
d.large_objects.push_back(static_cast<std::uint32_t>(i));
}
}
}
build_object_indexes(d);
d.discarded = false;
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+1 -1
View File
@@ -73,7 +73,7 @@ class view_iterator
NLOHMANN_VIEW_ALWAYS_INLINE View operator*() const noexcept
{
return View(m_doc, m_pos + m_value_offset);
return View(m_doc, View::navigation::value(m_pos + m_value_offset));
}
pointer operator->() const noexcept
+26 -12
View File
@@ -16,6 +16,7 @@
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
#include <nlohmann/detail/view/object_index.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -83,14 +84,20 @@ class short_key
/// the key node of the first member of an object with the given key, or
/// nullptr; most keys are rejected by their length, from the index alone
inline const node* find_member(const document_data& d, const node* object, const char* key, std::size_t n) noexcept
template<bool Editable>
const node* find_member(const document_data& d, const node* object, const char* key, std::size_t n) noexcept
{
const node* const end = document_data::child_end(object);
using nav = navigation<Editable>;
if (NLOHMANN_VIEW_UNLIKELY(object->extra != 0) && (!Editable || (object->flags & node_flags::moved) == 0))
{
return find_indexed(d, object, key, n); // a large object (whose members have not been edited)
}
const node* const end = nav::end(d, object);
const auto* const k = reinterpret_cast<const unsigned char*>(key); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
if (NLOHMANN_VIEW_LIKELY(n <= 16))
{
const short_key probe(k, n);
for (const node* m = document_data::first_child(object); m != end; m = document_data::after(m + 1))
for (const node* m = nav::first(d, object); m != end; m = document_data::after(m + 1))
{
if (m->len == n && probe.matches(reinterpret_cast<const unsigned char*>(d.str(*m)))) // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
{
@@ -99,7 +106,7 @@ inline const node* find_member(const document_data& d, const node* object, const
}
return nullptr;
}
for (const node* m = document_data::first_child(object); m != end; m = document_data::after(m + 1))
for (const node* m = nav::first(d, object); m != end; m = document_data::after(m + 1))
{
if (m->len == n && std::memcmp(d.str(*m), key, n) == 0)
{
@@ -109,10 +116,16 @@ inline const node* find_member(const document_data& d, const node* object, const
return nullptr;
}
/// the element of an array at an index below its size
inline const node* element_at(const node* array, std::size_t idx) noexcept
/// the entry of the element of an array at an index below its size (a link
/// in the moved sequences of editable documents)
template<bool Editable>
const node* element_at(const document_data& d, const node* array, std::size_t idx) noexcept
{
const node* e = document_data::first_child(array);
const node* e = navigation<Editable>::first(d, array);
if (Editable && (array->flags & node_flags::moved) != 0 && d.edits->moved_cap[array->off] != 0)
{
return e + idx; // a growable block: one link per element
}
for (std::size_t i = 0; i < idx; ++i)
{
e = document_data::after(e);
@@ -120,13 +133,14 @@ inline const node* element_at(const node* array, std::size_t idx) noexcept
return e;
}
/// the last element of a non-empty array, or the key of the last member of a
/// non-empty object
inline const node* last_child(const node* container) noexcept
/// the entry of the last element of a non-empty array, or the key of the
/// last member of a non-empty object
template<bool Editable>
const node* last_child(const document_data& d, const node* container) noexcept
{
const std::size_t value_offset = container->kind == static_cast<std::uint8_t>(value_t::object) ? 1 : 0;
const node* const end = document_data::child_end(container);
const node* last = document_data::first_child(container);
const node* const end = navigation<Editable>::end(d, container);
const node* last = navigation<Editable>::first(d, container);
for (const node* c = document_data::after(last + value_offset); c != end; c = document_data::after(c + value_offset))
{
last = c;
@@ -19,3 +19,8 @@
#undef NLOHMANN_VIEW_THROW
#undef NLOHMANN_VIEW_LITTLE_ENDIAN
#undef NLOHMANN_VIEW_REPEAT16
#undef NLOHMANN_VIEW_NEON
#undef NLOHMANN_VIEW_SSE2
#undef NLOHMANN_VIEW_SSSE3
#undef NLOHMANN_VIEW_VECTOR
#undef NLOHMANN_VIEW_VECTOR_UTF8
+34 -28
View File
@@ -34,17 +34,24 @@ iterative, so the nesting depth is limited by memory only, as for parse().
Without a lexer the handler records no source positions
(JSON_DIAGNOSTIC_POSITIONS).
*/
template<typename BasicJsonType>
template<typename BasicJsonType, bool Editable>
BasicJsonType materialize(const document_data& d, const node* n)
{
using string_t = typename BasicJsonType::string_t;
using sax_t = json_sax_dom_parser<BasicJsonType, iterator_input_adapter<const char*>>;
using nav = navigation<Editable>;
struct frame
{
const node* pos; ///< next element, or key of the next member
const node* end;
bool object;
};
BasicJsonType result;
sax_t sax(result, true);
const string_t no_token{};
// the ends of the open containers, and whether they are objects
std::vector<std::pair<const node*, bool>> open;
std::vector<frame> open;
for (;;)
{
switch (static_cast<value_t>(n->kind))
@@ -61,67 +68,66 @@ BasicJsonType materialize(const document_data& d, const node* n)
{
sax.start_array(n->len);
}
open.emplace_back(document_data::child_end(n), object);
n = document_data::first_child(n);
open.push_back(frame{nav::first(d, n), nav::end(d, n), object});
break;
}
case value_t::string:
{
string_t s(d.str(*n), n->len);
sax.string(s);
++n;
break;
}
case value_t::number_integer:
sax.number_integer(static_cast<typename BasicJsonType::number_integer_t>(static_cast<std::int64_t>(integer_bits(*n))));
++n;
break;
case value_t::number_unsigned:
sax.number_unsigned(static_cast<typename BasicJsonType::number_unsigned_t>(integer_bits(*n)));
++n;
break;
case value_t::number_float:
sax.number_float(float_value<typename BasicJsonType::number_float_t>(d.str(*n), *n), no_token);
++n;
sax.number_float(float_value<typename BasicJsonType::number_float_t>(d, *n), no_token);
break;
case value_t::boolean:
sax.boolean((n->flags & node_flags::is_true) != 0);
++n;
break;
case value_t::null:
case value_t::binary:
case value_t::discarded:
default:
sax.null();
++n;
break;
}
// the next value: close finished containers, then read the key
for (;;)
{
if (open.empty())
{
return result;
}
if (n != open.back().first)
frame& f = open.back();
if (f.pos == f.end)
{
break;
if (f.object)
{
sax.end_object();
}
else
{
sax.end_array();
}
open.pop_back();
continue;
}
if (open.back().second)
const node* entry = f.pos;
if (f.object)
{
sax.end_object();
string_t key(d.str(*entry), entry->len);
sax.key(key);
++entry;
}
else
{
sax.end_array();
}
open.pop_back();
}
if (open.back().second)
{
// the key of the next member
string_t key(d.str(*n), n->len);
sax.key(key);
++n;
n = nav::value(entry);
f.pos = document_data::after(entry);
break;
}
}
}
+25 -3
View File
@@ -33,20 +33,27 @@ static_assert(static_cast<std::uint8_t>(value_t::null) == 0 && static_cast<std::
struct node_flags
{
static constexpr std::uint8_t escaped = 1; ///< string payload lives in the decode arena, not the source
static constexpr std::uint8_t edited = 2; ///< string or number token lives in the edit arena (editable documents)
static constexpr std::uint8_t storage = 3; ///< mask: where a string or number token lives (index into document_data::base)
static constexpr std::uint8_t is_true = 4; ///< boolean value
static constexpr std::uint8_t moved = 8; ///< array/object: the elements live in a separate sequence (editable documents)
static constexpr std::uint8_t is_new = 16; ///< written by an edit: no source position
};
/// kind of an entry of an edited sequence that stands for a value stored
/// elsewhere (the address of the value's node is kept in the len/next bytes)
constexpr std::uint8_t kind_link = 10;
/// One entry of the flat index, in document order. An object's members are
/// stored as key node followed by the value's subtree. Integers keep their
/// converted 64-bit value in the len/next bytes (the node after a scalar is
/// always the next one, and the token length follows from `extra`).
struct node
{
std::uint8_t kind; ///< value_t
std::uint8_t kind; ///< value_t, or kind_link
std::uint8_t flags; ///< node_flags
std::uint16_t extra; ///< numbers: integer digits (low byte) and fraction digits (high byte), 255 = "many"; otherwise 0
std::uint32_t off; ///< source offset (string content, number token, literal, bracket); arena offset if node_flags::escaped
std::uint16_t extra; ///< numbers: integer digits (low byte) and fraction digits (high byte), 255 = "many"; objects: number of the hash index; otherwise 0
std::uint32_t off; ///< source offset (string content, number token, literal, bracket); arena offset if escaped/edited; number of the element sequence if moved
std::uint32_t len; ///< string: decoded bytes; float: token bytes; array/object: element count
std::uint32_t next; ///< array/object: number of nodes of the subtree (its extent in the enclosing sequence)
};
@@ -57,6 +64,21 @@ NLOHMANN_VIEW_ALWAYS_INLINE bool is_container(const node& n) noexcept
return static_cast<unsigned>(n.kind) - 1u <= 1u;
}
/// the value a link node stands for
NLOHMANN_VIEW_ALWAYS_INLINE const node* link_target(const node& n) noexcept
{
const node* t = nullptr;
std::memcpy(static_cast<void*>(&t), reinterpret_cast<const unsigned char*>(&n) + 8, sizeof(const node*)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
return t;
}
inline void make_link(node& n, const node* target) noexcept
{
n = node{};
n.kind = kind_link;
std::memcpy(reinterpret_cast<unsigned char*>(&n) + 8, static_cast<const void*>(&target), sizeof(const node*)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
}
/// the converted value of an integer node (stored in len/next)
NLOHMANN_VIEW_ALWAYS_INLINE std::uint64_t integer_bits(const node& n) noexcept
{
+182 -27
View File
@@ -8,12 +8,20 @@
#pragma once
#include <array> // array
#include <cfloat> // FLT_EVAL_METHOD
#include <cstddef> // size_t
#include <cstdint> // int64_t, uint64_t
#include <cstring> // memcpy
#include <limits> // numeric_limits
#include <string> // string
#include <type_traits> // integral_constant, is_same
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
#include <nlohmann/detail/view/scan.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -21,43 +29,76 @@ namespace detail
namespace view
{
/*!
@brief locate the decimal point and the end of the mantissa of a float token
Also checks that the token is a JSON number. Tokens of the parser and of edits
always are; an image loaded with image_check::bounds can hold any bytes, which
must not reach the conversion (it expects a well-formed token).
*/
inline bool float_token_layout(const char* first, const char* last, std::size_t& dot, std::size_t& mantissa_end) noexcept
{
const auto digit = [last](const char* q)
{
return q != last && is_digit(static_cast<unsigned char>(*q));
};
const char* p = first;
p += (p != last && *p == '-') ? 1 : 0;
if (!digit(p) || (*p == '0' && digit(p + 1)))
{
return false;
}
while (digit(p))
{
++p;
}
dot = std::string::npos;
if (p != last && *p == '.')
{
dot = static_cast<std::size_t>(p - first);
if (!digit(++p))
{
return false;
}
while (digit(p))
{
++p;
}
}
mantissa_end = static_cast<std::size_t>(p - first);
if (p != last && (*p == 'e' || *p == 'E'))
{
++p;
p += (p != last && (*p == '+' || *p == '-')) ? 1 : 0;
if (!digit(p))
{
return false;
}
while (digit(p))
{
++p;
}
}
return p == last;
}
/*!
@brief the value of the float token of a node, as parse() converts it
Uses the lexer's conversion (detail::convert_float_fast, then the locale-aware
strtod fallback), so that the values are bit-identical to parse(). The digit
layout recorded while parsing locates the decimal point and the exponent
without scanning the token.
strtod fallback), so that the values are bit-identical to parse(). A token
that is not a JSON number (only in a damaged image loaded with
image_check::bounds) yields 0.
*/
template<typename FloatType>
NLOHMANN_VIEW_NOINLINE FloatType float_value(const char* first, const node& n)
{
const char* const last = first + n.len;
const std::size_t neg = first[0] == '-' ? 1 : 0;
const std::size_t int_digits = n.extra & 0xFFu;
const std::size_t frac_digits = n.extra >> 8u;
std::size_t dot = std::string::npos;
std::size_t mantissa_end = n.len;
if (int_digits != 255 && frac_digits != 255)
std::size_t dot = 0;
std::size_t mantissa_end = 0;
if (NLOHMANN_VIEW_UNLIKELY(!float_token_layout(first, last, dot, mantissa_end)))
{
dot = frac_digits != 0 ? neg + int_digits : std::string::npos;
mantissa_end = neg + int_digits + (frac_digits != 0 ? 1 + frac_digits : 0);
}
else
{
// more digits than the layout records: locate them
for (std::size_t i = 0; i < n.len; ++i)
{
if (first[i] == '.')
{
dot = i;
}
else if (first[i] == 'e' || first[i] == 'E')
{
mantissa_end = i;
break;
}
}
return FloatType{};
}
FloatType v{};
if (!convert_float_fast(first, last, dot, mantissa_end, v))
@@ -68,6 +109,120 @@ NLOHMANN_VIEW_NOINLINE FloatType float_value(const char* first, const node& n)
return v;
}
/*!
@brief the double of a float token with at most 19 digits, from its layout
The digit layout recorded while parsing says where the integer digits, the
fraction digits, and the exponent are, so the digits are read eight at a
time without scanning. The result is correctly rounded (Clinger's fast path
where both operands are exact, else the Eisel-Lemire algorithm, which needs
no fallback for up to 19 digits), so it is the value parse() produces.
@param[in] p first character of the token
@param[in] e end of the token
@param[in] limit end of the readable memory (the source text)
*/
NLOHMANN_VIEW_ALWAYS_INLINE double layout_double(const unsigned char* p, const unsigned char* e, unsigned int_digits, unsigned frac_digits, const unsigned char* limit) noexcept
{
const bool negative = *p == '-';
p += negative ? 1 : 0;
std::uint64_t w = parse_upto19(p, int_digits, limit);
p += int_digits;
std::int64_t q = 0;
if (frac_digits != 0)
{
w = (w * int_pow10(frac_digits)) + parse_upto19(p + 1, frac_digits, limit);
p += 1 + frac_digits;
q = -static_cast<std::int64_t>(frac_digits);
}
if (p != e)
{
// [eE][+-]digits; huge exponents saturate (the parser rejected
// overflow). The token is not read beyond e, and the digits are taken
// as unsigned, so that a token that is not well-formed (a damaged
// image loaded with image_check::bounds) yields a wrong value, but no
// overflow.
++p;
const bool exp_negative = p != e && *p == '-';
p += (p != e && (*p == '-' || *p == '+')) ? 1 : 0;
std::int64_t exp_value = 0;
for (; p != e; ++p)
{
if (exp_value < 0x10000000)
{
exp_value = (exp_value * 10) + static_cast<unsigned char>(*p - '0');
}
}
q += exp_negative ? -exp_value : exp_value;
}
double result = 0;
if (w != 0)
{
#if !defined(FLT_EVAL_METHOD) || FLT_EVAL_METHOD == 0
static const std::array<double, 23> pow10 = {{1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11, 1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22}};
if (q >= -22 && q <= 22 && w <= (std::uint64_t{1} << 53))
{
// Clinger's fast path: both operands exact, one rounding
result = static_cast<double>(w);
result = q < 0 ? result / pow10[static_cast<std::size_t>(-q)] : result * pow10[static_cast<std::size_t>(q)];
return negative ? -result : result;
}
#endif
const std::uint64_t bits = eisel_lemire(q, w);
std::memcpy(&result, &bits, sizeof(result));
}
return negative ? -result : result;
}
/// the value of a float set by an edit: its token (the shortest round-trip
/// text, or "nan", "inf", "-inf") in the edit arena
template<typename FloatType>
NLOHMANN_VIEW_NOINLINE FloatType edited_float(const char* token, const node& n)
{
if (token[0] == 'n')
{
return std::numeric_limits<FloatType>::quiet_NaN();
}
if (token[0] == 'i' || (token[0] == '-' && token[1] == 'i'))
{
return token[0] == 'i' ? std::numeric_limits<FloatType>::infinity() : -std::numeric_limits<FloatType>::infinity();
}
return float_value<FloatType>(token, n);
}
/// the value of the float token of a node, as parse() converts it; doubles
/// with at most 19 digits are converted from the digit layout
template<typename FloatType>
FloatType float_value(const document_data& d, const node& n)
{
if (NLOHMANN_VIEW_UNLIKELY((n.flags & node_flags::storage) == node_flags::edited))
{
return edited_float<FloatType>(d.str(n), n);
}
return float_value<FloatType>(d, n, std::is_same<FloatType, double> {});
}
template<typename FloatType>
FloatType float_value(const document_data& d, const node& n, std::true_type /*double*/)
{
const unsigned int_digits = n.extra & 0xFFu;
const unsigned frac_digits = n.extra >> 8u;
if (NLOHMANN_VIEW_LIKELY(int_digits + frac_digits <= 19)) // (255 marks "many")
{
// (a float token not written by an edit is in the text)
const auto* const first = reinterpret_cast<const unsigned char*>(d.src + n.off); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
return layout_double(first, first + n.len, int_digits, frac_digits, reinterpret_cast<const unsigned char*>(d.src + d.size)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
}
return float_value<FloatType>(d.str(n), n);
}
template<typename FloatType>
FloatType float_value(const document_data& d, const node& n, std::false_type /*other*/)
{
return float_value<FloatType>(d.str(n), n);
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -0,0 +1,129 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef> // size_t
#include <cstdint> // uint32_t, uint64_t
#include <cstring> // memcmp
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
// Hash indexes of large objects, so that a lookup does not compare thousands
// of keys (as Boost.JSON switches from a linear search to a hash table for
// large objects). An object with document_data::index_min_members members or
// more gets an open-addressing table after parsing; its node stores the
// number of the table (1-based) in `extra`. A slot holds the offset of a key
// node from its object node (0: empty). Of duplicate keys, the first is kept,
// as for the linear search.
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
/// hash of a key: its bytes, eight at a time, in a fixed byte order
inline std::uint64_t key_hash(const char* s, std::size_t n) noexcept
{
std::uint64_t h = 0x9E3779B97F4A7C15u * (n + 1);
const auto* p = reinterpret_cast<const unsigned char*>(s); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
while (n >= 8)
{
h = (h ^ read_eight_bytes(p)) * 0xBF58476D1CE4E5B9u;
h ^= h >> 29u;
p += 8;
n -= 8;
}
std::uint64_t w = 0;
for (std::size_t i = 0; i < n; ++i)
{
w |= static_cast<std::uint64_t>(p[i]) << (8u * i);
}
h = (h ^ w) * 0x94D049BB133111EBu;
return h ^ (h >> 31u);
}
/// build the table of a large object
inline void build_object_index(document_data& d, node* obj)
{
if (d.indexes.size() >= 0xFFFFu)
{
return; // LCOV_EXCL_LINE (the number must fit `extra`; more large objects are searched linearly)
}
std::size_t cap = 16;
while (cap < 2 * static_cast<std::size_t>(obj->len))
{
cap *= 2;
}
const std::size_t start = d.index_slots.size();
d.index_slots.resize(start + cap, 0);
std::uint32_t* const slots = d.index_slots.data() + start;
const std::size_t mask = cap - 1;
for (const node* k = document_data::first_child(obj), *end = document_data::child_end(obj); k != end; k = document_data::after(k + 1))
{
const char* const key = d.str(*k);
std::size_t i = static_cast<std::size_t>(key_hash(key, k->len)) & mask;
bool duplicate = false;
while (slots[i] != 0)
{
const node* const other = obj + slots[i];
if (other->len == k->len && (k->len == 0 || std::memcmp(d.str(*other), key, k->len) == 0))
{
duplicate = true; // keep the first
break;
}
i = (i + 1) & mask;
}
if (!duplicate)
{
slots[i] = static_cast<std::uint32_t>(k - obj);
}
}
d.indexes.push_back(document_data::object_index{start, static_cast<std::uint32_t>(mask)});
obj->extra = static_cast<std::uint16_t>(d.indexes.size());
}
/// build the tables of the large objects the parser noted
inline void build_object_indexes(document_data& d)
{
for (const std::uint32_t i : d.large_objects)
{
build_object_index(d, d.tape + i);
}
}
/// the key node of the first member with this key of an indexed object, or
/// nullptr
inline const node* find_indexed(const document_data& d, const node* obj, const char* key, std::size_t n) noexcept
{
const document_data::object_index& ix = d.indexes[obj->extra - 1u];
const std::uint32_t* const slots = d.index_slots.data() + ix.start;
std::size_t i = static_cast<std::size_t>(key_hash(key, n)) & ix.mask;
for (;;)
{
const std::uint32_t s = slots[i];
if (s == 0)
{
return nullptr;
}
const node* const k = obj + s;
if (k->len == n && (n == 0 || std::memcmp(d.str(*k), key, n) == 0))
{
return k;
}
i = (i + 1) & ix.mask;
}
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+29 -5
View File
@@ -16,6 +16,7 @@
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/simd.hpp>
// Scanning primitives of the view's parser. The unrolled checks at fixed
// offsets follow yyjson (https://github.com/ibireme/yyjson, MIT license): the
@@ -62,9 +63,12 @@ NLOHMANN_VIEW_ALWAYS_INLINE std::uint16_t load16(const unsigned char* p) noexcep
/// Advance over plain string bytes and well-formed UTF-8. Stops at a quote,
/// a backslash, a control character, ill-formed UTF-8, or the end. The first
/// 16 bytes are checked one by one, so that the position advances by
/// constants in predicted branches (most strings are short); longer runs
/// continue eight bytes at a time.
/// bytes are checked one by one, so that the position advances by constants
/// in predicted branches: 16 for keys, whose lengths repeat from record to
/// record, and 8 for string values (Value) where a vector loop follows, as
/// their lengths vary more. Longer runs continue 16 bytes at a time with NEON
/// or SSE2, else eight bytes at a time.
template<bool Value = false>
NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned char* p, const unsigned char* e) noexcept
{
const std::uint8_t* plain = string_plain();
@@ -73,9 +77,23 @@ NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned
if (e - p >= 16)
{
#define NLOHMANN_VIEW_STEP(i) if (NLOHMANN_VIEW_LIKELY(plain[p[i]] != 0)) {} else { p += (i); goto stop; }
NLOHMANN_VIEW_REPEAT16(NLOHMANN_VIEW_STEP)
NLOHMANN_VIEW_STEP(0) NLOHMANN_VIEW_STEP(1) NLOHMANN_VIEW_STEP(2) NLOHMANN_VIEW_STEP(3)
NLOHMANN_VIEW_STEP(4) NLOHMANN_VIEW_STEP(5) NLOHMANN_VIEW_STEP(6) NLOHMANN_VIEW_STEP(7)
if (!Value || !NLOHMANN_VIEW_VECTOR)
{
NLOHMANN_VIEW_STEP(8) NLOHMANN_VIEW_STEP(9) NLOHMANN_VIEW_STEP(10) NLOHMANN_VIEW_STEP(11)
NLOHMANN_VIEW_STEP(12) NLOHMANN_VIEW_STEP(13) NLOHMANN_VIEW_STEP(14) NLOHMANN_VIEW_STEP(15)
p += 8;
}
#undef NLOHMANN_VIEW_STEP
p += 16;
p += 8;
#if NLOHMANN_VIEW_VECTOR
p = vector_plain_run(p, e);
if (p != e && plain[*p] == 0)
{
goto stop;
}
#else
while (e - p >= 8)
{
const std::uint64_t special = swar_string_special(read_eight_bytes(p));
@@ -86,6 +104,7 @@ NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned
}
p += 8;
}
#endif
continue;
}
while (p != e && plain[*p] != 0)
@@ -101,6 +120,10 @@ stop:
{
return p; // quote, backslash, or control character
}
#if NLOHMANN_VIEW_VECTOR_UTF8
// non-ASCII: the vector check, out of line
return scan_string_vector(p, e, plain);
#else
// non-ASCII: a run of well-formed sequences (the library's check, so
// that exactly what json::parse accepts is accepted)
do
@@ -113,6 +136,7 @@ stop:
p += n;
}
while (p != e && *p >= 0x80);
#endif
}
}
+839
View File
@@ -0,0 +1,839 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <algorithm> // max
#include <array> // array
#include <cmath> // isfinite
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t
#include <cstring> // memcpy, memset
#include <limits> // numeric_limits
#include <type_traits> // integral_constant
#include <vector> // vector
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
#include <nlohmann/detail/view/number.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
/// append-only output buffer: writes through a raw pointer into a string that
/// is resized ahead, and trimmed by finish()
template<typename StringType>
class output_buffer
{
public:
output_buffer(StringType& out, std::size_t estimate)
: m_out(sized(out, estimate))
, m_pos(&m_out[0])
, m_end(m_pos + m_out.size())
{}
void finish()
{
m_out.resize(static_cast<std::size_t>(m_pos - m_out.data()));
}
NLOHMANN_VIEW_ALWAYS_INLINE void reserve(std::size_t n)
{
if (NLOHMANN_VIEW_UNLIKELY(static_cast<std::size_t>(m_end - m_pos) < n))
{
grow(n);
}
}
NLOHMANN_VIEW_ALWAYS_INLINE void put(char c)
{
reserve(1);
*m_pos++ = c;
}
NLOHMANN_VIEW_ALWAYS_INLINE void put(const char* s, std::size_t n)
{
reserve(n);
std::memcpy(m_pos, s, n);
m_pos += n;
}
void put_repeated(char c, std::size_t n)
{
reserve(n);
std::memset(m_pos, c, n);
m_pos += n;
}
/// the write position and the end of the writable space, for a writer
/// that keeps the position in a local variable (set_cursor() hands it back)
char* cursor() const noexcept
{
return m_pos;
}
char* limit() const noexcept
{
return m_end;
}
void set_cursor(char* p) noexcept
{
m_pos = p;
}
private:
static StringType& sized(StringType& out, std::size_t estimate)
{
out.resize((std::max)(estimate, static_cast<std::size_t>(64)));
return out;
}
NLOHMANN_VIEW_NOINLINE void grow(std::size_t n)
{
const auto used = static_cast<std::size_t>(m_pos - m_out.data());
m_out.resize((std::max)(m_out.size() * 2, used + n + 256));
m_pos = &m_out[0] + used;
m_end = &m_out[0] + m_out.size();
}
StringType& m_out;
char* m_pos;
char* m_end;
};
/// The length of the run at s that dump() writes unchanged without
/// ensure_ascii: all bytes but quotes, backslashes, and control characters.
/// Unlike detail::string_bulk_run(), non-ASCII bytes are not validated: the
/// strings of a document are valid UTF-8 (a damaged image loaded with
/// image_check::bounds can have others, which are then written unchanged).
inline std::size_t plain_output_run(const unsigned char* s, std::size_t n) noexcept
{
constexpr std::uint64_t ones = 0x0101010101010101ull;
constexpr std::uint64_t high = 0x8080808080808080ull;
std::size_t i = 0;
for (; i + 8 <= n; i += 8)
{
const std::uint64_t v = read_eight_bytes(s + i);
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"'
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\'
const std::uint64_t stop = (((q - ones) & ~q) | ((b - ones) & ~b) | ((v - 0x2020202020202020ull) & ~v)) & high;
if (stop != 0)
{
// the lowest flagged byte is the first stop: borrows only flag bytes above a true one
return i + (static_cast<std::size_t>(count_trailing_zeros(stop)) / 8);
}
}
for (; i < n; ++i)
{
if (s[i] == '"' || s[i] == '\\' || s[i] < 0x20)
{
return i;
}
}
return n;
}
/// A stack that starts in a buffer of the caller (a local array) and moves to
/// the heap (a vector of the caller) only when that is full, so that dumps of
/// shallow documents need no allocation. The top is a pointer, as in
/// std::vector. The address of the stack never escapes (the growth gets the
/// vector and returns the new storage), so its pointers stay in registers.
template<typename T>
class small_stack
{
public:
small_stack(T* buffer, std::size_t capacity, std::vector<T>& heap) noexcept
: m_begin(buffer), m_top(buffer), m_end(buffer + capacity), m_heap(&heap)
{}
small_stack(const small_stack&) = delete;
small_stack(small_stack&&) = delete;
small_stack& operator=(const small_stack&) = delete;
small_stack& operator=(small_stack&&) = delete;
~small_stack() = default;
NLOHMANN_VIEW_ALWAYS_INLINE void push_back(const T& x)
{
if (NLOHMANN_VIEW_UNLIKELY(m_top == m_end))
{
const std::size_t used = size();
const std::size_t capacity = 2 * static_cast<std::size_t>(m_end - m_begin);
m_begin = grow(*m_heap, m_begin, used, capacity);
m_top = m_begin + used;
m_end = m_begin + capacity;
}
*m_top++ = x;
}
NLOHMANN_VIEW_ALWAYS_INLINE T& back() noexcept
{
return m_top[-1];
}
NLOHMANN_VIEW_ALWAYS_INLINE void pop_back() noexcept
{
--m_top;
}
NLOHMANN_VIEW_ALWAYS_INLINE bool empty() const noexcept
{
return m_top == m_begin;
}
NLOHMANN_VIEW_ALWAYS_INLINE std::size_t size() const noexcept
{
return static_cast<std::size_t>(m_top - m_begin);
}
private:
/// the used entries moved to heap storage of the given capacity
NLOHMANN_VIEW_NOINLINE static T* grow(std::vector<T>& heap, const T* begin, std::size_t used, std::size_t capacity)
{
std::vector<T> bigger(capacity);
std::copy(begin, begin + used, bigger.begin());
heap.swap(bigger);
return heap.data();
}
T* m_begin;
T* m_top;
T* m_end;
std::vector<T>* m_heap;
};
/// how the view's dump() writes a value
struct dump_style
{
bool pretty = false; ///< indent >= 0
std::size_t indent = 0; ///< characters per level
char indent_char = ' ';
bool ensure_ascii = false;
bool source_numbers = false; ///< copy number tokens from the source
};
/*!
@brief write a view's subtree as basic_json::dump() writes the value
The output of a subtree equals ordered_json::parse(text).dump() of it for
the same arguments (members in document order): strings are escaped by the
same rules, with the library's scanning kernels; floats are written with
the library's conversion; integers are copied from the source, where they
are canonical (except "-0", which parse() reads as 0). The walk is
iterative, so the nesting depth is limited by memory only.
*/
template<typename BasicJsonType, bool Editable>
class view_serializer
{
using nav = navigation<Editable>;
using string_t = typename BasicJsonType::string_t;
using number_float_t = typename BasicJsonType::number_float_t;
public:
view_serializer(const document_data& d, string_t& out, std::size_t estimate, const dump_style& style)
: m_doc(d), m_out(out, estimate), m_style(style)
{}
void dump(const node* root)
{
if (!m_style.pretty && !m_style.ensure_ascii)
{
if (m_style.source_numbers)
{
dump_compact<true>(root);
}
else
{
dump_compact<false>(root);
}
return;
}
struct frame
{
const node* pos; ///< next element, or key of the next member
const node* end;
bool object;
bool first; ///< nothing written yet
};
std::array<frame, 32> buffer; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
std::vector<frame> heap;
small_stack<frame> stack(buffer.data(), buffer.size(), heap);
const node* n = root;
for (;;)
{
// write the value at n
if (is_container(*n))
{
const bool object = n->kind == static_cast<std::uint8_t>(value_t::object);
if (n->len == 0)
{
m_out.put(object ? "{}" : "[]", 2);
}
else
{
m_out.put(object ? '{' : '[');
stack.push_back(frame{nav::first(m_doc, n), nav::end(m_doc, n), object, true});
}
}
else
{
write_scalar(*n);
}
// go to the next value: close finished containers, then separate
for (;;)
{
if (stack.empty())
{
m_out.finish();
return;
}
frame& f = stack.back();
if (f.pos == f.end)
{
const bool object = f.object;
stack.pop_back();
newline(stack.size());
m_out.put(object ? '}' : ']');
continue;
}
if (!f.first)
{
m_out.put(',');
}
f.first = false;
newline(stack.size());
if (f.object)
{
write_string(*f.pos);
if (m_style.pretty)
{
m_out.put(": ", 2);
}
else
{
m_out.put(':');
}
n = nav::value(f.pos + 1);
f.pos = document_data::after(f.pos + 1);
}
else
{
n = nav::value(f.pos);
f.pos = document_data::after(f.pos);
}
break;
}
}
}
private:
/*!
@brief the compact output without ensure_ascii (the default dump())
The same walk as dump(), with the write position in a local variable
(stores through char pointers would otherwise force a reload of the
buffer's members after each one), and with strings and number tokens of
the source copied by fixed-size moves of 32 bytes where the source has
that many bytes left, instead of a library call per token. The buffer
keeps 64 bytes of slack for the overshoot.
*/
/// a string that is not a plain string of the source (decoded, or written
/// by an edit), without ensure_ascii: runs without characters to escape
/// are copied
NLOHMANN_VIEW_NOINLINE void write_decoded(const node& n)
{
const auto* const s = reinterpret_cast<const unsigned char*>(m_doc.str(n)); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
m_out.put('"');
for (std::size_t i = 0; i < n.len;)
{
const std::size_t run = plain_output_run(s + i, n.len - i);
if (run != 0)
{
m_out.put(reinterpret_cast<const char*>(s + i), run); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
i += run;
continue;
}
write_codepoint<false>(s[i], s + i, 1); // a quote, a backslash, or a control character
++i;
}
m_out.put('"');
}
/// the copies of dump_compact() that are not fixed-size moves (long
/// strings, or near the end of the source); out of line, so that the
/// compiler does not merge the fixed-size moves into this call
NLOHMANN_VIEW_NOINLINE static void copy_long(char* to, const char* from, std::size_t n) noexcept
{
std::memcpy(to, from, n);
}
template<bool SourceNumbers>
void dump_compact(const node* root)
{
struct frame
{
const node* pos; ///< (editable documents) next element, or key of the next member
const node* end;
bool object;
};
std::array<frame, 32> buffer; // NOLINT(cppcoreguidelines-pro-type-member-init,hicpp-member-init): written before read
std::vector<frame> heap;
small_stack<frame> stack(buffer.data(), buffer.size(), heap);
const char* const src = m_doc.src;
const char* const src_end = src + m_doc.size;
char* w = m_out.cursor();
char* lim = m_out.limit();
// room for n bytes and the slack
const auto room = [&](std::size_t n)
{
if (NLOHMANN_VIEW_UNLIKELY(static_cast<std::size_t>(lim - w) < n + 64))
{
m_out.set_cursor(w);
m_out.reserve(n + 64);
w = m_out.cursor();
lim = m_out.limit();
}
};
// copy n bytes of the source (after room(n))
const auto copy = [&](const char* from, std::size_t n)
{
if (n <= 32 && src_end - from >= 32)
{
std::memcpy(w, from, 32);
}
else if (n <= 256 && src_end - from >= static_cast<std::ptrdiff_t>(n) + 32)
{
for (std::size_t i = 0; i < n; i += 32)
{
std::memcpy(w + i, from + i, 32);
}
}
else
{
copy_long(w, from, n);
}
w += n;
};
// a literal of n bytes (after room(n))
const auto literal = [&](const char* text, std::size_t n)
{
std::memcpy(w, text, n);
w += n;
};
// a string that is not a plain string of the source (out of line, so
// that the cursor stays in a register here)
const auto escaped = [&](const node & n)
{
m_out.set_cursor(w);
write_decoded(n);
w = m_out.cursor();
lim = m_out.limit();
};
// Read-only documents: the elements of a container follow it in the
// node array, so the walk goes through the array in order, and a
// frame only needs the end of its container. Editable documents: the
// elements of a moved container live elsewhere, so a frame keeps the
// position of the next element (see navigation).
// The innermost open container is kept in registers (cur; end ==
// nullptr: none), the stack holds the ones around it.
frame cur{nullptr, nullptr, false};
const node* n = root;
for (;;)
{
// write the value at n (read-only documents: and advance n)
bool opened = false;
switch (static_cast<value_t>(n->kind))
{
case value_t::string:
if ((n->flags & node_flags::storage) == 0)
{
room(n->len + 2);
*w++ = '"';
copy(src + n->off, n->len);
*w++ = '"';
}
else
{
escaped(*n);
}
break;
case value_t::number_integer:
case value_t::number_unsigned:
{
const std::uint32_t len = number_length(*n);
room(len);
if (Editable && (n->flags & node_flags::storage) != 0)
{
copy_long(w, m_doc.str(*n), len); // a canonical token written by an edit
w += len;
break;
}
const char* const token = src + n->off;
if (!SourceNumbers && NLOHMANN_VIEW_UNLIKELY(len == 2 && token[0] == '-' && token[1] == '0'))
{
*w++ = '0'; // parse() reads -0 as the integer 0
}
else
{
copy(token, len);
}
break;
}
case value_t::number_float:
if (SourceNumbers && (n->flags & node_flags::storage) != node_flags::edited)
{
room(n->len);
copy(src + n->off, n->len);
}
else
{
m_out.set_cursor(w);
write_float(float_value<number_float_t>(m_doc, *n));
w = m_out.cursor();
lim = m_out.limit();
}
break;
case value_t::boolean:
room(8);
if ((n->flags & node_flags::is_true) != 0)
{
literal("true", 4);
}
else
{
literal("false", 5);
}
break;
case value_t::object:
case value_t::array:
{
const bool object = n->kind == static_cast<std::uint8_t>(value_t::object);
room(8);
if (n->len == 0)
{
literal(object ? "{}" : "[]", 2);
}
else
{
*w++ = object ? '{' : '[';
stack.push_back(cur);
if (Editable)
{
cur = frame{nav::first(m_doc, n), nav::end(m_doc, n), object};
}
else
{
cur = frame{nullptr, n + n->next, object};
}
opened = true;
}
break;
}
case value_t::null:
room(8);
literal("null", 4);
break;
case value_t::binary: // LCOV_EXCL_LINE (not in a document)
case value_t::discarded: // LCOV_EXCL_LINE
default: // LCOV_EXCL_LINE
break; // LCOV_EXCL_LINE
}
if (!Editable)
{
++n; // the next node: the first element of an opened container, or the node after a scalar
}
// go to the next value: close finished containers, then separate
// (a container just opened has an element)
if (!opened)
{
for (;;)
{
if (cur.end == nullptr)
{
m_out.set_cursor(w);
m_out.finish();
return;
}
if ((Editable ? cur.pos : n) != cur.end)
{
break;
}
room(1);
*w++ = cur.object ? '}' : ']';
cur = stack.back();
stack.pop_back();
}
room(1);
*w++ = ',';
}
const node* const at = Editable ? cur.pos : n;
if (cur.object)
{
const node& key = *at;
if ((key.flags & node_flags::storage) == 0)
{
room(key.len + 3);
*w++ = '"';
copy(src + key.off, key.len);
w[0] = '"';
w[1] = ':';
w += 2;
}
else
{
escaped(key);
room(1);
*w++ = ':';
}
if (Editable)
{
n = nav::value(at + 1);
cur.pos = document_data::after(at + 1);
}
else
{
++n;
}
}
else if (Editable)
{
n = nav::value(at);
cur.pos = document_data::after(at);
}
}
}
void newline(std::size_t level)
{
if (m_style.pretty)
{
m_out.put('\n');
m_out.put_repeated(m_style.indent_char, level * m_style.indent);
}
}
void write_scalar(const node& n)
{
switch (static_cast<value_t>(n.kind))
{
case value_t::null:
m_out.put("null", 4);
break;
case value_t::boolean:
if ((n.flags & node_flags::is_true) != 0)
{
m_out.put("true", 4);
}
else
{
m_out.put("false", 5);
}
break;
case value_t::string:
write_string(n);
break;
case value_t::number_integer:
case value_t::number_unsigned:
{
const char* const token = m_doc.str(n);
const std::uint32_t len = number_length(n);
if (!m_style.source_numbers && len == 2 && token[0] == '-' && token[1] == '0')
{
m_out.put('0'); // parse() reads -0 as the integer 0
}
else
{
m_out.put(token, len);
}
break;
}
case value_t::number_float:
if (m_style.source_numbers && (n.flags & node_flags::storage) != node_flags::edited)
{
m_out.put(m_doc.str(n), n.len); // (a float set by an edit is written as with shortest)
}
else
{
write_float(float_value<number_float_t>(m_doc, n));
}
break;
case value_t::object: // LCOV_EXCL_LINE (containers are written by dump())
case value_t::array: // LCOV_EXCL_LINE
case value_t::binary: // LCOV_EXCL_LINE (not in a document)
case value_t::discarded: // LCOV_EXCL_LINE
default: // LCOV_EXCL_LINE
break; // LCOV_EXCL_LINE
}
}
/// as serializer::dump_float()
void write_float(number_float_t x)
{
if (!std::isfinite(x))
{
m_out.put("null", 4);
return;
}
write_float(x, std::integral_constant < bool,
(std::numeric_limits<number_float_t>::is_iec559 && std::numeric_limits<number_float_t>::digits == 24 && std::numeric_limits<number_float_t>::max_exponent == 128)
|| (std::numeric_limits<number_float_t>::is_iec559 && std::numeric_limits<number_float_t>::digits == 53 && std::numeric_limits<number_float_t>::max_exponent == 1024) > {});
}
void write_float(number_float_t x, std::true_type /*is_ieee_single_or_double*/)
{
std::array<char, 64> buf{};
const char* const end = ::nlohmann::detail::to_chars(buf.data(), buf.data() + buf.size(), x);
m_out.put(buf.data(), static_cast<std::size_t>(end - buf.data()));
}
void write_float(number_float_t x, std::false_type /*is_ieee_single_or_double*/)
{
// other types (e.g. long double) are rare: the library writes them
const string_t s = BasicJsonType(x).dump();
m_out.put(s.data(), s.size());
}
void write_string(const node& n)
{
const char* const s = m_doc.str(n);
m_out.put('"');
if ((n.flags & node_flags::storage) == 0 && !m_style.ensure_ascii)
{
// a string of the source without escape sequences has nothing to escape
m_out.put(s, n.len);
}
else if (m_style.ensure_ascii)
{
write_escaped<true>(reinterpret_cast<const unsigned char*>(s), n.len); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
}
else
{
write_escaped<false>(reinterpret_cast<const unsigned char*>(s), n.len); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
}
m_out.put('"');
}
/// as serializer::dump_escaped(); strings of a document are valid UTF-8,
/// except in a damaged image loaded with image_check::bounds, for which
/// this throws what basic_json::dump() throws for the string
template<bool EnsureAscii>
void write_escaped(const unsigned char* s, std::size_t n)
{
std::size_t i = 0;
while (i < n)
{
std::size_t run = 0;
if (!EnsureAscii)
{
run = string_bulk_run(s + i, n - i);
}
else if (is_ascii_copyable(s[i]))
{
run = find_ascii_copyable_run(s + i, n - i);
}
if (run != 0)
{
m_out.put(reinterpret_cast<const char*>(s + i), run); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
i += run;
continue;
}
std::uint32_t codepoint = s[i];
std::size_t len = 1;
if (codepoint >= 0x80)
{
len = validate_one_utf8(s + i, n - i);
if (NLOHMANN_VIEW_UNLIKELY(len == 0))
{
invalid_utf8(s, n);
return;
}
codepoint &= 0xFFu >> (len + 1);
for (std::size_t k = 1; k < len; ++k)
{
codepoint = (codepoint << 6u) | (s[i + k] & 0x3Fu);
}
}
write_codepoint<EnsureAscii>(codepoint, s + i, len);
i += len;
}
}
/// throw what basic_json::dump() throws for a string that is not valid UTF-8
NLOHMANN_VIEW_NOINLINE static void invalid_utf8(const unsigned char* s, std::size_t n)
{
const string_t dumped = BasicJsonType(string_t(reinterpret_cast<const char*>(s), n)).dump(); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
static_cast<void>(dumped);
}
template<bool EnsureAscii>
void write_codepoint(std::uint32_t codepoint, const unsigned char* bytes, std::size_t len)
{
switch (codepoint)
{
case 0x08:
m_out.put("\\b", 2);
return;
case 0x09:
m_out.put("\\t", 2);
return;
case 0x0A:
m_out.put("\\n", 2);
return;
case 0x0C:
m_out.put("\\f", 2);
return;
case 0x0D:
m_out.put("\\r", 2);
return;
case 0x22:
m_out.put("\\\"", 2);
return;
case 0x5C:
m_out.put("\\\\", 2);
return;
default:
break;
}
if (codepoint <= 0x1F || (EnsureAscii && codepoint >= 0x7F))
{
if (codepoint <= 0xFFFF)
{
write_u_escape(codepoint);
}
else
{
write_u_escape(0xD7C0u + (codepoint >> 10u));
write_u_escape(0xDC00u + (codepoint & 0x3FFu));
}
return;
}
m_out.put(reinterpret_cast<const char*>(bytes), len); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast) LCOV_EXCL_LINE (printable characters are copied in runs)
}
void write_u_escape(std::uint32_t u)
{
static constexpr const char* hex = "0123456789abcdef";
const std::array<char, 6> e = {{'\\', 'u', hex[(u >> 12u) & 0xFu], hex[(u >> 8u) & 0xFu], hex[(u >> 4u) & 0xFu], hex[u & 0xFu]}};
m_out.put(e.data(), e.size());
}
const document_data& m_doc;
output_buffer<string_t> m_out;
const dump_style m_style;
};
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+281
View File
@@ -0,0 +1,281 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-FileCopyrightText: 2018-2025 The simdjson authors <https://github.com/simdjson/simdjson>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint64_t
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
// Vector code for long runs of string bytes. NEON (AArch64) and SSE2 (x86-64)
// belong to the baseline instruction sets and are used by default. The vector
// UTF-8 check needs NEON, or SSSE3 if JSON_VIEW_USE_SSSE3 is defined: SSSE3 is
// not part of x86-64, so it must not depend on the flags of a translation unit
// (two translation units with different flags would have different
// definitions of the same inline functions). JSON_VIEW_NO_SIMD selects the
// portable code.
#if !defined(JSON_VIEW_NO_SIMD) && defined(__aarch64__) && (defined(__GNUC__) || defined(__clang__)) && NLOHMANN_VIEW_LITTLE_ENDIAN
#include <arm_neon.h>
#define NLOHMANN_VIEW_NEON 1
#else
#define NLOHMANN_VIEW_NEON 0
#endif
#if !defined(JSON_VIEW_NO_SIMD) && !NLOHMANN_VIEW_NEON && (defined(__SSE2__) || defined(_M_X64) || (defined(_M_IX86_FP) && _M_IX86_FP >= 2))
#include <emmintrin.h>
#define NLOHMANN_VIEW_SSE2 1
#else
#define NLOHMANN_VIEW_SSE2 0
#endif
#if NLOHMANN_VIEW_SSE2 && defined(JSON_VIEW_USE_SSSE3)
#include <tmmintrin.h>
#define NLOHMANN_VIEW_SSSE3 1 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#else
#define NLOHMANN_VIEW_SSSE3 0 // NOLINT(cppcoreguidelines-macro-to-enum,modernize-macro-to-enum)
#endif
#define NLOHMANN_VIEW_VECTOR (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSE2)
#define NLOHMANN_VIEW_VECTOR_UTF8 (NLOHMANN_VIEW_NEON || NLOHMANN_VIEW_SSSE3)
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
#if NLOHMANN_VIEW_VECTOR
/*!
@brief the first byte of a string run that is a quote, a backslash, a control
character, or not ASCII, 16 bytes per step
Stops at such a byte, or where fewer than 16 bytes are left (the caller tells
the two apart). A signed compare with 0x20 finds control characters and
non-ASCII bytes at once.
*/
NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* vector_plain_run(const unsigned char* p, const unsigned char* e) noexcept
{
while (e - p >= 16)
{
#if NLOHMANN_VIEW_NEON
const uint8x16_t in = vld1q_u8(p);
const uint8x16_t special = vorrq_u8(vorrq_u8(vceqq_u8(in, vdupq_n_u8('"')), vceqq_u8(in, vdupq_n_u8('\\'))),
vcltq_s8(vreinterpretq_s8_u8(in), vdupq_n_s8(0x20)));
// one nibble per byte (the usual NEON replacement of x86's movemask, see
// D. Kutenin, "Porting x86 vector bitmask optimizations to Arm NEON", 2022)
const std::uint64_t bits = vget_lane_u64(vreinterpret_u64_u8(vshrn_n_u16(vreinterpretq_u16_u8(special), 4)), 0);
if (bits != 0)
{
return p + (count_trailing_zeros(bits) >> 2u);
}
#else
const __m128i in = _mm_loadu_si128(static_cast<const __m128i*>(static_cast<const void*>(p)));
const __m128i special = _mm_or_si128(_mm_or_si128(_mm_cmpeq_epi8(in, _mm_set1_epi8('"')), _mm_cmpeq_epi8(in, _mm_set1_epi8('\\'))),
_mm_cmplt_epi8(in, _mm_set1_epi8(0x20)));
const auto bits = static_cast<std::uint64_t>(static_cast<unsigned>(_mm_movemask_epi8(special)));
if (bits != 0)
{
return p + count_trailing_zeros(bits);
}
#endif
p += 16;
}
return p;
}
#endif
#if NLOHMANN_VIEW_VECTOR_UTF8
/// Tables of the UTF-8 check of J. Keiser and D. Lemire, "Validating UTF-8 In
/// Less Than One Instruction Per Byte" (2021), as in simdjson ("lookup4"): each
/// maps a nibble (high and low nibble of the previous byte, high nibble of the
/// current byte) to the errors it allows; a byte pair is ill-formed if all
/// three have an error bit in common.
template<typename Dummy = void>
struct utf8_lookup4
{
static constexpr std::uint8_t too_short = 1u << 0u, too_long = 1u << 1u, overlong_3 = 1u << 2u, too_large = 1u << 3u;
static constexpr std::uint8_t surrogate = 1u << 4u, overlong_2 = 1u << 5u, too_large_1000 = 1u << 6u, overlong_4 = 1u << 6u;
static constexpr std::uint8_t two_conts = 1u << 7u, carry = too_short | too_long | two_conts;
static const std::array<std::uint8_t, 16> byte_1_high;
static const std::array<std::uint8_t, 16> byte_1_low;
static const std::array<std::uint8_t, 16> byte_2_high;
};
template<typename Dummy>
const std::array<std::uint8_t, 16> utf8_lookup4<Dummy>::byte_1_high =
{
{
too_long, too_long, too_long, too_long, too_long, too_long, too_long, too_long,
two_conts, two_conts, two_conts, two_conts,
too_short | overlong_2, too_short, too_short | overlong_3 | surrogate, too_short | too_large | too_large_1000 | overlong_4
}
};
template<typename Dummy>
const std::array<std::uint8_t, 16> utf8_lookup4<Dummy>::byte_1_low =
{
{
carry | overlong_3 | overlong_2 | overlong_4, carry | overlong_2, carry, carry,
carry | too_large, carry | too_large | too_large_1000, carry | too_large | too_large_1000, carry | too_large | too_large_1000,
carry | too_large | too_large_1000, carry | too_large | too_large_1000, carry | too_large | too_large_1000, carry | too_large | too_large_1000,
carry | too_large | too_large_1000, carry | too_large | too_large_1000 | surrogate, carry | too_large | too_large_1000, carry | too_large | too_large_1000
}
};
template<typename Dummy>
const std::array<std::uint8_t, 16> utf8_lookup4<Dummy>::byte_2_high =
{
{
too_short, too_short, too_short, too_short, too_short, too_short, too_short, too_short,
static_cast<std::uint8_t>(too_long | overlong_2 | two_conts | overlong_3 | too_large_1000 | overlong_4),
static_cast<std::uint8_t>(too_long | overlong_2 | two_conts | overlong_3 | too_large),
static_cast<std::uint8_t>(too_long | overlong_2 | two_conts | surrogate | too_large),
static_cast<std::uint8_t>(too_long | overlong_2 | two_conts | surrogate | too_large),
too_short, too_short, too_short, too_short
}
};
/// the end of scan_string_vector from block, where the vector loop stopped
/// (ill-formed UTF-8, or fewer than 16 bytes left): one byte or sequence at a
/// time, from the start of a sequence that crosses into the block
inline const unsigned char* scan_string_finish(const unsigned char* p, const unsigned char* block, const unsigned char* e, const std::uint8_t* plain) noexcept
{
for (int i = 1; i <= 3 && block - i >= p; ++i)
{
const unsigned char c = block[-i];
if (c < 0x80)
{
break;
}
if (c >= 0xC0)
{
const int len = 2 + static_cast<int>(c >= 0xE0) + static_cast<int>(c >= 0xF0);
if (len > i)
{
block -= i;
}
break;
}
}
for (p = block; p != e;)
{
if (*p < 0x80)
{
if (plain[*p] == 0)
{
return p;
}
++p;
continue;
}
const std::size_t n = validate_one_utf8(p, static_cast<std::size_t>(e - p));
if (n == 0)
{
return p;
}
p += n;
}
return p;
}
/*!
@brief the rest of a string from p (a character boundary), 16 bytes per step
The first quote, backslash, or control character is found with vector
compares, and the UTF-8 check covers the bytes up to it. Returns where the
string scan stops, like scan_string_run: before ill-formed UTF-8 and for the
last bytes of the input, the bytes are checked one sequence at a time. Out of
line, so that no constants of the check occupy registers in the parse loop.
*/
NLOHMANN_VIEW_NOINLINE inline const unsigned char* scan_string_vector(const unsigned char* p, const unsigned char* e, const std::uint8_t* plain) noexcept
{
using lookup = utf8_lookup4<>;
const unsigned char* block = p;
#if NLOHMANN_VIEW_NEON
const uint8x16_t t1h = vld1q_u8(lookup::byte_1_high.data());
const uint8x16_t t1l = vld1q_u8(lookup::byte_1_low.data());
const uint8x16_t t2h = vld1q_u8(lookup::byte_2_high.data());
uint8x16_t prev = vdupq_n_u8(0);
while (e - block >= 16)
{
const uint8x16_t in = vld1q_u8(block);
const uint8x16_t special = vorrq_u8(vorrq_u8(vceqq_u8(in, vdupq_n_u8('"')), vceqq_u8(in, vdupq_n_u8('\\'))), vcltq_u8(in, vdupq_n_u8(0x20)));
const uint8x16_t prev1 = vextq_u8(prev, in, 15);
const uint8x16_t sc = vandq_u8(vandq_u8(vqtbl1q_u8(t1h, vshrq_n_u8(prev1, 4)), vqtbl1q_u8(t1l, vandq_u8(prev1, vdupq_n_u8(0x0F)))), vqtbl1q_u8(t2h, vshrq_n_u8(in, 4)));
const uint8x16_t must23 = vorrq_u8(vqsubq_u8(vextq_u8(prev, in, 14), vdupq_n_u8(0xE0 - 0x80)), vqsubq_u8(vextq_u8(prev, in, 13), vdupq_n_u8(0xF0 - 0x80)));
const uint8x16_t err = veorq_u8(vandq_u8(must23, vdupq_n_u8(0x80)), sc);
const std::uint64_t special_bits = vget_lane_u64(vreinterpret_u64_u8(vshrn_n_u16(vreinterpretq_u16_u8(special), 4)), 0);
const std::uint64_t err_bits = vget_lane_u64(vreinterpret_u64_u8(vshrn_n_u16(vreinterpretq_u16_u8(vtstq_u8(err, err)), 4)), 0);
if (special_bits != 0)
{
// errors up to the special byte count (an incomplete sequence
// before a quote shows at the quote); the bytes after it do not
const unsigned k = static_cast<unsigned>(count_trailing_zeros(special_bits)) >> 2u;
const std::uint64_t upto = k == 15 ? ~std::uint64_t{0} :
(std::uint64_t{1} << (4u * (k + 1u))) - 1u;
if ((err_bits & upto) == 0)
{
return block + k;
}
break;
}
if (err_bits != 0)
{
break;
}
prev = in;
block += 16;
}
#else
// the same with SSSE3 (pshufb for the table lookups; nibbles from 16-bit
// shifts, as there are no byte shifts)
const __m128i t1h = _mm_loadu_si128(static_cast<const __m128i*>(static_cast<const void*>(lookup::byte_1_high.data())));
const __m128i t1l = _mm_loadu_si128(static_cast<const __m128i*>(static_cast<const void*>(lookup::byte_1_low.data())));
const __m128i t2h = _mm_loadu_si128(static_cast<const __m128i*>(static_cast<const void*>(lookup::byte_2_high.data())));
const __m128i nibble = _mm_set1_epi8(0x0F);
const __m128i zero = _mm_setzero_si128();
__m128i prev = zero;
while (e - block >= 16)
{
const __m128i in = _mm_loadu_si128(static_cast<const __m128i*>(static_cast<const void*>(block)));
const __m128i special = _mm_or_si128(_mm_or_si128(_mm_cmpeq_epi8(in, _mm_set1_epi8('"')), _mm_cmpeq_epi8(in, _mm_set1_epi8('\\'))),
_mm_cmpeq_epi8(_mm_subs_epu8(in, _mm_set1_epi8(0x1F)), zero)); // in < 0x20
const __m128i prev1 = _mm_alignr_epi8(in, prev, 15);
const __m128i sc = _mm_and_si128(_mm_and_si128(_mm_shuffle_epi8(t1h, _mm_and_si128(_mm_srli_epi16(prev1, 4), nibble)),
_mm_shuffle_epi8(t1l, _mm_and_si128(prev1, nibble))),
_mm_shuffle_epi8(t2h, _mm_and_si128(_mm_srli_epi16(in, 4), nibble)));
const __m128i must23 = _mm_or_si128(_mm_subs_epu8(_mm_alignr_epi8(in, prev, 14), _mm_set1_epi8(0xE0 - 0x80)),
_mm_subs_epu8(_mm_alignr_epi8(in, prev, 13), _mm_set1_epi8(0xF0 - 0x80)));
const __m128i err = _mm_xor_si128(_mm_and_si128(must23, _mm_set1_epi8(static_cast<char>(-128))), sc);
const auto special_bits = static_cast<unsigned>(_mm_movemask_epi8(special));
const auto err_bits = ~static_cast<unsigned>(_mm_movemask_epi8(_mm_cmpeq_epi8(err, zero))) & 0xFFFFu;
if (special_bits != 0)
{
const unsigned k = static_cast<unsigned>(count_trailing_zeros(static_cast<std::uint64_t>(special_bits)));
if ((err_bits & ((2u << k) - 1u)) == 0)
{
return block + k;
}
break;
}
if (err_bits != 0)
{
break;
}
prev = in;
block += 16;
}
#endif
return scan_string_finish(p, block, e, plain);
}
#endif
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+1 -1
View File
@@ -49,7 +49,7 @@ NLOHMANN_VIEW_ALWAYS_INLINE T arithmetic_value(const document_data& d, const nod
case value_t::number_integer:
return static_cast<T>(static_cast<typename BasicJsonType::number_integer_t>(static_cast<std::int64_t>(integer_bits(n))));
case value_t::number_float:
return static_cast<T>(float_value<typename BasicJsonType::number_float_t>(d.str(n), n));
return static_cast<T>(float_value<typename BasicJsonType::number_float_t>(d, n));
case value_t::boolean:
return static_cast<T>((n.flags & node_flags::is_true) != 0);
case value_t::null:
+398 -22
View File
@@ -25,10 +25,14 @@
#define INCLUDE_NLOHMANN_JSON_VIEW_HPP_
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t
#include <cstring> // memcpy, strlen
#include <iterator> // distance, input_iterator_tag, iterator_traits
#include <map> // map
#include <memory> // unique_ptr
#ifndef JSON_NO_IO
#include <ostream> // ostream
#endif
#include <string> // string
#include <tuple> // tuple_element, tuple_size
#include <type_traits> // decay, enable_if, integral_constant, is_arithmetic, is_base_of, is_integral, is_same, remove_cv, remove_extent
@@ -44,21 +48,27 @@
#endif
#include <nlohmann/detail/view/builder.hpp>
#include <nlohmann/detail/view/compare.hpp>
#include <nlohmann/detail/view/document_data.hpp>
#include <nlohmann/detail/view/edit.hpp>
#include <nlohmann/detail/view/edit_storage.hpp>
#include <nlohmann/detail/view/errors.hpp>
#include <nlohmann/detail/view/image.hpp>
#include <nlohmann/detail/view/input.hpp>
#include <nlohmann/detail/view/iterator.hpp>
#include <nlohmann/detail/view/lookup.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/materialize.hpp>
#include <nlohmann/detail/view/node.hpp>
#include <nlohmann/detail/view/object_index.hpp>
#include <nlohmann/detail/view/pointer.hpp>
#include <nlohmann/detail/view/serializer.hpp>
#include <nlohmann/detail/view/string_ref.hpp>
#include <nlohmann/detail/view/value.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
template<typename BasicJsonType>
template<typename BasicJsonType, bool Editable>
class basic_json_document;
/*!
@@ -67,11 +77,13 @@ class basic_json_document;
Trivially copyable (two pointers). Valid as long as the document is alive and
has not been re-parsed, and as long as a borrowed source text is alive.
*/
template<typename BasicJsonType>
template<typename BasicJsonType, bool Editable = false>
class basic_json_view
{
using node = detail::view::node;
using document_data = detail::view::document_data;
/// how the index is walked (with edits only for editable documents)
using navigation = detail::view::navigation<Editable>;
public:
using value_t = detail::value_t;
@@ -264,7 +276,7 @@ class basic_json_view
{
detail::view::throw_type_error(305, "cannot use operator[] with a numeric argument with ", type_name());
}
return idx < m_node->len ? basic_json_view(m_doc, detail::view::element_at(m_node, idx)) : basic_json_view();
return idx < m_node->len ? basic_json_view(m_doc, navigation::value(detail::view::element_at<Editable>(*m_doc, m_node, idx))) : basic_json_view();
}
/// (an int argument would be ambiguous between size_type and const char*)
@@ -320,7 +332,7 @@ class basic_json_view
{
detail::view::throw_out_of_range(401, detail::concat("array index ", std::to_string(idx), " is out of range"));
}
return basic_json_view(m_doc, detail::view::element_at(m_node, idx));
return basic_json_view(m_doc, navigation::value(detail::view::element_at<Editable>(*m_doc, m_node, idx)));
}
basic_json_view at(int idx) const
@@ -392,7 +404,7 @@ class basic_json_view
{
if (is_structured() && m_node->len != 0)
{
return basic_json_view(m_doc, detail::view::last_child(m_node) + (is_object() ? 1 : 0));
return basic_json_view(m_doc, navigation::value(detail::view::last_child<Editable>(*m_doc, m_node) + (is_object() ? 1 : 0)));
}
return front();
}
@@ -409,7 +421,7 @@ class basic_json_view
{
return end();
}
const node* const k = detail::view::find_member(*m_doc, m_node, key.data(), key.size());
const node* const k = detail::view::find_member<Editable>(*m_doc, m_node, key.data(), key.size());
return k != nullptr ? iterator(m_doc, k, true) : end();
}
@@ -426,7 +438,7 @@ class basic_json_view
/// whether this is an object with a member with this key
bool contains(string_view_t key) const
{
return is_object() && detail::view::find_member(*m_doc, m_node, key.data(), key.size()) != nullptr;
return is_object() && detail::view::find_member<Editable>(*m_doc, m_node, key.data(), key.size()) != nullptr;
}
bool contains(const char* key) const
@@ -473,7 +485,7 @@ class basic_json_view
{
if (NLOHMANN_VIEW_LIKELY(is_structured()))
{
return iterator(m_doc, document_data::first_child(m_node), is_object());
return iterator(m_doc, navigation::first(*m_doc, m_node), is_object());
}
return iterator(m_doc, m_node, false);
}
@@ -482,7 +494,7 @@ class basic_json_view
{
if (NLOHMANN_VIEW_LIKELY(is_structured()))
{
return iterator(m_doc, document_data::child_end(m_node), is_object());
return iterator(m_doc, navigation::end(*m_doc, m_node), is_object());
}
return iterator(m_doc, (is_null() || is_discarded()) ? m_node : m_node + 1, false);
}
@@ -547,6 +559,111 @@ class basic_json_view
return {m_doc->str(*m_node), detail::view::number_length(*m_node)};
}
///////////////////
// serialization //
///////////////////
/// how dump() writes numbers
enum class number_format
{
/// as basic_json::dump(): integers canonically, floats with the
/// library's shortest round-trip digits ("1.5", "100.0", "1e+100")
shortest,
/// the number text of the source as it is ("1.50", "1E2", "-0", all
/// digits of a long integer)
source,
};
/// the text of this value; with number_format::shortest, the output of
/// ordered_json::parse(text).dump() with the same arguments (members in
/// document order, all of them should a key occur more than once)
string_t dump(const int indent = -1, const char indent_char = ' ', const bool ensure_ascii = false,
const number_format numbers = number_format::shortest) const
{
string_t out;
if (m_node == nullptr)
{
out = "<discarded>"; // as basic_json::dump() of a discarded value
return out;
}
detail::view::dump_style style;
style.pretty = indent >= 0;
style.indent = indent >= 0 ? static_cast<std::size_t>(indent) : 0;
style.indent_char = indent_char;
style.ensure_ascii = ensure_ascii;
style.source_numbers = numbers == number_format::source;
// the compact text is about as long as the source text of the value;
// the compact writer keeps 64 bytes of slack, so that it does not grow
// the buffer just before the end
const std::size_t estimate = source_extent() + (style.pretty ? source_extent() / 2 : 0) + 160;
detail::view::view_serializer<BasicJsonType, Editable>(*m_doc, out, estimate, style).dump(m_node);
return out;
}
#ifndef JSON_NO_IO
/// as operator<< of basic_json: a stream width > 0 is the indentation,
/// the fill character the indentation character
friend std::ostream& operator<<(std::ostream& o, const basic_json_view& v)
{
const bool pretty = o.width() > 0;
const auto indentation = pretty ? o.width() : 0;
o.width(0);
const string_t s = v.dump(pretty ? static_cast<int>(indentation) : -1, o.fill());
return o.write(s.data(), static_cast<std::streamsize>(s.size()));
}
#endif
////////////////
// comparison //
////////////////
/// whether the values parse() would produce for two views are equal, as
/// by BasicJsonType's operator== (numbers by value, objects by their
/// members with duplicate keys resolved as parse() resolves them)
friend bool operator==(const basic_json_view& a, const basic_json_view& b)
{
return detail::view::equal<BasicJsonType>(side(a), side(b));
}
friend bool operator!=(const basic_json_view& a, const basic_json_view& b)
{
return !(a == b);
}
/// a view of an editable document compares with one of a read-only document
template < bool E, typename std::enable_if < E != Editable, int >::type = 0 >
friend bool operator==(const basic_json_view& a, const basic_json_view<BasicJsonType, E>& b)
{
return detail::view::equal<BasicJsonType>(side(a), detail::view::view_side<BasicJsonType, basic_json_view<BasicJsonType, E>>(b));
}
template < bool E, typename std::enable_if < E != Editable, int >::type = 0 >
friend bool operator!=(const basic_json_view& a, const basic_json_view<BasicJsonType, E>& b)
{
return !(a == b);
}
/// whether the value parse() would produce for a view equals a value
friend bool operator==(const basic_json_view& a, const BasicJsonType& j)
{
return detail::view::equal<BasicJsonType>(side(a), json_side_t(j));
}
friend bool operator==(const BasicJsonType& j, const basic_json_view& a)
{
return a == j;
}
friend bool operator!=(const basic_json_view& a, const BasicJsonType& j)
{
return !(a == j);
}
friend bool operator!=(const BasicJsonType& j, const basic_json_view& a)
{
return !(a == j);
}
/////////////////
// materialize //
/////////////////
@@ -559,7 +676,7 @@ class basic_json_view
{
return BasicJsonType(value_t::discarded);
}
return detail::view::materialize<BasicJsonType>(*m_doc, m_node);
return detail::view::materialize<BasicJsonType, Editable>(*m_doc, m_node);
}
/// byte offset of this value in the source text (for strings: of the
@@ -567,24 +684,55 @@ class basic_json_view
/// a discarded view and for strings with escapes, which are decoded
std::size_t source_offset() const noexcept
{
return m_node != nullptr && (m_node->flags & detail::view::node_flags::storage) == 0
return m_node != nullptr && (m_node->flags & (detail::view::node_flags::storage | detail::view::node_flags::moved | detail::view::node_flags::is_new)) == 0
? m_node->off : static_cast<std::size_t>(-1);
}
private:
template<typename> friend class basic_json_document;
template<typename, bool> friend class basic_json_document;
template<typename, bool> friend class basic_json_view;
template<typename, typename> friend class detail::view::editor;
friend iterator;
basic_json_view(const document_data* d, const node* n) noexcept
: m_doc(d), m_node(n)
{}
using json_side_t = detail::view::json_side<BasicJsonType, string_view_t>;
static detail::view::view_side<BasicJsonType, basic_json_view> side(const basic_json_view& v) noexcept
{
return detail::view::view_side<BasicJsonType, basic_json_view>(v);
}
/// the number of source bytes of this value (estimated for values with
/// decoded strings)
std::size_t source_extent() const noexcept
{
if (Editable && m_doc->edits != nullptr)
{
// positions of moved and new values are not source offsets
return m_node == m_doc->tape ? m_doc->size + m_doc->edits->text_used : 64;
}
const node* const next = document_data::after(m_node);
const bool in_source = (m_node->flags & detail::view::node_flags::storage) == 0;
if (!in_source)
{
return m_node->len;
}
if (next != m_doc->tape + m_doc->tape_size && (next->flags & detail::view::node_flags::storage) == 0 && next->off >= m_node->off)
{
return next->off - m_node->off;
}
return m_doc->size - m_node->off;
}
/// the value of the first member with this key, or a discarded view
/// (object required)
NLOHMANN_VIEW_ALWAYS_INLINE basic_json_view lookup(string_view_t key) const noexcept
{
const node* const k = detail::view::find_member(*m_doc, m_node, key.data(), key.size());
return k != nullptr ? basic_json_view(m_doc, k + 1) : basic_json_view();
const node* const k = detail::view::find_member<Editable>(*m_doc, m_node, key.data(), key.size());
return k != nullptr ? basic_json_view(m_doc, navigation::value(k + 1)) : basic_json_view();
}
// --- get() dispatch ---
@@ -675,7 +823,7 @@ Borrowed parses keep a pointer to the caller's text, which must outlive the
document. Owned parses (parse_copy, rvalue std::string, streams, and inputs
that are not contiguous byte ranges) keep their own copy.
*/
template<typename BasicJsonType>
template<typename BasicJsonType, bool Editable = false>
class basic_json_document
{
using document_data = detail::view::document_data;
@@ -684,7 +832,7 @@ class basic_json_document
"json_view supports 64-bit integer types only");
public:
using view_type = basic_json_view<BasicJsonType>;
using view_type = basic_json_view<BasicJsonType, Editable>;
using value_t = detail::value_t;
/// an empty (discarded) document
@@ -788,7 +936,7 @@ class basic_json_document
/// whether the document holds its own copy of the text
bool owns_source() const noexcept
{
return m_data && !m_data->owned.empty() && m_data->src == m_data->owned.data();
return m_data && ((!m_data->owned.empty() && m_data->src == m_data->owned.data()) || !m_data->owned_image.empty());
}
/// number of index nodes (values plus object keys)
@@ -797,7 +945,7 @@ class basic_json_document
return m_data ? m_data->tape_size : 0;
}
/// bytes held by the document (index, decoded strings, owned text)
/// bytes held by the document (index, decoded strings, owned text or image)
std::size_t memory_usage() const noexcept
{
if (!m_data)
@@ -806,7 +954,10 @@ class basic_json_document
}
return sizeof(document_data) + (m_data->inline_cap * sizeof(detail::view::node))
+ (m_data->tape != m_data->inline_tape ? m_data->tape_cap * sizeof(detail::view::node) : 0)
+ m_data->arena.capacity() + m_data->owned.capacity();
+ m_data->arena.capacity() + m_data->owned.capacity() + m_data->owned_image.capacity()
+ (m_data->indexes.capacity() * sizeof(document_data::object_index)) + (m_data->index_slots.capacity() * sizeof(std::uint32_t))
+ (m_data->large_objects.capacity() * sizeof(std::uint32_t))
+ (m_data->edits != nullptr ? m_data->edits->bytes : 0);
}
/// release unused capacity of the index and the decoded strings; like
@@ -823,9 +974,12 @@ class basic_json_document
// allocate everything first, so that an exception leaves the document
// unchanged
// (the decoded strings of a loaded image stay in the image)
const bool arena_in_use = d.base[1] == d.arena.data();
const bool shrink_arena = d.arena.capacity() > d.arena.size();
std::string arena(shrink_arena ? d.arena : std::string());
const bool shrink_tape = d.tape != d.inline_tape && d.tape_size != d.tape_cap;
std::string arena(shrink_arena && arena_in_use ? d.arena : std::string());
// (edits link to the nodes of the index, which then stays in place)
const bool shrink_tape = d.tape != d.inline_tape && d.tape_size != d.tape_cap && d.edits == nullptr;
const bool into_header = d.tape_size <= d.inline_cap;
node* fresh = (shrink_tape && !into_header) ? static_cast<node*>(::operator new (d.tape_size * sizeof(node))) : d.inline_tape;
@@ -839,13 +993,219 @@ class basic_json_document
if (shrink_arena)
{
d.arena.swap(arena);
d.base[1] = d.arena.data();
if (arena_in_use)
{
d.base[1] = d.arena.data();
}
}
}
////////////
// images //
////////////
/// how load() checks an image (full, bounds, or none)
using image_check = detail::view::image_check;
/// The document as an image that load() reads without parsing: the node
/// index, the text, and the decoded strings. An edited document is
/// written in its current state (floats that are not finite become null,
/// as in dump()).
std::vector<std::uint8_t> save() const
{
if (NLOHMANN_VIEW_UNLIKELY(!m_data || m_data->discarded))
{
detail::view::throw_type_error(320, "cannot save a discarded json_document");
}
return detail::view::save_image(*m_data);
}
/// Read an image written by save(). The image is borrowed: it must stay
/// alive and unchanged while the document is used.
NLOHMANN_VIEW_NODISCARD
static basic_json_document load(const std::uint8_t* image, std::size_t size, const image_check check = image_check::full)
{
basic_json_document d;
d.ensure_data(nullptr, 0);
detail::view::load_image(*d.m_data, image, size, check);
return d;
}
/// read an image (borrowed)
NLOHMANN_VIEW_NODISCARD
static basic_json_document load(const std::vector<std::uint8_t>& image, const image_check check = image_check::full)
{
return load(image.data(), image.size(), check);
}
/// read an image and keep it (no copy)
NLOHMANN_VIEW_NODISCARD
static basic_json_document load(std::vector<std::uint8_t>&& image, const image_check check = image_check::full)
{
basic_json_document d;
d.ensure_data(nullptr, 0);
d.m_data->owned_image = std::move(image);
detail::view::load_image(*d.m_data, d.m_data->owned_image.data(), d.m_data->owned_image.size(), check);
return d;
}
///////////
// edits //
///////////
// The source text is never written; new values go to storage owned by
// the document. A view keeps referring to the same value: after an
// assignment it sees the new value, and edits elsewhere do not affect it.
// A view of an erased value keeps its last value. An edit of an
// array/object invalidates the iterators over it. Values are accepted as
// views (of any document), BasicJsonType values, and everything
// BasicJsonType can be constructed from.
using string_view_t = typename view_type::string_view_t;
using json_pointer = typename BasicJsonType::json_pointer;
/// replace a value (a view of this document); returns a view of it
template<typename V>
view_type set(view_type target, V&& value)
{
return editor().set(target, std::forward<V>(value));
}
/// set a member (added if missing; a null value becomes an object);
/// returns a view of the member value
template<typename V>
view_type set(view_type object, string_view_t key, V&& value)
{
return editor().set(object, key, std::forward<V>(value));
}
/// assign an existing array element; returns a view of it
template < typename I, typename V, typename std::enable_if < std::is_integral<I>::value && !std::is_same<I, bool>::value, int >::type = 0 >
view_type set(view_type array, I idx, V && value)
{
return editor().set(array, index(idx), std::forward<V>(value));
}
/// set the value at a JSON pointer: its parent must exist; an object
/// member is set (added if missing), an array element assigned, and "-"
/// or the size of the array appends
template<typename V>
view_type set(const json_pointer& ptr, V&& value)
{
if (ptr.empty())
{
return set(root(), std::forward<V>(value));
}
const view_type parent = root().at(ptr.parent_pointer());
const auto& token = ptr.back();
if (parent.is_array())
{
const std::size_t idx = token == "-" ? parent.size() : pointer_index(token);
if (idx == parent.size())
{
return push_back(parent, std::forward<V>(value));
}
return set(parent, idx, std::forward<V>(value));
}
return set(parent, string_view_t(token.data(), token.size()), std::forward<V>(value));
}
/// append to an array (a null value becomes an array); returns a view of
/// the new element
template<typename V>
view_type push_back(view_type array, V&& value)
{
return editor().push_back(array, std::forward<V>(value));
}
/// insert into an array before position idx (idx <= size()); returns a
/// view of the new element
template < typename I, typename V, typename std::enable_if < std::is_integral<I>::value && !std::is_same<I, bool>::value, int >::type = 0 >
view_type insert(view_type array, I idx, V && value)
{
return editor().insert(array, index(idx), std::forward<V>(value));
}
/// remove all members with this key; returns their number
std::size_t erase(view_type object, string_view_t key)
{
return editor().erase(object, key);
}
/// remove an array element
template < typename I, typename std::enable_if < std::is_integral<I>::value && !std::is_same<I, bool>::value, int >::type = 0 >
void erase(view_type array, I idx)
{
editor().erase(array, index(idx));
}
/// remove the value at a JSON pointer; returns the number of removed
/// values
std::size_t erase(const json_pointer& ptr)
{
if (ptr.empty())
{
detail::view::throw_out_of_range(405, "JSON pointer has no parent");
}
const view_type parent = root().at(ptr.parent_pointer());
const auto& token = ptr.back();
if (parent.is_array())
{
erase(parent, pointer_index(token));
return 1;
}
return erase(parent, string_view_t(token.data(), token.size()));
}
private:
using input_kind = detail::view::input_kind;
detail::view::editor<BasicJsonType, view_type> editor()
{
static_assert(Editable, "only an editable document can be edited: use basic_json_document<BasicJsonType, true> (json_editable_document)");
if (NLOHMANN_VIEW_UNLIKELY(!m_data || m_data->discarded))
{
detail::view::throw_invalid_iterator(202, "view does not belong to this document");
}
return detail::view::editor<BasicJsonType, view_type>(*m_data);
}
/// an index (out_of_range.401 if negative)
template<typename I>
static std::size_t index(I idx)
{
return index(idx, std::is_signed<I> {});
}
template<typename I>
static std::size_t index(I idx, std::true_type /*signed*/)
{
if (idx < 0)
{
detail::view::throw_out_of_range(401, detail::concat("array index ", std::to_string(idx), " is out of range"));
}
return static_cast<std::size_t>(idx);
}
template<typename I>
static std::size_t index(I idx, std::false_type /*unsigned*/)
{
return static_cast<std::size_t>(idx);
}
/// the array index of a JSON pointer token (json_pointer's rules)
template<typename StringType>
static std::size_t pointer_index(const StringType& token)
{
std::size_t idx = 0;
const detail::view::index_status status = detail::view::array_index(token, idx);
if (status != detail::view::index_status::ok)
{
detail::view::throw_array_index_error(status, token);
}
return idx;
}
/// create the storage (sized for the input) on first use
void ensure_data(const char* src, std::size_t size)
{
@@ -872,10 +1232,16 @@ class basic_json_document
{
d.owned.clear();
}
d.owned_image.clear();
d.src = src;
d.size = size;
d.tape_size = 0;
d.edits.reset(); // (views of the previous text end here anyway)
d.base[2] = nullptr;
d.arena.clear();
d.indexes.clear();
d.index_slots.clear();
d.large_objects.clear();
d.discarded = true;
detail::view::parse_failure failure;
bool ok = false;
@@ -891,6 +1257,8 @@ class basic_json_document
{
d.base[0] = d.src;
d.base[1] = d.arena.data();
d.arena_size = d.arena.size();
detail::view::build_object_indexes(d);
d.discarded = false;
return;
}
@@ -1006,6 +1374,14 @@ using json_view = basic_json_view<json>;
using ordered_json_document = basic_json_document<ordered_json>;
/// a value of an ordered_json_document
using ordered_json_view = basic_json_view<ordered_json>;
/// an editable parsed JSON text for json
using json_editable_document = basic_json_document<json, true>;
/// a value of a json_editable_document
using json_editable_view = basic_json_view<json, true>;
/// an editable parsed JSON text for ordered_json
using ordered_json_editable_document = basic_json_document<ordered_json, true>;
/// a value of an ordered_json_editable_document
using ordered_json_editable_view = basic_json_view<ordered_json, true>;
NLOHMANN_JSON_NAMESPACE_END
File diff suppressed because it is too large Load Diff
+20
View File
@@ -324,6 +324,26 @@ json_test_add_test_for(src/unit-diagnostic-positions.cpp
MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force}
)
# the json_view parser again with the portable string scanning instead of
# NEON/SSE2, and on x86-64 with the SSSE3 UTF-8 check (JSON_VIEW_USE_SSSE3)
json_test_set_test_options(test-json_view_builder_portable
COMPILE_DEFINITIONS JSON_VIEW_NO_SIMD
)
json_test_add_test_for(src/unit-json_view_builder.cpp
NAME test-json_view_builder_portable
MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force}
)
if(CMAKE_SYSTEM_PROCESSOR MATCHES "^(x86_64|AMD64|amd64)$" AND NOT MSVC)
json_test_set_test_options(test-json_view_builder_ssse3
COMPILE_DEFINITIONS JSON_VIEW_USE_SSSE3
COMPILE_OPTIONS -mssse3
)
json_test_add_test_for(src/unit-json_view_builder.cpp
NAME test-json_view_builder_ssse3
MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force}
)
endif()
# *DO NOT* use json_test_set_test_options() below this line
#############################################################################
+4 -1
View File
@@ -10,7 +10,7 @@ CXXFLAGS += -std=c++11
CPPFLAGS += -I ../single_include
FUZZER_ENGINE = src/fuzzer-driver_afl.cpp
FUZZERS = parse_afl_fuzzer parse_bson_fuzzer parse_cbor_fuzzer parse_msgpack_fuzzer parse_ubjson_fuzzer parse_bjdata_fuzzer parse_bon8_fuzzer parse_json_view_fuzzer
FUZZERS = parse_afl_fuzzer parse_bson_fuzzer parse_cbor_fuzzer parse_msgpack_fuzzer parse_ubjson_fuzzer parse_bjdata_fuzzer parse_bon8_fuzzer parse_json_view_fuzzer json_view_image_fuzzer
fuzzers: $(FUZZERS)
parse_afl_fuzzer:
@@ -19,6 +19,9 @@ parse_afl_fuzzer:
parse_json_view_fuzzer:
$(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-parse_json_view.cpp -o $@
json_view_image_fuzzer:
$(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-json_view_image.cpp -o $@
parse_bson_fuzzer:
$(CXX) $(CXXFLAGS) $(CPPFLAGS) $(FUZZER_ENGINE) src/fuzzer-parse_bson.cpp -o $@
+1
View File
@@ -21,6 +21,7 @@ Micro-benchmarks for parsing, serialization and the binary formats, written with
| `ViewParseIndented` | as `ParseIndented`, with a reused `json_document` |
| `ViewAccept` | validate with `json_document::accept`; compare with `Accept` |
| `ViewMaterialize` | convert a parsed `json_document` into a `json` value |
| `ViewDump` | serialize a parsed `json_document`; compare with `Dump` |
The input files are those of [nativejson-benchmark](https://github.com/miloyip/nativejson-benchmark) (`canada`,
`citm_catalog`, `twitter`), a large `jeopardy` file, and number-heavy files (`floats`, `signed_ints`, ...).
+1
View File
@@ -0,0 +1 @@
build/
+79
View File
@@ -0,0 +1,79 @@
# json_view compared with other libraries
The in-tree benchmarks in [`tests/benchmarks`](../README.md) measure `json_document` against `json::parse` only. The
programs here compare it with [yyjson](https://github.com/ibireme/yyjson),
[simdjson](https://github.com/simdjson/simdjson), and [Boost.JSON](https://github.com/boostorg/json): the question
users ask when they pick a library. They are not built by CMake or run by CI.
## Reproducing the numbers
`compare.py` builds the programs against `include/` of this checkout, runs them, and writes the results together with
everything needed to reproduce them to `results/<date>-<host>.md` (and `.csv`): the date, the commit, the CPU, the
OS, the compiler, the flags, and the versions of all libraries.
```sh
python3 tests/benchmarks/json_view/compare.py --data <json_test_data directory> [--native] [--rounds 30]
```
- `--data` is the downloaded [test data](https://github.com/nlohmann/json_test_data), e.g. the `test_files` directory
of a CMake build directory. It needs `nativejson-benchmark/{twitter,citm_catalog,canada}.json` and
`jeopardy/jeopardy.json`.
- The other libraries come from the system: pkg-config, or Homebrew (`brew install yyjson simdjson boost`). With
`--download`, pinned releases are downloaded instead and checked against their SHA-256. Without Boost headers (or
with `--no-boost`), the Boost.JSON columns are skipped, and the results say so.
- `--corpus file...` adds files to the corpus benchmark, e.g. those of
[simdjson-data](https://github.com/simdjson/simdjson-data) or the
[yyjson benchmark](https://github.com/ibireme/yyjson_benchmark).
- Only the Python 3 standard library is used; a C++17 compiler is needed (`CXX` and `CC` are honored).
For numbers worth publishing, use a quiet machine (see [Getting stable numbers](../README.md#getting-stable-numbers)),
the default 30 rounds or more, and `--native` only if the other libraries were built for the same CPU.
### On GitHub-hosted runners
The workflow [json_view benchmarks](../../../.github/workflows/json_view_benchmarks.yml) runs `compare.py --download`
on demand: by hand (Actions → "json_view benchmarks" → "Run workflow"), on an x86-64 or AArch64 Ubuntu runner with GCC
or Clang, or when a pull request gets the label `benchmark`, on both architectures with GCC. The results appear as the
job summary and as an artifact. Shared runners are noisy, so these numbers show
where `json_view` stands on another architecture; they are not meant for publication.
## What is measured
`bench_view.cpp` runs four workloads on twitter, citm_catalog, canada, jeopardy, a single tweet (`status`), and a
JSON-RPC request (`rpc`):
| workload | what it does |
|---|---|
| parse | build and free a document |
| traverse | parse, then visit every value, convert every number, touch every string and key |
| select | parse, then read a few fields per record (e.g. id, user name, and retweet count of each tweet) |
| dump | serialize a parsed document (compact) |
`bench_corpus.cpp` runs parse, traverse, and dump on any list of files, so that no library is tuned to a handful of
documents.
`bench_edit.cpp` measures read-modify-write: parse, apply the same logical edits with each library's own API, and
serialize (compact). Workloads: `patch` (a handful of edits at fixed places) and `update` (edits in every record).
An editable `json_document` edits in place; yyjson copies its immutable document into a mutable one first
(`yyjson_doc_mut_copy`); Boost.JSON and `json::parse` build mutable DOMs; simdjson cannot edit a document. All
outputs are checked to describe the same value.
Before anything is timed, all engines must accept each document and agree on the traversal: the number of values, the
bytes of all strings and keys, and the sum of all numbers. All engines run interleaved in every round, and the best
round is reported, as time and as a factor of the `json_view` time (below 1 means faster than `json_view`).
The engines do not all offer the same features, which the numbers should be read with:
| engine | document | random access | editable | notes |
|---|---|---|---|---|
| `json_view` | immutable index into the text | yes | no | a fresh document per parse; "reused" parses into the same document |
| yyjson | immutable (`yyjson_read`) | yes | via a mutable copy | |
| simdjson DOM | immutable, parser reused | yes | no | |
| simdjson On-Demand | none: forward-only, lazy | no | no | only traverse and select |
| Boost.JSON | owning, mutable DOM | yes | yes | monotonic resource |
| `json::parse` | owning, mutable DOM | yes | yes | |
## Published results
Results are only published with the file `compare.py` wrote, which names the machine and the versions; see
`results/`. Numbers from one machine and compiler do not carry over to another: rerun the script.
+326
View File
@@ -0,0 +1,326 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Corpus benchmark: the read-only workloads of bench_view.cpp on any list of
// JSON files (for example the benchmark sets of simdjson and yyjson).
//
// ./bench_corpus [--rounds N] file...
//
// For every file, all engines must accept it and agree on a traversal (value
// count, string bytes, sum of numbers) before anything is timed. Workloads:
// parse (build and free a document), traverse (visit every value, convert
// every number), dump (compact), and for json_view also dump with the source
// number text. Results go to bench_corpus.csv.
#include <nlohmann/json_view.hpp>
#if JSON_VIEW_BENCH_BOOST
#include <boost/json.hpp>
#include <boost/json/src.hpp>
#endif
#include <simdjson.h>
#include <yyjson.h>
#include <algorithm>
#include <chrono>
#include <cmath>
#include <cstdio>
#include <cstring>
#include <fstream>
#include <functional>
#include <sstream>
#include <string>
#include <vector>
using nlohmann::json;
using nlohmann::json_document;
using nlohmann::json_view;
static volatile double g_sink;
struct stats
{
double num = 0;
std::size_t str = 0, nodes = 0;
};
static void walk(json_view v, stats& st)
{
++st.nodes;
switch (v.type())
{
case json::value_t::object:
for (auto it = v.begin(); it != v.end(); ++it)
{
st.str += it.key().size();
walk(*it, st);
}
break;
case json::value_t::array:
for (const json_view e : v)
{
walk(e, st);
}
break;
case json::value_t::string:
st.str += v.get_string().size();
break;
case json::value_t::number_integer:
st.num += static_cast<double>(v.get<std::int64_t>());
break;
case json::value_t::number_unsigned:
st.num += static_cast<double>(v.get<std::uint64_t>());
break;
case json::value_t::number_float:
st.num += v.get<double>();
break;
default:
break;
}
}
static void walk(yyjson_val* v, stats& st)
{
++st.nodes;
switch (yyjson_get_type(v))
{
case YYJSON_TYPE_OBJ:
{
std::size_t idx, max;
yyjson_val* k, * val;
yyjson_obj_foreach(v, idx, max, k, val)
{
st.str += yyjson_get_len(k);
walk(val, st);
}
break;
}
case YYJSON_TYPE_ARR:
{
std::size_t idx, max;
yyjson_val* val;
yyjson_arr_foreach(v, idx, max, val)
{
walk(val, st);
}
break;
}
case YYJSON_TYPE_STR:
st.str += yyjson_get_len(v);
break;
case YYJSON_TYPE_NUM:
st.num += yyjson_is_sint(v) ? static_cast<double>(yyjson_get_sint(v)) : yyjson_is_uint(v) ? static_cast<double>(yyjson_get_uint(v)) : yyjson_get_real(v);
break;
default:
break;
}
}
static void walk(simdjson::dom::element e, stats& st)
{
++st.nodes;
switch (e.type())
{
case simdjson::dom::element_type::OBJECT:
for (auto f : simdjson::dom::object(e))
{
st.str += f.key.size();
walk(f.value, st);
}
break;
case simdjson::dom::element_type::ARRAY:
for (auto c : simdjson::dom::array(e))
{
walk(c, st);
}
break;
case simdjson::dom::element_type::STRING:
st.str += std::string_view(e).size();
break;
case simdjson::dom::element_type::INT64:
st.num += static_cast<double>(int64_t(e));
break;
case simdjson::dom::element_type::UINT64:
st.num += static_cast<double>(uint64_t(e));
break;
case simdjson::dom::element_type::DOUBLE:
st.num += double(e);
break;
default:
break;
}
}
#if JSON_VIEW_BENCH_BOOST
static void walk(const boost::json::value& v, stats& st)
{
++st.nodes;
switch (v.kind())
{
case boost::json::kind::object:
for (const auto& kv : v.get_object())
{
st.str += kv.key().size();
walk(kv.value(), st);
}
break;
case boost::json::kind::array:
for (const auto& c : v.get_array())
{
walk(c, st);
}
break;
case boost::json::kind::string:
st.str += v.get_string().size();
break;
case boost::json::kind::int64:
st.num += static_cast<double>(v.get_int64());
break;
case boost::json::kind::uint64:
st.num += static_cast<double>(v.get_uint64());
break;
case boost::json::kind::double_:
st.num += v.get_double();
break;
default:
break;
}
}
#endif
static std::string slurp(const std::string& p)
{
std::ifstream f(p, std::ios::binary);
std::stringstream ss;
ss << f.rdbuf();
return ss.str();
}
static bool same(const stats& a, const stats& b)
{
return a.nodes == b.nodes && a.str == b.str && (a.num == b.num || std::fabs(a.num - b.num) <= 1e-9 * std::fabs(a.num));
}
int main(int argc, char** argv)
{
int rounds = 0; // 0: by size
std::vector<std::string> files;
for (int i = 1; i < argc; ++i)
{
if (std::strcmp(argv[i], "--rounds") == 0 && i + 1 < argc)
{
rounds = std::atoi(argv[++i]);
}
else
{
files.push_back(argv[i]);
}
}
std::FILE* csv = std::fopen("bench_corpus.csv", "w");
std::fprintf(csv, "file,bytes,workload,engine,ns\n");
simdjson::dom::parser sj;
for (const auto& path : files)
{
const std::string s = slurp(path);
const std::string name = path.substr(path.rfind('/') + 1);
const simdjson::padded_string ps(s);
// all engines must agree before timing
stats a, b, c, d;
const json_document doc = json_document::parse(s);
walk(doc.root(), a);
yyjson_doc* y = yyjson_read(s.data(), s.size(), 0);
auto sjr = sj.parse(ps);
#if JSON_VIEW_BENCH_BOOST
boost::json::parse_options opt;
opt.numbers = boost::json::number_precision::precise;
boost::json::monotonic_resource mr0;
const boost::json::value bv = boost::json::parse(s, &mr0, opt);
#endif
if (y == nullptr || sjr.error())
{
std::printf("%-34s skipped (an engine rejects it)\n", name.c_str());
yyjson_doc_free(y);
continue;
}
walk(yyjson_doc_get_root(y), b);
walk(sjr.value_unsafe(), c);
#if JSON_VIEW_BENCH_BOOST
walk(bv, d);
#else
d = a;
#endif
yyjson_doc_free(y);
const bool ok = same(a, b) && same(a, c) && same(a, d);
const int r = rounds > 0 ? rounds : static_cast<int>(std::max<std::size_t>(3, std::min<std::size_t>(60, 400000000 / (s.size() + 1))));
struct engine
{
std::string name;
std::function<void()> fn;
};
json_document vd = json_document::parse(s);
yyjson_doc* yd = yyjson_read(s.data(), s.size(), 0);
simdjson::dom::parser sjd;
const simdjson::dom::element se = sjd.parse(ps).value_unsafe();
const std::vector<std::pair<std::string, std::vector<engine>>> workloads =
{
{
"parse", {
{"json_view", [&] { auto x = json_document::parse(s); g_sink = static_cast<double>(x.node_count()); }},
{"yyjson", [&] { yyjson_doc* x = yyjson_read(s.data(), s.size(), 0); g_sink = static_cast<double>(yyjson_doc_get_val_count(x)); yyjson_doc_free(x); }},
{"simdjson DOM", [&] { auto e = sj.parse(ps).value_unsafe(); g_sink = e.is_object(); }},
#if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); g_sink = v.is_object(); }},
#endif
}
},
{
"traverse", {
{"json_view", [&] { auto x = json_document::parse(s); stats st; walk(x.root(), st); g_sink = st.num; }},
{"yyjson", [&] { yyjson_doc* x = yyjson_read(s.data(), s.size(), 0); stats st; walk(yyjson_doc_get_root(x), st); g_sink = st.num; yyjson_doc_free(x); }},
{"simdjson DOM", [&] { stats st; walk(sj.parse(ps).value_unsafe(), st); g_sink = st.num; }},
#if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); stats st; walk(v, st); g_sink = st.num; }},
#endif
}
},
{
"dump", {
{"json_view", [&] { std::string o = vd.root().dump(); g_sink = static_cast<double>(o.size()); }},
{"yyjson", [&] { std::size_t n = 0; char* o = yyjson_write(yd, 0, &n); g_sink = static_cast<double>(n); std::free(o); }},
{"simdjson DOM", [&] { std::string o = simdjson::to_string(se); g_sink = static_cast<double>(o.size()); }},
{"json_view (source numbers)", [&] { std::string o = vd.root().dump(-1, ' ', false, json_view::number_format::source); g_sink = static_cast<double>(o.size()); }},
}
},
};
std::printf("%-34s %9zu B%s\n", name.c_str(), s.size(), ok ? "" : " [ENGINES DISAGREE]");
for (const auto& wl : workloads)
{
std::vector<double> best(wl.second.size(), 1e300);
for (int i = 0; i < r; ++i)
{
for (std::size_t k = 0; k < wl.second.size(); ++k)
{
const auto t0 = std::chrono::steady_clock::now();
wl.second[k].fn();
best[k] = std::min(best[k], std::chrono::duration<double, std::nano>(std::chrono::steady_clock::now() - t0).count());
}
}
std::printf(" %-9s", wl.first.c_str());
for (std::size_t k = 0; k < wl.second.size(); ++k)
{
std::printf(" %s %.2f GB/s (%.2fx)", wl.second[k].name.c_str(), static_cast<double>(s.size()) / best[k], best[k] / best[0]);
std::fprintf(csv, "%s,%zu,%s,%s,%.1f\n", name.c_str(), s.size(), wl.first.c_str(), wl.second[k].name.c_str(), best[k]);
}
std::printf("\n");
std::fflush(stdout);
}
yyjson_doc_free(yd);
}
std::fclose(csv);
}
+575
View File
@@ -0,0 +1,575 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Read-modify-write benchmark: parse a document, apply the same logical edits
// with each library's own API, and serialize it (compact).
//
// json_view json_editable_document: edits in place, unchanged values stay in the index
// yyjson yyjson_read + yyjson_doc_mut_copy (the way to edit a parsed document)
// Boost.JSON parse into a mutable DOM (monotonic resource, precise numbers), serialize
// json::parse nlohmann::json today
// simdjson has no mutable document and is not part of this comparison.
//
// Workloads:
// patch a handful of edits at fixed places (scalars, a new member, a new array element)
// update edits in every record (twitter: 100 statuses, citm: 243 performances,
// canada: 480 rings, jeopardy: 216,930 questions): set scalars, erase a
// member, add a member (canada: replace the first point of every ring)
//
// Build: see README.md (same flags as bench_view.cpp).
#include <nlohmann/json_view.hpp>
#if JSON_VIEW_BENCH_BOOST
#include <boost/json.hpp>
#include <boost/json/src.hpp>
#endif
#include <yyjson.h>
#include <algorithm>
#include <chrono>
#include <cstdio>
#include <cstdlib>
#include <fstream>
#include <functional>
#include <sstream>
using nlohmann::json;
using nlohmann::json_editable_document;
using nlohmann::json_editable_view;
#if JSON_VIEW_BENCH_BOOST
namespace bj = boost::json;
#endif
static volatile std::size_t g_sink;
// ---------------- json_view ----------------
static std::string edit_view(const std::string& name, const std::string& s, bool update)
{
json_editable_document d = json_editable_document::parse(s);
const json_editable_view r = d.root();
if (name == "twitter")
{
if (update)
{
std::int64_t i = 0;
for (const json_editable_view st : r["statuses"])
{
d.set(st, "retweet_count", i++);
d.set(st, "favorited", true);
d.set(st, "text", "redacted");
d.erase(st, "entities");
d.set(st, "edited", true);
}
}
else
{
d.set(r["search_metadata"], "count", 200);
d.set(r["statuses"][0], "text", "patched");
d.set(r["statuses"][0]["user"], "followers_count", 1);
d.set(r["statuses"][99], "favorited", true);
d.set(r, "patched", true);
}
}
else if (name == "citm_catalog")
{
if (update)
{
for (const json_editable_view p : r["performances"])
{
d.set(p, "name", "performance");
d.set(p, "start", p["start"].get<std::int64_t>() + 1);
d.erase(p, "seatMapImage");
d.set(p, "edited", true);
}
}
else
{
d.set(r["events"]["138586341"], "name", "patched");
d.set(r["performances"][0], "start", 0);
d.set(r["venueNames"], "PLEYEL_PLEYEL", "Salle");
d.set(r, "patched", true);
}
}
else if (name == "canada")
{
const json_editable_view coords = r["features"][0]["geometry"]["coordinates"];
if (update)
{
for (const json_editable_view ring : coords)
{
d.set(ring, 0, json::array({0.5, 0.5}));
}
}
else
{
d.set(r["features"][0]["properties"], "name", "patched");
d.set(r, "type", "FeatureCollection2");
d.set(coords[0], 0, json::array({0.0, 0.0}));
}
}
else if (name == "jeopardy")
{
if (update)
{
for (const json_editable_view q : r)
{
d.set(q, "value", "$1");
d.erase(q, "air_date");
}
}
else
{
d.set(r[0], "value", "$0");
d.set(r[100000], "answer", "patched");
d.set(r[216929], "round", "x");
d.push_back(r, json::object({{"category", "NEW"}, {"value", "$5"}}));
}
}
else if (name == "status")
{
d.set(r, "retweet_count", 1);
d.set(r["user"], "name", "x");
}
else if (name == "rpc")
{
d.set(r, "id", 4);
d.set(r["params"], "subtrahend", 24);
}
return r.dump();
}
// ---------------- nlohmann::json ----------------
static std::string edit_json(const std::string& name, const std::string& s, bool update)
{
json r = json::parse(s);
if (name == "twitter")
{
if (update)
{
std::int64_t i = 0;
for (auto& st : r["statuses"])
{
st["retweet_count"] = i++;
st["favorited"] = true;
st["text"] = "redacted";
st.erase("entities");
st["edited"] = true;
}
}
else
{
r["search_metadata"]["count"] = 200;
r["statuses"][0]["text"] = "patched";
r["statuses"][0]["user"]["followers_count"] = 1;
r["statuses"][99]["favorited"] = true;
r["patched"] = true;
}
}
else if (name == "citm_catalog")
{
if (update)
{
for (auto& p : r["performances"])
{
p["name"] = "performance";
p["start"] = p["start"].get<std::int64_t>() + 1;
p.erase("seatMapImage");
p["edited"] = true;
}
}
else
{
r["events"]["138586341"]["name"] = "patched";
r["performances"][0]["start"] = 0;
r["venueNames"]["PLEYEL_PLEYEL"] = "Salle";
r["patched"] = true;
}
}
else if (name == "canada")
{
json& coords = r["features"][0]["geometry"]["coordinates"];
if (update)
{
for (auto& ring : coords)
{
ring[0] = json::array({0.5, 0.5});
}
}
else
{
r["features"][0]["properties"]["name"] = "patched";
r["type"] = "FeatureCollection2";
coords[0][0] = json::array({0.0, 0.0});
}
}
else if (name == "jeopardy")
{
if (update)
{
for (auto& q : r)
{
q["value"] = "$1";
q.erase("air_date");
}
}
else
{
r[0]["value"] = "$0";
r[100000]["answer"] = "patched";
r[216929]["round"] = "x";
r.push_back(json::object({{"category", "NEW"}, {"value", "$5"}}));
}
}
else if (name == "status")
{
r["retweet_count"] = 1;
r["user"]["name"] = "x";
}
else if (name == "rpc")
{
r["id"] = 4;
r["params"]["subtrahend"] = 24;
}
return r.dump();
}
// ---------------- yyjson ----------------
static std::string edit_yyjson(const std::string& name, const std::string& s, bool update)
{
yyjson_doc* idoc = yyjson_read(s.data(), s.size(), 0);
yyjson_mut_doc* d = yyjson_doc_mut_copy(idoc, nullptr);
yyjson_doc_free(idoc);
yyjson_mut_val* r = yyjson_mut_doc_get_root(d);
auto get = [](yyjson_mut_val * o, const char* k)
{
return yyjson_mut_obj_get(o, k);
};
if (name == "twitter")
{
yyjson_mut_val* sts = get(r, "statuses");
if (update)
{
std::size_t idx, max;
yyjson_mut_val* st;
std::int64_t i = 0;
yyjson_mut_arr_foreach(sts, idx, max, st)
{
yyjson_mut_set_sint(get(st, "retweet_count"), i++);
yyjson_mut_set_bool(get(st, "favorited"), true);
yyjson_mut_set_str(get(st, "text"), "redacted");
yyjson_mut_obj_remove_key(st, "entities");
yyjson_mut_obj_add_bool(d, st, "edited", true);
}
}
else
{
yyjson_mut_set_sint(get(get(r, "search_metadata"), "count"), 200);
yyjson_mut_val* s0 = yyjson_mut_arr_get(sts, 0);
yyjson_mut_set_str(get(s0, "text"), "patched");
yyjson_mut_set_sint(get(get(s0, "user"), "followers_count"), 1);
yyjson_mut_set_bool(get(yyjson_mut_arr_get(sts, 99), "favorited"), true);
yyjson_mut_obj_add_bool(d, r, "patched", true);
}
}
else if (name == "citm_catalog")
{
if (update)
{
std::size_t idx, max;
yyjson_mut_val* p;
yyjson_mut_arr_foreach(get(r, "performances"), idx, max, p)
{
yyjson_mut_set_str(get(p, "name"), "performance");
yyjson_mut_val* start = get(p, "start");
yyjson_mut_set_sint(start, yyjson_mut_get_sint(start) + 1);
yyjson_mut_obj_remove_key(p, "seatMapImage");
yyjson_mut_obj_add_bool(d, p, "edited", true);
}
}
else
{
yyjson_mut_set_str(get(get(get(r, "events"), "138586341"), "name"), "patched");
yyjson_mut_set_sint(get(yyjson_mut_arr_get(get(r, "performances"), 0), "start"), 0);
yyjson_mut_set_str(get(get(r, "venueNames"), "PLEYEL_PLEYEL"), "Salle");
yyjson_mut_obj_add_bool(d, r, "patched", true);
}
}
else if (name == "canada")
{
yyjson_mut_val* f0 = yyjson_mut_arr_get(get(r, "features"), 0);
yyjson_mut_val* coords = get(get(f0, "geometry"), "coordinates");
if (update)
{
static const double half[2] = {0.5, 0.5};
std::size_t idx, max;
yyjson_mut_val* ring;
yyjson_mut_arr_foreach(coords, idx, max, ring)
{
yyjson_mut_arr_replace(ring, 0, yyjson_mut_arr_with_real(d, half, 2));
}
}
else
{
static const double zero[2] = {0.0, 0.0};
yyjson_mut_set_str(get(get(f0, "properties"), "name"), "patched");
yyjson_mut_set_str(get(r, "type"), "FeatureCollection2");
yyjson_mut_arr_replace(yyjson_mut_arr_get(coords, 0), 0, yyjson_mut_arr_with_real(d, zero, 2));
}
}
else if (name == "jeopardy")
{
if (update)
{
std::size_t idx, max;
yyjson_mut_val* q;
yyjson_mut_arr_foreach(r, idx, max, q)
{
yyjson_mut_set_str(get(q, "value"), "$1");
yyjson_mut_obj_remove_key(q, "air_date");
}
}
else
{
yyjson_mut_set_str(get(yyjson_mut_arr_get(r, 0), "value"), "$0");
yyjson_mut_set_str(get(yyjson_mut_arr_get(r, 100000), "answer"), "patched");
yyjson_mut_set_str(get(yyjson_mut_arr_get(r, 216929), "round"), "x");
yyjson_mut_val* o = yyjson_mut_obj(d);
yyjson_mut_obj_add_str(d, o, "category", "NEW");
yyjson_mut_obj_add_str(d, o, "value", "$5");
yyjson_mut_arr_append(r, o);
}
}
else if (name == "status")
{
yyjson_mut_set_sint(get(r, "retweet_count"), 1);
yyjson_mut_set_str(get(get(r, "user"), "name"), "x");
}
else if (name == "rpc")
{
yyjson_mut_set_sint(get(r, "id"), 4);
yyjson_mut_set_sint(get(get(r, "params"), "subtrahend"), 24);
}
std::size_t n = 0;
char* out = yyjson_mut_write(d, 0, &n);
std::string result(out, n);
std::free(out);
yyjson_mut_doc_free(d);
return result;
}
#if JSON_VIEW_BENCH_BOOST
// ---------------- Boost.JSON ----------------
static std::string edit_boost(const std::string& name, const std::string& s, bool update)
{
bj::monotonic_resource mr;
bj::parse_options opt;
opt.numbers = bj::number_precision::precise; // correctly rounded, like the others
bj::value v = bj::parse(s, &mr, opt);
bj::object* const obj = v.if_object(); // nullptr for jeopardy (an array)
if (name == "twitter")
{
bj::array& sts = (*obj)["statuses"].as_array();
if (update)
{
std::int64_t i = 0;
for (auto& e : sts)
{
bj::object& st = e.as_object();
st["retweet_count"] = i++;
st["favorited"] = true;
st["text"] = "redacted";
st.erase("entities");
st["edited"] = true;
}
}
else
{
(*obj)["search_metadata"].as_object()["count"] = 200;
bj::object& s0 = sts[0].as_object();
s0["text"] = "patched";
s0["user"].as_object()["followers_count"] = 1;
sts[99].as_object()["favorited"] = true;
(*obj)["patched"] = true;
}
}
else if (name == "citm_catalog")
{
if (update)
{
for (auto& e : (*obj)["performances"].as_array())
{
bj::object& p = e.as_object();
p["name"] = "performance";
p["start"] = p["start"].as_int64() + 1;
p.erase("seatMapImage");
p["edited"] = true;
}
}
else
{
(*obj)["events"].as_object()["138586341"].as_object()["name"] = "patched";
(*obj)["performances"].as_array()[0].as_object()["start"] = 0;
(*obj)["venueNames"].as_object()["PLEYEL_PLEYEL"] = "Salle";
(*obj)["patched"] = true;
}
}
else if (name == "canada")
{
bj::object& f0 = (*obj)["features"].as_array()[0].as_object();
bj::array& coords = f0["geometry"].as_object()["coordinates"].as_array();
if (update)
{
for (auto& ring : coords)
{
ring.as_array()[0] = bj::array({0.5, 0.5});
}
}
else
{
f0["properties"].as_object()["name"] = "patched";
(*obj)["type"] = "FeatureCollection2";
coords[0].as_array()[0] = bj::array({0.0, 0.0});
}
}
else if (name == "jeopardy")
{
bj::array& a = v.as_array();
if (update)
{
for (auto& e : a)
{
bj::object& q = e.as_object();
q["value"] = "$1";
q.erase("air_date");
}
}
else
{
a[0].as_object()["value"] = "$0";
a[100000].as_object()["answer"] = "patched";
a[216929].as_object()["round"] = "x";
a.push_back(bj::object({{"category", "NEW"}, {"value", "$5"}}));
}
}
else if (name == "status")
{
(*obj)["retweet_count"] = 1;
(*obj)["user"].as_object()["name"] = "x";
}
else if (name == "rpc")
{
(*obj)["id"] = 4;
(*obj)["params"].as_object()["subtrahend"] = 24;
}
return bj::serialize(v);
}
#endif
// ---------------- harness ----------------
static std::string slurp(const std::string& p)
{
std::ifstream f(p, std::ios::binary);
std::stringstream ss;
ss << f.rdbuf();
return ss.str();
}
int main(int argc, char** argv)
{
if (argc < 2)
{
std::fprintf(stderr, "usage: %s <json_test_data directory> [rounds] [document]\n", argv[0]);
return 1;
}
const std::string T = std::string(argv[1]) + "/";
const int rounds = argc > 2 ? std::atoi(argv[2]) : 20;
const std::string only = argc > 3 ? argv[3] : "";
struct doc
{
std::string name, text;
int batch;
};
std::vector<doc> docs;
for (const char* f :
{"nativejson-benchmark/twitter.json", "nativejson-benchmark/citm_catalog.json", "nativejson-benchmark/canada.json", "jeopardy/jeopardy.json"
})
{
std::string n = std::string(f).substr(std::string(f).find('/') + 1);
docs.push_back({n.substr(0, n.size() - 5), slurp(T + f), 1});
}
docs.push_back({"status", json::parse(docs[0].text)["statuses"][0].dump(), 200});
docs.push_back({"rpc", R"({"jsonrpc": "2.0", "method": "subtract", "params": {"minuend": 42, "subtrahend": 23}, "id": 3})", 5000});
using fn = std::string (*)(const std::string&, const std::string&, bool);
const std::vector<std::pair<std::string, fn>> engines =
{
{"json_view", edit_view}, {"yyjson", edit_yyjson},
#if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", edit_boost},
#endif
{"json::parse", edit_json}
};
std::FILE* csv = std::fopen("bench_edit.csv", "w");
std::fprintf(csv, "doc,bytes,workload,engine,ns\n");
for (const auto& dc : docs)
{
if (!only.empty() && dc.name != only)
{
continue;
}
for (const bool update :
{
false, true
})
{
if (update && (dc.name == "status" || dc.name == "rpc"))
{
continue;
}
// all engines must produce the same value
const json expected = json::parse(edit_json(dc.name, dc.text, update));
bool ok = true;
for (const auto& e : engines)
{
ok = ok && json::parse(e.second(dc.name, dc.text, update)) == expected;
}
std::vector<double> best(engines.size(), 1e300);
const int r = dc.text.size() > 10000000 ? std::max(3, rounds / 4) : rounds;
for (int i = 0; i < r; ++i)
{
for (std::size_t k = 0; k < engines.size(); ++k)
{
const auto t0 = std::chrono::steady_clock::now();
for (int b = 0; b < dc.batch; ++b)
{
g_sink = engines[k].second(dc.name, dc.text, update).size();
}
const double ns = std::chrono::duration<double, std::nano>(std::chrono::steady_clock::now() - t0).count() / dc.batch;
best[k] = std::min(best[k], ns);
}
}
const char* wl = update ? "update" : "patch";
std::printf("%-13s %-7s %s", dc.name.c_str(), wl, ok ? "" : "[OUTPUT MISMATCH] ");
for (std::size_t k = 0; k < engines.size(); ++k)
{
const double us = best[k] / 1e3;
std::printf(" %s %.*fus (%.2fx)", engines[k].first.c_str(), us < 10 ? 3 : (us < 1000 ? 1 : 0), us, best[k] / best[0]);
std::fprintf(csv, "%s,%zu,%s,%s,%.1f\n", dc.name.c_str(), dc.text.size(), wl, engines[k].first.c_str(), best[k]);
}
std::printf("\n");
std::fflush(stdout);
}
}
std::fclose(csv);
}
+735
View File
@@ -0,0 +1,735 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// Same-feature-set benchmark: read-only JSON documents with random access.
//
// json_view nlohmann/json_view.hpp (fresh document per parse / reused)
// yyjson yyjson_read(): immutable document, random access
// simdjson DOM dom::parser (reused, as recommended): immutable, random access
// references (different feature sets):
// simdjson OD On-Demand: forward-only, lazy
// Boost.JSON owning, mutable DOM (monotonic resource)
// json::parse owning, mutable DOM (nlohmann today)
//
// Workloads: parse (build + free), traverse (visit everything, convert every
// number, touch every string and key), select (a few fields per document),
// dump (compact serialization of the parsed document).
// All engines run interleaved in every round; the best round is reported.
#include <nlohmann/json_view.hpp>
#if JSON_VIEW_BENCH_BOOST
#include <boost/json.hpp>
#include <boost/json/src.hpp>
#endif
#include <simdjson.h>
#include <yyjson.h>
#include <algorithm>
#include <chrono>
#include <cmath>
#include <cstdio>
#include <fstream>
#include <functional>
#include <map>
#include <sstream>
using nlohmann::json;
using nlohmann::json_document;
using nlohmann::json_view;
static volatile double g_sink;
struct stats
{
double num = 0;
std::size_t str = 0, nodes = 0;
};
// ---------------- traversal ----------------
static void walk(json_view v, stats& st)
{
++st.nodes;
switch (v.type())
{
case json::value_t::object:
for (auto it = v.begin(); it != v.end(); ++it)
{
st.str += it.key().size();
walk(*it, st);
}
break;
case json::value_t::array:
for (const json_view e : v)
{
walk(e, st);
}
break;
case json::value_t::string:
st.str += v.get_string().size();
break;
case json::value_t::number_integer:
st.num += static_cast<double>(v.get<std::int64_t>());
break;
case json::value_t::number_unsigned:
st.num += static_cast<double>(v.get<std::uint64_t>());
break;
case json::value_t::number_float:
st.num += v.get<double>();
break;
default:
break;
}
}
static void walk(const json& j, stats& st)
{
++st.nodes;
switch (j.type())
{
case json::value_t::object:
for (const auto& kv : j.get_ref<const json::object_t&>())
{
st.str += kv.first.size();
walk(kv.second, st);
}
break;
case json::value_t::array:
for (const auto& e : j.get_ref<const json::array_t&>())
{
walk(e, st);
}
break;
case json::value_t::string:
st.str += j.get_ref<const std::string&>().size();
break;
case json::value_t::number_integer:
st.num += static_cast<double>(*j.get_ptr<const json::number_integer_t*>());
break;
case json::value_t::number_unsigned:
st.num += static_cast<double>(*j.get_ptr<const json::number_unsigned_t*>());
break;
case json::value_t::number_float:
st.num += *j.get_ptr<const json::number_float_t*>();
break;
default:
break;
}
}
static void walk(yyjson_val* v, stats& st)
{
++st.nodes;
switch (yyjson_get_type(v))
{
case YYJSON_TYPE_OBJ:
{
std::size_t idx, max;
yyjson_val* k, * val;
yyjson_obj_foreach(v, idx, max, k, val)
{
st.str += yyjson_get_len(k);
walk(val, st);
}
break;
}
case YYJSON_TYPE_ARR:
{
std::size_t idx, max;
yyjson_val* val;
yyjson_arr_foreach(v, idx, max, val)
{
walk(val, st);
}
break;
}
case YYJSON_TYPE_STR:
st.str += yyjson_get_len(v);
break;
case YYJSON_TYPE_NUM:
if (yyjson_is_sint(v))
{
st.num += static_cast<double>(yyjson_get_sint(v));
}
else if (yyjson_is_uint(v))
{
st.num += static_cast<double>(yyjson_get_uint(v));
}
else
{
st.num += yyjson_get_real(v);
}
break;
default:
break;
}
}
static void walk(simdjson::dom::element e, stats& st)
{
++st.nodes;
switch (e.type())
{
case simdjson::dom::element_type::OBJECT:
for (auto f : simdjson::dom::object(e))
{
st.str += f.key.size();
walk(f.value, st);
}
break;
case simdjson::dom::element_type::ARRAY:
for (auto c : simdjson::dom::array(e))
{
walk(c, st);
}
break;
case simdjson::dom::element_type::STRING:
st.str += std::string_view(e).size();
break;
case simdjson::dom::element_type::INT64:
st.num += static_cast<double>(int64_t(e));
break;
case simdjson::dom::element_type::UINT64:
st.num += static_cast<double>(uint64_t(e));
break;
case simdjson::dom::element_type::DOUBLE:
st.num += double(e);
break;
default:
break;
}
}
static void walk_od(simdjson::ondemand::value v, stats& st)
{
++st.nodes;
switch (v.type())
{
case simdjson::ondemand::json_type::object:
for (auto f : v.get_object())
{
st.str += std::string_view(f.unescaped_key()).size();
walk_od(f.value(), st);
}
break;
case simdjson::ondemand::json_type::array:
for (auto c : v.get_array())
{
walk_od(c.value(), st);
}
break;
case simdjson::ondemand::json_type::string:
st.str += std::string_view(v.get_string()).size();
break;
case simdjson::ondemand::json_type::number:
{
simdjson::ondemand::number n = v.get_number();
switch (n.get_number_type())
{
case simdjson::ondemand::number_type::signed_integer:
st.num += static_cast<double>(n.get_int64());
break;
case simdjson::ondemand::number_type::unsigned_integer:
st.num += static_cast<double>(n.get_uint64());
break;
default:
st.num += n.get_double();
break;
}
break;
}
case simdjson::ondemand::json_type::boolean:
(void)bool(v.get_bool());
break;
default:
(void)v.is_null();
break;
}
}
#if JSON_VIEW_BENCH_BOOST
static void walk(const boost::json::value& v, stats& st)
{
++st.nodes;
switch (v.kind())
{
case boost::json::kind::object:
for (const auto& kv : v.get_object())
{
st.str += kv.key().size();
walk(kv.value(), st);
}
break;
case boost::json::kind::array:
for (const auto& c : v.get_array())
{
walk(c, st);
}
break;
case boost::json::kind::string:
st.str += v.get_string().size();
break;
case boost::json::kind::int64:
st.num += static_cast<double>(v.get_int64());
break;
case boost::json::kind::uint64:
st.num += static_cast<double>(v.get_uint64());
break;
case boost::json::kind::double_:
st.num += v.get_double();
break;
default:
break;
}
}
#endif
// ---------------- selective access ----------------
// twitter: per status id, user.screen_name, retweet_count
// citm: per performance id, eventId, #seatCategories; #events
// canada: type, features[0].geometry.type, #coordinates
// jeopardy: per question: round == "Final Jeopardy!", len(category)
// status (one tweet): id, user.screen_name, retweet_count
// rpc: method, params.minuend, id
static double pick(const std::string& name, json_view r)
{
double acc = 0;
if (name == "twitter" || name == "status")
{
auto one = [&](json_view s)
{
acc += static_cast<double>(s["id"].get<std::uint64_t>());
acc += static_cast<double>(s["user"]["screen_name"].get_string().size());
acc += static_cast<double>(s["retweet_count"].get<std::int64_t>());
};
if (name == "twitter")
{
for (const json_view s : r["statuses"])
{
one(s);
}
}
else
{
one(r);
}
}
else if (name == "citm_catalog")
{
for (const json_view p : r["performances"])
{
acc += static_cast<double>(p["id"].get<std::uint64_t>() + p["eventId"].get<std::uint64_t>() + p["seatCategories"].size());
}
acc += static_cast<double>(r["events"].size());
}
else if (name == "canada")
{
const json_view g = r["features"][0]["geometry"];
acc += static_cast<double>(r["type"].get_string().size() + g["type"].get_string().size() + g["coordinates"].size());
}
else if (name == "jeopardy")
{
for (const json_view q : r)
{
acc += q["round"].get_string() == "Final Jeopardy!" ? 1 : 0;
acc += static_cast<double>(q["category"].get_string().size());
}
}
else if (name == "rpc")
{
acc += static_cast<double>(r["method"].get_string().size());
acc += static_cast<double>(r["params"]["minuend"].get<std::int64_t>() + r["id"].get<std::int64_t>());
}
return acc;
}
static double pick(const std::string& name, const json& r)
{
double acc = 0;
if (name == "twitter" || name == "status")
{
auto one = [&](const json & s)
{
acc += static_cast<double>(s["id"].get<std::uint64_t>());
acc += static_cast<double>(s["user"]["screen_name"].get_ref<const std::string&>().size());
acc += static_cast<double>(s["retweet_count"].get<std::int64_t>());
};
if (name == "twitter")
{
for (const auto& s : r["statuses"])
{
one(s);
}
}
else
{
one(r);
}
}
else if (name == "citm_catalog")
{
for (const auto& p : r["performances"])
{
acc += static_cast<double>(p["id"].get<std::uint64_t>() + p["eventId"].get<std::uint64_t>() + p["seatCategories"].size());
}
acc += static_cast<double>(r["events"].size());
}
else if (name == "canada")
{
const json& g = r["features"][0]["geometry"];
acc += static_cast<double>(r["type"].get_ref<const std::string&>().size() + g["type"].get_ref<const std::string&>().size() + g["coordinates"].size());
}
else if (name == "jeopardy")
{
for (const auto& q : r)
{
acc += q["round"].get_ref<const std::string&>() == "Final Jeopardy!" ? 1 : 0;
acc += static_cast<double>(q["category"].get_ref<const std::string&>().size());
}
}
else if (name == "rpc")
{
acc += static_cast<double>(r["method"].get_ref<const std::string&>().size());
acc += static_cast<double>(r["params"]["minuend"].get<std::int64_t>() + r["id"].get<std::int64_t>());
}
return acc;
}
static double pick(const std::string& name, yyjson_val* r)
{
double acc = 0;
auto get = [](yyjson_val * o, const char* k)
{
return yyjson_obj_get(o, k);
};
if (name == "twitter" || name == "status")
{
auto one = [&](yyjson_val * s)
{
acc += static_cast<double>(yyjson_get_uint(get(s, "id")));
acc += static_cast<double>(yyjson_get_len(get(get(s, "user"), "screen_name")));
acc += static_cast<double>(yyjson_get_sint(get(s, "retweet_count")));
};
if (name == "twitter")
{
std::size_t idx, max;
yyjson_val* s;
yyjson_arr_foreach(get(r, "statuses"), idx, max, s)
{
one(s);
}
}
else
{
one(r);
}
}
else if (name == "citm_catalog")
{
std::size_t idx, max;
yyjson_val* p;
yyjson_arr_foreach(get(r, "performances"), idx, max, p)
{
acc += static_cast<double>(yyjson_get_uint(get(p, "id")) + yyjson_get_uint(get(p, "eventId")) + yyjson_arr_size(get(p, "seatCategories")));
}
acc += static_cast<double>(yyjson_obj_size(get(r, "events")));
}
else if (name == "canada")
{
yyjson_val* g = get(yyjson_arr_get(get(r, "features"), 0), "geometry");
acc += static_cast<double>(yyjson_get_len(get(r, "type")) + yyjson_get_len(get(g, "type")) + yyjson_arr_size(get(g, "coordinates")));
}
else if (name == "jeopardy")
{
std::size_t idx, max;
yyjson_val* q;
yyjson_arr_foreach(r, idx, max, q)
{
acc += yyjson_equals_str(get(q, "round"), "Final Jeopardy!") ? 1 : 0;
acc += static_cast<double>(yyjson_get_len(get(q, "category")));
}
}
else if (name == "rpc")
{
acc += static_cast<double>(yyjson_get_len(get(r, "method")));
acc += static_cast<double>(yyjson_get_sint(get(get(r, "params"), "minuend")) + yyjson_get_sint(get(r, "id")));
}
return acc;
}
static double pick(const std::string& name, simdjson::dom::element r)
{
double acc = 0;
if (name == "twitter" || name == "status")
{
auto one = [&](simdjson::dom::element s)
{
acc += static_cast<double>(uint64_t(s["id"]));
acc += static_cast<double>(std::string_view(s["user"]["screen_name"]).size());
acc += static_cast<double>(int64_t(s["retweet_count"]));
};
if (name == "twitter")
{
for (auto s : simdjson::dom::array(r["statuses"]))
{
one(s);
}
}
else
{
one(r);
}
}
else if (name == "citm_catalog")
{
for (auto p : simdjson::dom::array(r["performances"]))
{
acc += static_cast<double>(uint64_t(p["id"]) + uint64_t(p["eventId"]) + simdjson::dom::array(p["seatCategories"]).size());
}
acc += static_cast<double>(simdjson::dom::object(r["events"]).size());
}
else if (name == "canada")
{
auto g = r["features"].at(0)["geometry"];
acc += static_cast<double>(std::string_view(r["type"]).size() + std::string_view(g["type"]).size() + simdjson::dom::array(g["coordinates"]).size());
}
else if (name == "jeopardy")
{
for (auto q : simdjson::dom::array(r))
{
acc += std::string_view(q["round"]) == "Final Jeopardy!" ? 1 : 0;
acc += static_cast<double>(std::string_view(q["category"]).size());
}
}
else if (name == "rpc")
{
acc += static_cast<double>(std::string_view(r["method"]).size());
acc += static_cast<double>(int64_t(r["params"]["minuend"]) + int64_t(r["id"]));
}
return acc;
}
static double pick_od(const std::string& name, simdjson::ondemand::document& d)
{
double acc = 0;
if (name == "twitter" || name == "status")
{
auto one = [&](simdjson::ondemand::object s)
{
acc += static_cast<double>(uint64_t(s["id"]));
acc += static_cast<double>(std::string_view(s["user"]["screen_name"]).size());
acc += static_cast<double>(int64_t(s["retweet_count"]));
};
if (name == "twitter")
{
for (auto s : d["statuses"])
{
one(s.get_object());
}
}
else
{
one(d.get_object());
}
}
else if (name == "citm_catalog")
{
simdjson::ondemand::object ev = d["events"].get_object();
acc += static_cast<double>(ev.count_fields());
for (auto p : d["performances"])
{
simdjson::ondemand::object o = p.get_object();
const auto a = uint64_t(o["eventId"]) + uint64_t(o["id"]);
simdjson::ondemand::array sc = o["seatCategories"].get_array();
acc += static_cast<double>(a + sc.count_elements());
}
}
else if (name == "canada")
{
acc += static_cast<double>(std::string_view(d["type"]).size());
auto g = d["features"].at(0)["geometry"];
acc += static_cast<double>(std::string_view(g["type"]).size());
simdjson::ondemand::array co = g["coordinates"].get_array();
acc += static_cast<double>(co.count_elements());
}
else if (name == "jeopardy")
{
for (auto q : d)
{
simdjson::ondemand::object o = q.get_object();
acc += static_cast<double>(std::string_view(o["category"]).size());
acc += std::string_view(o["round"]) == "Final Jeopardy!" ? 1 : 0;
}
}
else if (name == "rpc")
{
acc += static_cast<double>(std::string_view(d["method"]).size());
acc += static_cast<double>(int64_t(d["params"]["minuend"]));
acc += static_cast<double>(int64_t(d["id"]));
}
return acc;
}
// ---------------- harness ----------------
static std::string slurp(const std::string& p)
{
std::ifstream f(p, std::ios::binary);
std::stringstream ss;
ss << f.rdbuf();
return ss.str();
}
struct engine
{
std::string name;
std::function<void()> fn;
};
int main(int argc, char** argv)
{
if (argc < 2)
{
std::fprintf(stderr, "usage: %s <json_test_data directory> [rounds] [document]\n", argv[0]);
return 1;
}
const std::string T = std::string(argv[1]) + "/";
const int rounds = argc > 2 ? std::atoi(argv[2]) : 30;
const std::string only = argc > 3 ? argv[3] : "";
struct doc
{
std::string name, text;
int batch;
};
std::vector<doc> docs;
for (const char* f :
{"nativejson-benchmark/twitter.json", "nativejson-benchmark/citm_catalog.json", "nativejson-benchmark/canada.json", "jeopardy/jeopardy.json"
})
{
std::string n = std::string(f).substr(std::string(f).find('/') + 1);
docs.push_back({n.substr(0, n.size() - 5), slurp(T + f), 1});
}
docs.push_back({"status", json::parse(docs[0].text)["statuses"][0].dump(), 200});
docs.push_back({"rpc", R"({"jsonrpc": "2.0", "method": "subtract", "params": {"minuend": 42, "subtrahend": 23}, "id": 3})", 5000});
// correctness cross-check of the workloads
for (const auto& dc : docs)
{
stats a, b, c, dd;
walk(json::parse(dc.text), a);
auto d = json_document::parse(dc.text);
walk(d.root(), b);
yyjson_doc* y = yyjson_read(dc.text.data(), dc.text.size(), 0);
walk(yyjson_doc_get_root(y), c);
simdjson::dom::parser p;
walk(p.parse(dc.text).value(), dd);
const bool ok = a.nodes == b.nodes && a.nodes == c.nodes && a.nodes == dd.nodes && a.str == b.str && a.str == c.str && a.str == dd.str
&& std::fabs(a.num - b.num) <= 1e-9 * std::fabs(a.num) && std::fabs(a.num - c.num) <= 1e-9 * std::fabs(a.num);
const double pa = pick(dc.name, json::parse(dc.text)), pb = pick(dc.name, d.root()), pc = pick(dc.name, yyjson_doc_get_root(y)), pd = pick(dc.name, p.parse(dc.text).value());
std::printf("check %-13s traverse %s select %s\n", dc.name.c_str(), ok ? "OK" : "MISMATCH", (pa == pb && pa == pc && pa == pd) ? "OK" : "MISMATCH");
yyjson_doc_free(y);
}
std::FILE* csv = std::fopen("bench_view.csv", "w");
std::fprintf(csv, "doc,bytes,workload,engine,ns\n");
json_document reused;
simdjson::dom::parser sj;
simdjson::ondemand::parser od;
for (const auto& dc : docs)
{
if (!only.empty() && dc.name != only)
{
continue;
}
const std::string& s = dc.text;
const simdjson::padded_string ps(s);
const std::string name = dc.name;
std::vector<std::pair<std::string, std::vector<engine>>> workloads;
workloads.push_back({"parse", {
{"json_view", [&] { auto d = json_document::parse(s); g_sink = static_cast<double>(d.node_count()); }},
{"json_view (reused)", [&] { reused.read(s); g_sink = static_cast<double>(reused.node_count()); }},
{"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); g_sink = static_cast<double>(yyjson_doc_get_val_count(d)); yyjson_doc_free(d); }},
{"simdjson DOM", [&] { auto e = sj.parse(ps).value_unsafe(); g_sink = e.is_object(); }},
#if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); g_sink = v.is_object(); }},
#endif
{"json::parse", [&] { json j = json::parse(s); g_sink = static_cast<double>(j.size()); }},
}});
workloads.push_back({"traverse", {
{"json_view", [&] { auto d = json_document::parse(s); stats st; walk(d.root(), st); g_sink = st.num; }},
{"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); stats st; walk(yyjson_doc_get_root(d), st); g_sink = st.num; yyjson_doc_free(d); }},
{"simdjson DOM", [&] { stats st; walk(sj.parse(ps).value_unsafe(), st); g_sink = st.num; }},
{"simdjson OD", [&] { auto d = od.iterate(ps).value_unsafe(); stats st; walk_od(d.get_value().value_unsafe(), st); g_sink = st.num; }},
#if JSON_VIEW_BENCH_BOOST
{"Boost.JSON", [&] { boost::json::monotonic_resource mr; auto v = boost::json::parse(s, &mr); stats st; walk(v, st); g_sink = st.num; }},
#endif
{"json::parse", [&] { json j = json::parse(s); stats st; walk(j, st); g_sink = st.num; }},
}});
workloads.push_back({"select", {
{"json_view", [&] { auto d = json_document::parse(s); g_sink = pick(name, d.root()); }},
{"yyjson", [&] { yyjson_doc* d = yyjson_read(s.data(), s.size(), 0); g_sink = pick(name, yyjson_doc_get_root(d)); yyjson_doc_free(d); }},
{"simdjson DOM", [&] { g_sink = pick(name, sj.parse(ps).value_unsafe()); }},
{"simdjson OD", [&] { auto d = od.iterate(ps).value_unsafe(); g_sink = pick_od(name, d); }},
{"json::parse", [&] { json j = json::parse(s); g_sink = pick(name, j); }},
}});
{
// serialization of an already parsed document
static json_document vd;
vd.read(s);
static yyjson_doc* yd = nullptr;
if (yd)
{
yyjson_doc_free(yd);
}
yd = yyjson_read(s.data(), s.size(), 0);
static simdjson::dom::parser sjd;
static simdjson::dom::element se;
se = sjd.parse(ps).value_unsafe();
static json jd;
jd = json::parse(s);
workloads.push_back({"dump", {
{"json_view", [&] { std::string o = vd.root().dump(); g_sink = static_cast<double>(o.size()); }},
{"yyjson", [&] { std::size_t n = 0; char* o = yyjson_write(yd, 0, &n); g_sink = static_cast<double>(n); std::free(o); }},
{"simdjson DOM", [&] { std::string o = simdjson::to_string(se); g_sink = static_cast<double>(o.size()); }},
{"json::parse", [&] { std::string o = jd.dump(); g_sink = static_cast<double>(o.size()); }},
}});
}
for (auto& wl : workloads)
{
std::vector<double> best(wl.second.size(), 1e300);
const int r = s.size() > 10000000 ? std::max(3, rounds / 5) : rounds;
for (int i = 0; i < r; ++i)
{
for (std::size_t k = 0; k < wl.second.size(); ++k)
{
const auto t0 = std::chrono::steady_clock::now();
for (int b = 0; b < dc.batch; ++b)
{
wl.second[k].fn();
}
const double ns = std::chrono::duration<double, std::nano>(std::chrono::steady_clock::now() - t0).count() / dc.batch;
best[k] = std::min(best[k], ns);
}
}
const double ref = best[0];
std::printf("%-13s %-9s", dc.name.c_str(), wl.first.c_str());
for (std::size_t k = 0; k < wl.second.size(); ++k)
{
const double us = best[k] / 1e3;
std::printf(" %s %s%s (%.2fx)", wl.second[k].name.c_str(), us >= 100 ? "" : "", (us >= 1000 ? std::to_string(static_cast<long>(us)) + "us" : (std::to_string(us).substr(0, 5) + "us")).c_str(), best[k] / ref);
std::fprintf(csv, "%s,%zu,%s,%s,%.1f\n", dc.name.c_str(), s.size(), wl.first.c_str(), wl.second[k].name.c_str(), best[k]);
}
std::printf("\n");
std::fflush(stdout);
}
}
std::fclose(csv);
}
+302
View File
@@ -0,0 +1,302 @@
#!/usr/bin/env python3
# __ _____ _____ _____
# __| | __| | | | JSON for Modern C++ (supporting code)
# | | |__ | | | | | | version 3.12.0
# |_____|_____|_____|_|___| https://github.com/nlohmann/json
#
# SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
# SPDX-License-Identifier: MIT
"""Compare json_view with yyjson, simdjson, Boost.JSON, and json::parse.
Builds bench_view.cpp, bench_corpus.cpp, and bench_edit.cpp against the include/ directory of
this checkout, runs them, and writes the results with everything needed to
reproduce them (date, commit, CPU, OS, compiler, library versions, flags) to
results/<date>-<host>.md and .csv next to this script.
The other libraries come from the system (--system, the default: pkg-config
or Homebrew) or are downloaded as pinned releases and checked against their
SHA-256 (--download). Boost.JSON is optional: without Boost headers, its
columns are skipped, and the results say so.
Only the Python 3 standard library is used; a C++17 compiler is needed.
"""
import argparse
import datetime
import hashlib
import os
import platform
import re
import shlex
import shutil
import subprocess
import sys
import tarfile
import urllib.request
HERE = os.path.dirname(os.path.abspath(__file__))
REPO = os.path.abspath(os.path.join(HERE, '..', '..', '..'))
# pinned releases for --download; the hashes are those of the archives
PINNED = {
'yyjson': {
'version': '0.13.0',
'url': 'https://github.com/ibireme/yyjson/archive/refs/tags/0.13.0.tar.gz',
'sha256': '34e0f62a2bc11ab20d601e8ca1cc2b2079503aa45119a19133d89d19b94a0fae',
'dir': 'yyjson-0.13.0',
},
'simdjson': {
'version': '4.6.11',
'url': 'https://github.com/simdjson/simdjson/archive/refs/tags/v4.6.11.tar.gz',
'sha256': '61d948fc24f0d793829ad658058e7597d064988a89b4607ea02e401a82df98ff',
'dir': 'simdjson-4.6.11',
},
'boost': {
'version': '1.92.0',
'url': 'https://archives.boost.io/release/1.92.0/source/boost_1_92_0.tar.gz',
'sha256': 'c4a3b310ddd2472416e091067166b0713be97c63f38c212c484ada022fd296ce',
'dir': 'boost_1_92_0',
},
}
# the documents of bench_view.cpp, relative to the json_test_data directory
DEFAULT_CORPUS = [
'nativejson-benchmark/twitter.json',
'nativejson-benchmark/citm_catalog.json',
'nativejson-benchmark/canada.json',
'jeopardy/jeopardy.json',
]
def run(cmd, **kwargs):
print('+ ' + ' '.join(shlex.quote(c) for c in cmd), flush=True)
return subprocess.run(cmd, check=True, **kwargs)
def output(cmd):
try:
return subprocess.run(cmd, check=True, capture_output=True, text=True).stdout.strip()
except (OSError, subprocess.CalledProcessError):
return ''
# ---------------------------------------------------------------------------
# libraries
# ---------------------------------------------------------------------------
class Library:
"""include directories, sources to compile, and linker flags of a library"""
def __init__(self, name, include=None, sources=None, link=None, version=''):
self.name = name
self.include = include or []
self.sources = sources or []
self.link = link or []
self.version = version
def header_version(path, pattern):
try:
with open(path, encoding='utf-8', errors='replace') as f:
m = re.search(pattern, f.read())
return m.group(1) if m else ''
except OSError:
return ''
def library_version(name, include_dirs):
patterns = {
'yyjson': ('yyjson.h', r'#define\s+YYJSON_VERSION_STRING\s+"([^"]+)"'),
'simdjson': ('simdjson.h', r'#define\s+SIMDJSON_VERSION\s+"?([0-9.]+)"?'),
'boost': (os.path.join('boost', 'version.hpp'), r'#define\s+BOOST_LIB_VERSION\s+"([^"]+)"'),
}
header, pattern = patterns[name]
for d in include_dirs:
v = header_version(os.path.join(d, header), pattern)
if v:
return v.replace('_', '.')
return ''
def system_library(name):
"""a library found with pkg-config or Homebrew, or None"""
flags = output(['pkg-config', '--cflags', '--libs', name]).split()
if flags:
include = [f[2:] for f in flags if f.startswith('-I')]
link = [f for f in flags if f.startswith('-L') or f.startswith('-l')]
libdirs = [f[2:] for f in link if f.startswith('-L')]
link += ['-Wl,-rpath,' + d for d in libdirs]
return Library(name, include, [], link, library_version(name, include))
prefix = output(['brew', '--prefix', name]) if shutil.which('brew') else ''
if prefix and os.path.isdir(os.path.join(prefix, 'include')):
include = [os.path.join(prefix, 'include')]
link = []
if name != 'boost':
lib = os.path.join(prefix, 'lib')
link = ['-L' + lib, '-l' + name, '-Wl,-rpath,' + lib]
return Library(name, include, [], link, library_version(name, include))
if name == 'boost':
for d in ['/usr/include', '/usr/local/include']:
if os.path.isfile(os.path.join(d, 'boost', 'json.hpp')):
return Library(name, [d], [], [], library_version(name, [d]))
return None
def download_library(name, work):
"""a pinned release, downloaded and checked, or an error"""
pin = PINNED[name]
archive = os.path.join(work, 'download', os.path.basename(pin['url']))
os.makedirs(os.path.dirname(archive), exist_ok=True)
if not os.path.isfile(archive):
print(f'downloading {pin["url"]}', flush=True)
urllib.request.urlretrieve(pin['url'], archive)
with open(archive, 'rb') as f:
digest = hashlib.sha256(f.read()).hexdigest()
if digest != pin['sha256']:
sys.exit(f'error: SHA-256 of {archive} is {digest}, expected {pin["sha256"]}')
src = os.path.join(work, 'download', pin['dir'])
if not os.path.isdir(src):
with tarfile.open(archive) as t:
# (the 'data' filter rejects links and paths outside the target where Python has it)
kwargs = {'filter': 'data'} if hasattr(tarfile, 'data_filter') else {}
t.extractall(os.path.join(work, 'download'), **kwargs) # noqa: S202 (checked archive)
if name == 'yyjson':
return Library(name, [os.path.join(src, 'src')], [os.path.join(src, 'src', 'yyjson.c')], [], pin['version'])
if name == 'simdjson':
single = os.path.join(src, 'singleheader')
return Library(name, [single], [os.path.join(single, 'simdjson.cpp')], [], pin['version'])
return Library(name, [src], [], [], pin['version'])
# ---------------------------------------------------------------------------
# machine description
# ---------------------------------------------------------------------------
def cpu_model():
if sys.platform == 'darwin':
return output(['sysctl', '-n', 'machdep.cpu.brand_string'])
try:
with open('/proc/cpuinfo', encoding='utf-8') as f:
for line in f:
if line.startswith('model name') or line.startswith('Model'):
return line.split(':', 1)[1].strip()
except OSError:
pass
# (AArch64 Linux: /proc/cpuinfo has no model name, lscpu knows it)
for line in output(['lscpu']).splitlines():
if line.startswith('Model name:'):
return line.split(':', 1)[1].strip()
return platform.processor() or platform.machine()
def git_commit():
commit = output(['git', '-C', REPO, 'rev-parse', '--short=12', 'HEAD'])
dirty = output(['git', '-C', REPO, 'status', '--porcelain', '--untracked-files=no'])
return commit + (' (with local changes)' if dirty else '')
# ---------------------------------------------------------------------------
# main
# ---------------------------------------------------------------------------
def main():
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument('--data', required=True, help='json_test_data directory (with nativejson-benchmark/ and jeopardy/)')
ap.add_argument('--download', action='store_true', help='use pinned downloads instead of system libraries')
ap.add_argument('--no-boost', action='store_true', help='skip Boost.JSON')
ap.add_argument('--native', action='store_true', help='compile for this CPU (-march=native / -mcpu=native)')
ap.add_argument('--rounds', type=int, default=30, help='rounds of bench_view (default: 30)')
ap.add_argument('--corpus', nargs='*', default=[], help='more files for bench_corpus')
ap.add_argument('--build-dir', default=os.path.join(HERE, 'build'), help='where to build (default: build/ next to this script)')
args = ap.parse_args()
cxx = os.environ.get('CXX', 'c++')
cc = os.environ.get('CC', 'cc')
os.makedirs(args.build_dir, exist_ok=True)
libs = {}
for name in ['yyjson', 'simdjson', 'boost']:
if name == 'boost' and args.no_boost:
continue
lib = download_library(name, args.build_dir) if args.download else system_library(name)
if lib is None and name != 'boost':
sys.exit(f'error: {name} not found; install it, or use --download')
if lib is not None:
libs[name] = lib
with_boost = 'boost' in libs
if not with_boost:
print('Boost.JSON not found: its columns are skipped', flush=True)
flags = ['-std=c++17', '-O3', '-DNDEBUG', f'-DJSON_VIEW_BENCH_BOOST={1 if with_boost else 0}']
if args.native:
flags.append('-mcpu=native' if platform.machine().lower() in ('arm64', 'aarch64') else '-march=native')
include = ['-I' + os.path.join(REPO, 'include')] + ['-I' + d for lib in libs.values() for d in lib.include]
link = [f for lib in libs.values() for f in lib.link]
# C sources of downloaded libraries are compiled once
objects = []
for lib in libs.values():
for src in lib.sources:
obj = os.path.join(args.build_dir, os.path.basename(src) + '.o')
compiler = cc if src.endswith('.c') else cxx
run([compiler] + (['-std=c++17'] if compiler == cxx else []) + ['-O3', '-DNDEBUG', '-c', src, '-o', obj]
+ ['-I' + d for d in lib.include])
objects.append(obj)
binaries = {}
for bench in ['bench_view', 'bench_corpus', 'bench_edit']:
exe = os.path.join(args.build_dir, bench)
run([cxx] + flags + include + [os.path.join(HERE, bench + '.cpp')] + objects + link + ['-o', exe])
binaries[bench] = exe
# run: bench_view on its documents, bench_corpus on those and the given files
corpus = [os.path.join(args.data, f) for f in DEFAULT_CORPUS] + args.corpus
outputs = {}
outputs['bench_view'] = run([binaries['bench_view'], args.data, str(args.rounds)], cwd=args.build_dir,
capture_output=True, text=True).stdout
outputs['bench_corpus'] = run([binaries['bench_corpus']] + corpus, cwd=args.build_dir,
capture_output=True, text=True).stdout
outputs['bench_edit'] = run([binaries['bench_edit'], args.data, str(max(1, args.rounds // 2))], cwd=args.build_dir,
capture_output=True, text=True).stdout
for name, text in outputs.items():
print(text)
# results with their metadata
now = datetime.datetime.now()
host = re.sub(r'[^A-Za-z0-9-]+', '-', platform.node().split('.')[0]) or 'host'
stem = os.path.join(HERE, 'results', f'{now:%Y-%m-%d}-{host}')
os.makedirs(os.path.dirname(stem), exist_ok=True)
meta = [
('date', f'{now:%Y-%m-%d %H:%M}'),
('commit', git_commit()),
('CPU', cpu_model()),
('OS', f'{platform.system()} {platform.release()} ({platform.machine()})'),
('compiler', output([cxx, '--version']).splitlines()[0] if output([cxx, '--version']) else cxx),
('flags', ' '.join(flags)),
('yyjson', libs['yyjson'].version),
('simdjson', libs['simdjson'].version),
('Boost.JSON', libs['boost'].version if with_boost else 'skipped (not found)'),
('libraries from', 'pinned downloads' if args.download else 'the system'),
('rounds', str(args.rounds)),
]
with open(stem + '.md', 'w', encoding='utf-8') as f:
f.write(f'# json_view comparison, {now:%Y-%m-%d}\n\n')
f.write('Generated by `tests/benchmarks/json_view/compare.py`; best of the interleaved rounds.\n\n')
f.write('| | |\n|---|---|\n')
for key, value in meta:
f.write(f'| {key} | {value} |\n')
for name, text in outputs.items():
f.write(f'\n## {name}\n\n```\n{text.rstrip()}\n```\n')
with open(stem + '.csv', 'w', encoding='utf-8') as out:
out.write(''.join(f'# {key}: {value}\n' for key, value in meta))
for name in ['bench_view', 'bench_corpus', 'bench_edit']:
path = os.path.join(args.build_dir, name + '.csv')
if os.path.isfile(path):
with open(path, encoding='utf-8') as f:
out.write(f'# {name}\n' + f.read())
print(f'results: {stem}.md, {stem}.csv')
if __name__ == '__main__':
main()
+34
View File
@@ -144,3 +144,37 @@ static void ViewMaterialize(benchmark::State& state, const char* filename)
state.SetBytesProcessed(state.iterations() * str.size());
}
JSON_VIEW_BENCHMARK_FILES(ViewMaterialize);
//////////////////////////////////////////////////////////////////////////////
// serialize a parsed document (compare with Dump)
//////////////////////////////////////////////////////////////////////////////
static void ViewDump(benchmark::State& state, const char* filename, int indent)
{
const std::string str = read_file(filename);
const json_document d = json_document::parse(str);
while (state.KeepRunning())
{
std::string output = d.root().dump(indent);
benchmark::DoNotOptimize(output);
}
state.SetBytesProcessed(state.iterations() * d.root().dump(indent).size());
}
BENCHMARK_CAPTURE(ViewDump, jeopardy / -, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", -1);
BENCHMARK_CAPTURE(ViewDump, jeopardy / 4, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json", 4);
BENCHMARK_CAPTURE(ViewDump, canada / -, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", -1);
BENCHMARK_CAPTURE(ViewDump, canada / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json", 4);
BENCHMARK_CAPTURE(ViewDump, citm_catalog / -, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", -1);
BENCHMARK_CAPTURE(ViewDump, citm_catalog / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", 4);
BENCHMARK_CAPTURE(ViewDump, twitter / -, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", -1);
BENCHMARK_CAPTURE(ViewDump, twitter / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", 4);
BENCHMARK_CAPTURE(ViewDump, floats / -, TEST_DATA_DIRECTORY "/regression/floats.json", -1);
BENCHMARK_CAPTURE(ViewDump, floats / 4, TEST_DATA_DIRECTORY "/regression/floats.json", 4);
BENCHMARK_CAPTURE(ViewDump, signed_ints / -, TEST_DATA_DIRECTORY "/regression/signed_ints.json", -1);
BENCHMARK_CAPTURE(ViewDump, signed_ints / 4, TEST_DATA_DIRECTORY "/regression/signed_ints.json", 4);
BENCHMARK_CAPTURE(ViewDump, unsigned_ints / -, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json", -1);
BENCHMARK_CAPTURE(ViewDump, unsigned_ints / 4, TEST_DATA_DIRECTORY "/regression/unsigned_ints.json", 4);
BENCHMARK_CAPTURE(ViewDump, small_signed_ints / -, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json", -1);
BENCHMARK_CAPTURE(ViewDump, small_signed_ints / 4, TEST_DATA_DIRECTORY "/regression/small_signed_ints.json", 4);
+6
View File
@@ -10,6 +10,12 @@ produces, and that a rejected input makes both parsers throw with an identical `
reuses the `corpus_json` corpus (or, for the `make fuzz_testing_json_view` target below, `tests/data/json_tests`) rather
than a format of its own.
`json_view_image_fuzzer` (`tests/src/fuzzer-json_view_image.cpp`) tests the images of `json_document` (`save()` and
`load()`). It uses each input twice: as an image, which `load()` must either reject with `parse_error.116` or read
safely (with `image_check::full`, the document must also serialize to the JSON it reads as), and as a JSON text, whose
image must load and serialize to the same text. A corpus of images can be made from JSON files with a small program
that calls `json_document::parse(text).save()`; plain JSON files work as well.
## Corpus creation
For most effective fuzzing, a [corpus](https://llvm.org/docs/LibFuzzer.html#corpus) should be provided. A corpus is a

Some files were not shown because too many files have changed in this diff Show More