Lookups find the first member of a repeated key again, so set(object, key,
value) assigns that member (keeping the key where it is) and drops the
others, as documented before, and a view taken from object[key] shows the
new value. This undoes the code and documentation changes of 5cbb8e6d6.
Lookups in edited objects (navigation of editable documents) stop at the
first match, like those of parsed objects.
The tests that commit added expect the first member from lookups now. In the
seeded differential test, the documents with repeated keys repeat them after
the real members; the reference holds the first members, which the edits
address and the lookups are compared with, while materialize() is compared
with what parse() makes of the document's dump.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An editable document reads the last member of a repeated key (as the
read-only document does), but set(object, key, value) assigned the first
one: a view taken from object[key] before the call did not show the new
value. It now assigns the last member's value, keeps the key at the
position of its first occurrence (where materialize() puts it), and drops
the other members.
The seeded differential test now also edits documents whose objects
repeat keys and compares every lookup (operator[], at, find, value,
contains, count, and JSON pointers) with the parsed basic_json value.
A targeted test covers duplicates before and after edits that move the
object, in large objects with an index, in moved arrays, and in values
copied from other documents.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Replacing an array or object that spans several nodes by a scalar looks
up its parent from the root (find_parent), which is linear in the size
of the document: setting every element of a 20,000 element array took
more than half a second. The parent only has to switch to links once;
afterward the extent of the replaced value no longer matters. Nodes that
an entry of a moved sequence links to are now marked (node_flags::linked),
and assign() skips the lookup for them.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
set(json_pointer, value) turned a null parent into an object for every
last reference token, so "/a/0" and "/a/-" below {"a":null} made
{"a":{"0":1}}, where basic_json's operator[](json_pointer) makes an
array. A null parent now becomes an array for "-" and for digit tokens
(filled with nulls up to the index) and an object otherwise. An invalid
index is reported before the parent changes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
append_text() doubled the capacity of the arena and rejected the result
if it exceeded 4 GiB, so that an arena of more than 2 GiB could not grow
even though the 32-bit offsets of nodes address 4 GiB - 1 bytes. The
capacity is now clamped to that limit (text_capacity()), and an append
is rejected only if the bytes themselves do not fit.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An editable document only holds valid UTF-8, but copy_scalar() copied the
strings and keys of a view of another document unchecked. A document of
a weaker check (or a borrowed text that changed after parsing) could
therefore bring ill-formed UTF-8 into it. The copy is checked now, with
the error that dump() reports for the string; copies within the same
document stay unchecked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Counting and copying the nodes of a view or a basic_json value into an
editable document recursed once per nesting level, so that a deeply
nested value overflowed the stack. The four functions now walk the value
with an explicit stack, as materialize() does.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_document gets a second template parameter, Editable
(false by default), plus the aliases json_editable_document,
json_editable_view, ordered_json_editable_document and
ordered_json_editable_view.
Editable documents can change values and structure without
rewriting the source text: set()/push_back() on values, keys,
array indices and JSON pointers; insert() before an array
element; erase() of an object key, array index or JSON pointer.
New values and element sequences go into edit storage that the
document owns and never moves, so views keep referring to their
value across edits and a parsed node never moves. Read-only
documents walk the plain node array and are unaffected.
Strings are checked for UTF-8 on entry, so dump() of an editable
document never throws type_error.316. Binary values cannot be
stored (type_error.319).
A seeded differential test applies random edits to an editable
document and to the equivalent ordered_json and compares both
after every step.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The dump and comparison tests moved to their own file in the dump pull
request; the hash index tests of this branch join them there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more return the first member of a
repeated key again, like the linear search of smaller objects. This undoes
the code change of a69542046; its test now expects the first member from
lookups (with and without a table) and the last value from materialize()
and basic_json::parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more now return the last member of
a repeated key, like the linear search of smaller objects and like
materialize() and basic_json::parse(). build_object_index let the first
occurrence win, so the same text gave different results depending on the
size of the object.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
shrink_to_fit() now trims the tables of large objects like the node array and the decoded strings, and the list of large objects is released as soon as the tables are built.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The key hash is not seeded, so keys chosen to collide made building the table quadratic (20,000 colliding keys took 470 ms to parse). A key may now sit at most 64 slots from its home slot; if a key would sit further away, the table is dropped and the object is searched linearly. Lookups stop after the same distance.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.
Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.
On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.
Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.
simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The MinGW linker of the Windows clang jobs cannot link object files with more
than 32767 sections ("relocation truncated to fit: IMAGE_REL_AMD64_REL32
against `.rdata'"). unit-json_view.cpp reaches that limit as the stack
grows, so its "json_view dump" and "json_view comparison" test cases move
into unit-json_view_dump.cpp. The test generator and has_duplicate_keys()
that both files use move into json_view_test_helpers.hpp.
The new file mentions JSON_HAS_CPP_17, so it is built for C++17 like the file
it was split from.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
dump() writes every member and == resolves duplicate keys as parse()
does, whereas lookups find the first member of a duplicate key.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The source extent of a value was read from the next node, falling back to the rest of the document when that node held a decoded string. dump() of a small value could thus allocate a buffer as large as the document. Skip a few such nodes, cap the fallback estimate, and shrink a buffer that is much larger than its output.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Add basic_json_view::dump() and the comparison operators, and read
floats from the parser's digit layout instead of rescanning the
token.
dump(indent, indent_char, ensure_ascii, number_format) writes a
value the way ordered_json::parse(text).dump() writes it for the
same arguments: members in document order, all of them should a
key occur more than once; strings escaped by the same rules, using
the library's scanning kernels; floats written with the library's
to_chars conversion, so the output equals basic_json's byte for
byte; integers copied from the source, where they are already
canonical, except -0, which parse() reads as 0. There is no
error_handler argument, because the view only holds valid UTF-8.
number_format::source copies numbers exactly as they appear in the
source (e.g. "1.50", "1E2", "-0"), which basic_json cannot provide.
operator<< takes the indentation from the stream width, as for
basic_json. The writer walks iteratively, so nesting depth is
limited by memory only.
operator== and operator!= compare two views, or a view and a
basic_json value in either order, by the rules basic_json's
operator== uses: numbers compare by value across their types,
objects compare by their members with duplicate keys resolved as
parse() resolves them, member order matters only where the object
type keeps one, and discarded views compare as discarded basic_json
values do, including under JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON.
Nothing is materialized except single scalars.
While parsing, the view now records where the integer digits, the
fraction digits, and the exponent of a float token are, so floats
and doubles with at most 19 digits are read from that layout with
the library's decimal_to_float() instead of rescanning the token.
Both round correctly, so the values are those of parse(). get<double>(),
materialize(), dump(), and the comparisons all use it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The same changes as on the dump branch, where they were first made, so
that this branch passes clang-tidy on its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Looking up the last member of a duplicate key cannot stop at a match, so
every lookup scanned the whole object (1.6 to 3.4 times slower for small
objects). operator[](key), at, find, value, contains, count, and JSON
pointer resolution return the first member again, as yyjson and simdjson
do; materialize() and get<map>() keep the last value, as parse().
The documentation says so in the feature page and on each lookup page,
and explains how to get the value parse() would give. The integer index
templates and the discarded chaining of operator[] stay.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A view is tested with is_discarded(); the lookups in at(), value() and
contains(json_pointer) and the unit tests no longer rely on the explicit
conversion to bool.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Lookups (operator[], at, find, contains, count, value, JSON pointers)
return the last member of a duplicate key, as materialize() and
parse() keep it.
* operator[] on a discarded view returns a discarded view instead of
throwing, so v["a"]["b"] is safe for a missing "a".
* operator[] and at() take any integer type (not only int and size_t),
fixing ambiguous calls with unsigned, long, std::int64_t, ...
* Fix the operator[] documentation, which claimed a discarded view for
a type mismatch where type_error.305 is thrown.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Give basic_json_view the read-only access functions of basic_json:
operator[] and at() with keys and indices, front()/back(), find(),
contains(), count(), begin()/end() and cbegin()/cend(), items()
with structured bindings from C++17 on, and type_name().
Exceptions have the ids and messages of the const functions of
basic_json. Where basic_json has undefined behavior the view
answers safely: operator[] with a missing key or an out-of-range
index returns a discarded view, and front()/back() of an empty
container throw invalid_iterator.214. Objects are iterated in
document order, and all members are visited; duplicate-key lookups
find the first member (as yyjson and simdjson do), while parse(),
materialize(), and the map conversions keep the last value, as
parse() does. Keys of up to 16 bytes are compared with two
overlapping loads.
Add value conversions: get<T>()/get_to() for arithmetic types,
bool, nullptr_t, strings (std::basic_string copied,
string_view_t without a copy), BasicJsonType, views, std::vector,
and maps with string keys; get_string() for the string without a
copy; number_token() for the number exactly as written in the
source; value() with keys and JSON pointers; and operator[]/at()/
contains() with JSON pointers. Everything else, including types
with from_json(), goes through materialize() of that subtree.
get<T>() of arithmetic types is inlined down to the conversion, so
reading an integer needs no call.
detail::json_pointer_access exposes a pointer's reference tokens
to code outside basic_json.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
const_text is const, so std::move does not move from it; the second call
is part of the const rvalue test.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The fuzzer only parsed a std::string with default options. Also parse an
exact-size byte vector, which has no NUL after its last byte and takes the
bounds-checked path, and derive ignore_comments and ignore_trailing_commas
from the first input byte, comparing against json::parse and json::accept
with the same options.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The check rejects inputs of 0xFFFFFFF0 bytes or more, but the exception
message and the documentation said 4 GiB. Name the limit once
(max_input_size), and state 4 GiB minus 16 bytes in the message and the
documentation. Test the limit with a container that only claims the size.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
reserve(n) and the growth of the index computed n * sizeof(node) without a
check, which wraps around on 32-bit targets for inputs of about 1 GiB and
allocates a too small array. Throw std::bad_alloc for a count beyond the
address space and clamp the growth step to it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-little-endian path of the node writer wrote the integer's native
word over len and next, so len got the high half there. Compose and split
the value explicitly (len is the low half, next the high half); the
little-endian path stays a plain memcpy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
auto v = json_document::parse(text).root() compiled and left the view
dangling. Delete the overload for rvalue documents, take a named document
in the tests, and document the lifetime rule.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
explicit operator bool meant "refers to a value", which silently differs
from what a basic_json converts to. Use !v.is_discarded() instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json_document::parse(std::move(const_string)) failed to compile with
"no matching read_kind": only a non-const rvalue std::string can be moved
from. Treat a const rvalue as a copied byte container.
parse() does not accept everything BasicJsonType::parse() does: a FILE*
and pointers to or arrays of wide characters are rejected at compile time.
List the supported inputs instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
parse(ptr, len) compiled: len converted to allow_exceptions, and ptr was
read as a C string, past the end of a buffer without a terminating NUL.
Delete the overloads of parse, parse_copy, accept, and read that take an
integer other than bool where the flags are expected.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).
A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.
The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.
basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.
detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.
A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-template overloads for double asserted binary64 doubles wherever json.hpp was included, so the library no longer compiled where double is not IEEE 754 binary64 (AVR, -fshort-double). A trait now picks Zmij for any binary64 type, including a long double of that format (MSVC, Apple Arm), and Grisu2 for the others. Remove the unused write_short_decimal, powers_of_ten_16, zmij::decimal, zmij::to_decimal, and shortest_digits(double).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Write doubles with the conversion of Zmij by Victor Zverovich (MIT),
ported to C++11 (detail/conversions/zmij.hpp). It finds the
shortest decimal that reads back as the same double, and the
closest one if there are several. Grisu2, used until now, is fast
but not always shortest: it sometimes writes a 17th digit where 16
suffice, or a last digit that is not the closest. The layout is
unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2.
Digits are converted eight at a time with the BCD conversion of
Xiang JunBo, as in Zmij, and written with one byte swap per eight
digits and fixed-size moves instead of per-digit loops. Leading and
trailing zeros are counted from those bytes. to_chars() uses a
local buffer when the caller's is shorter than the 41 bytes this
may write. The powers of ten come from the number-parsing table,
adjusted where it holds values rounded up, and extended with Zmij's
compressed tables beyond 10^308.
write_shortest() converts its 16 digits in one vector register
(SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the
decimal point inside the register, avoiding a store-forwarding
stall that cost about 25% of the time to write a double. dump()
writes floats and integers straight into the serializer's write
buffer instead of copying them from a member buffer, and small
integers eight digits at a time. read_eight_bytes() and
parse_eight_digits() are marked always-inline, which GCC had been
calling out of line in the number-parsing loops.
Of one million random doubles, about 0.14% are now written with
different digits, always to a value that still reads back as the
same double.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>