Replacing an array or object that spans several nodes by a scalar looks
up its parent from the root (find_parent), which is linear in the size
of the document: setting every element of a 20,000 element array took
more than half a second. The parent only has to switch to links once;
afterward the extent of the replaced value no longer matters. Nodes that
an entry of a moved sequence links to are now marked (node_flags::linked),
and assign() skips the lookup for them.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
append_text() doubled the capacity of the arena and rejected the result
if it exceeded 4 GiB, so that an arena of more than 2 GiB could not grow
even though the 32-bit offsets of nodes address 4 GiB - 1 bytes. The
capacity is now clamped to that limit (text_capacity()), and an append
is rejected only if the bytes themselves do not fit.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An editable document only holds valid UTF-8, but copy_scalar() copied the
strings and keys of a view of another document unchecked. A document of
a weaker check (or a borrowed text that changed after parsing) could
therefore bring ill-formed UTF-8 into it. The copy is checked now, with
the error that dump() reports for the string; copies within the same
document stay unchecked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
assign() rewrote the kind, length and extent of a slot before set_moved()
registered its new element sequence. set_moved() can throw
(std::bad_alloc from reserving the bookkeeping vectors) for a slot that
is not moved yet, which left a container without the moved flag that
showed its old children. The bookkeeping is now reserved first
(reserve_moved()), so that nothing after the first write to the slot can
throw. block_of() reserves before it marks anything for the same reason.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Counting and copying the nodes of a view or a basic_json value into an
editable document recursed once per nesting level, so that a deeply
nested value overflowed the stack. The four functions now walk the value
with an explicit stack, as materialize() does.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Value-initialize the const std::less in find_parent (clang 3.4/3.6 do not
implement DR 253), and test the Editable template argument through a
function to avoid MSVC C4127.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_document gets a second template parameter, Editable
(false by default), plus the aliases json_editable_document,
json_editable_view, ordered_json_editable_document and
ordered_json_editable_view.
Editable documents can change values and structure without
rewriting the source text: set()/push_back() on values, keys,
array indices and JSON pointers; insert() before an array
element; erase() of an object key, array index or JSON pointer.
New values and element sequences go into edit storage that the
document owns and never moves, so views keep referring to their
value across edits and a parsed node never moves. Read-only
documents walk the plain node array and are unaffected.
Strings are checked for UTF-8 on entry, so dump() of an editable
document never throws type_error.316. Binary values cannot be
stored (type_error.319).
A seeded differential test applies random edits to an editable
document and to the equivalent ordered_json and compares both
after every step.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Silence bugprone-casting-through-void for the SSE loads (a reinterpret_cast
would trip -Wcast-align=strict), use auto for a cast initialiser, add
parentheses to a mixed expression, and name bugprone-std-namespace-modification
in the NOLINTs of the tuple_size/tuple_element specialisations.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more return the first member of a
repeated key again, like the linear search of smaller objects. This undoes
the code change of a69542046; its test now expects the first member from
lookups (with and without a table) and the last value from materialize()
and basic_json::parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more now return the last member of
a repeated key, like the linear search of smaller objects and like
materialize() and basic_json::parse(). build_object_index let the first
occurrence win, so the same text gave different results depending on the
size of the object.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The key hash is not seeded, so keys chosen to collide made building the table quadratic (20,000 colliding keys took 470 ms to parse). A key may now sit at most 64 slots from its home slot; if a key would sit further away, the table is dropped and the object is searched linearly. Lookups stop after the same distance.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC ignores the target attribute in modules, so the SSSE3 dispatch is
disabled for the module interface (the check stays portable, SSE2 is kept).
Make the 8-vs-16 byte unrolling condition in scan_string_run a
preprocessor/template split to avoid a constant condition (C4127).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.
Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.
On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.
Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.
simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
finish() copied the whole output when the buffer was more than twice as
large as the result (citm dump +11%). The tighter source_extent()
estimate already keeps the buffer of a small value small.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The source extent of a value was read from the next node, falling back to the rest of the document when that node held a decoded string. dump() of a small value could thus allocate a buffer as large as the document. Skip a few such nodes, cap the fallback estimate, and shrink a buffer that is much larger than its output.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Add basic_json_view::dump() and the comparison operators, and read
floats from the parser's digit layout instead of rescanning the
token.
dump(indent, indent_char, ensure_ascii, number_format) writes a
value the way ordered_json::parse(text).dump() writes it for the
same arguments: members in document order, all of them should a
key occur more than once; strings escaped by the same rules, using
the library's scanning kernels; floats written with the library's
to_chars conversion, so the output equals basic_json's byte for
byte; integers copied from the source, where they are already
canonical, except -0, which parse() reads as 0. There is no
error_handler argument, because the view only holds valid UTF-8.
number_format::source copies numbers exactly as they appear in the
source (e.g. "1.50", "1E2", "-0"), which basic_json cannot provide.
operator<< takes the indentation from the stream width, as for
basic_json. The writer walks iteratively, so nesting depth is
limited by memory only.
operator== and operator!= compare two views, or a view and a
basic_json value in either order, by the rules basic_json's
operator== uses: numbers compare by value across their types,
objects compare by their members with duplicate keys resolved as
parse() resolves them, member order matters only where the object
type keeps one, and discarded views compare as discarded basic_json
values do, including under JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON.
Nothing is materialized except single scalars.
While parsing, the view now records where the integer digits, the
fraction digits, and the exponent of a float token are, so floats
and doubles with at most 19 digits are read from that layout with
the library's decimal_to_float() instead of rescanning the token.
Both round correctly, so the values are those of parse(). get<double>(),
materialize(), dump(), and the comparisons all use it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The same changes as on the dump branch, where they were first made, so
that this branch passes clang-tidy on its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
__forceinline makes MSVC report C4714 for function templates it cannot
inline (get<std::string>), an error under /WX. The forced inlining is
only a performance hint, so use plain inline there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Looking up the last member of a duplicate key cannot stop at a match, so
every lookup scanned the whole object (1.6 to 3.4 times slower for small
objects). operator[](key), at, find, value, contains, count, and JSON
pointer resolution return the first member again, as yyjson and simdjson
do; materialize() and get<map>() keep the last value, as parse().
The documentation says so in the feature page and on each lookup page,
and explains how to get the value parse() would give. The integer index
templates and the discarded chaining of operator[] stay.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Lookups (operator[], at, find, contains, count, value, JSON pointers)
return the last member of a duplicate key, as materialize() and
parse() keep it.
* operator[] on a discarded view returns a discarded view instead of
throwing, so v["a"]["b"] is safe for a missing "a".
* operator[] and at() take any integer type (not only int and size_t),
fixing ambiguous calls with unsigned, long, std::int64_t, ...
* Fix the operator[] documentation, which claimed a discarded view for
a type mismatch where type_error.305 is thrown.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Give basic_json_view the read-only access functions of basic_json:
operator[] and at() with keys and indices, front()/back(), find(),
contains(), count(), begin()/end() and cbegin()/cend(), items()
with structured bindings from C++17 on, and type_name().
Exceptions have the ids and messages of the const functions of
basic_json. Where basic_json has undefined behavior the view
answers safely: operator[] with a missing key or an out-of-range
index returns a discarded view, and front()/back() of an empty
container throw invalid_iterator.214. Objects are iterated in
document order, and all members are visited; duplicate-key lookups
find the first member (as yyjson and simdjson do), while parse(),
materialize(), and the map conversions keep the last value, as
parse() does. Keys of up to 16 bytes are compared with two
overlapping loads.
Add value conversions: get<T>()/get_to() for arithmetic types,
bool, nullptr_t, strings (std::basic_string copied,
string_view_t without a copy), BasicJsonType, views, std::vector,
and maps with string keys; get_string() for the string without a
copy; number_token() for the number exactly as written in the
source; value() with keys and JSON pointers; and operator[]/at()/
contains() with JSON pointers. Everything else, including types
with from_json(), goes through materialize() of that subtree.
get<T>() of arithmetic types is inlined down to the conversion, so
reading an integer needs no call.
detail::json_pointer_access exposes a pointer's reference tokens
to code outside basic_json.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
emit() moves the node array when it grows, and the subtraction read base
in the same expression, so the order was unspecified. MSVC Release builds
without forced inlining read the old base; the container index then
pointed outside the array and close() wrote out of bounds.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The check rejects inputs of 0xFFFFFFF0 bytes or more, but the exception
message and the documentation said 4 GiB. Name the limit once
(max_input_size), and state 4 GiB minus 16 bytes in the message and the
documentation. Test the limit with a container that only claims the size.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
reserve(n) and the growth of the index computed n * sizeof(node) without a
check, which wraps around on 32-bit targets for inputs of about 1 GiB and
allocates a too small array. Throw std::bad_alloc for a count beyond the
address space and clamp the growth step to it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-little-endian path of the node writer wrote the integer's native
word over len and next, so len got the high half there. Compose and split
the value explicitly (len is the low half, next the high half); the
little-endian path stays a plain memcpy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json_document::parse(std::move(const_string)) failed to compile with
"no matching read_kind": only a non-const rvalue std::string can be moved
from. Treat a const rvalue as a copied byte container.
parse() does not accept everything BasicJsonType::parse() does: a FILE*
and pointers to or arrays of wide characters are rejected at compile time.
List the supported inputs instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
parse(ptr, len) compiled: len converted to allow_exceptions, and ptr was
read as a C string, past the end of a buffer without a terminating NUL.
Delete the overloads of parse, parse_copy, accept, and read that take an
integer other than bool where the flags are expected.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Drop the empty braced NSDMIs of the std::string members: old Clang
rejects the defaulted constructor when it is used by a member
initializer before the end of the class definition.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).
A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.
The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.
basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.
detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.
A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-template overloads for double asserted binary64 doubles wherever json.hpp was included, so the library no longer compiled where double is not IEEE 754 binary64 (AVR, -fshort-double). A trait now picks Zmij for any binary64 type, including a long double of that format (MSVC, Apple Arm), and Grisu2 for the others. Remove the unused write_short_decimal, powers_of_ten_16, zmij::decimal, zmij::to_decimal, and shortest_digits(double).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Write doubles with the conversion of Zmij by Victor Zverovich (MIT),
ported to C++11 (detail/conversions/zmij.hpp). It finds the
shortest decimal that reads back as the same double, and the
closest one if there are several. Grisu2, used until now, is fast
but not always shortest: it sometimes writes a 17th digit where 16
suffice, or a last digit that is not the closest. The layout is
unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2.
Digits are converted eight at a time with the BCD conversion of
Xiang JunBo, as in Zmij, and written with one byte swap per eight
digits and fixed-size moves instead of per-digit loops. Leading and
trailing zeros are counted from those bytes. to_chars() uses a
local buffer when the caller's is shorter than the 41 bytes this
may write. The powers of ten come from the number-parsing table,
adjusted where it holds values rounded up, and extended with Zmij's
compressed tables beyond 10^308.
write_shortest() converts its 16 digits in one vector register
(SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the
decimal point inside the register, avoiding a store-forwarding
stall that cost about 25% of the time to write a double. dump()
writes floats and integers straight into the serializer's write
buffer instead of copying them from a member buffer, and small
integers eight digits at a time. read_eight_bytes() and
parse_eight_digits() are marked always-inline, which GCC had been
calling out of line in the number-parsing loops.
Of one million random doubles, about 0.14% are now written with
different digits, always to a value that still reads back as the
same double.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- pow5_table.hpp: pow5_128_largest_power was unused in this branch's
own code (GCC -Werror=unused-const-variable); tie it to the table
size with a static_assert instead of removing it, since a later
branch in the stack (json-view/23-zmij) uses it.
- number_parse.hpp: rename the local variable `copy` to `buffer` to
satisfy cpplint's build/include_what_you_use check.
- unit-class_lexer.cpp: extend the NOLINT list on the seeded mt19937
with bugprone-random-generator-seed, and parenthesize
`8 * sizeof(Bits) - 1` for clang-tidy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Give the library its own correctly rounded float converter for
binary32 and binary64 (IEEE 754), and speed up the lexer's string
and escape scanning.
The converter splits a number token into sign, significand, and
decimal exponent, then tries Clinger's fast path, then a templated
Eisel-Lemire step, and falls back to an exact big-integer digit
comparison for tokens with more than 19 significant digits whose two
candidate values round differently. This replaces std::from_chars
and strtod/strtof for both formats, so parsed values no longer
depend on the C/C++ library or the current locale. The strtold
fallback kept for other long double formats (x87, binary128) now
also copies a multi-byte decimal point correctly, fixing #5660.
eisel_lemire() and decimal_to_float() are always inlined so callers
keep the whole conversion in their hot loop.
The string-scanning kernels in string_scan.hpp find a stop byte with
the trailing-zero count of the SWAR mask instead of a byte loop, and
scalar_string_bulk_run() validates a run of multi-byte UTF-8
sequences one after another instead of re-searching after each one.
get_codepoint() decodes a contiguous \uXXXX escape with one table
lookup per byte instead of four range-checked get() calls; the
streaming path and all error positions are unchanged.
Adds 508 generated hard float-parsing cases with expected binary32
and binary64 bits, and kernel-comparison tests for the string scans
and the escape table against byte-by-byte references.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop the BON8 writer overflowing the stack on deep values
to_bon8() recursed once per nesting level, so a value the iterative
BON8 reader accepts (e.g. ~24k nested one-element arrays) crashed on
the way back out. #5781 bounded the CBOR, MessagePack, and UBJSON/BJData
writers, but BON8 was merged before it and was not covered.
Apply the same scheme: recurse for the first recursion_depth_limit()
levels, then finish the value with write_bon8_iterative, which keeps
the open containers on a heap stack (reusing binary_container_frame)
and writes the 0xFE closer when it leaves a container with more than
four elements. The output is byte-for-byte unchanged.
Fixes https://issues.oss-fuzz.com/issues/572238015
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Merge the array and object branches of the BON8 iterative writer
write_bon8_iterative and write_bon8_value_or_push handled arrays and
objects in separate branches that repeated the end-of-container check,
the 0xFE closer and the marker computation. Share those parts and branch
only where arrays and objects really differ. The output is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Hz7VJi1FTKr6gpseLErbbS
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Restore v3.12.0 support for custom object key types
Custom object_t types whose key_type is not string_t compiled with
v3.12.0 for several APIs that unreleased changes broke:
- to_bson failed for every custom key type (#5553 kept a const string_t*
to the key); the nested entry's header is now written where the entry
is found.
- Copying deep values (and parse, merge_patch, update, insert) required
operator== on keys (#5389); keys without one are now paired via find().
- to_cbor/to_msgpack required an implicit conversion to string_t (#5746,
#5328); keys without one go through a temporary basic_json again.
- at() required a conversion to string_t for its error message (#5727);
other keys are passed to concat() unchanged again.
The new unit-custom-object-key-type.cpp covers five key types with
different capabilities.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the with_object_t alias for the custom object key test types
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Avoid floating-point equality in custom key type test
GCC with -Werror=float-equal rejects comparing the double value with ==.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Write the head of nested BSON elements in one helper
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Suppress bugprone-return-const-ref-from-parameter in key_for_message
The reference is only passed to concat() within the full-expression that
holds the key, like the similar helpers in binary_writer.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make the value of the custom test key types private
clang-tidy (cppcoreguidelines-non-private-member-variables-in-classes)
rejects the protected member; the derived key types use a protected
accessor instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Pass keys with data() and size() unchanged into the at() miss message
key_for_message() converted every key that string_t can be constructed
from, so a miss on a string_t or string_view key copied it before
concat() copied it again. Keys that concat() can append through data()
and size() are now passed through; only other keys (string literals,
key types that just convert to string_t) are converted.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
write_bon8_value silently skipped discarded values, but the array/object
count marker still counted them, so [1, discarded] produced 82 91: a marker
announcing two elements followed by one. Throw type_error.321 like the other
binary writers do, at any nesting level.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Report one-character non-numeric array indices like longer ones
A JSON pointer reference token that is not a number but has only one
character (e.g. "/a/x") was reported as out_of_range.404 ("unresolved
reference token"), because the "is not a number" check only ran for
tokens longer than one character; "/a/xy" got parse_error.109. Both now
throw parse_error.109. "-" and the empty token are still reported as
out_of_range.404. As a consequence, value(json_pointer, default) on an
array now throws for "/x" as it already did for "/xy".
Also document why ordered_map::erase's destroy/placement-new loop on
pair<const Key, T> is kept despite [basic.life]/8 before C++20.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document the one-character array index change
Add 3.13.0 version-history entries to at, operator[], value, patch,
patch_inplace, and unflatten, and describe in exceptions.md which array
indices throw parse_error.109 and which out_of_range.404.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use and extend the detail helpers to remove duplicated code
Library:
- binary_reader: format bytes with hex_byte() instead of snprintf
- add throw_type_must_be() for the 20 copies of type_error.302
- binary_reader: add last_byte_error()/unexpected_byte() for the 32
"parse error at the last read byte" sites (replaces bon8_error)
- json_sax: add check_container_size() for out_of_range.408 and
diagnostic_positions::set_container_start/_end()
- json_pointer: add throw_no_parent() (405) and throw_unresolved() (404)
Tests:
- unit-class_parser uses the shared utils::SaxCountdown
- move SaxEventLogger (and its ExitAfter* variants) from unit-class_parser
and unit-deserialization into the new tests/src/test_sax.hpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Factor out more repeated error paths and test boilerplate
Library:
- add throw_cannot_use_with() for the 34 copies of type_error.304-312
"cannot use X with Y"
- iter_impl: add throw_cannot_get_value() (invalid_iterator.214)
- parser: add syntax_error() for the 12 parse_error.101 sites
- ordered_map: share the four at() bodies via at_impl()
- json_sax_dom_callback_parser: add pop_container() for end_object()
and end_array()
- binary_reader: build the two UBJSON/BJData length-type messages with
concat() and last_byte_error()
Tests:
- move same_value(), the NDEBUG guard, and step 0 (parse without
exceptions) of the seven fuzzer drivers into tests/src/fuzzer_common.hpp
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name the exception ids and address review comments
Add detail::exception_id, a scoped enum with one named enumerator per
documented exception id, and use it for every id in the library. The
create() functions get an overload for it; the int overloads stay for
user code.
Following the review of #5783: add binary_reader::invalid_byte() and
length_type_error(), basic_json::throw_subscript_wrong_type(), move the
fuzzer includes and the using-declaration into fuzzer_common.hpp, and
rename test_sax.hpp to sax_event_loggers.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Silence -Wweak-vtables for the SAX event loggers
The loggers moved from anonymous namespaces in the test files into
sax_event_loggers.hpp, so clang now warns that their vtables are
emitted in every translation unit.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Silence MSVC 2015 C4100 in parser::syntax_error for static SAX::parse_error
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Inline invalid_byte() into its call sites
The wrapper only fixed the message string of unexpected_byte(), which is
the same kind of per-argument helper that was declined for
throw_type_must_be(). Call unexpected_byte("invalid byte", ...) directly.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Unflatten in time and memory linear in the pointer depth
#5443 made unflatten() decide between arrays and objects independently
of the iteration order by collecting the pointer prefixes that have a
reference token 0 below them in a std::set<std::vector<string_t>>.
Every such prefix was stored as a copy of all its reference tokens, and
get_and_create() compared whole prefix vectors at every step, so
unflattening a pointer of depth d took time and memory quadratic in d:
a 10,000-level array pointer took 18 s and 1.3 GB, a 100,000-level one
did not finish.
The prefixes are now numbered nodes of a tree, so each is stored once
and get_and_create() follows the tree token by token. The result is
unchanged, including its independence of the iteration order.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Initialize prefix_tree members to satisfy -Weffc++
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move prefix_tree setup and child insertion into member functions
The constructor now creates the root node, add_child() inserts a
reference token below a prefix and returns the child's number, and
find_child() looks one up for get_and_create().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Flatten deeply nested values without recursing per nesting level
json_pointer::flatten() called itself once per nesting level, so
flatten() on a value nested deeply enough exhausted the call stack.
#5547 and #5548 fixed merge_patch() and diff() from #5393, but flatten()
was left out.
flatten() now walks the value with an explicit stack and keeps the path
in one buffer that grows and shrinks with it. It has a single code path
and no depth limit: the old version built a new path string per child,
so the iterative one is no slower on shallow values and much faster on
deep ones. The output, including the order of an ordered_json result,
is unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Construct flatten frames in place
Give the frame a constructor so both call sites can use emplace_back, as
suggested in the review; index starts at 0 for every frame.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>