Silence bugprone-casting-through-void for the SSE loads (a reinterpret_cast
would trip -Wcast-align=strict), use auto for a cast initialiser, add
parentheses to a mixed expression, and name bugprone-std-namespace-modification
in the NOLINTs of the tuple_size/tuple_element specialisations.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Mark the includes that are needed by design (json.hpp, the macro
unscope, <tuple> for the tuple_size specialization), take size_t from
<cstring>, and drop the forward declaration of basic_json_document that
the friend declaration makes redundant.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
__forceinline makes MSVC report C4714 for function templates it cannot
inline (get<std::string>), an error under /WX. The forced inlining is
only a performance hint, so use plain inline there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more return the first member of a
repeated key again, like the linear search of smaller objects. This undoes
the code change of a69542046; its test now expects the first member from
lookups (with and without a table) and the last value from materialize()
and basic_json::parse().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
finish() copied the whole output when the buffer was more than twice as
large as the result (citm dump +11%). The tighter source_extent()
estimate already keeps the buffer of a small value small.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Looking up the last member of a duplicate key cannot stop at a match, so
every lookup scanned the whole object (1.6 to 3.4 times slower for small
objects). operator[](key), at, find, value, contains, count, and JSON
pointer resolution return the first member again, as yyjson and simdjson
do; materialize() and get<map>() keep the last value, as parse().
The documentation says so in the feature page and on each lookup page,
and explains how to get the value parse() would give. The integer index
templates and the discarded chaining of operator[] stay.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Unflatten in time and memory linear in the pointer depth
#5443 made unflatten() decide between arrays and objects independently
of the iteration order by collecting the pointer prefixes that have a
reference token 0 below them in a std::set<std::vector<string_t>>.
Every such prefix was stored as a copy of all its reference tokens, and
get_and_create() compared whole prefix vectors at every step, so
unflattening a pointer of depth d took time and memory quadratic in d:
a 10,000-level array pointer took 18 s and 1.3 GB, a 100,000-level one
did not finish.
The prefixes are now numbered nodes of a tree, so each is stored once
and get_and_create() follows the tree token by token. The result is
unchanged, including its independence of the iteration order.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Initialize prefix_tree members to satisfy -Weffc++
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move prefix_tree setup and child insertion into member functions
The constructor now creates the root node, add_child() inserts a
reference token below a prefix and returns the child's number, and
find_child() looks one up for get_and_create().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Lookups in objects with 128 members or more now return the last member of
a repeated key, like the linear search of smaller objects and like
materialize() and basic_json::parse(). build_object_index let the first
occurrence win, so the same text gave different results depending on the
size of the object.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Flatten deeply nested values without recursing per nesting level
json_pointer::flatten() called itself once per nesting level, so
flatten() on a value nested deeply enough exhausted the call stack.
#5547 and #5548 fixed merge_patch() and diff() from #5393, but flatten()
was left out.
flatten() now walks the value with an explicit stack and keeps the path
in one buffer that grows and shrinks with it. It has a single code path
and no depth limit: the old version built a new path string per child,
so the iterative one is no slower on shallow values and much faster on
deep ones. The output, including the order of an ordered_json result,
is unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Construct flatten frames in place
Give the frame a constructor so both call sites can use emplace_back, as
suggested in the review; index starts at 0 for every frame.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep converted object keys alive while writing UBJSON and BJData
Since #5746, write_ubjson and write_ubjson_iterative pass each object key
to sanitize_utf8_for_write and keep the returned reference. When
object_t::key_type is not string_t but converts to it, the argument is a
temporary that is destroyed at the end of the statement, and the
function returns a reference to it in every case but a sanitized copy,
so the key bytes are read from a dead object (AddressSanitizer:
stack-use-after-scope). Default json and ordered_json are unaffected.
Bind the key to a named object_key_string_t first: a reference when
key_type is string_t, so no copy is added there, and a converted copy
otherwise. A deleted overload of sanitize_utf8_for_write for anything
other than string_t turns a recurrence into a compile error.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Suppress -Wunused-member-function for the converting_key test type
converting_key::data() is only called when JSON_DIAGNOSTICS is enabled.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy findings in the UBJSON/BJData converted-key fix
Suppress hicpp/modernize-use-equals-delete on the deleted
sanitize_utf8_for_write overload: it guards a private helper and must stay
private. Replace the C-style array in the new test with std::array.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A view is tested with is_discarded(); the lookups in at(), value() and
contains(json_pointer) and the unit tests no longer rely on the explicit
conversion to bool.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The check rejects inputs of 0xFFFFFFF0 bytes or more, but the exception
message and the documentation said 4 GiB. Name the limit once
(max_input_size), and state 4 GiB minus 16 bytes in the message and the
documentation. Test the limit with a container that only claims the size.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
reserve(n) and the growth of the index computed n * sizeof(node) without a
check, which wraps around on 32-bit targets for inputs of about 1 GiB and
allocates a too small array. Throw std::bad_alloc for a count beyond the
address space and clamp the growth step to it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-little-endian path of the node writer wrote the integer's native
word over len and next, so len got the high half there. Compose and split
the value explicitly (len is the low half, next the high half); the
little-endian path stays a plain memcpy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
auto v = json_document::parse(text).root() compiled and left the view
dangling. Delete the overload for rvalue documents, take a named document
in the tests, and document the lifetime rule.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
explicit operator bool meant "refers to a value", which silently differs
from what a basic_json converts to. Use !v.is_discarded() instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json_document::parse(std::move(const_string)) failed to compile with
"no matching read_kind": only a non-const rvalue std::string can be moved
from. Treat a const rvalue as a copied byte container.
parse() does not accept everything BasicJsonType::parse() does: a FILE*
and pointers to or arrays of wide characters are rejected at compile time.
List the supported inputs instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Lookups (operator[], at, find, contains, count, value, JSON pointers)
return the last member of a duplicate key, as materialize() and
parse() keep it.
* operator[] on a discarded view returns a discarded view instead of
throwing, so v["a"]["b"] is safe for a missing "a".
* operator[] and at() take any integer type (not only int and size_t),
fixing ambiguous calls with unsigned, long, std::int64_t, ...
* Fix the operator[] documentation, which claimed a discarded view for
a type mismatch where type_error.305 is thrown.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
shrink_to_fit() now trims the tables of large objects like the node array and the decoded strings, and the list of large objects is released as soon as the tables are built.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The key hash is not seeded, so keys chosen to collide made building the table quadratic (20,000 colliding keys took 470 ms to parse). A key may now sit at most 64 slots from its home slot; if a key would sit further away, the table is dropped and the object is searched linearly. Lookups stop after the same distance.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>