A view is tested with is_discarded(); the lookups in at(), value() and
contains(json_pointer) and the unit tests no longer rely on the explicit
conversion to bool.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The check rejects inputs of 0xFFFFFFF0 bytes or more, but the exception
message and the documentation said 4 GiB. Name the limit once
(max_input_size), and state 4 GiB minus 16 bytes in the message and the
documentation. Test the limit with a container that only claims the size.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
reserve(n) and the growth of the index computed n * sizeof(node) without a
check, which wraps around on 32-bit targets for inputs of about 1 GiB and
allocates a too small array. Throw std::bad_alloc for a count beyond the
address space and clamp the growth step to it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-little-endian path of the node writer wrote the integer's native
word over len and next, so len got the high half there. Compose and split
the value explicitly (len is the low half, next the high half); the
little-endian path stays a plain memcpy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
auto v = json_document::parse(text).root() compiled and left the view
dangling. Delete the overload for rvalue documents, take a named document
in the tests, and document the lifetime rule.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
explicit operator bool meant "refers to a value", which silently differs
from what a basic_json converts to. Use !v.is_discarded() instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json_document::parse(std::move(const_string)) failed to compile with
"no matching read_kind": only a non-const rvalue std::string can be moved
from. Treat a const rvalue as a copied byte container.
parse() does not accept everything BasicJsonType::parse() does: a FILE*
and pointers to or arrays of wide characters are rejected at compile time.
List the supported inputs instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Lookups (operator[], at, find, contains, count, value, JSON pointers)
return the last member of a duplicate key, as materialize() and
parse() keep it.
* operator[] on a discarded view returns a discarded view instead of
throwing, so v["a"]["b"] is safe for a missing "a".
* operator[] and at() take any integer type (not only int and size_t),
fixing ambiguous calls with unsigned, long, std::int64_t, ...
* Fix the operator[] documentation, which claimed a discarded view for
a type mismatch where type_error.305 is thrown.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
shrink_to_fit() now trims the tables of large objects like the node array and the decoded strings, and the list of large objects is released as soon as the tables are built.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The key hash is not seeded, so keys chosen to collide made building the table quadratic (20,000 colliding keys took 470 ms to parse). A key may now sit at most 64 slots from its home slot; if a key would sit further away, the table is dropped and the object is searched linearly. Lookups stop after the same distance.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The source extent of a value was read from the next node, falling back to the rest of the document when that node held a decoded string. dump() of a small value could thus allocate a buffer as large as the document. Skip a few such nodes, cap the fallback estimate, and shrink a buffer that is much larger than its output.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
parse(ptr, len) compiled: len converted to allow_exceptions, and ptr was
read as a C string, past the end of a buffer without a terminating NUL.
Delete the overloads of parse, parse_copy, accept, and read that take an
integer other than bool where the flags are expected.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-template overloads for double asserted binary64 doubles wherever json.hpp was included, so the library no longer compiled where double is not IEEE 754 binary64 (AVR, -fshort-double). A trait now picks Zmij for any binary64 type, including a long double of that format (MSVC, Apple Arm), and Grisu2 for the others. Remove the unused write_short_decimal, powers_of_ten_16, zmij::decimal, zmij::to_decimal, and shortest_digits(double).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Drop the empty braced NSDMIs of the std::string members: old Clang
rejects the defaulted constructor when it is used by a member
initializer before the end of the class definition.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC ignores the target attribute in modules, so the SSSE3 dispatch is
disabled for the module interface (the check stays portable, SSE2 is kept).
Make the 8-vs-16 byte unrolling condition in scan_string_run a
preprocessor/template split to avoid a constant condition (C4127).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- pow5_table.hpp: pow5_128_largest_power was unused in this branch's
own code (GCC -Werror=unused-const-variable); tie it to the table
size with a static_assert instead of removing it, since a later
branch in the stack (json-view/23-zmij) uses it.
- number_parse.hpp: rename the local variable `copy` to `buffer` to
satisfy cpplint's build/include_what_you_use check.
- unit-class_lexer.cpp: extend the NOLINT list on the seeded mt19937
with bugprone-random-generator-seed, and parenthesize
`8 * sizeof(Bits) - 1` for clang-tidy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make std::hash<basic_json> consistent with operator== for numbers
operator== converts between number_integer, number_unsigned, and
number_float before comparing, so json(0), json(0U), and json(0.0)
all compare equal. hash() folded the specific value_t into the
result for each of the three numeric cases, giving each a distinct
hash and breaking the standard Hash requirement that a == b implies
hash(a) == hash(b). A std::unordered_set could therefore hold all
three as separate elements even though they compare equal.
hash() now treats all three numeric variants the same way: it
converts the value to number_float_t and combines it with a single
shared type tag, so any two numbers operator== considers equal hash
identically regardless of which internal type actually holds them.
Updated the accompanying test to check this consistency directly
(including via an actual unordered_set) instead of asserting that 0,
0U, and 0.0 hash differently, since that assumption was the bug.
Also corrected the function's own doc comment and the std::hash API
docs, which described the old behavior as intended.
Fixes#5400
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove now-unused number_integer_t/number_unsigned_t typedefs in hash()
Merging the three numeric branches into one that only reads
number_float_t left these two aliases unused, which several CI
configurations treat as a build error under -Wunused-local-typedefs.
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mark the unordered_set in the hash regression test const
clang-tidy's misc-const-correctness check flagged it: the set is
never mutated after construction, only read via size().
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Normalize -0.0 in number hashes and test range ends
operator== compares numbers exactly since #5459, so equal numbers
share one value and convert to the same number_float_t. Update the
comment accordingly, map -0.0 to 0.0 before hashing (std::hash need
not do that), and test -0.0 and the ends of the integer ranges. Show
hash(0.0) in the docs example and note the change in the version
history.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Clarify hash documentation after review
- Say "may hash differently" for null, false, and numbers, since a
collision across types is possible.
- Name the storage types (signed integer, unsigned integer,
floating-point number) instead of example literals.
- Explain that the hash survives converting an integer to
number_float_t but not the lossy conversion back, and that unequal
numbers may share a hash.
- State that the example hash values are illustrative only and vary by
platform, compiler, compiler version, and library version.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Afonso Januário <afonso-januario@hotmail.com>
* Stop binary writers overflowing the stack on deep values
to_cbor, to_msgpack, and to_ubjson recurse once per nesting level.
The parser is iterative, so a value the library accepts can crash on
the way back out.
Keep the existing recursive path for the first 128 levels and finish
anything deeper on a heap stack. Output is unchanged. BSON is left
alone because its extra size walk is a separate change.
Rebased onto the value-type output sink. The heap frames now initialize
every member, which is what -Weffc++ was rejecting.
See #5392.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
(cherry picked from commit cf65ac438f)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Redesign the iterative binary writers around a shared recursion depth limit
Address the open review on the non-recursive CBOR/MessagePack/UBJSON/BJData
writers (#5518):
- Delete the CBOR array/object prefix helpers; both the recursive and
iterative paths call write_cbor_head(), which already existed on develop.
- MessagePack: share one write_msgpack_array_prefix()/write_msgpack_object_prefix()
helper per container kind between the recursive and iterative paths, both
going through to_msgpack_length() so an over-long container throws
out_of_range.412 identically either way.
- Reuse detail::recursion_depth_limit() instead of a separate constant, the
same bound serializer::dump() and write_bson_document() already use.
- Redesign the frames after bson_frame/dump_frame: only a container with
elements is ever pushed, its header is written at the point it is pushed,
and the iterator is set in the frame's constructor instead of a
default-then-assign two-step with a since-removed "started" flag. The
UBJSON frame keeps only the value pointer, the per-element prefix_required
flag, and the iterator; write_closer and is_object are no longer stored,
since the former is always !use_count (use_count is constant for the whole
document) and the latter follows from value->is_object().
- Factor the BJData ND-array shape check into is_bjdata_ndarray(), used by
both the recursive object case and the iterative pushing logic.
- Give the frame classes the GCC -Weffc++ treatment already used for
diff_frame: a noexcept converting constructor plus the five special members
defaulted with no explicit noexcept.
- Fix two @ref self-references in write_cbor/write_msgpack/write_ubjson's own
doc comments to point at the public to_cbor/to_msgpack/to_ubjson/to_bjdata
API instead.
- The iterative object-key write for CBOR/MessagePack now runs the same
strict-mode check_utf8() against the parent object as diagnostics context
that the recursive path already ran, so the two paths raise identical
diagnostics across the switch-over.
- Rewrite the tests: round trips instead of a bare size check, byte-exact
comparisons against the recursive output at depths around the bound, a
deep object and a BJData ND-array past the bound, a deep discarded value
(type_error.321), and the OSS-Fuzz 566583014 CBOR/MessagePack regression.
BSON is unaffected by this change; it already walks its documents
iteratively and is covered separately by #5553.
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: ayush-singh-0601 <singhayush062006@gmail.com>
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.
Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.
On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.
Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.
simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>