* Unflatten in time and memory linear in the pointer depth
#5443 made unflatten() decide between arrays and objects independently
of the iteration order by collecting the pointer prefixes that have a
reference token 0 below them in a std::set<std::vector<string_t>>.
Every such prefix was stored as a copy of all its reference tokens, and
get_and_create() compared whole prefix vectors at every step, so
unflattening a pointer of depth d took time and memory quadratic in d:
a 10,000-level array pointer took 18 s and 1.3 GB, a 100,000-level one
did not finish.
The prefixes are now numbered nodes of a tree, so each is stored once
and get_and_create() follows the tree token by token. The result is
unchanged, including its independence of the iteration order.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Initialize prefix_tree members to satisfy -Weffc++
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move prefix_tree setup and child insertion into member functions
The constructor now creates the root node, add_child() inserts a
reference token below a prefix and returns the child's number, and
find_child() looks one up for get_and_create().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Flatten deeply nested values without recursing per nesting level
json_pointer::flatten() called itself once per nesting level, so
flatten() on a value nested deeply enough exhausted the call stack.
#5547 and #5548 fixed merge_patch() and diff() from #5393, but flatten()
was left out.
flatten() now walks the value with an explicit stack and keeps the path
in one buffer that grows and shrinks with it. It has a single code path
and no depth limit: the old version built a new path string per child,
so the iterative one is no slower on shallow values and much faster on
deep ones. The output, including the order of an ordered_json result,
is unchanged.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Construct flatten frames in place
Give the frame a constructor so both call sites can use emplace_back, as
suggested in the review; index starts at 0 for every frame.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep converted object keys alive while writing UBJSON and BJData
Since #5746, write_ubjson and write_ubjson_iterative pass each object key
to sanitize_utf8_for_write and keep the returned reference. When
object_t::key_type is not string_t but converts to it, the argument is a
temporary that is destroyed at the end of the statement, and the
function returns a reference to it in every case but a sanitized copy,
so the key bytes are read from a dead object (AddressSanitizer:
stack-use-after-scope). Default json and ordered_json are unaffected.
Bind the key to a named object_key_string_t first: a reference when
key_type is string_t, so no copy is added there, and a converted copy
otherwise. A deleted overload of sanitize_utf8_for_write for anything
other than string_t turns a recurrence into a compile error.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Suppress -Wunused-member-function for the converting_key test type
converting_key::data() is only called when JSON_DIAGNOSTICS is enabled.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy findings in the UBJSON/BJData converted-key fix
Suppress hicpp/modernize-use-equals-delete on the deleted
sanitize_utf8_for_write overload: it guards a private helper and must stay
private. Replace the C-style array in the new test with std::array.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The check rejects inputs of 0xFFFFFFF0 bytes or more, but the exception
message and the documentation said 4 GiB. Name the limit once
(max_input_size), and state 4 GiB minus 16 bytes in the message and the
documentation. Test the limit with a container that only claims the size.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
reserve(n) and the growth of the index computed n * sizeof(node) without a
check, which wraps around on 32-bit targets for inputs of about 1 GiB and
allocates a too small array. Throw std::bad_alloc for a count beyond the
address space and clamp the growth step to it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-little-endian path of the node writer wrote the integer's native
word over len and next, so len got the high half there. Compose and split
the value explicitly (len is the low half, next the high half); the
little-endian path stays a plain memcpy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
auto v = json_document::parse(text).root() compiled and left the view
dangling. Delete the overload for rvalue documents, take a named document
in the tests, and document the lifetime rule.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
explicit operator bool meant "refers to a value", which silently differs
from what a basic_json converts to. Use !v.is_discarded() instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json_document::parse(std::move(const_string)) failed to compile with
"no matching read_kind": only a non-const rvalue std::string can be moved
from. Treat a const rvalue as a copied byte container.
parse() does not accept everything BasicJsonType::parse() does: a FILE*
and pointers to or arrays of wide characters are rejected at compile time.
List the supported inputs instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
parse(ptr, len) compiled: len converted to allow_exceptions, and ptr was
read as a C string, past the end of a buffer without a terminating NUL.
Delete the overloads of parse, parse_copy, accept, and read that take an
integer other than bool where the flags are expected.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The non-template overloads for double asserted binary64 doubles wherever json.hpp was included, so the library no longer compiled where double is not IEEE 754 binary64 (AVR, -fshort-double). A trait now picks Zmij for any binary64 type, including a long double of that format (MSVC, Apple Arm), and Grisu2 for the others. Remove the unused write_short_decimal, powers_of_ten_16, zmij::decimal, zmij::to_decimal, and shortest_digits(double).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* avoid allocating temporary basic_json for CBOR and MessagePack object keys
Signed-off-by: alexprabhat99 <alexpbara@gmail.com>
* add size() to the custom object key test type
UBJSON and BJData access object keys through size() and c_str()
directly, so the key type now provides both and the comment says why.
Signed-off-by: alexprabhat99 <alexpbara@gmail.com>
* address review: drop key size()/c_str(), test keys below the depth limit
Nothing in the library calls size() or c_str() on an object key, so the
test key type only keeps data(), which JSON_DIAGNOSTICS needs.
The CBOR and MessagePack custom key tests now also nest objects deeper
than detail::recursion_depth_limit(), so keys written by
write_cbor_iterative and write_msgpack_iterative are covered as well.
Signed-off-by: alexprabhat99 <alexpbara@gmail.com>
---------
Signed-off-by: alexprabhat99 <alexpbara@gmail.com>
Drop the empty braced NSDMIs of the std::string members: old Clang
rejects the defaulted constructor when it is used by a member
initializer before the end of the class definition.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- pow5_table.hpp: pow5_128_largest_power was unused in this branch's
own code (GCC -Werror=unused-const-variable); tie it to the table
size with a static_assert instead of removing it, since a later
branch in the stack (json-view/23-zmij) uses it.
- number_parse.hpp: rename the local variable `copy` to `buffer` to
satisfy cpplint's build/include_what_you_use check.
- unit-class_lexer.cpp: extend the NOLINT list on the seeded mt19937
with bugprone-random-generator-seed, and parenthesize
`8 * sizeof(Bits) - 1` for clang-tidy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make std::hash<basic_json> consistent with operator== for numbers
operator== converts between number_integer, number_unsigned, and
number_float before comparing, so json(0), json(0U), and json(0.0)
all compare equal. hash() folded the specific value_t into the
result for each of the three numeric cases, giving each a distinct
hash and breaking the standard Hash requirement that a == b implies
hash(a) == hash(b). A std::unordered_set could therefore hold all
three as separate elements even though they compare equal.
hash() now treats all three numeric variants the same way: it
converts the value to number_float_t and combines it with a single
shared type tag, so any two numbers operator== considers equal hash
identically regardless of which internal type actually holds them.
Updated the accompanying test to check this consistency directly
(including via an actual unordered_set) instead of asserting that 0,
0U, and 0.0 hash differently, since that assumption was the bug.
Also corrected the function's own doc comment and the std::hash API
docs, which described the old behavior as intended.
Fixes#5400
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove now-unused number_integer_t/number_unsigned_t typedefs in hash()
Merging the three numeric branches into one that only reads
number_float_t left these two aliases unused, which several CI
configurations treat as a build error under -Wunused-local-typedefs.
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mark the unordered_set in the hash regression test const
clang-tidy's misc-const-correctness check flagged it: the set is
never mutated after construction, only read via size().
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Normalize -0.0 in number hashes and test range ends
operator== compares numbers exactly since #5459, so equal numbers
share one value and convert to the same number_float_t. Update the
comment accordingly, map -0.0 to 0.0 before hashing (std::hash need
not do that), and test -0.0 and the ends of the integer ranges. Show
hash(0.0) in the docs example and note the change in the version
history.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Clarify hash documentation after review
- Say "may hash differently" for null, false, and numbers, since a
collision across types is possible.
- Name the storage types (signed integer, unsigned integer,
floating-point number) instead of example literals.
- Explain that the hash survives converting an integer to
number_float_t but not the lossy conversion back, and that unequal
numbers may share a hash.
- State that the example hash values are illustrative only and vary by
platform, compiler, compiler version, and library version.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Afonso Januário <afonso-januario@hotmail.com>
* Stop binary writers overflowing the stack on deep values
to_cbor, to_msgpack, and to_ubjson recurse once per nesting level.
The parser is iterative, so a value the library accepts can crash on
the way back out.
Keep the existing recursive path for the first 128 levels and finish
anything deeper on a heap stack. Output is unchanged. BSON is left
alone because its extra size walk is a separate change.
Rebased onto the value-type output sink. The heap frames now initialize
every member, which is what -Weffc++ was rejecting.
See #5392.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
(cherry picked from commit cf65ac438f)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Redesign the iterative binary writers around a shared recursion depth limit
Address the open review on the non-recursive CBOR/MessagePack/UBJSON/BJData
writers (#5518):
- Delete the CBOR array/object prefix helpers; both the recursive and
iterative paths call write_cbor_head(), which already existed on develop.
- MessagePack: share one write_msgpack_array_prefix()/write_msgpack_object_prefix()
helper per container kind between the recursive and iterative paths, both
going through to_msgpack_length() so an over-long container throws
out_of_range.412 identically either way.
- Reuse detail::recursion_depth_limit() instead of a separate constant, the
same bound serializer::dump() and write_bson_document() already use.
- Redesign the frames after bson_frame/dump_frame: only a container with
elements is ever pushed, its header is written at the point it is pushed,
and the iterator is set in the frame's constructor instead of a
default-then-assign two-step with a since-removed "started" flag. The
UBJSON frame keeps only the value pointer, the per-element prefix_required
flag, and the iterator; write_closer and is_object are no longer stored,
since the former is always !use_count (use_count is constant for the whole
document) and the latter follows from value->is_object().
- Factor the BJData ND-array shape check into is_bjdata_ndarray(), used by
both the recursive object case and the iterative pushing logic.
- Give the frame classes the GCC -Weffc++ treatment already used for
diff_frame: a noexcept converting constructor plus the five special members
defaulted with no explicit noexcept.
- Fix two @ref self-references in write_cbor/write_msgpack/write_ubjson's own
doc comments to point at the public to_cbor/to_msgpack/to_ubjson/to_bjdata
API instead.
- The iterative object-key write for CBOR/MessagePack now runs the same
strict-mode check_utf8() against the parent object as diagnostics context
that the recursive path already ran, so the two paths raise identical
diagnostics across the switch-over.
- Rewrite the tests: round trips instead of a bare size check, byte-exact
comparisons against the recursive output at depths around the bound, a
deep object and a BJData ND-array past the bound, a deep discarded value
(type_error.321), and the OSS-Fuzz 566583014 CBOR/MessagePack regression.
BSON is unaffected by this change; it already walks its documents
iteratively and is covered separately by #5553.
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: ayush-singh-0601 <singhayush062006@gmail.com>
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).
A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.
The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.
basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.
detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.
A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Write doubles with the conversion of Zmij by Victor Zverovich (MIT),
ported to C++11 (detail/conversions/zmij.hpp). It finds the
shortest decimal that reads back as the same double, and the
closest one if there are several. Grisu2, used until now, is fast
but not always shortest: it sometimes writes a 17th digit where 16
suffice, or a last digit that is not the closest. The layout is
unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2.
Digits are converted eight at a time with the BCD conversion of
Xiang JunBo, as in Zmij, and written with one byte swap per eight
digits and fixed-size moves instead of per-digit loops. Leading and
trailing zeros are counted from those bytes. to_chars() uses a
local buffer when the caller's is shorter than the 41 bytes this
may write. The powers of ten come from the number-parsing table,
adjusted where it holds values rounded up, and extended with Zmij's
compressed tables beyond 10^308.
write_shortest() converts its 16 digits in one vector register
(SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the
decimal point inside the register, avoiding a store-forwarding
stall that cost about 25% of the time to write a double. dump()
writes floats and integers straight into the serializer's write
buffer instead of copying them from a member buffer, and small
integers eight digits at a time. read_eight_bytes() and
parse_eight_digits() are marked always-inline, which GCC had been
calling out of line in the number-parsing loops.
Of one million random doubles, about 0.14% are now written with
different digits, always to a value that still reads back as the
same double.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Give the library its own correctly rounded float converter for
binary32 and binary64 (IEEE 754), and speed up the lexer's string
and escape scanning.
The converter splits a number token into sign, significand, and
decimal exponent, then tries Clinger's fast path, then a templated
Eisel-Lemire step, and falls back to an exact big-integer digit
comparison for tokens with more than 19 significant digits whose two
candidate values round differently. This replaces std::from_chars
and strtod/strtof for both formats, so parsed values no longer
depend on the C/C++ library or the current locale. The strtold
fallback kept for other long double formats (x87, binary128) now
also copies a multi-byte decimal point correctly, fixing #5660.
eisel_lemire() and decimal_to_float() are always inlined so callers
keep the whole conversion in their hot loop.
The string-scanning kernels in string_scan.hpp find a stop byte with
the trailing-zero count of the SWAR mask instead of a byte loop, and
scalar_string_bulk_run() validates a run of multi-byte UTF-8
sequences one after another instead of re-searching after each one.
get_codepoint() decodes a contiguous \uXXXX escape with one table
lookup per byte instead of four range-checked get() calls; the
streaming path and all error positions are unchanged.
Adds 508 generated hard float-parsing cases with expected binary32
and binary64 bits, and kernel-comparison tests for the string scans
and the escape table against byte-by-byte references.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI on develop after #5585
- test-diagnostics-optimized: -O3 makes GCC's -Winline and
-Wsuggest-attribute=pure/const warnings fire with the ci_test_gcc flag
set; turn them off for this test.
- test-diagnostics-optimized: suppress Clang's -Wexit-time-destructors for
the static table in to_json.
- Infer: raise pulse-max-disjuncts from 20 to 40. With the default,
Pulse loses the stored type in basic_json::replace_value() and reports
false null dereferences of get_ptr() results in unit-pointer_access.cpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Ignore Infer's false USE_AFTER_DELETE in ordered_map::erase
Infer's std::string model keeps the buffer of a moved-from string, so the
destroy-and-reconstruct loop in erase(first, last) looks like it destroys a
buffer twice.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mark throw_on_discarded()'s parameters as used without exceptions
With JSON_NOEXCEPTION, JSON_THROW expands to std::abort(), so Clang's
-Wunused-parameter breaks test-disabled_exceptions (since #5761).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the span_input_adapter sax_parse checks with deleted deprecated functions
The #5676 regression test (#5740) calls the deprecated
sax_parse(span_input_adapter&&, ...), which JSON_DELETE_DEPRECATED_FUNCTIONS
deletes, so ci_test_delete_deprecated_functions failed to build.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fall back to the first entry in test-diagnostics-optimized's to_json
clang-tidy (clang-analyzer-security.ArrayBound) flagged it->second for a
value not in the table. Use the same fallback as
NLOHMANN_JSON_SERIALIZE_ENUM; the test still fails with -Werror=array-bounds
on the headers from before #5585.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy findings in tests from #5762 and #5774
- unit-regression2.cpp (#5762): const/auto for the destroy() test values;
NOLINT the intended copy in check_destroy_edge_case().
- unit-serialization.cpp (#5774): build the expected strings with += instead
of chained operator+ (performance-inefficient-string-concatenation).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
MSVC warns with C5260 that kAlpha and kGamma have internal linkage when
the header is used as a header unit or through a module. JSON_INLINE_VARIABLE
makes them inline variables from C++17 on.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_BRACE_INIT_COPY_SEMANTICS enabled, single-element brace
initialization from a JSON value decided whether to copy the value or
build an object by inspecting the value's runtime shape: a two-element
array whose first element is a string, such as ["key", 42], was turned
into an object instead of being copied. This made the behavior depend
on the element's content, and it did not distinguish an existing value
of this shape from a nested braced pair written in the source, such as
the inner {"key", "value"} of {{"key", "value"}}.
json_ref now records whether it was constructed from a braced list
(true only for the std::initializer_list<json_ref> constructor used
for nested braced lists) or from a value. The initializer-list
constructor uses this to copy or move a single non-braced-list element
before deciding whether the list describes an object, so a JSON value
is always copied regardless of its shape, while a braced pair written
in the source still creates an object.
Fixes#5662.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>