Test the template options through enabled() so that conditions combined
with them are not constant, and mark the nav::value() null dereference
as a false positive like materialize.hpp does.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- pow5_table.hpp: pow5_128_largest_power was unused in this branch's
own code (GCC -Werror=unused-const-variable); tie it to the table
size with a static_assert instead of removing it, since a later
branch in the stack (json-view/23-zmij) uses it.
- number_parse.hpp: rename the local variable `copy` to `buffer` to
satisfy cpplint's build/include_what_you_use check.
- unit-class_lexer.cpp: extend the NOLINT list on the seeded mt19937
with bugprone-random-generator-seed, and parenthesize
`8 * sizeof(Bits) - 1` for clang-tidy.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make std::hash<basic_json> consistent with operator== for numbers
operator== converts between number_integer, number_unsigned, and
number_float before comparing, so json(0), json(0U), and json(0.0)
all compare equal. hash() folded the specific value_t into the
result for each of the three numeric cases, giving each a distinct
hash and breaking the standard Hash requirement that a == b implies
hash(a) == hash(b). A std::unordered_set could therefore hold all
three as separate elements even though they compare equal.
hash() now treats all three numeric variants the same way: it
converts the value to number_float_t and combines it with a single
shared type tag, so any two numbers operator== considers equal hash
identically regardless of which internal type actually holds them.
Updated the accompanying test to check this consistency directly
(including via an actual unordered_set) instead of asserting that 0,
0U, and 0.0 hash differently, since that assumption was the bug.
Also corrected the function's own doc comment and the std::hash API
docs, which described the old behavior as intended.
Fixes#5400
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove now-unused number_integer_t/number_unsigned_t typedefs in hash()
Merging the three numeric branches into one that only reads
number_float_t left these two aliases unused, which several CI
configurations treat as a build error under -Wunused-local-typedefs.
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mark the unordered_set in the hash regression test const
clang-tidy's misc-const-correctness check flagged it: the set is
never mutated after construction, only read via size().
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Normalize -0.0 in number hashes and test range ends
operator== compares numbers exactly since #5459, so equal numbers
share one value and convert to the same number_float_t. Update the
comment accordingly, map -0.0 to 0.0 before hashing (std::hash need
not do that), and test -0.0 and the ends of the integer ranges. Show
hash(0.0) in the docs example and note the change in the version
history.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Clarify hash documentation after review
- Say "may hash differently" for null, false, and numbers, since a
collision across types is possible.
- Name the storage types (signed integer, unsigned integer,
floating-point number) instead of example literals.
- Explain that the hash survives converting an integer to
number_float_t but not the lossy conversion back, and that unequal
numbers may share a hash.
- State that the example hash values are illustrative only and vary by
platform, compiler, compiler version, and library version.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Afonso Januário <afonso-januario@hotmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Afonso Januário <afonso-januario@hotmail.com>
* Stop binary writers overflowing the stack on deep values
to_cbor, to_msgpack, and to_ubjson recurse once per nesting level.
The parser is iterative, so a value the library accepts can crash on
the way back out.
Keep the existing recursive path for the first 128 levels and finish
anything deeper on a heap stack. Output is unchanged. BSON is left
alone because its extra size walk is a separate change.
Rebased onto the value-type output sink. The heap frames now initialize
every member, which is what -Weffc++ was rejecting.
See #5392.
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
(cherry picked from commit cf65ac438f)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Redesign the iterative binary writers around a shared recursion depth limit
Address the open review on the non-recursive CBOR/MessagePack/UBJSON/BJData
writers (#5518):
- Delete the CBOR array/object prefix helpers; both the recursive and
iterative paths call write_cbor_head(), which already existed on develop.
- MessagePack: share one write_msgpack_array_prefix()/write_msgpack_object_prefix()
helper per container kind between the recursive and iterative paths, both
going through to_msgpack_length() so an over-long container throws
out_of_range.412 identically either way.
- Reuse detail::recursion_depth_limit() instead of a separate constant, the
same bound serializer::dump() and write_bson_document() already use.
- Redesign the frames after bson_frame/dump_frame: only a container with
elements is ever pushed, its header is written at the point it is pushed,
and the iterator is set in the frame's constructor instead of a
default-then-assign two-step with a since-removed "started" flag. The
UBJSON frame keeps only the value pointer, the per-element prefix_required
flag, and the iterator; write_closer and is_object are no longer stored,
since the former is always !use_count (use_count is constant for the whole
document) and the latter follows from value->is_object().
- Factor the BJData ND-array shape check into is_bjdata_ndarray(), used by
both the recursive object case and the iterative pushing logic.
- Give the frame classes the GCC -Weffc++ treatment already used for
diff_frame: a noexcept converting constructor plus the five special members
defaulted with no explicit noexcept.
- Fix two @ref self-references in write_cbor/write_msgpack/write_ubjson's own
doc comments to point at the public to_cbor/to_msgpack/to_ubjson/to_bjdata
API instead.
- The iterative object-key write for CBOR/MessagePack now runs the same
strict-mode check_utf8() against the parent object as diagnostics context
that the recursive path already ran, so the two paths raise identical
diagnostics across the switch-over.
- Rewrite the tests: round trips instead of a bare size check, byte-exact
comparisons against the recursive output at depths around the bound, a
deep object and a BJData ND-array past the bound, a deep discarded value
(type_error.321), and the OSS-Fuzz 566583014 CBOR/MessagePack regression.
BSON is unaffected by this change; it already walks its documents
iteratively and is covered separately by #5553.
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: ayush-singh-0601 <singhayush062006@gmail.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: ayush-singh-0601 <singhayush062006@gmail.com>
Co-authored-by: ayush-singh-0601 <179524189+ayush-singh-0601@users.noreply.github.com>
* Add a natvis fallback visualizer for detail::json_default_base
Squashed onto develop from:
- Add a type in the natvis template for detail::json_default_base
- Document the json_default_base natvis fallback and regenerate natvis
- Match json_default_base in both its current and 3.12.0 namespace
Co-authored-by: Mihnea Magheru <sakuntalle@yahoo.com>
Signed-off-by: Mihnea Magheru <sakuntalle@yahoo.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Take the natvis template from the PR in the amalgamation check
The check ran generate_natvis.py from a develop checkout, which loads
nlohmann_json.natvis.j2 from its own directory. A PR that changes the
template was therefore checked against develop's template and always
failed. Copy the PR's template next to the develop script before
running it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Mihnea Magheru <sakuntalle@yahoo.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: Mihnea Magheru <sakuntalle@yahoo.com>
The default dump() (no indentation, no ensure_ascii) gets its own
writer that makes the same walk and produces the same output:
- the write position stays in a local variable instead of a
member, so the compiler keeps it in a register across stores
through aliasing char pointers;
- strings and number tokens are copied with fixed-size 32-byte
moves wherever enough source bytes remain, instead of one
memcpy call per token;
- the innermost open container lives in local variables; a stack
that starts as a local array of 32 entries holds the rest;
- unedited documents are walked through the node array in order,
and integer tokens are read from the source directly.
On top of that, float tokens of at most 15 significant digits are
written straight from their digits via zmij::to_shortest() and
write_shortest(), without converting to a double and back: such
decimals are farther apart than a double's rounding interval, so
the token's digits are the double's shortest digits. Tokens of
16+ digits, or edited values, still go through decimal_to_float().
The view's own NEON write_decimal() is removed in favor of the
shared writer, and the dump output now grows in 64 KiB steps
instead of being resized to its estimate at once.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
An image is a document stored so that loading it needs no
parsing: save() writes the node index, the text and the decoded
strings; the static load() reads an image written by save().
load() takes a pointer and size, a borrowed vector, or an owned
rvalue vector; the nodes are copied so they are aligned and can
be edited, while the text and decoded strings stay in the image.
image_check controls how much load() trusts the input: full
checks structure, bounds, strings and numbers, the parser's own
guarantees; bounds checks structure and bounds only; none skips
all checks, for images from a trusted source.
Layout is little-endian only ("NJVI" header, nodes, text, decoded
strings), following the idea of zero-copy formats such as
FlatBuffers and YaFF; the check follows FlatBuffers' Verifier.
New errors: parse_error.116 for a malformed image or a failed
check, type_error.320 for a discarded document or a big-endian
target.
A dedicated fuzzer and 6,000 seeded corruptions, checked under
ASan/UBSan, found and fixed two gaps: unchecked reserved header
fields, and unbounded null/boolean offsets that could make
dump() throw std::length_error.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
basic_json_document gets a second template parameter, Editable
(false by default), plus the aliases json_editable_document,
json_editable_view, ordered_json_editable_document and
ordered_json_editable_view.
Editable documents can change values and structure without
rewriting the source text: set()/push_back() on values, keys,
array indices and JSON pointers; insert() before an array
element; erase() of an object key, array index or JSON pointer.
New values and element sequences go into edit storage that the
document owns and never moves, so views keep referring to their
value across edits and a parsed node never moves. Read-only
documents walk the plain node array and are unaffected.
Strings are checked for UTF-8 on entry, so dump() of an editable
document never throws type_error.316. Binary values cannot be
stored (type_error.319).
A seeded differential test applies random edits to an editable
document and to the equivalent ordered_json and compares both
after every step.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Speed up json_view's parser with SIMD scanning and a hash table
for large objects.
Long runs of string bytes are scanned 16 bytes at a time with NEON
(AArch64, GCC and Clang) and SSE2 (x86-64), both baseline
instruction sets. Keys keep 16 table checks before the vector
loop, because their lengths repeat from record to record; string
values get 8, because their lengths vary more. Non-ASCII text is
validated 16 bytes at a time with simdjson's "lookup4" check
(Keiser and Lemire, 2021), with NEON on AArch64 and, on x86-64,
with SSSE3. SSSE3 is not part of baseline x86-64, so the check is
compiled for SSSE3 with a function attribute and used only where
CPUID reports it, which all x86-64 CPUs since about 2011 do; the
answer is cached in a statically initialized atomic, so there is
no guard of a local static and no global constructor. The same
input is accepted either way. JSON_VIEW_NO_SIMD selects the
portable code.
On x86-64, string runs are now checked vector-first: one SSE2
compare from the first byte finds the end of most keys and short
values, instead of a branch per byte for the first 8-16 bytes.
AArch64 keeps the byte-wise steps, where a NEON mask costs more and
the branches predict well. Entering an object or array no longer
stalls: open() stores the parent's frame field by field instead of
building it on the stack and reading it back with wider loads,
which waited for the narrower stores to retire.
Objects with 128 members or more get an open-addressing hash table
built when the object closes, so operator[], at(), find(),
contains(), count(), value(), and JSON pointers take constant time
on average in such objects; of duplicate keys, the first is kept,
as for the linear search. The idea comes from Boost.JSON.
simdjson is credited in simd.hpp's SPDX block, the README, and
license.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Add basic_json_view::dump() and the comparison operators, and read
floats from the parser's digit layout instead of rescanning the
token.
dump(indent, indent_char, ensure_ascii, number_format) writes a
value the way ordered_json::parse(text).dump() writes it for the
same arguments: members in document order, all of them should a
key occur more than once; strings escaped by the same rules, using
the library's scanning kernels; floats written with the library's
to_chars conversion, so the output equals basic_json's byte for
byte; integers copied from the source, where they are already
canonical, except -0, which parse() reads as 0. There is no
error_handler argument, because the view only holds valid UTF-8.
number_format::source copies numbers exactly as they appear in the
source (e.g. "1.50", "1E2", "-0"), which basic_json cannot provide.
operator<< takes the indentation from the stream width, as for
basic_json. The writer walks iteratively, so nesting depth is
limited by memory only.
operator== and operator!= compare two views, or a view and a
basic_json value in either order, by the rules basic_json's
operator== uses: numbers compare by value across their types,
objects compare by their members with duplicate keys resolved as
parse() resolves them, member order matters only where the object
type keeps one, and discarded views compare as discarded basic_json
values do, including under JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON.
Nothing is materialized except single scalars.
While parsing, the view now records where the integer digits, the
fraction digits, and the exponent of a float token are, so floats
and doubles with at most 19 digits are read from that layout with
the library's decimal_to_float() instead of rescanning the token.
Both round correctly, so the values are those of parse(). get<double>(),
materialize(), dump(), and the comparisons all use it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Give basic_json_view the read-only access functions of basic_json:
operator[] and at() with keys and indices, front()/back(), find(),
contains(), count(), begin()/end() and cbegin()/cend(), items()
with structured bindings from C++17 on, and type_name().
Exceptions have the ids and messages of the const functions of
basic_json. Where basic_json has undefined behavior the view
answers safely: operator[] with a missing key or an out-of-range
index returns a discarded view, and front()/back() of an empty
container throw invalid_iterator.214. Objects are iterated in
document order, and all members are visited; duplicate-key lookups
find the first member (as yyjson and simdjson do), while parse(),
materialize(), and the map conversions keep the last value, as
parse() does. Keys of up to 16 bytes are compared with two
overlapping loads.
Add value conversions: get<T>()/get_to() for arithmetic types,
bool, nullptr_t, strings (std::basic_string copied,
string_view_t without a copy), BasicJsonType, views, std::vector,
and maps with string keys; get_string() for the string without a
copy; number_token() for the number exactly as written in the
source; value() with keys and JSON pointers; and operator[]/at()/
contains() with JSON pointers. Everything else, including types
with from_json(), goes through materialize() of that subtree.
get<T>() of arithmetic types is inlined down to the conversion, so
reading an integer needs no call.
detail::json_pointer_access exposes a pointer's reference tokens
to code outside basic_json.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Add json_document and json_view, a read-only, zero-copy index of a
JSON text, as the first public slice of the zero-copy view (#5295).
A parse produces a flat array of 16-byte nodes in document order,
one per value and one per object key. Strings stay in the source
text; escaped strings are decoded into an arena. Integers are
converted while their digits are in the cache; floats keep only
their digit layout and are converted on read. Containers store the
size of their subtree, so a reader can step over one in constant
time. A document makes a handful of allocations, however many
values it has.
The parser accepts exactly what json::parse accepts, with every
combination of ignore_comments and ignore_trailing_commas, with and
without a trailing NUL, and under JSON_STRICT_NUL_HANDLING. It is
portable C++11 and does not depend on byte order.
basic_json_document adds parse, parse_copy, accept, read (reuses a
document's memory), root, is_discarded, source, owns_source,
node_count, memory_usage, and shrink_to_fit. basic_json_view adds
type, the is_* queries, operator bool, size, empty, materialize,
and source_offset. A parse error throws the same exception
basic_json::parse would throw for the same input, message and
position included.
detail::abi_config keeps JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON readable after json.hpp
undefines them, in the ABI namespace so they always match the
basic_json in use.
A NUL byte that ends a // comment is the end of the input, as in
parse() since #5696.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Write doubles with the conversion of Zmij by Victor Zverovich (MIT),
ported to C++11 (detail/conversions/zmij.hpp). It finds the
shortest decimal that reads back as the same double, and the
closest one if there are several. Grisu2, used until now, is fast
but not always shortest: it sometimes writes a 17th digit where 16
suffice, or a last digit that is not the closest. The layout is
unchanged (1.5, 100.0, 1e+100, -0.0); float keeps Grisu2.
Digits are converted eight at a time with the BCD conversion of
Xiang JunBo, as in Zmij, and written with one byte swap per eight
digits and fixed-size moves instead of per-digit loops. Leading and
trailing zeros are counted from those bytes. to_chars() uses a
local buffer when the caller's is shorter than the 41 bytes this
may write. The powers of ten come from the number-parsing table,
adjusted where it holds values rounded up, and extended with Zmij's
compressed tables beyond 10^308.
write_shortest() converts its 16 digits in one vector register
(SSE2 on x86-64, NEON on 64-bit Arm, both baseline) and inserts the
decimal point inside the register, avoiding a store-forwarding
stall that cost about 25% of the time to write a double. dump()
writes floats and integers straight into the serializer's write
buffer instead of copying them from a member buffer, and small
integers eight digits at a time. read_eight_bytes() and
parse_eight_digits() are marked always-inline, which GCC had been
calling out of line in the number-parsing loops.
Of one million random doubles, about 0.14% are now written with
different digits, always to a value that still reads back as the
same double.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Give the library its own correctly rounded float converter for
binary32 and binary64 (IEEE 754), and speed up the lexer's string
and escape scanning.
The converter splits a number token into sign, significand, and
decimal exponent, then tries Clinger's fast path, then a templated
Eisel-Lemire step, and falls back to an exact big-integer digit
comparison for tokens with more than 19 significant digits whose two
candidate values round differently. This replaces std::from_chars
and strtod/strtof for both formats, so parsed values no longer
depend on the C/C++ library or the current locale. The strtold
fallback kept for other long double formats (x87, binary128) now
also copies a multi-byte decimal point correctly, fixing #5660.
eisel_lemire() and decimal_to_float() are always inlined so callers
keep the whole conversion in their hot loop.
The string-scanning kernels in string_scan.hpp find a stop byte with
the trailing-zero count of the SWAR mask instead of a byte loop, and
scalar_string_bulk_run() validates a run of multi-byte UTF-8
sequences one after another instead of re-searching after each one.
get_codepoint() decodes a contiguous \uXXXX escape with one table
lookup per byte instead of four range-checked get() calls; the
streaming path and all error positions are unchanged.
Adds 508 generated hard float-parsing cases with expected binary32
and binary64 bits, and kernel-comparison tests for the string scans
and the escape table against byte-by-byte references.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix CI on develop after #5585
- test-diagnostics-optimized: -O3 makes GCC's -Winline and
-Wsuggest-attribute=pure/const warnings fire with the ci_test_gcc flag
set; turn them off for this test.
- test-diagnostics-optimized: suppress Clang's -Wexit-time-destructors for
the static table in to_json.
- Infer: raise pulse-max-disjuncts from 20 to 40. With the default,
Pulse loses the stored type in basic_json::replace_value() and reports
false null dereferences of get_ptr() results in unit-pointer_access.cpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Ignore Infer's false USE_AFTER_DELETE in ordered_map::erase
Infer's std::string model keeps the buffer of a moved-from string, so the
destroy-and-reconstruct loop in erase(first, last) looks like it destroys a
buffer twice.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mark throw_on_discarded()'s parameters as used without exceptions
With JSON_NOEXCEPTION, JSON_THROW expands to std::abort(), so Clang's
-Wunused-parameter breaks test-disabled_exceptions (since #5761).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the span_input_adapter sax_parse checks with deleted deprecated functions
The #5676 regression test (#5740) calls the deprecated
sax_parse(span_input_adapter&&, ...), which JSON_DELETE_DEPRECATED_FUNCTIONS
deletes, so ci_test_delete_deprecated_functions failed to build.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fall back to the first entry in test-diagnostics-optimized's to_json
clang-tidy (clang-analyzer-security.ArrayBound) flagged it->second for a
value not in the table. Use the same fallback as
NLOHMANN_JSON_SERIALIZE_ENUM; the test still fails with -Werror=array-bounds
on the headers from before #5585.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix clang-tidy findings in tests from #5762 and #5774
- unit-regression2.cpp (#5762): const/auto for the destroy() test values;
NOLINT the intended copy in check_destroy_edge_case().
- unit-serialization.cpp (#5774): build the expected strings with += instead
of chained operator+ (performance-inefficient-string-concatenation).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
MSVC warns with C5260 that kAlpha and kGamma have internal linkage when
the header is used as a header unit or through a module. JSON_INLINE_VARIABLE
makes them inline variables from C++17 on.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Check in the fuzzers that parsing without exceptions agrees
Each fuzzer driver now also parses its input with allow_exceptions =
false. That call must never throw a parse_error, must return a discarded
value where parsing with exceptions fails, and must return the same value
where it succeeds. Values are compared by their dump(), because NaN is
not equal to itself.
A plain !is_discarded() assertion, as suggested in #3642, would never
fail: the drivers parse with exceptions, so a result can never be
discarded.
tests/fuzzing.md describes the checks and notes that OSS-Fuzz and
CIFuzz already run LeakSanitizer, because their default address
sanitizer includes it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use JSON_HAS_RANGE_VIEW_CONVERSION in the range view regression tests
#5728 combined the JSON_HAS_RANGES and MinGW conditions into
JSON_HAS_RANGE_VIEW_CONVERSION, but three test guards still spelled
them out.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test the serializer's buffers at their boundaries
The dump() indent overflow survived full line coverage because the tests
grew its buffer by only one step. This adds tests that land exactly on,
and one past, the limits of the other two serializer buffers:
- write_buffer (1024 bytes): strings of 1023, 1024 and 1025 bytes at the
top level, and of 1022 and 1023 bytes inside an array, so that both
guards in put_string() are hit at their boundary. Each is checked for
dump() and for stream output.
- string_buffer (512 bytes, flushed when fewer than 13 bytes remain):
runs of two-byte escapes, and a surrogate pair written with 14 bytes of
room, right after a flush, and one escape later.
- The 8-byte bulk scan from the serializer side: 0 to 17 plain bytes
followed by a quote, a control character, or a non-ASCII character.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test the chunked string and binary reads of all binary formats
The binary readers read strings and binary values in chunks of 4096
bytes. Only CBOR tested lengths around that size. MessagePack, UBJSON,
BJData and BSON now round-trip lengths 0, 1, 4095, 4096, 4097, 8192 and
100000 from vector and pointer input, and must report a truncated
payload as a parse error.
UBJSON reads binary values as arrays of numbers, so it is tested with
strings only. BJData binary values reach the chunked read only in
Draft 3. BON8 decodes strings byte by byte and does not use this path.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document what is not covered by the API stability guarantee
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the API stability guarantee to the roadmap
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Note that exceptions to the API stability rules are documented in the release notes
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
OSS-Fuzz moved its issues from bugs.chromium.org to issues.oss-fuzz.com.
The old link now redirects to the new tracker but drops the project
filter, so it showed every project's issues.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_BRACE_INIT_COPY_SEMANTICS enabled, single-element brace
initialization from a JSON value decided whether to copy the value or
build an object by inspecting the value's runtime shape: a two-element
array whose first element is a string, such as ["key", 42], was turned
into an object instead of being copied. This made the behavior depend
on the element's content, and it did not distinguish an existing value
of this shape from a nested braced pair written in the source, such as
the inner {"key", "value"} of {{"key", "value"}}.
json_ref now records whether it was constructed from a braced list
(true only for the std::initializer_list<json_ref> constructor used
for nested braced lists) or from a value. The initializer-list
constructor uses this to copy or move a single non-braced-list element
before deciding whether the list describes an object, so a JSON value
is always copied regardless of its shape, while a braced pair written
in the source still creates an object.
Fixes#5662.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Create a value before giving it its type
Squashed onto develop from:
- Create a value before giving it its type
- Skip the failed-allocation test when exceptions are disabled
- Keep the created pointer rather than an uninitialized json_value
- Skip the vector<bool> failed-allocation check for VS 2015 with iterator debugging
- Test the remaining to_json overloads with a failing allocation
- Skip the to_json allocation-failure section on VS 2015 Debug
- Fix false GCC -Warray-bounds error with JSON_DIAGNOSTICS at -O3 (#5744)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Store the new value with a helper in all to_json constructors
Every external_constructor<>::construct now creates the new value first and
hands it to basic_json::replace_value(), which destroys the old value, stores
the new one before setting its type (as elsewhere in this PR), sets the
parents, and checks the invariant.
The std::vector<bool> and std::valarray overloads use array_t's range
constructor, which the other array overloads already rely on. Range views
keep their loop, as begin() and end() of a view may have different types.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Squashed onto develop from:
- Cut Unicode ill-formed byte sweeps to one representative prefix
- Pin unrelated bytes in the remaining ill-formed UTF-8 sweeps
- Speed up unit-unicode1
- Run the cheap binary format size tests unconditionally
- Compile unit-msgpack.cpp only once
- Check the JSON Pointer roundtrip for every code point again
- Cover every byte class in the ill-formed UTF-8 sweeps
- Merge the Unicode tests into unit-unicode.cpp
- Stop excluding the Unicode tests in CI
Signed-off-by: elix3r <157088510+22elix3r@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: elix3r <157088510+22elix3r@users.noreply.github.com>
Squashed onto develop from:
- Make the json.tar.xz release archive configure standalone
- Update the supported compiler list and the CITATION.cff repository link
- Add json_fwd.hpp to the Bazel singleheader-json target
- Document SwiftPM's #include <json.hpp> form and add CI coverage
- Migrate REUSE metadata from deprecated .reuse/dep5 to REUSE.toml
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add helper types to make it easier to create a basic_json type with modified template parameters
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Rename with_changed_*_t aliases to with_*_t and merge integer/unsigned aliases
Per review discussion on #3898 between gregmarr and nlohmann:
- rename with_changed_X_t to with_X_t for brevity
- replace the separate with_changed_integer_t/with_changed_unsigned_t
aliases with a single with_integers_t<NumberIntegerType2, NumberUnsignedType2>
- add @sa doc comment links for the upcoming documentation page
Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add documentation for the with_*_t member alias templates
Add docs/mkdocs/docs/api/basic_json/with_t.md documenting with_object_t,
with_array_t, with_string_t, with_boolean_t, with_integers_t, with_float_t,
with_allocator_t, with_json_serializer_t, with_binary_t and with_base_class_t,
with an accompanying example, and link the page from the basic_json member
types list and the mkdocs navigation.
Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add tests for the with_*_t member alias templates
Check with std::is_same that each with_*_t alias produces the expected
basic_json type, and that with_string_t keeps nlohmann::ordered_map as
the object type when used on ordered_json.
Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix with_t nav entry and document chaining of the with_*_t aliases
Indent the with_t entry in mkdocs.yml so it is listed under basic_json,
explain that the aliases can be chained and work on ordered_json, and
test both, including json::with_object_t<ordered_map> == ordered_json.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add docset entry for basic_json::with_t
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: barcode <barcode@example.com>
Co-authored-by: Raphael Grimm <1005058+barcode@users.noreply.github.com>
2026-10-07: add the scalar enum labels already handled by the default
branch so -Werror=switch-enum builds succeed. Keep the existing behavior
and synchronize the generated single header.
Signed-off-by: fhgffy <102001626+fhgffy@users.noreply.github.com>
The non-recursive destroy walk from #5762 picked an object's last child
via object_t::rbegin() and std::prev(end()). Neither is available for
every ObjectType: no_key_compare_map in unit-custom-object-type.cpp has
no rbegin(), so develop no longer compiles that test, and hash maps such
as std::unordered_map only have forward iterators.
The walk can take an object's children in any order, as long as it
finds the same child again while the object is not modified in between.
So objects with bidirectional iterators keep using their last child
(O(1) to remove from vector-based maps like ordered_map), and objects
with forward-only iterators use begin() instead. No reverse iteration
or rbegin() is needed any more, and the walk stays allocation-free.
Adds a forward-only ObjectType to the tests, destroyed both with mixed
nesting and 100000 levels deep.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Check for the expected separator before the lexer's token switch
After a key the parser expects ':', after a value usually ','. Test for
that character first instead of going through scan()'s switch, which
compiles to an indirect jump. Any other character takes the old path,
so tokens and error messages are unchanged.
Parsing 6.3% faster with GCC 15.2 and 2.7% with Clang 22.1 (geomean of
the ParseString, ParseFile and ParseIndented benchmarks).
Signed-off-by: Michiel van Slobbe <michiel.van.slobbe@gmail.com>
* Improvement: address PR comments
Signed-off-by: Michiel van Slobbe <michiel.van.slobbe@gmail.com>
* fix: address comments
Signed-off-by: Michiel van Slobbe <michiel.van.slobbe@gmail.com>
* Fix clang-tidy bugprone-signed-char-misuse in scan_expecting
Convert the expected separator through unsigned char before storing it as
char_int_type. The generated code is unchanged.
Signed-off-by: Michiel van Slobbe <michiel.van.slobbe@gmail.com>
* Use raw string literals in the separator comment tests
Signed-off-by: Michiel van Slobbe <michiel.van.slobbe@gmail.com>
---------
Signed-off-by: Michiel van Slobbe <michiel.van.slobbe@gmail.com>
Co-authored-by: Michiel van Slobbe <michiel.van.slobbe@gmail.com>
* Replace retired macOS 14 runner and test all available Xcode versions
GitHub retires the macos-14 image on 2026-11-02 (brownouts from
2026-10-05). Xcode 15 is not available on any remaining hosted
runner, so drop the macos-14 job and its documented compilers.
Also test the Xcode versions that the images provide but CI did not
use (26.1.1-26.3 on macos-15, 26.4.1-26.6 on a new macos-26 job), and
pin GCC 16 explicitly next to gcc:latest.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document new Xcode and GCC versions in the supported compilers table
Versions taken from the CI logs of this PR (Xcode 26.1.1-26.6) and
from the gcc:16 image (same digest as gcc:16.2.0 and gcc:latest).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the provided allocator in destroy() (#4842)
Uses the provided allocator to allocate the stack used to avoid
recursion in the destroy() implementation used by ~basic_json.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test that the destructor uses the provided allocator
Adds a regression test for #4842: destroying a nested array or object must allocate its temporary stack through the basic_json allocator, not std::allocator.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix allocation failure during JSON destruction
Signed-off-by: Michael Sam <michaelsam94@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make json_value::destroy() non-recursive and allocation-free
destroy() used to flatten a nested array/object into a heap-allocated
std::vector to avoid recursing per nesting level. That vector could
itself throw bad_alloc under memory pressure, and since it now used the
basic_json's own allocator (#4842), a failing allocator supplied by the
caller made this more likely, not less. An exception thrown from inside
~basic_json(), which is noexcept, terminates the program (#5135).
Replace the vector-based stack with a pointer-reversal walk that visits
the tree without recursing per level and without allocating anything:
cur is the array/object currently being emptied, prev is its parent
(or null at the top). A parent's last child slot doubles as storage for
that parent's own parent link while we are below it, so no extra memory
is needed. A child is only ever removed once it is a scalar or an empty
array/object, which neither allocates nor recurses more than one level
deep. take() moves m_data between these locals directly, bypassing
set_parents()/assert_invariant() (the former is O(#children) per call
under JSON_DIAGNOSTICS, which would make the walk quadratic otherwise).
This also removes the std::vector<basic_json, allocator_type> stack
added by #4842, so the extra allocations it introduced disappear along
with it.
Co-authored-by: Michael Sam <9461037+michaelsam94@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Test that destroy() performs no allocation, even under memory pressure
Update the #4842 regression test: it used to check that destroying a
nested array/object made at least one allocation through the provided
allocator (the old flattening stack). Now that destroy() does not
allocate at all, assert the opposite: zero allocations, deallocations
only.
Rework the #5135 regression test to use a dedicated failing/counting
allocator instead of overriding the process-wide ::operator new and
::operator delete, which affected every allocation in the whole
unit-regression2 binary rather than just the values under test. Keep
the original small repro as one case, and add deep (100000 levels) and
wide-and-deep nested array/object/ordered_json cases, all destroyed
while every further allocation is made to fail: the destructor must
complete without allocating, without throwing, and without leaking.
Co-authored-by: Michael Sam <9461037+michaelsam94@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Refactor destroy() for readability and add edge-case tests
Apply review feedback from Greg Marr on the json_value::destroy()
non-recursive, allocation-free destruction walk (#5135):
- last_child() now uses object->rbegin()->second instead of
std::prev(object->end())->second; pop_last_child() keeps
std::prev(end()) since erase() needs a forward iterator.
- is_empty_container() becomes has_no_children(), a switch that
returns true for every non-container type as well as empty
array/object, simplifying the "scalar or already-empty child"
check at the call site. The local variable `last` is renamed to
`cur_last_ref` for clarity.
- free_container() asserts the array/object is already empty before
freeing it, and the object branches assert the expected type.
- destroy(value_t t) is now a thin dispatcher to destroy_string(),
destroy_binary(), and destroy_container(t), each handling its own
"not initialized" check and sharing the simple cases first in the
switch.
- destroy_container() moves the top-level container into the local
stand-in via a plain swap of the json_value union, instead of a
manual copy plus clearing array/object by hand.
- The "cur has no children and there is no parent" case now frees
cur and returns immediately, so the main loop is a plain
while (true) with no trailing code after it.
Also adds edge-case tests for both json and ordered_json (mixes of
empty/non-empty arrays and objects, container children in first/last
position, single-element chains, top-level empty containers, and
destruction via erase()/assignment), plus a mixed-tree case in the
"destructor performs no allocation" test.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make the destroy() walk helpers private
They modify basic_json internals without maintaining its invariants and
are only meant for destroy_container(), so they no longer need to be
reachable from the rest of basic_json.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Signed-off-by: Michael Sam <michaelsam94@users.noreply.github.com>
Co-authored-by: Vesko Karaganev <vesko.karaganev@gmail.com>
Co-authored-by: Michael Sam <michaelsam94@users.noreply.github.com>
Co-authored-by: Michael Sam <9461037+michaelsam94@users.noreply.github.com>
* Throw type_error.321 when serializing discarded values to binary formats
The CBOR, MessagePack, UBJSON, BJData, and BSON writers silently
skipped the payload of a value_t::discarded value nested in an array
or object, while still writing its slot in the element/member count
(and, for BSON, its entry header), producing a binary document whose
declared size does not match what was actually written.
Throw type_error.321 instead, for a discarded value anywhere in the
tree, including at the top level.
Rewritten from the original PR against the current (non-recursive
option aside) binary_writer.hpp, which has changed substantially since
this was first proposed; the out_of_range.412 MessagePack size check
and unrelated test reformatting from that PR are dropped as out of
scope here.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document and test type_error.321 for discarded binary values
Add docs for the new exception (home/exceptions.md and the Exceptions
sections of to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson) and test
coverage for a discarded value nested in an array or object, nested
deeper, and (for UBJSON/BJData) inside an optimized same-type array,
for each of CBOR, MessagePack, UBJSON, BJData, and BSON. Adjust the
three pre-existing "discarded" tests that asserted the old silent
behavior (empty/short output) to expect type_error.321 instead.
Co-authored-by: ameliabarnabyhub <312084480+ameliabarnabyhub@users.noreply.github.com>
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Co-authored-by: ameliabarnabyhub <ameliabarnabyhub@users.noreply.github.com>
Co-authored-by: ameliabarnabyhub <312084480+ameliabarnabyhub@users.noreply.github.com>
* Fix CI on develop after #5600, #5607, and #5755
- binary_reader: cast the result of -1 - number back to number_integer_t,
because a number_integer_t narrower than int is promoted to int, which
GCC's -Warith-conversion rejects (ci_test_gcc, ci_test_standards_gcc)
- JSON_DELETE_DEPRECATED_FUNCTIONS: declare the deleted stream operators
as function templates at namespace scope, because GCC < 5 rejects deleted
friend functions ("can't initialize friend function") and Clang 7-9 report
a redefinition when a class template has a deleted friend function
- docs: give the examples of JSON_USE_OBJECTS_FOR_ENUM_KEYED_MAPS "Example:"
titles and add the page to the docset (style_check)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Ignore libstdc++'s <format> sign change in the sanitizer job
libstdc++ 14's <format> initializes a size_t parameter with -1 (GCC bug
119429), so every std::format call fails ci_test_clang_sanitizer under
-fsanitize=integer (test-std-format_cpp20). Exclude only the
implicit-integer-sign-change check and only that header via
-fsanitize-ignorelist.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix the AppVeyor (MSVC 2015-2019) build
- binary_reader: emit_signed/emit_unsigned pass integers that do not fit
the number types to emit_float as long double, because MSVC's <cmath>
has no integer overloads of std::isfinite (C2668 'fpclassify'), from #5607
- scalar comparisons: the friend operators take the JSON type for their
noexcept from their parameter, because MSVC 2015/2017 take basic_json as
the class template there (C3203) and MSVC 2019 16.0 does not see member
types or template parameters, from #5751
- unit-conversions2: skip the !is_nothrow_constructible static_assert for
std::optional on MSVC 2017, which evaluates the conditional noexcept as
true (C2607), from #5754
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Re-amalgamate
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Avoid MSVC 2015's C4800 for enums with underlying type bool
MSVC 2015 warns about any conversion to bool (C4800), even with an
explicit cast, so the enum conversions from #5754 (#5671) failed the
AppVeyor build with /WX. Convert to bool by comparing with zero via the
new detail::bool_aware_static_cast, and keep doctest from printing the
enum in the test.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip the #5650 allocator test on MSVC 2015 debug builds
MSVC 2015's debug STL constructs the containers' debug proxies through
the allocator in noexcept constructors, so countdown_allocator's failing
construction crashes test-allocator (SIGSEGV) instead of throwing
std::bad_alloc. Use the guard #5585 uses for the same reason.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Skip deleted-function detection checks on MSVC 2015
MSVC 2015 does not treat selecting a deleted function in decltype as a
substitution failure, so the detection traits in
unit-delete_deprecated_functions (#5755) and the integral-key checks in
unit-element_access2 (#5657) report deleted overloads as callable there.
Calling them still fails to compile. Skip those checks for
_MSC_VER < 1910, and use the stream operators for real in the runtime
section, so MSVC 2015 still compiles them with the macro set.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Split the Visual Studio 2017 AppVeyor jobs in two
The VS 2017 jobs hit AppVeyor's 60-minute limit per job while still
compiling the tests (77 of about 108 test targets after 58 minutes).
They pass /std:c++17 for everything anyway, so build only the C++17
variant of each test (JSON_TestStandards=17), and split the unit test
files across two jobs each with the new JSON_TestShard=<index>/<count>
option, which keeps every <count>-th test file starting at <index>.
The extra variants of single test files are built in shard 0 only.
CMAKE_OPTIONS is no longer quoted in appveyor.yml, so that it can hold
more than one option.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix std::terminate when converting std::optional with MSVC 2017
MSVC 2017 evaluates std::is_nothrow_assignable<json&, const T&> as true
even if T's to_json throws, so to_json(json&, const std::optional<T>&)
was noexcept there and the exception from #5642's test called
std::terminate instead of propagating. Make that conversion never
noexcept on MSVC 2017; all other compilers keep the exact condition.
The static_asserts on the condition are skipped for MSVC 2017; the
runtime check that the exception propagates still runs there.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Add JSON_DELETE_DEPRECATED_FUNCTIONS to delete the deprecated functions
Defining JSON_DELETE_DEPRECATED_FUNCTIONS to 1 (or the CMake option
JSON_DeleteDeprecatedFunctions) declares every deprecated function as
deleted instead of deprecated, so that code that is not ready for 4.0.0
no longer compiles. A deleted function still takes part in overload
resolution, so from_*(ptr, len) cannot silently bind len to the strict
parameter of from_*(InputType&&, bool); the roadmap now plans to keep
these overloads deleted in 4.0.0 instead of removing them.
The legacy discarded-value comparison is left to its own macro.
Also update the 4.0 roadmap: add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
and JSON_DELETE_DEPRECATED_FUNCTIONS to the macro table, add the
from_bjdata/from_bon8 (ptr, len) overloads to the deprecated functions,
document the macro in the migration guide, and fix the docs style check
findings (example titles, missing docset entry for JSON_STRICT_BINARY_UTF8).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Declare each deprecated function once and guard only its body
Instead of repeating every deprecated declaration in an
#if JSON_DELETE_DEPRECATED_FUNCTIONS branch, keep one declaration
(with its deprecation attribute) and switch only between "= delete;"
and the function body. Suggested by @gregmarr in the review.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
from_bon8(ptr, len) and from_bjdata(ptr, len) had no overload for a
pointer and a length, unlike from_cbor/from_msgpack/from_ubjson/
from_bson. The call instead bound to from_*(InputType&&, bool strict),
which read ptr as a NUL-terminated C string via strlen and silently
converted len to the strict flag. Data containing a 0x00 byte was cut
off there; data without one was read past the end of the buffer.
Add a deprecated (ptr, len, strict, allow_exceptions) overload for
each function that forwards to (ptr, ptr + len, ...), matching the
existing deprecated overloads of the other four binary readers. Since
neither function ever had this overload, the deprecation is declared
as of version 3.13.0, the next unreleased version, rather than the
version each function was originally added in.
Fixes#5648.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>