json_document::read is a member function, not POSIX read(), and the C
string overload requires null-terminated input like json::parse.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Conflicted only in tests/src/unit-class_lexer.cpp, where develop's #5737
lint fix (CAPTURE(x); -> CAPTURE(x)) collided with this PR's rewrite of
the Eisel-Lemire float tests; kept the PR's new tests and applied the
lint-fixed CAPTURE style. single_include regenerated via make amalgamate.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Re-amalgamate single_include
#5737 changed 13 headers under include/ but merged without the matching
single_include/nlohmann/json.hpp update, so the amalgamated header still
had, among others, the GCC C++20 -Wignored-attributes pragma block and
the clang -Wdocumentation push/pop that #5737 removed, the forwarding
from_json tuple/array helpers it replaced with const references, and
lacked the output_adapter char_traits changes it added.
Regenerated with `make amalgamate` (astyle 3.4.13). The diff is exactly
`git diff b54ed188ee5a89d671 -- include/` (164+/95-); json_fwd.hpp and
json_literals.hpp were already up to date.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Run the amalgamation check on pushes to develop
#5737 was merged 39 seconds after its last push, while its own
check_amalgamation run was still queued (earlier runs had been
cancelled by the concurrency group), so the stale single_include
reached develop without any failing check.
Also run the check on pushes to develop, without cancelling in-progress
develop runs. The "save" job (PR number/author for the comment
workflow) only runs for pull requests, the checkout falls back to
github.sha, and comment_check_amalgamation.yml only comments for
PR-triggered runs, since push runs have no PR and no "pr" artifact.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Drop stale LCOV_EXCL_LINE from the json_pointer out_of_range.410 throw
The comment said the size_type overflow check in array_index() is only
triggered on special platforms like 32-bit, and the throw was excluded
from coverage. On 64-bit platforms the check is true for SIZE_MAX
itself, and unit-json_pointer.cpp has asserted that case four times
since #5395, so the line is executed in the coverage job. Reword the
comment and remove the exclusion marker so the coverage report notices
if the tests stop reaching it.
Part of #5725
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Name all three C-array check aliases in the enum-macro NOLINTs
NLOHMANN_JSON_SERIALIZE_ENUM(_STRICT) suppressed the c-array warning
under modernize-avoid-c-arrays only, but clang-tidy emits the same
diagnostic under the aliases cppcoreguidelines-avoid-c-arrays and
hicpp-avoid-c-arrays too. Any user running those checks got a false
positive at every macro expansion, and our own tests needed a local
NOLINT at each call site to work around it.
Name all three aliases in the four macro comments instead, and drop
the now-redundant c-array names from the five test call-site NOLINTs.
Comment-only change; behavior, the public API, and the ABI do not
change. Ran make amalgamate.
Part of #5725
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Include doctest as a SYSTEM directory instead of disabling warnings for all tests
test_main added -Wno-deprecated and -Wno-float-equal as PUBLIC compile
options for every non-MSVC compiler, so they were applied to every
translation unit, library headers included, and silenced the CI
warnings meant to check the library's own -Wfloat-equal pragmas. The
only code that actually needed the suppression was the vendored
doctest.h, which was included as a normal (non-SYSTEM) directory.
Include thirdparty/doctest as SYSTEM for test_main, matching what
tests/abi/CMakeLists.txt already does, and drop the two suppressions
from both targets. Verified locally that unit-comparison,
unit-conversions and unit-constructor1 compile clean with
-Werror -Weverything and doctest as -isystem, and that CMake still
configures with JSON_BuildTests=ON.
Part of #5725
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove the no-op ci_clang_analyze target
ci_clang_analyze configured the build with the real compiler and only
then wrapped ninja with scan-build. scan-build intercepts compiles by
overriding CC/CXX, but build.ninja already had the compiler path baked
in from the configure step, so every run bypassed the analyzer: CI
logs show "No bugs found" after a normal build, never an analysis.
The job also used Debian's frozen clang-tools-14 rather than the
image's own clang, and CLANG_ANALYZER_CHECKS still named three
valist.* checkers that current clang merged into security.VAList.
ci_clang_tidy already runs every clang-analyzer-* check (via
.clang-tidy's "Checks: '*'") with warnings as errors, so nothing is
lost by removing the dead job. Delete ci_clang_analyze,
CLANG_ANALYZER_CHECKS and the SCAN_BUILD_TOOL lookup from
cmake/ci.cmake, drop it from the ubuntu.yml ci_static_analysis_clang
matrix, and drop the now-unused clang-tools apt package (iwyu stays
for ci_single_binaries). Reword quality_assurance.md and
assurance_case.md, which described the dead job as a working control,
to say the Clang Static Analyzer checks run through clang-tidy.
Verified that `cmake -DJSON_CI=ON` still configures cleanly and that
ci_clang_analyze no longer appears in the generated build or in any
CMake/workflow file.
Part of #5725
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Re-enable portability-template-virtual-member-function; remove redundant forwards
.clang-tidy disabled three checks "to get the CI going" (#4489,
2024-11-13): portability-template-virtual-member-function,
bugprone-use-after-move and its alias hicpp-invalid-access-moved.
portability-template-virtual-member-function only flagged
output_stream_adapter::write_character/write_characters; annotate
both with NOLINT and re-enable the check.
bugprone-use-after-move flagged several double forwards that have no
effect at runtime:
- from_json.hpp calls std::forward<BasicJsonType>(j).at(Idx) inside
pack expansions; at() has no ref-qualified overloads and always
returns an lvalue reference, so the forward is a no-op. Replace with
plain j.at(Idx) in all four places.
- the move constructor forwards the whole object to its base class
and then reads other's members. That is item 9 of #5724 (together
with its cppcheck suppressions) and is left to that change.
- input_adapters.hpp forwards the container twice on purpose, so the
begin/end iterator types match adapter_type; annotate with NOLINT
and a comment instead of changing behavior.
The check still flags the move constructor (see above) and two sites
in at(KeyType&&) (both overloads, json.hpp, in the throw's
string_t(std::forward<KeyType>(key)) after
find(std::forward<KeyType>(key))). Open PR #5689 rewrites that hunk,
so bugprone-use-after-move (and hicpp-invalid-access-moved)
stay disabled for now, with a comment explaining why; re-enable them
once #5689 and the #5724 move-constructor change have landed.
Also resolve the portability-avoid-pragma-once TODO: single_include
never has #pragma once (amalgamate.py strips it) and every supported
compiler accepts it in include/, so keep it disabled with an
explanatory comment instead of a TODO. Fix the stale "json.hpp,
around line 1265" comment in unit-class_parser.cpp, which now points
at the move constructor's actual line.
Behavior, the public API and the ABI do not change. Verified with
clang-tidy 22.1.8 that portability-template-virtual-member-function
now reports nothing, that bugprone-use-after-move/
hicpp-invalid-access-moved report only the known at(KeyType&&) and
move-constructor sites, and that unit-custom-base-class, unit-constructor1,
unit-conversions, unit-element_access2, unit-class_parser and
unit-diagnostic-positions (JSON_DIAGNOSTIC_POSITIONS=1) compile
under ASan/UBSan and pass with the same assertion counts as before.
Ran make amalgamate.
Part of #5725
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix stale and malformed NOLINT comments
json_sax.hpp named "-warnings-as-errors" in the NOLINT list on the two
JSON_ASSERT(false) lines; that is the suffix clang-tidy appends to a
diagnostic tag under WarningsAsErrors, not a check name, and every
other JSON_ASSERT(false) omits it.
unit-capacity.cpp carried 30 "// NOLINT(misc-const-correctness)"
comments on "json j = ...;" declarations that are all used with
non-const members afterwards, so the check has nothing to report
there.
unit-constructor2.cpp used a blanket "// NOLINT: access after move is
OK here" on a use-after-move that hides every check on the line;
naming bugprone-use-after-move and hicpp-invalid-access-moved keeps
the intent once those checks are re-enabled (#5724).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5725 item 10
* Remove stale .clang-tidy entries
-google-runtime-references disabled a check that neither clang-tidy
22.1.8 nor 23.1.2 lists under --list-checks -checks='*'; it was
removed upstream. The commented-out HeaderFilterRegex line has been
unused since the active HeaderFilterRegex was introduced in #2561
(2021).
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5725 item 11
* Remove the GCC C++20 -Wignored-attributes pragma in json.hpp
The pragma (added in #5164) claimed to work around the C++ modules
redefinition errors of #5103, but #5103 is about hard errors (e.g.
"redefinition of std::__is_constant_evaluated()", conflicting
std::integral_constant) that ignoring a warning cannot suppress; they
are traced to GCC PR 124430 and reproduce with <map> or <string>
instead of json.hpp too. A GCC 16.2 -std=gnu++20 -fmodules build
following #5103's repro steps still fails with the pragma in place,
and a build of all test TUs with GCC_CXXFLAGS (which enable
-Wignored-attributes) and the pragma removed produces no such
warning. The block only hid a warning class from GCC C++20 users
while suggesting #5103 was handled.
Overlaps #5610, whose hunks touch the closing half of this pragma to
insert the json_literals.hpp include.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5725 item 9
* Fix stale doxygen comments hidden by the -Wdocumentation pragma
macro_scope.hpp ignores -Wdocumentation and -Wdocumentation-unknown-command
for the whole library, which also hides genuine documentation mistakes:
- detail::unescape() documented "@return unescaped string" but returns
void and unescapes its argument in place; reworded to
"@param[in,out] s string to unescape in place" and dropped the
bogus @return.
- basic_json::get()'s copy-conversion overload wrote "converted to
@tparam ValueType" inside @return, which Doxygen and Clang parse as
a second, malformed @tparam; changed to "@a ValueType", matching the
two other get() overloads a few lines above that already use it.
This narrows the gap the -Wdocumentation pragma needs to cover; fully
replacing the Doxygen-only commands it also hides (item 2c) is left
for after #5267.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5725 item 2
* Fix -Wextra-semi-stmt at its actual source, not assert()
clang_flags.cmake blamed the global -Wno-extra-semi-stmt on assert(),
but assert() expands to an expression under glibc and libc++ and does
not trigger this warning. unit-assert_macro.cpp overrides JSON_ASSERT
with "{if (!(x)) ++assert_counter; }", a bare block followed by a
semicolon at every JSON_ASSERT(...) call site in the library; that
was the actual source of 151 of the 208 -Wextra-semi-stmt sites found
in a Clang 22 -Weverything sweep of the test suite with the flag
removed. Switched to the standard do/while(false) macro idiom, which
does not expand to a statement-plus-semicolon, and corrected the
comment to name the remaining source instead: vendored Doctest's
CAPTURE(x) shim, which already ends in a semicolon.
Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that unit-assert_macro.cpp now compiles without any
-Wextra-semi-stmt diagnostic.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5725 item 8 (step 1 of 2; step 2 covers the CAPTURE() call sites)
* Drop the redundant semicolon from CAPTURE() call sites; remove -Wno-extra-semi-stmt
doctest_compatibility.h defines CAPTURE(x) as DOCTEST_CAPTURE(x); (with
a trailing semicolon baked into the macro), specifically so call sites
do not need to add one themselves; most of the ~267 call sites already
follow that convention. The remaining 64 call sites across 20 files
wrote "CAPTURE(x);" anyway, turning into a statement plus an empty
statement and triggering -Wextra-semi-stmt. Dropped the redundant
semicolon at each of those sites.
With item 6 having already made vendored Doctest a SYSTEM include, and
this the last known source of -Wextra-semi-stmt findings, removed the
flag from clang_flags.cmake entirely.
Verified with clang++ -Wextra-semi-stmt (plus the file's other CI
ignores) that all 20 touched files, plus a file with no CAPTURE() use
(unit-json_pointer.cpp), compile without any -Wextra-semi-stmt
diagnostic.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5725 item 8 (step 2 of 2)
* Switch ci_static_analysis_clang off the frozen LLVM 22 dev image
ubuntu.yml pinned the clang-tidy/clang-tidy-sanitizer/single-binaries
job to silkeh/clang:dev, a tag last pushed 2026-02-18 that reports
"clang version 22.0.0 (...+20251015...)", a pre-release snapshot from
before the LLVM 22 release; the maintainer now updates dev-unstable,
22, and latest instead. Switched to silkeh/clang:22, matching the
other clang jobs on :latest.
Verified with clang-tidy 22.1.8 (the image's actual version) against
this repository's .clang-tidy and library headers what the release
image newly reports compared to :dev:
- readability-redundant-typename fires at ~250 sites across the
_cpp20-relevant conversion/to_chars headers; the library targets
C++11 and keeps the typenames, so the check is disabled in
.clang-tidy, matching how the file already handles checks that
don't fit a C++11 codebase.
- misc-anonymous-namespace-in-header fires on the two anonymous
namespaces in from_json.hpp and to_json.hpp; added the alias to
their existing NOLINT (cert-dcl59-cpp, fuchsia-header-anon-namespaces,
google-build-namespaces).
- bugprone-std-namespace-modification fires on every addition to
namespace std: the std::hash, std::formatter and std::swap
overloads in json.hpp, and the std::tuple_size/std::tuple_element
specializations in iteration_proxy.hpp (this last file is not named
in #5725's item 5, found by actually running clang-tidy 22.1.8
against the current tree). All six are legal, deliberate additions
to namespace std (explicit/partial specializations of std types, or
the pre-C++20 std::swap overload); annotated each with the check
name next to its existing cert-dcl58-cpp NOLINT.
- modernize-avoid-c-style-cast reported nothing new.
Also added clang++-22/21, clang-tidy-22/21, g++-16 and gcov-16 to the
find_program search lists in ci.cmake so a local "maximal warnings"
configure prefers the current toolchain version over an older one on
PATH.
#5725 item 5
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Regenerate cmake/gcc_flags.cmake for GCC 16.2.0
GCC_CXXFLAGS was generated for GCC 15.1.0, but ci_test_gcc and
ci_test_gcc_cxx{11..26} now run in gcc:latest, currently GCC 16.2.0,
so the "maximal warnings" job was missing warnings introduced since
15.1.0 while carrying entries GCC 16 treats as duplicates or no-ops.
Regenerated with https://github.com/nlohmann/gcc_flags (patched
locally to not crash on an option whose "-x c++ <opt> -" probe fails
before it reads stdin, e.g. -Wabi=; the tool otherwise raises
BrokenPipeError instead of recording the option as an error) run
against g++ 16.2.0 in the official gcc:16 Docker image, keeping the
documented -Wno-* exclusions and the same alphabetical placement
scheme as before.
Also added three GCC 16 warnings the generator cannot discover on its
own because it only probes value ranges/lists it finds in the -Q
option name itself, not in the enum choices --help=warnings documents
separately:
- -Wbidi-chars=any, -Wleading-whitespace=spaces: manually verified
these compile cleanly with g++ 16.2.0.
- -Wstrict-flex-arrays: deliberately NOT added, unlike the other two.
Without -fstrict-flex-arrays (which the library does not enable, as
it would change codegen for flexible array members), GCC prints
"'-Wstrict-flex-arrays' is ignored when '-fstrict-flex-arrays' is
not present" on every translation unit, and under our -Werror that
note itself aborts the build. This differs from the harmless
no-op warnings already kept in the file (-Whsa, -Wsynth,
-Wunreachable-code, -Wunsafe-loop-optimizations), which emit
nothing; #5725 item 7 named -Wstrict-flex-arrays as one of the
flags GCC 16 adds, but did not anticipate this failure mode.
Verified: compiled the library header and a representative set of
test translation units (including ones touched by items 1, 3, 8, 9,
10 of this issue) with the regenerated GCC_CXXFLAGS plus -Werror
under g++ 16.2.0 at -std=c++11 through -std=c++26, with zero warnings;
ran the full local test suite (129/129 passing, unrelated to this
compiler) as a regression check. CI must still confirm the actual
ci_test_gcc / ci_test_standards_gcc targets end to end, since this was
verified with direct g++ invocations rather than through the CMake/
CXXFLAGS environment-variable plumbing in ci.cmake.
#5725 item 7
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Avoid std::basic_string<CharType> for non-character output_adapter CharType
output_adapter<CharType, StringType> defaulted StringType to
std::basic_string<CharType>, and (with JSON_NO_IO undefined) always
declared a std::basic_ostream<CharType>&-taking constructor. For
CharType with no non-deprecated std::char_traits specialization (only
std::uint8_t is ever used this way, by the binary writers), simply
naming either type - as an unused default template argument, or as an
unused, never-called constructor's parameter type - instantiates
std::char_traits<CharType> merely to name it, which some standard
libraries mark deprecated: with the library-wide -Wdocumentation
pragma (item 2's other half, left for a later commit) temporarily
removed, an Apple clang 21 / libc++ TU calling json::to_cbor(j, vec)
with std::vector<std::uint8_t>& got one -Wdeprecated-declarations
warning per binary writer at the old output_adapters.hpp:193.
Replaced the eager std::basic_string<CharType> / std::basic_ostream
<CharType> defaults with a bool-tagged partial specialization (not
std::conditional, which requires naming both branches' types up
front regardless of which is selected, reproducing the same warning)
that only ever names std::basic_string<CharType> / std::basic_ostream
<CharType> when CharType is actually one of char, wchar_t, char16_t,
char32_t, or (with __cpp_lib_char8_t) char8_t. For any other
CharType, output_adapter's StringType and ostream-constructor
parameter fall back to two distinct empty placeholder types, kept
distinct so the two constructor overloads do not collide into a
single redeclaration.
Public API / behavior: passing a std::basic_string<std::uint8_t>& or
std::basic_ostream<std::uint8_t>& directly to a binary writer's
output_adapter now fails to compile instead of compiling with a
deprecation warning; this was neither documented nor tested. All
documented uses (std::vector<CharType>, std::basic_ostream<CharType>
and StringType for character CharType) are unaffected.
Verified with Apple clang 21 / libc++, with the two -Wdocumentation*
"ignored" pragma lines in macro_scope.hpp temporarily removed and
-std=c++11/c++20 plus the project's -Weverything flag set: calling
to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson/to_bon8 on a
std::vector<std::uint8_t> now produces no char_traits<unsigned char>
(or any other) deprecation warning, while the char-based string- and
ostream-adapter paths, and a to_cbor/from_cbor round trip, still
compile and run correctly; also verified with GCC 16.2.0. Ran the
full local test suite, including the binary-format unit tests
(unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bson,
unit-bon8, unit-binary_writer_sinks, unit-binary_formats,
unit-custom-binary-type): 129/129 passing.
#5725 item 2 (step a)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove the library-wide -Wdocumentation pragma; fix what it hid
macro_scope.hpp / macro_unscope.hpp pushed and popped a Clang
diagnostic region over the entire library that ignored -Wdocumentation
and -Wdocumentation-unknown-command. Removed both pragmas and fixed
every finding a full -Wdocumentation (which implies
-Wdocumentation-unknown-command and -Wdocumentation-deprecated-sync)
build reports, so the library now compiles clean under Clang's
documentation checks without a blanket suppression. Overlaps #5267,
which is still open and edits a nearby doc block (json.hpp's
get()/get_impl() @return, already fixed in the item 2 step (b) commit
of this branch); this commit does not touch that block again.
Unknown Doxygen alias commands (Doxyfile removed in #3071, so these
were never rendered by anything) rewritten as plain prose, keeping the
same information:
- @requirement REQ-JSON-01 / REQ-JSON-02 (iter_impl.hpp,
json_reverse_iterator.hpp): now "This class satisfies the following
concept requirements (REQ-JSON-0N):".
- @liveexample{prose,example-id} (three sites in json.hpp): kept the
prose, dropped the command wrapper and the trailing example-id
(docs/mkdocs/docs/examples/*.cpp still exist and are used directly
by the rendered docs, not through this in-header alias) and
unescaped the "\," commas that were only needed for the old alias's
comma-separated argument syntax.
- @complexity X (json.hpp x4, json_pointer.hpp x2, serializer.hpp x1):
now "Complexity: X".
Backslash sequences Clang's comment lexer tried to parse as commands,
escaped to render as literal backslashes:
- lexer.hpp get_codepoint(): two `\u` occurrences.
- binary_reader.hpp get_bson_cstr() / get_bson_cstr_bulk(): two
`\x00` occurrences.
- serializer.hpp: three `\uXXXX` occurrences (constructor @param,
append_codepoint_to_string_buffer() @brief, and the ensure_ascii
member comment).
One finding remained after all of the above: Clang reports
"declaration is marked with '@deprecated' command but does not have a
deprecation attribute" on the deprecated sax_parse(span_input_adapter&&, ...)
overload, even though JSON_HEDLEY_DEPRECATED_FOR does expand to
__attribute__((deprecated(...))) for Clang. Several isolated
reproductions of this exact declaration shape - doc comment,
template<>, two stacked __attribute__ macros, an overload set sharing
the name - did not reproduce the warning, so this looks like a
Clang comment/declaration-association quirk specific to this overload
inside the much larger basic_json class template, not an actual
documentation defect. Rather than keep the pragma library-wide for one
Clang false positive, added a tightly scoped
-Wdocumentation-deprecated-sync push/pop around just that overload.
Verified with Apple clang 21 and the project's actual -Weverything
flag set (cmake/clang_flags.cmake) on the full header at -std=c++11
and -std=c++20: zero -Wdocumentation* diagnostics. Also compiled
clean with GCC 16.2.0 (the pragmas are already __clang__-gated, so
this only confirms no unrelated breakage). Ran make check-amalgamation
and the full local test suite: 129/129 passing.
#5725 item 2 (step c)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Take the JSON value by const reference in the array and tuple from_json paths
Review feedback on #5737 (gregmarr): once the no-op std::forward calls
are gone, the forwarding references have no purpose. from_json_fn
passes the value as const BasicJsonType&, so these functions were only
ever instantiated with a const lvalue anyway.
The std::array, std::pair and std::tuple overloads of from_json and
their helpers now take const BasicJsonType& and pass j on unchanged.
Because the deduced BasicJsonType is now the plain type, tuple_type and
the static_assert name const BasicJsonType& explicitly, so the
reference checks are unchanged: get<std::tuple<const std::string&>>()
still works, and get<std::tuple<std::string&>>() still fails the same
static_assert. from_json_tuple_get_impl keeps its forwarding reference,
since tuple_type calls it through std::declval.
Behavior, the public API and the ABI do not change. unit-conversions,
unit-constructor1, unit-udt, unit-udt_macro, unit-regression1/2/3,
unit-deserialization, unit-noexcept, unit-items, unit-allocator,
unit-custom-object-type, unit-ordered_json2 and
unit-brace-init-copy-semantics pass at C++11, C++17 and C++20 with
unchanged assertion counts. Ran make amalgamate.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Run the README test case in JSON_FastTests jobs
The "README" test case was marked doctest::skip() when the tests
moved from Catch to doctest in 2019, where it replaced Catch's hidden
tag. It is not slow (17 assertions, about 0.00 s), but cmake/test.cmake
only passes --no-skip when JSON_FastTests is off, so the per-compiler
ci_test_*_cxxNN matrix, macOS, Windows Release/ARM, icpc, icpx and
nvhpc compiled the README examples without running them.
Drop the skip decorator so every job runs the case. Test-only change.
Part of #5713
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove stale clang ranges guards in unit-iterators2.cpp
The "algorithms" and "views" sections were guarded by clang/libstdc++
checks written for a clang 15 (04/2022) bug. The first guard's
condition contradicts its own comment: it skips clang+libc++ and
keeps clang+libstdc++. Both sections already sit inside
`#if JSON_HAS_RANGES`, which macro_scope.hpp excludes for the
toolchains these guards targeted, so the inner guards never let the
sections run on the platforms they meant to protect and are
redundant on the rest. Verified locally with Apple clang 21/libc++
and clang 16.0.6/libstdc++ 12 (Docker): both pass all 1355
assertions with the guards removed.
Part of #5713
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix copy-pasted CBOR half-float checks; enable stale encode checks
In the RFC 8949 Appendix A test case, the decode checks for
5.960464477539063e-8 (0xf9 0x00 0x01) and 0.00006103515625
(0xf9 0x04 0x00) were copy-pasted from the neighboring -4.0 example,
so those two half-float byte sequences were never actually decoded
and checked, and -4.0 was checked three times instead. The two
float32 encode checks for 100000.0 and 3.4028234663852886e+38 were
commented out before the writer supported emitting float32 and are
now verified to match byte for byte, so they are enabled. The
remaining commented-out half-precision to_cbor checks are collapsed
into a single explanatory comment, since the writer never emits
half-precision floats.
Part of #5713
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Assert on the result of STL container conversions in tests
The "object-like STL containers" and "array-like STL containers"
sections converted json values into std::map, unordered_map,
multimap, unordered_multimap, list, forward_list, array, valarray,
vector, deque, set and unordered_set and discarded the result, so
these ~60 conversions only proved that the code compiles and does
not throw; a conversion that dropped or reordered elements would
still pass. Bind each result and compare it against the expected
container. Also fix a copy-paste slip in the deque section
(`j2.get<std::deque<double>>()` instead of j3, so j3's doubles were
never converted to a deque), and remove the dead
`// CHECK(m5["one"] == "eins")` comments that referred to a variable
that did not exist by asserting the equivalent through the bound
result.
Part of #5713
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Deduplicate SaxCountdown and other test helpers across formats
SaxCountdown was copied byte-for-byte into six binary-format test
files (unit-cbor.cpp, unit-msgpack.cpp, unit-ubjson.cpp,
unit-bjdata.cpp, unit-bon8.cpp, unit-bson.cpp), about 370 redundant
lines. Move it into tests/src/sax_countdown.hpp (namespace utils,
alongside test_utils.hpp and round_trip_corpus.hpp) and include it
from all six.
trait_test_arg and the "value_in_range_of trait"
TEST_CASE_TEMPLATE_DEFINE were duplicated between unit-32bit.cpp and
unit-bjdata.cpp; the trait is a detail/meta trait, not specific to
either file. Move it into tests/src/value_in_range_of_test.hpp;
unit-32bit.cpp keeps its own include, since JSON_32bitTest=ONLY
builds only that file. Each file keeps its own
TEST_CASE_TEMPLATE_INVOKE list.
sax_no_exception and the "issue #2824" section were duplicated in
unit-regression2.cpp and unit-disabled_exceptions.cpp. Drop the copy
from unit-regression2.cpp; unit-disabled_exceptions.cpp already
covers the no-exceptions case that #2824 was about, and
ci_test_noexceptions reruns it.
No behavior change. Verified by building and running unit-cbor,
unit-msgpack, unit-ubjson, unit-bjdata, unit-bon8, unit-bson,
unit-32bit, unit-regression2 and unit-disabled_exceptions against
include/ (clang++ -std=c++11, ASan/UBSan where applicable); assertion
counts are unchanged from before the refactor.
Part of #5714
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove the unreferenced vendored libFuzzer
tests/thirdparty/Fuzzer (155 files, ~776 KB of vendored Apache-2.0
LLVM code from the 2016 OSS-Fuzz import) is not referenced by any
CMakeLists, Makefile or workflow: the fuzz drivers link against
-fsanitize=fuzzer or the repo's own
tests/src/fuzzer-driver_afl.cpp. Its vendored README only points at
llvm.org's own libFuzzer docs. Being dead code, it also adds noise
to the flawfinder code-scanning workflow, which scans the whole
tree. Remove the directory and its .reuse/dep5 entry.
Part of #5714
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove unreferenced 2016 benchmark and fuzz reports
tests/reports (1.6 MB) holds AFL status pages and plots from
2016-08-29 and 2016-10-02, and a nativejson-benchmark snapshot from
2016 with links to rawgit.com, which shut down in 2019. Nothing
references this directory: no doc, README section, script or
workflow points at it, and it describes a ten-years-old, pre-2.0
snapshot of the library.
Part of #5714
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Run the CBOR, MessagePack, BSON and BON8 round-trip invariants in CI
tests/src/round_trip_corpus.hpp exists so that the byte-stability
invariant the fuzzer drivers check also runs on a fixed corpus in CI,
instead of only at OSS-Fuzz. So far only the UBJSON and BJData drivers
had a matching unit test; the CBOR, MessagePack, BSON and BON8 drivers
assert the same invariant (assert(to_X(j2) == vec)) but nothing ran it
outside OSS-Fuzz.
Add "<FORMAT> round-trip invariants" test cases to unit-cbor.cpp,
unit-msgpack.cpp, unit-bson.cpp and unit-bon8.cpp, modeled on the
UBJSON case: seed j1 from the corpus (skipping values that do not
survive the format's own round trip, as the fuzzer drivers only ever
see values from_X() actually produced), then require from_X(to_X(j1))
not to throw and check to_X(j2) == to_X(j1). BSON only serializes
objects, so non-object corpus values are skipped. Update the comments
in round_trip_corpus.hpp and tests/fuzzing.md to name all six formats.
The stream-versus-contiguous check in the BON8 driver is left out, as
#5601 reworks it.
A local probe confirms no violations on the current corpus (CBOR 3849
checked, MessagePack 3909, BSON 2958, BON8 3841 - matching the counts
already recorded for this probe in the issue).
Closes#5714 item 1.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix fuzzer driver step lists to match the checks the code performs
The header comment of six of the seven binary-format fuzzer drivers
listed an invariant the code does not check: CBOR, MessagePack, BSON
and BON8 said "assert(j1 == j2)", but the code checks byte stability,
assert(to_X(j2) == vec). UBJSON and BJData still described the old
"assert(j1 == j2/j3/j4)" byte-exact check from before PR #5494 replaced
it with a use_size/use_type-aware round trip (UBJSON) and a
value-stability check (BJData); BJData's added paragraph already
explained the new check, but the step list above it did not.
Also remove a dead branch in fuzzer-parse_bson.cpp: from_bson() is
called with allow_exceptions = true, so it throws instead of returning
a discarded value, and the "if (j1.is_discarded()) return 0;" guard
could never trigger. Drop the unused <iostream> include from all seven
drivers and <sstream> from all but fuzzer-parse_bon8.cpp, which is the
only one that uses std::istringstream.
Overlaps #5601, which edits all seven drivers in the same hunks.
Closes#5714 item 4.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Silence the CMP0169 deprecation in cmake_fetch_content, fix stale guards
tests/cmake_fetch_content/project calls the single-argument
FetchContent_Populate(json) after FetchContent_Declare(), which CMake
3.30 deprecated as CMP0169. Since the project declares
cmake_minimum_required(VERSION 3.11...3.14), the policy stays unset,
so every configure with a current CMake prints the deprecation
warning. The test is kept on purpose: it is the only coverage of the
FetchContent_Populate + add_subdirectory pattern for CMake 3.11-3.13
users, which the docs still describe as supported. Explicitly set
CMP0169 to OLD, with a comment explaining why.
Also fix two stale version guards:
- tests/cmake_fetch_content/CMakeLists.txt guarded the test with
VERSION_GREATER "3.11.0", which is dead now that tests/CMakeLists.txt
requires CMake 3.13.
- tests/cmake_fetch_content2/CMakeLists.txt guarded with
VERSION_GREATER "3.14.0", which skips exactly 3.14.0, the first
version with FetchContent_MakeAvailable. Change it to
VERSION_GREATER_EQUAL "3.14".
Verified locally: `ctest -R cmake_fetch_content` passes with CMake
4.1, and the CMP0169 deprecation warning that appeared before this
change is gone.
Closes#5714 item 5.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make the CMake integration-test wrappers consistent
The six tests/cmake_* integration-test wrappers had drifted:
- Only cmake_import and cmake_import_minver forwarded
-A "${CMAKE_GENERATOR_PLATFORM}" to the inner configure, and none
forwarded -T "${CMAKE_GENERATOR_TOOLSET}". The Windows workflow
configures the outer build with -A Win32 -T ClangCL, so without
forwarding, the inner projects of cmake_add_subdirectory,
cmake_fetch_content, cmake_fetch_content2 and
cmake_target_include_directories built with the generator defaults
instead of matching the outer build's platform and toolset. Forward
both consistently from all six wrappers.
- cmake_fetch_content and cmake_fetch_content2 passed
-Dnlohmann_json_source to their inner projects, which never read it
(CMake warns "manually-specified variables were not used"); the
inner projects fetch their own copy of the library instead. Drop it.
- tests/CMakeLists.txt set JSON_FORCED_GLOBAL_COMPILE_OPTIONS from the
matching environment variable but never read the cache variable
again; the lines right below it read $ENV{JSON_FORCED_GLOBAL_COMPILE_OPTIONS}
directly, like the LINK_OPTIONS counterpart already does. Remove the
dead set().
This changes which platform and toolset the Win32 and ClangCL CI jobs
build the four newly-forwarding wrappers' inner projects with, which
may surface new failures there; CI has to confirm those jobs.
Verified locally with Ninja (empty -A ""/-T "" is accepted): all 12
cmake_* tests still pass, and the inner fetch_content configures no
longer warn about the unused variable.
Closes#5714 item 6.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Turn the #972 fifo_map regression test into a real test
The #972 regression test in unit-regression1.cpp only built a
my_json array from a string literal (the original crash) and had no
CHECK, so the fifo_map object type it exists to demonstrate was never
exercised. Meanwhile the docs recommend fifo_map for keeping object
keys in insertion order (object_order.md, template_parameters.md),
and nothing tested that recommendation.
Extend the section: after the original array assignment, parse an
object with my_json::parse() (not via the "..."_json UDL, which
returns a plain nlohmann::json and would exercise the cross-basic_json
conversion constructor instead of the parser's own key insertion -
and, as tried locally, does not keep fifo order for this stateful
comparator) and check that dump() keeps insertion order, and that it
survives erase() and inserting a new key.
Also narrow thirdparty/fifo_map off the include path of every other
test-* target: it was a PUBLIC include directory of test_main, even
though unit-regression1.cpp is its only user. Add a small
fifo_map_include INTERFACE library with that include directory and
attach it to test-regression1 only via json_test_set_test_options().
Verified locally (test-regression1_cpp11, default build and
-fsanitize=address,undefined): the new checks pass; `git grep fifo_map
tests` still only finds unit-regression1.cpp and the vendored header.
Closes#5714 item 7.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Document the vendored doctest.h patch; fix stale doctest_compatibility.h comments
tests/thirdparty/doctest/doctest.h is doctest 2.4.12, imported in
#4771. Two weeks later, #4801 hand-edited translateActiveException()
to declare "String res;" inside the translator loop instead of before
it, so a translator that does not match does not leave a previous
translator's result in "res" for the next iteration to see. Nothing
recorded this, so re-vendoring doctest.h from upstream would silently
drop the fix. Add a comment at the patched site naming the version,
the PR and the reason, so a future re-vendor knows to re-apply it.
Also fix two stale comments in doctest_compatibility.h:
- The DOCTEST_THREAD_LOCAL comment referenced Xcode 6/7, which is no
longer supported; reword it to explain why the define must stay
regardless (it keeps doctest's own thread_local usage out of the way
of the same Clang/MinGW crash that JSON_NO_THREAD_LOCAL works around
in the library, see ci_test_no_thread_local).
- The <iosfwd> include's comment justified it with tests that define
"private" as "public"; no test under tests/src does that any more
(removed by #2352). Reword the comment instead of dropping the
include, since confirming it is safe to drop needs the full CI
matrix including MSVC 2015+.
Verified locally that tests/src/unit-readme.cpp still builds and
passes 17/17 with these headers.
Closes#5714 item 9.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop compiling unit-wstring.cpp out entirely on classic ICC
tests/src/unit-wstring.cpp wrapped the whole file in
#ifndef __INTEL_COMPILER, with the comment "ICPC errors out on
multibyte character sequences in source files". The ci_icpc job
(intel/oneapi-hpckit:2023.2.1) still exists, so that job ran none of
the wstring/u16string/u32string input adapter tests, including the
malformed-input checks #5704 (open) extends.
Only 9 lines contained non-ASCII bytes: the three *_is_utf16()/
*_is_utf32() probe functions, and three std::wstring/u16string/
u32string literals plus their narrow-string dump() expectations.
Rewrite all of them with \u/\U escapes in the wide/u16/u32 literals
and \x escapes (split into separate string-literal tokens so a
following byte is never read as part of the same hex escape, e.g.
"\xE1\x83\x85" "a") in the narrow ones. Remove the
#ifndef __INTEL_COMPILER/#endif guard along with it.
The *_is_utf16()/*_is_utf32() probes compared a raw multibyte literal
against an escape-based one to detect a compiler that misreads the
source file's encoding; with no raw literals left to misread, the
comparison is now tautological, so drop the probes and the "if"
guards around each SECTION's body instead of leaving them in as dead
checks.
The same non-ASCII-in-source-and-in-a-narrow-comparison pattern
existed once more in unit-deserialization.cpp's "Using _json with
char8_t literals #4945" test: a raw emoji character in a u8R"(...)"
literal, guarded by a check_utf8() that returned false for ICC (same
reason) and for Windows without the active UTF-8 code page. Rewrite
the literal with a \U escape and compare it against a \x-escaped
expectation instead of a second raw literal, and drop check_utf8()
and the now-unused <windows.h> include along with the guard.
Verified locally (clang, -std=c++11 and -std=c++20,
-fsanitize=address,undefined, and a plain build): test-wstring keeps
18/18 assertions and unit-deserialization keeps 466/466 (c++11) and
477/477 (c++20) assertions, matching this branch before the change
exactly - no coverage was gained or lost, only the source-encoding
dependency was removed. ci_icpc has to confirm classic ICC actually
builds and passes test-wstring now; if it does not, that is a real
finding, not a reason to restore the guard.
Overlaps #5704 (open), which edits unit-wstring.cpp inside the
previously-guarded region (an include near the top, checks in the
invalid-string sections, and a new section at the end).
Closes#5713 item 5.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Share the diagnostic-position setter of the DOM SAX parsers
json_sax_dom_parser and json_sax_dom_callback_parser each had a private
copy of handle_diagnostic_positions_for_json_value(), identical except
for comments. Move the body into one static member function,
detail::diagnostic_positions::set_from_lexer(value, lexer), which both
classes call with their lexer pointer. basic_json befriends the new
struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position
members are private.
The discarded case is reached through the callback parser, so the
LCOV_EXCL markers that only the dom parser's copy had are gone. The
NOLINT on the unreachable default case loses the stray
"-warnings-as-errors", which is not a check name.
The start-position setup in start_object()/start_array() is left alone,
as #5706 is editing the callback parser's versions.
Behavior, the public API and the ABI are unchanged.
Part of #5712
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Correct the parser comments on recursion and skip_to_state_evaluation
The class documentation called the parser a recursive descent parser,
but sax_parse_internal() is a loop that keeps the open containers on an
explicit stack. The comment at the end of an array and of an object
said the flag is set to false while the code below it sets it to true.
Describe what the code does instead.
Comments only; behavior, the public API and the ABI are unchanged.
Part of #5712
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Update the discard_number_values comments to the current number path
The comments explaining the accept() shortcut in convert_number() and
the member documentation still argued in terms of strtoull()/strtoll()
and errno, which #5283 replaced with convert_integer(), and pointed at
scan_number() instead of convert_number(). They also did not say that
scan_number_bulk_contiguous() converts integers itself, so the shortcut
is only reached for input without bulk access, with
JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back.
Rewrite both comments to describe the digit-count check in front of
convert_integer(), keeping the 18-digit bound and the json_sax_acceptor
argument. The stale <cstdlib> comment is left for after #5616, which
edits that include block.
Comments only; behavior, the public API and the ABI are unchanged.
Part of #5712
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* List the UTF-8 validators instead of calling the DFA the only one
The documentation of decode() called the Hoehrmann DFA the single
source of truth for UTF-8 validation. It is used only by the serializer
and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text
strings). The lexer's scan_string() switch, validate_one_utf8() /
valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer)
and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges
on their own.
Replace the sentence with a list of the four validators, what each is
used for, and a note that they must accept the same sequences. Sharing
code between them was considered and dropped: it would save a few lines
in a validator that is entangled with BON8 pushback, and #5677 is
editing the BON8 byte path.
Comments only; behavior, the public API and the ABI are unchanged.
Part of #5712
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix stale doc comments and include lists in the input headers
input_adapters.hpp included <memory> and <numeric> for the removed
shared_ptr-based adapter design but used neither; it called
(std::min) without including <algorithm>. json_sax.hpp used
std::numeric_limits without including <limits>. Also corrected
comments that no longer matched the code: input_stream_adapter does
not skip the input's BOM (the lexer's skip_bom() does), the
span_input_adapter comment named the no-longer-existing
input_buffer_adapter type, lexer::get_string() does not reset the
token, binary_reader's get_number() doc opened with /* instead of
/*! (so Doxygen skipped it) and omitted BON8 from its endianness
note, and the UBJSON-binary-types note did not mention that BJData
'B' arrays are read as binary.
Left out: the lgtm suppression on lexer.hpp's scan_number() (in
#5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's
start_object/start_array docs (in draft #5267's hunk), per the
verdict's conflict list.
Part of #5712
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Deduplicate the strict-EOF/release_lookahead/error block in parser::parse()
json_sax_dom_callback_parser and json_sax_dom_parser branches of
parser::parse() ran the same ~25 lines after sax_parse_internal():
the strict-mode EOF check (raising parse_error.101 through the SAX
parser), release_lookahead() in non-strict mode, and mapping an
errored SAX parser to a discarded result. The two copies had already
drifted apart in formatting and in the second copy's "see above"
comment.
Add a private parse_dom(DomSax&, strict) member that runs this shared
sequence once and returns whether the SAX parser did not error; both
branches of parse() now only construct their DOM SAX parser, call
parse_dom(), and (for the callback parser) map a discarded top-level
value to null. sax_parse() is left untouched, since it only runs the
EOF check and release_lookahead() when sax_parse_internal() succeeded,
unlike parse(), which runs them unconditionally.
Behavior-preserving: same operations in the same order for both SAX
parser kinds. Verified with unit-class_parser (strict/non-strict,
callback and non-callback), unit-deserialization and
unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean
make amalgamate / make check-amalgamation diff.
Overlaps #5601, which touches the same lines.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5712 item 2
* Share the code point to UTF-8 encoding between the wide-string helpers and the lexer
The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three
times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32
code point, in the UTF-16 helper for both a BMP code unit and a valid
surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The
copies had drifted: the UTF-32 helper masked the leading bits of each
byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift
alone, even though both give the same result for a code point that is
already known to be in range.
Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder
that invokes a callable once per output byte, most significant byte
first. Use it in the three valid-code-point branches (UTF-32 code
points up to U+10FFFF, UTF-16 code units outside the surrogate range,
and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where
out forwards to add(). The UTF-16 helper's deliberate pass-through of
malformed surrogate units and the UTF-32 helper's 0xFF sentinel for
code points above U+10FFFF are untouched, since neither reaches the new
helper.
Behavior-preserving: same bytes in the same order for every valid code
point, verified with unit-class_lexer, unit-class_parser,
unit-deserialization, unit-wstring and the non-test-data parts of
unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse
microbenchmark that shows no change (about 73 ms either way, median of
3, 1M escape sequences). single_include/ regenerated with make
amalgamate; make check-amalgamation leaves a clean tree.
Overlaps #5704, which rewrites the wide_string_input_helper
specializations touched here.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
#5712 item 6
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Share scalar serialization between dump_internal and dump_value
dump_value()'s cases for string, binary, boolean, number_integer,
number_unsigned, number_float, discarded and null were a byte-for-byte
copy of dump_internal()'s (added together in #5285 for the iterative
fallback path). Any future change to scalar output had to be made in
both places, or the recursive and depth-limited paths would silently
start producing different bytes.
Extract the shared cases into a private dump_scalar() and have both
dump_internal() and dump_value() call it. Output is unchanged: dump(),
dump(4), dump(-1,' ',true) and the replace/ignore error_handler_t
variants are byte-identical over the json_test_data corpus before and
after, and dump() throughput on a scalar-heavy document is unaffected.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Drop serializer.hpp's dependency on binary_writer.hpp
The only use of binary_writer in serializer.hpp was
binary_writer<BasicJsonType, char>::to_char_type() to write the
U+FFFD replacement character's three bytes. With CharType=char this
is an identity conversion, so the include of binary_writer.hpp (and
transitively binary_reader.hpp) pulled in a large, unrelated header
for a no-op call.
Write the three bytes directly instead. serializer.hpp compiles
standalone with -Wall -Wextra -Werror, with and without
-funsigned-char, and unit-serialization's error_handler_t::replace
cases (with and without ensure_ascii) still pass. Moving
binary_writer's to_char_type/to_msgpack_length to its private section
is left as an optional follow-up.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix stale #include lines in the output headers
output_adapters.hpp included <algorithm> and <iterator> for std::copy
and std::back_inserter, which have not been used there since #3569
(2022). serializer.hpp included <algorithm> for std::reverse (also
unused), <cmath> for labs/isnan/signbit (only std::isfinite is used)
and <utility> for std::move (nothing from <utility> is used there),
while using std::next without including <iterator> at all, relying on
getting it transitively through output_adapters.hpp's own stale
<iterator>.
Drop the unused includes, add <iterator> for std::next, and correct
the remaining include comments. Both headers still compile standalone
with -Wall -Wextra -Werror.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove JSON_HEDLEY_NON_NULL(2) from write_characters() overrides
output_vector_adapter, output_stream_adapter and output_string_adapter
declared their write_characters(const CharType*, std::size_t) override
JSON_HEDLEY_NON_NULL(2), but binary_writer legitimately calls it with
a null pointer and length 0 for an empty string or binary value; the
type-erased call path only stayed silent under UBSan because the
static callee at those call sites is the unattributed virtual base.
A nonnull attribute on a definition lets GCC and Clang assume the
parameter is non-null inside the function body even when the call is
virtual, so this was latent undefined behavior, not just style.
Drop the attribute from the three overrides and document the
(nullptr, 0) contract on output_adapter_protocol::write_characters.
unit-cbor, unit-msgpack, unit-bson and unit-bon8 (which all exercise
empty binary/string payloads through the stream and vector/string
adapters) pass under -fsanitize=address,undefined,nonnull-attribute.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Update stale serializer doc comments to match the current implementation
dump_internal()'s doc block still described the pre-#5285/#5449
implementation: an escape_string() function that does not exist
(the function is dump_escaped), integer conversion "implicitly via
operator<<" (dump_integer actually uses a digit-pair lookup table),
and floating-point conversion via "%g" (IEEE-754 types go through
to_chars, others through snprintf). dump_value()'s comment said
elements are pushed for dump_internal to walk, but it is
dump_iteratively() that walks the stack. dump_escaped(), dump_integer()
and dump_float() each said they write "to output stream @a o", which
has not been true since the writer moved to write_buffer.
Doc-only change; no behavior, API or ABI impact.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Merge duplicate byte-to-hex helper and drop stale '| 0' promotions
serializer::hex_bytes() and binary_writer::hex_byte() had identical
bodies. Keep one, detail::hex_byte() in string_utils.hpp, and use it
from both. Also drop the `| 0` at the two serializer call sites
(hex_bytes(byte | 0) and hex_bytes(s.back() | 0)): #3088 (7440786b8)
added it so that `ss << std::hex << (byte | 0)` printed a number
rather than a char with the old stringstream writer; the int result
just narrows back to uint8_t now, so it was a no-op.
Behavior is unchanged: unit-serialization, unit-bon8 (whose
type_error.316 messages exercise this code) and unit-diagnostics
pass, and both headers still compile standalone with -Werror.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Trim two stale lint suppressions in serializer.hpp
dump_integer()'s `auto buffer_ptr = number_buffer.begin();` carried
NOLINT entries for cppcoreguidelines-pro-type-vararg and hicpp-vararg,
left over from the snprintf-based implementation (#3088); there is no
variadic call on that line, so keep only the qualified-auto
suppressions it actually needs. remove_sign()'s assert checked
`x < 0 && x < (std::numeric_limits<number_integer_t>::max)()) `with a
NOLINT(misc-redundant-expression) to hide it; the second conjunct is
always true once x < 0, and has been since 6ce2f35ba (2019), so
reduce the assert to `x < 0` and drop the suppression instead of
masking it.
Both are documentation-only changes to assertions/suppressions, not
behavior. The to_chars.hpp `#if 0` branch this item also flagged is
left alone, next to draft PR #5634's pending hunk.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move the Hoehrmann SPDX copyright line to string_utils.hpp
serializer.hpp carried the SPDX-FileCopyrightText line for Björn
Hoehrmann's UTF-8 decoder, but the decoder (decode() and the utf8d
table) has lived in string_utils.hpp since #5185 (d19f7f5dc);
serializer.hpp now only calls decode(). Move the copyright line to
where the code it covers actually is.
Part of #5709
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop calling std::localeconv() on every dump()
The serializer constructor snapshotted std::localeconv() into a
locale_chars member on every dump(), even though the only reader is
dump_float(number_float_t, std::false_type)'s snprintf path, taken
only for a number_float_t that is neither IEEE single nor double.
localeconv() is not required to be thread-safe with setlocale(), so
every dump() paid for and raced on a lookup that almost never mattered.
Remove locale_chars and the locale member. Right before the
thousands-separator/decimal-point fixups in the snprintf path, read
std::localeconv() into local thousands_sep/decimal_point variables
(null-checked, first byte only, as before) - the same way
lexer::get_decimal_point() already does since #5597. Output is
unchanged unless the locale changes during a single dump(); in that
case the fixups now match what snprintf just produced, instead of a
value snapshotted before the call.
Overlaps draft PR #5608, which touches the same constructor and
dump_float() lines to move this code into a new
dump_float_snprintf(); this lands the lookup change now as #5709 asks,
and #5608 can do the lookup inside dump_float_snprintf() when it
rebases.
Verification: the full json_test_data corpus (742 files, dump(),
dump(4) and dump(-1,' ',true)) is byte-identical to before the change
under the C locale. Added a test pinning the new per-conversion
lookup: it switches LC_NUMERIC mid-dump() (via a streambuf that
switches on its first write, after the serializer's write buffer has
been flushed once but before a later float is converted) and checks
the decimal point is still normalized using the locale active at
conversion time. On a platform where long double is IEEE-754 double
(e.g. 64-bit Arm), dump_float() takes the locale-independent
to_chars() path and the test is a no-op there; it is meaningful on a
platform where long double is extended precision (most x86 targets).
Part of #5709 item 3
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Diff deeply nested values without recursing per nesting level
diff() descended into both values once per nesting level, and compared
them with operator== on every level on the way, which recurses as well.
Values nested deeply enough - 25,000 levels on an 8 MiB stack - exhausted
the call stack and terminated the process, although parse() accepts
them without complaint. On such a chain the per-level comparisons and
path strings also made diff() quadratic in time and memory.
Both the recursion and operator== only descend as far as the source is
nested. So diff() first checks, recursing at most diff_depth_limit()
(128) levels, whether the source is nested more deeply than that. If not
- all but a vanishing minority of values - the recursive algorithm
diffs it exactly as before, now as diff_recursively(). Otherwise
diff_iteratively() walks the two values on an explicit stack, emitting
the same operations in the same order. It does not compare arrays and
objects with operator== up front (equal ones yield no operations
anyway), keeps the path in one buffer instead of a new string per
level, and hands every subtree that is not nested too deeply back to
diff_recursively(), so equal parts are still skipped quickly.
The check costs one pass over the source. On a 3,000-object document
that is about 30% of diffing two equal values (which is just an
operator== call), about 10% of diffing values that differ in a few
places, and noise when arrays change length. Once operator== no longer
recurses (#5390), the check can go.
Tests check that the patch reproduces the target at every depth up to
300, for json and ordered_json, including reordered members. They also
check the exact operation for a difference deep inside, and diff values
nested 100,000 levels deep.
Fixes#5393 for diff().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Make diff_frame a member struct that declares its special members
GCC's -Weffc++ (an error in CI) asks a class with pointer members, a
user constructor and a non-trivial destructor to declare its copy
constructor and copy assignment; diff_frame's vector and basic_json
members make its destructor non-trivial. Declare all five as defaulted,
which also satisfies clang-tidy's special-member-functions check. Leave
their exception specifications implicit: GCC 4.8 rejects an explicit
one that differs from the implicit one, as it does for flatten_task in
#5517.
The converting constructor cannot throw, and is now declared noexcept
for GCC's -Wnoexcept, which flags the emplace_back() under C++26
otherwise. The struct also moves from diff_iteratively() into the class,
like dump_frame in the serializer.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the shared recursion limit in diff()
diff_depth_limit() is gone in favor of detail::recursion_depth_limit().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Diff fewer nesting depths so the test does not time out under Valgrind
Checking every depth up to 300 made test-json_patch exceed the 1500 s ctest
timeout in ci_test_valgrind. Check the depths up to 16, those around the
recursion limit of 128, and 300 instead.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Mark the diff frame's value-initialized members for clang-tidy
The braces are kept for GCC's -Weffc++, as in json_sax.hpp.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Bound diff()'s descent with a depth count instead of scanning the source
Now that operator== no longer recurses (#5390), diff() can keep its per-level
equality shortcut all the way down. It diffs recursively for the first
detail::recursion_depth_limit() levels, as merge_patch() does, and hands
anything deeper to diff_iteratively(). The nesting_exceeds() scan, which
cost about 30% on equal documents, is gone, and diff() is on par with
develop again.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Note that the diff frame reference is invalidated by pop_back() too
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Keep diff()'s recursive levels small and its result elided
diff_recursively built every patch operation in place from initializer
lists. Unoptimized builds give each of those temporaries its own stack
slot, so every level of the bounded descent cost kilobytes of stack
(about 6 KB with clang -O0), and the 128 recursive levels overflowed the
1 MB stack of MSVC Debug in the "deeply nested values" test. The
operations and the key comparison of two objects are now built by
separate functions, which diff_iteratively shares, and both diff
functions append to one result instead of returning a patch per level
that the caller copies. With clang -O0, diffing values nested 300 levels
deep now peaks at about 190 KB of stack instead of 880 KB.
Since diff() now owns the only returned value, clang's -Wnrvo no longer
reports the returns of diff_recursively, which alternated between the
local patch and diff_iteratively's result.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Copy the diff frame's members instead of holding a reference to it
The loop in diff_iteratively held a reference to the top frame, which
enter() invalidates when it pushes and the end of the loop invalidates
when it pops. Nothing used it afterwards, but a later change could. As in
the other iterative walks, the members the loop reads are now copied out
as constants and the ones it advances are changed through stack.back().
The frame as a whole is not copied: it holds the common keys and the
"add" operations of an object.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The "check" job (Check amalgamation) runs develop's amalgamate.py and read
all configurations from the develop checkout, where config_json_view.json
does not exist until this stack lands, so it failed with
FileNotFoundError. Read that configuration from the pull request's
checkout; the tool itself stays develop's.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- ci_test_noexceptions: the helpers that compare the exceptions of
json_document::parse() and json::parse() catch them outside a
CHECK_THROWS, so with JSON_NOEXCEPTION the first parse error aborted
the test. Compile those comparisons only with exceptions, as
unit-class_parser.cpp does.
- ci_test_gcc: -Werror=unused-result for CHECK_THROWS_AS(json_document::
parse(...)); assign the result to a dummy document.
- ci_test_compilers_clang (3.6): `const json_view invalid;` needs a
user-provided default constructor there (CWG 253); value-initialize it.
- ci_test_single_header: json_view.hpp now exists as a single header and
contains the internal view headers, so unit-json_view_builder.cpp
includes it instead of the detail headers in that mode, and the test
is built again with the single header.
- Regenerate single_include/nlohmann/json_view.hpp for the builder change
merged from json-view/08-view-builder.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A new section describes the 16-byte node: its fields, how integers,
floats, and object members are stored, how views navigate without
pointers, and a worked example. The feature page and the pages of
basic_json_document and node_count link to it where they mention the
16 bytes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The 4 GiB limit and the fallback for an input that parse() accepts but
the view rejects (a bug) are excluded from the coverage; shrink_to_fit()
of an empty document is tested.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- the input dispatch takes byte ranges by const reference and reads the
size once (which also settles a finding of the static analyzer); input
adapters are taken by value
- the classification of inputs keeps its nested conditional operators, a
constant expression of C++11 (NOLINT)
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
benchmarks_view.cpp adds ViewParse, ViewRead (a reused document),
ViewParseIndented, ViewAccept, and ViewMaterialize on the files of
ParseString, so that each row can be read against the json::parse row of
the same file; benchmarks.cpp gains Accept (json::accept) as the
counterpart of ViewAccept. The view benchmarks are built only if the
header directory has json_view.hpp, so that older versions can still be
benchmarked.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- API pages for basic_json_document and basic_json_view, one per member,
and for the four aliases, each with an example
- features/json_view.md: the problem the view solves, ownership and
lifetime, what matches basic_json::parse() and what differs, and when
to choose json, ordered_json, SAX, or the view
- the examples show why one would use the view, not only how: borrowed
vs. owned input, reading a few fields and materializing one subtree,
reusing a document across many messages
- registered in the mkdocs navigation, llms.txt, the docset, the
exceptions page (out_of_range.416), architecture.md, the integration
page, and the README; the yyjson credit is added to the README and
license.md
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
fuzzer-parse_json_view.cpp checks for every input that
json_document::accept agrees with json::accept, that an accepted input
materializes to the value json::parse returns, and that a rejected input
makes both throw the same exception with the same message. It is built
like the other fuzzers (tests/Makefile, and the root Makefile's
fuzz_testing_json_view target, which starts from the JSON test corpus) and
listed in tests/fuzzing.md.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The nlohmann.json module includes json_view.hpp in its global module
fragment and exports basic_json_document, basic_json_view, and the four
aliases next to basic_json, json, and ordered_json. features/modules.md
lists them, and tests/module_cpp20 parses a document, so that a missing
export fails the ci_module_cpp20 job.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The new single header goes through the same checks and install steps as
json.hpp and json_fwd.hpp:
- cmake/ci.cmake: ci_test_amalgamation regenerates, formats, and compares
json_view.hpp as well
- check_amalgamation.yml: the pull request check does the same; it runs
develop's tools, so it needs config_json_view.json on develop first
- meson.build: installs single_include/nlohmann/json_view.hpp
- gen_bazel_build_file.cmake, BUILD.bazel: json_view.hpp joins the
single-header target; the glob of the other target already covers the
new headers
- labeler.yml: an "aspect: json_view" label for the header, its
detail/view headers, tests, and documentation
The CMake install rules for include/ and single_include/, the REUSE
catch-all, and Package.swift cover the new files without changes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
unit-json_view.cpp: type queries, size, and empty against basic_json;
materialize() against parse() (json and ordered_json, generated documents,
duplicate keys, 100,000 levels of nesting, parent pointers with
JSON_DIAGNOSTICS); parse errors and their messages equal to parse() for
malformed inputs and all option combinations; the overflow of a float
document (1e39, 3.4028236e38) as in parse(); NUL and BOM; borrowed and
owned inputs (strings, C strings, literals, vectors, string_view, streams,
wide strings, parse_copy, and iterator ranges over pointers, vectors,
strings, and lists); reuse with read(); moves; shrink_to_fit() of the index
and of the decoded strings; source offsets.
unit-json_view_macros.cpp includes the header without
JSON_TEST_KEEP_MACROS, as users do: the view must not depend on the macros
json.hpp undefines, and must not leak its own.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The public classes of the zero-copy view (#5295), in the new header
<nlohmann/json_view.hpp>:
- basic_json_document<BasicJsonType>: parse (borrowing contiguous byte
inputs, owning rvalue strings, streams, and other inputs), parse_copy,
accept, read, root, is_discarded, source, owns_source, node_count,
memory_usage, shrink_to_fit
- basic_json_view<BasicJsonType>: type and the is_* queries, size, empty,
materialize (the value parse() would produce, built by the same SAX
handler), source_offset
- the aliases json_document, json_view, ordered_json_document, and
ordered_json_view
A parse error throws the exception basic_json::parse would throw for the
same input: the library parser is run on the failing input, so messages,
positions, and exception ids are the same. Inputs of 4 GiB or more are
rejected with out_of_range.416.
The single header single_include/nlohmann/json_view.hpp keeps including
json.hpp; make amalgamate, check-amalgamation, include.zip, and release
handle it.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast (ci_test_gcc on Linux x86-64) rejected
static_cast<std::size_t>(guess + (guess / 4) + 64): the sum is a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Cast a named variable
instead, which GCC does not report. The build stopped at an earlier error
before, so the previous CI run did not show this one.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is
constant) for `TrailingCommas && cur() == ']'` and the like when the
option is off. Route the template arguments through a static enabled()
function, as json.hpp's nesting_depth_exhausted() does.
- ci_test_single_header compiled unit-json_view_builder.cpp against
single_include/, which does not contain the internal
nlohmann/detail/view headers. Build that test only with the multiple
headers.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Six members of string_ref (length, begin, end, operator[], operator!=,
and operator<<) were not reached before C++17, where string_ref is
std::string_view.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Since whole blocks of eight digits are read directly, parse_upto8() only
gets fewer than eight digits; its eight-digit case was dead code.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
A block of eight digits of a number token lies inside the input (the
digits were counted while scanning, or are recorded in the digit
layout), so parse_upto19() reads it without the bounds check of the
last, partial block. Traversing canada.json: -11% instructions, -6%
cycles.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json::parse ends the input at a NUL only between values (where it does
at all); inside a string, a NUL is a control character that must be
escaped. The view reported it as a missing closing quote. (Only the
error code differed: the exception comes from the library parser.)
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Numbers of more than 19 digits at the boundary of the largest double are
decided by the exact comparison with the midpoint in the overflow check.
The check uses the floating-point type of the document: with float, the
view rejects what parse() rejects (1e39, 3.4028236e38, the midpoint between
the largest float and 2^128), and double documents are not affected.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
- tables as std::array; the frames of the first 64 levels stay a C array
(not initialized on purpose, NOLINT)
- \u escapes are decoded with the library's hex_codepoint() instead of a
second table
- the parse failure is private, with an accessor; the special member
functions of the builder are all declared
- no nested conditional operators; explicit parentheses; a repeated
branch body merged; auto for casts
- the test's C arrays, fixed seed, and escaped literals are marked, as in
the other tests
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
With JSON_NOEXCEPTION, NLOHMANN_VIEW_THROW(e) was std::abort() alone, so
the parameters of the functions that build the exceptions were unused, a
warning that the builds with -Werror turn into an error. The exception is
now evaluated before std::abort(); the program ends anyway.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The builder parses JSON text in one pass into the node index: strings and
numbers stay in the source (escaped strings are decoded into an arena),
integers are converted while their digits are in cache, and floats keep
their digit layout for a later conversion. It accepts exactly what
json::parse accepts, for every combination of comments and trailing commas,
with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING.
Parse state lives in a local cursor whose address never escapes, so that it
stays in registers; out-of-line helpers (errors, regrowth, escapes,
comments) are members of the builder and get the positions they need. The
value dispatch is expanded once for array elements and once for member
values. Literals are compared with memcmp and words read in a fixed byte
order, so nothing depends on the platform's byte order. Error messages come
with the public classes.
Tests (unit-json_view_builder.cpp): accept/reject and values against
json::parse for handwritten, generated, and damaged documents under all
option combinations, from std::string and from exact-size buffers (no read
past the input under AddressSanitizer), deep nesting up to 100,000 levels,
NUL/BOM/whitespace cases, and the test-suite files.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
Internal parts of the zero-copy view (#5295), under detail/view and not
included by json.hpp, so that users of json.hpp compile nothing of it:
- macro_scope.hpp/macro_unscope.hpp: the few macros the view needs, under
its own prefix (json.hpp undefines its own at its end); the throw macro
honors JSON_NOEXCEPTION and JSON_THROW_USER like JSON_THROW
- string_ref.hpp: std::string_view from C++17 on, else a small stand-in
- node.hpp: the 16-byte node of the index; its kinds are value_t values
(checked by a static_assert)
- document_data.hpp: the storage of a parsed document (node array, decode
arena, owned input)
- scan.hpp: string and digit scanning with unrolled checks at fixed offsets
(after yyjson) and the library's SWAR and UTF-8 checks, independent of the
byte order
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
json.hpp undefines JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on
the library after it, such as the planned json_view.hpp, reads them from
detail::abi_config instead. The constants live in the ABI namespace, which
already encodes both settings, so they always match the basic_json in use.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.
json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.
Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
GCC -Werror=useless-cast rejected static_cast<std::size_t>(next() % n):
on 64-bit Linux std::uint64_t and std::size_t are the same type, while
the cast is needed where std::size_t is 32 bits wide. Draw the sizes from
a 32-bit value instead, which converts to std::size_t implicitly on every
platform.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
find_string_special() and find_ascii_copyable_run() test eight bytes at a
time, but located the stopping byte inside a word with a byte loop. The
lowest flagged byte of the SWAR tests is always a true hit (the borrows of
the subtractions can only flag bytes above one), so its index is now the
trailing-zero count of the mask; words are read in little-endian order on
every platform, so this does not depend on the byte order.
scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one
after another instead of searching for the next special byte in between,
which helps text in non-Latin scripts.
The kernels serve the lexer's contiguous fast path, the serializer, and the
binary formats. New tests compare all three with byte-by-byte reference
scans on 100,000 generated buffers at three alignments; the portable
fallback of count_trailing_zeros() was checked against the builtin.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
The strtold fallback, which is left only for long double formats that
are not binary64 (x87, binary128), substituted the first byte of the
locale's decimal point for '.'. Under a locale whose decimal point is
longer than one byte, such as fa_IR.UTF-8 or ar_EG.UTF-8 (U+066B),
strtold stopped there and the value was truncated at the decimal point.
A longer decimal point is now put into a copy of the token.
The test "locale with a multi-byte decimal point" now compares the long
double values with those of the "C" locale; with x87 long doubles it
failed before.
Fixes#5660.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple
arm64) are now converted by the library itself, correctly rounded and
independent of the locale and of the C and C++ libraries:
- The token is split into sign, significand w (at most 19 digits), and
decimal exponent q, using the positions of the decimal point and the
exponent that the scanners already recorded, so no character is
classified again.
- Clinger's fast path where w and 10^|q| are exact.
- Eisel-Lemire otherwise, now templated for binary32 and binary64.
- For tokens with more than 19 digits whose w and w + 1 round differently,
an exact big-integer comparison with the midpoint between the two
candidates (the digit comparison of fast_float, simplified).
This replaces the separate token walks of Clinger's fast path and of
Eisel-Lemire, the significant-digit gate that avoided the former, and, for
float and double, std::from_chars and the locale-aware strtod. std::from_chars
and strtold remain only for other long double formats (x87, binary128,
double-double) and for types that are not IEEE-754. Values are bit-identical
to before wherever the previous conversion was correctly rounded; tokens
converted in a locale with a multi-byte decimal point are now also exact.
Overflow still gives out_of_range.406, underflow a signed zero.
convert_float() is the entry point for other parsers of JSON text: it
converts like the lexer, without allocation for binary32/binary64.
Tests: exact-bit tests for double and float (ties, subnormal and overflow
boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for
binary32, the round trips of 200,000 doubles and 100,000 floats without
declines, 508 generated hard cases with the expected bits of both formats
(float_hard_cases.hpp) through the converter and both scanners, and
JSON-level overflow/underflow checks for double and float. The locale tests
now check the values in a locale with a multi-byte decimal point.
Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float
credit now names the digit comparison.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:11 +02:00
407 changed files with 15481 additions and 13889 deletions
@diff BUILD.bazel BUILD.bazel~ ||(echo"===================================================================\n BUILD.bazel is out of date! Please run 'make BUILD.bazel'.\n==================================================================="; mv BUILD.bazel~ BUILD.bazel ;false)
@@ -181,7 +188,7 @@ json.tar.xz:
# We use `-X` to make the resulting ZIP file reproducible, see
Header `<nlohmann/json_view.hpp>` adds `json_document`/`json_view`, a read-only, non-owning way to look at a parsed
JSON text: parsing builds a flat index (16 bytes per value) instead of a tree, strings and numbers stay in the source
text, and `materialize()` builds a `json` value for a subtree only when you actually need one. See
[Zero-copy JSON views](https://json.nlohmann.me/features/json_view/) for the details, including which inputs are
borrowed and which are copied.
## Customers
The library is used in multiple projects, applications, operating systems, etc. The list below is not exhaustive, but the result of an internet search. If you know further customers of the library, please let me know, see [contact](#contact).
@@ -1394,7 +1402,8 @@ THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR I
- The class contains a copy of [Hedley](https://nemequ.github.io/hedley/) from Evan Nemerson which is licensed as [CC0-1.0](https://creativecommons.org/publicdomain/zero/1.0/).
- The class contains parts of [Google Abseil](https://github.com/abseil/abseil-cpp) which is licensed under the [Apache 2.0 License](https://opensource.org/licenses/Apache-2.0).
- The view's parser (`<nlohmann/json_view.hpp>`) contains techniques and code adapted from [yyjson](https://github.com/ibireme/yyjson) by YaoYuan, which is licensed under the [MIT License](https://opensource.org/licenses/MIT) (see above): table-driven decoding of `\u` escapes and fixed-offset unrolled checks.
This function returns `#!cpp true` if and only if the value is a number, i.e. an integer, unsigned integer, or floating-point value. It is defined as `#!cpp is_number_integer() || is_number_float()`.
## Return value
`#!cpp true` if the type is a number, `#!cpp false` otherwise.
## Exception safety
No-throw guarantee: this function never throws exceptions.
## Complexity
Constant.
## Examples
??? example
The example below classifies several parsed documents by the type of their root value, without materializing any
This function returns `#!cpp true` if and only if the value is primitive, i.e. `#!json null`, a boolean, a number, or a string. It is defined as `#!cpp is_null() || is_string() || is_boolean() || is_number()`.
## Return value
`#!cpp true` if the type is primitive, `#!cpp false` otherwise.
## Exception safety
No-throw guarantee: this function never throws exceptions.
## Complexity
Constant.
## Examples
??? example
The example below classifies several parsed documents by the type of their root value, without materializing any
This function returns `#!cpp true` if and only if the value is structured, i.e. an array or an object. It is defined as `#!cpp is_array() || is_object()`.
## Return value
`#!cpp true` if the type is structured, `#!cpp false` otherwise.
## Exception safety
No-throw guarantee: this function never throws exceptions.
## Complexity
Constant.
## Examples
??? example
The example below classifies several parsed documents by the type of their root value, without materializing any
// views are trivially copyable handles (two pointers); the document owns
// the actual data
nlohmann::json_viewcopy=v;
std::cout<<static_cast<bool>(copy)<<'\n';
}
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.