mirror of
https://github.com/nlohmann/json.git
synced 2026-10-01 12:10:32 +00:00
f3e30fa5b7cefa9b63d348777bc52ad4600add9f
5207
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f3e30fa5b7 |
Merge branch 'develop' into claude/binary-utf8-roundtrip-5651
Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e5a89d671f |
Fix lint debt: enum-macro NOLINTs, doctest as SYSTEM, no-op analyzer (#5737)
* Drop stale LCOV_EXCL_LINE from the json_pointer out_of_range.410 throw The comment said the size_type overflow check in array_index() is only triggered on special platforms like 32-bit, and the throw was excluded from coverage. On 64-bit platforms the check is true for SIZE_MAX itself, and unit-json_pointer.cpp has asserted that case four times since #5395, so the line is executed in the coverage job. Reword the comment and remove the exclusion marker so the coverage report notices if the tests stop reaching it. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name all three C-array check aliases in the enum-macro NOLINTs NLOHMANN_JSON_SERIALIZE_ENUM(_STRICT) suppressed the c-array warning under modernize-avoid-c-arrays only, but clang-tidy emits the same diagnostic under the aliases cppcoreguidelines-avoid-c-arrays and hicpp-avoid-c-arrays too. Any user running those checks got a false positive at every macro expansion, and our own tests needed a local NOLINT at each call site to work around it. Name all three aliases in the four macro comments instead, and drop the now-redundant c-array names from the five test call-site NOLINTs. Comment-only change; behavior, the public API, and the ABI do not change. Ran make amalgamate. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Include doctest as a SYSTEM directory instead of disabling warnings for all tests test_main added -Wno-deprecated and -Wno-float-equal as PUBLIC compile options for every non-MSVC compiler, so they were applied to every translation unit, library headers included, and silenced the CI warnings meant to check the library's own -Wfloat-equal pragmas. The only code that actually needed the suppression was the vendored doctest.h, which was included as a normal (non-SYSTEM) directory. Include thirdparty/doctest as SYSTEM for test_main, matching what tests/abi/CMakeLists.txt already does, and drop the two suppressions from both targets. Verified locally that unit-comparison, unit-conversions and unit-constructor1 compile clean with -Werror -Weverything and doctest as -isystem, and that CMake still configures with JSON_BuildTests=ON. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the no-op ci_clang_analyze target ci_clang_analyze configured the build with the real compiler and only then wrapped ninja with scan-build. scan-build intercepts compiles by overriding CC/CXX, but build.ninja already had the compiler path baked in from the configure step, so every run bypassed the analyzer: CI logs show "No bugs found" after a normal build, never an analysis. The job also used Debian's frozen clang-tools-14 rather than the image's own clang, and CLANG_ANALYZER_CHECKS still named three valist.* checkers that current clang merged into security.VAList. ci_clang_tidy already runs every clang-analyzer-* check (via .clang-tidy's "Checks: '*'") with warnings as errors, so nothing is lost by removing the dead job. Delete ci_clang_analyze, CLANG_ANALYZER_CHECKS and the SCAN_BUILD_TOOL lookup from cmake/ci.cmake, drop it from the ubuntu.yml ci_static_analysis_clang matrix, and drop the now-unused clang-tools apt package (iwyu stays for ci_single_binaries). Reword quality_assurance.md and assurance_case.md, which described the dead job as a working control, to say the Clang Static Analyzer checks run through clang-tidy. Verified that `cmake -DJSON_CI=ON` still configures cleanly and that ci_clang_analyze no longer appears in the generated build or in any CMake/workflow file. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Re-enable portability-template-virtual-member-function; remove redundant forwards .clang-tidy disabled three checks "to get the CI going" (#4489, 2024-11-13): portability-template-virtual-member-function, bugprone-use-after-move and its alias hicpp-invalid-access-moved. portability-template-virtual-member-function only flagged output_stream_adapter::write_character/write_characters; annotate both with NOLINT and re-enable the check. bugprone-use-after-move flagged several double forwards that have no effect at runtime: - from_json.hpp calls std::forward<BasicJsonType>(j).at(Idx) inside pack expansions; at() has no ref-qualified overloads and always returns an lvalue reference, so the forward is a no-op. Replace with plain j.at(Idx) in all four places. - the move constructor forwards the whole object to its base class and then reads other's members. That is item 9 of #5724 (together with its cppcheck suppressions) and is left to that change. - input_adapters.hpp forwards the container twice on purpose, so the begin/end iterator types match adapter_type; annotate with NOLINT and a comment instead of changing behavior. The check still flags the move constructor (see above) and two sites in at(KeyType&&) (both overloads, json.hpp, in the throw's string_t(std::forward<KeyType>(key)) after find(std::forward<KeyType>(key))). Open PR #5689 rewrites that hunk, so bugprone-use-after-move (and hicpp-invalid-access-moved) stay disabled for now, with a comment explaining why; re-enable them once #5689 and the #5724 move-constructor change have landed. Also resolve the portability-avoid-pragma-once TODO: single_include never has #pragma once (amalgamate.py strips it) and every supported compiler accepts it in include/, so keep it disabled with an explanatory comment instead of a TODO. Fix the stale "json.hpp, around line 1265" comment in unit-class_parser.cpp, which now points at the move constructor's actual line. Behavior, the public API and the ABI do not change. Verified with clang-tidy 22.1.8 that portability-template-virtual-member-function now reports nothing, that bugprone-use-after-move/ hicpp-invalid-access-moved report only the known at(KeyType&&) and move-constructor sites, and that unit-custom-base-class, unit-constructor1, unit-conversions, unit-element_access2, unit-class_parser and unit-diagnostic-positions (JSON_DIAGNOSTIC_POSITIONS=1) compile under ASan/UBSan and pass with the same assertion counts as before. Ran make amalgamate. Part of #5725 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale and malformed NOLINT comments json_sax.hpp named "-warnings-as-errors" in the NOLINT list on the two JSON_ASSERT(false) lines; that is the suffix clang-tidy appends to a diagnostic tag under WarningsAsErrors, not a check name, and every other JSON_ASSERT(false) omits it. unit-capacity.cpp carried 30 "// NOLINT(misc-const-correctness)" comments on "json j = ...;" declarations that are all used with non-const members afterwards, so the check has nothing to report there. unit-constructor2.cpp used a blanket "// NOLINT: access after move is OK here" on a use-after-move that hides every check on the line; naming bugprone-use-after-move and hicpp-invalid-access-moved keeps the intent once those checks are re-enabled (#5724). Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 10 * Remove stale .clang-tidy entries -google-runtime-references disabled a check that neither clang-tidy 22.1.8 nor 23.1.2 lists under --list-checks -checks='*'; it was removed upstream. The commented-out HeaderFilterRegex line has been unused since the active HeaderFilterRegex was introduced in #2561 (2021). Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 11 * Remove the GCC C++20 -Wignored-attributes pragma in json.hpp The pragma (added in #5164) claimed to work around the C++ modules redefinition errors of #5103, but #5103 is about hard errors (e.g. "redefinition of std::__is_constant_evaluated()", conflicting std::integral_constant) that ignoring a warning cannot suppress; they are traced to GCC PR 124430 and reproduce with <map> or <string> instead of json.hpp too. A GCC 16.2 -std=gnu++20 -fmodules build following #5103's repro steps still fails with the pragma in place, and a build of all test TUs with GCC_CXXFLAGS (which enable -Wignored-attributes) and the pragma removed produces no such warning. The block only hid a warning class from GCC C++20 users while suggesting #5103 was handled. Overlaps #5610, whose hunks touch the closing half of this pragma to insert the json_literals.hpp include. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 9 * Fix stale doxygen comments hidden by the -Wdocumentation pragma macro_scope.hpp ignores -Wdocumentation and -Wdocumentation-unknown-command for the whole library, which also hides genuine documentation mistakes: - detail::unescape() documented "@return unescaped string" but returns void and unescapes its argument in place; reworded to "@param[in,out] s string to unescape in place" and dropped the bogus @return. - basic_json::get()'s copy-conversion overload wrote "converted to @tparam ValueType" inside @return, which Doxygen and Clang parse as a second, malformed @tparam; changed to "@a ValueType", matching the two other get() overloads a few lines above that already use it. This narrows the gap the -Wdocumentation pragma needs to cover; fully replacing the Doxygen-only commands it also hides (item 2c) is left for after #5267. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 2 * Fix -Wextra-semi-stmt at its actual source, not assert() clang_flags.cmake blamed the global -Wno-extra-semi-stmt on assert(), but assert() expands to an expression under glibc and libc++ and does not trigger this warning. unit-assert_macro.cpp overrides JSON_ASSERT with "{if (!(x)) ++assert_counter; }", a bare block followed by a semicolon at every JSON_ASSERT(...) call site in the library; that was the actual source of 151 of the 208 -Wextra-semi-stmt sites found in a Clang 22 -Weverything sweep of the test suite with the flag removed. Switched to the standard do/while(false) macro idiom, which does not expand to a statement-plus-semicolon, and corrected the comment to name the remaining source instead: vendored Doctest's CAPTURE(x) shim, which already ends in a semicolon. Verified with clang++ -Wextra-semi-stmt (plus the file's other CI ignores) that unit-assert_macro.cpp now compiles without any -Wextra-semi-stmt diagnostic. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 8 (step 1 of 2; step 2 covers the CAPTURE() call sites) * Drop the redundant semicolon from CAPTURE() call sites; remove -Wno-extra-semi-stmt doctest_compatibility.h defines CAPTURE(x) as DOCTEST_CAPTURE(x); (with a trailing semicolon baked into the macro), specifically so call sites do not need to add one themselves; most of the ~267 call sites already follow that convention. The remaining 64 call sites across 20 files wrote "CAPTURE(x);" anyway, turning into a statement plus an empty statement and triggering -Wextra-semi-stmt. Dropped the redundant semicolon at each of those sites. With item 6 having already made vendored Doctest a SYSTEM include, and this the last known source of -Wextra-semi-stmt findings, removed the flag from clang_flags.cmake entirely. Verified with clang++ -Wextra-semi-stmt (plus the file's other CI ignores) that all 20 touched files, plus a file with no CAPTURE() use (unit-json_pointer.cpp), compile without any -Wextra-semi-stmt diagnostic. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5725 item 8 (step 2 of 2) * Switch ci_static_analysis_clang off the frozen LLVM 22 dev image ubuntu.yml pinned the clang-tidy/clang-tidy-sanitizer/single-binaries job to silkeh/clang:dev, a tag last pushed 2026-02-18 that reports "clang version 22.0.0 (...+20251015...)", a pre-release snapshot from before the LLVM 22 release; the maintainer now updates dev-unstable, 22, and latest instead. Switched to silkeh/clang:22, matching the other clang jobs on :latest. Verified with clang-tidy 22.1.8 (the image's actual version) against this repository's .clang-tidy and library headers what the release image newly reports compared to :dev: - readability-redundant-typename fires at ~250 sites across the _cpp20-relevant conversion/to_chars headers; the library targets C++11 and keeps the typenames, so the check is disabled in .clang-tidy, matching how the file already handles checks that don't fit a C++11 codebase. - misc-anonymous-namespace-in-header fires on the two anonymous namespaces in from_json.hpp and to_json.hpp; added the alias to their existing NOLINT (cert-dcl59-cpp, fuchsia-header-anon-namespaces, google-build-namespaces). - bugprone-std-namespace-modification fires on every addition to namespace std: the std::hash, std::formatter and std::swap overloads in json.hpp, and the std::tuple_size/std::tuple_element specializations in iteration_proxy.hpp (this last file is not named in #5725's item 5, found by actually running clang-tidy 22.1.8 against the current tree). All six are legal, deliberate additions to namespace std (explicit/partial specializations of std types, or the pre-C++20 std::swap overload); annotated each with the check name next to its existing cert-dcl58-cpp NOLINT. - modernize-avoid-c-style-cast reported nothing new. Also added clang++-22/21, clang-tidy-22/21, g++-16 and gcov-16 to the find_program search lists in ci.cmake so a local "maximal warnings" configure prefers the current toolchain version over an older one on PATH. #5725 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Regenerate cmake/gcc_flags.cmake for GCC 16.2.0 GCC_CXXFLAGS was generated for GCC 15.1.0, but ci_test_gcc and ci_test_gcc_cxx{11..26} now run in gcc:latest, currently GCC 16.2.0, so the "maximal warnings" job was missing warnings introduced since 15.1.0 while carrying entries GCC 16 treats as duplicates or no-ops. Regenerated with https://github.com/nlohmann/gcc_flags (patched locally to not crash on an option whose "-x c++ <opt> -" probe fails before it reads stdin, e.g. -Wabi=; the tool otherwise raises BrokenPipeError instead of recording the option as an error) run against g++ 16.2.0 in the official gcc:16 Docker image, keeping the documented -Wno-* exclusions and the same alphabetical placement scheme as before. Also added three GCC 16 warnings the generator cannot discover on its own because it only probes value ranges/lists it finds in the -Q option name itself, not in the enum choices --help=warnings documents separately: - -Wbidi-chars=any, -Wleading-whitespace=spaces: manually verified these compile cleanly with g++ 16.2.0. - -Wstrict-flex-arrays: deliberately NOT added, unlike the other two. Without -fstrict-flex-arrays (which the library does not enable, as it would change codegen for flexible array members), GCC prints "'-Wstrict-flex-arrays' is ignored when '-fstrict-flex-arrays' is not present" on every translation unit, and under our -Werror that note itself aborts the build. This differs from the harmless no-op warnings already kept in the file (-Whsa, -Wsynth, -Wunreachable-code, -Wunsafe-loop-optimizations), which emit nothing; #5725 item 7 named -Wstrict-flex-arrays as one of the flags GCC 16 adds, but did not anticipate this failure mode. Verified: compiled the library header and a representative set of test translation units (including ones touched by items 1, 3, 8, 9, 10 of this issue) with the regenerated GCC_CXXFLAGS plus -Werror under g++ 16.2.0 at -std=c++11 through -std=c++26, with zero warnings; ran the full local test suite (129/129 passing, unrelated to this compiler) as a regression check. CI must still confirm the actual ci_test_gcc / ci_test_standards_gcc targets end to end, since this was verified with direct g++ invocations rather than through the CMake/ CXXFLAGS environment-variable plumbing in ci.cmake. #5725 item 7 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Avoid std::basic_string<CharType> for non-character output_adapter CharType output_adapter<CharType, StringType> defaulted StringType to std::basic_string<CharType>, and (with JSON_NO_IO undefined) always declared a std::basic_ostream<CharType>&-taking constructor. For CharType with no non-deprecated std::char_traits specialization (only std::uint8_t is ever used this way, by the binary writers), simply naming either type - as an unused default template argument, or as an unused, never-called constructor's parameter type - instantiates std::char_traits<CharType> merely to name it, which some standard libraries mark deprecated: with the library-wide -Wdocumentation pragma (item 2's other half, left for a later commit) temporarily removed, an Apple clang 21 / libc++ TU calling json::to_cbor(j, vec) with std::vector<std::uint8_t>& got one -Wdeprecated-declarations warning per binary writer at the old output_adapters.hpp:193. Replaced the eager std::basic_string<CharType> / std::basic_ostream <CharType> defaults with a bool-tagged partial specialization (not std::conditional, which requires naming both branches' types up front regardless of which is selected, reproducing the same warning) that only ever names std::basic_string<CharType> / std::basic_ostream <CharType> when CharType is actually one of char, wchar_t, char16_t, char32_t, or (with __cpp_lib_char8_t) char8_t. For any other CharType, output_adapter's StringType and ostream-constructor parameter fall back to two distinct empty placeholder types, kept distinct so the two constructor overloads do not collide into a single redeclaration. Public API / behavior: passing a std::basic_string<std::uint8_t>& or std::basic_ostream<std::uint8_t>& directly to a binary writer's output_adapter now fails to compile instead of compiling with a deprecation warning; this was neither documented nor tested. All documented uses (std::vector<CharType>, std::basic_ostream<CharType> and StringType for character CharType) are unaffected. Verified with Apple clang 21 / libc++, with the two -Wdocumentation* "ignored" pragma lines in macro_scope.hpp temporarily removed and -std=c++11/c++20 plus the project's -Weverything flag set: calling to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson/to_bon8 on a std::vector<std::uint8_t> now produces no char_traits<unsigned char> (or any other) deprecation warning, while the char-based string- and ostream-adapter paths, and a to_cbor/from_cbor round trip, still compile and run correctly; also verified with GCC 16.2.0. Ran the full local test suite, including the binary-format unit tests (unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bson, unit-bon8, unit-binary_writer_sinks, unit-binary_formats, unit-custom-binary-type): 129/129 passing. #5725 item 2 (step a) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the library-wide -Wdocumentation pragma; fix what it hid macro_scope.hpp / macro_unscope.hpp pushed and popped a Clang diagnostic region over the entire library that ignored -Wdocumentation and -Wdocumentation-unknown-command. Removed both pragmas and fixed every finding a full -Wdocumentation (which implies -Wdocumentation-unknown-command and -Wdocumentation-deprecated-sync) build reports, so the library now compiles clean under Clang's documentation checks without a blanket suppression. Overlaps #5267, which is still open and edits a nearby doc block (json.hpp's get()/get_impl() @return, already fixed in the item 2 step (b) commit of this branch); this commit does not touch that block again. Unknown Doxygen alias commands (Doxyfile removed in #3071, so these were never rendered by anything) rewritten as plain prose, keeping the same information: - @requirement REQ-JSON-01 / REQ-JSON-02 (iter_impl.hpp, json_reverse_iterator.hpp): now "This class satisfies the following concept requirements (REQ-JSON-0N):". - @liveexample{prose,example-id} (three sites in json.hpp): kept the prose, dropped the command wrapper and the trailing example-id (docs/mkdocs/docs/examples/*.cpp still exist and are used directly by the rendered docs, not through this in-header alias) and unescaped the "\," commas that were only needed for the old alias's comma-separated argument syntax. - @complexity X (json.hpp x4, json_pointer.hpp x2, serializer.hpp x1): now "Complexity: X". Backslash sequences Clang's comment lexer tried to parse as commands, escaped to render as literal backslashes: - lexer.hpp get_codepoint(): two `\u` occurrences. - binary_reader.hpp get_bson_cstr() / get_bson_cstr_bulk(): two `\x00` occurrences. - serializer.hpp: three `\uXXXX` occurrences (constructor @param, append_codepoint_to_string_buffer() @brief, and the ensure_ascii member comment). One finding remained after all of the above: Clang reports "declaration is marked with '@deprecated' command but does not have a deprecation attribute" on the deprecated sax_parse(span_input_adapter&&, ...) overload, even though JSON_HEDLEY_DEPRECATED_FOR does expand to __attribute__((deprecated(...))) for Clang. Several isolated reproductions of this exact declaration shape - doc comment, template<>, two stacked __attribute__ macros, an overload set sharing the name - did not reproduce the warning, so this looks like a Clang comment/declaration-association quirk specific to this overload inside the much larger basic_json class template, not an actual documentation defect. Rather than keep the pragma library-wide for one Clang false positive, added a tightly scoped -Wdocumentation-deprecated-sync push/pop around just that overload. Verified with Apple clang 21 and the project's actual -Weverything flag set (cmake/clang_flags.cmake) on the full header at -std=c++11 and -std=c++20: zero -Wdocumentation* diagnostics. Also compiled clean with GCC 16.2.0 (the pragmas are already __clang__-gated, so this only confirms no unrelated breakage). Ran make check-amalgamation and the full local test suite: 129/129 passing. #5725 item 2 (step c) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Take the JSON value by const reference in the array and tuple from_json paths Review feedback on #5737 (gregmarr): once the no-op std::forward calls are gone, the forwarding references have no purpose. from_json_fn passes the value as const BasicJsonType&, so these functions were only ever instantiated with a const lvalue anyway. The std::array, std::pair and std::tuple overloads of from_json and their helpers now take const BasicJsonType& and pass j on unchanged. Because the deduced BasicJsonType is now the plain type, tuple_type and the static_assert name const BasicJsonType& explicitly, so the reference checks are unchanged: get<std::tuple<const std::string&>>() still works, and get<std::tuple<std::string&>>() still fails the same static_assert. from_json_tuple_get_impl keeps its forwarding reference, since tuple_type calls it through std::declval. Behavior, the public API and the ABI do not change. unit-conversions, unit-constructor1, unit-udt, unit-udt_macro, unit-regression1/2/3, unit-deserialization, unit-noexcept, unit-items, unit-allocator, unit-custom-object-type, unit-ordered_json2 and unit-brace-init-copy-semantics pass at C++11, C++17 and C++20 with unchanged assertion counts. Ran make amalgamate. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
b54ed188e6 |
Remove test debt: dead guards, discarded results, and unreferenced files (#5732)
* Run the README test case in JSON_FastTests jobs The "README" test case was marked doctest::skip() when the tests moved from Catch to doctest in 2019, where it replaced Catch's hidden tag. It is not slow (17 assertions, about 0.00 s), but cmake/test.cmake only passes --no-skip when JSON_FastTests is off, so the per-compiler ci_test_*_cxxNN matrix, macOS, Windows Release/ARM, icpc, icpx and nvhpc compiled the README examples without running them. Drop the skip decorator so every job runs the case. Test-only change. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove stale clang ranges guards in unit-iterators2.cpp The "algorithms" and "views" sections were guarded by clang/libstdc++ checks written for a clang 15 (04/2022) bug. The first guard's condition contradicts its own comment: it skips clang+libc++ and keeps clang+libstdc++. Both sections already sit inside `#if JSON_HAS_RANGES`, which macro_scope.hpp excludes for the toolchains these guards targeted, so the inner guards never let the sections run on the platforms they meant to protect and are redundant on the rest. Verified locally with Apple clang 21/libc++ and clang 16.0.6/libstdc++ 12 (Docker): both pass all 1355 assertions with the guards removed. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix copy-pasted CBOR half-float checks; enable stale encode checks In the RFC 8949 Appendix A test case, the decode checks for 5.960464477539063e-8 (0xf9 0x00 0x01) and 0.00006103515625 (0xf9 0x04 0x00) were copy-pasted from the neighboring -4.0 example, so those two half-float byte sequences were never actually decoded and checked, and -4.0 was checked three times instead. The two float32 encode checks for 100000.0 and 3.4028234663852886e+38 were commented out before the writer supported emitting float32 and are now verified to match byte for byte, so they are enabled. The remaining commented-out half-precision to_cbor checks are collapsed into a single explanatory comment, since the writer never emits half-precision floats. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Assert on the result of STL container conversions in tests The "object-like STL containers" and "array-like STL containers" sections converted json values into std::map, unordered_map, multimap, unordered_multimap, list, forward_list, array, valarray, vector, deque, set and unordered_set and discarded the result, so these ~60 conversions only proved that the code compiles and does not throw; a conversion that dropped or reordered elements would still pass. Bind each result and compare it against the expected container. Also fix a copy-paste slip in the deque section (`j2.get<std::deque<double>>()` instead of j3, so j3's doubles were never converted to a deque), and remove the dead `// CHECK(m5["one"] == "eins")` comments that referred to a variable that did not exist by asserting the equivalent through the bound result. Part of #5713 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate SaxCountdown and other test helpers across formats SaxCountdown was copied byte-for-byte into six binary-format test files (unit-cbor.cpp, unit-msgpack.cpp, unit-ubjson.cpp, unit-bjdata.cpp, unit-bon8.cpp, unit-bson.cpp), about 370 redundant lines. Move it into tests/src/sax_countdown.hpp (namespace utils, alongside test_utils.hpp and round_trip_corpus.hpp) and include it from all six. trait_test_arg and the "value_in_range_of trait" TEST_CASE_TEMPLATE_DEFINE were duplicated between unit-32bit.cpp and unit-bjdata.cpp; the trait is a detail/meta trait, not specific to either file. Move it into tests/src/value_in_range_of_test.hpp; unit-32bit.cpp keeps its own include, since JSON_32bitTest=ONLY builds only that file. Each file keeps its own TEST_CASE_TEMPLATE_INVOKE list. sax_no_exception and the "issue #2824" section were duplicated in unit-regression2.cpp and unit-disabled_exceptions.cpp. Drop the copy from unit-regression2.cpp; unit-disabled_exceptions.cpp already covers the no-exceptions case that #2824 was about, and ci_test_noexceptions reruns it. No behavior change. Verified by building and running unit-cbor, unit-msgpack, unit-ubjson, unit-bjdata, unit-bon8, unit-bson, unit-32bit, unit-regression2 and unit-disabled_exceptions against include/ (clang++ -std=c++11, ASan/UBSan where applicable); assertion counts are unchanged from before the refactor. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the unreferenced vendored libFuzzer tests/thirdparty/Fuzzer (155 files, ~776 KB of vendored Apache-2.0 LLVM code from the 2016 OSS-Fuzz import) is not referenced by any CMakeLists, Makefile or workflow: the fuzz drivers link against -fsanitize=fuzzer or the repo's own tests/src/fuzzer-driver_afl.cpp. Its vendored README only points at llvm.org's own libFuzzer docs. Being dead code, it also adds noise to the flawfinder code-scanning workflow, which scans the whole tree. Remove the directory and its .reuse/dep5 entry. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove unreferenced 2016 benchmark and fuzz reports tests/reports (1.6 MB) holds AFL status pages and plots from 2016-08-29 and 2016-10-02, and a nativejson-benchmark snapshot from 2016 with links to rawgit.com, which shut down in 2019. Nothing references this directory: no doc, README section, script or workflow points at it, and it describes a ten-years-old, pre-2.0 snapshot of the library. Part of #5714 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Run the CBOR, MessagePack, BSON and BON8 round-trip invariants in CI tests/src/round_trip_corpus.hpp exists so that the byte-stability invariant the fuzzer drivers check also runs on a fixed corpus in CI, instead of only at OSS-Fuzz. So far only the UBJSON and BJData drivers had a matching unit test; the CBOR, MessagePack, BSON and BON8 drivers assert the same invariant (assert(to_X(j2) == vec)) but nothing ran it outside OSS-Fuzz. Add "<FORMAT> round-trip invariants" test cases to unit-cbor.cpp, unit-msgpack.cpp, unit-bson.cpp and unit-bon8.cpp, modeled on the UBJSON case: seed j1 from the corpus (skipping values that do not survive the format's own round trip, as the fuzzer drivers only ever see values from_X() actually produced), then require from_X(to_X(j1)) not to throw and check to_X(j2) == to_X(j1). BSON only serializes objects, so non-object corpus values are skipped. Update the comments in round_trip_corpus.hpp and tests/fuzzing.md to name all six formats. The stream-versus-contiguous check in the BON8 driver is left out, as #5601 reworks it. A local probe confirms no violations on the current corpus (CBOR 3849 checked, MessagePack 3909, BSON 2958, BON8 3841 - matching the counts already recorded for this probe in the issue). Closes #5714 item 1. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix fuzzer driver step lists to match the checks the code performs The header comment of six of the seven binary-format fuzzer drivers listed an invariant the code does not check: CBOR, MessagePack, BSON and BON8 said "assert(j1 == j2)", but the code checks byte stability, assert(to_X(j2) == vec). UBJSON and BJData still described the old "assert(j1 == j2/j3/j4)" byte-exact check from before PR #5494 replaced it with a use_size/use_type-aware round trip (UBJSON) and a value-stability check (BJData); BJData's added paragraph already explained the new check, but the step list above it did not. Also remove a dead branch in fuzzer-parse_bson.cpp: from_bson() is called with allow_exceptions = true, so it throws instead of returning a discarded value, and the "if (j1.is_discarded()) return 0;" guard could never trigger. Drop the unused <iostream> include from all seven drivers and <sstream> from all but fuzzer-parse_bon8.cpp, which is the only one that uses std::istringstream. Overlaps #5601, which edits all seven drivers in the same hunks. Closes #5714 item 4. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Silence the CMP0169 deprecation in cmake_fetch_content, fix stale guards tests/cmake_fetch_content/project calls the single-argument FetchContent_Populate(json) after FetchContent_Declare(), which CMake 3.30 deprecated as CMP0169. Since the project declares cmake_minimum_required(VERSION 3.11...3.14), the policy stays unset, so every configure with a current CMake prints the deprecation warning. The test is kept on purpose: it is the only coverage of the FetchContent_Populate + add_subdirectory pattern for CMake 3.11-3.13 users, which the docs still describe as supported. Explicitly set CMP0169 to OLD, with a comment explaining why. Also fix two stale version guards: - tests/cmake_fetch_content/CMakeLists.txt guarded the test with VERSION_GREATER "3.11.0", which is dead now that tests/CMakeLists.txt requires CMake 3.13. - tests/cmake_fetch_content2/CMakeLists.txt guarded with VERSION_GREATER "3.14.0", which skips exactly 3.14.0, the first version with FetchContent_MakeAvailable. Change it to VERSION_GREATER_EQUAL "3.14". Verified locally: `ctest -R cmake_fetch_content` passes with CMake 4.1, and the CMP0169 deprecation warning that appeared before this change is gone. Closes #5714 item 5. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Make the CMake integration-test wrappers consistent The six tests/cmake_* integration-test wrappers had drifted: - Only cmake_import and cmake_import_minver forwarded -A "${CMAKE_GENERATOR_PLATFORM}" to the inner configure, and none forwarded -T "${CMAKE_GENERATOR_TOOLSET}". The Windows workflow configures the outer build with -A Win32 -T ClangCL, so without forwarding, the inner projects of cmake_add_subdirectory, cmake_fetch_content, cmake_fetch_content2 and cmake_target_include_directories built with the generator defaults instead of matching the outer build's platform and toolset. Forward both consistently from all six wrappers. - cmake_fetch_content and cmake_fetch_content2 passed -Dnlohmann_json_source to their inner projects, which never read it (CMake warns "manually-specified variables were not used"); the inner projects fetch their own copy of the library instead. Drop it. - tests/CMakeLists.txt set JSON_FORCED_GLOBAL_COMPILE_OPTIONS from the matching environment variable but never read the cache variable again; the lines right below it read $ENV{JSON_FORCED_GLOBAL_COMPILE_OPTIONS} directly, like the LINK_OPTIONS counterpart already does. Remove the dead set(). This changes which platform and toolset the Win32 and ClangCL CI jobs build the four newly-forwarding wrappers' inner projects with, which may surface new failures there; CI has to confirm those jobs. Verified locally with Ninja (empty -A ""/-T "" is accepted): all 12 cmake_* tests still pass, and the inner fetch_content configures no longer warn about the unused variable. Closes #5714 item 6. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Turn the #972 fifo_map regression test into a real test The #972 regression test in unit-regression1.cpp only built a my_json array from a string literal (the original crash) and had no CHECK, so the fifo_map object type it exists to demonstrate was never exercised. Meanwhile the docs recommend fifo_map for keeping object keys in insertion order (object_order.md, template_parameters.md), and nothing tested that recommendation. Extend the section: after the original array assignment, parse an object with my_json::parse() (not via the "..."_json UDL, which returns a plain nlohmann::json and would exercise the cross-basic_json conversion constructor instead of the parser's own key insertion - and, as tried locally, does not keep fifo order for this stateful comparator) and check that dump() keeps insertion order, and that it survives erase() and inserting a new key. Also narrow thirdparty/fifo_map off the include path of every other test-* target: it was a PUBLIC include directory of test_main, even though unit-regression1.cpp is its only user. Add a small fifo_map_include INTERFACE library with that include directory and attach it to test-regression1 only via json_test_set_test_options(). Verified locally (test-regression1_cpp11, default build and -fsanitize=address,undefined): the new checks pass; `git grep fifo_map tests` still only finds unit-regression1.cpp and the vendored header. Closes #5714 item 7. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Document the vendored doctest.h patch; fix stale doctest_compatibility.h comments tests/thirdparty/doctest/doctest.h is doctest 2.4.12, imported in #4771. Two weeks later, #4801 hand-edited translateActiveException() to declare "String res;" inside the translator loop instead of before it, so a translator that does not match does not leave a previous translator's result in "res" for the next iteration to see. Nothing recorded this, so re-vendoring doctest.h from upstream would silently drop the fix. Add a comment at the patched site naming the version, the PR and the reason, so a future re-vendor knows to re-apply it. Also fix two stale comments in doctest_compatibility.h: - The DOCTEST_THREAD_LOCAL comment referenced Xcode 6/7, which is no longer supported; reword it to explain why the define must stay regardless (it keeps doctest's own thread_local usage out of the way of the same Clang/MinGW crash that JSON_NO_THREAD_LOCAL works around in the library, see ci_test_no_thread_local). - The <iosfwd> include's comment justified it with tests that define "private" as "public"; no test under tests/src does that any more (removed by #2352). Reword the comment instead of dropping the include, since confirming it is safe to drop needs the full CI matrix including MSVC 2015+. Verified locally that tests/src/unit-readme.cpp still builds and passes 17/17 with these headers. Closes #5714 item 9. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop compiling unit-wstring.cpp out entirely on classic ICC tests/src/unit-wstring.cpp wrapped the whole file in #ifndef __INTEL_COMPILER, with the comment "ICPC errors out on multibyte character sequences in source files". The ci_icpc job (intel/oneapi-hpckit:2023.2.1) still exists, so that job ran none of the wstring/u16string/u32string input adapter tests, including the malformed-input checks #5704 (open) extends. Only 9 lines contained non-ASCII bytes: the three *_is_utf16()/ *_is_utf32() probe functions, and three std::wstring/u16string/ u32string literals plus their narrow-string dump() expectations. Rewrite all of them with \u/\U escapes in the wide/u16/u32 literals and \x escapes (split into separate string-literal tokens so a following byte is never read as part of the same hex escape, e.g. "\xE1\x83\x85" "a") in the narrow ones. Remove the #ifndef __INTEL_COMPILER/#endif guard along with it. The *_is_utf16()/*_is_utf32() probes compared a raw multibyte literal against an escape-based one to detect a compiler that misreads the source file's encoding; with no raw literals left to misread, the comparison is now tautological, so drop the probes and the "if" guards around each SECTION's body instead of leaving them in as dead checks. The same non-ASCII-in-source-and-in-a-narrow-comparison pattern existed once more in unit-deserialization.cpp's "Using _json with char8_t literals #4945" test: a raw emoji character in a u8R"(...)" literal, guarded by a check_utf8() that returned false for ICC (same reason) and for Windows without the active UTF-8 code page. Rewrite the literal with a \U escape and compare it against a \x-escaped expectation instead of a second raw literal, and drop check_utf8() and the now-unused <windows.h> include along with the guard. Verified locally (clang, -std=c++11 and -std=c++20, -fsanitize=address,undefined, and a plain build): test-wstring keeps 18/18 assertions and unit-deserialization keeps 466/466 (c++11) and 477/477 (c++20) assertions, matching this branch before the change exactly - no coverage was gained or lost, only the source-encoding dependency was removed. ci_icpc has to confirm classic ICC actually builds and passes test-wstring now; if it does not, that is a real finding, not a reason to restore the guard. Overlaps #5704 (open), which edits unit-wstring.cpp inside the previously-guarded region (an include near the top, checks in the invalid-string sections, and a new section at the end). Closes #5713 item 5. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
cff0a61369 |
Share DOM SAX position handling; fix stale parser and lexer comments (#5731)
* Share the diagnostic-position setter of the DOM SAX parsers json_sax_dom_parser and json_sax_dom_callback_parser each had a private copy of handle_diagnostic_positions_for_json_value(), identical except for comments. Move the body into one static member function, detail::diagnostic_positions::set_from_lexer(value, lexer), which both classes call with their lexer pointer. basic_json befriends the new struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position members are private. The discarded case is reached through the callback parser, so the LCOV_EXCL markers that only the dom parser's copy had are gone. The NOLINT on the unreachable default case loses the stray "-warnings-as-errors", which is not a check name. The start-position setup in start_object()/start_array() is left alone, as #5706 is editing the callback parser's versions. Behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Correct the parser comments on recursion and skip_to_state_evaluation The class documentation called the parser a recursive descent parser, but sax_parse_internal() is a loop that keeps the open containers on an explicit stack. The comment at the end of an array and of an object said the flag is set to false while the code below it sets it to true. Describe what the code does instead. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Update the discard_number_values comments to the current number path The comments explaining the accept() shortcut in convert_number() and the member documentation still argued in terms of strtoull()/strtoll() and errno, which #5283 replaced with convert_integer(), and pointed at scan_number() instead of convert_number(). They also did not say that scan_number_bulk_contiguous() converts integers itself, so the shortcut is only reached for input without bulk access, with JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back. Rewrite both comments to describe the digit-count check in front of convert_integer(), keeping the 18-digit bound and the json_sax_acceptor argument. The stale <cstdlib> comment is left for after #5616, which edits that include block. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * List the UTF-8 validators instead of calling the DFA the only one The documentation of decode() called the Hoehrmann DFA the single source of truth for UTF-8 validation. It is used only by the serializer and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text strings). The lexer's scan_string() switch, validate_one_utf8() / valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer) and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges on their own. Replace the sentence with a list of the four validators, what each is used for, and a note that they must accept the same sequences. Sharing code between them was considered and dropped: it would save a few lines in a validator that is entangled with BON8 pushback, and #5677 is editing the BON8 byte path. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale doc comments and include lists in the input headers input_adapters.hpp included <memory> and <numeric> for the removed shared_ptr-based adapter design but used neither; it called (std::min) without including <algorithm>. json_sax.hpp used std::numeric_limits without including <limits>. Also corrected comments that no longer matched the code: input_stream_adapter does not skip the input's BOM (the lexer's skip_bom() does), the span_input_adapter comment named the no-longer-existing input_buffer_adapter type, lexer::get_string() does not reset the token, binary_reader's get_number() doc opened with /* instead of /*! (so Doxygen skipped it) and omitted BON8 from its endianness note, and the UBJSON-binary-types note did not mention that BJData 'B' arrays are read as binary. Left out: the lgtm suppression on lexer.hpp's scan_number() (in #5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's start_object/start_array docs (in draft #5267's hunk), per the verdict's conflict list. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the strict-EOF/release_lookahead/error block in parser::parse() json_sax_dom_callback_parser and json_sax_dom_parser branches of parser::parse() ran the same ~25 lines after sax_parse_internal(): the strict-mode EOF check (raising parse_error.101 through the SAX parser), release_lookahead() in non-strict mode, and mapping an errored SAX parser to a discarded result. The two copies had already drifted apart in formatting and in the second copy's "see above" comment. Add a private parse_dom(DomSax&, strict) member that runs this shared sequence once and returns whether the SAX parser did not error; both branches of parse() now only construct their DOM SAX parser, call parse_dom(), and (for the callback parser) map a discarded top-level value to null. sax_parse() is left untouched, since it only runs the EOF check and release_lookahead() when sax_parse_internal() succeeded, unlike parse(), which runs them unconditionally. Behavior-preserving: same operations in the same order for both SAX parser kinds. Verified with unit-class_parser (strict/non-strict, callback and non-callback), unit-deserialization and unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean make amalgamate / make check-amalgamation diff. Overlaps #5601, which touches the same lines. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 2 * Share the code point to UTF-8 encoding between the wide-string helpers and the lexer The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32 code point, in the UTF-16 helper for both a BMP code unit and a valid surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The copies had drifted: the UTF-32 helper masked the leading bits of each byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift alone, even though both give the same result for a code point that is already known to be in range. Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder that invokes a callable once per output byte, most significant byte first. Use it in the three valid-code-point branches (UTF-32 code points up to U+10FFFF, UTF-16 code units outside the surrogate range, and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where out forwards to add(). The UTF-16 helper's deliberate pass-through of malformed surrogate units and the UTF-32 helper's 0xFF sentinel for code points above U+10FFFF are untouched, since neither reaches the new helper. Behavior-preserving: same bytes in the same order for every valid code point, verified with unit-class_lexer, unit-class_parser, unit-deserialization, unit-wstring and the non-test-data parts of unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse microbenchmark that shows no change (about 73 ms either way, median of 3, 1M escape sequences). single_include/ regenerated with make amalgamate; make check-amalgamation leaves a clean tree. Overlaps #5704, which rewrites the wide_string_input_helper specializations touched here. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 6 --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
3926fcaac3 |
Deduplicate serializer dump code; fix stale includes, docs, and lint (#5729)
* Share scalar serialization between dump_internal and dump_value dump_value()'s cases for string, binary, boolean, number_integer, number_unsigned, number_float, discarded and null were a byte-for-byte copy of dump_internal()'s (added together in #5285 for the iterative fallback path). Any future change to scalar output had to be made in both places, or the recursive and depth-limited paths would silently start producing different bytes. Extract the shared cases into a private dump_scalar() and have both dump_internal() and dump_value() call it. Output is unchanged: dump(), dump(4), dump(-1,' ',true) and the replace/ignore error_handler_t variants are byte-identical over the json_test_data corpus before and after, and dump() throughput on a scalar-heavy document is unaffected. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Drop serializer.hpp's dependency on binary_writer.hpp The only use of binary_writer in serializer.hpp was binary_writer<BasicJsonType, char>::to_char_type() to write the U+FFFD replacement character's three bytes. With CharType=char this is an identity conversion, so the include of binary_writer.hpp (and transitively binary_reader.hpp) pulled in a large, unrelated header for a no-op call. Write the three bytes directly instead. serializer.hpp compiles standalone with -Wall -Wextra -Werror, with and without -funsigned-char, and unit-serialization's error_handler_t::replace cases (with and without ensure_ascii) still pass. Moving binary_writer's to_char_type/to_msgpack_length to its private section is left as an optional follow-up. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale #include lines in the output headers output_adapters.hpp included <algorithm> and <iterator> for std::copy and std::back_inserter, which have not been used there since #3569 (2022). serializer.hpp included <algorithm> for std::reverse (also unused), <cmath> for labs/isnan/signbit (only std::isfinite is used) and <utility> for std::move (nothing from <utility> is used there), while using std::next without including <iterator> at all, relying on getting it transitively through output_adapters.hpp's own stale <iterator>. Drop the unused includes, add <iterator> for std::next, and correct the remaining include comments. Both headers still compile standalone with -Wall -Wextra -Werror. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove JSON_HEDLEY_NON_NULL(2) from write_characters() overrides output_vector_adapter, output_stream_adapter and output_string_adapter declared their write_characters(const CharType*, std::size_t) override JSON_HEDLEY_NON_NULL(2), but binary_writer legitimately calls it with a null pointer and length 0 for an empty string or binary value; the type-erased call path only stayed silent under UBSan because the static callee at those call sites is the unattributed virtual base. A nonnull attribute on a definition lets GCC and Clang assume the parameter is non-null inside the function body even when the call is virtual, so this was latent undefined behavior, not just style. Drop the attribute from the three overrides and document the (nullptr, 0) contract on output_adapter_protocol::write_characters. unit-cbor, unit-msgpack, unit-bson and unit-bon8 (which all exercise empty binary/string payloads through the stream and vector/string adapters) pass under -fsanitize=address,undefined,nonnull-attribute. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Update stale serializer doc comments to match the current implementation dump_internal()'s doc block still described the pre-#5285/#5449 implementation: an escape_string() function that does not exist (the function is dump_escaped), integer conversion "implicitly via operator<<" (dump_integer actually uses a digit-pair lookup table), and floating-point conversion via "%g" (IEEE-754 types go through to_chars, others through snprintf). dump_value()'s comment said elements are pushed for dump_internal to walk, but it is dump_iteratively() that walks the stack. dump_escaped(), dump_integer() and dump_float() each said they write "to output stream @a o", which has not been true since the writer moved to write_buffer. Doc-only change; no behavior, API or ABI impact. Part of #5709 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Merge duplicate byte-to-hex helper and drop stale '| 0' promotions serializer::hex_bytes() and binary_writer::hex_byte() had identical bodies. Keep one, detail::hex_byte() in string_utils.hpp, and use it from both. Also drop the `| 0` at the two serializer call sites (hex_bytes(byte | 0) and hex_bytes(s.back() | 0)): #3088 ( |
||
|
|
18dd5663b0 |
Diff deeply nested values without recursing per nesting level (#5548)
* Diff deeply nested values without recursing per nesting level diff() descended into both values once per nesting level, and compared them with operator== on every level on the way, which recurses as well. Values nested deeply enough - 25,000 levels on an 8 MiB stack - exhausted the call stack and terminated the process, although parse() accepts them without complaint. On such a chain the per-level comparisons and path strings also made diff() quadratic in time and memory. Both the recursion and operator== only descend as far as the source is nested. So diff() first checks, recursing at most diff_depth_limit() (128) levels, whether the source is nested more deeply than that. If not - all but a vanishing minority of values - the recursive algorithm diffs it exactly as before, now as diff_recursively(). Otherwise diff_iteratively() walks the two values on an explicit stack, emitting the same operations in the same order. It does not compare arrays and objects with operator== up front (equal ones yield no operations anyway), keeps the path in one buffer instead of a new string per level, and hands every subtree that is not nested too deeply back to diff_recursively(), so equal parts are still skipped quickly. The check costs one pass over the source. On a 3,000-object document that is about 30% of diffing two equal values (which is just an operator== call), about 10% of diffing values that differ in a few places, and noise when arrays change length. Once operator== no longer recurses (#5390), the check can go. Tests check that the patch reproduces the target at every depth up to 300, for json and ordered_json, including reordered members. They also check the exact operation for a difference deep inside, and diff values nested 100,000 levels deep. Fixes #5393 for diff(). Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Make diff_frame a member struct that declares its special members GCC's -Weffc++ (an error in CI) asks a class with pointer members, a user constructor and a non-trivial destructor to declare its copy constructor and copy assignment; diff_frame's vector and basic_json members make its destructor non-trivial. Declare all five as defaulted, which also satisfies clang-tidy's special-member-functions check. Leave their exception specifications implicit: GCC 4.8 rejects an explicit one that differs from the implicit one, as it does for flatten_task in #5517. The converting constructor cannot throw, and is now declared noexcept for GCC's -Wnoexcept, which flags the emplace_back() under C++26 otherwise. The struct also moves from diff_iteratively() into the class, like dump_frame in the serializer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Use the shared recursion limit in diff() diff_depth_limit() is gone in favor of detail::recursion_depth_limit(). Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Diff fewer nesting depths so the test does not time out under Valgrind Checking every depth up to 300 made test-json_patch exceed the 1500 s ctest timeout in ci_test_valgrind. Check the depths up to 16, those around the recursion limit of 128, and 300 instead. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Mark the diff frame's value-initialized members for clang-tidy The braces are kept for GCC's -Weffc++, as in json_sax.hpp. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Bound diff()'s descent with a depth count instead of scanning the source Now that operator== no longer recurses (#5390), diff() can keep its per-level equality shortcut all the way down. It diffs recursively for the first detail::recursion_depth_limit() levels, as merge_patch() does, and hands anything deeper to diff_iteratively(). The nesting_exceeds() scan, which cost about 30% on equal documents, is gone, and diff() is on par with develop again. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Note that the diff frame reference is invalidated by pop_back() too Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Keep diff()'s recursive levels small and its result elided diff_recursively built every patch operation in place from initializer lists. Unoptimized builds give each of those temporaries its own stack slot, so every level of the bounded descent cost kilobytes of stack (about 6 KB with clang -O0), and the 128 recursive levels overflowed the 1 MB stack of MSVC Debug in the "deeply nested values" test. The operations and the key comparison of two objects are now built by separate functions, which diff_iteratively shares, and both diff functions append to one result instead of returning a patch per level that the caller copies. With clang -O0, diffing values nested 300 levels deep now peaks at about 190 KB of stack instead of 880 KB. Since diff() now owns the only returned value, clang's -Wnrvo no longer reports the returns of diff_recursively, which alternated between the local patch and diff_iteratively's result. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Copy the diff frame's members instead of holding a reference to it The loop in diff_iteratively held a reference to the top frame, which enter() invalidates when it pushes and the end of the loop invalidates when it pops. Nothing used it afterwards, but a later change could. As in the other iterative walks, the members the loop reads are now copied out as constants and the ones it advances are changed through stack.back(). The frame as a whole is not copied: it holds the common keys and the "add" operations of an object. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
a32f61eb98 | Fix update() and merge_patch() when the argument is *this or one of its members (#5678) | ||
|
|
1d675cdb46 |
Fix CI jobs that check less than they claim; move arm64 to GitHub (#5733)
* Fix the ci_cmake_flags wiring so every option is checked
The CMake 3.31.6 flag list referred to itself before it was defined,
so only JSON_BuildTests was checked with that version. The targets for
the CMake running the build ("_2") were created but never added to
ci_cmake_flags, and the three versions shared one build directory.
JSON_StrictNulHandling was not in the list at all.
Use the 3.5.0 list for 3.31.6, add JSON_StrictNulHandling, and create
one ci_cmake_flag_<flag> target per option for the running CMake with
its own build directory. Also use the function parameter in the
COMMENT, refresh the stale version comment, and let ci_clean remove
the downloaded cmake-<version> directories instead of the long-gone
cmake-3.5.0-Darwin64.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix ci_test_clang_libcxx_cxx* jobs silently building without warnings
CMake only seeds CMAKE_CXX_FLAGS from the CXXFLAGS environment variable
when the cache entry is unset, so the explicit -DCMAKE_CXX_FLAGS="-stdlib=libc++"
argument made it ignore CXXFLAGS="${CLANG_CXXFLAGS}" entirely. The six
ci_test_standards_clang (..., libcxx) jobs therefore compiled without
-Weverything/-Werror while their libstdc++ siblings did use them.
Pass -stdlib=libc++ through the same CXXFLAGS value instead of a separate
-D argument, and give the target its own build directory
(build_clang_libcxx_cxx${CXX_STANDARD}) so it no longer shares a CMake
cache with the libstdc++ variant. Suppress the resulting
-Wthread-safety-negative finding from libc++'s std::mutex annotations,
which fires on doctest's reporters in this translation unit only.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove the no-op AppVeyor with_win_header job
The with_win_header matrix entry patched Windows.h into
single_include/nlohmann/json.hpp before building, but JSON_MultipleHeaders
has defaulted to ON since #3532 (2022-06), so CMakeLists.txt points the
tests at include/ and the patched single header is never compiled. The
job has been a no-op VS2015 build since then.
Windows.h coverage already exists through tests/src/unit-windows_h.cpp
(#3631), which runs in every MSVC job. Delete the dead matrix entry and
its before_build steps, and cite unit-windows_h.cpp from the QA page.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove unused ci_oclint and ci_pvs_studio targets
No workflow invokes ci_oclint, ci_pvs_studio, or their tool discovery.
ci_oclint also had a side effect on every JSON_CI configure: it copied
the single header into src_single/all.cpp and added an add_executable()
for it without EXCLUDE_FROM_ALL, so a plain build compiled a 1.2 MB
translation unit that only that unused target consumed. ci_pvs_studio
duplicates the Makefile's pvs_studio target, which is kept.
Also drop the duplicate --check-level=exhaustive flag passed twice to
the same ci_cppcheck invocation.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop Dependabot from proposing astyle bumps
astyle is deliberately pinned at 3.4.13 because newer versions reformat
unrelated lines and this version defines the formatting that
check_amalgamation.yml enforces. Without an ignore rule, Dependabot
keeps opening PRs for every new astyle release (most recently #4580,
#4942, #5445, #5448), each of which fails the amalgamation check and
gets closed unmerged.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move Linux arm64 CI from dead Cirrus CI to ubuntu-24.04-arm
Cirrus CI stopped reporting check runs on develop sometime after
|
||
|
|
c261578431 |
Deduplicate binary reader/writer helpers and fix stale comments (#5730)
* Fix stale and missing comments in binary_writer The doc block of write_number() ended up above the byte_swap() helpers added in #5286, about 80 lines from the function. It was also a plain comment that Doxygen skips, said "write a number to output input", and left BON8 out of the big-endian formats. Move it back onto write_number() as a /*! block and fix the text. write_bson() documented "@pre j.type() == value_t::object", but it throws type_error.317 for every other type, and to_bson() relies on that. Document the exception instead. Explain why the CBOR binary subtype is always written with a 0xD8..0xDB head and never in the one-byte tag form: binary_reader with cbor_tag_handler_t::store only keeps those heads as a subtype, so switching to write_cbor_head() would break round trips for subtypes 0..23. Also fix the grammar of the to_char_type comment. Comments only; no change in behavior, API or ABI. Part of #5710 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Merge the duplicated UBJSON/BJData integer marker ladders write_number_with_ubjson_prefix() (unsigned and signed overloads) and ubjson_prefix() (number_integer and number_unsigned cases) each picked the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their own independent if/else ladder, and the values beyond 64 bits were handled by a second, tag-dispatched pair of ladders. An optimized container announces the marker of its first element via ubjson_prefix() and then writes every element through write_number_with_ubjson_prefix(), so the two had to be kept in lockstep by hand across four call sites. Replace all of that with one ubjson_integer_prefix() built on value_in_range_of<T>, and one write_ubjson_integer_payload() that writes the value (or, for 'H', the decimal digits) for a given marker. write_number_with_ubjson_prefix() and ubjson_prefix() keep their signatures and now just call these two helpers. Behavior, the public API and the ABI are unchanged. Verified with a new regression test covering scalars and $-optimized arrays/objects at every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing to_ubjson/to_bjdata output before and after over the json_test_data corpus (bit-identical). Part of #5710 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove dead get_char parameters in binary_reader The non-recursive rewrite of the binary readers (#5505, #5506, #5507) left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char parameters dead: parse_cbor_internal() has one caller and it always passes true, and parse_ubjson_internal() has one caller and it always uses the true default. Both parameters, and the @param docs describing the "reuse the last character" mode they used to select, no longer correspond to anything. Drop both parameters, initialise fetch/prefix unconditionally, and update the two call sites in sax_parse(). parse_cbor_value()'s and get_ubjson_string()'s own get_char parameters are unrelated and are left alone; both still have a false caller. Also delete a stray `@return whether a valid MessagePack value was passed to the SAX parser` doxygen block that sits directly above parse_msgpack_value()'s real doc comment, a leftover of the same rewrite. Behavior, the public API and the ABI are unchanged; these are private members of detail::binary_reader. Verified by compiling with -Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and unit-msgpack (offline, against the stubbed test_data.hpp). Part of #5711 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Share the IEEE half-precision decoder between CBOR and BJData binary_reader had two ~45-line copies of the IEEE 754 half-precision decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is normalised, the two blocks were identical except for the byte order used to assemble the 16-bit half (CBOR is big endian, BJData is little endian). Any future change to half-float decoding had to be made and kept in sync in both places. Add one get_half_float(format, little_endian) helper that does the two get()/unexpect_eof() reads, assembles the half in the requested byte order, decodes it per RFC 8949 Appendix D, and calls sax->number_float. Both cases now just call it with their byte order; the BJData case keeps its bjdata-only guard. Behavior, the public API and the ABI are unchanged. Verified with a scratch probe comparing the old and new decoders bit-for-bit (NaN by isnan()) over all 65536 wire byte pairs, in both formats, and by running unit-cbor and unit-bjdata (offline, against the stubbed test_data.hpp). Part of #5711 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the MessagePack unsigned-integer writer ladder The number_integer (non-negative branch) and number_unsigned cases in write_msgpack() each held their own copy of the fixint/uint8/16/32/64 ladder, kept in lockstep only by a comment ("we used the code from the value_t::number_unsigned case here"). Both copies mixed union members: the signed copy compared number_unsigned but wrote number_integer, and vice versa. Extract write_msgpack_unsigned(std::uint64_t), mirroring how write_cbor_head() already avoids the same duplication for CBOR, and call it from both cases. Each case now reads only its own active union member. Output bytes are unchanged for the default 64-bit number types. #5710 item 3 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Unify float marker selection and fix the long double compile error Four formats picked between a float32 and float64 marker through four different helper styles: dummy-argument overloads for CBOR and MessagePack, an std::is_same template for BON8, and a runtime if-chain on input_format_t for write_compact_float(). With number_float_t set to long double, to_cbor, to_msgpack and to_ubjson failed inside the library with "call to 'get_cbor_float_prefix' is ambiguous", while to_bson kept working because write_bson_double() takes a plain double. Change write_compact_float() to take the two marker bytes directly (each of its three callers already knows them at compile time) instead of an input_format_t it only forwarded, and delete the now-unused get_cbor_float_prefix(), get_msgpack_float_prefix(), get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn the two get_ubjson_float_prefix() overloads into one template. Both write_compact_float() and get_ubjson_float_prefix() now report an unsupported number_float_t with a static_assert naming the requirement, rather than an ambiguous-overload error; the assert lives in the function body, not the class scope, so to_bson with long double is unaffected. Verified with a probe basic_json<..., long double>: to_bson still compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail to compile with the new static_assert message. This changes the text of an existing compile error for users with an unsupported number_float_t (documented as a public-API-visible change in #5710). #5710 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the BJData ndarray writer's dtype dispatch and drop <map> write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on every call just to translate the _ArrayType_ name to a dtype marker (the only reason binary_writer.hpp included <map>), then mapped dtype to C++ type twice more: once as a switch for the range-check pass and once as a separate if/else chain for the write pass, with nothing checking that the two agreed. The caller also ran three at() lookups, and the callee called value.at(key) about ten more times for the same three members. Replace the map with bjdata_ndarray_type_marker(), a plain string comparison chain (a C++11 constexpr function cannot contain a switch, so this mirrors binary_reader's own static table style). Replace the switch/if-chain pair with one write_bjdata_ndarray_elements() that switches on dtype once and calls a per-type helper - write_bjdata_ndarray_element<T>() for the eight integer dtypes and write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag selecting the range check or the actual write, so the two passes can no longer disagree on the type. _ArrayType_, _ArraySize_ and _ArrayData_ are now looked up once into references, and the four header marker bytes ('[', '$', '#') are written through to_char_type() like the rest of the UBJSON/BJData writer. The 'd' (single-precision) rule is left exactly as before, since #5707 is expected to change it separately. Verified byte-for-byte identical output before/after for every dtype (including the Draft 2/Draft 3 'byte' fallback and the use_count/ use_type combinations) via a standalone probe, plus round-tripping through from_bjdata(). Overlaps #5707, which is expected to touch the 'd' dtype case, and #5518, which is expected to move the write_bjdata_ndarray() call site. #5710 item 4 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Assert that write_bson_document() consumes every calc_bson_sizes() entry calc_bson_sizes() and write_bson_document() are a hand-synchronized pair of passes over the same object/array tree, introduced by #5553: the size pass appends to nested_sizes in visiting order, and the write pass consumes the table by position with nested_sizes[next_size++]. Nothing checked that the write pass consumed the whole table. If a future change touched only one of the two passes - for example to skip or reject an entry - every later size prefix in the document would be silently wrong. Add JSON_ASSERT(next_size == nested_sizes.size()) where write_bson_document() returns, so such a future drift between the two passes is caught immediately (JSON_ASSERT expands to nothing in release builds using assert(), and the fuzzers/tests already build with it enabled). The two passes agree today, so this changes nothing observable; it only guards against the risk described in #5710 item 5. Extracting a shared stepper for the two passes (the second half of the proposed change) is left for a follow-up: it only saves ~30 lines and the issue asks for it only if the result reads clearly, which needs more room to get right than a mechanical cleanup pass allows. #5710 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Make the BJData lookup tables static functions instead of members binary_reader held bjd_optimized_type_markers and bjd_types_map as non-static const members (12 string_t objects for the type-name table), built and destroyed on every from_cbor/from_msgpack/from_bson/ from_ubjson/from_bon8/from_bjdata call even though only from_bjdata ever reads them. They also needed the #define/decltype/#undef workaround from #3637 and two NOLINTNEXTLINE suppressions, and binary_writer already carries the same two lists in another form (is_bjdata_excluded_type_marker() and a local std::map in write_bjdata_ndarray(), the latter removed by the item-4 commit), so the excluded-marker lists could drift apart. Replace bjd_optimized_type_markers with static constexpr is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain as binary_writer's is_bjdata_excluded_type_marker(). Replace bjd_types_map with a non-constexpr static bjd_type_name(char_int_type) switch returning nullptr for an unknown marker (a C++11 constexpr function cannot contain a switch). Delete both JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the NOLINTNEXTLINE suppressions, detail::make_array() (no longer used anywhere), and the now-unused <algorithm> and <array> includes. Update the two call sites (the ND-array excluded-type check and the _ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's "LUT arrays are sorted" section, which only checked the two tables' internal ordering, with a check of all 12 type names and all 8 excluded markers against both new functions. #5711 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read CBOR's 1/2/4/8-byte argument through one helper parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte big-endian unsigned integer" ladder four times over: - twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore branch and once, nearly identically, in the ::store branch (~90 lines to read one integer); - twice more for container lengths, once for array heads 0x98-0x9B and once for map heads 0xB8-0xBB, where the 1/2-byte forms called enter_array()/enter_object() directly and the 4/8-byte forms additionally went through get_cbor_container_size(). Add get_cbor_argument(std::uint64_t&), reading the width selected by current & 0x1F via the same get_number() calls as before (so EOF is reported exactly as before), and route all four sites through it: - 0xD8-0xDB now read the argument once per branch instead of switching on `current` a second time; behavior split cleanly from embedded tags 0xC0-0xD7 (tag value in the head, no argument to read), which is now its own case block that no longer has to fall into the ::store switch's "default" case to reach the same tag_pending = true; return true; outcome. - 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always going through get_cbor_container_size() (harmless for 1/2-byte lengths, which already always fit). Verified byte-for-byte identical behavior before/after with a standalone probe covering embedded and multi-byte tags under all three tag_handler_t settings, a tag over a byte string (subtype path), truncated tag/length arguments of every width, and array/map lengths of every width, including the out_of_range.408 "excessive size" case: same exceptions, same messages, same chars_read, same successful results. Left the string/byte-string length ladders in get_cbor_string()/ get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is expected to touch them separately. Overlaps #5601 (adds a branch right above the embedded-tag case) and #5607 (touches the integer cases 0x18-0x1B, which share this ladder's shape in separate hunks). #5711 item 2 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Add leave_container() to match enter_container() Every container is opened through enter_container(), whose docs promise that a check placed there runs before every start event. The close side had no equivalent: the same "container_stack.pop_back(); dispatch to end_object() or end_array()" sequence was written out separately in BSON, CBOR, MessagePack, UBJSON/BJData and BON8, each copying the pattern of keeping an is_object flag around the pop_back() that would otherwise invalidate a reference to it. A check needed on close would have had to be added in five places, and a sixth copy could go unnoticed. Add leave_container() next to enter_container(), doing the same pop-then-dispatch, and replace the five sites with it. Each site keeps its own surrounding logic (BSON's check_bson_document_size() call before popping, MessagePack's is_object copy used again below, UBJSON/BJData's remaining-container handling after popping, BON8's top used again below); only the repeated pop/dispatch line pair is now shared. Verified all six binary-format unit suites and unit-regression2's deep-nesting tests (dependent count/reuse count and the bjdata ndarray depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan. Overlaps #5601, which is expected to add a sixth close site in its own skip loop; that site can route through leave_container() too once it lands. #5711 item 4 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop passing the input format to sax_parse() when the reader already has it binary_reader's constructor stores the format in the input_format member, and sax_parse(format, sax_, strict, tag_handler) took the same value again purely to dispatch on it. Every in-tree caller passed the same value both times (all 16 from_cbor/from_msgpack/from_ubjson/ from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three public basic_json::sax_parse() overloads), so nothing was broken today, but a caller of the detail class directly (only reachable via JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could pass a mismatched pair - say bjdata to the constructor and ubjson to sax_parse - and dispatch on one format while applying the other format's rules; the default-constructed input_format_t::json reader would additionally hit JSON_ASSERT(false) in exception_message() on its first error. Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the existing overload with the stored input_format, and switch every caller to it: the 16 from_*() sites (keeping their `// cppcheck-suppress[accessMoved]` comments) and the three basic_json::sax_parse() overloads, all of which already had the format available from their own `format` parameter. The four-argument overload is kept for anyone still calling it, now with JSON_ASSERT(format == input_format) so a mismatch fails immediately in a debug build (assert-enabled binaries, including the fuzzers and test suite) instead of misbehaving; verified with a probe that constructs a reader for one format and calls the explicit overload with another, which aborts on that assertion as expected. Removing or asserting against the constructor's input_format_t::json default, which would affect direct detail users, is left as a separate decision per #5711 item 5. Overlaps #5601, which is expected to add an AllowRecovery template parameter to sax_parse() and touch these same call sites in json.hpp. #5711 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate UBJSON/BJData signed-count handling, drop dead ndarray checks get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently sized signed integer and then repeated the same "reject negative with error 113" check; only 'L' additionally checked value_in_range_of for the out_of_range.408 case. Any change to that error path had to be made four times. Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read, the negative check and the range check once, and route all four markers through it. The range check is a no-op for 'i'/'I'/'l' (their values always fit std::size_t) and only live for 'L' on a 32-bit std::size_t target, matching today's behavior exactly. In the ndarray dimension-product loop, the preceding loop already returns early on any zero dimension and result starts at 1, so `i > 0` in the pre-multiplication overflow check was always true, and `result == 0` in the post-multiplication check could not be reached either: two positive factors whose product does not overflow (as the pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&` and narrow the post-check to `result == npos`, the one case the pre-check cannot rule out (an exact, non-overflowing match with the sentinel reserved for unknown-size containers), with a comment explaining why. Verified byte-for-byte identical behavior before/after with a standalone probe covering negative counts for every marker, a matching positive count, and ndarray inputs, plus the full unit-ubjson and unit-bjdata suites (same assertion counts as before this change). Overlaps #5601 (rewrites the four parse_error calls and the overflow checks touched here) and #5607/#5707 (touch neighboring lines in the same functions). #5711 item 6 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Drop redundant format parameter and dummy float argument (review) binary_reader::sax_parse(format, ...) only ever had to equal the format given to the constructor, which it asserted. With every caller already on the format-less overload, remove the four-argument overload and dispatch on the stored input_format directly. binary_reader is a detail class, so this is not a public API change. get_ubjson_float_prefix() took a value only to deduce its type; make the type an explicit template argument instead. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
35802e78d6 |
Remove dead Makefile targets; document macro_builder; tidy serve_header (#5735)
* Remove dead doctest help entry and pretty_format target from Makefile The top-level Makefile still carried three leftovers: - The help text listed a "doctest" target that was removed in #4560, so "make doctest" fails with "No rule to make target". The example check now runs as "make check_output -C docs". - "pretty_format" ran clang-format on all sources, but .clang-format was deleted in #4573, so the target reformatted everything in the default LLVM style, against the Artistic Style formatting that "make pretty" applies and CI enforces. - "clean" removed benchmarks/files/numbers/*.json, a directory that no longer exists since the benchmarks moved to tests/benchmarks (#3462). Only maintainer tooling changes; the library is not affected. Part of #5717 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Document tools/macro_builder and tidy up serve_header.py tools/macro_builder generates the NLOHMANN_JSON_EXPAND, NLOHMANN_JSON_GET_MACRO and NLOHMANN_JSON_PASTE* macros in macro_scope.hpp, but nothing referred to it. Add a README that explains what it generates, how to run it and where the output goes, and which dependent tables (NLOHMANN_JSON_DOUBLE_PASTE, NLOHMANN_JSON_TYPE_BODY) are maintained by hand. Point to it from a comment above NLOHMANN_JSON_EXPAND. The generator itself is unchanged; following the README reproduces the header byte for byte. In serve_header.py, drop the LGTM suppression (LGTM.com shut down in 2022), replace the """.""" placeholder docstrings with real ones, and import socket and ssl at module level. DualStackServer.server_bind uses socket, which was only imported under __main__; when the module was imported instead, the NameError was swallowed and IPV6_V6ONLY was not cleared. The header change is a comment only; behavior, API and ABI are unchanged. Part of #5717 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Hash every release_files artifact, not a hardcoded subset The `release` target signed and copied json_fwd.hpp into release_files alongside json.hpp, but the shasum line that writes hashes.txt only listed json.hpp, include.zip and json.tar.xz. Users could not verify the published json_fwd.hpp against hashes.txt. Hash every file in release_files except the .asc signatures instead of naming files by hand, so a newly shipped header (such as the json_literals.hpp that #5610 adds to this target) cannot be missed again. Only affects the generated hashes.txt release artifact; the library itself is unaffected. Overlaps #5610, which touches the same lines to add json_literals.hpp to the release target. #5717 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the broken fuzz_testing* Makefile targets fuzz_testing and fuzz_testing_{bon8,bson,cbor,msgpack,ubjson} seeded fuzz-testing/testcases from tests/data, which was removed in |
||
|
|
73d116547a |
Reject ill-formed UTF-8 in CBOR/UBJSON/BJData/BSON writers, accept it in MessagePack reader
to_cbor(), to_ubjson(), to_bjdata(), and to_bson() wrote a string or object
key with ill-formed UTF-8 byte for byte, but from_cbor()/from_ubjson()/
from_bjdata()/from_bson() reject such strings with parse_error.113 since
|
||
|
|
791cd88dfc |
Fix raw-TeX formulas and other documentation infrastructure debt (#5736)
* Fix math formulas rendering as raw TeX in the published docs The privacy plugin self-hosts MathJax 2.7.0 but drops its ?config=TeX-MML-AM_CHTML query string, so the rehosted script loads no input jax and the 10 formulas across 6 pages render as raw TeX to readers. Remove pymdownx.arithmatex and the MathJax extra_javascript entry, and rewrite the formulas in plain HTML (<sup>, <i>) instead. This also drops a nine-year-old third-party script from every page. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix publish_documentation triggers and persist unneeded git credentials publish_documentation.yml only triggered on docs/mkdocs/** pushes, but the site also embeds .github/CODE_OF_CONDUCT.md, CONTRIBUTING.md, SECURITY.md, cmake/{clang,gcc}_flags.cmake, .clang-tidy, tools/astyle/.astylerc and tests/fmt_formatter/project/main.cpp via pymdownx.snippets, so changes to those files never republished the site. Extend the path filter to cover them, and switch runs-on from the long-pinned ubuntu-22.04 to ubuntu-latest to match ci_test_documentation. Also add persist-credentials: false to the checkouts in ci_icpx and ci_nvhpc (ubuntu.yml) and msvc-vs2026/msvc-arm64 (windows.yml), none of which pushes with git, matching every other checkout in these workflows. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix example build: broken debug echo, deprecations hidden for everything docs/Makefile's debug echo used a space instead of a comma in $(call cxx_standard ...), so it always printed an empty standard. Every example was also compiled with -Wno-deprecated-declarations, which would silently hide an accidental deprecated-API call in any of them. Factor the duplicated compile flags into EXAMPLE_CPPFLAGS/ EXAMPLE_WARNFLAGS, build with -Werror=deprecated-declarations by default, and only allow the three examples that intentionally document deprecated API (the DEPRECATED_EXAMPLES list) to suppress it. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix check_structure.py NOLINT parsing and an off-by-one report line `line.strip("<!-- NOLINT")` followed by `.strip(" -->")` strips any of the characters in those sets from both ends, not a literal prefix; it only happened to work for "Examples". A NOLINT'd section name starting with N, O, L, I or T (e.g. "Notes", "Template parameters", "Iterator invalidation", "Literals") was silently mangled, so the suppression did not apply and the checker could report a spurious missing/misordered section. Parse the comment with a regex instead. The same fragile strip() pattern was used for heading text; replace it with a plain prefix slice. Also fix the admonition_title report, which used the 0-based line index while every other report in the file uses lineno+1. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix docset README's non-existent make target and stale fallback URL The README told readers to run `make nlohmann_json.docset`, but the Makefile's targets are `all`, `JSON_for_Modern_C++.docset` and `install_docset_zeal`; the documented command has failed with "No rule to make target" since #2967 (2021). Point the README at the real target and folder name. Info.plist's DashDocSetFallbackURL also still pointed at the old nlohmann.github.io/json/ URL instead of the canonical https://json.nlohmann.me/ from mkdocs.yml's site_url. Leave list_missing_pages/list_removed_paths alone: they may become redundant once #5638's check_docset() lands, which is a follow-up. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove two leftover Doxygen-era .link files from the examples directory parse__iterator_pair.link and parse__pointers.link each held only a Wandbox "online" permalink from the old Doxygen docs. #3071 deleted every other .link file in 2021; these two came in through a parallel PR (#3100) and were never referenced by any page, script or config. Part of #5718 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix MacPorts CMake example to include its own snippet files The MacPorts "Example: CMake" block included integration/homebrew/example.cpp and integration/homebrew/CMakeLists.txt instead of the MacPorts files right next to it, a copy-paste slip from the Homebrew section. Nothing referenced integration/macports/CMakeLists.txt as a result. The page rendered correctly only because the homebrew, macports and vcpkg/CMakeLists.txt snippets are byte-identical, so a future edit to the MacPorts files would not have shown up on the page. Part of #5718 item 6 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Do not persist git credentials in publish_documentation's checkout The checkout step in publish_documentation.yml left the default persist-credentials: true, so GITHUB_TOKEN stayed writable in .git/config for the rest of the job (zizmor's artipacked finding). The Deploy documentation step authenticates through its own github_token input to peaceiris/actions-gh-pages and does not push with the checked-out credentials, so persist-credentials: false is safe here, matching every other checkout in the workflow set. Overlaps #5638, which edits this same checkout step (adds fetch-depth: 0); expect a rebase conflict there. Part of #5718 item 2 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Replace list_missing_pages/list_removed_paths with a comm(1)-based diff The docset Makefile's list_missing_pages ran one sqlite3 query per mkdocs page, and list_removed_paths nested a loop over all mkdocs pages inside a loop over all docset index paths (O(n*m) shell iteration). Issue #5718 item 5 suggested removing or reducing these targets once #5638's check_docset() lands, but that PR is still open and covers only API pages and macros, not the full page set these targets check. Replace the loops with two sorted path lists (DOCSET_PAGE_PATHS from mkdocs' markdown sources, DOCSET_INDEX_PATHS from the built docset index) compared with a single comm(1) call each, verified to produce output identical to the old loops against the current docSet.dsidx. The sed expression used '#' as its delimiter, which GNU Make reads as a comment character even inside a variable assignment, truncating the line and orphaning the closing paren of $(shell ...) ("unterminated call to function 'shell': missing ')'"). Use '@' as the delimiter instead. Part of #5718 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
7d7055ec50 |
Fix stack overflow converting deep values between specializations (#5723)
* Fix stack overflow converting deep values between specializations Constructing a basic_json from another specialization (json to ordered_json or back, also via get<ordered_json>()) converted every container with its range constructor, which calls the converting constructor for each element. The call stack therefore grew with every nesting level, and a value nested some 30,000 levels deep overflowed it. The conversion now bounds its descent the way the copy constructor does since #5387: the first 128 levels are converted exactly as before, and below that convert_iteratively() finishes the value with an explicit stack. It builds each container bottom-up from its converted elements with the container's range constructor, so member order and keys that become equal are handled as before, and it gives a value its type only once its container exists, so an exception leaves nothing behind that cannot be destroyed. Parents (JSON_DIAGNOSTICS) and positions (JSON_DIAGNOSTIC_POSITIONS) are set for every value. Converting a null value no longer resets its positions: the constructor assigned null to a value that already was null, which swapped in the positions of the temporary. Fixes #5650. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Explain why converting null keeps positions and why next is a reference Review feedback on #5723 (gregmarr): clarify in comments that the converting constructor has already copied the positions of val, which the null case keeps like every other case, and that next must be a reference into pending so that ++next advances the stored iterator. Comments only; no code change. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Refer to recursion_depth_limit() in the convert_structured() docs The comment still named nesting_depth_limit, which #5637 removed on develop in favor of detail::recursion_depth_limit(). Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Advance the pending iterator through pending.back() and shorten the null comment Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
3b6ae43c53 |
Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions (#5598)
* Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions basic_json can be constructed from std::tuple<json&>, which it turns into a one-element array. Because of this, std::tuple picks its converting constructor that converts the whole source tuple instead of the element-wise one. As a result, std::tuple<const json&> built from std::forward_as_tuple(j) binds to a temporary (a compile error with libc++, a dangling reference with other standard libraries), and std::tuple<json> built the same way holds [j] instead of a copy of j. The new opt-in macro JSON_DISABLE_TUPLE_REFERENCE_CONVERSION (CMake option JSON_DisableTupleReferenceConversion) removes the conversion from a one-element tuple holding a reference to the same basic_json type, so std::tuple converts element-wise. It is off by default, so existing behavior is unchanged. It does not change any function body and therefore is not part of the ABI tag. Fixes #2226 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert one-element tuples to arrays on every compiler to_json for std::tuple assigns a braced list, j = { std::get<Idx>(t)... }. With a single element that is itself a basic_json, Apple clang 15 and 16 treat j = {x} as a copy of x, so std::tuple<json>{true} became true instead of [true]. The macOS jobs (Xcode 15.1, 16.1) failed the new checks in unit-disable-tuple-reference-conversion and unit-regression2. The one-element overload that already handles JSON_BRACE_INIT_COPY_SEMANTICS builds the array (or object, for a [string, value] element) explicitly, the same way the initializer-list constructor does. Use it unconditionally. The output is unchanged on compilers that already wrapped the element. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Skip json reference tuple tests on clang < 4 and GCC < 5 ci_test_compilers_gcc_old (4.8) and ci_test_compilers_clang (3.4) could not compile the new tuple tests. Creating a std::tuple of basic_json references, e.g. std::forward_as_tuple(j), makes these compilers instantiate basic_json's conversion operator for libstdc++'s internal tuple bases, which fails hard. This happens with and without JSON_DISABLE_TUPLE_REFERENCE_CONVERSION, so it is a limitation of these compilers, not of the new option. Tested with the CI images: clang 3.4 to 3.9 and GCC 4.8 and 4.9 fail, clang 4, 5, and 6 and GCC 5 and 6 compile all cases. Skip only the checks that create such tuples; the is_constructible checks still run. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6a073dbae4 |
Give operator>> a strong exception-safety guarantee (#5695)
operator>> parsed directly into its basic_json& target, so a parse error left the target holding whatever was parsed before the error instead of its previous value. With JSON_DIAGNOSTICS=1, that partial value also violated the class invariant, because the parent pointers of an array or object's elements are only set when the container is closed, which a failed parse never reaches; copying such a value then aborted in assert_invariant(). Fix it the way basic_json::parse() already handles this: parse into a temporary and move it into the target only once parsing succeeds, so the target is left unchanged if an exception is thrown. Fixes #5652. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
fdcc569eee |
Use only documented StringType members in json_pointer (#5692)
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.
Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.
Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).
Fixes #5666.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
6fc0d501f3 |
Move the user-defined string literals to <nlohmann/json_literals.hpp> and add JSON_NO_AUTOMATIC_UDLS (#5610)
* Add JSON_NO_UDLS to leave out the user-defined string literals The bodies of operator""_json and operator""_json_pointer call the parser, so every translation unit including the library instantiates it, even if it never parses anything. Defining JSON_NO_UDLS leaves the literals out entirely, which saves 15-35% compile time for such translation units (#5294). Nothing changes if the macro is not defined. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Mention JSON_NO_UDLS in the list of exported module symbols Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the user-defined string literals to <nlohmann/json_literals.hpp> Following the review in #5294, the literals now live in their own header instead of being removed entirely: <nlohmann/json.hpp> includes it at the end unless JSON_NO_AUTOMATIC_UDLS (renamed from JSON_NO_UDLS) is defined, so a project can opt out globally and include the header only where the literals are used. The header only uses public and standard macros, because the library's internal macros are undefined at the end of json.hpp and the amalgamation inlines macro_scope.hpp only once. For the same reason, the library no longer defines and undefines JSON_USE_GLOBAL_UDLS, so a user's definition is still visible to the header. The single-header copy is identical to the multi-header one, as it only includes <nlohmann/json.hpp>. The module always exports the literals. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: include cycle, GCC 4.8 literal operator spacing, and global UDLs off in the JSON_NO_AUTOMATIC_UDLS test - Suppress clang-tidy misc-header-include-cycle on the intentional mutual include of json.hpp and json_literals.hpp. - Use operator"" _json with a space for GCC 4.8 in the test's detection aliases, as the header does. - Only test the global literal operators when JSON_USE_GLOBAL_UDLS is on. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Declare the literal operators through a local macro The GCC 4.8 spacing condition was repeated for both operator definitions and the global using-declarations. NLOHMANN_JSON_LITERAL_OPERATOR(suffix) now selects operator""##suffix or operator"" suffix in one place and is undefined at the end of json_literals.hpp. Suggested by gregmarr in review. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
67435c9c7e |
Give a deep copy its type only after its container exists (#5721)
When copying a value nested deeper than 128 levels, and an allocation fails while an inner array or object is being copied, the partially built copy ended up with an element typed array/object but holding a null pointer. That element was already a fully constructed member of its parent's container, so destroying the parent during stack unwinding dereferenced the null pointer (release builds) or failed assert_invariant() (debug builds), instead of letting std::bad_alloc reach the caller. copy_iteratively() set a pending worklist element's type right after popping it, before the next loop iteration created its container in copy_array_level()/copy_object_level(). Move that type assignment into those two functions, right after the container is successfully created, and drop the premature one in copy_iteratively(), so a half-built element stays a null value - as copy_shallow()'s comment already promised - until it can safely hold one. Add a regression test to tests/src/unit-allocator.cpp that copies a value nested 130 levels deep (both arrays and objects, with a std::map- and an ordered_map-backed object_t) and fails every allocation of the copy in turn: each attempt must throw std::bad_alloc without crashing, and the source must stay unchanged. Fixes #5640. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
2ea6d8c127 |
Require the found key to equal the looked-up key when comparing objects (#5720)
For an object type whose comparator treats unequal keys as equivalent (for example a std::map with a case-insensitive comparator), compare_iteratively() looked up a mismatched left key in the right object with find(), which uses the object's own comparator, and accepted whatever entry it found without checking that the keys are actually equal. A case-insensitive comparator then found "KEY" for "key", so two objects nested past the recursion bound (or at every depth with JSON_NO_THREAD_LOCAL) could compare equal even though the object type's own operator== - and basic_json itself, below the bound - consider them different. Accept the found entry only if its key equals (not just compares equivalent to) the looked-up key. Fixes #5655. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
66877675b1 |
Check the iterator range for binary values in basic_json(first, last) (#5719)
basic_json(first, last) treated value_t::binary like the structured types (array, object) in the range check, so it always copied the whole binary value regardless of the iterators, even for an empty range such as (b.end(), b.end()). The other primitive types (number, boolean, string) already reject such a range with invalid_iterator.204, and erase(first, last) already does the same for binary values, so this made the constructor inconsistent with both. Move case value_t::binary into the group of checked primitive types. Also update the two matching passages in basic_json.md that describe overload 7, and add a version-history note. Fixes #5670. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
7fd6895788 |
Hide a discarded container's content from the parser callback (#5706)
When a parser callback rejects an object's or array's start event,
json_sax_dom_callback_parser kept calling it for everything inside
that container anyway: nested keys, values, and the start/end events
of containers below it. This contradicts parser_callback_t's own
documentation, which promises that discarding a container at its
start event also hides its content from the callback.
The same code path also kept a full copy of every key inside such a
discarded container in key_stack until the whole parse finished,
because the early return for values that are not stored skipped the
matching pop. Filtering out a large subtree is the main reason to use
a callback, so this made peak memory during the parse scale with the
size of the very subtree the callback was trying to skip.
Fix start_object(), start_array(), and key() so that a container
whose own start event was discarded, or that is nested inside one, is
never handed to the callback, and no longer pushes onto the key
stacks. A container whose start event was accepted but whose key was
rejected still gets its content reported, as documented ("the
callback is still called for the associated value, but its return
value has no further effect"); only its own bookkeeping is skipped
since it will not be stored.
Fixes #5643.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
6ae17630a4 |
Reject integral keys for contains(), find(), and count() at compile time (#5705)
j.contains(0), j.find(0), and j.count(0) used to compile: the literal 0 is a null pointer constant, so it converts to a null const char*, and the overloads taking const typename object_t::key_type& accepted it by constructing a std::string from that null pointer, which is undefined behavior (a crash with both libc++ and libstdc++). value(0, default_value) had the same problem in C++11, where the object comparator is not transparent. Add deleted overloads for integral arguments to contains(), find() (const and non-const), count(), and value() so that these calls are compile errors in every supported language mode instead of crashing. Calls with string, string_view, json_pointer, and size-typed element access (at(), operator[](), erase()) are unaffected. Fixes #5657. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
5bd766aa50 |
Move to_bson's binary subtype check into calc_bson_sizes (#5703)
to_bson() rejected a binary value's subtype above 255 (out_of_range.415) in write_bson_binary(), which only has the binary_t, not the basic_json value that holds it, so the exception was created with no JSON_DIAGNOSTICS context even though the equivalent to_msgpack() check names the value's path. The check also ran after the document size, all preceding elements, and this element's header and length had already reached the output adapter, so a caller-provided std::vector or std::string ended up holding a truncated document. calc_bson_sizes() already walks every value before anything is written, to size embedded documents and arrays and to reject invalid keys (out_of_range.409) up front. The subtype check now runs there instead, in calc_bson_binary_size(), which is given the basic_json value so the exception can use it as context. The now-redundant check in write_bson_binary() is removed, since calc_bson_sizes() always throws first if any binary value in the document has an oversized subtype. Fixes #5675. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
4bb1b14b06 |
Trim the compiler-appended NUL from wide/UTF string literals too (#5702)
With JSON_STRICT_NUL_HANDLING defined to 1, parsing a wide, UTF-16, UTF-32, or (C++20) UTF-8 string literal (e.g. json::parse(L"[1]")) failed with parse_error.101 at the terminating NUL of the literal, and accept() returned false. The array overload of input_adapter() only dropped the compiler-added trailing '\0' for arrays of char, so for wchar_t, char16_t, char32_t, and char8_t arrays that terminator was passed to the parser as data, which the macro then rejected. Broaden the trimming to every character type that a string literal can use (char, wchar_t, char16_t, char32_t, and, since C++20, char8_t). Arrays of any other element type (unsigned char, std::uint8_t, ...), as used for CBOR/MessagePack, are unaffected: a trailing zero byte there is still read as genuine data. Fixes #5658. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
bfe0f32d71 |
Fix std::terminate and null pointer access in input_stream_adapter (#5699)
Parsing from a std::istream crashed in two unusual but valid stream states, both in input_stream_adapter: - With eofbit in the stream's exceptions() mask, get_character() sets eofbit via is->clear(), which throws std::ios_base::failure. While that exception unwinds, ~input_stream_adapter() called clear() again to reset eofbit, which is still set and still in the exception mask, so it throws a second time out of the (implicitly noexcept) destructor and std::terminate() is called. The destructor now only calls clear() if a bit other than eofbit remains set, so the first exception can propagate normally. - For an std::istream without a stream buffer (rdbuf() == nullptr, e.g. std::istream(nullptr)), the constructor stored the null pointer without checking it, and get_character() dereferenced it. input_adapter(std::istream&) now throws parse_error.101 for such a stream, the same as it already does for a null FILE* or char*. Added regression tests to unit-deserialization.cpp and, for the JSON_PRECISE_STREAM_POSITION variant of get_character(), to unit-precise-stream-position.cpp; both crashed before this fix. Documented the two exceptions in parse.md and operator_gtgt.md. Fixes #5646. Signed-off-by: Niels Lohmann <mail@nlohmann.me> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
44a88d85be |
Fix NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's from_json message (#5698)
from_json built its out_of_range.410 message with "..." + j.dump(). If the unmatched value is (or contains) a string with invalid UTF-8, that dump() itself throws type_error.316, so the caller got type_error.316 instead of the documented out_of_range.410; such strings can reach get<Enum>() unvalidated, e.g. from from_cbor()/from_msgpack(). With a custom string_t, j.dump() returns that type, and "const char*" + string_t does not compile unless the type happens to provide operator+, so the macro failed to compile for such types. Build the message with detail::concat(), which appends any type exposing data()/size() and always yields a std::string, and dump with error_handler_t::replace so building the message itself cannot throw. Added regression tests: an invalid-UTF-8 case in the existing strict-enum test in unit-conversions.cpp, and a strict-enum use with alt_string (the custom string_t from unit-alt-string.cpp) to cover the compile failure. Fixes #5667. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
5b11a0282c |
Exchange the CustomBaseClass subobject in basic_json::swap() (#5697)
basic_json::swap() (and the friend swap() and the pre-C++20 std::swap overload that forward to it) only exchanged m_data.m_type/m_data.m_value, leaving each value's json_base_class_t subobject in place. This is inconsistent with the copy and move constructors and copy assignment, which all carry the base class along with the value, so after a.swap(b) any metadata stored in a CustomBaseClass ended up attached to the wrong value. Algorithms that mix swap() with moves, such as std::sort, scrambled the metadata across the whole container. Fix the member swap() to also exchange the json_base_class_t subobject and extend the noexcept specifications of swap() and the friend swap() accordingly. Fixes #5653. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
bfea6f36d3 |
Fix to_msgpack() reading the inactive number union member (#5694)
* Fix to_msgpack() reading the inactive number union member basic_json stores number_integer and number_unsigned in a union, and number_unsigned_t only has to be at least as wide as number_integer_t (with the default types, both are 64-bit and have the same representation). When number_integer_t is narrower, write_msgpack() read the wrong union member in two places: - The number_unsigned case wrote number_integer's bits instead of number_unsigned's, silently writing the wrong value whenever it did not fit in number_integer_t. - The number_integer case (non-negative branch) picked the encoded width by comparing number_unsigned's bits, which is undefined behavior, though the value written was still number_integer's, so at worst a too-wide encoding was chosen. Read the active member in both cases, like the other binary writers (CBOR, UBJSON, BJData, BSON, BON8) already do. Fixes #5644. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Cast number_integer to number_unsigned_t only once in to_msgpack() Addresses review comment by @gregmarr. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e444a66276 |
Copy values before inserting an initializer list into an array (#5693)
* Copy values before inserting an initializer list into an array
insert(pos, {...}) inserted wrong values when the initializer list
contained const references to elements of the array being inserted
into. json_ref stores only a pointer for a const lvalue, so the
initializer_list_t range passed straight to the array's range insert
aliased the array's own storage; std::vector::insert(pos, first, last)
may move or shift elements before copying from that range, so the
source elements were already stale by the time they were read
(different wrong results on libc++ and libstdc++).
Copy the referenced values into a temporary array_t first, then move
that temporary into place, so the source range never aliases the
array being modified.
Fixes #5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Use the reserve_array helper in the initializer_list insert fix
The previous commit called array_t::reserve() directly on the
temporary buffer used to copy an ilist's values before inserting.
std::deque, a documented ArrayType (tests/src/unit-custom-array-type.cpp),
has no reserve(), so insert(pos, initializer_list) no longer compiled
for it. Use the existing detail::reserve_array() SFINAE helper (already
used by the SAX DOM parser) instead, which leaves array types without
reserve() untouched.
Added a regression check that deque_json::insert(pos, {...}) compiles
and handles the aliasing case from #5656.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
---------
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
1151826508 |
Take diff()'s fast path unless the object type reorders members (#5691)
* Take diff()'s fast path unless the object type reorders members For every object type except an insertion-ordered one like ordered_map, diff() no longer produced a member-by-member patch when target had a key that sorts before a key the two objects share: it fell through to the slow path, which removes every member of source and re-adds every member of target, instead of just adding the new key. #5465 added an order check to require the fast path to also reproduce target's member order, needed because ordered_map's patch()-driven "add" appends a new member at the end. The check compared the common keys' order between source and target and also required that every added key come after every common key in target's order ("new_keys_form_suffix"). The comment above it argued this check is always true for std::map, and that reasoning is correct for the order of the common keys themselves, but not for new_keys_form_suffix: a std::map iterates in sorted key order, so a new key that sorts before an existing common key is enumerated between common keys, making new_keys_form_suffix false even though std::map's own key order does not need reordering at all - it places every member itself, regardless of insertion history, so a member-by-member diff already reproduces target's iteration order. Only require the order check for an object type that keeps insertion order, using the same detail::is_ordered_map trait the library already uses to recognize such an object type in set_parent(). Every other object type - std::map in key order, a hash map in an order its operator== ignores - always takes the fast path. Added a regression test to unit-json_patch.cpp: the issue's example now yields a single "add" op for json, while ordered_json still takes the slow path to reproduce target's member order. Fixes #5639. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Only track the target key order in diff() for insertion-ordered objects common_keys_target_order and new_keys_form_suffix are only read when object_t keeps its members in insertion order; skip building them otherwise. Addresses review comment by @gregmarr. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6218aa212b |
Throw std::length_error for operator[](SIZE_MAX) instead of corrupting the array (#5687)
For idx == SIZE_MAX, the non-const array operator[] computed the new size as idx + 1, which wraps to 0. resize(0) then emptied the array, and the subsequent operator[](idx) on the now-empty vector wrote one element before its buffer. Every other too-large index (e.g. SIZE_MAX - 1) already went through resize(), which throws std::length_error and leaves the array unchanged; SIZE_MAX was the one value for which the overflow bypassed that safety net. Add a guard that throws std::length_error before computing idx + 1 when idx is the largest representable size_type value, so the array is left unchanged, matching the exception vector::resize() already throws for smaller (but still too large) indices. Fixes #5647. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
42f89e3130 |
Let to_json(std::optional<T>) propagate exceptions from T's to_json (#5684)
* Let to_json(std::optional<T>) propagate exceptions from T's to_json The overload was marked noexcept even though its body assigns *opt to the JSON value, which calls T's to_json (or allocates for std::string, std::vector, or json). Any exception from there -- a user-defined to_json reporting an error, or std::bad_alloc -- called std::terminate() instead of propagating. The noexcept also made basic_json's converting constructor noexcept(true) for std::optional<T>, so json j = opt; could not report the error either. Fixes #5642. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: make to_json(std::optional<T>) conditionally noexcept GCC's -Wnoexcept (an error in ci_test_gcc) fired at to_json_fn's noexcept(noexcept(to_json(j, val))): after dropping the unconditional noexcept, to_json(std::optional<int>) had no exception specification although GCC could prove its body cannot throw. It also made json(std::optional<int>) lose its noexcept. Declare the overload noexcept exactly when assigning the contained value to the JSON value is (std::is_nothrow_assignable<BasicJsonType&, const T&>), which is what the body does. std::optional<int> is noexcept again; a T whose to_json may throw still propagates the exception. Static assertions in the test check both cases. The test's throwing to_json triggered -Wmissing-prototypes and -Wmissing-noreturn (clang) and -Wmissing-declarations and -Wsuggest-attribute=noreturn (GCC). Move the type and its to_json into an anonymous namespace, mark the function [[noreturn]], and compile them only without JSON_NOEXCEPTION, like the test that uses them. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6d7845d207 |
Fix deprecated json_pointer/string operator== warning in value() (#5683)
value(KeyType&&, default) is constrained on is_comparable_with_object_key, which passes KeyType as a reference. is_comparable's dispatch on is_json_pointer_of<A, B> only matches a json_pointer as a plain type or a plain reference, so a const-qualified reference (as produced when KeyType is deduced from a json_pointer argument) fell through to is_comparable_no_json_pointer, which instantiates the deprecated json_pointer/string comparison operators. This made ordered_json's transparent comparator (and any transparent comparator on a custom string type) warn under -Wdeprecated-declarations when calling value(json_pointer, default), even though no such comparison is ever performed. at() was already fixed for this in #5289, which does not use is_comparable_with_object_key. Strip references and cv-qualifiers with uncvref_t before the is_json_pointer_of dispatch, so any reference-to-json_pointer is recognized regardless of qualifiers. Fixes #5664. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
678fd3017b |
Fix element path for map/unordered_map JSON_DIAGNOSTICS errors (#5681)
When converting a JSON array to std::map or std::unordered_map with a non-string key, each element must itself be a [key, value] array. If an element is not an array, from_json() threw type_error 302 with the outer array's value (&j) as the exception context, so with JSON_DIAGNOSTICS enabled the message pointed at the whole array instead of the offending element (e.g. "(/outer/m)" instead of "(/outer/m/2)"), even though the message text already described the element's type. Both from_json() overloads now pass the element (&p) as the context, so the reported JSON Pointer matches the type named in the message, the same way std::vector<std::vector<T>> and similar conversions already do. Fixes #5668. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
fc4c9c3446 |
Fix clear() to also reset the subtype of a binary value (#5680)
* Fix clear() to also reset the subtype of a binary value clear() on a binary value cleared the bytes but left the subtype untouched, so the result was not equal to a default-constructed binary value even though the documentation says clear() has the same effect as *this = basic_json(type()). The fix calls byte_container_with_subtype::clear_subtype() alongside the existing clear() call. Extended the "filled binary" clear() test in unit-modifiers.cpp with a case that uses a subtype, since the existing cases only covered binary values without one. Fixes #5669. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the table alignment in clear.md Addresses review comment by @gregmarr. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
68beba727c |
Fix from_json() for enums with underlying type bool (#5679)
get_arithmetic_value() rejects boolean_t, so the default from_json()
for enums failed to compile for an enum whose underlying type is
bool (e.g. enum class Flag : bool { off, on }), even though the
matching to_json() serializes such enums as an unsigned number.
Read the underlying value through number_unsigned_t in that case,
matching what to_json() writes, then cast back to the underlying
type before constructing the enum.
Fixes #5671.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
0050d0f7a4 |
Share one nesting depth limit between all bounded descents (#5637)
Copying and comparing stopped their descent at basic_json::nesting_depth_limit(), while serializing, hashing and merging used detail::recursion_depth_limit(). Both were 128, but nothing kept them equal. The thread-local count now tests against detail::recursion_depth_limit() as well, and a static_assert keeps the limit small enough for the byte that holds the count. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6a8a7735ed |
Do not throw in contains() for an empty array reference token (#5614)
* Do not throw in contains() for an empty array reference token json_pointer::contains() rejected malformed array indices, but an empty reference token (e.g. "/a/" where "a" is an array, or "/" on an array) passed every check and reached array_index(), which throws out_of_range.404. contains() must not throw (cf. #5395), so it now returns false for an empty token. at() still throws out_of_range.404. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: bind j_nested_const by reference in the json_pointer test clang-tidy (ci_clang_tidy) flagged the new test with performance-unnecessary-copy-initialization: the local copy j_nested_const of j_nested is never modified. Bind it as a const reference instead; it still exercises the const overloads of at() and contains(). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
65af260117 |
Move instead of deep-copy ordered_json values when an object grows (#5609)
* Move instead of deep-copy ordered_json values when an object grows ordered_map keeps its elements in a std::vector<std::pair<const Key, T>>. With a std::string key, that pair is not nothrow move constructible (the const key has to be copied), so std::vector copies every element when it reallocates. For ordered_json, this deep-copies every member value an object already holds, including whole nested subtrees, on each growth step. Grow the storage in ordered_map instead, copying the keys and moving the values. This happens in two phases, so the strong exception guarantee is kept without try/catch. The first phase may throw, but only touches a temporary buffer: it copies the keys, value-initializes the values, and constructs the new element. The second phase moves the values (noexcept) and swaps the buffers. Because the new element is constructed before any value is moved, arguments that refer to elements of the container stay valid, as with std::vector. Types that cannot take this path keep the std::vector behavior. Parsing into ordered_json (ParseStringOrdered, Apple M1 Max, clang -O3): twitter 3.20 -> 1.70 ms, citm_catalog 7.73 -> 3.67 ms, jeopardy 219 -> 177 ms, canada unchanged. The number of allocations for twitter and citm_catalog drops by two thirds. Also add ParseStringOrdered rows to the benchmarks, and document the growth behavior and the exception safety of ordered_map. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: skip std::pair noexcept assumptions on EDG-based compilers ci_icpc and ci_nvhpc failed to compile unit-ordered_map.cpp: the static assertion that std::pair<const std::string, ordered_json> is not nothrow move-constructible fails there. The EDG front end (Intel icpc 2021.10, NVIDIA nvc++ 25.5) considers the defaulted move constructor of std::pair<const Key, T> noexcept even if copying Key can throw. With these compilers, std::vector already moves such elements itself when it grows, and ordered_map correctly leaves growing to it. The same misjudgement makes std::vector call std::terminate when a key copy throws during growth, so the exception-safety test with throwing_key would abort on these compilers as well. Skip the static assertion and the exception-safety section when __EDG__ is defined. Verified with icpc 2021.10 (-std=gnu++11) and nvc++ 25.5 (C++11 and C++17) on Compiler Explorer: unit-ordered_map and unit-disabled_exceptions build and pass. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
d268eaa693 |
Convert floats with Eisel-Lemire when std::from_chars is unavailable (#5617)
* Convert floats with Eisel-Lemire when std::from_chars is unavailable Float tokens that Clinger's fast path cannot convert (e.g. the 17-digit coordinates of canada.json) went to strtod unless std::from_chars was available. It is not used in C++11/14, and not with libc++, which does not define __cpp_lib_to_chars. The Eisel-Lemire algorithm (after fast_float's compute_float) now converts them with integer arithmetic, correctly rounded for any token with at most 19 significant digits. Longer tokens are truncated; the result is used if w and w + 1 round alike, else strtod decides as before. Overflow still yields infinity (out_of_range.406). The table of powers of five (fast_float's) lives in pow5_table.hpp; a unit test recomputes every entry with big-integer arithmetic. Further tests: known values generated with Python (whose float() is correctly rounded), 200,000 round trips through to_chars, and the 128-bit multiplication and leading-zero count against big-integer references (both with and without a 128-bit type). Checked against strtod on 6.5 million tokens, among them 60,000 exact halfway cases: no difference. json::parse on canada.json: -8.6% (C++11), -7.6% (C++17, Apple clang); other files unchanged. Compile time of a TU including json.hpp: +0.7%. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix CI: unused parse result and find() == npos in the Eisel-Lemire tests GCC (-Werror=unused-result) rejected CHECK_THROWS_WITH_AS(json::parse(...)) because parse() is [[nodiscard]]; assign the result to a dummy json as the other tests do. clang-tidy flagged longer.find('.') == npos with abseil-string-find-str-contains; store the position in a variable first. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
5b9ffae8bb |
Move the float conversion chain out of the lexer (#5616)
lexer::convert_number() converted float tokens with std::from_chars (when available), Clinger's fast path, and the locale-aware strtod fallback, all as lexer members. They are now free functions in number_parse.hpp: - convert_float_fast(): std::from_chars, then Clinger's fast path, skipped when the mantissa has too many significant digits - convert_float_locale_aware(): strtof/strtod/strtold with the decimal point of the current locale, retried when the locale changed (#5198) so that other code converting JSON number tokens gets the same values. No change in behavior; the lexer no longer includes <clocale> and <cstdlib>. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
04f6ddd227 |
Keep external headers as #include when amalgamating (#5615)
A header that builds on json.hpp (such as the planned json_view.hpp) must not inline json.hpp: its single-header version would contain a second copy of the library, and that copy would change with every library change. The optional config key "external" lists include paths that are kept as #include directives. Only the first directive per path is kept; repeated ones are commented out, as the tool already does for inlined headers. config_json_view.json uses it for json_view.hpp; the existing configs do not set it, and json.hpp and json_fwd.hpp regenerate byte-identically. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
633de8e44b |
Fix CI: clang-tidy and GCC -Wnoexcept in the locale test (#5613)
#5597 was merged before all of its CI jobs had run, and two of them fail on develop now, and so on every pull request: - ci_clang_tidy: cert-err33-c for the two std::setlocale(LC_NUMERIC, "C") calls whose result was discarded. Check the result, like the other resets in the file. - ci_test_standards_gcc (20) with GCC 16: -Wnoexcept for the two parser callbacks, which cannot throw but were not declared noexcept. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
fc03b9912e |
Look up the locale decimal point at conversion time, not lexer construction (#5597)
* Look up the locale decimal point at conversion time, not lexer construction The lexer read localeconv()->decimal_point once in its constructor and wrote that character into token_buffer in place of '.'. The strtod fallback then used the locale current at conversion time, so an LC_NUMERIC change in between (parser callback, SAX handler, another thread) truncated the value in release builds and fired the endptr assertion in debug builds. token_buffer now always holds '.'. Only the strtof/strtod/strtold fallback depends on the locale: it looks up the decimal point right before the call, restores '.' afterwards, and repeats the conversion if the locale changed in between. As a side effect, std::from_chars and Clinger's fast path now also apply under locales whose decimal point is not '.'. Fixes #5198 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop the strtod retry loop when the decimal point is unchanged convert_float_locale_aware() repeated the conversion until strtod consumed the whole token, assuming an early stop can only mean a locale change. Under a locale whose decimal point is not a single character (e.g. the two-byte U+066B of ar_EG.UTF-8, ar_SA.UTF-8, or fa_IR.UTF-8, all available on macOS), the in-place substitution can never succeed, so parsing any float that reaches the strtod fallback (for example 3.14159265358979323846 at C++11) hung forever. Before this branch, the same input was truncated. Retry only if the decimal point changed since the previous attempt; otherwise keep the value strtod parsed so far, as before. Add a test that parses such numbers under a multi-byte decimal point locale; it hangs without this change. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix -Weffc++ errors in the #5198 locale test GCC's -Weffc++ (an error in ci_test_gcc and ci_test_standards_gcc) rejected LocaleSwitchingSax: it has a pointer data member but does not declare its copy operations, and its vectors are not initialized in the member initializer list. Store the locale name as a std::string and give the vectors brace initializers, like SaxEventLogger in unit-deserialization.cpp. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
9e1a09eec0 |
Name the key type when rejecting non-string CBOR/MessagePack map keys (#5594)
* Name the key type when rejecting non-string CBOR/MessagePack map keys CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings, so such maps are rejected. The error so far was the one for a malformed string (e.g. "expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xC0" for a nil key), which does not tell the user what went wrong. Report the type of the key instead: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0 The exception id (parse_error.113) and type are unchanged. Malformed string keys and a missing key keep their previous messages. Document the restriction on the CBOR and MessagePack pages. Refs #2766, #3381 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Point the MessagePack key note to the spec's profile section The note linked to "Serialization: type to format conversion", which says nothing about key types. Restricting map keys to strings is only mentioned in the "Profile" section (under "Future discussion") as an example of a JSON-compatible profile, so link there and describe it as such instead of as a permission. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
373005f7ac |
Fix MSVC: avoid reserving by faked size in MessagePack size tests (#5604)
The "Size above uint32" tests for arrays and objects fake a container size of 2^32 and expect to_msgpack() to throw out_of_range.412. But to_msgpack(j) first reserves binary_reserve_hint(j) bytes, which is size + 1 for arrays and 2 * size + 1 for objects, i.e. 4 or 8 GiB. Linux and macOS overcommit, so the reservation succeeds; on Windows it throws std::bad_alloc before the size check is reached (seen with msvc-vs2026 Debug x64 on the object test). Write into a caller-owned vector instead, so nothing is reserved. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
de6acd651e |
Fix CI: wrap an overlong line in the cbor_tag_handler_t documentation (#5602)
#5559 added a 224-character line to cbor_tag_handler_t.md; the documentation style check allows at most 160. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
56ddcb65f0 |
Fix documentation style check: wrap long line in cbor_tag_handler_t.md (#5603)
The line added in #5559 exceeded the 160-character limit enforced by docs/mkdocs/scripts/check_structure.py, breaking the documentation build. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
509c07041f |
Fix CI: clang-tidy and clang/libstdc++ 10 in the MessagePack size tests (#5599)
The tests added by #5515 fail two ways on develop: - clang-tidy reports the size() overrides of huge_string and huge_binary (readability-convert-member-functions-to-static) and the non-const test value (misc-const-correctness); mark them like the #5584 types - clang with libstdc++ 10 cannot compile the file for C++17: the std::filesystem::path conversion considered for huge_string, a class derived from std::string, is ambiguous. Guard it with JSON_TEST_BEYOND_UINT32_STRING, which #5584 introduced for the same reason, and define that macro before both test blocks. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
1e101ecac1 |
Add BON8 support (#2998)
* Add BON8 support Add to_bon8/from_bon8 and input_format_t::bon8 for BON8, a binary format that uses the byte values that cannot begin a UTF-8 character as type markers, so strings need no length prefix. It is the most compact of the supported binary formats on the benchmark files. The reader is non-recursive like the other binary readers. A string ends at the first byte that cannot continue it, so the reader hands the one or two bytes it reads past a string back to the value that follows. The writer produces the canonical representation of the specification, except for NFC normalization; its output is identical to that of the reference implementation (HikoGUI) on all files of the test data. The round-trip tests need the .bon8 files of json_test_data 3.2.0. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Address review comments - Reuse detail::validate_one_utf8 to check strings in to_bon8; the error now names the first byte of the invalid sequence. - Document that to_bon8 leaves bytes in the output adapter on an exception, and that string_open is only an output of write_bon8_marker. - Explain why the pushback buffer of the BON8 reader cannot overflow. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Select the BON8 float prefix by type get_bon8_float_prefix only depends on the type of its argument, so make the type a template parameter instead of passing an unused value. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Rename a test variable that Flawfinder mistakes for read() Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures - compare the float in write_bon8_float with number_float_t constants, so GCC does not warn about a float-to-double conversion - mark check_bon8_utf8's context as used when exceptions are disabled - choose the compact float prefix in a helper rather than with nested conditional operators (clang-tidy) - use auto for the cast in the BON8 integer reader (clang-tidy) - write the int32 minimum test values as long long literals (MSVC C4146) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Amalgamate Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BON8 strings in bulk from contiguous input - copy the valid UTF-8 of a string in one step when the input is contiguous (twitter.json is read in 1.68 instead of 2.52 ms, jeopardy.json in 196 instead of 297 ms, close to CBOR and MessagePack) - share the new valid_utf8_prefix() with the writer's UTF-8 check, which now skips ASCII 8 bytes at a time - let the fuzzer check that contiguous and stream input give the same value or error, and test both paths in the unit tests - clarify that a second 0xFF after a string is an empty string Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Link the BON8 functions from the other binary format pages Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Name the bulk scan flag after the input, not BON8 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read BSON keys in bulk from contiguous input BSON keys (and array indices) are C-style strings, which were read byte by byte. For contiguous input they are now read up to their \x00-byte in one step, using the same bulk_scan flag as BON8 strings: twitter.json is read in 1.46 instead of 2.01 ms, citm_catalog.json in 2.93 instead of 3.33 ms, jeopardy.json in 182 instead of 207 ms. canada.json, whose keys are almost all one-digit array indices, takes 2 % longer. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix the BON8 CI failures of the bulk-read tests - skip the contiguous-versus-stream tests of BON8 strings and BSON keys when exceptions are disabled: they catch the parse errors of invalid input, and without exceptions the library aborts instead - use static_cast for the int64 test value (google-readability-casting) Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Move the explicit basic_json instantiation into its own test file Linking test-regression3_cpp20 with clang and MinGW failed with "relocation truncated to fit: IMAGE_REL_AMD64_REL32 against `.rdata'", as test-regression2 did before #5511. The explicit instantiation of basic_json<> for #4825 compiles every member function, including the BON8 reader and writer, into that object, and it was already close to the limit (2,226,104 bytes on develop, 2,234,960 with BON8; clang -O1, C++20). Give the instantiation a file of its own: unit-regression3 is now 1,594,736 bytes and unit-explicit_instantiation 1,095,064. The new file mentions JSON_HAS_CPP_17 and JSON_HAS_CPP_20 so it keeps being built for the C++17 standard the regression was about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert the bytes of the BON8 test strings explicitly The str() helper constructed a std::string from a byte range, which converts each unsigned char implicitly; -fsanitize=integer reports that for bytes of 0x80 and above (ci_test_clang_sanitizer). Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |