mirror of
https://github.com/nlohmann/json.git
synced 2026-10-03 13:10:33 +00:00
f579e7e05507a649208d8abe7281ec6b17a62368
1170
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d675cdb46 |
Fix CI jobs that check less than they claim; move arm64 to GitHub (#5733)
* Fix the ci_cmake_flags wiring so every option is checked
The CMake 3.31.6 flag list referred to itself before it was defined,
so only JSON_BuildTests was checked with that version. The targets for
the CMake running the build ("_2") were created but never added to
ci_cmake_flags, and the three versions shared one build directory.
JSON_StrictNulHandling was not in the list at all.
Use the 3.5.0 list for 3.31.6, add JSON_StrictNulHandling, and create
one ci_cmake_flag_<flag> target per option for the running CMake with
its own build directory. Also use the function parameter in the
COMMENT, refresh the stale version comment, and let ci_clean remove
the downloaded cmake-<version> directories instead of the long-gone
cmake-3.5.0-Darwin64.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Fix ci_test_clang_libcxx_cxx* jobs silently building without warnings
CMake only seeds CMAKE_CXX_FLAGS from the CXXFLAGS environment variable
when the cache entry is unset, so the explicit -DCMAKE_CXX_FLAGS="-stdlib=libc++"
argument made it ignore CXXFLAGS="${CLANG_CXXFLAGS}" entirely. The six
ci_test_standards_clang (..., libcxx) jobs therefore compiled without
-Weverything/-Werror while their libstdc++ siblings did use them.
Pass -stdlib=libc++ through the same CXXFLAGS value instead of a separate
-D argument, and give the target its own build directory
(build_clang_libcxx_cxx${CXX_STANDARD}) so it no longer shares a CMake
cache with the libstdc++ variant. Suppress the resulting
-Wthread-safety-negative finding from libc++'s std::mutex annotations,
which fires on doctest's reporters in this translation unit only.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove the no-op AppVeyor with_win_header job
The with_win_header matrix entry patched Windows.h into
single_include/nlohmann/json.hpp before building, but JSON_MultipleHeaders
has defaulted to ON since #3532 (2022-06), so CMakeLists.txt points the
tests at include/ and the patched single header is never compiled. The
job has been a no-op VS2015 build since then.
Windows.h coverage already exists through tests/src/unit-windows_h.cpp
(#3631), which runs in every MSVC job. Delete the dead matrix entry and
its before_build steps, and cite unit-windows_h.cpp from the QA page.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Remove unused ci_oclint and ci_pvs_studio targets
No workflow invokes ci_oclint, ci_pvs_studio, or their tool discovery.
ci_oclint also had a side effect on every JSON_CI configure: it copied
the single header into src_single/all.cpp and added an add_executable()
for it without EXCLUDE_FROM_ALL, so a plain build compiled a 1.2 MB
translation unit that only that unused target consumed. ci_pvs_studio
duplicates the Makefile's pvs_studio target, which is kept.
Also drop the duplicate --check-level=exhaustive flag passed twice to
the same ci_cppcheck invocation.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Stop Dependabot from proposing astyle bumps
astyle is deliberately pinned at 3.4.13 because newer versions reformat
unrelated lines and this version defines the formatting that
check_amalgamation.yml enforces. Without an ignore rule, Dependabot
keeps opening PRs for every new astyle release (most recently #4580,
#4942, #5445, #5448), each of which fails the amalgamation check and
gets closed unmerged.
Part of #5715
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
* Move Linux arm64 CI from dead Cirrus CI to ubuntu-24.04-arm
Cirrus CI stopped reporting check runs on develop sometime after
|
||
|
|
c261578431 |
Deduplicate binary reader/writer helpers and fix stale comments (#5730)
* Fix stale and missing comments in binary_writer The doc block of write_number() ended up above the byte_swap() helpers added in #5286, about 80 lines from the function. It was also a plain comment that Doxygen skips, said "write a number to output input", and left BON8 out of the big-endian formats. Move it back onto write_number() as a /*! block and fix the text. write_bson() documented "@pre j.type() == value_t::object", but it throws type_error.317 for every other type, and to_bson() relies on that. Document the exception instead. Explain why the CBOR binary subtype is always written with a 0xD8..0xDB head and never in the one-byte tag form: binary_reader with cbor_tag_handler_t::store only keeps those heads as a subtype, so switching to write_cbor_head() would break round trips for subtypes 0..23. Also fix the grammar of the to_char_type comment. Comments only; no change in behavior, API or ABI. Part of #5710 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Merge the duplicated UBJSON/BJData integer marker ladders write_number_with_ubjson_prefix() (unsigned and signed overloads) and ubjson_prefix() (number_integer and number_unsigned cases) each picked the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their own independent if/else ladder, and the values beyond 64 bits were handled by a second, tag-dispatched pair of ladders. An optimized container announces the marker of its first element via ubjson_prefix() and then writes every element through write_number_with_ubjson_prefix(), so the two had to be kept in lockstep by hand across four call sites. Replace all of that with one ubjson_integer_prefix() built on value_in_range_of<T>, and one write_ubjson_integer_payload() that writes the value (or, for 'H', the decimal digits) for a given marker. write_number_with_ubjson_prefix() and ubjson_prefix() keep their signatures and now just call these two helpers. Behavior, the public API and the ABI are unchanged. Verified with a new regression test covering scalars and $-optimized arrays/objects at every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing to_ubjson/to_bjdata output before and after over the json_test_data corpus (bit-identical). Part of #5710 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove dead get_char parameters in binary_reader The non-recursive rewrite of the binary readers (#5505, #5506, #5507) left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char parameters dead: parse_cbor_internal() has one caller and it always passes true, and parse_ubjson_internal() has one caller and it always uses the true default. Both parameters, and the @param docs describing the "reuse the last character" mode they used to select, no longer correspond to anything. Drop both parameters, initialise fetch/prefix unconditionally, and update the two call sites in sax_parse(). parse_cbor_value()'s and get_ubjson_string()'s own get_char parameters are unrelated and are left alone; both still have a false caller. Also delete a stray `@return whether a valid MessagePack value was passed to the SAX parser` doxygen block that sits directly above parse_msgpack_value()'s real doc comment, a leftover of the same rewrite. Behavior, the public API and the ABI are unchanged; these are private members of detail::binary_reader. Verified by compiling with -Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and unit-msgpack (offline, against the stubbed test_data.hpp). Part of #5711 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Share the IEEE half-precision decoder between CBOR and BJData binary_reader had two ~45-line copies of the IEEE 754 half-precision decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is normalised, the two blocks were identical except for the byte order used to assemble the 16-bit half (CBOR is big endian, BJData is little endian). Any future change to half-float decoding had to be made and kept in sync in both places. Add one get_half_float(format, little_endian) helper that does the two get()/unexpect_eof() reads, assembles the half in the requested byte order, decodes it per RFC 8949 Appendix D, and calls sax->number_float. Both cases now just call it with their byte order; the BJData case keeps its bjdata-only guard. Behavior, the public API and the ABI are unchanged. Verified with a scratch probe comparing the old and new decoders bit-for-bit (NaN by isnan()) over all 65536 wire byte pairs, in both formats, and by running unit-cbor and unit-bjdata (offline, against the stubbed test_data.hpp). Part of #5711 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the MessagePack unsigned-integer writer ladder The number_integer (non-negative branch) and number_unsigned cases in write_msgpack() each held their own copy of the fixint/uint8/16/32/64 ladder, kept in lockstep only by a comment ("we used the code from the value_t::number_unsigned case here"). Both copies mixed union members: the signed copy compared number_unsigned but wrote number_integer, and vice versa. Extract write_msgpack_unsigned(std::uint64_t), mirroring how write_cbor_head() already avoids the same duplication for CBOR, and call it from both cases. Each case now reads only its own active union member. Output bytes are unchanged for the default 64-bit number types. #5710 item 3 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Unify float marker selection and fix the long double compile error Four formats picked between a float32 and float64 marker through four different helper styles: dummy-argument overloads for CBOR and MessagePack, an std::is_same template for BON8, and a runtime if-chain on input_format_t for write_compact_float(). With number_float_t set to long double, to_cbor, to_msgpack and to_ubjson failed inside the library with "call to 'get_cbor_float_prefix' is ambiguous", while to_bson kept working because write_bson_double() takes a plain double. Change write_compact_float() to take the two marker bytes directly (each of its three callers already knows them at compile time) instead of an input_format_t it only forwarded, and delete the now-unused get_cbor_float_prefix(), get_msgpack_float_prefix(), get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn the two get_ubjson_float_prefix() overloads into one template. Both write_compact_float() and get_ubjson_float_prefix() now report an unsupported number_float_t with a static_assert naming the requirement, rather than an ambiguous-overload error; the assert lives in the function body, not the class scope, so to_bson with long double is unaffected. Verified with a probe basic_json<..., long double>: to_bson still compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail to compile with the new static_assert message. This changes the text of an existing compile error for users with an unsupported number_float_t (documented as a public-API-visible change in #5710). #5710 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the BJData ndarray writer's dtype dispatch and drop <map> write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on every call just to translate the _ArrayType_ name to a dtype marker (the only reason binary_writer.hpp included <map>), then mapped dtype to C++ type twice more: once as a switch for the range-check pass and once as a separate if/else chain for the write pass, with nothing checking that the two agreed. The caller also ran three at() lookups, and the callee called value.at(key) about ten more times for the same three members. Replace the map with bjdata_ndarray_type_marker(), a plain string comparison chain (a C++11 constexpr function cannot contain a switch, so this mirrors binary_reader's own static table style). Replace the switch/if-chain pair with one write_bjdata_ndarray_elements() that switches on dtype once and calls a per-type helper - write_bjdata_ndarray_element<T>() for the eight integer dtypes and write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag selecting the range check or the actual write, so the two passes can no longer disagree on the type. _ArrayType_, _ArraySize_ and _ArrayData_ are now looked up once into references, and the four header marker bytes ('[', '$', '#') are written through to_char_type() like the rest of the UBJSON/BJData writer. The 'd' (single-precision) rule is left exactly as before, since #5707 is expected to change it separately. Verified byte-for-byte identical output before/after for every dtype (including the Draft 2/Draft 3 'byte' fallback and the use_count/ use_type combinations) via a standalone probe, plus round-tripping through from_bjdata(). Overlaps #5707, which is expected to touch the 'd' dtype case, and #5518, which is expected to move the write_bjdata_ndarray() call site. #5710 item 4 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Assert that write_bson_document() consumes every calc_bson_sizes() entry calc_bson_sizes() and write_bson_document() are a hand-synchronized pair of passes over the same object/array tree, introduced by #5553: the size pass appends to nested_sizes in visiting order, and the write pass consumes the table by position with nested_sizes[next_size++]. Nothing checked that the write pass consumed the whole table. If a future change touched only one of the two passes - for example to skip or reject an entry - every later size prefix in the document would be silently wrong. Add JSON_ASSERT(next_size == nested_sizes.size()) where write_bson_document() returns, so such a future drift between the two passes is caught immediately (JSON_ASSERT expands to nothing in release builds using assert(), and the fuzzers/tests already build with it enabled). The two passes agree today, so this changes nothing observable; it only guards against the risk described in #5710 item 5. Extracting a shared stepper for the two passes (the second half of the proposed change) is left for a follow-up: it only saves ~30 lines and the issue asks for it only if the result reads clearly, which needs more room to get right than a mechanical cleanup pass allows. #5710 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Make the BJData lookup tables static functions instead of members binary_reader held bjd_optimized_type_markers and bjd_types_map as non-static const members (12 string_t objects for the type-name table), built and destroyed on every from_cbor/from_msgpack/from_bson/ from_ubjson/from_bon8/from_bjdata call even though only from_bjdata ever reads them. They also needed the #define/decltype/#undef workaround from #3637 and two NOLINTNEXTLINE suppressions, and binary_writer already carries the same two lists in another form (is_bjdata_excluded_type_marker() and a local std::map in write_bjdata_ndarray(), the latter removed by the item-4 commit), so the excluded-marker lists could drift apart. Replace bjd_optimized_type_markers with static constexpr is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain as binary_writer's is_bjdata_excluded_type_marker(). Replace bjd_types_map with a non-constexpr static bjd_type_name(char_int_type) switch returning nullptr for an unknown marker (a C++11 constexpr function cannot contain a switch). Delete both JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the NOLINTNEXTLINE suppressions, detail::make_array() (no longer used anywhere), and the now-unused <algorithm> and <array> includes. Update the two call sites (the ND-array excluded-type check and the _ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's "LUT arrays are sorted" section, which only checked the two tables' internal ordering, with a check of all 12 type names and all 8 excluded markers against both new functions. #5711 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Read CBOR's 1/2/4/8-byte argument through one helper parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte big-endian unsigned integer" ladder four times over: - twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore branch and once, nearly identically, in the ::store branch (~90 lines to read one integer); - twice more for container lengths, once for array heads 0x98-0x9B and once for map heads 0xB8-0xBB, where the 1/2-byte forms called enter_array()/enter_object() directly and the 4/8-byte forms additionally went through get_cbor_container_size(). Add get_cbor_argument(std::uint64_t&), reading the width selected by current & 0x1F via the same get_number() calls as before (so EOF is reported exactly as before), and route all four sites through it: - 0xD8-0xDB now read the argument once per branch instead of switching on `current` a second time; behavior split cleanly from embedded tags 0xC0-0xD7 (tag value in the head, no argument to read), which is now its own case block that no longer has to fall into the ::store switch's "default" case to reach the same tag_pending = true; return true; outcome. - 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always going through get_cbor_container_size() (harmless for 1/2-byte lengths, which already always fit). Verified byte-for-byte identical behavior before/after with a standalone probe covering embedded and multi-byte tags under all three tag_handler_t settings, a tag over a byte string (subtype path), truncated tag/length arguments of every width, and array/map lengths of every width, including the out_of_range.408 "excessive size" case: same exceptions, same messages, same chars_read, same successful results. Left the string/byte-string length ladders in get_cbor_string()/ get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is expected to touch them separately. Overlaps #5601 (adds a branch right above the embedded-tag case) and #5607 (touches the integer cases 0x18-0x1B, which share this ladder's shape in separate hunks). #5711 item 2 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Add leave_container() to match enter_container() Every container is opened through enter_container(), whose docs promise that a check placed there runs before every start event. The close side had no equivalent: the same "container_stack.pop_back(); dispatch to end_object() or end_array()" sequence was written out separately in BSON, CBOR, MessagePack, UBJSON/BJData and BON8, each copying the pattern of keeping an is_object flag around the pop_back() that would otherwise invalidate a reference to it. A check needed on close would have had to be added in five places, and a sixth copy could go unnoticed. Add leave_container() next to enter_container(), doing the same pop-then-dispatch, and replace the five sites with it. Each site keeps its own surrounding logic (BSON's check_bson_document_size() call before popping, MessagePack's is_object copy used again below, UBJSON/BJData's remaining-container handling after popping, BON8's top used again below); only the repeated pop/dispatch line pair is now shared. Verified all six binary-format unit suites and unit-regression2's deep-nesting tests (dependent count/reuse count and the bjdata ndarray depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan. Overlaps #5601, which is expected to add a sixth close site in its own skip loop; that site can route through leave_container() too once it lands. #5711 item 4 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Stop passing the input format to sax_parse() when the reader already has it binary_reader's constructor stores the format in the input_format member, and sax_parse(format, sax_, strict, tag_handler) took the same value again purely to dispatch on it. Every in-tree caller passed the same value both times (all 16 from_cbor/from_msgpack/from_ubjson/ from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three public basic_json::sax_parse() overloads), so nothing was broken today, but a caller of the detail class directly (only reachable via JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could pass a mismatched pair - say bjdata to the constructor and ubjson to sax_parse - and dispatch on one format while applying the other format's rules; the default-constructed input_format_t::json reader would additionally hit JSON_ASSERT(false) in exception_message() on its first error. Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the existing overload with the stored input_format, and switch every caller to it: the 16 from_*() sites (keeping their `// cppcheck-suppress[accessMoved]` comments) and the three basic_json::sax_parse() overloads, all of which already had the format available from their own `format` parameter. The four-argument overload is kept for anyone still calling it, now with JSON_ASSERT(format == input_format) so a mismatch fails immediately in a debug build (assert-enabled binaries, including the fuzzers and test suite) instead of misbehaving; verified with a probe that constructs a reader for one format and calls the explicit overload with another, which aborts on that assertion as expected. Removing or asserting against the constructor's input_format_t::json default, which would affect direct detail users, is left as a separate decision per #5711 item 5. Overlaps #5601, which is expected to add an AllowRecovery template parameter to sax_parse() and touch these same call sites in json.hpp. #5711 item 5 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate UBJSON/BJData signed-count handling, drop dead ndarray checks get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently sized signed integer and then repeated the same "reject negative with error 113" check; only 'L' additionally checked value_in_range_of for the out_of_range.408 case. Any change to that error path had to be made four times. Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read, the negative check and the range check once, and route all four markers through it. The range check is a no-op for 'i'/'I'/'l' (their values always fit std::size_t) and only live for 'L' on a 32-bit std::size_t target, matching today's behavior exactly. In the ndarray dimension-product loop, the preceding loop already returns early on any zero dimension and result starts at 1, so `i > 0` in the pre-multiplication overflow check was always true, and `result == 0` in the post-multiplication check could not be reached either: two positive factors whose product does not overflow (as the pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&` and narrow the post-check to `result == npos`, the one case the pre-check cannot rule out (an exact, non-overflowing match with the sentinel reserved for unknown-size containers), with a comment explaining why. Verified byte-for-byte identical behavior before/after with a standalone probe covering negative counts for every marker, a matching positive count, and ndarray inputs, plus the full unit-ubjson and unit-bjdata suites (same assertion counts as before this change). Overlaps #5601 (rewrites the four parse_error calls and the overflow checks touched here) and #5607/#5707 (touch neighboring lines in the same functions). #5711 item 6 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Drop redundant format parameter and dummy float argument (review) binary_reader::sax_parse(format, ...) only ever had to equal the format given to the constructor, which it asserted. With every caller already on the format-less overload, remove the four-argument overload and dispatch on the stored input_format directly. binary_reader is a detail class, so this is not a public API change. get_ubjson_float_prefix() took a value only to deduce its type; make the type an explicit template argument instead. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
35802e78d6 |
Remove dead Makefile targets; document macro_builder; tidy serve_header (#5735)
* Remove dead doctest help entry and pretty_format target from Makefile The top-level Makefile still carried three leftovers: - The help text listed a "doctest" target that was removed in #4560, so "make doctest" fails with "No rule to make target". The example check now runs as "make check_output -C docs". - "pretty_format" ran clang-format on all sources, but .clang-format was deleted in #4573, so the target reformatted everything in the default LLVM style, against the Artistic Style formatting that "make pretty" applies and CI enforces. - "clean" removed benchmarks/files/numbers/*.json, a directory that no longer exists since the benchmarks moved to tests/benchmarks (#3462). Only maintainer tooling changes; the library is not affected. Part of #5717 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Document tools/macro_builder and tidy up serve_header.py tools/macro_builder generates the NLOHMANN_JSON_EXPAND, NLOHMANN_JSON_GET_MACRO and NLOHMANN_JSON_PASTE* macros in macro_scope.hpp, but nothing referred to it. Add a README that explains what it generates, how to run it and where the output goes, and which dependent tables (NLOHMANN_JSON_DOUBLE_PASTE, NLOHMANN_JSON_TYPE_BODY) are maintained by hand. Point to it from a comment above NLOHMANN_JSON_EXPAND. The generator itself is unchanged; following the README reproduces the header byte for byte. In serve_header.py, drop the LGTM suppression (LGTM.com shut down in 2022), replace the """.""" placeholder docstrings with real ones, and import socket and ssl at module level. DualStackServer.server_bind uses socket, which was only imported under __main__; when the module was imported instead, the NameError was swallowed and IPV6_V6ONLY was not cleared. The header change is a comment only; behavior, API and ABI are unchanged. Part of #5717 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Hash every release_files artifact, not a hardcoded subset The `release` target signed and copied json_fwd.hpp into release_files alongside json.hpp, but the shasum line that writes hashes.txt only listed json.hpp, include.zip and json.tar.xz. Users could not verify the published json_fwd.hpp against hashes.txt. Hash every file in release_files except the .asc signatures instead of naming files by hand, so a newly shipped header (such as the json_literals.hpp that #5610 adds to this target) cannot be missed again. Only affects the generated hashes.txt release artifact; the library itself is unaffected. Overlaps #5610, which touches the same lines to add json_literals.hpp to the release target. #5717 item 1 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Remove the broken fuzz_testing* Makefile targets fuzz_testing and fuzz_testing_{bon8,bson,cbor,msgpack,ubjson} seeded fuzz-testing/testcases from tests/data, which was removed in |
||
|
|
7d7055ec50 |
Fix stack overflow converting deep values between specializations (#5723)
* Fix stack overflow converting deep values between specializations Constructing a basic_json from another specialization (json to ordered_json or back, also via get<ordered_json>()) converted every container with its range constructor, which calls the converting constructor for each element. The call stack therefore grew with every nesting level, and a value nested some 30,000 levels deep overflowed it. The conversion now bounds its descent the way the copy constructor does since #5387: the first 128 levels are converted exactly as before, and below that convert_iteratively() finishes the value with an explicit stack. It builds each container bottom-up from its converted elements with the container's range constructor, so member order and keys that become equal are handled as before, and it gives a value its type only once its container exists, so an exception leaves nothing behind that cannot be destroyed. Parents (JSON_DIAGNOSTICS) and positions (JSON_DIAGNOSTIC_POSITIONS) are set for every value. Converting a null value no longer resets its positions: the constructor assigned null to a value that already was null, which swapped in the positions of the temporary. Fixes #5650. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Explain why converting null keeps positions and why next is a reference Review feedback on #5723 (gregmarr): clarify in comments that the converting constructor has already copied the positions of val, which the null case keeps like every other case, and that next must be a reference into pending so that ++next advances the stored iterator. Comments only; no code change. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Refer to recursion_depth_limit() in the convert_structured() docs The comment still named nesting_depth_limit, which #5637 removed on develop in favor of detail::recursion_depth_limit(). Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Advance the pending iterator through pending.back() and shorten the null comment Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
dcc81f43f0 |
Add insert() and erase() to editable documents
- insert(array, index, value): insert before an element (index <= size) - erase(object, key): remove all members with the key; returns their number - erase(array, index): remove an element - erase(json_pointer): remove the member or element a pointer names The errors are those of basic_json (type_error.307/309, out_of_range.401/403/405). A view of an erased value keeps its last value, and views of other values keep referring to them when elements move. Tests: the differential test now also inserts and erases members and elements, directly and through JSON pointers; plus the errors, views across inserts and erasures, duplicate keys, and large objects. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
383ce0b040 |
Cover the edit storage in tests
Test edits of an empty document and assignments through a view of a value that is no longer part of the document; copy the entries of a block through deref() (one path for links and values); mark the 4 GiB limit and the returns of find_parent() that no document reaches. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
51ee239b3c |
Address the clang-tidy findings of the editable documents
Pick the overloads of encode() with a first_true trait instead of nested conditionals, name the pointer type in the copies of links, mark the owning pointers of the edit storage, and compare doubles by their bits in the tests. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6e50766dc3 |
Add editable documents with set() and push_back()
basic_json_document<BasicJsonType, true> (json_editable_document,
ordered_json_editable_document) can be edited:
- set(view, value): replace a value
- set(object, key, value): assign a member, or add it (a null becomes an
object); with duplicate keys, the first is assigned and the others go
- set(array, index, value): assign an element
- set(json_pointer, value): the member, element, or ("-", or the size of
the array) the end of an array a pointer names
- push_back(array, value): append (a null becomes an array)
Values are views (of any document, copied), BasicJsonType values, and
everything BasicJsonType can be constructed from. The source text is never
written, and the parsed index never moves: new values and element
sequences go to storage owned by the document (edit_storage.hpp), so views
stay valid, and a view keeps referring to its value (after an assignment,
it sees the new one). Read-only documents are unchanged; editing one does
not compile.
Errors are those of basic_json where the operation corresponds
(type_error.305/308, out_of_range.401/403/405, parse_error.106/109); a view
of another document is invalid_iterator.202. Strings are checked for UTF-8
when they enter the document, with the type_error.316 that
basic_json::dump() throws for the same string, so that a document only
holds valid UTF-8. Binary values cannot be stored (the new
type_error.319), and edits of 4 GiB or more end with out_of_range.416.
Views of editable and read-only documents compare with each other.
Tests (unit-json_view_edit.cpp): random assignments, member and element
changes, copies within and between documents, and pushes, applied to an
ordered_json_editable_document and to the ordered_json value; after every
edit both must serialize (also indented and with ensure_ascii),
materialize, compare, and read back the same. Further: the errors, strings
that stay valid while the edit arena grows, numbers (NaN, infinities,
extremes; number_format::source), nulls that become containers, the root
replaced, duplicate keys, values of other documents, large objects, and
documents reused with read().
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
d4fab0971f |
Walk the index of json_view through a navigation policy
Preparation for editable documents, without a change in behavior: views and documents get a template parameter Editable (false by default), and every walk over the index (iterators, lookups, dump(), materialize()) goes through detail::view::navigation<Editable>. For read-only documents it is the plain node array, as before, so they compile without any of the edit handling. For editable documents it also follows the representation of edits, which this commit defines: - node flags `edited` (a string or number token in the edit arena), `moved` (the elements of an array/object live in a separate sequence), and `is_new` (no source position), and link nodes (kind_link) that stand for a value stored elsewhere - document_data::edit_state: the moved sequences, the storage of new values, and the edit arena materialize() now keeps a frame per open container instead of returning to the end of a closed one, as the serializer does, so that it can follow moved sequences. Floats whose token lives in the edit arena (also "nan", "inf", "-inf") are converted out of line. dump() copies only strings of the source without escaping, and shrink_to_fit() leaves the node array in place once there are edits, as they link into it. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
eede67ca92 |
Fix old clang: do not declare the defaulted document_data() noexcept
With the nested struct object_index, clang 4 (and, by the same bug, the clang 3.x of ci_test_compilers_clang) rejects the explicitly noexcept defaulted constructor: "default member initializer for 'indexes' needed within definition of enclosing class 'document_data' outside of member functions". Nothing depends on the constructor being noexcept, so let it take the implicit exception specification. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
82b31f31b8 |
Fix CI: useless casts of the key hash of the view's object index
GCC -Werror=useless-cast on Linux x86-64 rejects static_cast<std::size_t>(key_hash(...)): the call returns a std::uint64_t prvalue, the same type as std::size_t there, while the cast is needed where std::size_t is 32 bits wide. Store the hash in a variable and cast that, which GCC does not report. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
904c8c6710 |
Index the large objects of json_view
Lookups in objects are linear, as for ordered_json. Objects with 128 members or more now get a hash table after parsing (open addressing; the first of duplicate keys is kept, as for the linear search), so that operator[], at(), find(), contains(), count(), value(), and JSON pointers take constant time on average in them; the idea of switching to a hash table for large objects is Boost.JSON's. The parser notes such objects when it closes them (out of line, so that the parse loop only has a call for it), and the object node keeps the number of its table. Looking up each key of an object with 10,000 members: 59.8 ms -> 0.16 ms. Parsing (json_document::parse, best of 7, separate processes): most files within 1%; canada +5%, mesh.pretty +3%, citm +3%. Tests: objects with 127, 128, 129, and 10,000 members (escaped, empty, and duplicate keys, missing keys, comparisons), nested large objects, and documents reused with read(). Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
1f0c3be6f3 |
Address the clang-tidy findings of the SIMD scan
Hold the UTF-8 lookup tables in std::array, compute the length of a sequence without nested conditionals, and use std::array in the tests. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
c3b51addbf |
Scan the strings of json_view with NEON and SSE2
Long runs of string bytes are scanned 16 at a time with NEON (AArch64, with GCC and Clang) and SSE2 (x86-64): both belong to the baseline instruction sets. A signed compare with 0x20 finds control characters and non-ASCII bytes at once. Keys keep 16 table checks before the vector loop (their lengths repeat from record to record, so the branches predict well); string values have 8, as their lengths vary more. Non-ASCII text is validated 16 bytes at a time with the "lookup4" check of simdjson (J. Keiser and D. Lemire, "Validating UTF-8 In Less Than One Instruction Per Byte", 2021): with NEON, and on x86-64 with SSSE3 if JSON_VIEW_USE_SSSE3 is defined (SSSE3 is not part of x86-64, and the code must not depend on the flags of a translation unit). JSON_VIEW_NO_SIMD selects the portable code. The vector code sits in detail/view/simd.hpp; the same input is accepted either way. json_document::parse, best of 7 runs in separate processes (M1 Max): poet.json (CJK text) -72%, random.json -25%, twitter.json -22%, gsoc-2018.json -20%, semanticscholar -19%, github_events -11%, apache_builds -9.5%, canada/citm -5/-6%; lottie +4%, tree-pretty +2.5%. Tests: every two-byte sequence and three- and four-byte sequences with continuation bytes at the edges of their ranges, at every offset around the vector blocks of keys and values, cut short, and long runs of text with a damaged byte, against json::accept and json::parse. CMake builds the parser tests again with JSON_VIEW_NO_SIMD, and on x86-64 with JSON_VIEW_USE_SSSE3 and -mssse3; the macros are documented. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
ecf9df7c45 |
Convert the floats of json_view from the digit layout
The parser records where the integer digits, the fraction digits, and the exponent of a float token are. For floats and doubles with at most 19 digits, the value is now read from that layout: the digits eight at a time, without scanning the token, and rounded by the library's conversion core (detail::decimal_to_float(): Clinger's fast path where both operands are exact, else the Eisel-Lemire algorithm, which needs no fallback for up to 19 digits). It rounds correctly, so the values are those of parse(); other tokens and types keep the library's conversion of the whole token. get<double>(), materialize(), dump(), and comparisons use it. Traversing canada.json (111,000 floats, every number converted): 0.95 -> 1.29 GB/s. Tests add tokens around the limits (19 and 20 digits, 2^53, 10^22, and those of float) to the bit-for-bit comparison with parse(). Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
a7fa8d04e8 |
Address the cpplint findings of json_view's comparisons
compare.hpp includes <string> (build/include_what_you_use). Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
677507137b |
Address the clang-tidy findings of the comparisons
Separate the comparison of discarded values from the other types, so that the conditional chain has no repeated branch bodies, and mark the deliberate comparisons of views with empty containers in the tests. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e842f3a68f |
Add comparisons to json_view
basic_json_view gains operator== and operator!= with other views and with basic_json values. Two views are equal if the values parse() would produce for them are equal by basic_json's operator==: numbers compare by value across their types, and objects by their members, with duplicate keys resolved as parse() resolves them (the last value, at the position of the first key). Objects are compared in member order if the object type keeps an order (ordered_json), by key otherwise, as basic_json does. Discarded views compare as discarded basic_json values do, which follows JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON. Nothing is materialized except single numbers, and the walk is iterative. Tests compare the results for pairs of 1,200 generated documents (also written differently: sorted keys, canonical numbers) with those of basic_json, for json and ordered_json, plus numbers, duplicate keys, member order, discarded values, and 100,000 levels of nesting. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
bfb2b0cb48 |
Mark the cases of the view's serializer that tests cannot reach
Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
8cccae029b |
Address the clang-tidy findings of dump()
The output buffer initializes its members in the initializer list, and the escaping has no nested conditional operators; the test marks a fixed seed. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
83d72fd7a7 |
Add dump() to json_view
basic_json_view::dump(indent, indent_char, ensure_ascii, number_format)
writes the text of a value as ordered_json::parse(text).dump() writes it
for the same arguments: members in document order (all of them, should a
key occur more than once), strings escaped by the same rules and with the
library's scanning kernels, floats with the library's conversion, and
integers copied from the source, where they are canonical except "-0".
With number_format::source, numbers are copied as they appear in the
source ("1.50", "1E2", "-0", all digits of long integers). operator<<
takes the indentation from the stream width, as for basic_json.
The writer (detail/view/serializer.hpp) writes through a raw pointer into
a string sized from the source extent of the value, and walks the index
iteratively, so the nesting depth is limited by memory only.
Tests compare the output of 2,000 generated documents with
ordered_json::dump() for several indentations and ensure_ascii, strings
with every kind of escape, numbers (5,000 random doubles, float as
number_float_t), duplicate keys, 100,000 levels of nesting, and streams.
ViewDump joins the benchmarks.
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
|
||
|
|
da971a52c8 |
Fix CI: useless cast in the array index check of the view's JSON pointers
GCC -Werror=useless-cast on Linux x86-64 rejected static_cast<std::uint64_t>((std::numeric_limits<std::size_t>::max)()), as both are the same type there. Compare without the cast: std::size_t converts to std::uint64_t implicitly on every platform. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
a7256deba3 |
Inline get() of arithmetic values of json_view
get<T>() of arithmetic types is inlined down to the conversion, so that its checks of the node kind merge with those of the caller, and reading an integer needs no call. Traversing every value: citm_catalog -6%, marine_ik -5%, numbers and twitter -3%, mesh -2.5%, canada -1% (and more above the float conversion from the digit layout: citm_catalog -14%, marine_ik -11%). Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
3a4eddf4ac |
Address the clang-tidy findings of values and JSON pointers
get_string() and number_token() return braced lists; the test compares floats by their bit patterns instead of with memcmp, uses std::any_of, and marks a fixed seed, a default member initializer (needed by GCC's -Weffc++), and a string search. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
0650a48659 |
Add values and JSON pointers to json_view
basic_json_view gains get<T>(), get_to(), value() with keys and JSON pointers, and operator[], at(), and contains() with JSON pointers, plus two functions basic_json has no counterpart for: - get_string(): the string without a copy (a string_view into the source, or into the decoded strings for strings with escapes) - number_token(): the text of a number as it appears in the source get<T>() converts arithmetic types, strings (also string_view_t), std::nullptr_t, std::vector, maps with string keys, and views directly; floats are converted from the digit layout recorded by the parser with the library's conversion chain, so the values are bit-identical to parse(). Other types, including user types with from_json(), go through materialize(). The exceptions are those of basic_json, message included. Where const basic_json has undefined behavior (a missing key or an index out of range with operator[] and a JSON pointer), the result is a discarded view; value() returns the default wherever basic_json catches out_of_range, and contains() never throws. Array indices of JSON pointers follow json_pointer's rules (parse_error.106/109, out_of_range.404/410). Tests compare the conversions of 2,000 generated documents, 20,000 float tokens (double and float, bit for bit), and every JSON pointer of 1,000 documents with basic_json, and the exceptions for malformed pointers. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
29bb5c48b8 |
Give code outside basic_json the reference tokens of a json_pointer
detail::json_pointer_access returns the reference tokens of a pointer, so that code resolving pointers without a basic_json value (such as the zero-copy view) does not have to parse to_string() again. No change in behavior. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
c2d7177d30 |
Address the clang-tidy findings of element access and iteration
Marks the default initializer of the item's index string (needed by GCC's -Weffc++) and, in the test, an escaped literal and a comparison of find() with end(), which is what the test is about. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
7dae258350 |
Add element access, lookup, and iteration to json_view
basic_json_view gains the read-only access functions of basic_json: operator[] and at() with keys and indices, front(), back(), find(), contains(), count(), begin()/end(), items() (with structured bindings from C++17 on), and type_name(). They throw the exceptions (ids and messages) that the const functions of basic_json throw; where basic_json has undefined behavior (operator[] with a missing key or an index out of range, front()/back() of an empty container), the view returns a discarded view or throws invalid_iterator.214. Objects are iterated in document order, and all members are visited. With duplicate keys, lookups find the first member, so that a lookup can stop at the first match; parse() keeps the last value. Keys of up to 16 bytes are compared with two overlapping loads instead of memcmp, and most keys are rejected by their length alone, from the index. The iterators and items live in detail/view/iterator.hpp, the lookups in detail/view/lookup.hpp. Tests compare every element and member of 2,000 generated documents with ordered_json, keys of every length around the load sizes, the exceptions against const basic_json, and the iterators. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
7da943c5c4 |
Name JSON types without a basic_json value
basic_json::type_name() now calls detail::value_type_name(value_t), so that code which reports types without a basic_json value at hand, such as the zero-copy view, uses the same names in its exception messages. No change in behavior. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
20d0723b67 |
Address the cpplint findings of json_view's materialize()
materialize.hpp includes <string> (build/include_what_you_use). Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
06bebe2af6 |
Mark the code of json_view that tests cannot reach
The 4 GiB limit and the fallback for an input that parse() accepts but the view rejects (a bug) are excluded from the coverage; shrink_to_fit() of an empty document is tested. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
723cf14be4 |
Address the clang-tidy findings of json_document and json_view
- the input dispatch takes byte ranges by const reference and reads the size once (which also settles a finding of the static analyzer); input adapters are taken by value - the classification of inputs keeps its nested conditional operators, a constant expression of C++11 (NOLINT) - the test's C arrays, fixed seed, and escaped literals are marked, as in the other tests Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
cf352d4ef5 |
Add json_document and json_view: parse, accept, types, materialize
The public classes of the zero-copy view (#5295), in the new header <nlohmann/json_view.hpp>: - basic_json_document<BasicJsonType>: parse (borrowing contiguous byte inputs, owning rvalue strings, streams, and other inputs), parse_copy, accept, read, root, is_discarded, source, owns_source, node_count, memory_usage, shrink_to_fit - basic_json_view<BasicJsonType>: type and the is_* queries, size, empty, materialize (the value parse() would produce, built by the same SAX handler), source_offset - the aliases json_document, json_view, ordered_json_document, and ordered_json_view A parse error throws the exception basic_json::parse would throw for the same input: the library parser is run on the failing input, so messages, positions, and exception ids are the same. Inputs of 4 GiB or more are rejected with out_of_range.416. The single header single_include/nlohmann/json_view.hpp keeps including json.hpp; make amalgamate, check-amalgamation, include.zip, and release handle it. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
3b6ae43c53 |
Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions (#5598)
* Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions basic_json can be constructed from std::tuple<json&>, which it turns into a one-element array. Because of this, std::tuple picks its converting constructor that converts the whole source tuple instead of the element-wise one. As a result, std::tuple<const json&> built from std::forward_as_tuple(j) binds to a temporary (a compile error with libc++, a dangling reference with other standard libraries), and std::tuple<json> built the same way holds [j] instead of a copy of j. The new opt-in macro JSON_DISABLE_TUPLE_REFERENCE_CONVERSION (CMake option JSON_DisableTupleReferenceConversion) removes the conversion from a one-element tuple holding a reference to the same basic_json type, so std::tuple converts element-wise. It is off by default, so existing behavior is unchanged. It does not change any function body and therefore is not part of the ABI tag. Fixes #2226 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Convert one-element tuples to arrays on every compiler to_json for std::tuple assigns a braced list, j = { std::get<Idx>(t)... }. With a single element that is itself a basic_json, Apple clang 15 and 16 treat j = {x} as a copy of x, so std::tuple<json>{true} became true instead of [true]. The macOS jobs (Xcode 15.1, 16.1) failed the new checks in unit-disable-tuple-reference-conversion and unit-regression2. The one-element overload that already handles JSON_BRACE_INIT_COPY_SEMANTICS builds the array (or object, for a [string, value] element) explicitly, the same way the initializer-list constructor does. Use it unconditionally. The output is unchanged on compilers that already wrapped the element. Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Skip json reference tuple tests on clang < 4 and GCC < 5 ci_test_compilers_gcc_old (4.8) and ci_test_compilers_clang (3.4) could not compile the new tuple tests. Creating a std::tuple of basic_json references, e.g. std::forward_as_tuple(j), makes these compilers instantiate basic_json's conversion operator for libstdc++'s internal tuple bases, which fails hard. This happens with and without JSON_DISABLE_TUPLE_REFERENCE_CONVERSION, so it is a limitation of these compilers, not of the new option. Tested with the CI images: clang 3.4 to 3.9 and GCC 4.8 and 4.9 fail, clang 4, 5, and 6 and GCC 5 and 6 compile all cases. Skip only the checks that create such tuples; the is_constructible checks still run. Signed-off-by: Niels Lohmann <mail@nlohmann.me> --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e96e2982a5 |
Fix CI: useless cast in the growth of the view's node index
GCC -Werror=useless-cast (ci_test_gcc on Linux x86-64) rejected static_cast<std::size_t>(guess + (guess / 4) + 64): the sum is a std::uint64_t prvalue, the same type as std::size_t there, while the cast is needed where std::size_t is 32 bits wide. Cast a named variable instead, which GCC does not report. The build stopped at an earlier error before, so the previous CI run did not show this one. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
0a64e6c99b |
Fix CI: MSVC C4127 in the view builder and the single-header test build
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is constant) for `TrailingCommas && cur() == ']'` and the like when the option is off. Route the template arguments through a static enabled() function, as json.hpp's nesting_depth_exhausted() does. - ci_test_single_header compiled unit-json_view_builder.cpp against single_include/, which does not contain the internal nlohmann/detail/view headers. Build that test only with the multiple headers. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
2fb66ba859 |
Remove the eight-digit case of the view's digit parser
Since whole blocks of eight digits are read directly, parse_upto8() only gets fewer than eight digits; its eight-digit case was dead code. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
ab49a7b1bf |
Address the cpplint findings of the view's parser
The exponent of the overflow check is an std::int64_t instead of a long (runtime/int). Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
154f240022 |
Read whole blocks of eight digits of json_view directly
A block of eight digits of a number token lies inside the input (the digits were counted while scanning, or are recorded in the digit layout), so parse_upto19() reads it without the bounds check of the last, partial block. Traversing canada.json: -11% instructions, -6% cycles. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
8d3280de1e |
Classify a NUL inside a string of json_view as a control character
json::parse ends the input at a NUL only between values (where it does at all); inside a string, a NUL is a control character that must be escaped. The view reported it as a missing closing quote. (Only the error code differed: the exception comes from the library parser.) Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
9b0b4ff26d |
Address the clang-tidy findings of the view's parser
- tables as std::array; the frames of the first 64 levels stay a C array (not initialized on purpose, NOLINT) - \u escapes are decoded with the library's hex_codepoint() instead of a second table - the parse failure is private, with an accessor; the special member functions of the builder are all declared - no nested conditional operators; explicit parentheses; a repeated branch body merged; auto for casts - the test's C arrays, fixed seed, and escaped literals are marked, as in the other tests Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
694bfd1b4e |
Keep the arguments of the view's throw helpers used without exceptions
With JSON_NOEXCEPTION, NLOHMANN_VIEW_THROW(e) was std::abort() alone, so the parameters of the functions that build the exceptions were unused, a warning that the builds with -Werror turn into an error. The exception is now evaluated before std::abort(); the program ends anyway. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
ad82821c89 |
Add the one-pass parser of json_view (internal)
The builder parses JSON text in one pass into the node index: strings and numbers stay in the source (escaped strings are decoded into an arena), integers are converted while their digits are in cache, and floats keep their digit layout for a later conversion. It accepts exactly what json::parse accepts, for every combination of comments and trailing commas, with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING. Parse state lives in a local cursor whose address never escapes, so that it stays in registers; out-of-line helpers (errors, regrowth, escapes, comments) are members of the builder and get the positions they need. The value dispatch is expanded once for array elements and once for member values. Literals are compared with memcmp and words read in a fixed byte order, so nothing depends on the platform's byte order. Error messages come with the public classes. Tests (unit-json_view_builder.cpp): accept/reject and values against json::parse for handwritten, generated, and damaged documents under all option combinations, from std::string and from exact-size buffers (no read past the input under AddressSanitizer), deep nesting up to 100,000 levels, NUL/BOM/whitespace cases, and the test-suite files. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
1bf0b1b6c2 |
Add the node index and scanning primitives of json_view (internal)
Internal parts of the zero-copy view (#5295), under detail/view and not included by json.hpp, so that users of json.hpp compile nothing of it: - macro_scope.hpp/macro_unscope.hpp: the few macros the view needs, under its own prefix (json.hpp undefines its own at its end); the throw macro honors JSON_NOEXCEPTION and JSON_THROW_USER like JSON_THROW - string_ref.hpp: std::string_view from C++17 on, else a small stand-in - node.hpp: the 16-byte node of the index; its kinds are value_t values (checked by a static_assert) - document_data.hpp: the storage of a parsed document (node array, decode arena, owned input) - scan.hpp: string and digit scanning with unrolled checks at fixed offsets (after yyjson) and the library's SWAR and UTF-8 checks, independent of the byte order Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
b88e5f9107 |
Keep the behavior-changing configuration readable after json.hpp
json.hpp undefines JSON_STRICT_NUL_HANDLING and JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on the library after it, such as the planned json_view.hpp, reads them from detail::abi_config instead. The constants live in the ABI namespace, which already encodes both settings, so they always match the basic_json in use. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
d23803fd32 |
Decode \u escapes with a table in the lexer
get_codepoint() read the four hex digits of a \u escape with four calls to get(), each classified by a chain of range comparisons. For contiguous input, get_codepoint_bulk() now decodes them with one lookup per byte (hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a 256-entry table maps a byte to its value, or 0xFF for anything else, and an invalid digit shows in the OR of the four values. It then skips the four bytes and updates the position counters as four get() calls would. If a digit is invalid or fewer than four bytes are left, it changes nothing and the existing loop runs, so errors are reported with the same message and position as before. json::parse, best of 5 runs in separate processes (M1 Max): the escaped twitter.json (every non-ASCII character as \u) -13.6%, all other files within 0.3%. Tests compare the contiguous and the streaming path (value or exception message) for valid escapes, surrogate pairs, truncated and invalid digits at every position, and 3,000 seeded random escapes. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
e91fdad877 |
Find the stop byte of a string run without a byte loop
find_string_special() and find_ascii_copyable_run() test eight bytes at a time, but located the stopping byte inside a word with a byte loop. The lowest flagged byte of the SWAR tests is always a true hit (the borrows of the subtractions can only flag bytes above one), so its index is now the trailing-zero count of the mask; words are read in little-endian order on every platform, so this does not depend on the byte order. scalar_string_bulk_run() validates a run of multi-byte UTF-8 sequences one after another instead of searching for the next special byte in between, which helps text in non-Latin scripts. The kernels serve the lexer's contiguous fast path, the serializer, and the binary formats. New tests compare all three with byte-by-byte reference scans on 100,000 generated buffers at three alignments; the portable fallback of count_trailing_zeros() was checked against the builtin. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
9c71689715 |
Convert long doubles under a multi-byte decimal point completely
The strtold fallback, which is left only for long double formats that are not binary64 (x87, binary128), substituted the first byte of the locale's decimal point for '.'. Under a locale whose decimal point is longer than one byte, such as fa_IR.UTF-8 or ar_EG.UTF-8 (U+066B), strtold stopped there and the value was truncated at the decimal point. A longer decimal point is now put into a copy of the token. The test "locale with a multi-byte decimal point" now compares the long double values with those of the "C" locale; with x87 long doubles it failed before. Fixes #5660. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
44ec53c77b |
Convert float and double with the library's own correctly rounded parser
float, double, and long double where it is IEEE-754 binary64 (MSVC, Apple arm64) are now converted by the library itself, correctly rounded and independent of the locale and of the C and C++ libraries: - The token is split into sign, significand w (at most 19 digits), and decimal exponent q, using the positions of the decimal point and the exponent that the scanners already recorded, so no character is classified again. - Clinger's fast path where w and 10^|q| are exact. - Eisel-Lemire otherwise, now templated for binary32 and binary64. - For tokens with more than 19 digits whose w and w + 1 round differently, an exact big-integer comparison with the midpoint between the two candidates (the digit comparison of fast_float, simplified). This replaces the separate token walks of Clinger's fast path and of Eisel-Lemire, the significant-digit gate that avoided the former, and, for float and double, std::from_chars and the locale-aware strtod. std::from_chars and strtold remain only for other long double formats (x87, binary128, double-double) and for types that are not IEEE-754. Values are bit-identical to before wherever the previous conversion was correctly rounded; tokens converted in a locale with a multi-byte decimal point are now also exact. Overflow still gives out_of_range.406, underflow a signed zero. convert_float() is the entry point for other parsers of JSON text: it converts like the lexer, without allocation for binary32/binary64. Tests: exact-bit tests for double and float (ties, subnormal and overflow boundaries, huge exponents, more digits than any midpoint), Eisel-Lemire for binary32, the round trips of 200,000 doubles and 100,000 floats without declines, 508 generated hard cases with the expected bits of both formats (float_hard_cases.hpp) through the converter and both scanners, and JSON-level overflow/underflow checks for double and float. The locale tests now check the values in a locale with a multi-byte decimal point. Docs: the statements that parsing uses strtod/strtof/strtold; the fast_float credit now names the digit comparison. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |
||
|
|
6a073dbae4 |
Give operator>> a strong exception-safety guarantee (#5695)
operator>> parsed directly into its basic_json& target, so a parse error left the target holding whatever was parsed before the error instead of its previous value. With JSON_DIAGNOSTICS=1, that partial value also violated the class invariant, because the parent pointers of an array or object's elements are only set when the container is closed, which a failed parse never reaches; copying such a value then aborted in assert_invariant(). Fix it the way basic_json::parse() already handles this: parse into a temporary and move it into the target only once parsing succeeds, so the target is left unchanged if an exception is thrown. Fixes #5652. Signed-off-by: Niels Lohmann <mail@nlohmann.me> |