Compare commits

...
Author SHA1 Message Date
Niels Lohmann 7e8e8e219b Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/input/binary_reader.hpp
#	include/nlohmann/json.hpp
#	single_include/nlohmann/json.hpp
2026-09-30 22:53:03 +02:00
Niels Lohmann c261578431 Deduplicate binary reader/writer helpers and fix stale comments (#5730)
* Fix stale and missing comments in binary_writer

The doc block of write_number() ended up above the byte_swap() helpers
added in #5286, about 80 lines from the function. It was also a plain
comment that Doxygen skips, said "write a number to output input", and
left BON8 out of the big-endian formats. Move it back onto
write_number() as a /*! block and fix the text.

write_bson() documented "@pre j.type() == value_t::object", but it
throws type_error.317 for every other type, and to_bson() relies on
that. Document the exception instead.

Explain why the CBOR binary subtype is always written with a 0xD8..0xDB
head and never in the one-byte tag form: binary_reader with
cbor_tag_handler_t::store only keeps those heads as a subtype, so
switching to write_cbor_head() would break round trips for subtypes
0..23.

Also fix the grammar of the to_char_type comment. Comments only; no
change in behavior, API or ABI.

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Merge the duplicated UBJSON/BJData integer marker ladders

write_number_with_ubjson_prefix() (unsigned and signed overloads) and
ubjson_prefix() (number_integer and number_unsigned cases) each picked
the UBJSON/BJData integer marker (i, U, I, u, l, m, L, M, H) with their
own independent if/else ladder, and the values beyond 64 bits were
handled by a second, tag-dispatched pair of ladders. An optimized
container announces the marker of its first element via ubjson_prefix()
and then writes every element through write_number_with_ubjson_prefix(),
so the two had to be kept in lockstep by hand across four call sites.

Replace all of that with one ubjson_integer_prefix() built on
value_in_range_of<T>, and one write_ubjson_integer_payload() that
writes the value (or, for 'H', the decimal digits) for a given marker.
write_number_with_ubjson_prefix() and ubjson_prefix() keep their
signatures and now just call these two helpers.

Behavior, the public API and the ABI are unchanged. Verified with a
new regression test covering scalars and $-optimized arrays/objects at
every int8/uint8/int16/uint16/int32/uint32/int64/uint64 boundary for
to_ubjson/to_bjdata (both use_size/use_type settings), and by diffing
to_ubjson/to_bjdata output before and after over the json_test_data
corpus (bit-identical).

Part of #5710

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove dead get_char parameters in binary_reader

The non-recursive rewrite of the binary readers (#5505, #5506, #5507)
left parse_cbor_internal()'s and parse_ubjson_internal()'s get_char
parameters dead: parse_cbor_internal() has one caller and it always
passes true, and parse_ubjson_internal() has one caller and it always
uses the true default. Both parameters, and the @param docs describing
the "reuse the last character" mode they used to select, no longer
correspond to anything.

Drop both parameters, initialise fetch/prefix unconditionally, and
update the two call sites in sax_parse(). parse_cbor_value()'s and
get_ubjson_string()'s own get_char parameters are unrelated and are
left alone; both still have a false caller.

Also delete a stray `@return whether a valid MessagePack value was
passed to the SAX parser` doxygen block that sits directly above
parse_msgpack_value()'s real doc comment, a leftover of the same
rewrite.

Behavior, the public API and the ABI are unchanged; these are private
members of detail::binary_reader. Verified by compiling with
-Wunused-parameter and running unit-cbor, unit-ubjson, unit-bjdata and
unit-msgpack (offline, against the stubbed test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Share the IEEE half-precision decoder between CBOR and BJData

binary_reader had two ~45-line copies of the IEEE 754 half-precision
decoder: CBOR's case 0xF9 and BJData's case 'h'. Once formatting is
normalised, the two blocks were identical except for the byte order
used to assemble the 16-bit half (CBOR is big endian, BJData is little
endian). Any future change to half-float decoding had to be made and
kept in sync in both places.

Add one get_half_float(format, little_endian) helper that does the two
get()/unexpect_eof() reads, assembles the half in the requested byte
order, decodes it per RFC 8949 Appendix D, and calls sax->number_float.
Both cases now just call it with their byte order; the BJData case
keeps its bjdata-only guard.

Behavior, the public API and the ABI are unchanged. Verified with a
scratch probe comparing the old and new decoders bit-for-bit (NaN by
isnan()) over all 65536 wire byte pairs, in both formats, and by
running unit-cbor and unit-bjdata (offline, against the stubbed
test_data.hpp).

Part of #5711

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the MessagePack unsigned-integer writer ladder

The number_integer (non-negative branch) and number_unsigned cases in
write_msgpack() each held their own copy of the fixint/uint8/16/32/64
ladder, kept in lockstep only by a comment ("we used the code from the
value_t::number_unsigned case here"). Both copies mixed union members:
the signed copy compared number_unsigned but wrote number_integer, and
vice versa.

Extract write_msgpack_unsigned(std::uint64_t), mirroring how
write_cbor_head() already avoids the same duplication for CBOR, and
call it from both cases. Each case now reads only its own active
union member. Output bytes are unchanged for the default 64-bit
number types.

#5710 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Unify float marker selection and fix the long double compile error

Four formats picked between a float32 and float64 marker through four
different helper styles: dummy-argument overloads for CBOR and
MessagePack, an std::is_same template for BON8, and a runtime if-chain
on input_format_t for write_compact_float(). With number_float_t set
to long double, to_cbor, to_msgpack and to_ubjson failed inside the
library with "call to 'get_cbor_float_prefix' is ambiguous", while
to_bson kept working because write_bson_double() takes a plain double.

Change write_compact_float() to take the two marker bytes directly
(each of its three callers already knows them at compile time) instead
of an input_format_t it only forwarded, and delete the now-unused
get_cbor_float_prefix(), get_msgpack_float_prefix(),
get_bon8_float_prefix() and get_compact_float_prefix() helpers. Turn
the two get_ubjson_float_prefix() overloads into one template. Both
write_compact_float() and get_ubjson_float_prefix() now report an
unsupported number_float_t with a static_assert naming the requirement,
rather than an ambiguous-overload error; the assert lives in the
function body, not the class scope, so to_bson with long double is
unaffected.

Verified with a probe basic_json<..., long double>: to_bson still
compiles and round-trips, while to_cbor/to_msgpack/to_ubjson now fail
to compile with the new static_assert message.

This changes the text of an existing compile error for users with an
unsupported number_float_t (documented as a public-API-visible change
in #5710).

#5710 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the BJData ndarray writer's dtype dispatch and drop <map>

write_bjdata_ndarray() built a 12-entry std::map<string_t, CharType> on
every call just to translate the _ArrayType_ name to a dtype marker
(the only reason binary_writer.hpp included <map>), then mapped dtype
to C++ type twice more: once as a switch for the range-check pass and
once as a separate if/else chain for the write pass, with nothing
checking that the two agreed. The caller also ran three at() lookups,
and the callee called value.at(key) about ten more times for the same
three members.

Replace the map with bjdata_ndarray_type_marker(), a plain string
comparison chain (a C++11 constexpr function cannot contain a switch,
so this mirrors binary_reader's own static table style). Replace the
switch/if-chain pair with one write_bjdata_ndarray_elements() that
switches on dtype once and calls a per-type helper -
write_bjdata_ndarray_element<T>() for the eight integer dtypes and
write_bjdata_ndarray_float_element() for 'd' - with a dry_run flag
selecting the range check or the actual write, so the two passes can
no longer disagree on the type. _ArrayType_, _ArraySize_ and
_ArrayData_ are now looked up once into references, and the four
header marker bytes ('[', '$', '#') are written through to_char_type()
like the rest of the UBJSON/BJData writer.

The 'd' (single-precision) rule is left exactly as before, since #5707
is expected to change it separately.

Verified byte-for-byte identical output before/after for every dtype
(including the Draft 2/Draft 3 'byte' fallback and the use_count/
use_type combinations) via a standalone probe, plus round-tripping
through from_bjdata().

Overlaps #5707, which is expected to touch the 'd' dtype case, and
#5518, which is expected to move the write_bjdata_ndarray() call site.

#5710 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Assert that write_bson_document() consumes every calc_bson_sizes() entry

calc_bson_sizes() and write_bson_document() are a hand-synchronized
pair of passes over the same object/array tree, introduced by #5553:
the size pass appends to nested_sizes in visiting order, and the write
pass consumes the table by position with nested_sizes[next_size++].
Nothing checked that the write pass consumed the whole table. If a
future change touched only one of the two passes - for example to skip
or reject an entry - every later size prefix in the document would be
silently wrong.

Add JSON_ASSERT(next_size == nested_sizes.size()) where
write_bson_document() returns, so such a future drift between the two
passes is caught immediately (JSON_ASSERT expands to nothing in
release builds using assert(), and the fuzzers/tests already build
with it enabled). The two passes agree today, so this changes nothing
observable; it only guards against the risk described in #5710 item 5.

Extracting a shared stepper for the two passes (the second half of the
proposed change) is left for a follow-up: it only saves ~30 lines and
the issue asks for it only if the result reads clearly, which needs
more room to get right than a mechanical cleanup pass allows.

#5710 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make the BJData lookup tables static functions instead of members

binary_reader held bjd_optimized_type_markers and bjd_types_map as
non-static const members (12 string_t objects for the type-name table),
built and destroyed on every from_cbor/from_msgpack/from_bson/
from_ubjson/from_bon8/from_bjdata call even though only from_bjdata
ever reads them. They also needed the #define/decltype/#undef
workaround from #3637 and two NOLINTNEXTLINE suppressions, and
binary_writer already carries the same two lists in another form
(is_bjdata_excluded_type_marker() and a local std::map in
write_bjdata_ndarray(), the latter removed by the item-4 commit), so
the excluded-marker lists could drift apart.

Replace bjd_optimized_type_markers with static constexpr
is_bjd_excluded_optimized_type(char_int_type), using the same ||-chain
as binary_writer's is_bjdata_excluded_type_marker(). Replace
bjd_types_map with a non-constexpr static bjd_type_name(char_int_type)
switch returning nullptr for an unknown marker (a C++11 constexpr
function cannot contain a switch). Delete both
JSON_BINARY_READER_MAKE_* macros, the bjd_type pair alias, the
NOLINTNEXTLINE suppressions, detail::make_array() (no longer used
anywhere), and the now-unused <algorithm> and <array> includes.

Update the two call sites (the ND-array excluded-type check and the
_ArrayType_ lookup) accordingly, and replace unit-bjdata.cpp's
"LUT arrays are sorted" section, which only checked the two tables'
internal ordering, with a check of all 12 type names and all 8
excluded markers against both new functions.

#5711 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Read CBOR's 1/2/4/8-byte argument through one helper

parse_cbor_internal() hand-wrote the same "read a 1/2/4/8-byte
big-endian unsigned integer" ladder four times over:
- twice for tag numbers 0xD8-0xDB, once in the tag_handler::ignore
  branch and once, nearly identically, in the ::store branch (~90
  lines to read one integer);
- twice more for container lengths, once for array heads 0x98-0x9B and
  once for map heads 0xB8-0xBB, where the 1/2-byte forms called
  enter_array()/enter_object() directly and the 4/8-byte forms
  additionally went through get_cbor_container_size().

Add get_cbor_argument(std::uint64_t&), reading the width selected by
current & 0x1F via the same get_number() calls as before (so EOF is
reported exactly as before), and route all four sites through it:
- 0xD8-0xDB now read the argument once per branch instead of switching
  on `current` a second time; behavior split cleanly from embedded tags
  0xC0-0xD7 (tag value in the head, no argument to read), which is now
  its own case block that no longer has to fall into the ::store
  switch's "default" case to reach the same tag_pending = true; return
  true; outcome.
- 0x98-0x9B and 0xB8-0xBB collapse into one case block each, always
  going through get_cbor_container_size() (harmless for 1/2-byte
  lengths, which already always fit).

Verified byte-for-byte identical behavior before/after with a
standalone probe covering embedded and multi-byte tags under all three
tag_handler_t settings, a tag over a byte string (subtype path),
truncated tag/length arguments of every width, and array/map lengths
of every width, including the out_of_range.408 "excessive size" case:
same exceptions, same messages, same chars_read, same successful
results.

Left the string/byte-string length ladders in get_cbor_string()/
get_cbor_binary() untouched, as noted in #5711 item 2, since #5325 is
expected to touch them separately.

Overlaps #5601 (adds a branch right above the embedded-tag case) and
#5607 (touches the integer cases 0x18-0x1B, which share this ladder's
shape in separate hunks).

#5711 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Add leave_container() to match enter_container()

Every container is opened through enter_container(), whose docs
promise that a check placed there runs before every start event. The
close side had no equivalent: the same
"container_stack.pop_back(); dispatch to end_object() or end_array()"
sequence was written out separately in BSON, CBOR, MessagePack,
UBJSON/BJData and BON8, each copying the pattern of keeping an
is_object flag around the pop_back() that would otherwise invalidate
a reference to it. A check needed on close would have had to be added
in five places, and a sixth copy could go unnoticed.

Add leave_container() next to enter_container(), doing the same
pop-then-dispatch, and replace the five sites with it. Each site keeps
its own surrounding logic (BSON's check_bson_document_size() call
before popping, MessagePack's is_object copy used again below,
UBJSON/BJData's remaining-container handling after popping, BON8's
top used again below); only the repeated pop/dispatch line pair is
now shared.

Verified all six binary-format unit suites and unit-regression2's
deep-nesting tests (dependent count/reuse count and the bjdata ndarray
depth cases) still pass, compiled with -Wall -Wextra and ASan/UBSan.

Overlaps #5601, which is expected to add a sixth close site in its own
skip loop; that site can route through leave_container() too once it
lands.

#5711 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop passing the input format to sax_parse() when the reader already has it

binary_reader's constructor stores the format in the input_format
member, and sax_parse(format, sax_, strict, tag_handler) took the same
value again purely to dispatch on it. Every in-tree caller passed the
same value both times (all 16 from_cbor/from_msgpack/from_ubjson/
from_bjdata/from_bon8/from_bson call sites in json.hpp, and the three
public basic_json::sax_parse() overloads), so nothing was broken
today, but a caller of the detail class directly (only reachable via
JSON_PRIVATE_UNLESS_TESTED, as unit-bjdata.cpp already does) could
pass a mismatched pair - say bjdata to the constructor and ubjson to
sax_parse - and dispatch on one format while applying the other
format's rules; the default-constructed input_format_t::json reader
would additionally hit JSON_ASSERT(false) in exception_message() on
its first error.

Add sax_parse(json_sax_t*, bool, cbor_tag_handler_t) forwarding to the
existing overload with the stored input_format, and switch every
caller to it: the 16 from_*() sites (keeping their
`// cppcheck-suppress[accessMoved]` comments) and the three
basic_json::sax_parse() overloads, all of which already had the format
available from their own `format` parameter. The four-argument overload
is kept for anyone still calling it, now with
JSON_ASSERT(format == input_format) so a mismatch fails immediately
in a debug build (assert-enabled binaries, including the fuzzers and
test suite) instead of misbehaving; verified with a probe that
constructs a reader for one format and calls the explicit overload
with another, which aborts on that assertion as expected.

Removing or asserting against the constructor's input_format_t::json
default, which would affect direct detail users, is left as a separate
decision per #5711 item 5.

Overlaps #5601, which is expected to add an AllowRecovery template
parameter to sax_parse() and touch these same call sites in json.hpp.

#5711 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate UBJSON/BJData signed-count handling, drop dead ndarray checks

get_ubjson_size_value()'s 'i'/'I'/'l'/'L' cases each read a differently
sized signed integer and then repeated the same "reject negative with
error 113" check; only 'L' additionally checked value_in_range_of for
the out_of_range.408 case. Any change to that error path had to be
made four times.

Add get_ubjson_signed_count<SignedType>(std::size_t&), doing the read,
the negative check and the range check once, and route all four
markers through it. The range check is a no-op for 'i'/'I'/'l' (their
values always fit std::size_t) and only live for 'L' on a 32-bit
std::size_t target, matching today's behavior exactly.

In the ndarray dimension-product loop, the preceding loop already
returns early on any zero dimension and result starts at 1, so `i > 0`
in the pre-multiplication overflow check was always true, and
`result == 0` in the post-multiplication check could not be reached
either: two positive factors whose product does not overflow (as the
pre-check already guarantees) cannot be zero. Drop the dead `i > 0 &&`
and narrow the post-check to `result == npos`, the one case the
pre-check cannot rule out (an exact, non-overflowing match with the
sentinel reserved for unknown-size containers), with a comment
explaining why.

Verified byte-for-byte identical behavior before/after with a
standalone probe covering negative counts for every marker, a matching
positive count, and ndarray inputs, plus the full unit-ubjson and
unit-bjdata suites (same assertion counts as before this change).

Overlaps #5601 (rewrites the four parse_error calls and the overflow
checks touched here) and #5607/#5707 (touch neighboring lines in the
same functions).

#5711 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Drop redundant format parameter and dummy float argument (review)

binary_reader::sax_parse(format, ...) only ever had to equal the format
given to the constructor, which it asserted. With every caller already
on the format-less overload, remove the four-argument overload and
dispatch on the stored input_format directly. binary_reader is a
detail class, so this is not a public API change.

get_ubjson_float_prefix() took a value only to deduce its type; make
the type an explicit template argument instead.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:48:32 +02:00
Niels Lohmann 35802e78d6 Remove dead Makefile targets; document macro_builder; tidy serve_header (#5735)
* Remove dead doctest help entry and pretty_format target from Makefile

The top-level Makefile still carried three leftovers:

- The help text listed a "doctest" target that was removed in #4560,
  so "make doctest" fails with "No rule to make target". The example
  check now runs as "make check_output -C docs".
- "pretty_format" ran clang-format on all sources, but .clang-format
  was deleted in #4573, so the target reformatted everything in the
  default LLVM style, against the Artistic Style formatting that
  "make pretty" applies and CI enforces.
- "clean" removed benchmarks/files/numbers/*.json, a directory that no
  longer exists since the benchmarks moved to tests/benchmarks (#3462).

Only maintainer tooling changes; the library is not affected.

Part of #5717

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Document tools/macro_builder and tidy up serve_header.py

tools/macro_builder generates the NLOHMANN_JSON_EXPAND,
NLOHMANN_JSON_GET_MACRO and NLOHMANN_JSON_PASTE* macros in
macro_scope.hpp, but nothing referred to it. Add a README that explains
what it generates, how to run it and where the output goes, and which
dependent tables (NLOHMANN_JSON_DOUBLE_PASTE, NLOHMANN_JSON_TYPE_BODY)
are maintained by hand. Point to it from a comment above
NLOHMANN_JSON_EXPAND. The generator itself is unchanged; following the
README reproduces the header byte for byte.

In serve_header.py, drop the LGTM suppression (LGTM.com shut down in
2022), replace the """.""" placeholder docstrings with real ones, and
import socket and ssl at module level. DualStackServer.server_bind uses
socket, which was only imported under __main__; when the module was
imported instead, the NameError was swallowed and IPV6_V6ONLY was not
cleared.

The header change is a comment only; behavior, API and ABI are
unchanged.

Part of #5717

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Hash every release_files artifact, not a hardcoded subset

The `release` target signed and copied json_fwd.hpp into release_files
alongside json.hpp, but the shasum line that writes hashes.txt only
listed json.hpp, include.zip and json.tar.xz. Users could not verify
the published json_fwd.hpp against hashes.txt.

Hash every file in release_files except the .asc signatures instead
of naming files by hand, so a newly shipped header (such as the
json_literals.hpp that #5610 adds to this target) cannot be missed
again.

Only affects the generated hashes.txt release artifact; the library
itself is unaffected.

Overlaps #5610, which touches the same lines to add json_literals.hpp
to the release target.

#5717 item 1

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove the broken fuzz_testing* Makefile targets

fuzz_testing and fuzz_testing_{bon8,bson,cbor,msgpack,ubjson} seeded
fuzz-testing/testcases from tests/data, which was removed in dbf1a1f41
(2020) when the test data moved to the external json_test_data repo.
The find command found nothing, but the pipeline's exit status was
that of xargs, so the recipe still reported success with an empty
corpus, and the printed afl-fuzz command would refuse to start.

The recipes were also six near-identical copies with unquoted -name
patterns, used the legacy CXX=afl-clang++, and fuzzing-start/stop were
missing from both the help output and .PHONY. tests/fuzzing.md already
documents the working flow (download json_test_data, then
`make -C tests fuzzers`), so replace the six broken targets and their
help lines with a single pointer to that document instead of trying
to keep six copies of a fragile shell pipeline in sync.

This does not affect OSS-Fuzz, which builds through tests/Makefile.

Overlaps #5621, which adds a seventh copy of the same broken line for
fuzz_testing_json_view.

#5717 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove stale Travis comment above check-amalgamation

check-amalgamation carried "Note: this target is called by Travis",
left over from before the project switched off Travis CI. The prior
Makefile cleanup commit removed the other stale Travis-era leftovers
(the doctest help entry, pretty_format, and the benchmarks/ path in
clean) but missed this comment.

#5717 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale install/usage instructions in the vendored amalgamate README

tools/amalgamate/README.md is the unmodified upstream text and no
longer matches how the tool is used here:

- It named a Bitbucket origin that no longer exists; CHANGES.md
  already tracks the GitHub mirror commit this copy is based on.
- It asked for Python 2.7, but CI and the Makefile run the script
  with python3.
- It told readers to run ./test.sh (not vendored) and install to
  /usr/local/bin; in this repository the tool runs through
  `make amalgamate`.
- Its usage synopsis showed `-v` taking no argument, but the script's
  own argparser requires `choices=["yes", "no"]`, so that form fails
  with "argument -v/--verbose: expected one argument". The Makefile
  calls it as `--verbose=yes`.
- It pointed at test/source.c.json and test/include.h.json, which are
  not vendored; the configs actually used are config_json.json and
  config_json_fwd.json.

Rewrote only the Installing and Using sections to match; left the
"Here be dragons" caveats and the rest of the vendored code untouched
to avoid diverging further from upstream.

Overlaps #5615, which edits amalgamate.py, this README and CHANGES.md.

#5717 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Derive generate_natvis.py's ABI tag list and version from abi_macros.hpp

generate_natvis.py hard-coded abi_tags = ['_diag', '_ldvcmp', '_dp',
'_bics', '_psp', '_snul'] and required --version on the command line.
The source of truth is include/nlohmann/detail/abi_macros.hpp: the
NLOHMANN_JSON_ABI_TAG_* defines, the argument order of
NLOHMANN_JSON_ABI_TAGS_CONCAT, and NLOHMANN_JSON_VERSION_MAJOR/MINOR/
PATCH. Nothing checked that the copies stayed in sync, and they have
drifted apart before: _dp was added in #4517 but missed here until
#5544, and 3.11.3 shipped json_abi_v3_11_2 namespaces (#4340).

Parse the tag list (in NLOHMANN_JSON_ABI_TAGS_CONCAT order) and the
version from abi_macros.hpp instead of hard-coding them. Make
--version optional (falling back to the parsed version) and default
the output directory to the repository root the script lives in.

Add a "natvis" Makefile target that runs the script, and extend
check-amalgamation to regenerate nlohmann_json.natvis and fail on a
diff, the same way it already does for the amalgamated headers and
BUILD.bazel. Wire the same regeneration into check_amalgamation.yml,
using the tool copy checked out from develop (as the workflow already
does for amalgamate.py) and installing jinja2 from
tools/generate_natvis/requirements.txt. Update the tool's README to
say it must be re-run after adding an ABI tag or bumping the version.

Verified: a run against develop produces no diff (with either the
default or an explicit --version 3.12.0); adding a dummy
NLOHMANN_JSON_ABI_TAG_* without a matching #define makes the script
fail loudly instead of silently omitting the tag; xmllint --noout
passes on the regenerated file; and running the script from a
directory other than the one being checked (simulating the workflow's
separate tool checkout) against this repository root also produces no
diff.

Overlaps #5600, which added _ekmo to the same hand-written abi_tags
line and regenerated the file.

#5717 item 3

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Strip the leading "./" find(1) prefix from release hashes.txt entries

bc6e7db72 (#5717 item 1) switched the release target's shasum line from
naming files by hand to $$(find . -type f -not -name '*.asc' | sort),
so a newly shipped header is hashed automatically. Run from inside
release_files, that find prints paths as "./json.hpp" instead of
"json.hpp", so hashes.txt lists "./json.hpp" etc. instead of the plain
filenames it always used. shasum -c still verifies "./json.hpp" fine,
but it is a needless cosmetic regression for anyone reading the file
or matching it against release notes.

Strip the "./" prefix with sed before sorting, keeping the filenames
exactly as before while still hashing every artifact automatically.

Review fix for #5717 item 1 (PR #5735).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix check_amalgamation.yml: pass --version to generate_natvis.py

bff45f111 (#5717 item 3) made --version optional in
tools/generate_natvis/generate_natvis.py and wired the workflow's new
"Regenerate nlohmann_json.natvis" step to call it without --version,
relying on the script deriving the version from abi_macros.hpp itself.

But NATVIS_TOOL_DIR is checked out from develop, the same way TOOL_DIR
already is for amalgamate.py, precisely so an in-flight PR's tooling
changes cannot mark themselves clean. Until this PR (or an equivalent)
merges to develop, that checkout is the old generate_natvis.py, whose
--version argument is still required=True. The new step's invocation
of "generate_natvis.py $MAIN_DIR" (no --version) then fails argparse
on this PR's own CI run with "the following arguments are required:
--version", before the check ever gets to compare output.

Extract the version from $MAIN_DIR's own abi_macros.hpp in the
workflow and always pass it as --version. That satisfies the old
script's required argument and is accepted as an explicit override by
the new one, so the step behaves the same whether NATVIS_TOOL_DIR holds
the pre- or post-merge tool, and stays correct for later PRs that bump
the version.

Verified by running the workflow step's shell logic locally against
both the pre-#5717 generate_natvis.py (checked out at 633de8e44) and
the new one: both produce the identical nlohmann_json.natvis as the
committed file.

Review fix for #5717 item 3 (PR #5735).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Extend tools/macro_builder to also generate DOUBLE_PASTE and TYPE_BODY

tools/macro_builder only emitted NLOHMANN_JSON_EXPAND..PASTE64, so two
other tables that scale with the same max_args stayed hand-maintained
with nothing checking them: NLOHMANN_JSON_DOUBLE_PASTE (added by hand
in #4563 for the *_WITH_NAMES macros) and the 64-slot dispatch table of
NLOHMANN_JSON_TYPE_BODY (the #4041 zero-member/one-or-more-member
switch). Both tables pass one macro name per slot to the same
NLOHMANN_JSON_GET_MACRO dispatch as PASTE, so they can drift out of
sync with max_args exactly the way _dp did in the ABI tag list fixed
by #5544.

Extend main.cpp with build_double_paste_code() (same recursive-doubling
shape as build_paste_code(), but DOUBLE_PASTE consumes two arguments
per member, so an even slot index falls back to the next lower odd
DOUBLE_PASTE<N>) and build_type_body_table() (max_args - 1 MEMBERS
slots and one trailing EMPTY slot, 8 per line, matching how it is
written by hand today). Add a "type_body" argument that selects the
TYPE_BODY block, since it lives at a separate location in
macro_scope.hpp from the EXPAND..DOUBLE_PASTE63 block; plain invocation
is unchanged apart from covering the extended range. No longer emit
the tool's old trailing blank line, so its output is directly diffable
without post-processing.

Verified with c++ -std=c++11: running the tool (with and without
"type_body") and piping the raw output through the pinned astyle
reproduces both blocks of the current macro_scope.hpp byte for byte.
tests/src/unit-udt_macro.cpp (all NLOHMANN_DEFINE_TYPE_*/_WITH_NAMES/
zero-member variants) passes unchanged under -std=c++11 and -std=c++17
with -fsanitize=address,undefined.

Add a "macro_builder_check" Makefile target that builds main.cpp,
regenerates both blocks into a scratch directory inside the repository
(astyle's --project lookup needs the target files under the same tree
as .astylerc, unlike an external /tmp directory), and diffs them
against the corresponding ranges of macro_scope.hpp; wire it into
check-amalgamation next to the natvis check. Wire the same regeneration
into check_amalgamation.yml, splicing the (still unindented) generated
blocks back into the PR's own macro_scope.hpp before the existing
astyle/amalgamation step runs, so that step's own tree-wide astyle
pass both indents them and folds any drift into the amalgamation
patch/diff the workflow already produces.

Unlike amalgamate.py and generate_natvis.py, this step builds
tools/macro_builder/main.cpp from the pull request's own checkout
($MAIN_DIR) rather than a separate checkout of tools/ at develop: this
tool has no independent source of truth to regenerate against (its
README documents that it must reproduce macro_scope.hpp byte for
byte), so a develop-pinned copy would only reproduce the
generate_natvis.py trap fixed in a previous commit on this branch,
where a PR that teaches the tool to cover more of the file fails its
own CI until that PR merges and updates the develop copy.

Add tools/macro_builder/README.md documentation for both new tables
and the two-invocation usage, and a short pointer comment above
NLOHMANN_JSON_TYPE_BODY (the EXPAND pointer already covered the first
block; extended its wording to include DOUBLE_PASTE63).

Closes #5717 item 5 in full, completing what the documentation-only
"Document tools/macro_builder..." commit already on this branch left
open (that commit's README/pointer-comment half stands; it also covers
item 7).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:21:27 +02:00
Niels Lohmann 791cd88dfc Fix raw-TeX formulas and other documentation infrastructure debt (#5736)
* Fix math formulas rendering as raw TeX in the published docs

The privacy plugin self-hosts MathJax 2.7.0 but drops its
?config=TeX-MML-AM_CHTML query string, so the rehosted script loads no
input jax and the 10 formulas across 6 pages render as raw TeX to
readers. Remove pymdownx.arithmatex and the MathJax extra_javascript
entry, and rewrite the formulas in plain HTML (<sup>, <i>) instead.
This also drops a nine-year-old third-party script from every page.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix publish_documentation triggers and persist unneeded git credentials

publish_documentation.yml only triggered on docs/mkdocs/** pushes, but
the site also embeds .github/CODE_OF_CONDUCT.md, CONTRIBUTING.md,
SECURITY.md, cmake/{clang,gcc}_flags.cmake, .clang-tidy,
tools/astyle/.astylerc and tests/fmt_formatter/project/main.cpp via
pymdownx.snippets, so changes to those files never republished the
site. Extend the path filter to cover them, and switch runs-on from
the long-pinned ubuntu-22.04 to ubuntu-latest to match
ci_test_documentation.

Also add persist-credentials: false to the checkouts in ci_icpx and
ci_nvhpc (ubuntu.yml) and msvc-vs2026/msvc-arm64 (windows.yml), none
of which pushes with git, matching every other checkout in these
workflows.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix example build: broken debug echo, deprecations hidden for everything

docs/Makefile's debug echo used a space instead of a comma in
$(call cxx_standard ...), so it always printed an empty standard.
Every example was also compiled with -Wno-deprecated-declarations,
which would silently hide an accidental deprecated-API call in any of
them. Factor the duplicated compile flags into EXAMPLE_CPPFLAGS/
EXAMPLE_WARNFLAGS, build with -Werror=deprecated-declarations by
default, and only allow the three examples that intentionally
document deprecated API (the DEPRECATED_EXAMPLES list) to suppress it.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix check_structure.py NOLINT parsing and an off-by-one report line

`line.strip("<!-- NOLINT")` followed by `.strip(" -->")` strips any of
the characters in those sets from both ends, not a literal prefix; it
only happened to work for "Examples". A NOLINT'd section name
starting with N, O, L, I or T (e.g. "Notes", "Template parameters",
"Iterator invalidation", "Literals") was silently mangled, so the
suppression did not apply and the checker could report a spurious
missing/misordered section. Parse the comment with a regex instead.
The same fragile strip() pattern was used for heading text; replace it
with a plain prefix slice. Also fix the admonition_title report, which
used the 0-based line index while every other report in the file uses
lineno+1.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix docset README's non-existent make target and stale fallback URL

The README told readers to run `make nlohmann_json.docset`, but the
Makefile's targets are `all`, `JSON_for_Modern_C++.docset` and
`install_docset_zeal`; the documented command has failed with "No
rule to make target" since #2967 (2021). Point the README at the real
target and folder name. Info.plist's DashDocSetFallbackURL also still
pointed at the old nlohmann.github.io/json/ URL instead of the
canonical https://json.nlohmann.me/ from mkdocs.yml's site_url.

Leave list_missing_pages/list_removed_paths alone: they may become
redundant once #5638's check_docset() lands, which is a follow-up.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove two leftover Doxygen-era .link files from the examples directory

parse__iterator_pair.link and parse__pointers.link each held only a
Wandbox "online" permalink from the old Doxygen docs. #3071 deleted
every other .link file in 2021; these two came in through a parallel
PR (#3100) and were never referenced by any page, script or config.

Part of #5718

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix MacPorts CMake example to include its own snippet files

The MacPorts "Example: CMake" block included
integration/homebrew/example.cpp and
integration/homebrew/CMakeLists.txt instead of the MacPorts files
right next to it, a copy-paste slip from the Homebrew section. Nothing
referenced integration/macports/CMakeLists.txt as a result. The page
rendered correctly only because the homebrew, macports and
vcpkg/CMakeLists.txt snippets are byte-identical, so a future edit to
the MacPorts files would not have shown up on the page.

Part of #5718 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Do not persist git credentials in publish_documentation's checkout

The checkout step in publish_documentation.yml left the default
persist-credentials: true, so GITHUB_TOKEN stayed writable in
.git/config for the rest of the job (zizmor's artipacked finding). The
Deploy documentation step authenticates through its own github_token
input to peaceiris/actions-gh-pages and does not push with the
checked-out credentials, so persist-credentials: false is safe here,
matching every other checkout in the workflow set.

Overlaps #5638, which edits this same checkout step (adds
fetch-depth: 0); expect a rebase conflict there.

Part of #5718 item 2

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Replace list_missing_pages/list_removed_paths with a comm(1)-based diff

The docset Makefile's list_missing_pages ran one sqlite3 query per
mkdocs page, and list_removed_paths nested a loop over all mkdocs pages
inside a loop over all docset index paths (O(n*m) shell iteration).
Issue #5718 item 5 suggested removing or reducing these targets once
#5638's check_docset() lands, but that PR is still open and covers only
API pages and macros, not the full page set these targets check.
Replace the loops with two sorted path lists (DOCSET_PAGE_PATHS from
mkdocs' markdown sources, DOCSET_INDEX_PATHS from the built docset
index) compared with a single comm(1) call each, verified to produce
output identical to the old loops against the current docSet.dsidx.

The sed expression used '#' as its delimiter, which GNU Make reads as
a comment character even inside a variable assignment, truncating the
line and orphaning the closing paren of $(shell ...) ("unterminated
call to function 'shell': missing ')'"). Use '@' as the delimiter
instead.

Part of #5718 item 5

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 22:15:20 +02:00
Niels Lohmann 7d7055ec50 Fix stack overflow converting deep values between specializations (#5723)
* Fix stack overflow converting deep values between specializations

Constructing a basic_json from another specialization (json to
ordered_json or back, also via get<ordered_json>()) converted every
container with its range constructor, which calls the converting
constructor for each element. The call stack therefore grew with every
nesting level, and a value nested some 30,000 levels deep overflowed it.

The conversion now bounds its descent the way the copy constructor does
since #5387: the first 128 levels are converted exactly as before, and
below that convert_iteratively() finishes the value with an explicit
stack. It builds each container bottom-up from its converted elements
with the container's range constructor, so member order and keys that
become equal are handled as before, and it gives a value its type only
once its container exists, so an exception leaves nothing behind that
cannot be destroyed. Parents (JSON_DIAGNOSTICS) and positions
(JSON_DIAGNOSTIC_POSITIONS) are set for every value.

Converting a null value no longer resets its positions: the constructor
assigned null to a value that already was null, which swapped in the
positions of the temporary.

Fixes #5650.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Explain why converting null keeps positions and why next is a reference

Review feedback on #5723 (gregmarr): clarify in comments that the
converting constructor has already copied the positions of val, which
the null case keeps like every other case, and that next must be a
reference into pending so that ++next advances the stored iterator.

Comments only; no code change.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Refer to recursion_depth_limit() in the convert_structured() docs

The comment still named nesting_depth_limit, which #5637 removed on
develop in favor of detail::recursion_depth_limit().

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Advance the pending iterator through pending.back() and shorten the null comment

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:36:04 +02:00
Niels Lohmann 3b6ae43c53 Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions (#5598)
* Add JSON_DISABLE_TUPLE_REFERENCE_CONVERSION to fix std::tuple conversions

basic_json can be constructed from std::tuple<json&>, which it turns into
a one-element array. Because of this, std::tuple picks its converting
constructor that converts the whole source tuple instead of the
element-wise one. As a result, std::tuple<const json&> built from
std::forward_as_tuple(j) binds to a temporary (a compile error with libc++,
a dangling reference with other standard libraries), and std::tuple<json>
built the same way holds [j] instead of a copy of j.

The new opt-in macro JSON_DISABLE_TUPLE_REFERENCE_CONVERSION (CMake option
JSON_DisableTupleReferenceConversion) removes the conversion from a
one-element tuple holding a reference to the same basic_json type, so
std::tuple converts element-wise. It is off by default, so existing
behavior is unchanged. It does not change any function body and therefore
is not part of the ABI tag.

Fixes #2226

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Convert one-element tuples to arrays on every compiler

to_json for std::tuple assigns a braced list, j = { std::get<Idx>(t)... }.
With a single element that is itself a basic_json, Apple clang 15 and 16
treat j = {x} as a copy of x, so std::tuple<json>{true} became true
instead of [true]. The macOS jobs (Xcode 15.1, 16.1) failed the new
checks in unit-disable-tuple-reference-conversion and unit-regression2.

The one-element overload that already handles
JSON_BRACE_INIT_COPY_SEMANTICS builds the array (or object, for a
[string, value] element) explicitly, the same way the initializer-list
constructor does. Use it unconditionally. The output is unchanged on
compilers that already wrapped the element.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Skip json reference tuple tests on clang < 4 and GCC < 5

ci_test_compilers_gcc_old (4.8) and ci_test_compilers_clang (3.4) could
not compile the new tuple tests. Creating a std::tuple of basic_json
references, e.g. std::forward_as_tuple(j), makes these compilers
instantiate basic_json's conversion operator for libstdc++'s internal
tuple bases, which fails hard. This happens with and without
JSON_DISABLE_TUPLE_REFERENCE_CONVERSION, so it is a limitation of these
compilers, not of the new option.

Tested with the CI images: clang 3.4 to 3.9 and GCC 4.8 and 4.9 fail,
clang 4, 5, and 6 and GCC 5 and 6 compile all cases. Skip only the
checks that create such tuples; the is_constructible checks still run.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:53:56 +02:00
Niels Lohmann 5f1727cef2 Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	tests/src/unit-alt-string.cpp
2026-09-30 20:11:28 +02:00
Niels Lohmann c16dd7e4f5 Fix CI: do not hide base class methods in the #3989 test SAX parsers
clang-tidy's bugprone-derived-method-shadowing-base-method rejected the
test SAX parsers that derive from json_sax_dom_parser (or the test's
SaxEventLogger) and redefine their non-virtual event functions. The
recovering DOM parsers of unit-class_parser.cpp and unit-regression2.cpp
now hold a json_sax_dom_parser and forward to it, SaxEventLogger gets a
flag to recover from errors instead of a derived class, and the one
function unit-alt-string.cpp redefines is marked as intended.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 04:50:11 +02:00
Niels Lohmann 792853d725 Fix CI: clang-tidy and JSON_NOEXCEPTION in the #3989 changes
ci_clang_tidy flagged nested conditional operators in the BON8 skip
code and in remove_incomplete_utf8_sequence()
(readability-avoid-nested-conditional-operator), and two branches with
the same body in recover_string() (bugprone-branch-clone). Use if
chains instead of the nested conditionals and merge the two branches
into one condition; the short-circuit order is unchanged.

ci_test_noexceptions aborted in the #3989 regression test: it compares
the first recovered error with the message of the exception that
from_cbor() and friends throw, and under JSON_NOEXCEPTION that call
aborts instead of throwing. Guard binary_error_message() and its uses
with #if !defined(JSON_NOEXCEPTION), as other tests in the suite do.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 16:42:57 +02:00
Niels Lohmann 4bdf1b7e74 Fix CI: no global std::string in the #3989 regression test
The U+FFFD constant was a global std::string, which the clang CI build
rejects with -Wglobal-constructors. It is now a function.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 09:27:46 +02:00
Niels Lohmann b1e9d98e41 Merge remote-tracking branch 'origin/claude/fix-issue-3989-db7e45' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:31:07 +02:00
Niels Lohmann f855d257df Fix CI: keep raw strings with backslashes out of test macros for MSVC
MSVC stringizes the arguments of doctest's CHECK() so that a raw string
literal becomes an ordinary one, and then reported the "\q" in the new
alt_string recovery test as warning C4129, an error with /WX. The input
is now a variable. The same pattern with "\u0000" in the parser's
recovery test is replaced by a JSON value built from a std::string.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-29 03:30:39 +02:00
Niels Lohmann 3e683e9c04 Merge branch 'develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:22:25 +02:00
Niels Lohmann d1d84ed9af Fix Flawfinder: format the BSON element type without snprintf
Moving the report of an unsupported BSON element type into
skip_unsupported_bson_element() moved its snprintf() call, which
Flawfinder then reported as a new CWE-134 finding. The two hexadecimal
digits are now computed directly; the message is unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 20:36:28 +02:00
Niels Lohmann de8529f99b Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>

# Conflicts:
#	include/nlohmann/detail/input/binary_reader.hpp
#	single_include/nlohmann/json.hpp
2026-09-28 17:53:51 +02:00
Niels Lohmann 677794f076 Fix CI: recover numbers with the string operations every string_t has
recover_number() used back(), pop_back(), front(), and
find_first_of(const char*), which the minimal alt_string of
unit-alt-string.cpp does not provide. Since the binary readers recover
UBJSON/BJData high-precision numbers with it, from_ubjson() instantiated
it, too, and the test no longer compiled. It now uses only size(),
operator[], and resize(), and unit-alt-string.cpp recovers from errors
in JSON text, UBJSON, and CBOR.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 12:17:11 +02:00
Niels Lohmann 437a95cfdb Repair complete items in binary formats when parse_error() returns true (#3989)
When the SAX parser asks to recover, the binary readers now repair an
item whose end is known and read on after it, as RFC 8949, Section 5.3
describes for CBOR:

- CBOR: tags are ignored, and simple values other than false, true, and
  null become null (RFC 8949, Section 6.1); a negative integer below the
  range of number_integer_t becomes the nearest floating-point number.
- Strings that are not valid UTF-8 get U+FFFD for each ill-formed
  sequence, as in JSON text; so does a UBJSON/BJData char above 0x7F.
- UBJSON/BJData high-precision numbers keep their longest valid
  beginning (via the lexer's recover_token()), or become infinity.
- Members whose key is not a string are skipped (CBOR, MessagePack,
  BON8), like members without a key in JSON text.
- BSON elements of types the library does not read (ObjectId, datetime,
  decimal128, ...) become null; a string without its terminator and a
  document whose size does not match are kept.

Where the end of an item is unknown, reading stops as before, except
that BSON skips to the end of the document, whose size it knows.

The value read before such an error is now completed by the reader from
its container stack, as the JSON parser does, instead of by a proxy SAX
parser, which is removed. Like the parser, binary_reader gets an
AllowRecovery template parameter, so that from_*() compile without the
new code.

Tests: a table of repairs, numbers out of range, errors that stop, and
a sweep over changed and removed bytes of eight encodings that checks
balanced events and that the first error is the one from_*() reports.
All fuzzers now run a recovering checker; the binary ones also check
that it reports an error exactly when from_*() fails.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 21:40:59 +02:00
Niels Lohmann c8735246d0 Merge remote-tracking branch 'origin/develop' into claude/fix-issue-3989-db7e45
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 21:02:38 +02:00
Niels Lohmann 9adb510a0d Recover from parse errors when parse_error() returns true (#3989)
The return value of json_sax::parse_error() was documented both as "must
return false" and as "whether the parsing should continue", and the code
just passed it on. For JSON text, parsing stopped anyway, but sax_parse()
could report success for invalid input. The binary readers read on after
the error, looping forever on a CBOR indefinite-length array without its
end.

Now false stops parsing, and true recovers from the error:

- JSON text is repaired with the smallest local edit (insert a missing
  ',' or ':', remove a stray token, keep the readable part of a broken
  string or number, null for a value that cannot be read, close the
  innermost container at a wrong closing bracket and all of them at the
  end of the input), and parsing continues. The SAX events stay balanced,
  every key is followed by exactly one value, and each token is reported
  at most once.
- The binary formats cannot resynchronize, so they stop, but complete
  the value read so far.

sax_parse() returns false after any error. parse(), accept(), and the
from_*() functions never recover and compile to the same code as before.

Supersedes #4522.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 20:16:42 +02:00
76 changed files with 9727 additions and 2392 deletions
+56
View File
@@ -35,6 +35,7 @@ jobs:
MAIN_DIR: ${{ github.workspace }}/main MAIN_DIR: ${{ github.workspace }}/main
INCLUDE_DIR: ${{ github.workspace }}/main/single_include/nlohmann INCLUDE_DIR: ${{ github.workspace }}/main/single_include/nlohmann
TOOL_DIR: ${{ github.workspace }}/tools/tools/amalgamate TOOL_DIR: ${{ github.workspace }}/tools/tools/amalgamate
NATVIS_TOOL_DIR: ${{ github.workspace }}/tools/tools/generate_natvis
steps: steps:
- name: Harden Runner - name: Harden Runner
@@ -61,6 +62,48 @@ jobs:
python3 -mvenv venv python3 -mvenv venv
venv/bin/pip3 install -r $MAIN_DIR/tools/astyle/requirements.txt venv/bin/pip3 install -r $MAIN_DIR/tools/astyle/requirements.txt
- name: Install generate_natvis dependencies
run: pip3 install -r $NATVIS_TOOL_DIR/requirements.txt
- name: Regenerate the tools/macro_builder tables in macro_scope.hpp
run: |
cd $MAIN_DIR
# Built from this PR's own tools/macro_builder/main.cpp, not a
# develop checkout: unlike amalgamate.py and generate_natvis.py,
# this tool has no other source of truth to check against (its
# own README documents that it must reproduce macro_scope.hpp
# byte for byte), so there is nothing to gain from checking out a
# separate copy, and doing so would make this step fail on a PR
# that adds support for a new dispatch table until that PR itself
# merges to develop, the same way generate_natvis.py's --version
# requirement briefly did.
TMPDIR=$(mktemp -d ./macro_builder_check.XXXXXX)
c++ -std=c++11 tools/macro_builder/main.cpp -o "$TMPDIR/macro_builder"
"$TMPDIR/macro_builder" > "$TMPDIR/paste.hpp"
"$TMPDIR/macro_builder" type_body > "$TMPDIR/type_body.hpp"
# Splice the (still unindented) generated blocks back into
# macro_scope.hpp; the astyle pass below indents their
# continuation lines the same way it does for the rest of
# include/, so a correctly regenerated file comes out unchanged.
awk -v newfile="$TMPDIR/paste.hpp" '
BEGIN { while ((getline line < newfile) > 0) { new = new line "\n" } }
/^#define NLOHMANN_JSON_EXPAND\( x \) x$/ { printf "%s", new; skip=1 }
skip && /^#define NLOHMANN_JSON_DOUBLE_PASTE63\(/ { skip=0; next }
skip { next }
{ print }
' include/nlohmann/detail/macro_scope.hpp > "$TMPDIR/macro_scope_1.hpp"
awk -v newfile="$TMPDIR/type_body.hpp" '
BEGIN { while ((getline line < newfile) > 0) { new = new line "\n" } }
/^#define NLOHMANN_JSON_TYPE_BODY\(Prefix, \.\.\.\)/ { printf "%s", new; skip=1 }
skip && /^[[:space:]]*NLOHMANN_JSON_TYPE_BODY_SENTINEL\)\)$/ { skip=0; next }
skip { next }
{ print }
' "$TMPDIR/macro_scope_1.hpp" > "$TMPDIR/macro_scope_2.hpp"
mv "$TMPDIR/macro_scope_2.hpp" include/nlohmann/detail/macro_scope.hpp
rm -rf "$TMPDIR"
- name: Regenerate amalgamation, formatting, and BUILD.bazel - name: Regenerate amalgamation, formatting, and BUILD.bazel
run: | run: |
cd $MAIN_DIR cd $MAIN_DIR
@@ -88,6 +131,19 @@ jobs:
${{ github.workspace }}/venv/bin/astyle --project=tools/astyle/.astylerc --suffix=none --quiet \ ${{ github.workspace }}/venv/bin/astyle --project=tools/astyle/.astylerc --suffix=none --quiet \
$(find $SOURCE_DIRS -type f \( -name '*.hpp' -o -name '*.cpp' -o -name '*.cu' \) -not -path 'tests/thirdparty/*' -not -path 'tests/abi/include/nlohmann/*' | sort) $(find $SOURCE_DIRS -type f \( -name '*.hpp' -o -name '*.cpp' -o -name '*.cu' \) -not -path 'tests/thirdparty/*' -not -path 'tests/abi/include/nlohmann/*' | sort)
- name: Regenerate nlohmann_json.natvis
run: |
cd $MAIN_DIR
# Pass --version explicitly so this step also works with the tool
# copy from develop before this repository's own generate_natvis.py
# learns to derive the version itself: the older script requires
# --version, and the newer one accepts it as an explicit override.
ABI_MACROS=include/nlohmann/detail/abi_macros.hpp
VERSION_MAJOR=$(grep -m1 'define NLOHMANN_JSON_VERSION_MAJOR' $ABI_MACROS | grep -o '[0-9]\+')
VERSION_MINOR=$(grep -m1 'define NLOHMANN_JSON_VERSION_MINOR' $ABI_MACROS | grep -o '[0-9]\+')
VERSION_PATCH=$(grep -m1 'define NLOHMANN_JSON_VERSION_PATCH' $ABI_MACROS | grep -o '[0-9]\+')
python3 $NATVIS_TOOL_DIR/generate_natvis.py --version "$VERSION_MAJOR.$VERSION_MINOR.$VERSION_PATCH" $MAIN_DIR
- name: Build patch and check for differences - name: Build patch and check for differences
id: diff id: diff
run: | run: |
+13 -1
View File
@@ -7,6 +7,16 @@ on:
- develop - develop
paths: paths:
- docs/mkdocs/** - docs/mkdocs/**
# the site also embeds these files via pymdownx.snippets
# (mkdocs.yml sets restrict_base_path: false for this)
- .clang-tidy
- .github/CODE_OF_CONDUCT.md
- .github/CONTRIBUTING.md
- .github/SECURITY.md
- cmake/clang_flags.cmake
- cmake/gcc_flags.cmake
- tests/fmt_formatter/project/main.cpp
- tools/astyle/.astylerc
workflow_dispatch: workflow_dispatch:
# we don't want to have concurrent jobs, and we don't want to cancel running jobs to avoid broken publications # we don't want to have concurrent jobs, and we don't want to cancel running jobs to avoid broken publications
@@ -23,7 +33,7 @@ jobs:
contents: write contents: write
if: github.repository == 'nlohmann/json' if: github.repository == 'nlohmann/json'
runs-on: ubuntu-22.04 runs-on: ubuntu-latest
steps: steps:
- name: Harden Runner - name: Harden Runner
uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1 uses: step-security/harden-runner@e14015d583714f6e62063499dc959a02595150a1 # v2.21.1
@@ -31,6 +41,8 @@ jobs:
egress-policy: audit egress-policy: audit
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Install virtual environment - name: Install virtual environment
run: make install_venv -C docs/mkdocs run: make install_venv -C docs/mkdocs
+5 -1
View File
@@ -100,7 +100,7 @@ jobs:
container: ubuntu:focal container: ubuntu:focal
strategy: strategy:
matrix: matrix:
target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_disableenumserialization, ci_test_skiplibraryversioncheck, ci_test_simdutf, ci_test_strict_nul_handling, ci_test_no_thread_local] target: [ci_cmake_flags, ci_test_diagnostics, ci_test_diagnostic_positions, ci_test_noexceptions, ci_test_noimplicitconversions, ci_test_legacycomparison, ci_test_noglobaludls, ci_test_disableenumserialization, ci_test_disabletuplereferenceconversion, ci_test_skiplibraryversioncheck, ci_test_simdutf, ci_test_strict_nul_handling, ci_test_no_thread_local]
steps: steps:
- name: Install build-essential - name: Install build-essential
run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev run: apt-get update ; apt-get install -y build-essential unzip wget git libssl-dev
@@ -346,6 +346,8 @@ jobs:
container: intel/oneapi-hpckit:latest container: intel/oneapi-hpckit:latest
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2 uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2
- name: Run CMake - name: Run CMake
@@ -358,6 +360,8 @@ jobs:
container: nvcr.io/nvidia/nvhpc:25.5-devel-cuda12.9-ubuntu22.04 container: nvcr.io/nvidia/nvhpc:25.5-devel-cuda12.9-ubuntu22.04
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2 uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2
- name: Run CMake - name: Run CMake
+4
View File
@@ -87,6 +87,8 @@ jobs:
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Get latest CMake and ninja - name: Get latest CMake and ninja
uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2 uses: lukka/get-cmake@fffaaafeea488556c2c12dad60690008bc1caacb # v4.4.2
- name: Set extra CXX_FLAGS for latest std_version - name: Set extra CXX_FLAGS for latest std_version
@@ -123,6 +125,8 @@ jobs:
steps: steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
persist-credentials: false
- name: Run CMake (Release) - name: Run CMake (Release)
run: cmake -S . -B build -G "Visual Studio 18 2026" -A ARM64 -DJSON_BuildTests=On -DCMAKE_CXX_FLAGS="/W4 /WX" run: cmake -S . -B build -G "Visual Studio 18 2026" -A ARM64 -DJSON_BuildTests=On -DCMAKE_CXX_FLAGS="/W4 /WX"
if: matrix.build_type == 'Release' if: matrix.build_type == 'Release'
+6
View File
@@ -55,6 +55,7 @@ option(JSON_Diagnostic_Positions "Enable diagnostic positions." OFF)
option(JSON_GlobalUDLs "Place user-defined string literals in the global namespace." ON) option(JSON_GlobalUDLs "Place user-defined string literals in the global namespace." ON)
option(JSON_ImplicitConversions "Enable implicit conversions." ON) option(JSON_ImplicitConversions "Enable implicit conversions." ON)
option(JSON_DisableEnumSerialization "Disable default integer enum serialization." OFF) option(JSON_DisableEnumSerialization "Disable default integer enum serialization." OFF)
option(JSON_DisableTupleReferenceConversion "Disable conversion from a one-element tuple of a JSON reference." OFF)
option(JSON_LegacyDiscardedValueComparison "Enable legacy discarded value comparison." OFF) option(JSON_LegacyDiscardedValueComparison "Enable legacy discarded value comparison." OFF)
option(JSON_Install "Install CMake targets during install step." ${MAIN_PROJECT}) option(JSON_Install "Install CMake targets during install step." ${MAIN_PROJECT})
option(JSON_MultipleHeaders "Use non-amalgamated version of the library." ON) option(JSON_MultipleHeaders "Use non-amalgamated version of the library." ON)
@@ -101,6 +102,10 @@ if (JSON_DisableEnumSerialization)
message(STATUS "Enum integer serialization is disabled (JSON_DISABLE_ENUM_SERIALIZATION=1)") message(STATUS "Enum integer serialization is disabled (JSON_DISABLE_ENUM_SERIALIZATION=1)")
endif() endif()
if (JSON_DisableTupleReferenceConversion)
message(STATUS "Tuple reference conversion is disabled (JSON_DISABLE_TUPLE_REFERENCE_CONVERSION=1)")
endif()
if (JSON_LegacyDiscardedValueComparison) if (JSON_LegacyDiscardedValueComparison)
message(STATUS "Legacy discarded value comparison enabled (JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1)") message(STATUS "Legacy discarded value comparison enabled (JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1)")
endif() endif()
@@ -143,6 +148,7 @@ target_compile_definitions(
$<$<NOT:$<BOOL:${JSON_GlobalUDLs}>>:JSON_USE_GLOBAL_UDLS=0> $<$<NOT:$<BOOL:${JSON_GlobalUDLs}>>:JSON_USE_GLOBAL_UDLS=0>
$<$<NOT:$<BOOL:${JSON_ImplicitConversions}>>:JSON_USE_IMPLICIT_CONVERSIONS=0> $<$<NOT:$<BOOL:${JSON_ImplicitConversions}>>:JSON_USE_IMPLICIT_CONVERSIONS=0>
$<$<BOOL:${JSON_DisableEnumSerialization}>:JSON_DISABLE_ENUM_SERIALIZATION=1> $<$<BOOL:${JSON_DisableEnumSerialization}>:JSON_DISABLE_ENUM_SERIALIZATION=1>
$<$<BOOL:${JSON_DisableTupleReferenceConversion}>:JSON_DISABLE_TUPLE_REFERENCE_CONVERSION=1>
$<$<BOOL:${JSON_Diagnostics}>:JSON_DIAGNOSTICS=1> $<$<BOOL:${JSON_Diagnostics}>:JSON_DIAGNOSTICS=1>
$<$<BOOL:${JSON_Diagnostic_Positions}>:JSON_DIAGNOSTIC_POSITIONS=1> $<$<BOOL:${JSON_Diagnostic_Positions}>:JSON_DIAGNOSTIC_POSITIONS=1>
$<$<BOOL:${JSON_LegacyDiscardedValueComparison}>:JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1> $<$<BOOL:${JSON_LegacyDiscardedValueComparison}>:JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON=1>
+32 -83
View File
@@ -1,4 +1,4 @@
.PHONY: pretty clean ChangeLog.md release update_hedley update_hedley_undef BUILD.bazel .PHONY: pretty clean ChangeLog.md release update_hedley update_hedley_undef BUILD.bazel natvis macro_builder_check
########################################################################## ##########################################################################
# configuration # configuration
@@ -24,6 +24,9 @@ AMALGAMATED_FWD_FILE=single_include/nlohmann/json_fwd.hpp
# json_literals.hpp only includes <nlohmann/json.hpp>, so it is copied verbatim # json_literals.hpp only includes <nlohmann/json.hpp>, so it is copied verbatim
AMALGAMATED_LITERALS_FILE=single_include/nlohmann/json_literals.hpp AMALGAMATED_LITERALS_FILE=single_include/nlohmann/json_literals.hpp
# the header with the argument-counting macros generated by tools/macro_builder
MACRO_SCOPE_HPP=include/nlohmann/detail/macro_scope.hpp
########################################################################## ##########################################################################
# documentation of the Makefile's targets # documentation of the Makefile's targets
@@ -36,13 +39,9 @@ all:
@echo "ChangeLog.md - generate ChangeLog file" @echo "ChangeLog.md - generate ChangeLog file"
@echo "check-amalgamation - check whether sources have been amalgamated and BUILD.bazel is up to date" @echo "check-amalgamation - check whether sources have been amalgamated and BUILD.bazel is up to date"
@echo "clean - remove built files" @echo "clean - remove built files"
@echo "doctest - compile example files and check their output" @echo "fuzzing - see tests/fuzzing.md for how to build and run the fuzzers"
@echo "fuzz_testing - prepare fuzz testing of the JSON parser" @echo "macro_builder_check - check that macro_scope.hpp matches tools/macro_builder's output"
@echo "fuzz_testing_bon8 - prepare fuzz testing of the BON8 parser" @echo "natvis - regenerate nlohmann_json.natvis from the current ABI tags and version"
@echo "fuzz_testing_bson - prepare fuzz testing of the BSON parser"
@echo "fuzz_testing_cbor - prepare fuzz testing of the CBOR parser"
@echo "fuzz_testing_msgpack - prepare fuzz testing of the MessagePack parser"
@echo "fuzz_testing_ubjson - prepare fuzz testing of the UBJSON parser"
@echo "pretty - beautify code with Artistic Style" @echo "pretty - beautify code with Artistic Style"
@echo "run_benchmarks - build and run benchmarks" @echo "run_benchmarks - build and run benchmarks"
@echo "update_hedley - download Hedley and regenerate hedley.hpp / hedley_undef.hpp" @echo "update_hedley - download Hedley and regenerate hedley.hpp / hedley_undef.hpp"
@@ -61,74 +60,6 @@ run_benchmarks:
cd cmake-build-benchmarks ; ./json_benchmarks cd cmake-build-benchmarks ; ./json_benchmarks
##########################################################################
# fuzzing
##########################################################################
# the overall fuzz testing target
fuzz_testing:
rm -fr fuzz-testing
mkdir -p fuzz-testing fuzz-testing/testcases fuzz-testing/out
$(MAKE) parse_afl_fuzzer -C tests CXX=afl-clang++
mv tests/parse_afl_fuzzer fuzz-testing/fuzzer
find tests/data/json_tests -size -5k -name *json | xargs -I{} cp "{}" fuzz-testing/testcases
@echo "Execute: afl-fuzz -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer"
fuzz_testing_bon8:
rm -fr fuzz-testing
mkdir -p fuzz-testing fuzz-testing/testcases fuzz-testing/out
$(MAKE) parse_bon8_fuzzer -C tests CXX=afl-clang++
mv tests/parse_bon8_fuzzer fuzz-testing/fuzzer
find tests/data -size -5k -name *.bon8 | xargs -I{} cp "{}" fuzz-testing/testcases
@echo "Execute: afl-fuzz -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer"
fuzz_testing_bson:
rm -fr fuzz-testing
mkdir -p fuzz-testing fuzz-testing/testcases fuzz-testing/out
$(MAKE) parse_bson_fuzzer -C tests CXX=afl-clang++
mv tests/parse_bson_fuzzer fuzz-testing/fuzzer
find tests/data -size -5k -name *.bson | xargs -I{} cp "{}" fuzz-testing/testcases
@echo "Execute: afl-fuzz -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer"
fuzz_testing_cbor:
rm -fr fuzz-testing
mkdir -p fuzz-testing fuzz-testing/testcases fuzz-testing/out
$(MAKE) parse_cbor_fuzzer -C tests CXX=afl-clang++
mv tests/parse_cbor_fuzzer fuzz-testing/fuzzer
find tests/data -size -5k -name *.cbor | xargs -I{} cp "{}" fuzz-testing/testcases
@echo "Execute: afl-fuzz -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer"
fuzz_testing_msgpack:
rm -fr fuzz-testing
mkdir -p fuzz-testing fuzz-testing/testcases fuzz-testing/out
$(MAKE) parse_msgpack_fuzzer -C tests CXX=afl-clang++
mv tests/parse_msgpack_fuzzer fuzz-testing/fuzzer
find tests/data -size -5k -name *.msgpack | xargs -I{} cp "{}" fuzz-testing/testcases
@echo "Execute: afl-fuzz -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer"
fuzz_testing_ubjson:
rm -fr fuzz-testing
mkdir -p fuzz-testing fuzz-testing/testcases fuzz-testing/out
$(MAKE) parse_ubjson_fuzzer -C tests CXX=afl-clang++
mv tests/parse_ubjson_fuzzer fuzz-testing/fuzzer
find tests/data -size -5k -name *.ubjson | xargs -I{} cp "{}" fuzz-testing/testcases
@echo "Execute: afl-fuzz -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer"
fuzzing-start:
afl-fuzz -S fuzzer1 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer > /dev/null &
afl-fuzz -S fuzzer2 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer > /dev/null &
afl-fuzz -S fuzzer3 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer > /dev/null &
afl-fuzz -S fuzzer4 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer > /dev/null &
afl-fuzz -S fuzzer5 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer > /dev/null &
afl-fuzz -S fuzzer6 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer > /dev/null &
afl-fuzz -S fuzzer7 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer > /dev/null &
afl-fuzz -M fuzzer0 -i fuzz-testing/testcases -o fuzz-testing/out fuzz-testing/fuzzer
fuzzing-stop:
-killall fuzzer
-killall afl-fuzz
########################################################################## ##########################################################################
# Static analysis # Static analysis
########################################################################## ##########################################################################
@@ -158,10 +89,6 @@ install_astyle:
pretty: install_astyle pretty: install_astyle
$(ASTYLE) --project=tools/astyle/.astylerc $(SRCS) $(TESTS_SRCS) $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_LITERALS_FILE) docs/mkdocs/docs/examples/*.cpp $(ASTYLE) --project=tools/astyle/.astylerc $(SRCS) $(TESTS_SRCS) $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_LITERALS_FILE) docs/mkdocs/docs/examples/*.cpp
# call the Clang-Format on all source files
pretty_format:
for FILE in $(SRCS) $(TESTS_SRCS) $(AMALGAMATED_FILE) docs/mkdocs/docs/examples/*.cpp; do echo $$FILE; clang-format -i $$FILE; done
# create single header files and pretty print # create single header files and pretty print
amalgamate: $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_LITERALS_FILE) amalgamate: $(AMALGAMATED_FILE) $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_LITERALS_FILE)
$(MAKE) pretty $(MAKE) pretty
@@ -178,8 +105,26 @@ $(AMALGAMATED_FWD_FILE): $(SRCS)
$(AMALGAMATED_LITERALS_FILE): include/nlohmann/json_literals.hpp $(AMALGAMATED_LITERALS_FILE): include/nlohmann/json_literals.hpp
cp include/nlohmann/json_literals.hpp $(AMALGAMATED_LITERALS_FILE) cp include/nlohmann/json_literals.hpp $(AMALGAMATED_LITERALS_FILE)
# regenerate nlohmann_json.natvis from the ABI tags and version in include/nlohmann/detail/abi_macros.hpp
natvis:
python3 tools/generate_natvis/generate_natvis.py .
# regenerate the two tools/macro_builder blocks of $(MACRO_SCOPE_HPP) (see its README.md) and diff against the
# checked-in header; phony, because it never writes $(MACRO_SCOPE_HPP) itself
macro_builder_check:
@set -e; \
TMPDIR=$$(mktemp -d ./macro_builder_check.XXXXXX); \
trap 'rm -rf "$$TMPDIR"' EXIT; \
$(CXX) -std=c++11 tools/macro_builder/main.cpp -o "$$TMPDIR/macro_builder"; \
"$$TMPDIR/macro_builder" > "$$TMPDIR/paste.hpp"; \
"$$TMPDIR/macro_builder" type_body > "$$TMPDIR/type_body.hpp"; \
$(ASTYLE) --project=tools/astyle/.astylerc --suffix=none --quiet "$$TMPDIR/paste.hpp" "$$TMPDIR/type_body.hpp"; \
sed -n '/^#define NLOHMANN_JSON_EXPAND( x ) x$$/,/^#define NLOHMANN_JSON_DOUBLE_PASTE63(/p' $(MACRO_SCOPE_HPP) > "$$TMPDIR/paste_actual.hpp"; \
sed -n '/^#define NLOHMANN_JSON_TYPE_BODY(Prefix, \.\.\.)/,/^ NLOHMANN_JSON_TYPE_BODY_SENTINEL))$$/p' $(MACRO_SCOPE_HPP) > "$$TMPDIR/type_body_actual.hpp"; \
diff "$$TMPDIR/paste.hpp" "$$TMPDIR/paste_actual.hpp" || (echo "===================================================================\n $(MACRO_SCOPE_HPP) (NLOHMANN_JSON_EXPAND..NLOHMANN_JSON_DOUBLE_PASTE63) is out of date!\n Regenerate it, see tools/macro_builder/README.md.\n===================================================================" ; exit 1); \
diff "$$TMPDIR/type_body.hpp" "$$TMPDIR/type_body_actual.hpp" || (echo "===================================================================\n $(MACRO_SCOPE_HPP) (NLOHMANN_JSON_TYPE_BODY) is out of date!\n Regenerate it, see tools/macro_builder/README.md.\n===================================================================" ; exit 1)
# check if file single_include/nlohmann/json.hpp has been amalgamated from the nlohmann sources # check if file single_include/nlohmann/json.hpp has been amalgamated from the nlohmann sources
# Note: this target is called by Travis
check-amalgamation: check-amalgamation:
@mv $(AMALGAMATED_FILE) $(AMALGAMATED_FILE)~ @mv $(AMALGAMATED_FILE) $(AMALGAMATED_FILE)~
@mv $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_FWD_FILE)~ @mv $(AMALGAMATED_FWD_FILE) $(AMALGAMATED_FWD_FILE)~
@@ -195,6 +140,11 @@ check-amalgamation:
@$(MAKE) BUILD.bazel @$(MAKE) BUILD.bazel
@diff BUILD.bazel BUILD.bazel~ || (echo "===================================================================\n BUILD.bazel is out of date! Please run 'make BUILD.bazel'.\n===================================================================" ; mv BUILD.bazel~ BUILD.bazel ; false) @diff BUILD.bazel BUILD.bazel~ || (echo "===================================================================\n BUILD.bazel is out of date! Please run 'make BUILD.bazel'.\n===================================================================" ; mv BUILD.bazel~ BUILD.bazel ; false)
@mv BUILD.bazel~ BUILD.bazel @mv BUILD.bazel~ BUILD.bazel
@mv nlohmann_json.natvis nlohmann_json.natvis~
@$(MAKE) natvis
@diff nlohmann_json.natvis nlohmann_json.natvis~ || (echo "===================================================================\n nlohmann_json.natvis is out of date! Please run 'make natvis'.\n===================================================================" ; mv nlohmann_json.natvis~ nlohmann_json.natvis ; false)
@mv nlohmann_json.natvis~ nlohmann_json.natvis
@$(MAKE) macro_builder_check
# generate the Bazel BUILD file; phony, because a removed header would not trigger a rebuild # generate the Bazel BUILD file; phony, because a removed header would not trigger a rebuild
BUILD.bazel: BUILD.bazel:
@@ -246,7 +196,7 @@ release: include.zip json.tar.xz
cp $(AMALGAMATED_FWD_FILE) release_files cp $(AMALGAMATED_FWD_FILE) release_files
cp $(AMALGAMATED_LITERALS_FILE) release_files cp $(AMALGAMATED_LITERALS_FILE) release_files
mv $(AMALGAMATED_FILE).asc $(AMALGAMATED_FWD_FILE).asc $(AMALGAMATED_LITERALS_FILE).asc json.tar.xz json.tar.xz.asc include.zip include.zip.asc release_files mv $(AMALGAMATED_FILE).asc $(AMALGAMATED_FWD_FILE).asc $(AMALGAMATED_LITERALS_FILE).asc json.tar.xz json.tar.xz.asc include.zip include.zip.asc release_files
cd release_files ; shasum -a 256 json.hpp include.zip json.tar.xz > hashes.txt cd release_files ; shasum -a 256 $$(find . -type f -not -name '*.asc' | sed 's|^\./||' | sort) > hashes.txt
########################################################################## ##########################################################################
@@ -256,7 +206,6 @@ release: include.zip json.tar.xz
# clean up # clean up
clean: clean:
rm -fr fuzz fuzz-testing *.dSYM tests/*.dSYM rm -fr fuzz fuzz-testing *.dSYM tests/*.dSYM
rm -fr benchmarks/files/numbers/*.json
rm -fr cmake-build-benchmarks fuzz-testing cmake-build-pvs-studio release_files rm -fr cmake-build-benchmarks fuzz-testing cmake-build-pvs-studio release_files
$(MAKE) clean -Cdocs $(MAKE) clean -Cdocs
+2 -2
View File
@@ -496,7 +496,7 @@ bool key(string_t& val);
bool parse_error(std::size_t position, const std::string& last_token, const detail::exception& ex); bool parse_error(std::size_t position, const std::string& last_token, const detail::exception& ex);
``` ```
The return value of each function determines whether parsing should proceed. The return value of each function determines whether parsing should proceed. For `parse_error`, returning `true` [recovers from the error](https://json.nlohmann.me/features/parsing/error_recovery/): the parser repairs the input and continues.
To implement your own SAX handler, proceed as follows: To implement your own SAX handler, proceed as follows:
@@ -504,7 +504,7 @@ To implement your own SAX handler, proceed as follows:
2. Create an object of your SAX interface class, e.g. `my_sax`. 2. Create an object of your SAX interface class, e.g. `my_sax`.
3. Call `bool json::sax_parse(input, &my_sax)`; where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface. 3. Call `bool json::sax_parse(input, &my_sax)`; where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface.
Note the `sax_parse` function only returns a `bool` indicating the result of the last executed SAX event. It does not return a `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error -- it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file [`json_sax.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/json_sax.hpp). Note the `sax_parse` function only returns a `bool` indicating whether the input was parsed without errors and no SAX event returned `false`. It does not return a `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error -- it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file [`json_sax.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/json_sax.hpp).
### STL-like access ### STL-like access
+14
View File
@@ -276,6 +276,20 @@ add_custom_target(ci_test_disableenumserialization
COMMENT "Compile and test with enum serialization disabled" COMMENT "Compile and test with enum serialization disabled"
) )
###############################################################################
# Disable conversion from a one-element tuple of a JSON reference.
###############################################################################
add_custom_target(ci_test_disabletuplereferenceconversion
COMMAND ${CMAKE_COMMAND}
-DCMAKE_BUILD_TYPE=Debug -GNinja
-DJSON_BuildTests=ON -DJSON_FastTests=ON -DJSON_DisableTupleReferenceConversion=ON
-S${PROJECT_SOURCE_DIR} -B${PROJECT_BINARY_DIR}/build_disabletuplereferenceconversion
COMMAND ${CMAKE_COMMAND} --build ${PROJECT_BINARY_DIR}/build_disabletuplereferenceconversion
COMMAND cd ${PROJECT_BINARY_DIR}/build_disabletuplereferenceconversion && ${CMAKE_CTEST_COMMAND} --parallel ${N} --output-on-failure
COMMENT "Compile and test with tuple reference conversion disabled"
)
############################################################################### ###############################################################################
# Skip the multiple-inclusion library version check. # Skip the multiple-inclusion library version check.
############################################################################### ###############################################################################
+16 -5
View File
@@ -11,20 +11,31 @@ EXAMPLES = $(wildcard mkdocs/docs/examples/*.cpp)
cxx_standard = $(lastword c++11 $(filter c++%, $(subst ., ,$1))) cxx_standard = $(lastword c++11 $(filter c++%, $(subst ., ,$1)))
# common compile flags for the stand-alone example files
EXAMPLE_CPPFLAGS = -I $(SRCDIR) -DJSON_USE_GLOBAL_UDLS=0
EXAMPLE_WARNFLAGS = -Werror=deprecated-declarations
# examples that document deprecated API and are allowed to use it
DEPRECATED_EXAMPLES = $(addprefix mkdocs/docs/examples/, \
json_pointer__operator__equal_stringtype \
json_pointer__operator__notequal_stringtype \
json_pointer__operator_string_t)
$(DEPRECATED_EXAMPLES:=.output) $(DEPRECATED_EXAMPLES:=.test): EXAMPLE_WARNFLAGS = -Wno-deprecated-declarations
# create output from a stand-alone example file # create output from a stand-alone example file
%.output: %.cpp %.output: %.cpp
@echo "standard $(call cxx_standard $(<:.cpp=))" @echo "standard $(call cxx_standard,$(<:.cpp=))"
$(MAKE) $(<:.cpp=) \ $(MAKE) $(<:.cpp=) \
CPPFLAGS="-I $(SRCDIR) -DJSON_USE_GLOBAL_UDLS=0" \ CPPFLAGS="$(EXAMPLE_CPPFLAGS)" \
CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) -Wno-deprecated-declarations" CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) $(EXAMPLE_WARNFLAGS)"
./$(<:.cpp=) > $@ ./$(<:.cpp=) > $@
rm $(<:.cpp=) rm $(<:.cpp=)
# compare created output with current output of the example files # compare created output with current output of the example files
%.test: %.cpp %.test: %.cpp
$(MAKE) $(<:.cpp=) \ $(MAKE) $(<:.cpp=) \
CPPFLAGS="-I $(SRCDIR) -DJSON_USE_GLOBAL_UDLS=0" \ CPPFLAGS="$(EXAMPLE_CPPFLAGS)" \
CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) -Wno-deprecated-declarations" CXXFLAGS="-std=$(call cxx_standard,$(<:.cpp=)) $(EXAMPLE_WARNFLAGS)"
./$(<:.cpp=) > $@ ./$(<:.cpp=) > $@
diff $@ $(<:.cpp=.output) diff $@ $(<:.cpp=.output)
rm $(<:.cpp=) $@ rm $(<:.cpp=) $@
+1 -1
View File
@@ -13,7 +13,7 @@
<key>dashIndexFilePath</key> <key>dashIndexFilePath</key>
<string>index.html</string> <string>index.html</string>
<key>DashDocSetFallbackURL</key> <key>DashDocSetFallbackURL</key>
<string>https://nlohmann.github.io/json/</string> <string>https://json.nlohmann.me/</string>
<key>isJavaScriptEnabled</key> <key>isJavaScriptEnabled</key>
<true/> <true/>
</dict> </dict>
+12 -22
View File
@@ -52,35 +52,25 @@ install_docset_zeal: JSON_for_Modern_C++.docset
mkdir -p $$docset_root; \ mkdir -p $$docset_root; \
cp -r JSON_for_Modern_C++.docset $$docset_root/ cp -r JSON_for_Modern_C++.docset $$docset_root/
# both targets below compare the docset search index with the mkdocs page
# set. They share the same normalization (docs/foo/index.md and
# docs/foo.md both become foo/index.html, the URL mkdocs itself would
# give the page; the top-level index.md is excluded, as it is not part
# of the hand-curated docSet.sql) and use comm(1) on two sorted lists
# instead of running a sqlite3 query, or an O(n*m) nested shell loop,
# once per page.
DOCSET_INDEX_PATHS=$(shell sqlite3 docSet.dsidx "SELECT DISTINCT path FROM searchIndex" | sort)
DOCSET_PAGE_PATHS=$(shell echo '$(MKDOCS_PAGES)' | tr ' ' '\n' | grep -v '^index\.md$$' | $(SED) -E 's@/index\.md$$@/index.html@; s@\.md$$@/index.html@' | sort)
# list mkdocs pages missing from the docset index # list mkdocs pages missing from the docset index
.PHONY: list_missing_pages .PHONY: list_missing_pages
list_missing_pages: docSet.dsidx list_missing_pages: docSet.dsidx
@for page in $(MKDOCS_PAGES); do \ @comm -23 <(echo '$(DOCSET_PAGE_PATHS)' | tr ' ' '\n') <(echo '$(DOCSET_INDEX_PATHS)' | tr ' ' '\n')
case "$$page" in \
*/index.md) path=$${page/\/index.md/} ;; \
*) path=$${page/.md/} ;; \
esac; \
if [ "x$$page" != "xindex.md" -a "x$$(sqlite3 docSet.dsidx "SELECT COUNT(*) FROM searchIndex WHERE path='$$path/index.html'")" = "x0" ]; then \
echo $$page; \
fi \
done
# list paths in the docset index without a corresponding mkdocs page # list paths in the docset index without a corresponding mkdocs page
.PHONY: list_removed_paths .PHONY: list_removed_paths
list_removed_paths: docSet.dsidx list_removed_paths: docSet.dsidx
@for path in $$(sqlite3 docSet.dsidx "SELECT path FROM searchIndex"); do \ @comm -13 <(echo '$(DOCSET_PAGE_PATHS)' | tr ' ' '\n') <(echo '$(DOCSET_INDEX_PATHS)' | tr ' ' '\n')
page=$${path/\/index.html/.md}; \
page_index=$${path/index.html/index.md}; \
page_found=0; \
for p in $(MKDOCS_PAGES); do \
if [ "x$$p" = "x$$page" -o "x$$p" = "x$$page_index" ]; then \
page_found=1; \
fi \
done; \
if [ "x$$page_found" = "x0" ]; then \
echo $$path; \
fi \
done
.PHONY: clean .PHONY: clean
clean: clean:
+3 -2
View File
@@ -7,10 +7,11 @@ documentation browsers like [Dash](https://kapeli.com/dash), [Velocity](https://
The docset can be created with The docset can be created with
```sh ```sh
make nlohmann_json.docset make JSON_for_Modern_C++.docset
``` ```
The generated folder `nlohmann_json.docset` can then be opened in the documentation browser. The generated folder `JSON_for_Modern_C++.docset` can then be opened in the documentation browser. `make all` builds a
`JSON_for_Modern_C++.tgz` archive instead, and `make install_docset_zeal` installs the docset for Zeal directly.
A recent version is also part of the [Dash user contributions](https://github.com/Kapeli/Dash-User-Contributions/tree/master/docsets/JSON_for_Modern_C%2B%2B). A recent version is also part of the [Dash user contributions](https://github.com/Kapeli/Dash-User-Contributions/tree/master/docsets/JSON_for_Modern_C%2B%2B).
+1
View File
@@ -210,6 +210,7 @@ INSERT INTO searchIndex(name, type, path) VALUES ('JSON_CATCH_USER', 'Macro', 'a
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_DIAGNOSTICS', 'Macro', 'api/macros/json_diagnostics/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('JSON_DIAGNOSTICS', 'Macro', 'api/macros/json_diagnostics/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_DIAGNOSTIC_POSITIONS', 'Macro', 'api/macros/json_diagnostic_positions/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('JSON_DIAGNOSTIC_POSITIONS', 'Macro', 'api/macros/json_diagnostic_positions/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_DISABLE_ENUM_SERIALIZATION', 'Macro', 'api/macros/json_disable_enum_serialization/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('JSON_DISABLE_ENUM_SERIALIZATION', 'Macro', 'api/macros/json_disable_enum_serialization/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_DISABLE_TUPLE_REFERENCE_CONVERSION', 'Macro', 'api/macros/json_disable_tuple_reference_conversion/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_HAS_CPP_11', 'Macro', 'api/macros/json_has_cpp_11/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('JSON_HAS_CPP_11', 'Macro', 'api/macros/json_has_cpp_11/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_HAS_CPP_14', 'Macro', 'api/macros/json_has_cpp_11/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('JSON_HAS_CPP_14', 'Macro', 'api/macros/json_has_cpp_11/index.html');
INSERT INTO searchIndex(name, type, path) VALUES ('JSON_HAS_CPP_17', 'Macro', 'api/macros/json_has_cpp_11/index.html'); INSERT INTO searchIndex(name, type, path) VALUES ('JSON_HAS_CPP_17', 'Macro', 'api/macros/json_has_cpp_11/index.html');
@@ -159,6 +159,8 @@ basic_json(basic_json&& other) noexcept;
- `CompatibleType` is not `basic_json` (to avoid hijacking copy/move constructors), - `CompatibleType` is not `basic_json` (to avoid hijacking copy/move constructors),
- `CompatibleType` is not a different `basic_json` type (i.e. with different template arguments) - `CompatibleType` is not a different `basic_json` type (i.e. with different template arguments)
- `CompatibleType` is not a `basic_json` nested type (e.g., `json_pointer`, `iterator`, etc.) - `CompatibleType` is not a `basic_json` nested type (e.g., `json_pointer`, `iterator`, etc.)
- if [`JSON_DISABLE_TUPLE_REFERENCE_CONVERSION`](../macros/json_disable_tuple_reference_conversion.md) is defined
to `1`: `CompatibleType` is not a one-element `std::tuple` holding a reference to `basic_json`
- `json_serializer<U>` (with `U = uncvref_t<CompatibleType>`) has a `to_json(basic_json_t&, CompatibleType&&)` - `json_serializer<U>` (with `U = uncvref_t<CompatibleType>`) has a `to_json(basic_json_t&, CompatibleType&&)`
method method
@@ -51,7 +51,7 @@ range will yield over/underflow when used in a constructor. During deserializati
will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md). will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states: [RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are > Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are
> interoperable in the sense that implementations will agree exactly on their numeric values. > interoperable in the sense that implementations will agree exactly on their numeric values.
As this range is a subrange of the exactly supported range [INT64_MIN, INT64_MAX], this class's integer type is As this range is a subrange of the exactly supported range [INT64_MIN, INT64_MAX], this class's integer type is
@@ -52,7 +52,7 @@ when used in a constructor. During deserialization, too large or small integer n
as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md). as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states: [RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are > Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are
> interoperable in the sense that implementations will agree exactly on their numeric values. > interoperable in the sense that implementations will agree exactly on their numeric values.
As this range is a subrange (when considered in conjunction with the `number_integer_t` type) of the exactly supported As this range is a subrange (when considered in conjunction with the `number_integer_t` type) of the exactly supported
@@ -54,7 +54,7 @@ classDiagram
## Notes ## Notes
For an input with $n$ bytes, 1 is the index of the first character and $n+1$ is the index of the terminating null byte For an input with <i>n</i> bytes, 1 is the index of the first character and <i>n</i>+1 is the index of the terminating null byte
or the end of file. This also holds true when reading a byte vector for binary formats. or the end of file. This also holds true when reading a byte vector for binary formats.
## Examples ## Examples
+4 -1
View File
@@ -90,7 +90,9 @@ The SAX event lister must follow the interface of [`json_sax`](../json_sax/index
## Return value ## Return value
return value of the last processed SAX event `#!cpp true` if the input was parsed without errors and no SAX event returned `#!cpp false`; `#!cpp false` otherwise.
In particular, the result is `#!cpp false` for input with errors, even if the SAX parser recovered from all of them
(see [error recovery](../../features/parsing/error_recovery.md)).
## Exception safety ## Exception safety
@@ -138,6 +140,7 @@ A UTF-8 byte order mark is silently ignored.
- Ignoring comments via `ignore_comments` added in version 3.9.0. - Ignoring comments via `ignore_comments` added in version 3.9.0.
- Added `ignore_trailing_commas` in version 3.13.0. - Added `ignore_trailing_commas` in version 3.13.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0. - Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Recovering from parse errors (see [`parse_error`](../json_sax/parse_error.md)) added in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0. - Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- `JSON_PRECISE_STREAM_POSITION` added in version 3.13.0 to optionally leave a `#!cpp std::istream` positioned right - `JSON_PRECISE_STREAM_POSITION` added in version 3.13.0 to optionally leave a `#!cpp std::istream` positioned right
after the parsed value when `strict` is `#!cpp false`. after the parsed value when `strict` is `#!cpp false`.
+2 -1
View File
@@ -7,7 +7,8 @@ struct json_sax;
This class describes the SAX interface used by [sax_parse](../basic_json/sax_parse.md). Each function is called in This class describes the SAX interface used by [sax_parse](../basic_json/sax_parse.md). Each function is called in
different situations while the input is parsed. The boolean return value informs the parser whether to continue different situations while the input is parsed. The boolean return value informs the parser whether to continue
processing the input. processing the input; for [`parse_error`](parse_error.md), it decides whether to
[recover from the error](../../features/parsing/error_recovery.md).
## Template parameters ## Template parameters
+24 -1
View File
@@ -21,7 +21,14 @@ A parse error occurred.
## Return value ## Return value
Whether parsing should proceed (**must return `#!cpp false`**). Whether to recover from the error:
- `#!cpp false` stops parsing.
- `#!cpp true` recovers from the error: the error is repaired and parsing continues. If that is not possible, which
happens in the binary formats when the end of the item with the error is unknown, the value read so far is completed
and parsing stops. See [error recovery](../../features/parsing/error_recovery.md) for how errors are repaired.
Either way, [`sax_parse`](../basic_json/sax_parse.md) returns `#!cpp false`.
## Examples ## Examples
@@ -39,6 +46,22 @@ Whether parsing should proceed (**must return `#!cpp false`**).
--8<-- "examples/sax_parse.output" --8<-- "examples/sax_parse.output"
``` ```
??? example
The example below shows how a SAX parser recovers from errors.
```cpp
--8<-- "examples/sax_parse__error_recovery.cpp"
```
Output:
```
--8<-- "examples/sax_parse__error_recovery.output"
```
## Version history ## Version history
- Added in version 3.2.0. - Added in version 3.2.0.
- Returning `#!cpp true` recovers from the error since version 3.13.0; before, parsing stopped, but the result of
[`sax_parse`](../basic_json/sax_parse.md) could be wrong.
+1
View File
@@ -53,6 +53,7 @@ header. See also the [macro overview page](../../features/macros.md).
- [**JSON_BRACE_INIT_COPY_SEMANTICS**](json_brace_init_copy_semantics.md) - opt in to copy/move semantics for single-element brace initialization - [**JSON_BRACE_INIT_COPY_SEMANTICS**](json_brace_init_copy_semantics.md) - opt in to copy/move semantics for single-element brace initialization
- [**JSON_DISABLE_ENUM_SERIALIZATION**](json_disable_enum_serialization.md) - switch off default serialization/deserialization functions for enums - [**JSON_DISABLE_ENUM_SERIALIZATION**](json_disable_enum_serialization.md) - switch off default serialization/deserialization functions for enums
- [**JSON_DISABLE_TUPLE_REFERENCE_CONVERSION**](json_disable_tuple_reference_conversion.md) - switch off conversion from a one-element tuple of a JSON reference
- [**JSON_USE_IMPLICIT_CONVERSIONS**](json_use_implicit_conversions.md) - control implicit conversions - [**JSON_USE_IMPLICIT_CONVERSIONS**](json_use_implicit_conversions.md) - control implicit conversions
## Comparison behavior ## Comparison behavior
@@ -0,0 +1,114 @@
# JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
```cpp
#define JSON_DISABLE_TUPLE_REFERENCE_CONVERSION /* value */
```
When defined to `1`, a `basic_json` value can no longer be constructed from a one-element `std::tuple` whose element is
a reference to that `basic_json` type, such as `std::tuple<json&>`, `std::tuple<const json&>`, or `std::tuple<json&&>`.
These are the tuples created by `std::forward_as_tuple(j)`.
## Default definition
The default value is `0` (disabled — existing behavior is preserved).
```cpp
#define JSON_DISABLE_TUPLE_REFERENCE_CONVERSION 0
```
## Notes
!!! note "Background"
By default, `basic_json` can be constructed from any `std::tuple` whose elements can be converted to JSON; the result
is an array. This includes `std::tuple<json&>`, which becomes a one-element array.
`std::tuple` only converts another tuple element by element if its element type cannot be constructed from the whole
source tuple. Because `json` *can* be constructed from `std::tuple<json&>`, `std::tuple` instead converts the whole
tuple into a single `json` value. This has two surprising effects:
```cpp
json j = true;
// rejected by some standard libraries (e.g., libc++); with others, the
// reference binds to a temporary that is destroyed right away
std::tuple<const json&> t1(std::forward_as_tuple(j));
// compiles, but std::get<0>(t2) is [true], not true
std::tuple<json> t2(std::forward_as_tuple(j));
```
Enabling this macro removes the conversion, so both tuples are converted element by element: `std::get<0>(t1)`
refers to `j`, and `std::get<0>(t2)` is a copy of `j` (see [#2226](https://github.com/nlohmann/json/issues/2226)).
!!! warning "Opt-in only"
This macro must be defined **before** including `<nlohmann/json.hpp>`. Defining it after the include has no effect.
!!! note "Affected conversions"
Only one-element tuples holding a reference to the **same** `basic_json` type are affected. Constructing a JSON value
from them no longer compiles:
```cpp
json j = true;
json a = std::forward_as_tuple(j); // error with the macro enabled
json b = json::array({j}); // use this instead: [true]
```
Tuples holding a JSON value (`std::make_tuple(j)`), tuples with more than one element, and tuples holding references
to other types (including other `basic_json` specializations) are converted to arrays as before.
!!! hint "CMake option"
This behavior can also be controlled with the CMake option
[`JSON_DisableTupleReferenceConversion`](../../integration/cmake.md#json_disabletuplereferenceconversion)
(`OFF` by default) which defines `JSON_DISABLE_TUPLE_REFERENCE_CONVERSION` accordingly.
## Examples
??? example "Default behavior (macro not defined)"
```cpp
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
json j = true;
std::tuple<json> t(std::forward_as_tuple(j));
// std::get<0>(t) is [true] -- the whole tuple was converted
}
```
??? example "Conversion disabled (macro defined to 1)"
```cpp
#define JSON_DISABLE_TUPLE_REFERENCE_CONVERSION 1
#include <nlohmann/json.hpp>
using json = nlohmann::json;
int main()
{
json j = true;
std::tuple<json> t(std::forward_as_tuple(j));
// std::get<0>(t) is true -- a copy of j
std::tuple<const json&> r(std::forward_as_tuple(j));
// std::get<0>(r) refers to j
}
```
## See also
- [**basic_json(CompatibleType&&)**](../basic_json/basic_json.md) - the affected constructor
- [:simple-cmake: JSON_DisableTupleReferenceConversion](../../integration/cmake.md#json_disabletuplereferenceconversion) -
CMake option to control the macro
## Version history
- Added in version 3.13.0.
@@ -1 +0,0 @@
<a target="_blank" href="https://wandbox.org/permlink/hUJYo1HWmfTBLMGn"><b>online</b></a>
@@ -1 +0,0 @@
<a target="_blank" href="https://wandbox.org/permlink/AWbpa8e1xRV3y4MM"><b>online</b></a>
@@ -0,0 +1,43 @@
#include <iostream>
#include <iomanip>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
// a SAX parser that creates a JSON value like json::parse does, but that
// recovers from parse errors instead of stopping at the first one
class recovering_parser : public nlohmann::detail::json_sax_dom_parser<json>
{
public:
explicit recovering_parser(json& result)
: nlohmann::detail::json_sax_dom_parser<json>(result, false)
{}
bool parse_error(std::size_t position,
const std::string& /*last_token*/,
const json::exception& ex)
{
std::cout << "byte " << position << ": " << ex.what() << '\n';
// repair the input and continue
return true;
}
};
int main()
{
// JSON text with several mistakes that ends too early
const std::string text = R"({
"name": "Hello World",
"tags": ["a" "b",],
"valid": tru,
"size": 1.,
"nested": {"x": 1)";
json result;
recovering_parser sax(result);
const bool valid = json::sax_parse(text, &sax);
std::cout << "\nvalid JSON: " << std::boolalpha << valid << '\n'
<< std::setw(4) << result << std::endl;
}
@@ -0,0 +1,19 @@
byte 49: [json.exception.parse_error.101] parse error at line 3, column 20: syntax error while parsing array - unexpected string literal; expected ']'
byte 51: [json.exception.parse_error.101] parse error at line 3, column 22: syntax error while parsing value - unexpected ']'; expected '[', '{', or a literal
byte 70: [json.exception.parse_error.101] parse error at line 4, column 17: syntax error while parsing value - invalid literal; last read: '"valid": tru,'
byte 86: [json.exception.parse_error.101] parse error at line 5, column 15: syntax error while parsing value - invalid number; expected digit after '.'; last read: '1.,'
byte 109: [json.exception.parse_error.101] parse error at line 6, column 22: syntax error while parsing object - unexpected end of input; expected '}'
valid JSON: false
{
"name": "Hello World",
"nested": {
"x": 1
},
"size": 1,
"tags": [
"a",
"b"
],
"valid": null
}
@@ -61,7 +61,7 @@ The library uses the following mapping from JSON values types to BJData types ac
The following values can **not** be converted to a BJData value: The following values can **not** be converted to a BJData value:
- strings with more than 18446744073709551615 bytes, i.e., $2^{64}-1$ bytes (theoretical) - strings with more than 18446744073709551615 bytes, i.e., 2<sup>64</sup>-1 bytes (theoretical)
!!! info "Unused BJData markers" !!! info "Unused BJData markers"
+7
View File
@@ -83,6 +83,13 @@ When defined, default parse and serialize functions for enums are excluded and h
See [full documentation of `JSON_DISABLE_ENUM_SERIALIZATION`](../api/macros/json_disable_enum_serialization.md). See [full documentation of `JSON_DISABLE_ENUM_SERIALIZATION`](../api/macros/json_disable_enum_serialization.md).
## `JSON_DISABLE_TUPLE_REFERENCE_CONVERSION`
When defined to `1`, a JSON value can no longer be created from a one-element `std::tuple` holding a reference to a JSON
value, such as the result of `std::forward_as_tuple(j)`. This lets `std::tuple` convert such tuples element-wise.
See [full documentation of `JSON_DISABLE_TUPLE_REFERENCE_CONVERSION`](../api/macros/json_disable_tuple_reference_conversion.md).
## `JSON_NO_AUTOMATIC_UDLS` ## `JSON_NO_AUTOMATIC_UDLS`
When defined, `<nlohmann/json.hpp>` does not include `<nlohmann/json_literals.hpp>` with the user-defined string literals When defined, `<nlohmann/json.hpp>` does not include `<nlohmann/json_literals.hpp>` with the user-defined string literals
@@ -0,0 +1,121 @@
# Error Recovery
By default, parsing stops at the first error. With the [SAX interface](sax_interface.md), you can instead ask the
parser to *recover*: to repair the error and continue, so that you get as much as possible out of malformed input, for
instance a file that was cut off, JSON edited by hand, or the output of a language model.
## Recovering from errors
The SAX parser's [`parse_error`](../../api/json_sax/parse_error.md) function is called for every error. Its return value
decides what happens next:
- `#!cpp false` stops parsing. This is what the SAX parsers of the library do, so [`parse`](../../api/basic_json/parse.md)
and [`accept`](../../api/basic_json/accept.md) never recover.
- `#!cpp true` repairs the error and continues parsing.
When recovering, the SAX parser still receives well-formed events: every `start_object` or `start_array` is followed by
the matching `end_object` or `end_array`, and every `key` is followed by exactly one value. A SAX parser that creates a
JSON value, such as the one in the example below, therefore gets a complete value. Parsing always ends, and
[`sax_parse`](../../api/basic_json/sax_parse.md) returns `#!cpp false` for input that is not valid JSON, even if every
error was repaired. Each token is reported at most once, and the SAX parser can stop at any error by returning
`#!cpp false`.
!!! example
The example below derives a SAX parser from the library's parser for `json` values (`json_sax_dom_parser`),
and recovers from all errors.
```cpp
--8<-- "examples/sax_parse__error_recovery.cpp"
```
Output:
```
--8<-- "examples/sax_parse__error_recovery.output"
```
## How errors are repaired
Each error is repaired with the smallest local edit: a missing separator is inserted, a stray token is removed, what can
be read of a broken string or number is kept, and a value that cannot be read at all becomes `#!json null`.
| Mistake | Repair | Example | Result |
|---------------------------|--------------------------------------------------------------------------------|------------------------------------------|----------------------------|
| missing `,` or `:` | inserted | `#!json [1 2]`, `#!json {"a" 1}` | `[1,2]`, `{"a":1}` |
| missing value | `#!json null` for an object key or between commas in an array | `#!json {"a":}`, `#!json [1,,2]` | `{"a":null}`, `[1,null,2]` |
| trailing comma | removed | `#!json [1,2,]` | `[1,2]` |
| broken string | invalid escapes and bytes are replaced (see below); a line break ends the string | `#!json ["a\qb"]` | `["aqb"]` |
| broken number | the longest valid beginning is kept | `#!json [1., 2e+]` | `[1,2]` |
| unreadable value | `#!json null` | `#!json [1, NaN, tru]` | `[1,null,null]` |
| number too large | passed as infinity, together with its text | `#!json [1e999]` | infinity (see below) |
| stray `:` | removed | `#!json ["a":1]` | `["a",1]` |
| member without a key | skipped up to the next `,` or `}` | `#!json {1:2, "b":3}` | `{"b":3}` |
| wrong closing bracket | closes the innermost array or object | `#!json {"a":[1,2}, "b":3}` | `{"a":[1,2],"b":3}` |
| input ends too early | all open arrays and objects are closed | `#!json {"a":[1,2` | `{"a":[1,2]}` |
| text before the value | skipped | `#!json )]}'{"a":1}` | `{"a":1}` |
In a string, an unknown escape like `\q` stands for the escaped character (`q`), as in JavaScript. An invalid `\u`
escape, a lone surrogate, and ill-formed UTF-8 are each replaced by U+FFFD (REPLACEMENT CHARACTER), and control
characters are kept. A string without its closing quote ends at the next line break or at the end of the input.
The input after the top-level value is not repaired: as without recovery, it is reported as an error, and parsing stops.
## Binary formats
The binary formats ([BJData](../binary_formats/bjdata.md), [BON8](../binary_formats/bon8.md),
[BSON](../binary_formats/bson.md), [CBOR](../binary_formats/cbor.md), [MessagePack](../binary_formats/messagepack.md),
and [UBJSON](../binary_formats/ubjson.md)) have no delimiters to find the next value by. So what can be repaired depends
on whether the end of the item with the error is known, a distinction that
[RFC 8949, Section 5.3](https://www.rfc-editor.org/rfc/rfc8949.html#section-5.3) makes for CBOR, too.
If the item is complete, but cannot be passed on as it is, it is replaced, and parsing continues after it:
| Mistake | Formats | Repair |
|---------------------------------------------------------------------|-----------------------------------------|-------------------------------------------------------------------------|
| tag | CBOR | ignored |
| simple value other than `false`, `true`, and `null`, like undefined | CBOR | `#!json null` |
| negative integer below the range of `number_integer_t` | CBOR | the nearest floating-point number |
| string that is not valid UTF-8 | BJData, BSON, CBOR, MessagePack, UBJSON | each ill-formed sequence becomes U+FFFD |
| character (`C`) that is not ASCII | BJData, UBJSON | U+FFFD |
| invalid high-precision number (`H`) | BJData, UBJSON | the longest valid beginning is kept, as for JSON text, or `#!json null` |
| high-precision number too large | BJData, UBJSON | passed as infinity, together with its text |
| object key that is not a string | BON8, CBOR, MessagePack | the member is skipped |
| element of a type the library does not read, like ObjectId or date | BSON | `#!json null` |
| string without its terminator | BSON | kept |
| document whose size does not match its content | BSON | kept |
CBOR tags and simple values are repaired as [RFC 8949, Section 6.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-6.1)
suggests for converting CBOR to JSON. Note that [`sax_parse`](../../api/basic_json/sax_parse.md) has no parameter for
CBOR tags, so every tag is an error there; when recovering, tags are ignored like with
[`cbor_tag_handler_t::ignore`](../../api/basic_json/cbor_tag_handler_t.md).
After any other error, the end of the item is unknown: the input ended, a byte is not a valid type marker, or a size
cannot be right. Parsing then stops, and the value read so far is completed: a key that waits for its value gets
`#!json null`, and all open arrays and objects are closed. This keeps everything before the error of an input that was
cut off. The exception is BSON, which stores the size of every document: an element whose end is unknown gets
`#!json null`, the rest of its document is skipped, and parsing continues after the document.
## Limitations
- A repair is a guess. For example, `#!json {"a" "b": 1}` could be meant as `#!json {"a": "b"}` or as
`#!json {"a": null, "b": 1}`; it is repaired to the former. Treat recovered values as a best effort, and check the
reported errors.
- A closing bracket always closes the innermost array or object. If a bracket is missing rather than wrong, the
repair differs from the intention: `#!json {"a": {"b": [1, 2}, "c": 3}` is repaired to
`#!json {"a": {"b": [1, 2], "c": 3}}`, although `#!json {"a": {"b": [1, 2]}, "c": 3}` may have been meant.
- Keys without quotes, and strings in single quotes, are not supported; such members are skipped.
- In the binary formats, a member that is skipped because its key is not a string is lost, and so are the elements of a
BSON document after one whose end is unknown.
- A number that is too large for `number_float_t` is passed as positive or negative infinity. The SAX parser's
`number_float` also gets the number's text, but a JSON value cannot store it, and
[`dump`](../../api/basic_json/dump.md) serializes infinity as `#!json null`.
- When parsing is not strict (see [`sax_parse`](../../api/basic_json/sax_parse.md)), a repair may read parts of the
input after the value, for instance of the next value in a stream of concatenated values.
## See also
- [SAX interface](sax_interface.md) - implement a custom SAX handler
- [`parse_error`](../../api/json_sax/parse_error.md) - the SAX event for parse errors
- [`sax_parse`](../../api/basic_json/sax_parse.md) - generate SAX events
- [parsing and exceptions](parse_exceptions.md) - control error handling
+2 -1
View File
@@ -65,7 +65,7 @@ You can influence a DOM parse without switching to the SAX interface by passing
When the input is not valid JSON, the `parse` function throws an exception by default. If exceptions are undesired or When the input is not valid JSON, the `parse` function throws an exception by default. If exceptions are undesired or
unavailable, the parser can instead return a discarded value, or [`accept`](../../api/basic_json/accept.md) can be used unavailable, the parser can instead return a discarded value, or [`accept`](../../api/basic_json/accept.md) can be used
to only check whether an input is valid JSON. See [parsing and exceptions](parse_exceptions.md) for the available to only check whether an input is valid JSON. See [parsing and exceptions](parse_exceptions.md) for the available
options. options. To get as much as possible out of malformed input, a SAX parser can [recover from errors](error_recovery.md).
## See also ## See also
@@ -76,3 +76,4 @@ options.
- [parser callbacks](parser_callbacks.md) - influence the parsing by a callback function - [parser callbacks](parser_callbacks.md) - influence the parsing by a callback function
- [SAX interface](sax_interface.md) - implement a custom SAX handler - [SAX interface](sax_interface.md) - implement a custom SAX handler
- [parsing and exceptions](parse_exceptions.md) - control error handling - [parsing and exceptions](parse_exceptions.md) - control error handling
- [error recovery](error_recovery.md) - get as much as possible out of malformed input
@@ -64,7 +64,8 @@ bool parse_error(std::size_t position,
const json::exception& ex); const json::exception& ex);
``` ```
The return value indicates whether the parsing should continue, so the function should usually return `#!cpp false`. The return value decides whether to stop parsing (`#!cpp false`) or to repair the error and continue
(`#!cpp true`); see [error recovery](error_recovery.md) for the latter.
??? example ??? example
@@ -60,7 +60,8 @@ bool key(string_t& val);
bool parse_error(std::size_t position, const std::string& last_token, const json::exception& ex); bool parse_error(std::size_t position, const std::string& last_token, const json::exception& ex);
``` ```
The return value of each function determines whether parsing should proceed. The return value of each function determines whether parsing should proceed. For `parse_error`, returning
`#!cpp true` [recovers from the error](error_recovery.md).
To implement your own SAX handler, proceed as follows: To implement your own SAX handler, proceed as follows:
@@ -68,7 +69,7 @@ To implement your own SAX handler, proceed as follows:
2. Create an object of your SAX interface class, e.g. `my_sax`. 2. Create an object of your SAX interface class, e.g. `my_sax`.
3. Call `#!cpp bool json::sax_parse(input, &my_sax);` where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface. 3. Call `#!cpp bool json::sax_parse(input, &my_sax);` where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface.
Note the `sax_parse` function only returns a `#!cpp bool` indicating the result of the last executed SAX event. It does not return `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error - it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file `json_sax.hpp`. Note the `sax_parse` function only returns a `#!cpp bool` indicating whether the input was parsed without errors and no SAX event returned `#!cpp false`. It does not return `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error - it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file `json_sax.hpp`.
## See also ## See also
+1 -1
View File
@@ -273,7 +273,7 @@ When the default type is used, the maximal unsigned integer number that can be s
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states: [RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are interoperable in the sense that implementations will agree exactly on their numeric values. > Note that when such software is used, numbers that are integers and are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are interoperable in the sense that implementations will agree exactly on their numeric values.
As this range is a subrange of the exactly supported range [`INT64_MIN`, `INT64_MAX`], this class's integer type is interoperable. As this range is a subrange of the exactly supported range [`INT64_MIN`, `INT64_MAX`], this class's integer type is interoperable.
@@ -48,7 +48,7 @@ On number interoperability, the following remarks are made:
for numeric magnitude and precision than is widely available. for numeric magnitude and precision than is widely available.
Note that when such software is used, numbers that are integers and Note that when such software is used, numbers that are integers and
are in the range $[-2^{53}+1, 2^{53}-1]$ are interoperable in the are in the range [-2<sup>53</sup>+1, 2<sup>53</sup>-1] are interoperable in the
sense that implementations will agree exactly on their numeric sense that implementations will agree exactly on their numeric
values. values.
@@ -95,9 +95,9 @@ This is the same behavior as the code `#!c double x = 3.141592653589793238462643
!!! success "Interoperability" !!! success "Interoperability"
- The library is interoperable with respect to the specification, because its supported range $[-2^{63}, 2^{64}-1]$ is - The library is interoperable with respect to the specification, because its supported range [-2<sup>63</sup>, 2<sup>64</sup>-1] is
larger than the described range $[-2^{53}+1, 2^{53}-1]$. larger than the described range [-2<sup>53</sup>+1, 2<sup>53</sup>-1].
- All integers outside the range $[-2^{63}, 2^{64}-1]$, as well as floating-point numbers are stored as `double`. - All integers outside the range [-2<sup>63</sup>, 2<sup>64</sup>-1], as well as floating-point numbers are stored as `double`.
This also concurs with the specification above. This also concurs with the specification above.
### Zeros ### Zeros
+6
View File
@@ -169,6 +169,12 @@ Enable position diagnostics by defining macro [`JSON_DIAGNOSTIC_POSITIONS`](../a
Disable default `enum` serialization by defining the macro Disable default `enum` serialization by defining the macro
[`JSON_DISABLE_ENUM_SERIALIZATION`](../api/macros/json_disable_enum_serialization.md). This option is `OFF` by default. [`JSON_DISABLE_ENUM_SERIALIZATION`](../api/macros/json_disable_enum_serialization.md). This option is `OFF` by default.
### `JSON_DisableTupleReferenceConversion`
Disable the conversion from a one-element `std::tuple` holding a reference to a JSON value by defining the macro
[`JSON_DISABLE_TUPLE_REFERENCE_CONVERSION`](../api/macros/json_disable_tuple_reference_conversion.md). This option is
`OFF` by default.
### `JSON_FastTests` ### `JSON_FastTests`
Skip expensive/slow test suites. This option is `OFF` by default. Depends on `JSON_BuildTests`. Skip expensive/slow test suites. This option is `OFF` by default. Depends on `JSON_BuildTests`.
@@ -678,11 +678,11 @@ to install the [nlohmann-json](https://ports.macports.org/port/nlohmann-json/) p
1. Create the following files: 1. Create the following files:
```cpp title="example.cpp" ```cpp title="example.cpp"
--8<-- "integration/homebrew/example.cpp" --8<-- "integration/macports/example.cpp"
``` ```
```cmake title="CMakeLists.txt" ```cmake title="CMakeLists.txt"
--8<-- "integration/homebrew/CMakeLists.txt" --8<-- "integration/macports/CMakeLists.txt"
``` ```
2. Install the package: 2. Install the package:
+2 -4
View File
@@ -87,6 +87,7 @@ nav:
- features/object_order.md - features/object_order.md
- Parsing: - Parsing:
- features/parsing/index.md - features/parsing/index.md
- features/parsing/error_recovery.md
- features/parsing/json_lines.md - features/parsing/json_lines.md
- features/parsing/parse_exceptions.md - features/parsing/parse_exceptions.md
- features/parsing/parser_callbacks.md - features/parsing/parser_callbacks.md
@@ -287,6 +288,7 @@ nav:
- 'JSON_DIAGNOSTICS': api/macros/json_diagnostics.md - 'JSON_DIAGNOSTICS': api/macros/json_diagnostics.md
- 'JSON_DIAGNOSTIC_POSITIONS': api/macros/json_diagnostic_positions.md - 'JSON_DIAGNOSTIC_POSITIONS': api/macros/json_diagnostic_positions.md
- 'JSON_DISABLE_ENUM_SERIALIZATION': api/macros/json_disable_enum_serialization.md - 'JSON_DISABLE_ENUM_SERIALIZATION': api/macros/json_disable_enum_serialization.md
- 'JSON_DISABLE_TUPLE_REFERENCE_CONVERSION': api/macros/json_disable_tuple_reference_conversion.md
- 'JSON_HAS_CPP_11, JSON_HAS_CPP_14, JSON_HAS_CPP_17, JSON_HAS_CPP_20': api/macros/json_has_cpp_11.md - 'JSON_HAS_CPP_11, JSON_HAS_CPP_14, JSON_HAS_CPP_17, JSON_HAS_CPP_20': api/macros/json_has_cpp_11.md
- 'JSON_HAS_EXPERIMENTAL_FILESYSTEM, JSON_HAS_FILESYSTEM': api/macros/json_has_filesystem.md - 'JSON_HAS_EXPERIMENTAL_FILESYSTEM, JSON_HAS_FILESYSTEM': api/macros/json_has_filesystem.md
- 'JSON_HAS_RANGES': api/macros/json_has_ranges.md - 'JSON_HAS_RANGES': api/macros/json_has_ranges.md
@@ -350,7 +352,6 @@ markdown_extensions:
- toc: - toc:
permalink: true permalink: true
- md_in_html - md_in_html
- pymdownx.arithmatex
- pymdownx.betterem: - pymdownx.betterem:
smart_enable: all smart_enable: all
- pymdownx.caret - pymdownx.caret
@@ -442,6 +443,3 @@ plugins:
extra_css: extra_css:
- css/custom.css - css/custom.css
extra_javascript:
- https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.0/MathJax.js?config=TeX-MML-AM_CHTML
+5 -5
View File
@@ -79,9 +79,9 @@ def check_structure() -> None:
report("whitespace/line_length", f"{file}:{lineno+1} ({current_section})", f"line is too long ({len(line)} vs. 160 chars)") report("whitespace/line_length", f"{file}:{lineno+1} ({current_section})", f"line is too long ({len(line)} vs. 160 chars)")
# sections in `<!-- NOLINT -->` comments are treated as present # sections in `<!-- NOLINT -->` comments are treated as present
if line.startswith("<!-- NOLINT"): nolint_match = re.match(r"<!--\s*NOLINT\s+(.*?)\s*-->", line)
current_section = line.strip("<!-- NOLINT") if nolint_match:
current_section = current_section.strip(" -->") current_section = nolint_match.group(1)
existing_sections.append(current_section) existing_sections.append(current_section)
# check if sections are correct # check if sections are correct
@@ -97,7 +97,7 @@ def check_structure() -> None:
if len(unexpected): if len(unexpected):
report("style/numbering", f"{file}:{lineno} ({current_section})", f'unexpected overloads: {", ".join([f"({x})" for x in unexpected])}') report("style/numbering", f"{file}:{lineno} ({current_section})", f'unexpected overloads: {", ".join([f"({x})" for x in unexpected])}')
current_section = line.strip("## ") current_section = line[3:]
existing_sections.append(current_section) existing_sections.append(current_section)
if current_section in expected_sections: if current_section in expected_sections:
@@ -141,7 +141,7 @@ def check_structure() -> None:
# check that non-example admonitions have titles # check that non-example admonitions have titles
untitled_admonition = re.match(r"^(\?\?\?|!!!) ([^ ]+)$", line) untitled_admonition = re.match(r"^(\?\?\?|!!!) ([^ ]+)$", line)
if untitled_admonition and untitled_admonition.group(2) != "example": if untitled_admonition and untitled_admonition.group(2) != "example":
report("style/admonition_title", f"{file}:{lineno} ({current_section})", f'"{untitled_admonition.group(2)}" admonitions should have a title') report("style/admonition_title", f"{file}:{lineno+1} ({current_section})", f'"{untitled_admonition.group(2)}" admonitions should have a title')
previous_line = line previous_line = line
@@ -471,11 +471,13 @@ inline void to_json_tuple_impl(BasicJsonType& j, const Tuple& t, index_sequence<
j = { std::get<Idx>(t)... }; j = { std::get<Idx>(t)... };
} }
#if JSON_BRACE_INIT_COPY_SEMANTICS // A one-element braced list does not reliably wrap its element: with
// JSON_BRACE_INIT_COPY_SEMANTICS makes a one-element braced list copy its // JSON_BRACE_INIT_COPY_SEMANTICS it copies it, which would serialize
// element instead of wrapping it, which would serialize std::tuple<int>{5} as 5 // std::tuple<int>{5} as 5 rather than [5], and some compilers (e.g., Apple clang
// rather than [5]. Build what the default deduction builds instead: an object // 15 and 16) copy an element that is itself a basic_json even without it, so
// if the element is a [string, value] pair, a one-element array otherwise. // std::tuple<json>{true} became true rather than [true]. Build what the default
// deduction builds instead: an object if the element is a [string, value] pair,
// a one-element array otherwise.
template<typename BasicJsonType, typename Tuple> template<typename BasicJsonType, typename Tuple>
inline void to_json_tuple_impl(BasicJsonType& j, const Tuple& t, index_sequence<0> /*unused*/) inline void to_json_tuple_impl(BasicJsonType& j, const Tuple& t, index_sequence<0> /*unused*/)
{ {
@@ -493,7 +495,6 @@ inline void to_json_tuple_impl(BasicJsonType& j, const Tuple& t, index_sequence<
j = BasicJsonType::array({std::move(element)}); j = BasicJsonType::array({std::move(element)});
} }
} }
#endif
template<typename BasicJsonType, typename Tuple> template<typename BasicJsonType, typename Tuple>
inline void to_json_tuple_impl(BasicJsonType& j, const Tuple& /*unused*/, index_sequence<> /*unused*/) inline void to_json_tuple_impl(BasicJsonType& j, const Tuple& /*unused*/, index_sequence<> /*unused*/)
File diff suppressed because it is too large Load Diff
+9 -4
View File
@@ -131,7 +131,9 @@ struct json_sax
@param[in] position the position in the input where the error occurs @param[in] position the position in the input where the error occurs
@param[in] last_token the last read token @param[in] last_token the last read token
@param[in] ex an exception object describing the error @param[in] ex an exception object describing the error
@return whether parsing should proceed (must return false) @return whether to recover from the error: false stops parsing; true
repairs the error and continues, or, if that is not possible,
stops after completing the value read so far
*/ */
virtual bool parse_error(std::size_t position, virtual bool parse_error(std::size_t position,
const std::string& last_token, const std::string& last_token,
@@ -186,9 +188,12 @@ a pointer to the respective array or object for each recursion depth.
After successful parsing, the value that is passed by reference to the After successful parsing, the value that is passed by reference to the
constructor contains the parsed value. constructor contains the parsed value.
@tparam BasicJsonType the JSON type @tparam BasicJsonType the JSON type
@tparam InputAdapterType the input adapter of the lexer that can be passed to
the constructor to record diagnostic positions; it
does not matter if no lexer is passed
*/ */
template<typename BasicJsonType, typename InputAdapterType> template<typename BasicJsonType, typename InputAdapterType = string_input_adapter_type>
class json_sax_dom_parser class json_sax_dom_parser
{ {
public: public:
@@ -505,7 +510,7 @@ class json_sax_dom_parser
lexer_t* m_lexer_ref = nullptr; lexer_t* m_lexer_ref = nullptr;
}; };
template<typename BasicJsonType, typename InputAdapterType> template<typename BasicJsonType, typename InputAdapterType = string_input_adapter_type>
class json_sax_dom_callback_parser class json_sax_dom_callback_parser
{ {
public: public:
+590 -1
View File
@@ -10,6 +10,7 @@
#include <array> // array #include <array> // array
#include <cstddef> // size_t #include <cstddef> // size_t
#include <cstdint> // uint8_t
#include <cstdio> // snprintf #include <cstdio> // snprintf
#include <initializer_list> // initializer_list #include <initializer_list> // initializer_list
#include <string> // char_traits, string #include <string> // char_traits, string
@@ -437,8 +438,16 @@ class lexer : public lexer_base<BasicJsonType>
if (0xD800 <= codepoint1 && codepoint1 <= 0xDBFF) if (0xD800 <= codepoint1 && codepoint1 <= 0xDBFF)
{ {
// expect next \uxxxx entry // expect next \uxxxx entry
if (JSON_HEDLEY_LIKELY(get() == '\\' && get() == 'u')) if (JSON_HEDLEY_LIKELY(get() == '\\'))
{ {
if (JSON_HEDLEY_UNLIKELY(get() != 'u'))
{
// current is the character escaped by the backslash
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
string_error_resume = resume_kind::escaped_character;
return token_type::parse_error;
}
const int codepoint2 = get_codepoint(); const int codepoint2 = get_codepoint();
if (JSON_HEDLEY_UNLIKELY(codepoint2 == -1)) if (JSON_HEDLEY_UNLIKELY(codepoint2 == -1))
@@ -463,7 +472,11 @@ class lexer : public lexer_base<BasicJsonType>
} }
else else
{ {
// the second escape was read completely and is a
// code point of its own
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF"; error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
string_error_resume = resume_kind::after_escape;
string_error_codepoint = codepoint2;
return token_type::parse_error; return token_type::parse_error;
} }
} }
@@ -477,7 +490,9 @@ class lexer : public lexer_base<BasicJsonType>
{ {
if (JSON_HEDLEY_UNLIKELY(0xDC00 <= codepoint1 && codepoint1 <= 0xDFFF)) if (JSON_HEDLEY_UNLIKELY(0xDC00 <= codepoint1 && codepoint1 <= 0xDFFF))
{ {
// the escape was read completely
error_message = "invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF"; error_message = "invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF";
string_error_resume = resume_kind::after_escape;
return token_type::parse_error; return token_type::parse_error;
} }
} }
@@ -2130,6 +2145,573 @@ scan_number_done:
} }
} }
/////////////////////
// error recovery
/////////////////////
/*!
@brief make the best of the token that scan() rejected
Called by the parser after scan() returned token_type::parse_error and the
SAX parser asked to recover from the error (see #3989). Keeps what can be
read of the token and skips the rest:
- A string keeps its characters. An unknown escape stands for the escaped
character itself (as in JavaScript), an invalid `\u` escape and ill-formed
UTF-8 become U+FFFD, and a control character is kept. A line break or the
end of the input ends a string that lacks its closing quote.
- A number keeps its longest valid prefix, e.g. `1` for `1.` or `1e+`.
- A block comment that is not closed runs to the end of the input.
- Anything else is skipped.
The rest of an invalid token is skipped up to the next delimiter
(whitespace, a structural character, or a quote). A delimiter that the
invalid token consumed is returned to the input, so that the next scan()
reads it.
@return token_type::value_string or a number token type if a string or a
number could be read, token_type::end_of_input for a block comment
that is not closed, token_type::uninitialized otherwise
*/
token_type recover_token()
{
const resume_kind resume = string_error_resume;
const int codepoint = string_error_codepoint;
string_error_resume = resume_kind::character;
string_error_codepoint = -1;
if (error_message_starts_with("invalid string"))
{
return recover_string(resume, codepoint);
}
if (error_message_starts_with("invalid number"))
{
return recover_number();
}
if (error_message_starts_with("invalid comment; missing"))
{
// the comment runs to the end of the input
return token_type::end_of_input;
}
skip_to_delimiter();
return token_type::uninitialized;
}
/*!
@brief return the token that scan() read last to the input, so that the
next scan() reads it again
Called by the parser when recovering from an error. The token must be a
single character (',', ':', '[', ']', '{', or '}') or the end of the
input, and scan() must have read it last.
*/
void unget_token()
{
JSON_ASSERT(!next_unget);
unget();
}
/*!
@brief let the token string for the next error begin at the current character
The token string of an error reaches back to the beginning of the last
string or number. After an error, the parser calls this function so that
the next error does not report (and, with many errors, copy) everything
read since then.
*/
void restart_token_string()
{
restart_token_string_impl(std::integral_constant<bool, lazy_token_string> {});
}
private:
/// how recover_string() continues after the error scan_string() reported
enum class resume_kind : std::uint8_t
{
/// current is the next character of the string (or the end of input)
character,
/// current is the character escaped by the preceding backslash
escaped_character,
/// current is the last character of a complete escape
after_escape
};
/// whether error_message begins with @a prefix
bool error_message_starts_with(const char* prefix) const noexcept
{
const char* message = error_message;
while (*prefix != '\0')
{
if (*message++ != *prefix++)
{
return false;
}
}
return true;
}
/// whether current ends an invalid token (see recover_token())
bool current_is_delimiter() const noexcept
{
switch (current)
{
case ' ':
case '\t':
case '\n':
case '\r':
case '[':
case ']':
case '{':
case '}':
case ',':
case ':':
case '\"':
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
case char_traits<char_type>::eof():
return true;
case '/':
return ignore_comments;
default:
return false;
}
}
/// skip the rest of an invalid token and return its delimiter to the input
void skip_to_delimiter()
{
while (!current_is_delimiter())
{
get();
}
if (current != char_traits<char_type>::eof())
{
unget();
}
}
/// append U+FFFD REPLACEMENT CHARACTER to token_buffer
void add_replacement_character()
{
add(0xEF);
add(0xBF);
add(0xBD);
}
/// append the UTF-8 encoding of @a codepoint (not a surrogate) to token_buffer
void add_codepoint(const int codepoint)
{
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
const auto cp = static_cast<unsigned int>(codepoint);
if (cp < 0x80)
{
add(static_cast<char_int_type>(cp));
}
else if (cp <= 0x7FF)
{
add(static_cast<char_int_type>(0xC0u | (cp >> 6u)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
else if (cp <= 0xFFFF)
{
add(static_cast<char_int_type>(0xE0u | (cp >> 12u)));
add(static_cast<char_int_type>(0x80u | ((cp >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
else
{
add(static_cast<char_int_type>(0xF0u | (cp >> 18u)));
add(static_cast<char_int_type>(0x80u | ((cp >> 12u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | ((cp >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
}
/// append a code point read from a `\u` escape; a surrogate becomes U+FFFD
void add_escaped_codepoint(const int codepoint)
{
if (0xD800 <= codepoint && codepoint <= 0xDFFF)
{
add_replacement_character();
}
else
{
add_codepoint(codepoint);
}
}
/*!
@brief remove an incomplete UTF-8 sequence from the end of token_buffer
next_byte_in_range() adds the bytes of a sequence as it checks them, so
when it rejects a byte, the beginning of the sequence is already in
token_buffer, which otherwise holds only complete sequences.
@return whether an incomplete sequence was removed
*/
bool remove_incomplete_utf8_sequence()
{
std::size_t lead = token_buffer.size();
std::size_t continuation_bytes = 0;
while (lead > 0 && continuation_bytes < 3
&& (static_cast<unsigned char>(token_buffer[lead - 1]) & 0xC0u) == 0x80u)
{
--lead;
++continuation_bytes;
}
if (lead == 0)
{
return false;
}
const auto lead_byte = static_cast<unsigned char>(token_buffer[lead - 1]);
std::size_t expected = 0;
if (lead_byte >= 0xF0)
{
expected = 3;
}
else if (lead_byte >= 0xE0)
{
expected = 2;
}
else if (lead_byte >= 0xC0)
{
expected = 1;
}
if (continuation_bytes >= expected)
{
return false;
}
token_buffer.resize(lead - 1);
return true;
}
/*!
@brief read the UTF-8 sequence that begins with current, which is not ASCII
@return whether the next character must be read; false if current still
needs to be handled, because it does not belong to the sequence
*/
bool recover_utf8_sequence()
{
// the number of continuation bytes and the range of the first one;
// see the ranges in scan_string()
std::size_t count = 0;
char_int_type low = 0x80;
char_int_type high = 0xBF;
if (current >= 0xC2 && current <= 0xDF)
{
count = 1;
}
else if (current >= 0xE0 && current <= 0xEF)
{
count = 2;
low = (current == 0xE0) ? 0xA0 : 0x80;
high = (current == 0xED) ? 0x9F : 0xBF;
}
else if (current >= 0xF0 && current <= 0xF4)
{
count = 3;
low = (current == 0xF0) ? 0x90 : 0x80;
high = (current == 0xF4) ? 0x8F : 0xBF;
}
else
{
// an ill-formed byte
add_replacement_character();
return true;
}
const std::size_t start = token_buffer.size();
add(current);
for (std::size_t i = 0; i < count; ++i)
{
get();
if (current < low || current > high)
{
token_buffer.resize(start);
add_replacement_character();
return false;
}
add(current);
low = 0x80;
high = 0xBF;
}
return true;
}
/*!
@brief read the low surrogate that must follow the high surrogate @a high
@return whether the next character must be read; false if current still
needs to be handled
*/
bool recover_low_surrogate(int high)
{
while (true)
{
if (get() != '\\')
{
add_replacement_character();
return false;
}
if (get() != 'u')
{
add_replacement_character();
// not 'u', so this does not come back here
return recover_escape();
}
const int low = get_codepoint();
if (low == -1)
{
add_replacement_character();
return false;
}
if (0xDC00 <= low && low <= 0xDFFF)
{
add_codepoint(static_cast<int>((static_cast<unsigned int>(high) << 10u)
+ static_cast<unsigned int>(low) - 0x35FDC00u));
return true;
}
// high has no low surrogate
add_replacement_character();
if (low < 0xD800 || low > 0xDBFF)
{
add_codepoint(low);
return true;
}
// another high surrogate
high = low;
}
}
/*!
@brief read the escape whose backslash was read; current is the escaped character
@return whether the next character must be read; false if current still
needs to be handled
*/
bool recover_escape()
{
switch (current)
{
case '\"':
add('\"');
return true;
case '\\':
add('\\');
return true;
case '/':
add('/');
return true;
case 'b':
add('\b');
return true;
case 'f':
add('\f');
return true;
case 'n':
add('\n');
return true;
case 'r':
add('\r');
return true;
case 't':
add('\t');
return true;
case 'u':
{
const int codepoint = get_codepoint();
if (codepoint == -1)
{
add_replacement_character();
return false;
}
if (0xD800 <= codepoint && codepoint <= 0xDBFF)
{
return recover_low_surrogate(codepoint);
}
add_escaped_codepoint(codepoint);
return true;
}
// an unknown escape stands for the escaped character
default:
return false;
}
}
/*!
@brief read the rest of a string after scan_string() rejected it
token_buffer holds what scan_string() read before the error. See
recover_token() for how errors are repaired.
@param[in] resume how to continue, see resume_kind
@param[in] codepoint for a high surrogate followed by an escape of another
code point: that code point; -1 otherwise
*/
token_type recover_string(const resume_kind resume, const int codepoint)
{
// whether the next character must be read before it can be handled
bool fetch = false;
if (error_message_starts_with("invalid string: surrogate")
|| error_message_starts_with("invalid string: '\\u'")
|| (error_message_starts_with("invalid string: ill-formed UTF-8")
&& remove_incomplete_utf8_sequence()))
{
add_replacement_character();
}
switch (resume)
{
case resume_kind::escaped_character:
fetch = recover_escape();
break;
case resume_kind::after_escape:
if (0xD800 <= codepoint && codepoint <= 0xDBFF)
{
fetch = recover_low_surrogate(codepoint);
}
else
{
if (codepoint != -1)
{
add_escaped_codepoint(codepoint);
}
fetch = true;
}
break;
case resume_kind::character:
default:
break;
}
while (true)
{
if (fetch)
{
get();
}
fetch = true;
switch (current)
{
case '\"':
// a line break or the end of the input ends a string that
// lacks its closing quote
case '\n':
case '\r':
case char_traits<char_type>::eof():
return token_type::value_string;
#if !JSON_STRICT_NUL_HANDLING
case '\0':
// the end of the input, see scan()
unget();
return token_type::value_string;
#endif
case '\\':
get();
fetch = recover_escape();
break;
default:
if (current < 0x80)
{
// including control characters
add(current);
}
else
{
fetch = recover_utf8_sequence();
}
break;
}
}
}
/*!
@brief keep the longest valid prefix of a number that scan_number() rejected
token_buffer holds the characters scan_number() accepted before the error,
so the prefix ends at its last digit.
*/
token_type recover_number()
{
// only size(), operator[], and resize() are used, which every string
// type the library supports provides
std::size_t length = token_buffer.size();
while (length != 0 && (token_buffer[length - 1] < '0' || token_buffer[length - 1] > '9'))
{
--length;
}
token_buffer.resize(length);
if (length == 0)
{
skip_to_delimiter();
return token_type::uninitialized;
}
if (decimal_point_position >= length)
{
decimal_point_position = std::string::npos;
}
std::size_t exponent = std::string::npos;
for (std::size_t i = 0; i < length; ++i)
{
if (token_buffer[i] == 'e' || token_buffer[i] == 'E')
{
exponent = i;
break;
}
}
const std::size_t mantissa_end = (exponent == std::string::npos) ? length : exponent;
token_type number_type = token_type::value_unsigned;
if (decimal_point_position != std::string::npos || exponent != std::string::npos)
{
number_type = token_type::value_float;
}
else if (token_buffer[0] == '-')
{
number_type = token_type::value_integer;
}
const token_type result = convert_number(number_type, mantissa_end);
skip_to_delimiter();
return result;
}
/// seekable adapter: the token string begins at current, which was consumed
void restart_token_string_impl(std::true_type /*lazy*/) noexcept
{
const std::size_t consumed = ia.get_consumed_count();
token_string_start = (consumed > 0 && current != char_traits<char_type>::eof()) ? consumed - 1 : consumed;
}
/// streaming adapter: the token string begins at current; a character
/// that was put back is copied again when it is read again
void restart_token_string_impl(std::false_type /*lazy*/)
{
token_string.clear();
if (!next_unget && current != char_traits<char_type>::eof())
{
token_string.push_back(char_traits<char_type>::to_char_type(current));
}
}
private: private:
/// input adapter /// input adapter
InputAdapterType ia; InputAdapterType ia;
@@ -2170,6 +2752,13 @@ scan_number_done:
/// a description of occurred lexer errors /// a description of occurred lexer errors
const char* error_message = ""; const char* error_message = "";
/// how recover_token() continues a string that scan_string() rejected;
/// set only on the error paths that need more than error_message
resume_kind string_error_resume = resume_kind::character;
/// the code point of the second escape when a high surrogate is followed
/// by an escape that is not a low surrogate; -1 otherwise
int string_error_codepoint = -1;
// number values // number values
number_integer_t value_integer = 0; number_integer_t value_integer = 0;
number_unsigned_t value_unsigned = 0; number_unsigned_t value_unsigned = 0;
+663 -51
View File
@@ -98,7 +98,7 @@ class parser
if (callback) if (callback)
{ {
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer); json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
sax_parse_internal(&sdp); sax_parse_internal<false>(&sdp);
if (strict) if (strict)
{ {
@@ -135,7 +135,7 @@ class parser
else else
{ {
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer); json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
sax_parse_internal(&sdp); sax_parse_internal<false>(&sdp);
if (strict) if (strict)
{ {
@@ -173,26 +173,59 @@ class parser
bool accept(const bool strict = true) bool accept(const bool strict = true)
{ {
json_sax_acceptor<BasicJsonType> sax_acceptor; json_sax_acceptor<BasicJsonType> sax_acceptor;
return sax_parse(&sax_acceptor, strict); return sax_parse_impl<false>(&sax_acceptor, strict);
} }
/*!
@brief public SAX interface
If the SAX parser's parse_error() returns true, the parser recovers from
the error: it repairs the input and continues (see #3989).
@param[in] sax the SAX parser
@param[in] strict whether to expect the last token to be EOF
@return whether the input was parsed without errors and no SAX event
returned false
*/
template<typename SAX> template<typename SAX>
JSON_HEDLEY_NON_NULL(2) JSON_HEDLEY_NON_NULL(2)
bool sax_parse(SAX* sax, const bool strict = true) bool sax_parse(SAX* sax, const bool strict = true)
{
return sax_parse_impl<true>(sax, strict);
}
private:
/// what sax_parse_internal() does after an object key was expected
enum class next_step : std::uint8_t
{
/// stop parsing
stop,
/// parse a value that begins with last_token
parse_value,
/// evaluate the state of the innermost container, which reads
/// last_token again
evaluate_state
};
template<bool AllowRecovery, typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse_impl(SAX* sax, const bool strict)
{ {
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {}; (void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
const bool result = sax_parse_internal(sax); const bool result = sax_parse_internal<AllowRecovery>(sax);
if (result) if (result)
{ {
if (strict) if (strict)
{ {
// strict mode: next byte must be EOF // strict mode: next byte must be EOF; after recovering from an
if (get_token() != token_type::end_of_input) // error, the end of the input may already have been read
if (last_token != token_type::end_of_input && get_token() != token_type::end_of_input)
{ {
return sax->parse_error(m_lexer.get_position(), // the value is complete, so there is nothing to recover
m_lexer.get_token_string(), static_cast<void>(report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr)); std::integral_constant<bool, AllowRecovery> {}));
return false;
} }
} }
else else
@@ -203,14 +236,23 @@ class parser
} }
} }
return result; return result && !error_reported;
} }
private: /*!
template<typename SAX> @brief parse a JSON value and pass it to a SAX parser
@tparam AllowRecovery whether to recover from an error if the SAX parser's
parse_error() returns true; false for the SAX parsers
of parse() and accept(), which never do, so that no
code for recovering is generated for them
*/
template<bool AllowRecovery, typename SAX>
JSON_HEDLEY_NON_NULL(2) JSON_HEDLEY_NON_NULL(2)
bool sax_parse_internal(SAX* sax) bool sax_parse_internal(SAX* sax)
{ {
const std::integral_constant<bool, AllowRecovery> allow_recovery{};
// stack to remember the hierarchy of structured values we are parsing // stack to remember the hierarchy of structured values we are parsing
// true = array; false = object // true = array; false = object
std::vector<bool> states; std::vector<bool> states;
@@ -241,12 +283,18 @@ class parser
break; break;
} }
// parse key // remember we are now inside an object
states.push_back(false);
// parse key (the steps of parse_key(), which are
// repeated here and below for speed)
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string)) if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{ {
return sax->parse_error(m_lexer.get_position(), if (!continue_after(key_error(sax, allow_recovery, false), skip_to_state_evaluation))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr)); return false;
}
continue;
} }
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string()))) if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{ {
@@ -256,14 +304,13 @@ class parser
// parse separator (:) // parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator)) if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{ {
return sax->parse_error(m_lexer.get_position(), if (!continue_after(key_error(sax, allow_recovery, true), skip_to_state_evaluation))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr)); return false;
}
continue;
} }
// remember we are now inside an object
states.push_back(false);
// parse values // parse values
get_token(); get_token();
continue; continue;
@@ -299,9 +346,11 @@ class parser
if (JSON_HEDLEY_UNLIKELY(!std::isfinite(res))) if (JSON_HEDLEY_UNLIKELY(!std::isfinite(res)))
{ {
return sax->parse_error(m_lexer.get_position(), if (!overflow_error(sax, res, allow_recovery))
m_lexer.get_token_string(), {
out_of_range::create(406, concat("number overflow parsing '", m_lexer.get_token_string(), '\''), nullptr)); return false;
}
break;
} }
if (JSON_HEDLEY_UNLIKELY(!sax->number_float(res, m_lexer.get_string()))) if (JSON_HEDLEY_UNLIKELY(!sax->number_float(res, m_lexer.get_string())))
@@ -369,23 +418,63 @@ class parser
case token_type::parse_error: case token_type::parse_error:
{ {
// using "uninitialized" to avoid an "expected" message // using "uninitialized" to avoid an "expected" message
return sax->parse_error(m_lexer.get_position(), if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::uninitialized, "value"), nullptr), allow_recovery))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::uninitialized, "value"), nullptr)); return false;
}
// recover: keep what can be read of the token
recover_token();
if (last_token != token_type::uninitialized)
{
// a string or a number
continue;
}
if (states.empty())
{
// look for the value after the garbage
if (!skip_to_value())
{
return false;
}
continue;
}
// nothing could be read
if (JSON_HEDLEY_UNLIKELY(!sax->null()))
{
return false;
}
break;
} }
case token_type::end_of_input: case token_type::end_of_input:
{ {
if (JSON_HEDLEY_UNLIKELY(m_lexer.get_position().chars_read_total == 1)) if (JSON_HEDLEY_UNLIKELY(m_lexer.get_position().chars_read_total == 1))
{ {
return sax->parse_error(m_lexer.get_position(), // there is nothing to recover
m_lexer.get_token_string(), static_cast<void>(report_error(sax, parse_error::create(101, m_lexer.get_position(),
parse_error::create(101, m_lexer.get_position(), "attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr), allow_recovery));
"attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr)); return false;
} }
return sax->parse_error(m_lexer.get_position(), if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr), allow_recovery))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr)); return false;
}
// recover: the input ends where a value is missing
if (states.empty())
{
// there is no value
return false;
}
if (!recover_missing_value(sax, states))
{
return false;
}
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
continue;
} }
case token_type::uninitialized: case token_type::uninitialized:
case token_type::end_array: case token_type::end_array:
@@ -395,9 +484,35 @@ class parser
case token_type::literal_or_value: case token_type::literal_or_value:
default: // the last token was unexpected default: // the last token was unexpected
{ {
return sax->parse_error(m_lexer.get_position(), if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr), allow_recovery))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr)); return false;
}
// recover
if (states.empty())
{
// look for the value after the garbage
if (!skip_to_value())
{
return false;
}
continue;
}
if (last_token == token_type::name_separator)
{
// a stray ':'; the value may follow
get_token();
continue;
}
if (!recover_missing_value(sax, states))
{
return false;
}
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
continue;
} }
} }
} }
@@ -447,9 +562,30 @@ class parser
continue; continue;
} }
return sax->parse_error(m_lexer.get_position(), if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_array, "array"), nullptr), allow_recovery))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_array, "array"), nullptr)); return false;
}
// recover
if (last_token == token_type::end_of_input)
{
// the input ends inside the array
return close_containers(sax, states);
}
if (last_token == token_type::end_object)
{
// a wrong closing bracket closes the innermost container
if (JSON_HEDLEY_UNLIKELY(!sax->end_array()))
{
return false;
}
states.pop_back();
skip_to_state_evaluation = true;
}
// otherwise, a missing ',' (or a stray ':', which value
// parsing drops): the next value begins here
continue;
} }
// states.back() is false -> object // states.back() is false -> object
@@ -466,11 +602,12 @@ class parser
// parse key // parse key
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string)) if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{ {
return sax->parse_error(m_lexer.get_position(), if (!continue_after(key_error(sax, allow_recovery, false), skip_to_state_evaluation))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr)); return false;
}
continue;
} }
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string()))) if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{ {
return false; return false;
@@ -479,9 +616,11 @@ class parser
// parse separator (:) // parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator)) if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{ {
return sax->parse_error(m_lexer.get_position(), if (!continue_after(key_error(sax, allow_recovery, true), skip_to_state_evaluation))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr)); return false;
}
continue;
} }
// parse values // parse values
@@ -508,12 +647,479 @@ class parser
continue; continue;
} }
return sax->parse_error(m_lexer.get_position(), if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_object, "object"), nullptr), allow_recovery))
m_lexer.get_token_string(), {
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_object, "object"), nullptr)); return false;
}
// recover
if (last_token == token_type::end_of_input)
{
// the input ends inside the object
return close_containers(sax, states);
}
if (last_token == token_type::end_array)
{
// a wrong closing bracket closes the innermost container
if (JSON_HEDLEY_UNLIKELY(!sax->end_object()))
{
return false;
}
states.pop_back();
skip_to_state_evaluation = true;
continue;
}
if (!continue_after(recover_member(sax, allow_recovery), skip_to_state_evaluation))
{
return false;
}
} }
} }
/*!
@brief continue sax_parse_internal() after a recovery
@return whether to continue parsing
*/
bool continue_after(const next_step step, bool& skip_to_state_evaluation)
{
if (step == next_step::evaluate_state)
{
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
}
return step != next_step::stop;
}
/// the parser for parse() and accept() never recovers: stop parsing
static std::false_type continue_after(std::false_type /*step*/, bool& /*skip_to_state_evaluation*/) noexcept
{
return {};
}
/*!
@brief parse an object key and the name separator (:) after it
last_token is the token where the key is expected. sax_parse_internal()
repeats these steps rather than calling this function, which is used
when recovering from an error.
@return next_step::parse_value if the value follows, with last_token its
first token; next_step::evaluate_state if the object's state is
to be evaluated after recovering from an error; next_step::stop
to stop parsing
*/
template<typename SAX>
next_step parse_key(SAX* sax)
{
const std::true_type allow_recovery{};
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
return key_error(sax, allow_recovery, false);
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
return next_step::stop;
}
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
return key_error(sax, allow_recovery, true);
}
// the value begins with the next token
get_token();
return next_step::parse_value;
}
/*!
@brief report a number that is too large for number_float_t, and recover
from the error by passing the value on; the SAX parser gets the
number's text as well
This is a separate function, as reading other numbers is measurably
slower if the error is handled where they are read.
@param[in] sax the SAX parser
@param[in] value the value that is not finite
@return whether to continue parsing
*/
template<typename SAX, typename AllowRecovery>
bool overflow_error(SAX* sax, const number_float_t value, AllowRecovery allow_recovery)
{
if (!report_error(sax, out_of_range::create(406, concat("number overflow parsing '", m_lexer.get_token_string(), '\''), nullptr), allow_recovery))
{
return false;
}
return sax->number_float(value, m_lexer.get_string());
}
/*!
@brief report a missing key, or a missing name separator (:) after the
key; the parser for parse() and accept() never recovers
@param[in] key_read whether the key was read, so that the name separator
is missing
@return std::false_type, see report_error()
*/
template<typename SAX>
std::false_type key_error(SAX* sax, std::false_type allow_recovery, const bool key_read)
{
return report_error(sax, parse_error::create(101, m_lexer.get_position(), key_read
? exception_message(token_type::name_separator, "object separator")
: exception_message(token_type::value_string, "object key"), nullptr), allow_recovery);
}
/*!
@brief report a missing key, or a missing name separator (:) after the
key, and recover from it
@param[in] key_read whether the key was read, so that the name separator
is missing
*/
template<typename SAX>
next_step key_error(SAX* sax, std::true_type allow_recovery, const bool key_read)
{
if (!key_read)
{
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr), allow_recovery))
{
return next_step::stop;
}
return recover_key(sax);
}
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr), allow_recovery))
{
return next_step::stop;
}
return recover_name_separator(sax);
}
/////////////////////
// error recovery
/////////////////////
/*
The functions below repair an error after the SAX parser's parse_error()
returned true (see #3989). Each mistake is repaired by the smallest local
edit: a missing ',' or ':' is inserted, a stray token is removed, what can
be read of an invalid string or number is kept (see
lexer::recover_token()), a missing value becomes null, a wrong closing
bracket closes the innermost container, and the end of the input closes
all of them. The events stay balanced, and every key() is followed by
exactly one value.
A repair hands a token to the state evaluation, by returning it to the
lexer (lexer::unget_token()) so that the state evaluation reads it again,
only if it is ',', ']', '}', or the end of the input. The state evaluation
hands a token to value or key parsing only if it is none of them, so a
token is never handed back and forth. Every other step reads a token or
closes a container, so parsing always ends.
*/
/*!
@brief report an error to the SAX parser; the parser for parse() and
accept() never recovers
@return std::false_type rather than false: its value is known where the
function is called even if the call is not inlined, so the code
for recovering is not generated
*/
template<typename SAX, typename Exception>
std::false_type report_error(SAX* sax, const Exception& ex, std::false_type /*allow_recovery*/)
{
error_reported = true;
static_cast<void>(sax->parse_error(m_lexer.get_position(), m_lexer.get_token_string(), ex));
return {};
}
/*!
@brief report an error to the SAX parser
@return whether to recover from the error
*/
template<typename SAX, typename Exception>
bool report_error(SAX* sax, const Exception& ex, std::true_type /*allow_recovery*/)
{
const std::size_t position = m_lexer.get_position().chars_read_total;
if (error_reported && position == last_error_position && last_token == last_error_token)
{
// a repair handed on the token of the error it repaired; the
// token was reported already, and the SAX parser asked to recover
return true;
}
error_reported = true;
last_error_position = position;
last_error_token = last_token;
if (!sax->parse_error(m_lexer.get_position(), m_lexer.get_token_string(), ex))
{
return false;
}
// the token string of the next error begins here
m_lexer.restart_token_string();
return true;
}
/*!
@brief keep what can be read of the token that the lexer rejected
The error was reported for the rejected token, so it is not reported again
for the token it is repaired to (see lexer::recover_token()).
*/
token_type recover_token()
{
last_token = m_lexer.recover_token();
last_error_position = m_lexer.get_position().chars_read_total;
last_error_token = last_token;
return last_token;
}
/// pass the end events of all open containers
template<typename SAX>
bool close_containers(SAX* sax, std::vector<bool>& states)
{
while (!states.empty())
{
const bool is_array = states.back();
states.pop_back();
if (JSON_HEDLEY_UNLIKELY(is_array ? !sax->end_array() : !sax->end_object()))
{
return false;
}
}
return true;
}
/*!
@brief read tokens until one begins a value, skipping everything before
the top-level value
@return whether a value begins with last_token
*/
bool skip_to_value()
{
while (true)
{
switch (get_token())
{
case token_type::begin_array:
case token_type::begin_object:
case token_type::literal_false:
case token_type::literal_null:
case token_type::literal_true:
case token_type::value_float:
case token_type::value_integer:
case token_type::value_string:
case token_type::value_unsigned:
return true;
case token_type::end_of_input:
return false;
case token_type::parse_error:
recover_token();
if (last_token != token_type::uninitialized)
{
return true;
}
break;
case token_type::uninitialized:
case token_type::end_array:
case token_type::end_object:
case token_type::name_separator:
case token_type::value_separator:
case token_type::literal_or_value:
default:
break;
}
}
}
/*!
@brief skip the rest of an object member that cannot be read
Reads tokens, beginning with last_token, until a ',', '}', or ']' that is
not inside a container that begins in the skipped tokens, or the end of
the input.
*/
void skip_member()
{
std::size_t depth = 0;
while (true)
{
switch (last_token)
{
case token_type::begin_array:
case token_type::begin_object:
++depth;
break;
case token_type::end_array:
case token_type::end_object:
if (depth == 0)
{
return;
}
--depth;
break;
case token_type::value_separator:
if (depth == 0)
{
return;
}
break;
case token_type::end_of_input:
return;
case token_type::parse_error:
recover_token();
break;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::name_separator:
case token_type::literal_or_value:
default:
break;
}
get_token();
}
}
/*!
@brief pass a value where it is missing
last_token is ',', ']', '}', or the end of the input, where a value was
expected. In an object, the key gets null; in an array, a ',' where a
value is missing stands for null (as in JavaScript), while an array that
ends there just ends.
*/
template<typename SAX>
bool recover_missing_value(SAX* sax, const std::vector<bool>& states)
{
JSON_ASSERT(!states.empty());
if (!states.back() || last_token == token_type::value_separator)
{
return sax->null();
}
return true;
}
/// recover from a missing key; last_token is where it was expected
template<typename SAX>
next_step recover_key(SAX* sax)
{
switch (last_token)
{
case token_type::value_separator:
case token_type::end_object:
case token_type::end_array:
case token_type::end_of_input:
// no member: the object's state handles the token
return next_step::evaluate_state;
case token_type::parse_error:
recover_token();
if (last_token == token_type::value_string)
{
// a key that could be repaired
return parse_key(sax);
}
skip_member();
return next_step::evaluate_state;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::begin_array:
case token_type::begin_object:
case token_type::name_separator:
case token_type::literal_or_value:
default:
// a member without a key
skip_member();
return next_step::evaluate_state;
}
}
/// recover from a missing name separator (:) after the key; last_token
/// is where it was expected
template<typename SAX>
next_step recover_name_separator(SAX* sax)
{
switch (last_token)
{
case token_type::value_separator:
case token_type::end_object:
case token_type::end_array:
case token_type::end_of_input:
// the value is missing as well
return sax->null() ? next_step::evaluate_state : next_step::stop;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::begin_array:
case token_type::begin_object:
case token_type::name_separator:
case token_type::parse_error:
case token_type::literal_or_value:
default:
// a missing ':'; the value begins here
return next_step::parse_value;
}
}
/// recover from a token after an object member that is neither ',' nor
/// '}' (nor ']' or the end of the input, which the caller handles)
template<typename SAX>
next_step recover_member(SAX* sax, std::true_type /*allow_recovery*/)
{
if (last_token == token_type::parse_error)
{
recover_token();
}
if (last_token == token_type::value_string)
{
// a missing ','; the next key begins here
return parse_key(sax);
}
skip_member();
return next_step::evaluate_state;
}
/// the parser for parse() and accept() never recovers (and does not come
/// here, as report_error() returned false)
template<typename SAX>
std::false_type recover_member(SAX* /*sax*/, std::false_type /*allow_recovery*/) const noexcept
{
return {};
}
/// get next token from lexer /// get next token from lexer
token_type get_token() token_type get_token()
{ {
@@ -560,6 +1166,12 @@ class parser
const bool allow_exceptions = true; const bool allow_exceptions = true;
/// whether trailing commas in objects and arrays should be ignored (true) or signaled as errors (false) /// whether trailing commas in objects and arrays should be ignored (true) or signaled as errors (false)
const bool ignore_trailing_commas = false; const bool ignore_trailing_commas = false;
/// whether an error was reported to the SAX parser
bool error_reported = false;
/// the position of the last reported error
std::size_t last_error_position = 0;
/// the token of the last reported error
token_type last_error_token = token_type::uninitialized;
}; };
} // namespace detail } // namespace detail
+9
View File
@@ -359,6 +359,7 @@ void templated_json_throw(ExceptionType exception)
// Macros to simplify conversion from/to types // Macros to simplify conversion from/to types
// NLOHMANN_JSON_EXPAND to NLOHMANN_JSON_DOUBLE_PASTE63 are generated by tools/macro_builder (see its README.md)
#define NLOHMANN_JSON_EXPAND( x ) x #define NLOHMANN_JSON_EXPAND( x ) x
#define NLOHMANN_JSON_GET_MACRO(_1, _2, _3, _4, _5, _6, _7, _8, _9, _10, _11, _12, _13, _14, _15, _16, _17, _18, _19, _20, _21, _22, _23, _24, _25, _26, _27, _28, _29, _30, _31, _32, _33, _34, _35, _36, _37, _38, _39, _40, _41, _42, _43, _44, _45, _46, _47, _48, _49, _50, _51, _52, _53, _54, _55, _56, _57, _58, _59, _60, _61, _62, _63, _64, NAME,...) NAME #define NLOHMANN_JSON_GET_MACRO(_1, _2, _3, _4, _5, _6, _7, _8, _9, _10, _11, _12, _13, _14, _15, _16, _17, _18, _19, _20, _21, _22, _23, _24, _25, _26, _27, _28, _29, _30, _31, _32, _33, _34, _35, _36, _37, _38, _39, _40, _41, _42, _43, _44, _45, _46, _47, _48, _49, _50, _51, _52, _53, _54, _55, _56, _57, _58, _59, _60, _61, _62, _63, _64, NAME,...) NAME
#define NLOHMANN_JSON_PASTE(...) NLOHMANN_JSON_EXPAND(NLOHMANN_JSON_GET_MACRO(__VA_ARGS__, \ #define NLOHMANN_JSON_PASTE(...) NLOHMANN_JSON_EXPAND(NLOHMANN_JSON_GET_MACRO(__VA_ARGS__, \
@@ -621,6 +622,10 @@ void templated_json_throw(ExceptionType exception)
// arguments, so dispatching on Type,BaseType,member... directly would run out // arguments, so dispatching on Type,BaseType,member... directly would run out
// one slot early and cap the derived-type macros at 62 members instead of the // one slot early and cap the derived-type macros at 62 members instead of the
// 63 that NLOHMANN_JSON_PASTE supports. // 63 that NLOHMANN_JSON_PASTE supports.
//
// The slot table below (down to the closing NLOHMANN_JSON_TYPE_BODY_SENTINEL))
// is generated by tools/macro_builder (see its README.md; run with the
// "type_body" argument).
#define NLOHMANN_JSON_TYPE_BODY(Prefix, ...) NLOHMANN_JSON_EXPAND(NLOHMANN_JSON_GET_MACRO(__VA_ARGS__, \ #define NLOHMANN_JSON_TYPE_BODY(Prefix, ...) NLOHMANN_JSON_EXPAND(NLOHMANN_JSON_GET_MACRO(__VA_ARGS__, \
Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, \ Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, \
Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, \ Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, Prefix ## MEMBERS, \
@@ -915,3 +920,7 @@ void templated_json_throw(ExceptionType exception)
#ifndef JSON_DISABLE_ENUM_SERIALIZATION #ifndef JSON_DISABLE_ENUM_SERIALIZATION
#define JSON_DISABLE_ENUM_SERIALIZATION 0 #define JSON_DISABLE_ENUM_SERIALIZATION 0
#endif #endif
#ifndef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
#define JSON_DISABLE_TUPLE_REFERENCE_CONVERSION 0
#endif
@@ -25,6 +25,7 @@
#undef JSON_INLINE_VARIABLE #undef JSON_INLINE_VARIABLE
#undef JSON_NO_UNIQUE_ADDRESS #undef JSON_NO_UNIQUE_ADDRESS
#undef JSON_DISABLE_ENUM_SERIALIZATION #undef JSON_DISABLE_ENUM_SERIALIZATION
#undef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
#ifndef JSON_TEST_KEEP_MACROS #ifndef JSON_TEST_KEEP_MACROS
#undef JSON_CATCH #undef JSON_CATCH
@@ -9,7 +9,6 @@
#pragma once #pragma once
#include <array> // array
#include <cstddef> // size_t #include <cstddef> // size_t
#include <type_traits> // conditional, enable_if, false_type, integral_constant, is_constructible, is_integral, is_same, remove_cv, remove_reference, true_type #include <type_traits> // conditional, enable_if, false_type, integral_constant, is_constructible, is_integral, is_same, remove_cv, remove_reference, true_type
#include <utility> // index_sequence, make_index_sequence, index_sequence_for #include <utility> // index_sequence, make_index_sequence, index_sequence_for
@@ -161,11 +160,5 @@ struct static_const
constexpr T static_const<T>::value; constexpr T static_const<T>::value;
#endif #endif
template<typename T, typename... Args>
constexpr std::array<T, sizeof...(Args)> make_array(Args&& ... args)
{
return std::array<T, sizeof...(Args)> {{static_cast<T>(std::forward<Args>(args))...}};
}
} // namespace detail } // namespace detail
NLOHMANN_JSON_NAMESPACE_END NLOHMANN_JSON_NAMESPACE_END
@@ -636,6 +636,18 @@ template<typename BasicJsonType, typename CompatibleType>
struct is_compatible_type struct is_compatible_type
: is_compatible_type_impl<BasicJsonType, CompatibleType> {}; : is_compatible_type_impl<BasicJsonType, CompatibleType> {};
// a one-element std::tuple holding a reference to BasicJsonType, as created by
// std::forward_as_tuple(j); see JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
template<typename BasicJsonType, typename T>
struct is_basic_json_reference_tuple : std::false_type {};
template<typename BasicJsonType, typename T>
struct is_basic_json_reference_tuple<BasicJsonType, std::tuple<T>>
{
static constexpr bool value =
std::is_reference<T>::value && std::is_same<uncvref_t<T>, BasicJsonType>::value;
};
template<typename BasicJsonType, typename CompatibleArrayType> template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_binary_type struct is_compatible_binary_type
{ {
File diff suppressed because it is too large Load Diff
+74
View File
@@ -12,6 +12,7 @@
#include <cstddef> // size_t #include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t #include <cstdint> // uint8_t, uint32_t
#include <string> // string, to_string #include <string> // string, to_string
#include <utility> // move
#include <nlohmann/detail/abi_macros.hpp> #include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/macro_scope.hpp> #include <nlohmann/detail/macro_scope.hpp>
@@ -133,5 +134,78 @@ inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noex
return state == UTF8_ACCEPT; return state == UTF8_ACCEPT;
} }
/*!
@brief append U+FFFD REPLACEMENT CHARACTER, encoded in UTF-8
@param[in,out] s the string to append to
*/
template<typename StringType>
inline void append_replacement_character(StringType& s)
{
s.push_back(static_cast<typename StringType::value_type>(0xEFu));
s.push_back(static_cast<typename StringType::value_type>(0xBFu));
s.push_back(static_cast<typename StringType::value_type>(0xBDu));
}
/*!
@brief replace ill-formed UTF-8 with U+FFFD REPLACEMENT CHARACTER
Each maximal subpart of an ill-formed sequence becomes one U+FFFD, as the
Unicode Standard recommends (Section 3.9, "U+FFFD Substitution of Maximal
Subparts"), and as the parser for JSON text does when it recovers from errors.
@param[in,out] s the string to repair
@param[in] first index of the first byte to repair; the bytes before it are
assumed to be valid UTF-8 that ends on a code point boundary
*/
template<typename StringType>
inline void replace_invalid_utf8(StringType& s, const std::size_t first = 0)
{
StringType result = s;
result.resize(first);
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
// the first byte of the sequence being decoded
std::size_t sequence_start = first;
std::size_t i = first;
while (i < s.size())
{
switch (decode(state, codepoint, static_cast<std::uint8_t>(s[i])))
{
case UTF8_ACCEPT:
for (++i; sequence_start < i; ++sequence_start)
{
result.push_back(s[sequence_start]);
}
break;
case UTF8_REJECT:
append_replacement_character(result);
// the byte that made the sequence ill-formed begins the next
// one, unless it began this one
if (i == sequence_start)
{
++i;
}
state = UTF8_ACCEPT;
sequence_start = i;
break;
default: // in the middle of a sequence
++i;
break;
}
}
// a sequence that the string ends in the middle of
if (state != UTF8_ACCEPT)
{
append_replacement_character(result);
}
s = std::move(result);
}
} // namespace detail } // namespace detail
NLOHMANN_JSON_NAMESPACE_END NLOHMANN_JSON_NAMESPACE_END
+253 -68
View File
@@ -145,7 +145,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
friend class ::nlohmann::detail::iter_impl; friend class ::nlohmann::detail::iter_impl;
template<typename BasicJsonType, typename CharType, typename OutputSinkType> template<typename BasicJsonType, typename CharType, typename OutputSinkType>
friend class ::nlohmann::detail::binary_writer; friend class ::nlohmann::detail::binary_writer;
template<typename BasicJsonType, typename InputType, typename SAX> template<typename BasicJsonType, typename InputType, typename SAX, bool AllowRecovery>
friend class ::nlohmann::detail::binary_reader; friend class ::nlohmann::detail::binary_reader;
template<typename BasicJsonType, typename InputAdapterType> template<typename BasicJsonType, typename InputAdapterType>
friend class ::nlohmann::detail::json_sax_dom_parser; friend class ::nlohmann::detail::json_sax_dom_parser;
@@ -906,11 +906,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/*! /*!
@brief how many levels the operation going on in this thread has descended into @brief how many levels the operation going on in this thread has descended into
Copying a value and comparing two values share this count. The library never Copying a value, converting one from another specialization, and comparing
nests one inside the other - copying a value does not compare one, and two values share this count. The library never nests one of them inside
comparing two values does not copy them - and where user code nests them another - none of them does either of the other two on the way - and where
anyway, sharing the count only ends a descent sooner than it had to, which user code nests them anyway, sharing the count only ends a descent sooner
costs a little speed and is never wrong. than it had to, which costs a little speed and is never wrong.
A byte is enough: the count never exceeds the limit by more than the single A byte is enough: the count never exceeds the limit by more than the single
level that notices the limit has been reached. level that notices the limit has been reached.
@@ -1267,6 +1267,222 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
copy_iteratively(src); copy_iteratively(src);
} }
/*!
@brief convert the value @a val of another specialization into this null
value; @a val must be neither an object nor an array
Converting such a value never descends, so both ways of converting an
object or an array (@ref convert_structured) leave their elements of this
kind to the converting constructor, which leaves them to this.
*/
template<typename BasicJsonType>
void convert_leaf(const BasicJsonType& val)
{
using other_boolean_t = typename BasicJsonType::boolean_t;
using other_number_float_t = typename BasicJsonType::number_float_t;
using other_number_integer_t = typename BasicJsonType::number_integer_t;
using other_number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using other_string_t = typename BasicJsonType::string_t;
using other_binary_t = typename BasicJsonType::binary_t;
switch (val.type())
{
case value_t::boolean:
JSONSerializer<other_boolean_t>::to_json(*this, val.template get<other_boolean_t>());
break;
case value_t::number_float:
JSONSerializer<other_number_float_t>::to_json(*this, val.template get<other_number_float_t>());
break;
case value_t::number_integer:
JSONSerializer<other_number_integer_t>::to_json(*this, val.template get<other_number_integer_t>());
break;
case value_t::number_unsigned:
JSONSerializer<other_number_unsigned_t>::to_json(*this, val.template get<other_number_unsigned_t>());
break;
case value_t::string:
JSONSerializer<other_string_t>::to_json(*this, val.template get_ref<const other_string_t&>());
break;
case value_t::binary:
JSONSerializer<other_binary_t>::to_json(*this, val.template get_ref<const other_binary_t&>());
break;
case value_t::null:
// m_data.m_type is already value_t::null
break;
case value_t::discarded:
m_data.m_type = value_t::discarded;
break;
case value_t::object: // LCOV_EXCL_LINE
case value_t::array: // LCOV_EXCL_LINE
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
}
/// scratch space for the converted elements of the arrays that
/// @ref convert_iteratively has yet to create
using convert_scratch_t = std::vector<basic_json, AllocatorType<basic_json>>;
/*!
@brief create the object or array @a val converted into this null value
Its converted elements are the last `val.size()` entries of @a elements (an
array) or of @a members (an object); they are moved into the container in
one go and then removed.
*/
template<typename BasicJsonType>
void convert_level(const BasicJsonType& val, convert_scratch_t& elements, copy_scratch_t& members)
{
if (val.is_object())
{
const auto first = members.end() - static_cast<typename copy_scratch_t::difference_type>(val.size());
m_data.m_value.object = create<object_t>(std::make_move_iterator(first),
std::make_move_iterator(members.end()));
// only now that the object exists may this stop being a null value
m_data.m_type = value_t::object;
members.erase(first, members.end());
}
else
{
const auto first = elements.end() - static_cast<typename convert_scratch_t::difference_type>(val.size());
m_data.m_value.array = create<array_t>(std::make_move_iterator(first),
std::make_move_iterator(elements.end()));
// only now that the array exists may this stop being a null value
m_data.m_type = value_t::array;
elements.erase(first, elements.end());
}
set_parents();
}
/*!
@brief convert the object or array @a val of another specialization into
this null value without recursing
The containers whose conversion has begun are kept on an explicit stack
rather than on the call stack. Unlike @ref copy_iteratively, this builds
every container from the bottom up: all its elements are converted first,
and the container is then created from them in one go, the way the range
constructor that converts the levels above the bound does. The two object
types need not enumerate their members in the same order, so the members
could not be paired up by position anyway, and building from a range keeps
what the range constructor does with keys that become equal on conversion.
Every value is complete before it is handed on, and a container gets its
type only once it exists, so whatever throws, every value left behind can
be destroyed.
*/
template<typename BasicJsonType>
void convert_iteratively(const BasicJsonType& val)
{
using other_const_iterator = typename BasicJsonType::const_iterator;
// the containers whose conversion has begun, innermost last, each with
// its element to convert next
std::vector<std::pair<const BasicJsonType*, other_const_iterator>> pending;
// the converted elements of the pending arrays and the converted
// members of the pending objects, those of the innermost one last
convert_scratch_t elements;
copy_scratch_t members;
pending.emplace_back(&val, val.cbegin());
for (;;)
{
const BasicJsonType& container = *pending.back().first;
// a copy, as descending below can reallocate pending; the
// iterator kept in pending is only advanced through pending.back()
const other_const_iterator next = pending.back().second;
if (next != container.cend())
{
if (next->is_structured())
{
// convert its elements first; next stays where it is until
// the converted container is handed back to this one
pending.emplace_back(&*next, next->cbegin());
continue;
}
// the converting constructor does not descend into this value
if (container.is_object())
{
members.emplace_back(next.key(), *next);
}
else
{
elements.emplace_back(*next);
}
++pending.back().second;
continue;
}
// all elements of the container are converted: create it
pending.pop_back();
if (pending.empty())
{
convert_level(container, elements, members);
return;
}
basic_json converted;
converted.convert_level(container, elements, members);
#if JSON_DIAGNOSTIC_POSITIONS
converted.start_position = container.start_pos();
converted.end_position = container.end_pos();
#endif
// hand it to the container it is an element of
if (pending.back().first->is_object())
{
members.emplace_back(pending.back().second.key(), std::move(converted));
}
else
{
elements.push_back(std::move(converted));
}
++pending.back().second;
}
}
/*!
@brief convert the object or array @a val of another specialization into
this null value
Converting a container converts its elements, so a value nested deeply
enough used to exhaust the call stack. The descent is bounded here as in
@ref copy_structured: the first `detail::recursion_depth_limit()` levels
are converted by the containers' range constructors, just as they always were,
and anything below that is converted without the call stack by
@ref convert_iteratively.
@sa https://github.com/nlohmann/json/issues/5650
*/
template<typename BasicJsonType>
void convert_structured(const BasicJsonType& val)
{
const nesting_depth_guard guard;
if (JSON_HEDLEY_LIKELY(guard.okay()))
{
// every element comes back to the converting constructor
if (val.is_object())
{
using other_object_t = typename BasicJsonType::object_t;
JSONSerializer<other_object_t>::to_json(*this, val.template get_ref<const other_object_t&>());
}
else
{
using other_array_t = typename BasicJsonType::array_t;
JSONSerializer<other_array_t>::to_json(*this, val.template get_ref<const other_array_t&>());
}
return;
}
convert_iteratively(val);
}
/// the result of comparing two values, including values that cannot be /// the result of comparing two values, including values that cannot be
/// ordered at all, such as a discarded value or a NaN /// ordered at all, such as a discarded value or a NaN
@@ -1596,7 +1812,12 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
template < typename CompatibleType, template < typename CompatibleType,
typename U = detail::uncvref_t<CompatibleType>, typename U = detail::uncvref_t<CompatibleType>,
detail::enable_if_t < detail::enable_if_t <
!detail::is_basic_json<U>::value && detail::is_compatible_type<basic_json_t, U>::value, int > = 0 > !detail::is_basic_json<U>::value && detail::is_compatible_type<basic_json_t, U>::value
#if JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
// see https://github.com/nlohmann/json/issues/2226
&& !detail::is_basic_json_reference_tuple<basic_json_t, U>::value
#endif
, int > = 0 >
basic_json(CompatibleType && val) noexcept(noexcept( // NOLINT(bugprone-forwarding-reference-overload,bugprone-exception-escape) basic_json(CompatibleType && val) noexcept(noexcept( // NOLINT(bugprone-forwarding-reference-overload,bugprone-exception-escape)
JSONSerializer<U>::to_json(std::declval<basic_json_t&>(), JSONSerializer<U>::to_json(std::declval<basic_json_t&>(),
std::forward<CompatibleType>(val)))) std::forward<CompatibleType>(val))))
@@ -1617,49 +1838,13 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
end_position(val.end_pos()) end_position(val.end_pos())
#endif #endif
{ {
using other_boolean_t = typename BasicJsonType::boolean_t; if (val.is_structured())
using other_number_float_t = typename BasicJsonType::number_float_t;
using other_number_integer_t = typename BasicJsonType::number_integer_t;
using other_number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using other_string_t = typename BasicJsonType::string_t;
using other_object_t = typename BasicJsonType::object_t;
using other_array_t = typename BasicJsonType::array_t;
using other_binary_t = typename BasicJsonType::binary_t;
switch (val.type())
{ {
case value_t::boolean: convert_structured(val);
JSONSerializer<other_boolean_t>::to_json(*this, val.template get<other_boolean_t>()); }
break; else
case value_t::number_float: {
JSONSerializer<other_number_float_t>::to_json(*this, val.template get<other_number_float_t>()); convert_leaf(val);
break;
case value_t::number_integer:
JSONSerializer<other_number_integer_t>::to_json(*this, val.template get<other_number_integer_t>());
break;
case value_t::number_unsigned:
JSONSerializer<other_number_unsigned_t>::to_json(*this, val.template get<other_number_unsigned_t>());
break;
case value_t::string:
JSONSerializer<other_string_t>::to_json(*this, val.template get_ref<const other_string_t&>());
break;
case value_t::object:
JSONSerializer<other_object_t>::to_json(*this, val.template get_ref<const other_object_t&>());
break;
case value_t::array:
JSONSerializer<other_array_t>::to_json(*this, val.template get_ref<const other_array_t&>());
break;
case value_t::binary:
JSONSerializer<other_binary_t>::to_json(*this, val.template get_ref<const other_binary_t&>());
break;
case value_t::null:
*this = nullptr;
break;
case value_t::discarded:
m_data.m_type = value_t::discarded;
break;
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
} }
JSON_ASSERT(m_data.m_type == val.type()); JSON_ASSERT(m_data.m_type == val.type());
@@ -5034,7 +5219,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
return format == input_format_t::json return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict) ? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(format, sax, strict); : detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(sax, strict);
} }
/// @brief generate SAX events (iterator pair, or iterator+sentinel pair for C++20 ranges support) /// @brief generate SAX events (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -5051,7 +5236,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
return format == input_format_t::json return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict) ? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(format, sax, strict); : detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(sax, strict);
} }
/// @brief generate SAX events /// @brief generate SAX events
@@ -5073,7 +5258,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict) ? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(format, sax, strict); : detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(sax, strict);
} }
#ifndef JSON_NO_IO #ifndef JSON_NO_IO
/// @brief deserialize from stream /// @brief deserialize from stream
@@ -5371,7 +5556,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5391,7 +5576,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5420,7 +5605,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(input_format_t::cbor, &sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::cbor).sax_parse(&sdp, strict, tag_handler)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5438,7 +5623,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5457,7 +5642,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5484,7 +5669,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(input_format_t::msgpack, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::msgpack).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5502,7 +5687,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5521,7 +5706,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5548,7 +5733,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(input_format_t::ubjson, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::ubjson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5566,7 +5751,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5585,7 +5770,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(input_format_t::bjdata, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bjdata).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5603,7 +5788,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bon8).sax_parse(input_format_t::bon8, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bon8).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5622,7 +5807,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bon8).sax_parse(input_format_t::bon8, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bon8).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5640,7 +5825,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::forward<InputType>(i)); auto ia = detail::input_adapter(std::forward<InputType>(i));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5659,7 +5844,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
basic_json result; basic_json result;
auto ia = detail::input_adapter(std::move(first), std::move(last)); auto ia = detail::input_adapter(std::move(first), std::move(last));
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
@@ -5686,7 +5871,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = i.get(); auto ia = i.get();
detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions); detail::json_sax_dom_parser<basic_json, decltype(ia)> sdp(result, allow_exceptions);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(input_format_t::bson, &sdp, strict)) // cppcheck-suppress[accessMoved] if (!binary_reader<decltype(ia)>(std::move(ia), input_format_t::bson).sax_parse(&sdp, strict)) // cppcheck-suppress[accessMoved]
{ {
result = value_t::discarded; result = value_t::discarded;
} }
File diff suppressed because it is too large Load Diff
+1
View File
@@ -47,6 +47,7 @@ inline namespace json_literals
namespace detail namespace detail
{ {
using NLOHMANN_JSON_NAMESPACE::detail::json_sax_dom_callback_parser; using NLOHMANN_JSON_NAMESPACE::detail::json_sax_dom_callback_parser;
using NLOHMANN_JSON_NAMESPACE::detail::json_sax_dom_parser;
using NLOHMANN_JSON_NAMESPACE::detail::unknown_size; using NLOHMANN_JSON_NAMESPACE::detail::unknown_size;
} // namespace detail } // namespace detail
+12
View File
@@ -45,6 +45,10 @@ dumps is stable under exactly the same values that break operator==.
The unit tests run the same checks on a fixed corpus (see the "BJData round-trip The unit tests run the same checks on a fixed corpus (see the "BJData round-trip
invariants" test case), so keep both in sync. invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bjdata() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers. drivers.
*/ */
@@ -59,6 +63,8 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG" #error "the fuzzer drivers must be built without NDEBUG"
#endif #endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json; using json = nlohmann::json;
// value-stable comparison for the round-trip checks below; see the note // value-stable comparison for the round-trip checks below; see the note
@@ -71,11 +77,15 @@ static bool is_value_stable(const json& lhs, const json& rhs)
// see http://llvm.org/docs/LibFuzzer.html // see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{ {
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bjdata).errors == 0;
try try
{ {
// step 1: parse input // step 1: parse input
std::vector<uint8_t> const vec1(data, data + size); std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bjdata(vec1); json const j1 = json::from_bjdata(vec1);
assert(recovered_without_errors);
try try
{ {
@@ -109,6 +119,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&) catch (const json::parse_error&)
{ {
// parse errors are ok, because input may be random bytes // parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
} }
catch (const json::type_error&) catch (const json::type_error&)
{ {
@@ -117,6 +128,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&) catch (const json::out_of_range&)
{ {
// out of range errors may happen if provided sizes are excessive // out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
} }
// return 0 - non-zero return values are reserved for future use // return 0 - non-zero return values are reserved for future use
+12
View File
@@ -19,6 +19,10 @@ It also checks that reading the data from a stream, which reads strings byte by
byte, gives the same value or error as reading it from contiguous memory, which byte, gives the same value or error as reading it from contiguous memory, which
copies strings in bulk. copies strings in bulk.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bon8() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers. drivers.
*/ */
@@ -33,6 +37,8 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG" #error "the fuzzer drivers must be built without NDEBUG"
#endif #endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json; using json = nlohmann::json;
namespace namespace
@@ -56,6 +62,9 @@ std::string read_bon8(InputType&& input)
// see http://llvm.org/docs/LibFuzzer.html // see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{ {
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bon8).errors == 0;
// contiguous and stream input must be read alike // contiguous and stream input must be read alike
{ {
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size)); std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
@@ -67,6 +76,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
// step 1: parse input // step 1: parse input
std::vector<uint8_t> const vec1(data, data + size); std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bon8(vec1); json const j1 = json::from_bon8(vec1);
assert(recovered_without_errors);
try try
{ {
@@ -88,6 +98,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&) catch (const json::parse_error&)
{ {
// parse errors are ok, because input may be random bytes // parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
} }
catch (const json::type_error&) catch (const json::type_error&)
{ {
@@ -96,6 +107,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&) catch (const json::out_of_range&)
{ {
// out of range errors may happen if provided sizes are excessive // out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
} }
// return 0 - non-zero return values are reserved for future use // return 0 - non-zero return values are reserved for future use
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_bson(vec) - j2 = from_bson(vec)
- assert(j1 == j2) - assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers. drivers.
*/ */
@@ -29,16 +33,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG" #error "the fuzzer drivers must be built without NDEBUG"
#endif #endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json; using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html // see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{ {
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bson).errors == 0;
try try
{ {
// step 1: parse input // step 1: parse input
std::vector<uint8_t> const vec1(data, data + size); std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bson(vec1); json const j1 = json::from_bson(vec1);
assert(recovered_without_errors);
if (j1.is_discarded()) if (j1.is_discarded())
{ {
@@ -65,6 +75,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&) catch (const json::parse_error&)
{ {
// parse errors are ok, because input may be random bytes // parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
} }
catch (const json::type_error&) catch (const json::type_error&)
{ {
@@ -73,6 +84,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&) catch (const json::out_of_range&)
{ {
// out of range errors can occur during parsing, too // out of range errors can occur during parsing, too
assert(!recovered_without_errors);
} }
// return 0 - non-zero return values are reserved for future use // return 0 - non-zero return values are reserved for future use
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_cbor(vec) - j2 = from_cbor(vec)
- assert(j1 == j2) - assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_cbor() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers. drivers.
*/ */
@@ -29,16 +33,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG" #error "the fuzzer drivers must be built without NDEBUG"
#endif #endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json; using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html // see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{ {
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::cbor).errors == 0;
try try
{ {
// step 1: parse input // step 1: parse input
std::vector<uint8_t> const vec1(data, data + size); std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_cbor(vec1); json const j1 = json::from_cbor(vec1);
assert(recovered_without_errors);
try try
{ {
@@ -60,6 +70,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&) catch (const json::parse_error&)
{ {
// parse errors are ok, because input may be random bytes // parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
} }
catch (const json::type_error&) catch (const json::type_error&)
{ {
@@ -68,6 +79,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&) catch (const json::out_of_range&)
{ {
// out of range errors can occur during parsing, too // out of range errors can occur during parsing, too
assert(!recovered_without_errors);
} }
// return 0 - non-zero return values are reserved for future use // return 0 - non-zero return values are reserved for future use
+14
View File
@@ -16,6 +16,10 @@ array data, it performs the following steps:
- s2 = serialize(j2) - s2 = serialize(j2)
- assert(s1 == s2) - assert(s1 == s2)
Furthermore, it parses data with a SAX parser that recovers from every error
and checks that the events are balanced, that parsing ends, and that valid
input is parsed without errors (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers. drivers.
*/ */
@@ -23,6 +27,7 @@ drivers.
#include <cassert> #include <cassert>
#include <iostream> #include <iostream>
#include <sstream> #include <sstream>
#include <string>
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
// the round-trip checks below are assertions; NDEBUG would compile them away // the round-trip checks below are assertions; NDEBUG would compile them away
@@ -30,11 +35,20 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG" #error "the fuzzer drivers must be built without NDEBUG"
#endif #endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json; using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html // see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{ {
// step 0: recover from all errors, reading from memory and from a stream
{
const auto checker = check_recovering_parse(data, size, json::input_format_t::json);
assert(checker.events <= (4 * size) + 4);
assert((checker.errors == 0) == json::accept(data, data + size));
}
try try
{ {
// step 1: parse input // step 1: parse input
+12
View File
@@ -15,6 +15,10 @@ array data, it performs the following steps:
- j2 = from_msgpack(vec) - j2 = from_msgpack(vec)
- assert(j1 == j2) - assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_msgpack() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers. drivers.
*/ */
@@ -29,16 +33,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG" #error "the fuzzer drivers must be built without NDEBUG"
#endif #endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json; using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html // see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{ {
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::msgpack).errors == 0;
try try
{ {
// step 1: parse input // step 1: parse input
std::vector<uint8_t> const vec1(data, data + size); std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_msgpack(vec1); json const j1 = json::from_msgpack(vec1);
assert(recovered_without_errors);
try try
{ {
@@ -60,6 +70,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&) catch (const json::parse_error&)
{ {
// parse errors are ok, because input may be random bytes // parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
} }
catch (const json::type_error&) catch (const json::type_error&)
{ {
@@ -68,6 +79,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&) catch (const json::out_of_range&)
{ {
// out of range errors may happen if provided sizes are excessive // out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
} }
// return 0 - non-zero return values are reserved for future use // return 0 - non-zero return values are reserved for future use
+12
View File
@@ -24,6 +24,10 @@ array data, it performs the following steps:
The unit tests run the same checks on a fixed corpus (see the "UBJSON round-trip The unit tests run the same checks on a fixed corpus (see the "UBJSON round-trip
invariants" test case), so keep both in sync. invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_ubjson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers. drivers.
*/ */
@@ -38,16 +42,22 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG" #error "the fuzzer drivers must be built without NDEBUG"
#endif #endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json; using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html // see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size) extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{ {
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::ubjson).errors == 0;
try try
{ {
// step 1: parse input // step 1: parse input
std::vector<uint8_t> const vec1(data, data + size); std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_ubjson(vec1); json const j1 = json::from_ubjson(vec1);
assert(recovered_without_errors);
try try
{ {
@@ -79,6 +89,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&) catch (const json::parse_error&)
{ {
// parse errors are ok, because input may be random bytes // parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
} }
catch (const json::type_error&) catch (const json::type_error&)
{ {
@@ -87,6 +98,7 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&) catch (const json::out_of_range&)
{ {
// out of range errors may happen if provided sizes are excessive // out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
} }
// return 0 - non-zero return values are reserved for future use // return 0 - non-zero return values are reserved for future use
+154
View File
@@ -0,0 +1,154 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cassert>
#include <cstddef>
#include <cstdint>
#include <sstream>
#include <string>
#include <vector>
#include <nlohmann/json.hpp>
namespace
{
// a SAX parser that recovers from every error and checks that the events are
// balanced and that every key is followed by exactly one value
class recovering_checker : public nlohmann::json_sax<nlohmann::json>
{
public:
bool null() override
{
return value();
}
bool boolean(bool /*val*/) override
{
return value();
}
bool number_integer(number_integer_t /*val*/) override
{
return value();
}
bool number_unsigned(number_unsigned_t /*val*/) override
{
return value();
}
bool number_float(number_float_t /*val*/, const string_t& /*s*/) override
{
return value();
}
bool string(string_t& /*val*/) override
{
return value();
}
bool binary(binary_t& /*val*/) override
{
return value();
}
bool start_object(std::size_t /*elements*/) override
{
value();
stack.push_back('o');
return true;
}
bool key(string_t& /*val*/) override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.back() = 'v';
return true;
}
bool end_object() override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.pop_back();
return true;
}
bool start_array(std::size_t /*elements*/) override
{
value();
stack.push_back('a');
return true;
}
bool end_array() override
{
++events;
assert(!stack.empty() && stack.back() == 'a');
stack.pop_back();
return true;
}
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const nlohmann::detail::exception& /*ex*/) override
{
++errors;
return true;
}
bool complete() const
{
return stack.empty();
}
std::size_t events = 0;
std::size_t errors = 0;
private:
bool value()
{
++events;
if (!stack.empty())
{
// an array element, or the value of a key
assert(stack.back() != 'o');
if (stack.back() == 'v')
{
stack.back() = 'o';
}
}
return true;
}
// 'a' for an array, 'o' for an object that expects a key, 'v' for an
// object that expects the value of a key
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
};
/// parses @a data with a recovering_checker from memory and from a stream,
/// checks that both see the same, that the events are balanced, and that the
/// number of errors is bounded, and returns the checker (see #3989)
inline recovering_checker check_recovering_parse(const std::uint8_t* data, const std::size_t size, const nlohmann::json::input_format_t format)
{
recovering_checker checker;
const bool ok = nlohmann::json::sax_parse(data, data + size, &checker, format);
assert(checker.complete());
assert(checker.errors <= size + 1);
assert(ok == (checker.errors == 0));
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
recovering_checker stream_checker;
assert(nlohmann::json::sax_parse(stream, &stream_checker, format) == ok);
assert(stream_checker.complete());
assert(stream_checker.events == checker.events);
assert(stream_checker.errors == checker.errors);
return checker;
}
} // namespace
+71
View File
@@ -479,6 +479,77 @@ TEST_CASE("deep copy uses the provided allocator")
CHECK(copy == j); CHECK(copy == j);
} }
namespace
{
// the number of constructions countdown_allocator lets happen, including the
// one that fails; 0 means none ever fails
std::size_t constructions_until_failure = 0;
template<class T>
struct countdown_allocator : std::allocator<T>
{
using std::allocator<T>::allocator;
template<class U, class... Args>
void construct(U* p, Args&& ... args)
{
if (constructions_until_failure != 0 && --constructions_until_failure == 0)
{
throw std::bad_alloc();
}
::new (static_cast<void*>(p)) U(std::forward<Args>(args)...);
}
template <class U>
struct rebind
{
using other = countdown_allocator<U>;
};
};
} // namespace
TEST_CASE("converting a deeply nested value from another specialization fails cleanly (#5650)")
{
using countdown_json = nlohmann::basic_json<std::map,
std::vector,
std::string,
bool,
std::int64_t,
std::uint64_t,
double,
countdown_allocator>;
// deeper than the 128 levels the converting constructor descends into, so
// that failures land on both sides of the bound - or, built with
// JSON_NO_THREAD_LOCAL, all in the iterative conversion
json j = {1, "two", {{"three", 3}}};
for (std::size_t i = 0; i < 150; ++i)
{
j = json{{"a", json::array({j, "sibling"})}};
}
// Fail every construction in turn. Each failure has to reach the caller,
// and everything built until then has to be destroyed cleanly.
std::size_t failures = 0;
for (std::size_t n = 1;; ++n)
{
constructions_until_failure = n;
try
{
const countdown_json converted = j;
constructions_until_failure = 0;
CHECK(converted.dump() == j.dump());
break;
}
catch (const std::bad_alloc&)
{
++failures;
}
}
CHECK(failures > 0);
}
namespace namespace
{ {
template<class T> template<class T>
+40
View File
@@ -419,6 +419,46 @@ TEST_CASE("alternative string type")
CHECK(j2.dump() == R"({"/foo/0":"bar","/foo/1":"baz"})"); CHECK(j2.dump() == R"({"/foo/0":"bar","/foo/1":"baz"})");
} }
SECTION("error recovery")
{
// a SAX parser that recovers from every error (see #3989)
struct recovering_parser : nlohmann::detail::json_sax_dom_parser<alt_json>
{
explicit recovering_parser(alt_json& j)
: nlohmann::detail::json_sax_dom_parser<alt_json>(j, false)
{}
// sax_parse() calls the SAX parser's own parse_error(), so hiding
// the one of the base class is what recovering takes
// NOLINTNEXTLINE(bugprone-derived-method-shadowing-base-method)
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const nlohmann::detail::exception& /*unused*/)
{
++errors;
return true;
}
std::size_t errors = 0;
};
alt_json j;
recovering_parser sax(j);
// not inside CHECK(): MSVC reads the escape in a stringized raw string
const std::string input = R"([1., "a\qb", tru, {"k" 2}])";
CHECK(!alt_json::sax_parse(input, &sax));
CHECK(sax.errors == 4);
CHECK(j.dump() == R"([1,"aqb",null,{"k":2}])");
// a UBJSON high-precision number, a CBOR key that is not a string
alt_json u;
recovering_parser ubjson_sax(u);
CHECK(!alt_json::sax_parse(std::vector<std::uint8_t> {'[', 'H', 'i', 2, '1', '.', ']'}, &ubjson_sax, alt_json::input_format_t::ubjson));
CHECK(u.dump() == "[1]");
alt_json c;
recovering_parser cbor_sax(c);
CHECK(!alt_json::sax_parse(std::vector<std::uint8_t> {0xA2, 0x01, 0x02, 0x61, 'a', 0x03}, &cbor_sax, alt_json::input_format_t::cbor));
CHECK(c.dump() == R"({"a":3})");
}
SECTION("strict enum") SECTION("strict enum")
{ {
// regression test for #5667: NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's from_json // regression test for #5667: NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's from_json
+30 -3
View File
@@ -210,15 +210,42 @@ TEST_CASE_TEMPLATE_INVOKE(value_in_range_of_test, \
TEST_CASE("BJData") TEST_CASE("BJData")
{ {
SECTION("binary_reader BJData LUT arrays are sorted") SECTION("binary_reader BJData lookup tables")
{ {
std::vector<std::uint8_t> const data; std::vector<std::uint8_t> const data;
auto ia = nlohmann::detail::input_adapter(data); auto ia = nlohmann::detail::input_adapter(data);
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) // NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
nlohmann::detail::binary_reader<json, decltype(ia)> const br{std::move(ia), json::input_format_t::bjdata}; nlohmann::detail::binary_reader<json, decltype(ia)> const br{std::move(ia), json::input_format_t::bjdata};
CHECK(std::is_sorted(br.bjd_optimized_type_markers.begin(), br.bjd_optimized_type_markers.end())); // the excluded optimized-type markers must match binary_writer's
CHECK(std::is_sorted(br.bjd_types_map.begin(), br.bjd_types_map.end())); // is_bjdata_excluded_type_marker(), which encodes the same 8 markers
for (const char marker :
{'[', '{', 'S', 'H', 'T', 'F', 'N', 'Z'
})
{
CHECK(br.is_bjd_excluded_optimized_type(marker));
}
for (const char marker :
{'U', 'i', 'u', 'I', 'm', 'l', 'M', 'L', 'd', 'D', 'C', 'B', 'x'
})
{
CHECK(!br.is_bjd_excluded_optimized_type(marker));
}
// every dtype marker must round-trip to its ND-array type name
const std::vector<std::pair<char, std::string>> types
{
{'B', "byte"}, {'C', "char"}, {'D', "double"}, {'I', "int16"},
{'L', "int64"}, {'M', "uint64"}, {'U', "uint8"}, {'d', "single"},
{'i', "int8"}, {'l', "int32"}, {'m', "uint32"}, {'u', "uint16"}
};
for (const auto& type : types)
{
const char* name = br.bjd_type_name(type.first);
REQUIRE(name != nullptr);
CHECK(std::string(name) == type.second);
}
CHECK(br.bjd_type_name('x') == nullptr);
} }
SECTION("individual values") SECTION("individual values")
+583 -1
View File
@@ -143,11 +143,13 @@ class SaxEventLogger
{ {
errored = true; errored = true;
events.push_back("parse_error(" + std::to_string(position) + ")"); events.push_back("parse_error(" + std::to_string(position) + ")");
return false; return recover;
} }
std::vector<std::string> events {}; // NOLINT(readability-redundant-member-init) std::vector<std::string> events {}; // NOLINT(readability-redundant-member-init)
bool errored = false; bool errored = false;
/// whether parse_error() asks the parser to recover from the error (see #3989)
bool recover = false;
}; };
class SaxCountdown : public nlohmann::json::json_sax_t class SaxCountdown : public nlohmann::json::json_sax_t
@@ -2937,3 +2939,583 @@ TEST_CASE("diagnostic positions: value lifetime, input adapters, and SAX")
} }
} }
#endif #endif
namespace
{
/// builds a value like json::parse(), but asks the parser to recover from
/// errors (see #3989), and checks that the events it receives are balanced
class RecoveringDomParser
{
public:
explicit RecoveringDomParser(json& j, std::size_t max_errors_ = static_cast<std::size_t>(-1))
: dom(j, false)
, max_errors(max_errors_)
{}
bool null()
{
value();
return dom.null();
}
bool boolean(bool val)
{
value();
return dom.boolean(val);
}
bool number_integer(json::number_integer_t val)
{
value();
return dom.number_integer(val);
}
bool number_unsigned(json::number_unsigned_t val)
{
value();
return dom.number_unsigned(val);
}
bool number_float(json::number_float_t val, const std::string& s)
{
value();
return dom.number_float(val, s);
}
bool string(std::string& val)
{
value();
return dom.string(val);
}
bool binary(json::binary_t& val)
{
value();
return dom.binary(val);
}
bool start_object(std::size_t elements)
{
value();
stack.push_back('o');
return dom.start_object(elements);
}
bool key(std::string& val)
{
++events;
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.back() = 'v';
return dom.key(val);
}
bool end_object()
{
++events;
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_object();
}
bool start_array(std::size_t elements)
{
value();
stack.push_back('a');
return dom.start_array(elements);
}
bool end_array()
{
++events;
if (stack.empty() || stack.back() != 'a')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_array();
}
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& ex)
{
errors.emplace_back(ex.what());
return errors.size() < max_errors;
}
/// whether the events were balanced and every key was followed by a value
bool balanced() const
{
return well_formed && stack.empty();
}
/// builds the value
nlohmann::detail::json_sax_dom_parser<json> dom;
std::vector<std::string> errors {}; // NOLINT(readability-redundant-member-init)
std::size_t events = 0;
/// the open containers: 'a' for an array, 'o' for an object that expects
/// a key, 'v' for an object that expects the value of a key
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
bool well_formed = true;
std::size_t max_errors;
private:
/// a value is passed: it is an array element, or the value of a key
void value()
{
++events;
if (!stack.empty())
{
if (stack.back() == 'v')
{
stack.back() = 'o';
}
else if (stack.back() == 'o')
{
// a value without a key
well_formed = false;
}
}
}
};
struct RecoveryResult
{
json value;
std::vector<std::string> errors;
std::size_t events;
bool ok;
bool balanced;
};
template<typename InputType>
RecoveryResult parse_recovering(InputType&& input, const bool strict = true,
const bool ignore_comments = false, const bool ignore_trailing_commas = false)
{
json j;
RecoveringDomParser sax(j);
const bool ok = json::sax_parse(std::forward<InputType>(input), &sax, json::input_format_t::json,
strict, ignore_comments, ignore_trailing_commas);
return {j, sax.errors, sax.events, ok, sax.balanced()};
}
/// stops after a number of events, but recovers from errors
class RecoveringCountdown : public SaxCountdown
{
public:
using SaxCountdown::SaxCountdown;
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const json::exception& /*ex*/) override
{
return true;
}
};
/// a repaired input: the value it is repaired to, and the number of errors
struct Repair
{
const char* input;
const char* expected;
std::size_t errors;
};
} // namespace
TEST_CASE("parser error recovery (#3989)")
{
SECTION("repairs")
{
const std::vector<Repair> repairs =
{
// a missing separator is inserted
{"[1 2]", "[1,2]", 1},
{R"({"a":1 "b":2})", R"({"a":1,"b":2})", 1},
{R"({"a" 1})", R"({"a":1})", 1},
{"[1 tru 2]", "[1,null,2]", 2},
{R"({"a" "b": 1})", R"({"a":"b"})", 2},
// a missing value is null in an object; in an array, a ',' stands
// for null, while an array that ends there just ends
{R"({"a":})", R"({"a":null})", 1},
{R"({"a"})", R"({"a":null})", 1},
{R"({"a","b":1})", R"({"a":null,"b":1})", 1},
{"[1,,2]", "[1,null,2]", 1},
{"[,1]", "[null,1]", 1},
{"[1,]", "[1]", 1},
{"[1,2,3,]", "[1,2,3]", 1},
{R"({"a":1,})", R"({"a":1})", 1},
// a broken string keeps what can be read
{R"(["a\qb"])", R"(["aqb"])", 1},
{R"({"na\me":1})", R"({"name":1})", 1},
{"[\"\xFF\"]", R"(["\uFFFD"])", 1},
{"[\"a\xC3(\"]", R"(["a\uFFFD("])", 1},
{"[\"\xE2\x82\"]", R"(["\uFFFD"])", 1},
{"[\"\xC3\\\\\", 1]", R"(["\uFFFD\\",1])", 1},
{R"(["\u12"])", R"(["\uFFFD"])", 1},
{R"(["\u12G4"])", R"(["\uFFFDG4"])", 1},
{R"(["\uDC00x"])", R"(["\uFFFDx"])", 1},
{R"(["\uD800x"])", R"(["\uFFFDx"])", 1},
{R"(["\uD800\u0041"])", R"(["\uFFFDA"])", 1},
{R"(["\uD800\uD800\uDC00"])", R"(["\uFFFD\uD800\uDC00"])", 1},
{R"(["\uD800\uD800\uD800x"])", R"(["\uFFFD\uFFFD\uFFFDx"])", 1},
{
R"(["\uD800\"x", 1])", R"(["\uFFFD\"x",1])", 1
},
{R"(["\uD800\q"])", R"(["\uFFFDq"])", 1},
{"[\"a\tb\"]", R"(["a\tb"])", 1},
{R"(["a\qb\u0041\x"])", R"(["aqbAx"])", 1},
// a broken number keeps its longest valid prefix
{"[1.]", "[1]", 1},
{"[-2.]", "[-2]", 1},
{"[1.5e]", "[1.5]", 1},
{"[1e+]", "[1]", 1},
{"[1.x2, 3]", "[1,3]", 1},
// what cannot be read at all is null
{"[1,NaN,3]", "[1,null,3]", 1},
{"[tru]", "[null]", 1},
{"[-]", "[null]", 1},
{R"({"a":Infinity})", R"({"a":null})", 1},
// a stray token is dropped
{"[:1]", "[1]", 1},
{R"(["a":1])", R"(["a",1])", 1},
{R"({"a"::1})", R"({"a":1})", 1},
// a member that cannot be read is skipped
{R"({1:2,"b":3})", R"({"b":3})", 1},
{R"({"a":1 2})", R"({"a":1})", 1},
{R"({,"a":1})", R"({"a":1})", 1},
{R"({"a":1,,"b":2})", R"({"a":1,"b":2})", 1},
{"{a:1}", "{}", 1},
{R"({"a":1 [1,{"b":2}], "c":3})", R"({"a":1,"c":3})", 1},
{R"([{1}, "a"])", R"([{},"a"])", 1},
// a wrong closing bracket closes the innermost container
{R"({"a":[1,2}, "b":3})", R"({"a":[1,2],"b":3})", 1},
{R"([{"a":1], 2])", R"([{"a":1},2])", 1},
{"{]", "{}", 1},
{"[}", "[]", 1},
// the end of the input closes all containers
{R"({"a":[1,2)", R"({"a":[1,2]})", 1},
{"[", "[]", 1},
{"{", "{}", 1},
{R"({"a")", R"({"a":null})", 1},
{R"({"a":)", R"({"a":null})", 1},
{"[1,", "[1]", 1},
{"[[[1", "[[[1]]]", 1},
{
R"(["abc)", R"(["abc"])", 2
},
{"[1,tr", "[1,null]", 2},
{"\"abc", "\"abc\"", 1},
{"[\"ab\ncd\"]", R"(["ab",null,"]"])", 4},
// what comes before the top-level value is skipped
{")]}'\n{\"a\":1}", R"({"a":1})", 1},
{R"(data: {"a":1})", R"({"a":1})", 1},
{"\xEF\xBB[1]", "[1]", 1},
// what comes after it is an error that ends parsing
{R"({"a":1}})", R"({"a":1})", 1},
{"[1}]", "[1]", 2},
{"[1] [2]", "[1]", 1},
};
for (const auto& repair : repairs)
{
CAPTURE(repair.input);
const auto result = parse_recovering(std::string(repair.input));
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.value == json::parse(repair.expected));
CHECK(result.errors.size() == repair.errors);
}
}
SECTION("number overflow")
{
const auto result = parse_recovering(std::string("[1e999,-1e999]"));
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.errors.size() == 2);
CHECK(result.errors[0] == "[json.exception.out_of_range.406] number overflow parsing '1e999'");
REQUIRE(result.value.size() == 2);
CHECK(result.value[0].is_number_float());
CHECK(result.value[0].get<double>() == std::numeric_limits<double>::infinity());
CHECK(result.value[1].get<double>() == -std::numeric_limits<double>::infinity());
// the SAX parser gets the number's text
SaxEventLogger logger;
logger.recover = true;
CHECK(!json::sax_parse("1e999", &logger));
CHECK(logger.events == std::vector<std::string>({"parse_error(5)", "number_float(1e999)"}));
}
SECTION("nothing to recover")
{
for (const std::string s :
{
"", " ", "]", "tru", "NaN", ",:", "/* comment"
})
{
CAPTURE(s);
const auto result = parse_recovering(s, true, true);
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.events == 0);
CHECK(result.value == nullptr);
CHECK(result.errors.size() == 1);
}
}
SECTION("error messages")
{
// the first error is reported as without recovery
for (const std::string s :
{
"[1 2]", R"({"a":1 "b":2})", R"({"a" 1})", R"({"a":})", "[1,]", "[1.]",
R"(["a\qb"])", "[1e999]", "{1:2}", R"({"a":[1,2}})", "[1,", "[1] [2]", "{a:1}"
})
{
CAPTURE(s);
const auto result = parse_recovering(s);
REQUIRE(!result.errors.empty());
json _;
CHECK_THROWS_WITH_STD_STR(_ = json::parse(s), result.errors.front());
}
// the token of an error begins where the previous error was
const auto result = parse_recovering(std::string("[tru, fals, nul]"));
CHECK(result.errors == std::vector<std::string>(
{
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: '[tru,'",
"[json.exception.parse_error.101] parse error at line 1, column 11: syntax error while parsing value - invalid literal; last read: ', fals,'",
"[json.exception.parse_error.101] parse error at line 1, column 16: syntax error while parsing value - invalid literal; last read: ', nul]'"
}));
CHECK(result.value == json::parse("[null,null,null]"));
}
SECTION("events")
{
// see #4522
SaxEventLogger logger;
logger.recover = true;
CHECK(!json::sax_parse(R"([{1}, "a"])", &logger));
CHECK(logger.events == std::vector<std::string>(
{
"start_array()", "start_object()", "parse_error(3)", "end_object()", "string(a)", "end_array()"
}));
}
SECTION("options")
{
SECTION("strict")
{
const auto result = parse_recovering(std::string("[1 2] [3]"), false);
CHECK(!result.ok);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
}
SECTION("ignore_trailing_commas")
{
for (const std::string s :
{
"[1,]", R"({"a":1,})", "[[1,],]"
})
{
CAPTURE(s);
const auto result = parse_recovering(s, true, false, true);
CHECK(result.ok);
CHECK(result.errors.empty());
}
auto result = parse_recovering(std::string("[1,,]"), true, false, true);
CHECK(result.value == json::parse("[1,null]"));
CHECK(result.errors.size() == 1);
result = parse_recovering(std::string(R"({"a":1,,})"), true, false, true);
CHECK(result.value == json::parse(R"({"a":1})"));
CHECK(result.errors.size() == 1);
}
SECTION("ignore_comments")
{
auto result = parse_recovering(std::string("[1 /* one */ 2]"), true, true);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
// a comment that is not closed runs to the end of the input, which
// is not reported again
result = parse_recovering(std::string("[1, 2 /* unterminated"), true, true);
CHECK(result.balanced);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
// a '/' that does not begin a comment is garbage
result = parse_recovering(std::string("[1, /x, 2]"), true, true);
CHECK(result.balanced);
CHECK(result.value == json::parse("[1,null,2]"));
CHECK(result.errors.size() == 1);
}
}
SECTION("null bytes")
{
// a null byte ends the input, unless JSON_STRICT_NUL_HANDLING is set
const auto result = parse_recovering(std::string("[1,\0x", 5));
CHECK(result.balanced);
CHECK(!result.ok);
#ifdef JSON_TEST_STRICT_NUL_HANDLING_ENABLED
CHECK(result.value == json::parse("[1,null]"));
#else
CHECK(result.value == json::parse("[1]"));
CHECK(result.errors.size() == 1);
#endif
const auto in_string = parse_recovering(std::string("[\"a\0b\"]", 7));
CHECK(in_string.balanced);
#ifdef JSON_TEST_STRICT_NUL_HANDLING_ENABLED
CHECK(in_string.value == json::array({std::string("a\0b", 3)}));
#else
CHECK(in_string.value == json::parse(R"(["a"])"));
#endif
}
SECTION("the SAX parser stops recovering")
{
json j;
RecoveringDomParser sax(j, 2);
CHECK(!json::sax_parse("[1 2 3 4 5]", &sax));
CHECK(sax.errors.size() == 2);
// an error at a delimiter that an invalid token consumed is reported
// to the SAX parser, too
json j2;
RecoveringDomParser sax2(j2, 2);
CHECK(!json::sax_parse("[tru}, 1]", &sax2));
CHECK(sax2.errors.size() == 2);
}
SECTION("an event stops parsing during a repair")
{
// start_object() and key() are passed, then null() for the missing
// value returns false
RecoveringCountdown countdown(2);
CHECK(!json::sax_parse(R"({"a":})", &countdown));
// the end of the input: end_array() for the second array returns false
RecoveringCountdown countdown2(4);
CHECK(!json::sax_parse("[[1", &countdown2));
}
SECTION("input adapters")
{
// the lexer reads contiguous and streaming input differently, and it
// puts back a character that ended an invalid token
for (const std::string s :
{
"[1 2]", "[tru}, 1]", R"({"a" "b\q", "c":[1.x, 2}})", "[\"\xFF\xC3(\", -, 1e+]", "{a:1,\"b\":2", ")]}' [1]"
})
{
CAPTURE(s);
const auto reference = parse_recovering(s);
CHECK(reference.balanced);
const auto from_c_string = parse_recovering(s.c_str());
CHECK(from_c_string.value == reference.value);
CHECK(from_c_string.errors == reference.errors);
const std::list<char> l(s.begin(), s.end());
json j;
RecoveringDomParser sax(j);
CHECK(!json::sax_parse(l.begin(), l.end(), &sax));
CHECK(j == reference.value);
CHECK(sax.errors == reference.errors);
std::istringstream ss(s);
const auto from_stream = parse_recovering(ss);
CHECK(from_stream.value == reference.value);
CHECK(from_stream.errors == reference.errors);
}
}
SECTION("long runs of errors")
{
// no error may copy all the input read before it
const auto closing = parse_recovering("[" + std::string(100000, '}'));
CHECK(closing.balanced);
CHECK(closing.value == json::array());
const auto garbage = parse_recovering("[" + std::string(100000, 'x') + "]");
CHECK(garbage.balanced);
CHECK(garbage.errors.size() == 1);
const auto commas = parse_recovering("{" + std::string(100000, ',') + "}");
CHECK(commas.balanced);
CHECK(commas.value == json::object());
}
SECTION("mutations of valid input")
{
// whatever the input, the events are balanced, every error is reported
// at most once, and valid input is parsed as usual
const std::vector<std::string> documents =
{
R"({"name": "value", "list": [1, -2.5, true, null, {"x": [[]]}], "e": "\u00e9"})",
R"([{"a": [1, 2, {"b": "c"}]}, [], {}, "\ud83d\ude00", 1e10])",
"{\"\xC3\xA9\": \"\xF0\x9F\x98\x80\"}",
R"( {"k" : [ "v" , 0 ] } )",
};
// each character that can be inserted, including a null byte
const std::string insertions("[]{},:\"x\\\0\xFF", 11);
std::vector<std::string> inputs;
for (const auto& doc : documents)
{
for (std::size_t i = 0; i <= doc.size(); ++i)
{
inputs.push_back(doc.substr(0, i));
if (i < doc.size())
{
inputs.push_back(doc.substr(0, i) + doc.substr(i + 1));
}
for (const char c : insertions)
{
inputs.push_back(doc.substr(0, i) + c + doc.substr(i));
}
}
}
for (const auto& s : inputs)
{
CAPTURE(s);
const auto result = parse_recovering(s);
CHECK(result.balanced);
CHECK(result.errors.size() <= s.size() + 1);
CHECK(result.events <= (4 * s.size()) + 4);
if (json::accept(s))
{
CHECK(result.ok);
CHECK(result.errors.empty());
CHECK(result.value == json::parse(s));
}
else
{
CHECK(!result.ok);
CHECK(!result.errors.empty());
}
}
}
}
+52
View File
@@ -141,6 +141,58 @@ TEST_CASE("Better diagnostics with positions")
check_objects(300); check_objects(300);
} }
SECTION("converting keeps the positions of nested values (#5650)")
{
// Values nested deeper than the converting constructor's descent bound
// are converted without the call stack, on a path that has to carry the
// positions of every value over itself. Objects and arrays take turns,
// and the innermost value is null, which used to lose its positions.
const auto check_conversion = [](std::size_t depth)
{
CAPTURE(depth)
std::string text;
std::string closing;
for (std::size_t i = 0; i < depth; ++i)
{
text += (i % 2 == 0) ? "[12, " : R"({"b":1, "a":)";
closing += (i % 2 == 0) ? ']' : '}';
}
text += "null";
text.append(closing.rbegin(), closing.rend());
const json original = json::parse(text);
const nlohmann::ordered_json converted = original;
const json* o = &original;
const nlohmann::ordered_json* c = &converted;
for (std::size_t level = 0; level <= depth; ++level)
{
CAPTURE(level)
REQUIRE(c->start_pos() == o->start_pos());
REQUIRE(c->end_pos() == o->end_pos());
if (level < depth)
{
// the number beside the value nested next
const json& o_number = o->is_object() ? o->at("b") : o->at(0);
const nlohmann::ordered_json& c_number = c->is_object() ? c->at("b") : c->at(0);
REQUIRE(c_number.start_pos() == o_number.start_pos());
REQUIRE(c_number.end_pos() == o_number.end_pos());
o = o->is_object() ? &o->at("a") : &o->at(1);
c = c->is_object() ? &c->at("a") : &c->at(1);
}
}
};
check_conversion(1);
check_conversion(127);
check_conversion(128);
check_conversion(129);
check_conversion(300);
}
SECTION("JSON patch add to primitive parent (#4292)") SECTION("JSON patch add to primitive parent (#4292)")
{ {
// the JSON Patch "add" target /foo/bar/baz has a string parent // the JSON Patch "add" target /foo/bar/baz has a string parent
+30
View File
@@ -341,6 +341,36 @@ TEST_CASE("Regression tests for extended diagnostics")
} }
} }
SECTION("Regression test for issue #5650 - converting keeps the parents of nested values")
{
// A value nested deeper than the converting constructor's descent bound
// is converted without the call stack. Every container that path creates
// has to have the parents of its children set, or the JSON Pointer in the
// diagnostic is cut short. Objects and arrays take turns.
const std::size_t pairs = 150;
json j = "not a number";
std::string pointer;
for (std::size_t i = 0; i < pairs; ++i)
{
j = json{{"a", json::array({j})}};
pointer += "/a/0";
}
const nlohmann::ordered_json converted = j;
const nlohmann::ordered_json* inner = &converted;
for (std::size_t i = 0; i < pairs; ++i)
{
inner = &inner->at("a").at(0);
}
std::string const expected = "[json.exception.type_error.302] (" + pointer + ") type must be number, but is string";
int i = 0;
CHECK_THROWS_WITH_AS(i = inner->get<int>(), expected.c_str(), nlohmann::ordered_json::type_error);
CHECK(i == 0);
}
SECTION("Regression test for issue #5668 - wrong path for std::map/unordered_map with non-string keys") SECTION("Regression test for issue #5668 - wrong path for std::map/unordered_map with non-string keys")
{ {
// a map with non-string keys is read from an array of [key, value] arrays; // a map with non-string keys is read from an array of [key, value] arrays;
@@ -0,0 +1,93 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// This file tests the opt-in JSON_DISABLE_TUPLE_REFERENCE_CONVERSION, so it
// defines the macro itself rather than relying on a -D flag, and runs in every
// build.
#ifdef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
#undef JSON_DISABLE_TUPLE_REFERENCE_CONVERSION
#endif
#define JSON_DISABLE_TUPLE_REFERENCE_CONVERSION 1
#include <nlohmann/json.hpp>
using nlohmann::json;
using nlohmann::ordered_json;
#include <string>
#include <tuple>
#include <type_traits>
#include <utility>
// clang before 4 and GCC before 5 cannot create a std::tuple of basic_json
// references at all, with or without JSON_DISABLE_TUPLE_REFERENCE_CONVERSION:
// the tuple constructors make them instantiate basic_json's conversion operator
// for libstdc++'s internal tuple bases, which fails hard
#if (defined(__clang__) && __clang_major__ < 4) || (!defined(__clang__) && defined(__GNUC__) && __GNUC__ < 5)
#define SKIP_TESTS_FOR_JSON_REFERENCE_TUPLES
#endif
TEST_CASE("JSON_DISABLE_TUPLE_REFERENCE_CONVERSION")
{
SECTION("json is not constructible from a one-element tuple of a json reference")
{
CHECK_FALSE(std::is_constructible<json, std::tuple<json&>>::value);
CHECK_FALSE(std::is_constructible<json, std::tuple<const json&>>::value);
CHECK_FALSE(std::is_constructible < json, std::tuple < json && >>::value);
CHECK_FALSE(std::is_constructible<json, const std::tuple<json&>&>::value);
CHECK_FALSE(std::is_constructible<ordered_json, std::tuple<ordered_json&>>::value);
}
#ifndef SKIP_TESTS_FOR_JSON_REFERENCE_TUPLES
SECTION("issue #2226 - tuple<const json&> from tuple<json&> keeps the reference")
{
json j = true;
const std::tuple<const json&> tup(std::forward_as_tuple(j));
CHECK(&std::get<0>(tup) == &j);
}
SECTION("tuple<json> from tuple<json&> copies the element")
{
const json j = {{"key", "value"}};
const std::tuple<json> t1(std::forward_as_tuple(j));
CHECK(std::get<0>(t1) == j);
json j2 = "text";
const std::tuple<json> t2(std::forward_as_tuple(std::move(j2)));
CHECK(std::get<0>(t2) == "text");
}
#endif
SECTION("other tuple conversions are not affected")
{
const json j = true;
// one-element tuple holding a json value
CHECK(json(std::make_tuple(j)) == json::array({true}));
// tuples with more than one element, even when holding references
int i = 1;
#ifndef SKIP_TESTS_FOR_JSON_REFERENCE_TUPLES
CHECK(json(std::forward_as_tuple(i, j)) == json::array({1, true}));
CHECK(json(std::forward_as_tuple(j, j)) == json::array({true, true}));
#endif
// one-element tuples holding references to other types
std::string s = "text";
CHECK(json(std::forward_as_tuple(s)) == json::array({"text"}));
CHECK(json(std::forward_as_tuple(i)) == json::array({1}));
#ifndef SKIP_TESTS_FOR_JSON_REFERENCE_TUPLES
// a reference to a different basic_json specialization
ordered_json oj = true;
CHECK(json(std::forward_as_tuple(oj)) == json::array({true}));
#endif
}
}
+128
View File
@@ -13,6 +13,7 @@ using nlohmann::json;
#include <algorithm> #include <algorithm>
#include <string> #include <string>
#include <vector>
TEST_CASE("tests on very large JSONs") TEST_CASE("tests on very large JSONs")
{ {
@@ -53,6 +54,24 @@ const json* innermost_value(const json& j, std::size_t& depth)
return current; return current;
} }
// The text of a value nested depth levels deep around the number 0. Level i is
// an array if pattern[i % pattern.size()] is '[', and otherwise an object with
// the single member "a", which every object type enumerates in the same order.
std::string nested_text(std::size_t depth, const std::string& pattern)
{
std::string text;
std::string closing;
for (std::size_t i = 0; i < depth; ++i)
{
const bool array = pattern[i % pattern.size()] == '[';
text += array ? "[" : "{\"a\":";
closing += array ? ']' : '}';
}
text += '0';
text.append(closing.rbegin(), closing.rend());
return text;
}
} // namespace } // namespace
TEST_CASE("tests on deeply nested JSONs") TEST_CASE("tests on deeply nested JSONs")
@@ -224,5 +243,114 @@ TEST_CASE("tests on deeply nested JSONs")
CHECK(*innermost_value(j, unused) == 0); CHECK(*innermost_value(j, unused) == 0);
} }
} }
SECTION("issue #5650 - stack overflow converting between specializations")
{
const std::vector<std::string> patterns = {"[", "{", "[{"};
SECTION("json to ordered_json")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const json j = json::parse(text);
const nlohmann::ordered_json converted = j;
CHECK(converted.dump() == text);
}
}
SECTION("ordered_json to json")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const nlohmann::ordered_json o = nlohmann::ordered_json::parse(text);
const json converted = o;
CHECK(converted.dump() == text);
}
}
SECTION("get<ordered_json>()")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const json j = json::parse(text);
CHECK(j.get<nlohmann::ordered_json>().dump() == text);
}
}
SECTION("depths around the bound of the recursive descent")
{
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d);
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(d, pattern);
const json j = json::parse(text);
const nlohmann::ordered_json converted = j;
CHECK(converted.dump() == text);
const json back = converted;
CHECK(back.dump() == text);
}
}
}
SECTION("values below the bound are converted as values above it")
{
// Bury a value below the bound, where it is converted without the
// call stack, and compare it with the same value converted on its
// own by the containers' range constructors. Its objects have
// members that the two object types enumerate in different orders.
const auto bury = [](nlohmann::ordered_json value)
{
for (std::size_t i = 0; i < 200; ++i)
{
value = nlohmann::ordered_json::array({std::move(value)});
}
return value;
};
const auto dig = [](const json & value)
{
const json* current = &value;
for (std::size_t i = 0; i < 200; ++i)
{
current = &current->at(0);
}
return current;
};
nlohmann::ordered_json value = nlohmann::ordered_json::object();
value["z"] = {1, -2, 3U, 4.5, true, nullptr, "six", nlohmann::ordered_json::binary({7, 8}, 9),
nlohmann::ordered_json::binary({10}), nlohmann::ordered_json::array(), nlohmann::ordered_json::object()
};
value["y"] = {{"x", {{"w", 1}, {"v", 2}}}, {"u", {3, {{"t", 4}, {"s", 5}}}}};
value["r"] = nlohmann::ordered_json::array({nlohmann::ordered_json(nlohmann::ordered_json::value_t::discarded)});
const json converted_above = value;
const json buried = bury(value);
const json& converted_below = *dig(buried);
CHECK(converted_below.dump() == converted_above.dump());
CHECK(converted_below.at("z").at(7).get_binary().subtype() == 9);
CHECK_FALSE(converted_below.at("z").at(8).get_binary().has_subtype());
CHECK(converted_below.at("r").at(0).is_discarded());
// a discarded value is never equal to anything, so compare the rest
value.erase("r");
const json without_discarded_above = value;
const json without_discarded_buried = bury(value);
CHECK(*dig(without_discarded_buried) == without_discarded_above);
}
}
} }
+609
View File
@@ -18,6 +18,19 @@
// for some reason including this after the json header leads to linker errors with VS 2017... // for some reason including this after the json header leads to linker errors with VS 2017...
#include <locale> #include <locale>
// skip tests if JSON_DISABLE_TUPLE_REFERENCE_CONVERSION=1 (#2226)
#if defined(JSON_DISABLE_TUPLE_REFERENCE_CONVERSION) && (JSON_DISABLE_TUPLE_REFERENCE_CONVERSION == 1)
#define SKIP_TESTS_FOR_TUPLE_REFERENCE_CONVERSION
#endif
// clang before 4 and GCC before 5 cannot create a std::tuple of basic_json
// references at all, with or without JSON_DISABLE_TUPLE_REFERENCE_CONVERSION:
// the tuple constructors make them instantiate basic_json's conversion operator
// for libstdc++'s internal tuple bases, which fails hard
#if (defined(__clang__) && __clang_major__ < 4) || (!defined(__clang__) && defined(__GNUC__) && __GNUC__ < 5)
#define SKIP_TESTS_FOR_JSON_REFERENCE_TUPLES
#endif
#define JSON_TESTS_PRIVATE #define JSON_TESTS_PRIVATE
#include <nlohmann/json.hpp> #include <nlohmann/json.hpp>
using json = nlohmann::json; using json = nlohmann::json;
@@ -28,6 +41,7 @@ using ordered_json = nlohmann::ordered_json;
#include <cstdio> #include <cstdio>
#include <list> #include <list>
#include <tuple>
#include <type_traits> #include <type_traits>
#include <utility> #include <utility>
@@ -542,6 +556,20 @@ TEST_CASE("regression tests 2")
))); )));
} }
#ifndef SKIP_TESTS_FOR_TUPLE_REFERENCE_CONVERSION
SECTION("issue #2226 - std::tuple dangling reference - implicit conversion")
{
// by default, a one-element tuple holding a json reference converts to
// a one-element array; JSON_DISABLE_TUPLE_REFERENCE_CONVERSION removes
// this conversion (see unit-disable-tuple-reference-conversion.cpp)
const json j = true;
CHECK(std::is_constructible<json, std::tuple<const json&>>::value);
#ifndef SKIP_TESTS_FOR_JSON_REFERENCE_TUPLES
CHECK(json(std::forward_as_tuple(j)) == json::array({true}));
#endif
}
#endif
SECTION("PR #2181 - regression bug with lvalue") SECTION("PR #2181 - regression bug with lvalue")
{ {
// see https://github.com/nlohmann/json/pull/2181#issuecomment-653326060 // see https://github.com/nlohmann/json/pull/2181#issuecomment-653326060
@@ -870,4 +898,585 @@ TEST_CASE("regression test - excessive binary container size honors allow_except
CHECK(json::from_cbor(std::vector<std::uint8_t> {0x9b, 0, 0, 0, 0, 0, 0, 0, 0x02}, true, false).is_discarded()); CHECK(json::from_cbor(std::vector<std::uint8_t> {0x9b, 0, 0, 0, 0, 0, 0, 0, 0x02}, true, false).is_discarded());
} }
namespace
{
/// builds a value from SAX events, asks the parser to recover from its first
/// 100 errors, and checks that the events are balanced (see #3989)
class RecoveringParser
{
public:
explicit RecoveringParser(json& j)
: dom(j, false)
{}
bool null()
{
value();
return dom.null();
}
bool boolean(bool val)
{
value();
return dom.boolean(val);
}
bool number_integer(json::number_integer_t val)
{
value();
return dom.number_integer(val);
}
bool number_unsigned(json::number_unsigned_t val)
{
value();
return dom.number_unsigned(val);
}
bool number_float(json::number_float_t val, const std::string& s)
{
value();
return dom.number_float(val, s);
}
bool string(std::string& val)
{
value();
return dom.string(val);
}
bool binary(json::binary_t& val)
{
value();
return dom.binary(val);
}
bool start_object(std::size_t elements)
{
value();
stack.push_back('o');
return dom.start_object(elements);
}
bool key(std::string& val)
{
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.back() = 'v';
return dom.key(val);
}
bool end_object()
{
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_object();
}
bool start_array(std::size_t elements)
{
value();
stack.push_back('a');
return dom.start_array(elements);
}
bool end_array()
{
if (stack.empty() || stack.back() != 'a')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_array();
}
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& ex)
{
messages.emplace_back(ex.what());
// a limit, so that a reader that does not stop fails the test
// instead of making it hang
return ++errors < 100;
}
/// whether the events were balanced and every key was followed by a value
bool balanced() const
{
return well_formed && stack.empty();
}
/// builds the value
nlohmann::detail::json_sax_dom_parser<json> dom;
std::size_t errors = 0;
std::vector<std::string> messages {}; // NOLINT(readability-redundant-member-init)
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
bool well_formed = true;
private:
void value()
{
if (!stack.empty())
{
if (stack.back() == 'v')
{
stack.back() = 'o';
}
else if (stack.back() == 'o')
{
well_formed = false;
}
}
}
};
struct BinaryParseResult
{
json value;
std::size_t errors;
std::vector<std::string> messages;
bool ok;
bool balanced;
};
BinaryParseResult parse_binary_recovering(const std::vector<std::uint8_t>& input, const json::input_format_t format)
{
json j;
RecoveringParser sax(j);
const bool ok = json::sax_parse(input, &sax, format);
return {j, sax.errors, sax.messages, ok, sax.balanced()};
}
#if !defined(JSON_NOEXCEPTION)
/// the message of the exception that reading @a input into a JSON value
/// throws, or an empty string if reading succeeds
std::string binary_error_message(const std::vector<std::uint8_t>& input, const json::input_format_t format)
{
try
{
json _;
switch (format)
{
case json::input_format_t::cbor:
_ = json::from_cbor(input);
break;
case json::input_format_t::msgpack:
_ = json::from_msgpack(input);
break;
case json::input_format_t::ubjson:
_ = json::from_ubjson(input);
break;
case json::input_format_t::bjdata:
_ = json::from_bjdata(input);
break;
case json::input_format_t::bson:
_ = json::from_bson(input);
break;
case json::input_format_t::bon8:
_ = json::from_bon8(input);
break;
case json::input_format_t::json:
default:
break;
}
}
catch (const json::exception& e)
{
return e.what();
}
return "";
}
#endif
/// a BSON element: its type, its name, and its value
std::vector<std::uint8_t> bson_element(const std::uint8_t type, const std::string& name, const std::vector<std::uint8_t>& value)
{
std::vector<std::uint8_t> result = {type};
result.insert(result.end(), name.begin(), name.end());
result.push_back(0x00);
result.insert(result.end(), value.begin(), value.end());
return result;
}
/// a BSON document of the given elements; @a size_offset is added to the
/// size it declares
std::vector<std::uint8_t> bson_document(const std::vector<std::vector<std::uint8_t>>& elements, const int size_offset = 0)
{
std::vector<std::uint8_t> body;
for (const auto& element : elements)
{
body.insert(body.end(), element.begin(), element.end());
}
const auto size = static_cast<std::uint32_t>(static_cast<int>(body.size()) + 5 + size_offset);
std::vector<std::uint8_t> result = {static_cast<std::uint8_t>(size & 0xFFu), static_cast<std::uint8_t>((size >> 8u) & 0xFFu),
static_cast<std::uint8_t>((size >> 16u) & 0xFFu), static_cast<std::uint8_t>((size >> 24u) & 0xFFu)
};
result.insert(result.end(), body.begin(), body.end());
result.push_back(0x00);
return result;
}
/// a BSON int32 value
std::vector<std::uint8_t> bson_int32(const std::int32_t value)
{
const auto u = static_cast<std::uint32_t>(value);
return {static_cast<std::uint8_t>(u & 0xFFu), static_cast<std::uint8_t>((u >> 8u) & 0xFFu),
static_cast<std::uint8_t>((u >> 16u) & 0xFFu), static_cast<std::uint8_t>((u >> 24u) & 0xFFu)};
}
/// a BSON string value, whose length is @a length_offset off
std::vector<std::uint8_t> bson_string(const std::string& value, const std::int32_t length_offset = 0)
{
auto result = bson_int32(static_cast<std::int32_t>(value.size() + 1) + length_offset);
result.insert(result.end(), value.begin(), value.end());
result.push_back(0x00);
return result;
}
/// @a count bytes of value 0xAB
std::vector<std::uint8_t> bytes(const std::size_t count)
{
return std::vector<std::uint8_t>(count, 0xAB);
}
template<typename... Parts>
std::vector<std::uint8_t> concatenated(const std::vector<std::uint8_t>& first, const Parts& ... rest)
{
std::vector<std::uint8_t> result = first;
for (const auto& part : std::initializer_list<std::vector<std::uint8_t>> {rest...})
{
result.insert(result.end(), part.begin(), part.end());
}
return result;
}
/// U+FFFD REPLACEMENT CHARACTER
std::string replacement_character()
{
return "\xEF\xBF\xBD";
}
} // namespace
TEST_CASE("regression test - #3989 SAX parse_error() returning true")
{
SECTION("binary formats complete what was read before the input ends")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}};
const std::vector<std::pair<json::input_format_t, std::vector<std::uint8_t>>> encodings =
{
{json::input_format_t::cbor, json::to_cbor(j)},
{json::input_format_t::msgpack, json::to_msgpack(j)},
{json::input_format_t::ubjson, json::to_ubjson(j)},
{json::input_format_t::ubjson, json::to_ubjson(j, true, true)},
{json::input_format_t::bjdata, json::to_bjdata(j)},
{json::input_format_t::bjdata, json::to_bjdata(j, true, true)},
{json::input_format_t::bson, json::to_bson(j)},
{json::input_format_t::bon8, json::to_bon8(j)},
};
for (const auto& encoding : encodings)
{
const auto format = encoding.first;
const auto& bytes = encoding.second;
CAPTURE(format);
// every prefix is truncated input
for (std::size_t length = 0; length < bytes.size(); ++length)
{
CAPTURE(length);
const auto result = parse_binary_recovering(std::vector<std::uint8_t>(bytes.begin(), bytes.begin() + static_cast<std::ptrdiff_t>(length)), format);
CHECK(!result.ok);
CHECK(result.errors == 1);
CHECK(result.balanced);
}
// the complete input is read as usual (binary values do not
// round-trip through every format, so compare with a plain parse)
json expected;
nlohmann::detail::json_sax_dom_parser<json> dom(expected);
CHECK(json::sax_parse(bytes, &dom, format));
const auto complete = parse_binary_recovering(bytes, format);
CHECK(complete.ok);
CHECK(complete.errors == 0);
CHECK(complete.value == expected);
// a byte after the value
auto trailing_bytes = bytes;
trailing_bytes.push_back(0x01);
const auto trailing = parse_binary_recovering(trailing_bytes, format);
CHECK(!trailing.ok);
CHECK(trailing.errors == 1);
CHECK(trailing.value == expected);
}
}
SECTION("containers without an end")
{
// these made the readers loop, or read on, after the error
const auto cbor_array = parse_binary_recovering({0x9F}, json::input_format_t::cbor);
CHECK(cbor_array.errors == 1);
CHECK(cbor_array.value == json::array());
const auto cbor_map = parse_binary_recovering({0xBF, 0x61, 'a'}, json::input_format_t::cbor);
CHECK(cbor_map.errors == 1);
CHECK(cbor_map.value == json({{"a", nullptr}}));
const auto msgpack_array = parse_binary_recovering({0xDD, 0xFF, 0xFF, 0xFF, 0xFF}, json::input_format_t::msgpack);
CHECK(msgpack_array.errors == 1);
CHECK(msgpack_array.value == json::array());
const auto msgpack_map = parse_binary_recovering({0x81, 0xA1, 'a', 0x92, 0x01}, json::input_format_t::msgpack);
CHECK(msgpack_map.errors == 1);
CHECK(msgpack_map.value == json({{"a", {1}}}));
}
SECTION("BJData ndarray")
{
// a 2x3 int8 array with two of its six elements; the annotated array
// format opens an object and two arrays of its own
const auto result = parse_binary_recovering({'[', '$', 'i', '#', '[', '$', 'i', '#', 'i', 2, 2, 3, 1, 2}, json::input_format_t::bjdata);
CHECK(result.errors == 1);
CHECK(result.balanced);
CHECK(result.value == json({{"_ArrayType_", "int8"}, {"_ArraySize_", {2, 3}}, {"_ArrayData_", {1, 2}}}));
}
SECTION("binary formats repair items whose end is known")
{
struct Repair
{
json::input_format_t format;
std::vector<std::uint8_t> input;
json expected;
std::size_t errors;
};
const std::vector<Repair> repairs =
{
// CBOR: tags are ignored (here tag 1 and the self-describe tag 55799)
{json::input_format_t::cbor, {0x82, 0xC1, 0x05, 0xD9, 0xD9, 0xF7, 0x06}, {5, 6}, 2},
// CBOR: undefined and other simple values become null
{json::input_format_t::cbor, {0x84, 0xF7, 0xE0, 0xF8, 0x20, 0x01}, {nullptr, nullptr, nullptr, 1}, 3},
// CBOR: ill-formed UTF-8 becomes U+FFFD, also in keys
{json::input_format_t::cbor, {0xA1, 0x61, 0xFF, 0x62, 0xC3, 0x28}, {{replacement_character(), replacement_character() + "("}}, 2},
// CBOR: members whose key is not a string are skipped, whatever their key and value
{json::input_format_t::cbor, {0xA4, 0x01, 0x02, 0x82, 0x01, 0x02, 0xA1, 0x61, 'x', 0x9F, 0xFF, 0xC1, 0x01, 0x5F, 0x41, 0x00, 0xFF, 0x61, 'a', 0x03}, {{"a", 3}}, 3},
{json::input_format_t::cbor, {0xBF, 0xF5, 0xBF, 0x61, 'x', 0x7F, 0x61, 'y', 0xFF, 0xFF, 0x61, 'a', 0x03, 0xFF}, {{"a", 3}}, 1},
// MessagePack: members whose key is not a string are skipped
{json::input_format_t::msgpack, {0x84, 0x01, 0x02, 0x81, 0xA1, 'x', 0x01, 0x92, 0x01, 0x02, 0xD4, 0x01, 0x02, 0xC0, 0xA1, 'a', 0x04}, {{"a", 4}}, 3},
// MessagePack: ill-formed UTF-8 becomes U+FFFD
{json::input_format_t::msgpack, {0x92, 0xA2, 0xC3, 0x28, 0xA3, 0xE2, 0x82, 'x'}, {replacement_character() + "(", replacement_character() + "x"}, 2},
// UBJSON: a char that is not ASCII becomes U+FFFD
{json::input_format_t::ubjson, {'[', 'C', 0x80, 'C', 'A', ']'}, {replacement_character(), "A"}, 1},
// UBJSON: the longest beginning of a high-precision number is kept
{json::input_format_t::ubjson, {'[', 'H', 'i', 5, '1', '2', 'a', 'b', 'c', 'H', 'i', 2, '1', '.', 'H', 'i', 3, 'a', 'b', 'c', 'H', 'i', 3, '4', '.', '5', ']'}, {12, 1, nullptr, 4.5}, 3},
// BJData, too
{json::input_format_t::bjdata, {'[', 'C', 0xFF, 'H', 'i', 2, '-', '1', 'H', 'i', 2, '-', 'x', ']'}, {replacement_character(), -1, nullptr}, 2},
// BON8: members whose key is not a string are skipped
{json::input_format_t::bon8, {0x89, 0x91, 0x92, 0xC9, 0x40, 0x82, 0x91, 0x92, 0x61, 0x93}, {{"a", 3}}, 2},
{json::input_format_t::bon8, {0x8B, 0x91, 0x85, 0x91, 0xFE, 0xFA, 0x8B, 'x', 0x91, 0xFE, 0x61, 0x93, 0xFE}, {{"a", 3}}, 2},
// BSON: elements of types the library does not read become null
{
json::input_format_t::bson, bson_document(
{
bson_element(0x07, "_id", bytes(12)), // ObjectId
bson_element(0x09, "date", bytes(8)), // UTC datetime
bson_element(0x13, "decimal", bytes(16)), // 128-bit decimal
bson_element(0x0B, "regex", {'a', '+', 0, 'i', 0}), // regular expression
bson_element(0x0D, "code", bson_string("f()")), // JavaScript code
bson_element(0x0E, "symbol", bson_string("s")), // symbol
bson_element(0x0C, "pointer", concatenated(bson_string("c"), bytes(12))), // DBPointer
bson_element(0x0F, "scope", concatenated(bson_int32(15), bson_string("g"), bson_document({}))), // code with scope
bson_element(0x06, "undefined", {}), // undefined
bson_element(0xFF, "min", {}), // min key
bson_element(0x7F, "max", {}), // max key
bson_element(0x10, "z", bson_int32(7)),
}),
{{"_id", nullptr}, {"date", nullptr}, {"decimal", nullptr}, {"regex", nullptr}, {"code", nullptr}, {"symbol", nullptr}, {"pointer", nullptr}, {"scope", nullptr}, {"undefined", nullptr}, {"min", nullptr}, {"max", nullptr}, {"z", 7}},
11
},
// BSON: an element of an unknown type becomes null, and the rest of its document is skipped
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x04, "array", bson_document({bson_element(0x10, "0", bson_int32(1)), bson_element(0x42, "1", bytes(3))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"a", 1}, {"x", nullptr}}}, {"array", {1, nullptr}}, {"after", 3}},
2
},
// BSON: so does a string or byte array whose length cannot be right
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x02, "s", bson_string("abc", -10)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x03, "bin", bson_document({bson_element(0x05, "b", concatenated(bson_int32(-1), bytes(1))), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"s", nullptr}}}, {"bin", {{"b", nullptr}}}, {"after", 3}},
2
},
// BSON: a string without its terminator, and a document whose size does not match, are kept
{
json::input_format_t::bson, bson_document(
{
bson_element(0x02, "s", {2, 0, 0, 0, 'a', 'X'}),
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1))}, 1)),
}),
{{"s", "a"}, {"inner", {{"a", 1}}}},
2
},
};
for (const auto& repair : repairs)
{
CAPTURE(repair.format);
CAPTURE(repair.input);
const auto result = parse_binary_recovering(repair.input, repair.format);
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.errors == repair.errors);
CHECK(result.value == repair.expected);
REQUIRE(!result.messages.empty());
#if !defined(JSON_NOEXCEPTION)
// the first error is the one reported without recovering; under
// JSON_NOEXCEPTION, reading without recovering aborts instead of
// throwing, so there is no message to compare with
CHECK(result.messages.front() == binary_error_message(repair.input, repair.format));
#endif
}
}
SECTION("binary formats repair numbers that are out of range")
{
// CBOR: a negative integer below the range of number_integer_t
const auto cbor = parse_binary_recovering({0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value.is_number_float());
CHECK(cbor.value.get<double>() == -18446744073709551616.0);
// UBJSON: a high-precision number too large for number_float_t
const auto ubjson = parse_binary_recovering({'H', 'i', 5, '1', 'e', '9', '9', '9'}, json::input_format_t::ubjson);
CHECK(ubjson.errors == 1);
CHECK(ubjson.value.is_number_float());
CHECK(std::isinf(ubjson.value.get<double>()));
}
SECTION("binary formats stop where the end of an item is not known")
{
// a byte that begins no item
const auto cbor = parse_binary_recovering({0x82, 0x01, 0x1C, 0x02}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value == json({1}));
// a key that is no item: the unused MessagePack byte, a CBOR break
// in a map of known size, and the end of a BON8 container
const auto msgpack = parse_binary_recovering({0x82, 0xA1, 'a', 0x01, 0xC1, 0x02}, json::input_format_t::msgpack);
CHECK(msgpack.errors == 1);
CHECK(msgpack.value == json({{"a", 1}}));
const auto cbor_break = parse_binary_recovering({0xA2, 0x61, 'a', 0x01, 0xFF, 0x02}, json::input_format_t::cbor);
CHECK(cbor_break.errors == 1);
CHECK(cbor_break.value == json({{"a", 1}}));
const auto bon8 = parse_binary_recovering({0x88, 0x61, 0x91, 0xFE}, json::input_format_t::bon8);
CHECK(bon8.errors == 1);
CHECK(bon8.value == json({{"a", 1}}));
// a skipped member that the input ends in
const auto truncated = parse_binary_recovering({0xA2, 0x01, 0x82, 0x01}, json::input_format_t::cbor);
CHECK(truncated.errors == 2);
CHECK(truncated.balanced);
CHECK(truncated.value == json::object());
// a BSON element of an unknown type in a document whose size cannot be right
const auto bson = parse_binary_recovering(bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3))}, -10), json::input_format_t::bson);
CHECK(bson.errors == 1);
CHECK(bson.value == json({{"a", 1}, {"x", nullptr}}));
}
SECTION("changed bytes in binary input")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}, {"i", "\xC3\xA4"}};
const std::vector<std::pair<json::input_format_t, std::vector<std::uint8_t>>> encodings =
{
{json::input_format_t::cbor, json::to_cbor(j)},
{json::input_format_t::msgpack, json::to_msgpack(j)},
{json::input_format_t::ubjson, json::to_ubjson(j)},
{json::input_format_t::ubjson, json::to_ubjson(j, true, true)},
{json::input_format_t::bjdata, json::to_bjdata(j)},
{json::input_format_t::bjdata, json::to_bjdata(j, true, true)},
{json::input_format_t::bson, json::to_bson(j)},
{json::input_format_t::bon8, json::to_bon8(j)},
};
const std::vector<std::uint8_t> replacements = {0x00, 0x01, 0x7F, 0x80, 0xC1, 0xD9, 0xE0, 0xF7, 0xFE, 0xFF};
for (const auto& encoding : encodings)
{
const auto format = encoding.first;
const auto& original = encoding.second;
CAPTURE(format);
std::vector<std::vector<std::uint8_t>> inputs;
for (std::size_t position = 0; position < original.size(); ++position)
{
for (const auto replacement : replacements)
{
auto changed = original;
changed[position] = replacement;
inputs.push_back(changed);
}
auto removed = original;
removed.erase(removed.begin() + static_cast<std::ptrdiff_t>(position));
inputs.push_back(removed);
}
for (const auto& input : inputs)
{
CAPTURE(input);
const auto result = parse_binary_recovering(input, format);
CHECK(result.balanced);
CHECK(result.errors <= input.size() + 1);
#if !defined(JSON_NOEXCEPTION)
// an error is reported exactly if reading into a JSON value
// fails, and the first one is the same (under JSON_NOEXCEPTION,
// that reading aborts instead of throwing)
const auto message = binary_error_message(input, format);
CHECK(result.ok == message.empty());
if (!result.ok && result.errors < 100)
{
CHECK(result.messages.front() == message);
}
#endif
}
}
}
SECTION("JSON text")
{
// the parser stopped, but reported success
json j;
RecoveringParser sax(j);
CHECK(!json::sax_parse("[1,2,3,]", &sax));
CHECK(sax.errors == 1);
CHECK(j == json({1, 2, 3}));
}
SECTION("the SAX parsers of the library stop")
{
json _;
CHECK(json::from_cbor(std::vector<std::uint8_t> {0x9F}, true, false).is_discarded());
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<std::uint8_t> {0x9F}), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing CBOR value: unexpected end of input", json::parse_error&);
CHECK(json::parse("[1,2,3,]", nullptr, false).is_discarded());
CHECK(!json::accept("[1,2,3,]"));
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP DOCTEST_CLANG_SUPPRESS_WARNING_POP
+221
View File
@@ -3033,3 +3033,224 @@ TEST_CASE("UBJSON optimized array of unsigned integers beyond int64")
CHECK(json::to_ubjson(j, true, true) == expected); CHECK(json::to_ubjson(j, true, true) == expected);
CHECK(json::from_ubjson(expected) == j); CHECK(json::from_ubjson(expected) == j);
} }
namespace
{
// the bytes that follow the marker of an integer: the value in the width of
// the marker (big endian for UBJSON, little endian for BJData), or, for a
// high-precision number, the length and the decimal digits
std::vector<std::uint8_t> integer_payload(const char marker, const json& value, const bool little_endian)
{
std::size_t width = 0;
switch (marker)
{
case 'i':
case 'U':
width = 1;
break;
case 'I':
case 'u':
width = 2;
break;
case 'l':
case 'm':
width = 4;
break;
case 'L':
case 'M':
width = 8;
break;
default:
{
const std::string digits = value.dump();
std::vector<std::uint8_t> result = {'i', static_cast<std::uint8_t>(digits.size())};
for (const char c : digits)
{
result.push_back(static_cast<std::uint8_t>(c));
}
return result;
}
}
const std::uint64_t bits = value.is_number_unsigned()
? value.get<std::uint64_t>()
: static_cast<std::uint64_t>(value.get<std::int64_t>());
std::vector<std::uint8_t> result(width);
for (std::size_t i = 0; i < width; ++i)
{
result[little_endian ? i : width - 1 - i] = static_cast<std::uint8_t>(bits >> (8 * i));
}
return result;
}
json i64(const std::int64_t v)
{
return v;
}
json u64(const std::uint64_t v)
{
return v;
}
} // namespace
TEST_CASE("UBJSON and BJData integer markers at every range edge")
{
// An optimized container announces the marker of its values after `$` and
// then writes every value without a marker, so the marker the writer
// announces and the width it writes must match for every value. This
// checks both for the values around each edge of the integer types, as
// scalars and as the values of optimized arrays and objects.
struct integer_case
{
json value;
char ubjson; // expected UBJSON marker
char bjdata; // expected BJData marker
};
const std::int64_t int64_min = (std::numeric_limits<std::int64_t>::min)();
const std::int64_t int64_max = (std::numeric_limits<std::int64_t>::max)();
const std::uint64_t uint64_max = (std::numeric_limits<std::uint64_t>::max)();
const std::vector<integer_case> cases =
{
// int8
{i64(-129), 'I', 'I'},
{i64(-128), 'i', 'i'},
{i64(-127), 'i', 'i'},
{i64(-1), 'i', 'i'},
{i64(0), 'i', 'i'},
{u64(0), 'i', 'i'},
{i64(126), 'i', 'i'},
{i64(127), 'i', 'i'},
{u64(127), 'i', 'i'},
{i64(128), 'U', 'U'},
{u64(128), 'U', 'U'},
// uint8
{i64(254), 'U', 'U'},
{i64(255), 'U', 'U'},
{u64(255), 'U', 'U'},
{i64(256), 'I', 'I'},
{u64(256), 'I', 'I'},
// int16
{i64(-32769), 'l', 'l'},
{i64(-32768), 'I', 'I'},
{i64(-32767), 'I', 'I'},
{i64(32766), 'I', 'I'},
{i64(32767), 'I', 'I'},
{u64(32767), 'I', 'I'},
{i64(32768), 'l', 'u'},
{u64(32768), 'l', 'u'},
// uint16 (BJData only)
{i64(65534), 'l', 'u'},
{i64(65535), 'l', 'u'},
{u64(65535), 'l', 'u'},
{i64(65536), 'l', 'l'},
{u64(65536), 'l', 'l'},
// int32
{i64(-2147483649LL), 'L', 'L'},
{i64(-2147483648LL), 'l', 'l'},
{i64(-2147483647LL), 'l', 'l'},
{i64(2147483646LL), 'l', 'l'},
{i64(2147483647LL), 'l', 'l'},
{u64(2147483647ULL), 'l', 'l'},
{i64(2147483648LL), 'L', 'm'},
{u64(2147483648ULL), 'L', 'm'},
// uint32 (BJData only)
{i64(4294967294LL), 'L', 'm'},
{i64(4294967295LL), 'L', 'm'},
{u64(4294967295ULL), 'L', 'm'},
{i64(4294967296LL), 'L', 'L'},
{u64(4294967296ULL), 'L', 'L'},
// int64
{i64(int64_min), 'L', 'L'},
{i64(int64_min + 1), 'L', 'L'},
{i64(int64_max - 1), 'L', 'L'},
{i64(int64_max), 'L', 'L'},
{u64(static_cast<std::uint64_t>(int64_max)), 'L', 'L'},
// uint64 (BJData only; UBJSON writes a high-precision number)
{u64(static_cast<std::uint64_t>(int64_max) + 1), 'H', 'M'},
{u64(uint64_max - 1), 'H', 'M'},
{u64(uint64_max), 'H', 'M'},
};
for (const auto& c : cases)
{
for (const bool bjdata :
{
false, true
})
{
const char marker = bjdata ? c.bjdata : c.ubjson;
const std::vector<std::uint8_t> payload = integer_payload(marker, c.value, bjdata);
const auto to_binary = [bjdata](const json & j, const bool use_size, const bool use_type)
{
return bjdata ? json::to_bjdata(j, use_size, use_type) : json::to_ubjson(j, use_size, use_type);
};
const auto from_binary = [bjdata](const std::vector<std::uint8_t>& v)
{
return bjdata ? json::from_bjdata(v) : json::from_ubjson(v);
};
INFO("value = " << c.value.dump() << (c.value.is_number_unsigned() ? " (unsigned)" : "") << ", format = " << (bjdata ? "BJData" : "UBJSON"));
// scalar
std::vector<std::uint8_t> expected = {static_cast<std::uint8_t>(marker)};
expected.insert(expected.end(), payload.begin(), payload.end());
for (const bool use_size :
{
false, true
})
{
CHECK(to_binary(c.value, use_size, false) == expected);
}
CHECK(from_binary(expected) == c.value);
const json arr = {c.value, c.value, c.value};
// array without count or type: every value has its marker
expected = {'['};
for (int i = 0; i < 3; ++i)
{
expected.push_back(static_cast<std::uint8_t>(marker));
expected.insert(expected.end(), payload.begin(), payload.end());
}
expected.push_back(']');
CHECK(to_binary(arr, false, false) == expected);
CHECK(from_binary(expected) == arr);
// array with count: every value has its marker
expected = {'[', '#', 'i', 3};
for (int i = 0; i < 3; ++i)
{
expected.push_back(static_cast<std::uint8_t>(marker));
expected.insert(expected.end(), payload.begin(), payload.end());
}
CHECK(to_binary(arr, true, false) == expected);
CHECK(from_binary(expected) == arr);
// array with type and count: the marker once, then the payloads
expected = {'[', '$', static_cast<std::uint8_t>(marker), '#', 'i', 3};
for (int i = 0; i < 3; ++i)
{
expected.insert(expected.end(), payload.begin(), payload.end());
}
CHECK(to_binary(arr, true, true) == expected);
CHECK(from_binary(expected) == arr);
// object with type and count: the marker once, then key and payload
const json obj = {{"a", c.value}, {"b", c.value}};
expected = {'{', '$', static_cast<std::uint8_t>(marker), '#', 'i', 2};
for (const char key :
{'a', 'b'
})
{
expected.push_back('i');
expected.push_back(1);
expected.push_back(static_cast<std::uint8_t>(key));
expected.insert(expected.end(), payload.begin(), payload.end());
}
CHECK(to_binary(obj, true, true) == expected);
CHECK(from_binary(expected) == obj);
}
}
}
+15 -11
View File
@@ -1,8 +1,8 @@
# amalgamate.py - Amalgamate C source and header files # amalgamate.py - Amalgamate C source and header files
Origin: https://bitbucket.org/erikedlund/amalgamate Origin: https://github.com/edlund/amalgamate (formerly hosted at
https://bitbucket.org/erikedlund/amalgamate, which no longer exists; see
Mirror: https://github.com/edlund/amalgamate `CHANGES.md` for the upstream commit this copy is based on)
`amalgamate.py` aims to make it easy to use SQLite-style C source and header `amalgamate.py` aims to make it easy to use SQLite-style C source and header
amalgamation in projects. amalgamation in projects.
@@ -41,21 +41,22 @@ results.
## Installing amalgamate.py ## Installing amalgamate.py
Python v.2.7.0 or higher is required. Python 3 is required.
`amalgamate.py` can be tested and installed using the following commands: In this repository, `amalgamate.py` is not installed separately; it is run in
place through `make amalgamate`, which calls it once for `json.hpp` and once
./test.sh && sudo -k cp ./amalgamate.py /usr/local/bin/ for `json_fwd.hpp` (see the root `Makefile`).
## Using amalgamate.py ## Using amalgamate.py
amalgamate.py [-v] -c path/to/config.json -s path/to/source/dir \ amalgamate.py -c path/to/config.json -s path/to/source/dir \
[-p path/to/prologue.(c|h)] [-p path/to/prologue.(c|h)] [--verbose=yes|no]
* The `-c, --config` option should specify the path to a JSON config file which * The `-c, --config` option should specify the path to a JSON config file which
lists the source files, include paths and where to write the resulting lists the source files, include paths and where to write the resulting
amalgamation. Have a look at `test/source.c.json` and `test/include.h.json` amalgamation. `config_json.json` and `config_json_fwd.json` in this
to see two examples. directory are the configs used for `json.hpp` and `json_fwd.hpp`; each
sets `target`, `sources` and `include_paths`.
The optional `external` list names include paths that are kept as `#include` The optional `external` list names include paths that are kept as `#include`
directives instead of being inlined, e.g. `["nlohmann/json.hpp"]` for a header directives instead of being inlined, e.g. `["nlohmann/json.hpp"]` for a header
@@ -68,3 +69,6 @@ Python v.2.7.0 or higher is required.
* The `-p, --prologue` option should specify the path to a file which will be * The `-p, --prologue` option should specify the path to a file which will be
added to the beginning of the amalgamation. It is optional. added to the beginning of the amalgamation. It is optional.
* The `-v, --verbose` option takes `yes` or `no` (for example
`--verbose=yes`, as used by the Makefile). It is optional.
+19 -1
View File
@@ -2,8 +2,26 @@
Generate the Natvis debugger visualization file for all supported namespace combinations. Generate the Natvis debugger visualization file for all supported namespace combinations.
The ABI tag list and the library version are parsed from
`include/nlohmann/detail/abi_macros.hpp`, so this script must be re-run (via
`make natvis`) whenever an `NLOHMANN_JSON_ABI_TAG_*` macro is added to that
file or the library version is bumped — otherwise the committed
`nlohmann_json.natvis` drifts from the header it visualizes, and
`make check-amalgamation` fails.
## Usage ## Usage
```shell ```shell
./generate_natvis.py --version X.Y.Z output_directory/ make natvis
``` ```
or, equivalently:
```shell
./generate_natvis.py [--version X.Y.Z] [repository_root/]
```
`--version` and the output/repository-root directory both default to values
derived from this script's own location, so they only need to be given
explicitly when generating a Natvis file for a different checkout or a
version other than the one in `abi_macros.hpp`.
+58 -4
View File
@@ -7,21 +7,75 @@ import os
import re import re
import sys import sys
# Directory of the repository, assuming this script stays at
# tools/generate_natvis/generate_natvis.py. Used only as the default value
# for the "output" argument below.
REPO_ROOT = os.path.normpath(os.path.join(sys.path[0], '..', '..'))
def semver(v): def semver(v):
if not re.fullmatch(r'\d+\.\d+\.\d+', v): if not re.fullmatch(r'\d+\.\d+\.\d+', v):
raise ValueError raise ValueError
return v return v
def abi_info(repo_root):
"""Derive the ABI tag list (in NLOHMANN_JSON_ABI_TAGS_CONCAT order) and the
library version from <repo_root>/include/nlohmann/detail/abi_macros.hpp,
so this script cannot drift from the header it visualizes."""
abi_macros_hpp = os.path.join(repo_root, 'include', 'nlohmann', 'detail', 'abi_macros.hpp')
with open(abi_macros_hpp) as f:
content = f.read()
# find the NLOHMANN_JSON_ABI_TAGS_CONCAT(...) invocation that lists the
# NLOHMANN_JSON_ABI_TAG_* identifiers in order (not its own #define, which
# only names its formal parameters a, b, c, ...)
tag_idents = None
for args in re.findall(r'NLOHMANN_JSON_ABI_TAGS_CONCAT\(\s*(.*?)\)', content, re.S):
idents = re.findall(r'NLOHMANN_JSON_ABI_TAG_\w+', args)
if idents:
tag_idents = idents
break
if not tag_idents:
raise ValueError(f'could not find NLOHMANN_JSON_ABI_TAGS_CONCAT(...) in {abi_macros_hpp}')
abi_tags = []
for ident in tag_idents:
# each tag is #define'd to its suffix (e.g. _diag) when the matching
# JSON_* option is enabled, and to nothing in the #else branch; only
# the non-empty definition matches here
match = re.search(r'#define\s+' + re.escape(ident) + r'\s+(_\w+)\s*\n', content)
if not match:
raise ValueError(f'could not find a non-empty #define for {ident} in {abi_macros_hpp}')
abi_tags.append(match.group(1))
version = {}
for part in ('MAJOR', 'MINOR', 'PATCH'):
match = re.search(r'#define\s+NLOHMANN_JSON_VERSION_' + part + r'\s+(\d+)', content)
if not match:
raise ValueError(f'could not find NLOHMANN_JSON_VERSION_{part} in {abi_macros_hpp}')
version[part] = match.group(1)
return abi_tags, '{MAJOR}.{MINOR}.{PATCH}'.format(**version)
if __name__ == '__main__': if __name__ == '__main__':
parser = argparse.ArgumentParser() parser = argparse.ArgumentParser()
parser.add_argument('--version', required=True, type=semver, help='Library version number') parser.add_argument('--version', type=semver,
parser.add_argument('output', help='Output directory for nlohmann_json.natvis') help='Library version number (default: parsed from '
'include/nlohmann/detail/abi_macros.hpp below "output")')
parser.add_argument('output', nargs='?', default=REPO_ROOT,
help='Repository root: where include/nlohmann/detail/abi_macros.hpp is '
'read from and where nlohmann_json.natvis is written '
'(default: the repository root this script lives in)')
args = parser.parse_args() args = parser.parse_args()
derived_tags, derived_version = abi_info(args.output)
namespaces = ['nlohmann'] namespaces = ['nlohmann']
abi_prefix = 'json_abi' abi_prefix = 'json_abi'
abi_tags = ['_diag', '_ldvcmp', '_dp', '_bics', '_psp', '_snul'] abi_tags = derived_tags
version = '_v' + args.version.replace('.', '_') version = '_v' + (args.version or derived_version).replace('.', '_')
inline_namespaces = [] inline_namespaces = []
# generate all combinations of inline namespace names # generate all combinations of inline namespace names
+51
View File
@@ -0,0 +1,51 @@
# macro_builder
Generates the argument-counting macros behind the `NLOHMANN_DEFINE_TYPE_*` and `NLOHMANN_DEFINE_DERIVED_TYPE_*`
macros in [`include/nlohmann/detail/macro_scope.hpp`](../../include/nlohmann/detail/macro_scope.hpp):
- `NLOHMANN_JSON_EXPAND`
- `NLOHMANN_JSON_GET_MACRO`, which selects a macro by the number of its arguments (64 slots)
- `NLOHMANN_JSON_PASTE`, which calls a function-like macro for each member, and its helpers `NLOHMANN_JSON_PASTE2`
to `NLOHMANN_JSON_PASTE64`
- `NLOHMANN_JSON_DOUBLE_PASTE`, which the `*_WITH_NAMES` macros use to call a function-like macro for each
(JSON name, member) pair, and its helpers `NLOHMANN_JSON_DOUBLE_PASTE3` to `NLOHMANN_JSON_DOUBLE_PASTE63`
- the slot table of `NLOHMANN_JSON_TYPE_BODY`, which dispatches `NLOHMANN_DEFINE_TYPE_*(Type)` (no further
arguments) to the zero-member implementation and every other argument count to the one-or-more-member
implementation
The number of slots (`max_args` in [`main.cpp`](main.cpp)) sets the member limit of these macros.
`NLOHMANN_JSON_PASTE` and `NLOHMANN_JSON_TYPE_BODY` take the function/prefix as their first argument, so 64 slots
allow 63 members; `NLOHMANN_JSON_DOUBLE_PASTE` additionally consumes its members two at a time (name, member), so
it only defines the odd helpers up to `NLOHMANN_JSON_DOUBLE_PASTE63`.
## Usage
From the project root:
```shell
c++ -std=c++11 tools/macro_builder/main.cpp -o macro_builder
./macro_builder
./macro_builder type_body
```
1. Run `./macro_builder` (no arguments). In `include/nlohmann/detail/macro_scope.hpp`, replace the lines from
`#define NLOHMANN_JSON_EXPAND( x ) x` to the `#define NLOHMANN_JSON_DOUBLE_PASTE63(...)` line with the output,
without its trailing empty line.
2. Run `./macro_builder type_body`. Replace the lines from `#define NLOHMANN_JSON_TYPE_BODY(Prefix, ...)` to the
`NLOHMANN_JSON_TYPE_BODY_SENTINEL))` line with the output.
3. Run `make amalgamate`. It updates `single_include/nlohmann/json.hpp` and runs `make pretty`, which indents the
continuation lines that the tool writes unindented.
With an unchanged `main.cpp`, these steps reproduce both blocks of `macro_scope.hpp` byte for byte. `make
macro_builder_check` (also run by CI, see `.github/workflows/check_amalgamation.yml`) automates this: it builds
`main.cpp`, regenerates both blocks, and fails on a diff against the checked-in header.
## Maintained by hand
The tool does not generate everything that depends on the number of slots. When changing `max_args`, also update:
- the documented limit of 63 members in `docs/mkdocs/docs` and the tests at that limit in
`tests/src/unit-udt_macro.cpp`
All three tables pass one macro name per slot to `NLOHMANN_JSON_GET_MACRO`, so they need exactly as many entries as
it has slots.
+69 -4
View File
@@ -1,10 +1,14 @@
#include <cstdlib> #include <cstdlib>
#include <iostream> #include <iostream>
#include <sstream> #include <sstream>
#include <string>
using namespace std; using namespace std;
void build_code(int max_args) // Builds NLOHMANN_JSON_EXPAND, NLOHMANN_JSON_GET_MACRO, and the
// NLOHMANN_JSON_PASTE / NLOHMANN_JSON_PASTE2..PASTE<max_args> dispatch table
// and recursive definitions.
string build_paste_code(int max_args)
{ {
stringstream ss; stringstream ss;
ss << "#define NLOHMANN_JSON_EXPAND( x ) x" << endl; ss << "#define NLOHMANN_JSON_EXPAND( x ) x" << endl;
@@ -30,14 +34,75 @@ void build_code(int max_args)
ss << "v" << i-1 << ")" << endl; ss << "v" << i-1 << ")" << endl;
} }
cout << ss.str() << endl; return ss.str();
}
// Builds the NLOHMANN_JSON_DOUBLE_PASTE dispatch table and recursive
// definitions used by the *_WITH_NAMES macros. Its GET_MACRO dispatch reuses
// the same max_args slots as NLOHMANN_JSON_PASTE, but DOUBLE_PASTE consumes
// its arguments two at a time (name, member), so an even slot count falls
// back to the next lower odd NLOHMANN_JSON_DOUBLE_PASTE<N>.
string build_double_paste_code(int max_args)
{
stringstream ss;
ss << "#define NLOHMANN_JSON_DOUBLE_PASTE(...) NLOHMANN_JSON_EXPAND(NLOHMANN_JSON_GET_MACRO(__VA_ARGS__, \\" << endl;
for (int i = max_args ; i > 1 ; i--)
{
int k = (i % 2 == 1) ? i : i - 1;
ss << "NLOHMANN_JSON_DOUBLE_PASTE" << k << ", \\" << endl;
}
ss << "NLOHMANN_JSON_DOUBLE_PASTE1)(__VA_ARGS__))" << endl;
ss << "#define NLOHMANN_JSON_DOUBLE_PASTE3(func, v1, v2) func(v1, v2)" << endl;
for (int k = 5 ; k <= max_args - 1 ; k += 2)
{
ss << "#define NLOHMANN_JSON_DOUBLE_PASTE" << k << "(func, ";
for (int j = 1 ; j < k - 1 ; j++)
ss << "v" << j << ", ";
ss << "v" << k - 1 << ") NLOHMANN_JSON_DOUBLE_PASTE3(func, v1, v2) NLOHMANN_JSON_DOUBLE_PASTE" << k - 2 << "(func, ";
for (int j = 3 ; j < k - 1 ; j++)
ss << "v" << j << ", ";
ss << "v" << k - 1 << ")" << endl;
}
return ss.str();
}
// Builds the NLOHMANN_JSON_TYPE_BODY dispatch table: max_args - 1 slots
// selecting the *_MEMBERS implementation and a final slot selecting
// *_EMPTY, so NLOHMANN_DEFINE_TYPE_*(Type) with no further arguments still
// resolves (issue #4041).
string build_type_body_table(int max_args)
{
stringstream ss;
ss << "#define NLOHMANN_JSON_TYPE_BODY(Prefix, ...) NLOHMANN_JSON_EXPAND(NLOHMANN_JSON_GET_MACRO(__VA_ARGS__, \\" << endl;
const int per_line = 8;
for (int i = 1 ; i <= max_args ; i++)
{
ss << (i == max_args ? "Prefix ## EMPTY" : "Prefix ## MEMBERS") << ", ";
if (i % per_line == 0)
ss << "\\" << endl;
}
ss << "NLOHMANN_JSON_TYPE_BODY_SENTINEL))" << endl;
return ss.str();
} }
int main(int argc, char** argv) int main(int argc, char** argv)
{ {
int max_args = 64; int max_args = 64;
build_code(max_args);
// With "type_body", print only the NLOHMANN_JSON_TYPE_BODY dispatch
// table (a separate insertion point in macro_scope.hpp); otherwise
// print the EXPAND/GET_MACRO/PASTE/DOUBLE_PASTE block that precedes it.
if (argc > 1 && string(argv[1]) == "type_body")
{
cout << build_type_body_table(max_args);
}
else
{
cout << build_paste_code(max_args) << build_double_paste_code(max_args);
}
return 0; return 0;
} }
+11 -11
View File
@@ -5,6 +5,8 @@ import logging
import os import os
import re import re
import shutil import shutil
import socket
import ssl
import sys import sys
import subprocess import subprocess
@@ -34,7 +36,7 @@ JSON_VERSION_RE = re.compile(r'\s*#\s*define\s+NLOHMANN_JSON_VERSION_MAJOR\s+')
class ExitHandler(logging.StreamHandler): class ExitHandler(logging.StreamHandler):
def __init__(self, level): def __init__(self, level):
""".""" """Exit the process on log records at or above level."""
super().__init__() super().__init__()
self.level = level self.level = level
@@ -54,7 +56,7 @@ def is_project_root(test_dir='.'):
class DirectoryEventBucket: class DirectoryEventBucket:
def __init__(self, callback, delay=1.2, threshold=0.8): def __init__(self, callback, delay=1.2, threshold=0.8):
""".""" """Batch directory events and pass their common path to callback."""
self.delay = delay self.delay = delay
self.threshold = timedelta(seconds=threshold) self.threshold = timedelta(seconds=threshold)
self.callback = callback self.callback = callback
@@ -99,7 +101,7 @@ class WorkTree:
make_command = 'make' make_command = 'make'
def __init__(self, root_dir, tree_dir): def __init__(self, root_dir, tree_dir):
""".""" """Track the working tree at tree_dir and its amalgamated header."""
self.root_dir = root_dir self.root_dir = root_dir
self.tree_dir = tree_dir self.tree_dir = tree_dir
self.rel_dir = os.path.relpath(tree_dir, root_dir) self.rel_dir = os.path.relpath(tree_dir, root_dir)
@@ -114,11 +116,11 @@ class WorkTree:
self.build_time = t.strftime(DATETIME_FORMAT) self.build_time = t.strftime(DATETIME_FORMAT)
def __hash__(self): def __hash__(self):
""".""" """Hash by working tree directory."""
return hash((self.tree_dir)) return hash((self.tree_dir))
def __eq__(self, other): def __eq__(self, other):
""".""" """Compare by working tree directory."""
if not isinstance(other, type(self)): if not isinstance(other, type(self)):
return NotImplemented return NotImplemented
return self.tree_dir == other.tree_dir return self.tree_dir == other.tree_dir
@@ -150,7 +152,7 @@ class WorkTree:
class WorkTrees(FileSystemEventHandler): class WorkTrees(FileSystemEventHandler):
def __init__(self, root_dir): def __init__(self, root_dir):
""".""" """Find the working trees below root_dir and watch it for changes."""
super().__init__() super().__init__()
self.root_dir = root_dir self.root_dir = root_dir
self.trees = set([]) self.trees = set([])
@@ -250,11 +252,11 @@ class WorkTrees(FileSystemEventHandler):
self.observer.stop() self.observer.stop()
self.observer.join() self.observer.join()
class HeaderRequestHandler(SimpleHTTPRequestHandler): # lgtm[py/missing-call-to-init] class HeaderRequestHandler(SimpleHTTPRequestHandler):
cors_origins = DEFAULT_CORS_ORIGINS cors_origins = DEFAULT_CORS_ORIGINS
def __init__(self, request, client_address, server): def __init__(self, request, client_address, server):
""".""" """Handle a request for a header below the working trees' root directory."""
self.worktrees = server.worktrees self.worktrees = server.worktrees
self.worktree = None self.worktree = None
try: try:
@@ -336,7 +338,7 @@ class HeaderRequestHandler(SimpleHTTPRequestHandler): # lgtm[py/missing-call-to-
class DualStackServer(ThreadingHTTPServer): class DualStackServer(ThreadingHTTPServer):
def __init__(self, addr, worktrees): def __init__(self, addr, worktrees):
""".""" """Serve the headers of worktrees on addr."""
self.worktrees = worktrees self.worktrees = worktrees
super().__init__(addr, HeaderRequestHandler) super().__init__(addr, HeaderRequestHandler)
@@ -349,8 +351,6 @@ class DualStackServer(ThreadingHTTPServer):
if __name__ == '__main__': if __name__ == '__main__':
import argparse import argparse
import ssl
import socket
import yaml import yaml
# exit code # exit code