Compare commits

..
Author SHA1 Message Date
Niels Lohmann 855155d1ab Regenerate single_include after merging develop
The merge commit kept develop's single_include/nlohmann/json.hpp because
make amalgamate saw it as up to date.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:34:15 +02:00
Niels Lohmann 5fc53e3e0a Merge branch 'develop' into techdebt/5708-meta-cleanup
Conflicts:
- include/nlohmann/detail/conversions/from_json.hpp: kept the shared from_json_pair_array_to_map() body; develop's #5681 element-path fix (&p instead of &j) is already applied there
- include/nlohmann/detail/json_pointer.hpp: kept the parse_array_index()-based contains(); it already returns false for an empty token (#5614) and only uses documented StringType members (#5692)
- include/nlohmann/detail/macro_scope.hpp: combined develop's #5698 NLOHMANN_JSON_SERIALIZE_ENUM_STRICT message with the ::nlohmann::detail::templated_json_throw qualification
- single_include/nlohmann/json.hpp: regenerated with make amalgamate

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:30:46 +02:00
Niels Lohmann 6a073dbae4 Give operator>> a strong exception-safety guarantee (#5695)
operator>> parsed directly into its basic_json& target, so a parse
error left the target holding whatever was parsed before the error
instead of its previous value. With JSON_DIAGNOSTICS=1, that partial
value also violated the class invariant, because the parent pointers
of an array or object's elements are only set when the container is
closed, which a failed parse never reaches; copying such a value then
aborted in assert_invariant().

Fix it the way basic_json::parse() already handles this: parse into a
temporary and move it into the target only once parsing succeeds, so
the target is left unchanged if an exception is thrown.

Fixes #5652.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:34 +02:00
Niels Lohmann fdcc569eee Use only documented StringType members in json_pointer (#5692)
contains(const json_pointer&) and operator/=(std::size_t) (and hence
operator/(std::size_t)) used string_t operations that the StringType
template parameter documentation explicitly does not require:
comparing string_t with a const char* literal, c_str(), and
constructibility from std::string. This made both functions fail to
compile for a conforming custom StringType, even though the
documentation's own reference StringType satisfies the requirements.

Fix contains() to compare individual chars ('0'..'9') instead of
comparing string_t with const char* literals, and to call data()
(documented to be null-terminated) instead of c_str(). Fix
operator/=(std::size_t) to build the array-index token via the
existing detail::to_string<StringType> helper (ADL int_to_string() or
assignment from std::to_string()) instead of via std::to_string()
directly, matching how diff(), items(), and std::hash already convert
a std::size_t to a StringType.

Add regression tests to tests/src/unit-alt-string.cpp: contains() for
present/missing keys and indices, "-", a leading zero, and a
non-numeric token on an array, plus json_pointer::operator/(std::size_t).

Fixes #5666.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:13:31 +02:00
Niels Lohmann f30ee4bcfe Unify json_pointer's three array-index parsers
array_index(), contains() and get_checked_or_null() each re-implemented
the RFC 6901 array-index rules and the size_type range check: array_index()
does the canonical parse and throws; contains() (which must not throw,
#5395) re-validates every digit by hand and runs its own strtoull/ERANGE
check before calling array_index() anyway, parsing every array token
twice; get_checked_or_null() wraps array_index() in JSON_TRY/
JSON_INTERNAL_CATCH (detail::out_of_range&) to turn an unrepresentable
index into "not found".

Add a single private, noexcept parse_array_index(s, idx) returning an
array_index_status (ok / leading_zero / not_a_number / unresolved /
exceeds_size_type). array_index() becomes a thin wrapper mapping each
status to the existing parse_error.106/109 or out_of_range.404/410;
contains() and get_checked_or_null() switch on the status directly. This
removes contains()'s digit-validation loop and its second strtoull call,
and get_checked_or_null()'s JSON_TRY/JSON_INTERNAL_CATCH.

Bugfix as a consequence: get_checked_or_null()'s JSON_TRY/
JSON_INTERNAL_CATCH was dead code under JSON_NOEXCEPTION (JSON_TRY
expands to "if(true)" and the catch to "if(false)", so JSON_THROW's
std::abort() ran unconditionally), meaning value() and contains() would
abort instead of returning the default/false for an out-of-range-sized
or oversized array index when exceptions are disabled (#5672). Switching
on parse_array_index()'s return value instead of relying on an actual
throw/catch fixes this: get_checked_or_null() now returns nullptr for
array_index_status::unresolved/exceeds_size_type in every build
configuration, and still calls JSON_THROW (aborting under
JSON_NOEXCEPTION, as before) only for a malformed index
(leading_zero/not_a_number), matching its documented @throw list.

All existing error ids, messages and diagnostic paths are unchanged; a
few reference tokens that used to fail contains()'s manual per-character
validation (e.g. "1a") now fail via array_index_status::unresolved
instead, with no observable difference since contains() only returns
bool.

Adds regression tests to unit-element_access2.cpp's "access on array
type" section covering value() with an index that exceeds size_type and
one with a trailing non-digit, both of which must yield the default
value rather than abort/throw.

Public API: no change.

Overlaps #5700, #5614 and #5692, which touch the contains() and
get_checked_or_null() array hunks; this change replaces those hunks with
calls into the new shared parser, so a rebase will need to re-apply
their token-handling changes (e.g. the empty-token case) on top of the
switch statements here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 4
2026-09-30 18:27:29 +02:00
Niels Lohmann e32de7bb8c De-duplicate from_json.hpp's map and array-fallback bodies
Several from_json() overload pairs in from_json.hpp were copies of each
other, so a fix has to be applied twice (as #5681 already does):

- from_json(..., std::map&) and from_json(..., std::unordered_map&) for
  non-string keys had identical 16-line bodies: array check, m.clear(),
  pair check loop, m.emplace(...). Route both through a new
  from_json_pair_array_to_map(j, m) helper.
- The from_json_array_impl priority_tag<1> and priority_tag<0> fallbacks
  ran the same std::transform/std::inserter loop, differing only in
  ret.reserve(j.size()). Merge them into one body and, modeled on the
  existing from_json_object_reserve, add a from_json_array_reserve pair
  so the reserve() call is only made for ConstructibleArrayType that
  support it.

Error ids (type_error.302), messages, diagnostic paths ((at(0)/at(1))
and behavior for types with/without reserve() are unchanged; only the
duplication is removed.

Public API: no change.

Overlaps #5681, which changes the "&j" to "&p" line in both map bodies;
the shared helper here should make that a one-line change instead of two
on rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 5
2026-09-30 18:16:32 +02:00
Niels Lohmann 93231c6c2f Factor the repeated JSON_HAS_RANGES/MinGW guard into one macro
The std::ranges view conversion (excluded on MinGW because of its
incomplete C++20 ranges support, #4916) was gated by the same
#if JSON_HAS_RANGES && !defined(__MINGW32__) condition at seven
independent sites in to_json.hpp and type_traits.hpp, with the MinGW
rationale duplicated in two of them and missing from the rest. Since the
sites come in matching pairs (one enables is_compatible_range_view and a
view-based overload, the other adds the exclusion to the
plain-array-type overload), a drift between any pair would produce an
ambiguous or missing overload on exactly one platform.

Add JSON_HAS_RANGE_VIEW_CONVERSION next to JSON_HAS_RANGES in
macro_scope.hpp, combining both conditions with the #4916 reasoning in
one place, #undef it in macro_unscope.hpp, and use it at all seven
sites. This does not fold the MinGW check into JSON_HAS_RANGES itself:
JSON_HAS_RANGES is user-overridable and also gates the
enable_borrowed_range specialization in iteration_proxy.hpp, which is
not excluded on MinGW.

No behavior or public API change: JSON_HAS_RANGE_VIEW_CONVERSION expands
to exactly the condition that was previously written out at each site.

Overlaps #5585, #5600 and #3575, which touch the same to_json.hpp and
type_traits.hpp lines; the change here is a mechanical
search-and-replace of the guard condition and should rebase cleanly.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 11
2026-09-30 18:12:58 +02:00
Niels Lohmann d254f62764 Move templated_json_throw into nlohmann::detail
templated_json_throw() was defined in macro_scope.hpp, which is included
outside NLOHMANN_JSON_NAMESPACE_BEGIN, so the helper leaked into the
global namespace as ::templated_json_throw with no ABI tag. Unqualified
lookup in NLOHMANN_JSON_SERIALIZE_ENUM_STRICT could then bind to a
same-named function declared in the user's own namespace instead, which
fails to compile with Clang ("does not name a template").

Move the helper next to the exception classes in exceptions.hpp, inside
nlohmann::detail, and call it qualified as
::nlohmann::detail::templated_json_throw<...>(...) from both macro
expansion sites. Rewrite the doc comment to give the real reason for the
helper (JSON_THROW may expand to code that discards its argument, e.g.
when exceptions are disabled) and fix the "supress" typo.

templated_json_throw was never released (added by #5151 after v3.12.0),
so it can be moved freely.

Adds a regression test that expands NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
inside a namespace declaring its own templated_json_throw.

Public API: no change (::templated_json_throw was an unreleased,
unintentional global-namespace leak with no callers relying on its
location).

Overlaps #5698, which rewrites the same two macro call lines; the
overlapping hunks are small and should be trivial to reconcile on
rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 2
2026-09-30 18:08:45 +02:00
Niels Lohmann 23aa069e63 Support any-rank C arrays in from_json, not just rank 1-4
from_json() for C arrays had four hand-unrolled overloads (rank 1-4,
added incrementally in #4262), each with its own nested loops. to_json()
already handles any rank recursively, so a rank-5+ C array could be
serialized but not read back with get_to()/get<>().

Replace the four overloads with one from_json() SFINAE-constrained on
get<remove_all_extents<T>::type>() existing, forwarding to a pair of
mutually recursive from_json_c_array_element() helpers: one assigns a
non-array element via get<T>(), the other loops over a array element and
recurses one dimension at a time. Each dimension still goes through at(),
so type_error.304/out_of_range.401 stay unchanged; ranks 1-4 keep their
existing behavior and semantics.

Adds rank-5 round-trip and mismatched-shape tests to unit-conversions.cpp.

Public API: additive only (rank 5+ C arrays become readable).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5708 item 1
2026-09-30 18:00:16 +02:00
Niels Lohmann b5fad30499 Fix misplaced and stale comments in JSON_HAS_RANGES and conversions
The JSON_HAS_RANGES feature-detection block had its libc++ comment
sitting above the clang+libstdc++ branch it does not describe,
leaving the libc++ branch uncommented and the clang+libstdc++ branch
without its own rationale. Move each comment to sit under its own
branch, and give the clang+libstdc++ branch (added in issue 5161) its
own one-line reason referencing that issue instead of reusing the
libc++ branch's comment. Also fix a duplicated-word typo ("in large
in large cpp files") in from_json.hpp, drop two unanswered 2017
design questions left as comments in type_traits.hpp and
from_json.hpp that no longer reflect open questions, and correct
NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's @since tag from 3.12.0 to
3.13.0, the release it was actually introduced in.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:11:36 +02:00
Niels Lohmann a9cc4243b7 Fix tautological clause in iter_impl's iterator category assertion
The static_assert meant to check the LegacyBidirectionalIterator
named requirement had a first clause comparing
std::bidirectional_iterator_tag to itself, which is always true and
checks nothing; only array_t::iterator was actually being checked,
despite the message claiming object iterators were checked too.
Drop the tautological clause, reword the message to describe what
is actually checked, and note that object_t may use a forward-only
iterator as long as reverse iteration and operator-- are unused.
The check is intentionally not extended to object_t::iterator, since
that would reject object types with forward-only iterators that
compile and work correctly today.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:54:59 +02:00
Niels Lohmann 03319da529 Remove duplicate const overload of json_pointer::get_checked
The const and non-const get_checked() overloads had byte-identical
50-line bodies, differing only in the signature. The remaining
template deduces a const-qualified BasicJsonType for const callers,
so at(), the out_of_range::create() calls and the bounds check all
still work.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:53:14 +02:00
Niels Lohmann ae9ef9dde1 Simplify is_ordered_map to reuse has_capacity
is_ordered_map re-detected capacity() with a C++03 sizeof/vararg
trick right after has_capacity did the same detection through
is_detected. For ordered_map, the old trick took the address of
std::vector::capacity, which [namespace.std]/6 makes unspecified.
Reuse has_capacity instead, which removes the unspecified-behavior
pointer-to-std-member and two NOLINT suppressions.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:50:52 +02:00
Niels Lohmann f9d4b51455 Remove unused would_call_std_* from NLOHMANN_CAN_CALL_STD_FUNC_IMPL
Besides detail::result_of_begin/end, which is_range and iterator_t use,
the macro defined a namespace detail2 with a tag type, a catch-all
overload and would_call_std_begin/end, plus would_call_std_begin/end
structs directly in namespace nlohmann. Nothing has used them since
they were added in #3020.

Reduce the macro to its detail part. Without the trailing struct the
';' after the two invocations would be an empty declaration that
-Wextra-semi flags, so drop it. macro_scope.hpp included
meta/detected.hpp only for this macro; all users of detected.hpp
include it (or type_traits.hpp) themselves, so remove the include.

Behavior and ABI are unchanged. The undocumented, untested and unused
names nlohmann::would_call_std_begin, nlohmann::would_call_std_end and
namespace nlohmann::detail2 are no longer declared.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:46:04 +02:00
Niels Lohmann 31883d06a8 Replace meta/logic.hpp with a disjunction trait
meta/logic.hpp added a second set of type-level boolean helpers
(cxpr_and, cxpr_or, cxpr_not, ...) next to the existing conjunction
and negation in type_traits.hpp. It was used only by one static_assert
in from_json_tuple_impl, two of its templates were never used, and it
was the only header without the license banner and relied on
transitive includes for <type_traits>.

Add the missing disjunction next to conjunction and negation, use the
three in the static_assert, and delete logic.hpp together with its
BUILD.bazel entry. same_sign now uses disjunction as well, which
resolves the 2022 TODO waiting for such a trait.

The static_assert accepts and rejects the same types as before. Only
names in nlohmann::detail change; behavior, public API and ABI are
unchanged.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:42:12 +02:00
Niels Lohmann 4c208728e7 Remove unused is_sax and is_detected_convertible
detail::is_sax had no user: the parser and the binary reader only use
is_sax_static_asserts, so is_sax was a second, unchecked copy of the
SAX event list. is_sax_static_asserts asserted boolean(bool) twice in
a row, and detail::is_detected_convertible was never used anywhere.

Remove all three and include <cstddef> for size_t instead of <cstdint>.
Only names in nlohmann::detail are removed; behavior, public API and ABI
are unchanged. The diagnostics for an incomplete SAX handler are the
same, apart from the duplicated boolean() message.

Part of #5708

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:38:54 +02:00
50 changed files with 1000 additions and 7526 deletions
-1
View File
@@ -53,7 +53,6 @@ cc_library(
"include/nlohmann/detail/meta/detected.hpp",
"include/nlohmann/detail/meta/identity_tag.hpp",
"include/nlohmann/detail/meta/is_sax.hpp",
"include/nlohmann/detail/meta/logic.hpp",
"include/nlohmann/detail/meta/std_fs.hpp",
"include/nlohmann/detail/meta/type_traits.hpp",
"include/nlohmann/detail/meta/void_t.hpp",
+2 -2
View File
@@ -496,7 +496,7 @@ bool key(string_t& val);
bool parse_error(std::size_t position, const std::string& last_token, const detail::exception& ex);
```
The return value of each function determines whether parsing should proceed. For `parse_error`, returning `true` [recovers from the error](https://json.nlohmann.me/features/parsing/error_recovery/): the parser repairs the input and continues.
The return value of each function determines whether parsing should proceed.
To implement your own SAX handler, proceed as follows:
@@ -504,7 +504,7 @@ To implement your own SAX handler, proceed as follows:
2. Create an object of your SAX interface class, e.g. `my_sax`.
3. Call `bool json::sax_parse(input, &my_sax)`; where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface.
Note the `sax_parse` function only returns a `bool` indicating whether the input was parsed without errors and no SAX event returned `false`. It does not return a `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error -- it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file [`json_sax.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/json_sax.hpp).
Note the `sax_parse` function only returns a `bool` indicating the result of the last executed SAX event. It does not return a `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error -- it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file [`json_sax.hpp`](https://github.com/nlohmann/json/blob/develop/include/nlohmann/detail/input/json_sax.hpp).
### STL-like access
+1 -4
View File
@@ -90,9 +90,7 @@ The SAX event lister must follow the interface of [`json_sax`](../json_sax/index
## Return value
`#!cpp true` if the input was parsed without errors and no SAX event returned `#!cpp false`; `#!cpp false` otherwise.
In particular, the result is `#!cpp false` for input with errors, even if the SAX parser recovered from all of them
(see [error recovery](../../features/parsing/error_recovery.md)).
return value of the last processed SAX event
## Exception safety
@@ -140,7 +138,6 @@ A UTF-8 byte order mark is silently ignored.
- Ignoring comments via `ignore_comments` added in version 3.9.0.
- Added `ignore_trailing_commas` in version 3.13.0.
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
- Recovering from parse errors (see [`parse_error`](../json_sax/parse_error.md)) added in version 3.13.0.
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
- `JSON_PRECISE_STREAM_POSITION` added in version 3.13.0 to optionally leave a `#!cpp std::istream` positioned right
after the parsed value when `strict` is `#!cpp false`.
+1 -2
View File
@@ -7,8 +7,7 @@ struct json_sax;
This class describes the SAX interface used by [sax_parse](../basic_json/sax_parse.md). Each function is called in
different situations while the input is parsed. The boolean return value informs the parser whether to continue
processing the input; for [`parse_error`](parse_error.md), it decides whether to
[recover from the error](../../features/parsing/error_recovery.md).
processing the input.
## Template parameters
+1 -24
View File
@@ -21,14 +21,7 @@ A parse error occurred.
## Return value
Whether to recover from the error:
- `#!cpp false` stops parsing.
- `#!cpp true` recovers from the error: the error is repaired and parsing continues. If that is not possible, which
happens in the binary formats when the end of the item with the error is unknown, the value read so far is completed
and parsing stops. See [error recovery](../../features/parsing/error_recovery.md) for how errors are repaired.
Either way, [`sax_parse`](../basic_json/sax_parse.md) returns `#!cpp false`.
Whether parsing should proceed (**must return `#!cpp false`**).
## Examples
@@ -46,22 +39,6 @@ Either way, [`sax_parse`](../basic_json/sax_parse.md) returns `#!cpp false`.
--8<-- "examples/sax_parse.output"
```
??? example
The example below shows how a SAX parser recovers from errors.
```cpp
--8<-- "examples/sax_parse__error_recovery.cpp"
```
Output:
```
--8<-- "examples/sax_parse__error_recovery.output"
```
## Version history
- Added in version 3.2.0.
- Returning `#!cpp true` recovers from the error since version 3.13.0; before, parsing stopped, but the result of
[`sax_parse`](../basic_json/sax_parse.md) could be wrong.
+6
View File
@@ -18,6 +18,10 @@ Deserializes an input stream to a JSON value.
the stream `i`
## Exception safety
Strong guarantee: if an exception is thrown, there are no changes in `j`.
## Exceptions
- Throws [`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101) in case of an unexpected token, or if
@@ -125,3 +129,5 @@ being read.
the stream; planned to become the default in version 4.0.0.
- Fixed a null pointer dereference for an `std::istream` without a stream buffer (now throws `parse_error.101`), and a
crash (`std::terminate`) when `i` has `eofbit` in its exception mask, in version 3.13.0.
- Changed to the strong exception safety guarantee in version 3.13.0: `j` is no longer left with a partially parsed
value if parsing throws.
@@ -1,43 +0,0 @@
#include <iostream>
#include <iomanip>
#include <nlohmann/json.hpp>
using json = nlohmann::json;
// a SAX parser that creates a JSON value like json::parse does, but that
// recovers from parse errors instead of stopping at the first one
class recovering_parser : public nlohmann::detail::json_sax_dom_parser<json>
{
public:
explicit recovering_parser(json& result)
: nlohmann::detail::json_sax_dom_parser<json>(result, false)
{}
bool parse_error(std::size_t position,
const std::string& /*last_token*/,
const json::exception& ex)
{
std::cout << "byte " << position << ": " << ex.what() << '\n';
// repair the input and continue
return true;
}
};
int main()
{
// JSON text with several mistakes that ends too early
const std::string text = R"({
"name": "Hello World",
"tags": ["a" "b",],
"valid": tru,
"size": 1.,
"nested": {"x": 1)";
json result;
recovering_parser sax(result);
const bool valid = json::sax_parse(text, &sax);
std::cout << "\nvalid JSON: " << std::boolalpha << valid << '\n'
<< std::setw(4) << result << std::endl;
}
@@ -1,19 +0,0 @@
byte 49: [json.exception.parse_error.101] parse error at line 3, column 20: syntax error while parsing array - unexpected string literal; expected ']'
byte 51: [json.exception.parse_error.101] parse error at line 3, column 22: syntax error while parsing value - unexpected ']'; expected '[', '{', or a literal
byte 70: [json.exception.parse_error.101] parse error at line 4, column 17: syntax error while parsing value - invalid literal; last read: '"valid": tru,'
byte 86: [json.exception.parse_error.101] parse error at line 5, column 15: syntax error while parsing value - invalid number; expected digit after '.'; last read: '1.,'
byte 109: [json.exception.parse_error.101] parse error at line 6, column 22: syntax error while parsing object - unexpected end of input; expected '}'
valid JSON: false
{
"name": "Hello World",
"nested": {
"x": 1
},
"size": 1,
"tags": [
"a",
"b"
],
"valid": null
}
@@ -1,121 +0,0 @@
# Error Recovery
By default, parsing stops at the first error. With the [SAX interface](sax_interface.md), you can instead ask the
parser to *recover*: to repair the error and continue, so that you get as much as possible out of malformed input, for
instance a file that was cut off, JSON edited by hand, or the output of a language model.
## Recovering from errors
The SAX parser's [`parse_error`](../../api/json_sax/parse_error.md) function is called for every error. Its return value
decides what happens next:
- `#!cpp false` stops parsing. This is what the SAX parsers of the library do, so [`parse`](../../api/basic_json/parse.md)
and [`accept`](../../api/basic_json/accept.md) never recover.
- `#!cpp true` repairs the error and continues parsing.
When recovering, the SAX parser still receives well-formed events: every `start_object` or `start_array` is followed by
the matching `end_object` or `end_array`, and every `key` is followed by exactly one value. A SAX parser that creates a
JSON value, such as the one in the example below, therefore gets a complete value. Parsing always ends, and
[`sax_parse`](../../api/basic_json/sax_parse.md) returns `#!cpp false` for input that is not valid JSON, even if every
error was repaired. Each token is reported at most once, and the SAX parser can stop at any error by returning
`#!cpp false`.
!!! example
The example below derives a SAX parser from the library's parser for `json` values (`json_sax_dom_parser`),
and recovers from all errors.
```cpp
--8<-- "examples/sax_parse__error_recovery.cpp"
```
Output:
```
--8<-- "examples/sax_parse__error_recovery.output"
```
## How errors are repaired
Each error is repaired with the smallest local edit: a missing separator is inserted, a stray token is removed, what can
be read of a broken string or number is kept, and a value that cannot be read at all becomes `#!json null`.
| Mistake | Repair | Example | Result |
|---------------------------|--------------------------------------------------------------------------------|------------------------------------------|----------------------------|
| missing `,` or `:` | inserted | `#!json [1 2]`, `#!json {"a" 1}` | `[1,2]`, `{"a":1}` |
| missing value | `#!json null` for an object key or between commas in an array | `#!json {"a":}`, `#!json [1,,2]` | `{"a":null}`, `[1,null,2]` |
| trailing comma | removed | `#!json [1,2,]` | `[1,2]` |
| broken string | invalid escapes and bytes are replaced (see below); a line break ends the string | `#!json ["a\qb"]` | `["aqb"]` |
| broken number | the longest valid beginning is kept | `#!json [1., 2e+]` | `[1,2]` |
| unreadable value | `#!json null` | `#!json [1, NaN, tru]` | `[1,null,null]` |
| number too large | passed as infinity, together with its text | `#!json [1e999]` | infinity (see below) |
| stray `:` | removed | `#!json ["a":1]` | `["a",1]` |
| member without a key | skipped up to the next `,` or `}` | `#!json {1:2, "b":3}` | `{"b":3}` |
| wrong closing bracket | closes the innermost array or object | `#!json {"a":[1,2}, "b":3}` | `{"a":[1,2],"b":3}` |
| input ends too early | all open arrays and objects are closed | `#!json {"a":[1,2` | `{"a":[1,2]}` |
| text before the value | skipped | `#!json )]}'{"a":1}` | `{"a":1}` |
In a string, an unknown escape like `\q` stands for the escaped character (`q`), as in JavaScript. An invalid `\u`
escape, a lone surrogate, and ill-formed UTF-8 are each replaced by U+FFFD (REPLACEMENT CHARACTER), and control
characters are kept. A string without its closing quote ends at the next line break or at the end of the input.
The input after the top-level value is not repaired: as without recovery, it is reported as an error, and parsing stops.
## Binary formats
The binary formats ([BJData](../binary_formats/bjdata.md), [BON8](../binary_formats/bon8.md),
[BSON](../binary_formats/bson.md), [CBOR](../binary_formats/cbor.md), [MessagePack](../binary_formats/messagepack.md),
and [UBJSON](../binary_formats/ubjson.md)) have no delimiters to find the next value by. So what can be repaired depends
on whether the end of the item with the error is known, a distinction that
[RFC 8949, Section 5.3](https://www.rfc-editor.org/rfc/rfc8949.html#section-5.3) makes for CBOR, too.
If the item is complete, but cannot be passed on as it is, it is replaced, and parsing continues after it:
| Mistake | Formats | Repair |
|---------------------------------------------------------------------|-----------------------------------------|-------------------------------------------------------------------------|
| tag | CBOR | ignored |
| simple value other than `false`, `true`, and `null`, like undefined | CBOR | `#!json null` |
| negative integer below the range of `number_integer_t` | CBOR | the nearest floating-point number |
| string that is not valid UTF-8 | BJData, BSON, CBOR, MessagePack, UBJSON | each ill-formed sequence becomes U+FFFD |
| character (`C`) that is not ASCII | BJData, UBJSON | U+FFFD |
| invalid high-precision number (`H`) | BJData, UBJSON | the longest valid beginning is kept, as for JSON text, or `#!json null` |
| high-precision number too large | BJData, UBJSON | passed as infinity, together with its text |
| object key that is not a string | BON8, CBOR, MessagePack | the member is skipped |
| element of a type the library does not read, like ObjectId or date | BSON | `#!json null` |
| string without its terminator | BSON | kept |
| document whose size does not match its content | BSON | kept |
CBOR tags and simple values are repaired as [RFC 8949, Section 6.1](https://www.rfc-editor.org/rfc/rfc8949.html#section-6.1)
suggests for converting CBOR to JSON. Note that [`sax_parse`](../../api/basic_json/sax_parse.md) has no parameter for
CBOR tags, so every tag is an error there; when recovering, tags are ignored like with
[`cbor_tag_handler_t::ignore`](../../api/basic_json/cbor_tag_handler_t.md).
After any other error, the end of the item is unknown: the input ended, a byte is not a valid type marker, or a size
cannot be right. Parsing then stops, and the value read so far is completed: a key that waits for its value gets
`#!json null`, and all open arrays and objects are closed. This keeps everything before the error of an input that was
cut off. The exception is BSON, which stores the size of every document: an element whose end is unknown gets
`#!json null`, the rest of its document is skipped, and parsing continues after the document.
## Limitations
- A repair is a guess. For example, `#!json {"a" "b": 1}` could be meant as `#!json {"a": "b"}` or as
`#!json {"a": null, "b": 1}`; it is repaired to the former. Treat recovered values as a best effort, and check the
reported errors.
- A closing bracket always closes the innermost array or object. If a bracket is missing rather than wrong, the
repair differs from the intention: `#!json {"a": {"b": [1, 2}, "c": 3}` is repaired to
`#!json {"a": {"b": [1, 2], "c": 3}}`, although `#!json {"a": {"b": [1, 2]}, "c": 3}` may have been meant.
- Keys without quotes, and strings in single quotes, are not supported; such members are skipped.
- In the binary formats, a member that is skipped because its key is not a string is lost, and so are the elements of a
BSON document after one whose end is unknown.
- A number that is too large for `number_float_t` is passed as positive or negative infinity. The SAX parser's
`number_float` also gets the number's text, but a JSON value cannot store it, and
[`dump`](../../api/basic_json/dump.md) serializes infinity as `#!json null`.
- When parsing is not strict (see [`sax_parse`](../../api/basic_json/sax_parse.md)), a repair may read parts of the
input after the value, for instance of the next value in a stream of concatenated values.
## See also
- [SAX interface](sax_interface.md) - implement a custom SAX handler
- [`parse_error`](../../api/json_sax/parse_error.md) - the SAX event for parse errors
- [`sax_parse`](../../api/basic_json/sax_parse.md) - generate SAX events
- [parsing and exceptions](parse_exceptions.md) - control error handling
+1 -2
View File
@@ -65,7 +65,7 @@ You can influence a DOM parse without switching to the SAX interface by passing
When the input is not valid JSON, the `parse` function throws an exception by default. If exceptions are undesired or
unavailable, the parser can instead return a discarded value, or [`accept`](../../api/basic_json/accept.md) can be used
to only check whether an input is valid JSON. See [parsing and exceptions](parse_exceptions.md) for the available
options. To get as much as possible out of malformed input, a SAX parser can [recover from errors](error_recovery.md).
options.
## See also
@@ -76,4 +76,3 @@ options. To get as much as possible out of malformed input, a SAX parser can [re
- [parser callbacks](parser_callbacks.md) - influence the parsing by a callback function
- [SAX interface](sax_interface.md) - implement a custom SAX handler
- [parsing and exceptions](parse_exceptions.md) - control error handling
- [error recovery](error_recovery.md) - get as much as possible out of malformed input
@@ -64,8 +64,7 @@ bool parse_error(std::size_t position,
const json::exception& ex);
```
The return value decides whether to stop parsing (`#!cpp false`) or to repair the error and continue
(`#!cpp true`); see [error recovery](error_recovery.md) for the latter.
The return value indicates whether the parsing should continue, so the function should usually return `#!cpp false`.
??? example
@@ -60,8 +60,7 @@ bool key(string_t& val);
bool parse_error(std::size_t position, const std::string& last_token, const json::exception& ex);
```
The return value of each function determines whether parsing should proceed. For `parse_error`, returning
`#!cpp true` [recovers from the error](error_recovery.md).
The return value of each function determines whether parsing should proceed.
To implement your own SAX handler, proceed as follows:
@@ -69,7 +68,7 @@ To implement your own SAX handler, proceed as follows:
2. Create an object of your SAX interface class, e.g. `my_sax`.
3. Call `#!cpp bool json::sax_parse(input, &my_sax);` where the first parameter can be any input like a string or an input stream and the second parameter is a pointer to your SAX interface.
Note the `sax_parse` function only returns a `#!cpp bool` indicating whether the input was parsed without errors and no SAX event returned `#!cpp false`. It does not return `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error - it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file `json_sax.hpp`.
Note the `sax_parse` function only returns a `#!cpp bool` indicating the result of the last executed SAX event. It does not return `json` value - it is up to you to decide what to do with the SAX events. Furthermore, no exceptions are thrown in case of a parse error - it is up to you what to do with the exception object passed to your `parse_error` implementation. Internally, the SAX interface is used for the DOM parser (class `json_sax_dom_parser`) as well as the acceptor (`json_sax_acceptor`), see file `json_sax.hpp`.
## See also
@@ -389,6 +389,7 @@ using array_t = ArrayType<basic_json, AllocatorType<basic_json>>;
| Functionality | Additional requirement |
|-----------------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| [`diff`](../../api/basic_json/diff.md), [`items`](../../api/basic_json/items.md), [`std::hash`](../../api/basic_json/std_hash.md) | conversion of a `#!cpp std::size_t` to `StringType`: either assignability from the result of `#!cpp std::to_string`, or an ADL overload `#!cpp void int_to_string(StringType&, std::size_t)` |
| [`operator/(std::size_t)`](../../api/json_pointer/operator_slash.md) | the same conversion of a `#!cpp std::size_t` to `StringType` as `diff`, `items`, and `std::hash` above |
| [`std::hash<basic_json>`](../../api/basic_json/std_hash.md) | additionally a specialization of `#!cpp std::hash<StringType>` |
| [`to_bson`](../../api/basic_json/to_bson.md) | `find(value_type)` and `npos` |
| [`parse`](../../api/basic_json/parse.md) from a `string_t` | the input adapters must accept it; otherwise pass a character range |
-1
View File
@@ -87,7 +87,6 @@ nav:
- features/object_order.md
- Parsing:
- features/parsing/index.md
- features/parsing/error_recovery.md
- features/parsing/json_lines.md
- features/parsing/parse_exceptions.md
- features/parsing/parser_callbacks.md
@@ -27,7 +27,6 @@
#include <nlohmann/detail/meta/identity_tag.hpp>
#include <nlohmann/detail/meta/std_fs.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/meta/logic.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/value_t.hpp>
@@ -211,62 +210,29 @@ inline void from_json(const BasicJsonType& j, std::valarray<T>& l)
});
}
// element is not itself a C array: read it directly
template<typename BasicJsonType, typename T>
auto from_json_c_array_element(const BasicJsonType& j, T& e)
-> decltype(e = j.template get<T>(), void())
{
e = j.template get<T>();
}
// element is itself a C array: recurse one dimension at a time, so any rank is supported
template<typename BasicJsonType, typename T, std::size_t N>
auto from_json(const BasicJsonType& j, T (&arr)[N]) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
-> decltype(j.template get<T>(), void())
void from_json_c_array_element(const BasicJsonType& j, T (&arr)[N]) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
{
for (std::size_t i = 0; i < N; ++i)
{
arr[i] = j.at(i).template get<T>();
from_json_c_array_element(j.at(i), arr[i]);
}
}
template<typename BasicJsonType, typename T, std::size_t N1, std::size_t N2>
auto from_json(const BasicJsonType& j, T (&arr)[N1][N2]) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
-> decltype(j.template get<T>(), void())
template<typename BasicJsonType, typename T, std::size_t N>
auto from_json(const BasicJsonType& j, T (&arr)[N]) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
-> decltype(j.template get<typename std::remove_all_extents<T>::type>(), void())
{
for (std::size_t i1 = 0; i1 < N1; ++i1)
{
for (std::size_t i2 = 0; i2 < N2; ++i2)
{
arr[i1][i2] = j.at(i1).at(i2).template get<T>();
}
}
}
template<typename BasicJsonType, typename T, std::size_t N1, std::size_t N2, std::size_t N3>
auto from_json(const BasicJsonType& j, T (&arr)[N1][N2][N3]) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
-> decltype(j.template get<T>(), void())
{
for (std::size_t i1 = 0; i1 < N1; ++i1)
{
for (std::size_t i2 = 0; i2 < N2; ++i2)
{
for (std::size_t i3 = 0; i3 < N3; ++i3)
{
arr[i1][i2][i3] = j.at(i1).at(i2).at(i3).template get<T>();
}
}
}
}
template<typename BasicJsonType, typename T, std::size_t N1, std::size_t N2, std::size_t N3, std::size_t N4>
auto from_json(const BasicJsonType& j, T (&arr)[N1][N2][N3][N4]) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
-> decltype(j.template get<T>(), void())
{
for (std::size_t i1 = 0; i1 < N1; ++i1)
{
for (std::size_t i2 = 0; i2 < N2; ++i2)
{
for (std::size_t i3 = 0; i3 < N3; ++i3)
{
for (std::size_t i4 = 0; i4 < N4; ++i4)
{
arr[i1][i2][i3][i4] = j.at(i1).at(i2).at(i3).at(i4).template get<T>();
}
}
}
}
from_json_c_array_element(j, arr);
}
template<typename BasicJsonType>
@@ -286,20 +252,33 @@ auto from_json_array_impl(const BasicJsonType& j, std::array<T, N>& arr,
}
}
// reserve() is called through this pair (modeled on from_json_object_reserve)
// so from_json_array_impl below has a single body for both ConstructibleArrayType
// that support reserve() and those that don't.
template<typename ConstructibleArrayType>
auto from_json_array_reserve(ConstructibleArrayType& arr, typename ConstructibleArrayType::size_type size, priority_tag<1> /*unused*/)
-> decltype(arr.reserve(size), void())
{
arr.reserve(size);
}
template<typename ConstructibleArrayType>
inline void from_json_array_reserve(ConstructibleArrayType& /*arr*/, std::size_t /*size*/, priority_tag<0> /*unused*/)
{}
template<typename BasicJsonType, typename ConstructibleArrayType,
enable_if_t<
std::is_assignable<ConstructibleArrayType&, ConstructibleArrayType>::value,
int> = 0>
auto from_json_array_impl(const BasicJsonType& j, ConstructibleArrayType& arr, priority_tag<1> /*unused*/)
-> decltype(
arr.reserve(std::declval<typename ConstructibleArrayType::size_type>()),
j.template get<typename ConstructibleArrayType::value_type>(),
void())
{
using std::end;
ConstructibleArrayType ret;
ret.reserve(j.size());
from_json_array_reserve(ret, j.size(), priority_tag<1> {});
std::transform(j.begin(), j.end(),
std::inserter(ret, end(ret)), [](const BasicJsonType & i)
{
@@ -310,27 +289,6 @@ auto from_json_array_impl(const BasicJsonType& j, ConstructibleArrayType& arr, p
arr = std::move(ret);
}
template<typename BasicJsonType, typename ConstructibleArrayType,
enable_if_t<
std::is_assignable<ConstructibleArrayType&, ConstructibleArrayType>::value,
int> = 0>
inline void from_json_array_impl(const BasicJsonType& j, ConstructibleArrayType& arr,
priority_tag<0> /*unused*/)
{
using std::end;
ConstructibleArrayType ret;
std::transform(
j.begin(), j.end(), std::inserter(ret, end(ret)),
[](const BasicJsonType & i)
{
// get<BasicJsonType>() returns *this, this won't call a from_json
// method when value_type is BasicJsonType
return i.template get<typename ConstructibleArrayType::value_type>();
});
arr = std::move(ret);
}
template < typename BasicJsonType, typename ConstructibleArrayType,
enable_if_t <
is_constructible_array_type<BasicJsonType, ConstructibleArrayType>::value&&
@@ -433,9 +391,7 @@ inline void from_json(const BasicJsonType& j, ConstructibleObjectType& obj)
}
// overload for arithmetic types, not chosen for basic_json template arguments
// (BooleanType, etc.); note: Is it really necessary to provide explicit
// overloads for boolean_t etc. in case of a custom BooleanType which is not
// an arithmetic type?
// (BooleanType, etc.)
template < typename BasicJsonType, typename ArithmeticType,
enable_if_t <
std::is_arithmetic<ArithmeticType>::value&&
@@ -531,7 +487,7 @@ inline void from_json_tuple_impl(BasicJsonType&& j, std::pair<A1, A2>& p, priori
template<typename BasicJsonType, typename... Args>
std::tuple<Args...> from_json_tuple_impl(BasicJsonType&& j, identity_tag<std::tuple<Args...>> /*unused*/, priority_tag<2> /*unused*/)
{
static_assert(cxpr_and<cxpr_or<cxpr_not<std::is_reference<Args>>, is_compatible_reference_type<BasicJsonType, Args>>...>::value,
static_assert(conjunction<disjunction<negation<std::is_reference<Args>>, is_compatible_reference_type<BasicJsonType, Args>>...>::value,
"Can not return a tuple containing references to types not contained in a Json, try Json::get_to()");
return from_json_tuple_impl_base<1, Args...>(std::forward<BasicJsonType>(j), index_sequence_for<Args...> {});
}
@@ -554,10 +510,10 @@ auto from_json(BasicJsonType&& j, TupleRelated&& t)
return from_json_tuple_impl(std::forward<BasicJsonType>(j), std::forward<TupleRelated>(t), priority_tag<3> {});
}
template < typename BasicJsonType, typename Key, typename Value, typename Compare, typename Allocator,
typename = enable_if_t < !std::is_constructible <
typename BasicJsonType::string_t, Key >::value >>
inline void from_json(const BasicJsonType& j, std::map<Key, Value, Compare, Allocator>& m)
// shared body for std::map/std::unordered_map with a non-string Key: both
// containers are read from an array of [key, value] pairs the same way
template<typename BasicJsonType, typename MapType>
inline void from_json_pair_array_to_map(const BasicJsonType& j, MapType& m)
{
if (JSON_HEDLEY_UNLIKELY(!j.is_array()))
{
@@ -570,33 +526,29 @@ inline void from_json(const BasicJsonType& j, std::map<Key, Value, Compare, Allo
{
JSON_THROW(type_error::create(302, concat("type must be array, but is ", p.type_name()), &p));
}
m.emplace(p.at(0).template get<Key>(), p.at(1).template get<Value>());
m.emplace(p.at(0).template get<typename MapType::key_type>(), p.at(1).template get<typename MapType::mapped_type>());
}
}
template < typename BasicJsonType, typename Key, typename Value, typename Compare, typename Allocator,
typename = enable_if_t < !std::is_constructible <
typename BasicJsonType::string_t, Key >::value >>
inline void from_json(const BasicJsonType& j, std::map<Key, Value, Compare, Allocator>& m)
{
from_json_pair_array_to_map(j, m);
}
template < typename BasicJsonType, typename Key, typename Value, typename Hash, typename KeyEqual, typename Allocator,
typename = enable_if_t < !std::is_constructible <
typename BasicJsonType::string_t, Key >::value >>
inline void from_json(const BasicJsonType& j, std::unordered_map<Key, Value, Hash, KeyEqual, Allocator>& m)
{
if (JSON_HEDLEY_UNLIKELY(!j.is_array()))
{
JSON_THROW(type_error::create(302, concat("type must be array, but is ", j.type_name()), &j));
}
m.clear();
for (const auto& p : j)
{
if (JSON_HEDLEY_UNLIKELY(!p.is_array()))
{
JSON_THROW(type_error::create(302, concat("type must be array, but is ", p.type_name()), &p));
}
m.emplace(p.at(0).template get<Key>(), p.at(1).template get<Value>());
}
from_json_pair_array_to_map(j, m);
}
#if JSON_HAS_FILESYSTEM || JSON_HAS_EXPERIMENTAL_FILESYSTEM
// Workaround for MSVC 19.51 (and possibly later): in large in large cpp files, the compiler may fail to resolve with generic has_from_json (issue #4996)
// Workaround for MSVC 19.51 (and possibly later): in large cpp files, the compiler may fail to resolve with generic has_from_json (issue #4996)
template<typename BasicJsonType>
struct has_from_json<BasicJsonType, std_fs::path, void> : std::true_type {};
@@ -178,7 +178,7 @@ struct external_constructor<value_t::array>
template < typename BasicJsonType, typename CompatibleArrayType,
enable_if_t < !std::is_same<CompatibleArrayType, typename BasicJsonType::array_t>::value
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#if JSON_HAS_RANGE_VIEW_CONVERSION
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
, int > = 0 >
@@ -222,9 +222,7 @@ struct external_constructor<value_t::array>
j.assert_invariant();
}
// std::ranges does not work properly on MinGW due to incomplete C++20 support
// see https://github.com/nlohmann/json/issues/4916
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#if JSON_HAS_RANGE_VIEW_CONVERSION
template<typename BasicJsonType, typename CompatibleArrayType,
enable_if_t<is_compatible_range_view<std::remove_cvref_t<CompatibleArrayType>>::value, int> = 0>
static void construct(BasicJsonType& j, CompatibleArrayType && arr)
@@ -379,7 +377,7 @@ template < typename BasicJsonType, typename CompatibleArrayType,
!std::is_same<typename BasicJsonType::binary_t, CompatibleArrayType>::value&&
!is_compatible_binary_type<BasicJsonType, CompatibleArrayType>::value&&
!is_basic_json<CompatibleArrayType>::value
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#if JSON_HAS_RANGE_VIEW_CONVERSION
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
,
@@ -389,7 +387,7 @@ inline void to_json(BasicJsonType& j, const CompatibleArrayType& arr)
external_constructor<value_t::array>::construct(j, arr);
}
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#if JSON_HAS_RANGE_VIEW_CONVERSION
template < typename BasicJsonType, typename T,
enable_if_t < is_compatible_range_view<std::remove_cvref_t<T>>::value
&& !is_compatible_string_type<BasicJsonType, std::remove_cvref_t<T>>::value
+21
View File
@@ -286,6 +286,27 @@ class other_error : public exception
other_error(int id_, const char* what_arg) : exception(id_, what_arg) {}
};
/*!
@brief helper function to call JSON_THROW from a template
@note JSON_THROW is a macro that, depending on the JSON_THROW_USER /
JSON_TRY_USER / JSON_NOEXCEPTION configuration, may expand to code
that does not reference its argument (e.g. `std::abort()`), which
would trigger a compilation error if the argument's type depends on
a template parameter that is otherwise unused. Wrapping the call in
a templated function avoids this and gives the compiler a single
place to see the (possibly unused) parameter.
*/
template<typename ExceptionType>
void templated_json_throw(ExceptionType exception)
{
JSON_THROW(exception);
// JSON_THROW may expand to code that discards its argument (e.g. when
// exceptions are disabled) - the cast below avoids an unused-parameter
// warning with -Werror in that case
(void)exception;
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
File diff suppressed because it is too large Load Diff
+4 -9
View File
@@ -131,9 +131,7 @@ struct json_sax
@param[in] position the position in the input where the error occurs
@param[in] last_token the last read token
@param[in] ex an exception object describing the error
@return whether to recover from the error: false stops parsing; true
repairs the error and continues, or, if that is not possible,
stops after completing the value read so far
@return whether parsing should proceed (must return false)
*/
virtual bool parse_error(std::size_t position,
const std::string& last_token,
@@ -188,12 +186,9 @@ a pointer to the respective array or object for each recursion depth.
After successful parsing, the value that is passed by reference to the
constructor contains the parsed value.
@tparam BasicJsonType the JSON type
@tparam InputAdapterType the input adapter of the lexer that can be passed to
the constructor to record diagnostic positions; it
does not matter if no lexer is passed
@tparam BasicJsonType the JSON type
*/
template<typename BasicJsonType, typename InputAdapterType = string_input_adapter_type>
template<typename BasicJsonType, typename InputAdapterType>
class json_sax_dom_parser
{
public:
@@ -510,7 +505,7 @@ class json_sax_dom_parser
lexer_t* m_lexer_ref = nullptr;
};
template<typename BasicJsonType, typename InputAdapterType = string_input_adapter_type>
template<typename BasicJsonType, typename InputAdapterType>
class json_sax_dom_callback_parser
{
public:
+1 -590
View File
@@ -10,7 +10,6 @@
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint8_t
#include <cstdio> // snprintf
#include <initializer_list> // initializer_list
#include <string> // char_traits, string
@@ -438,16 +437,8 @@ class lexer : public lexer_base<BasicJsonType>
if (0xD800 <= codepoint1 && codepoint1 <= 0xDBFF)
{
// expect next \uxxxx entry
if (JSON_HEDLEY_LIKELY(get() == '\\'))
if (JSON_HEDLEY_LIKELY(get() == '\\' && get() == 'u'))
{
if (JSON_HEDLEY_UNLIKELY(get() != 'u'))
{
// current is the character escaped by the backslash
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
string_error_resume = resume_kind::escaped_character;
return token_type::parse_error;
}
const int codepoint2 = get_codepoint();
if (JSON_HEDLEY_UNLIKELY(codepoint2 == -1))
@@ -472,11 +463,7 @@ class lexer : public lexer_base<BasicJsonType>
}
else
{
// the second escape was read completely and is a
// code point of its own
error_message = "invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF";
string_error_resume = resume_kind::after_escape;
string_error_codepoint = codepoint2;
return token_type::parse_error;
}
}
@@ -490,9 +477,7 @@ class lexer : public lexer_base<BasicJsonType>
{
if (JSON_HEDLEY_UNLIKELY(0xDC00 <= codepoint1 && codepoint1 <= 0xDFFF))
{
// the escape was read completely
error_message = "invalid string: surrogate U+DC00..U+DFFF must follow U+D800..U+DBFF";
string_error_resume = resume_kind::after_escape;
return token_type::parse_error;
}
}
@@ -2145,573 +2130,6 @@ scan_number_done:
}
}
/////////////////////
// error recovery
/////////////////////
/*!
@brief make the best of the token that scan() rejected
Called by the parser after scan() returned token_type::parse_error and the
SAX parser asked to recover from the error (see #3989). Keeps what can be
read of the token and skips the rest:
- A string keeps its characters. An unknown escape stands for the escaped
character itself (as in JavaScript), an invalid `\u` escape and ill-formed
UTF-8 become U+FFFD, and a control character is kept. A line break or the
end of the input ends a string that lacks its closing quote.
- A number keeps its longest valid prefix, e.g. `1` for `1.` or `1e+`.
- A block comment that is not closed runs to the end of the input.
- Anything else is skipped.
The rest of an invalid token is skipped up to the next delimiter
(whitespace, a structural character, or a quote). A delimiter that the
invalid token consumed is returned to the input, so that the next scan()
reads it.
@return token_type::value_string or a number token type if a string or a
number could be read, token_type::end_of_input for a block comment
that is not closed, token_type::uninitialized otherwise
*/
token_type recover_token()
{
const resume_kind resume = string_error_resume;
const int codepoint = string_error_codepoint;
string_error_resume = resume_kind::character;
string_error_codepoint = -1;
if (error_message_starts_with("invalid string"))
{
return recover_string(resume, codepoint);
}
if (error_message_starts_with("invalid number"))
{
return recover_number();
}
if (error_message_starts_with("invalid comment; missing"))
{
// the comment runs to the end of the input
return token_type::end_of_input;
}
skip_to_delimiter();
return token_type::uninitialized;
}
/*!
@brief return the token that scan() read last to the input, so that the
next scan() reads it again
Called by the parser when recovering from an error. The token must be a
single character (',', ':', '[', ']', '{', or '}') or the end of the
input, and scan() must have read it last.
*/
void unget_token()
{
JSON_ASSERT(!next_unget);
unget();
}
/*!
@brief let the token string for the next error begin at the current character
The token string of an error reaches back to the beginning of the last
string or number. After an error, the parser calls this function so that
the next error does not report (and, with many errors, copy) everything
read since then.
*/
void restart_token_string()
{
restart_token_string_impl(std::integral_constant<bool, lazy_token_string> {});
}
private:
/// how recover_string() continues after the error scan_string() reported
enum class resume_kind : std::uint8_t
{
/// current is the next character of the string (or the end of input)
character,
/// current is the character escaped by the preceding backslash
escaped_character,
/// current is the last character of a complete escape
after_escape
};
/// whether error_message begins with @a prefix
bool error_message_starts_with(const char* prefix) const noexcept
{
const char* message = error_message;
while (*prefix != '\0')
{
if (*message++ != *prefix++)
{
return false;
}
}
return true;
}
/// whether current ends an invalid token (see recover_token())
bool current_is_delimiter() const noexcept
{
switch (current)
{
case ' ':
case '\t':
case '\n':
case '\r':
case '[':
case ']':
case '{':
case '}':
case ',':
case ':':
case '\"':
#if !JSON_STRICT_NUL_HANDLING
case '\0':
#endif
case char_traits<char_type>::eof():
return true;
case '/':
return ignore_comments;
default:
return false;
}
}
/// skip the rest of an invalid token and return its delimiter to the input
void skip_to_delimiter()
{
while (!current_is_delimiter())
{
get();
}
if (current != char_traits<char_type>::eof())
{
unget();
}
}
/// append U+FFFD REPLACEMENT CHARACTER to token_buffer
void add_replacement_character()
{
add(0xEF);
add(0xBF);
add(0xBD);
}
/// append the UTF-8 encoding of @a codepoint (not a surrogate) to token_buffer
void add_codepoint(const int codepoint)
{
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
const auto cp = static_cast<unsigned int>(codepoint);
if (cp < 0x80)
{
add(static_cast<char_int_type>(cp));
}
else if (cp <= 0x7FF)
{
add(static_cast<char_int_type>(0xC0u | (cp >> 6u)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
else if (cp <= 0xFFFF)
{
add(static_cast<char_int_type>(0xE0u | (cp >> 12u)));
add(static_cast<char_int_type>(0x80u | ((cp >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
else
{
add(static_cast<char_int_type>(0xF0u | (cp >> 18u)));
add(static_cast<char_int_type>(0x80u | ((cp >> 12u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | ((cp >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (cp & 0x3Fu)));
}
}
/// append a code point read from a `\u` escape; a surrogate becomes U+FFFD
void add_escaped_codepoint(const int codepoint)
{
if (0xD800 <= codepoint && codepoint <= 0xDFFF)
{
add_replacement_character();
}
else
{
add_codepoint(codepoint);
}
}
/*!
@brief remove an incomplete UTF-8 sequence from the end of token_buffer
next_byte_in_range() adds the bytes of a sequence as it checks them, so
when it rejects a byte, the beginning of the sequence is already in
token_buffer, which otherwise holds only complete sequences.
@return whether an incomplete sequence was removed
*/
bool remove_incomplete_utf8_sequence()
{
std::size_t lead = token_buffer.size();
std::size_t continuation_bytes = 0;
while (lead > 0 && continuation_bytes < 3
&& (static_cast<unsigned char>(token_buffer[lead - 1]) & 0xC0u) == 0x80u)
{
--lead;
++continuation_bytes;
}
if (lead == 0)
{
return false;
}
const auto lead_byte = static_cast<unsigned char>(token_buffer[lead - 1]);
std::size_t expected = 0;
if (lead_byte >= 0xF0)
{
expected = 3;
}
else if (lead_byte >= 0xE0)
{
expected = 2;
}
else if (lead_byte >= 0xC0)
{
expected = 1;
}
if (continuation_bytes >= expected)
{
return false;
}
token_buffer.resize(lead - 1);
return true;
}
/*!
@brief read the UTF-8 sequence that begins with current, which is not ASCII
@return whether the next character must be read; false if current still
needs to be handled, because it does not belong to the sequence
*/
bool recover_utf8_sequence()
{
// the number of continuation bytes and the range of the first one;
// see the ranges in scan_string()
std::size_t count = 0;
char_int_type low = 0x80;
char_int_type high = 0xBF;
if (current >= 0xC2 && current <= 0xDF)
{
count = 1;
}
else if (current >= 0xE0 && current <= 0xEF)
{
count = 2;
low = (current == 0xE0) ? 0xA0 : 0x80;
high = (current == 0xED) ? 0x9F : 0xBF;
}
else if (current >= 0xF0 && current <= 0xF4)
{
count = 3;
low = (current == 0xF0) ? 0x90 : 0x80;
high = (current == 0xF4) ? 0x8F : 0xBF;
}
else
{
// an ill-formed byte
add_replacement_character();
return true;
}
const std::size_t start = token_buffer.size();
add(current);
for (std::size_t i = 0; i < count; ++i)
{
get();
if (current < low || current > high)
{
token_buffer.resize(start);
add_replacement_character();
return false;
}
add(current);
low = 0x80;
high = 0xBF;
}
return true;
}
/*!
@brief read the low surrogate that must follow the high surrogate @a high
@return whether the next character must be read; false if current still
needs to be handled
*/
bool recover_low_surrogate(int high)
{
while (true)
{
if (get() != '\\')
{
add_replacement_character();
return false;
}
if (get() != 'u')
{
add_replacement_character();
// not 'u', so this does not come back here
return recover_escape();
}
const int low = get_codepoint();
if (low == -1)
{
add_replacement_character();
return false;
}
if (0xDC00 <= low && low <= 0xDFFF)
{
add_codepoint(static_cast<int>((static_cast<unsigned int>(high) << 10u)
+ static_cast<unsigned int>(low) - 0x35FDC00u));
return true;
}
// high has no low surrogate
add_replacement_character();
if (low < 0xD800 || low > 0xDBFF)
{
add_codepoint(low);
return true;
}
// another high surrogate
high = low;
}
}
/*!
@brief read the escape whose backslash was read; current is the escaped character
@return whether the next character must be read; false if current still
needs to be handled
*/
bool recover_escape()
{
switch (current)
{
case '\"':
add('\"');
return true;
case '\\':
add('\\');
return true;
case '/':
add('/');
return true;
case 'b':
add('\b');
return true;
case 'f':
add('\f');
return true;
case 'n':
add('\n');
return true;
case 'r':
add('\r');
return true;
case 't':
add('\t');
return true;
case 'u':
{
const int codepoint = get_codepoint();
if (codepoint == -1)
{
add_replacement_character();
return false;
}
if (0xD800 <= codepoint && codepoint <= 0xDBFF)
{
return recover_low_surrogate(codepoint);
}
add_escaped_codepoint(codepoint);
return true;
}
// an unknown escape stands for the escaped character
default:
return false;
}
}
/*!
@brief read the rest of a string after scan_string() rejected it
token_buffer holds what scan_string() read before the error. See
recover_token() for how errors are repaired.
@param[in] resume how to continue, see resume_kind
@param[in] codepoint for a high surrogate followed by an escape of another
code point: that code point; -1 otherwise
*/
token_type recover_string(const resume_kind resume, const int codepoint)
{
// whether the next character must be read before it can be handled
bool fetch = false;
if (error_message_starts_with("invalid string: surrogate")
|| error_message_starts_with("invalid string: '\\u'")
|| (error_message_starts_with("invalid string: ill-formed UTF-8")
&& remove_incomplete_utf8_sequence()))
{
add_replacement_character();
}
switch (resume)
{
case resume_kind::escaped_character:
fetch = recover_escape();
break;
case resume_kind::after_escape:
if (0xD800 <= codepoint && codepoint <= 0xDBFF)
{
fetch = recover_low_surrogate(codepoint);
}
else
{
if (codepoint != -1)
{
add_escaped_codepoint(codepoint);
}
fetch = true;
}
break;
case resume_kind::character:
default:
break;
}
while (true)
{
if (fetch)
{
get();
}
fetch = true;
switch (current)
{
case '\"':
// a line break or the end of the input ends a string that
// lacks its closing quote
case '\n':
case '\r':
case char_traits<char_type>::eof():
return token_type::value_string;
#if !JSON_STRICT_NUL_HANDLING
case '\0':
// the end of the input, see scan()
unget();
return token_type::value_string;
#endif
case '\\':
get();
fetch = recover_escape();
break;
default:
if (current < 0x80)
{
// including control characters
add(current);
}
else
{
fetch = recover_utf8_sequence();
}
break;
}
}
}
/*!
@brief keep the longest valid prefix of a number that scan_number() rejected
token_buffer holds the characters scan_number() accepted before the error,
so the prefix ends at its last digit.
*/
token_type recover_number()
{
// only size(), operator[], and resize() are used, which every string
// type the library supports provides
std::size_t length = token_buffer.size();
while (length != 0 && (token_buffer[length - 1] < '0' || token_buffer[length - 1] > '9'))
{
--length;
}
token_buffer.resize(length);
if (length == 0)
{
skip_to_delimiter();
return token_type::uninitialized;
}
if (decimal_point_position >= length)
{
decimal_point_position = std::string::npos;
}
std::size_t exponent = std::string::npos;
for (std::size_t i = 0; i < length; ++i)
{
if (token_buffer[i] == 'e' || token_buffer[i] == 'E')
{
exponent = i;
break;
}
}
const std::size_t mantissa_end = (exponent == std::string::npos) ? length : exponent;
token_type number_type = token_type::value_unsigned;
if (decimal_point_position != std::string::npos || exponent != std::string::npos)
{
number_type = token_type::value_float;
}
else if (token_buffer[0] == '-')
{
number_type = token_type::value_integer;
}
const token_type result = convert_number(number_type, mantissa_end);
skip_to_delimiter();
return result;
}
/// seekable adapter: the token string begins at current, which was consumed
void restart_token_string_impl(std::true_type /*lazy*/) noexcept
{
const std::size_t consumed = ia.get_consumed_count();
token_string_start = (consumed > 0 && current != char_traits<char_type>::eof()) ? consumed - 1 : consumed;
}
/// streaming adapter: the token string begins at current; a character
/// that was put back is copied again when it is read again
void restart_token_string_impl(std::false_type /*lazy*/)
{
token_string.clear();
if (!next_unget && current != char_traits<char_type>::eof())
{
token_string.push_back(char_traits<char_type>::to_char_type(current));
}
}
private:
/// input adapter
InputAdapterType ia;
@@ -2752,13 +2170,6 @@ scan_number_done:
/// a description of occurred lexer errors
const char* error_message = "";
/// how recover_token() continues a string that scan_string() rejected;
/// set only on the error paths that need more than error_message
resume_kind string_error_resume = resume_kind::character;
/// the code point of the second escape when a high surrogate is followed
/// by an escape that is not a low surrogate; -1 otherwise
int string_error_codepoint = -1;
// number values
number_integer_t value_integer = 0;
number_unsigned_t value_unsigned = 0;
+51 -663
View File
@@ -98,7 +98,7 @@ class parser
if (callback)
{
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
sax_parse_internal<false>(&sdp);
sax_parse_internal(&sdp);
if (strict)
{
@@ -135,7 +135,7 @@ class parser
else
{
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
sax_parse_internal<false>(&sdp);
sax_parse_internal(&sdp);
if (strict)
{
@@ -173,59 +173,26 @@ class parser
bool accept(const bool strict = true)
{
json_sax_acceptor<BasicJsonType> sax_acceptor;
return sax_parse_impl<false>(&sax_acceptor, strict);
return sax_parse(&sax_acceptor, strict);
}
/*!
@brief public SAX interface
If the SAX parser's parse_error() returns true, the parser recovers from
the error: it repairs the input and continues (see #3989).
@param[in] sax the SAX parser
@param[in] strict whether to expect the last token to be EOF
@return whether the input was parsed without errors and no SAX event
returned false
*/
template<typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse(SAX* sax, const bool strict = true)
{
return sax_parse_impl<true>(sax, strict);
}
private:
/// what sax_parse_internal() does after an object key was expected
enum class next_step : std::uint8_t
{
/// stop parsing
stop,
/// parse a value that begins with last_token
parse_value,
/// evaluate the state of the innermost container, which reads
/// last_token again
evaluate_state
};
template<bool AllowRecovery, typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse_impl(SAX* sax, const bool strict)
{
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
const bool result = sax_parse_internal<AllowRecovery>(sax);
const bool result = sax_parse_internal(sax);
if (result)
{
if (strict)
{
// strict mode: next byte must be EOF; after recovering from an
// error, the end of the input may already have been read
if (last_token != token_type::end_of_input && get_token() != token_type::end_of_input)
// strict mode: next byte must be EOF
if (get_token() != token_type::end_of_input)
{
// the value is complete, so there is nothing to recover
static_cast<void>(report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr),
std::integral_constant<bool, AllowRecovery> {}));
return false;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr));
}
}
else
@@ -236,23 +203,14 @@ class parser
}
}
return result && !error_reported;
return result;
}
/*!
@brief parse a JSON value and pass it to a SAX parser
@tparam AllowRecovery whether to recover from an error if the SAX parser's
parse_error() returns true; false for the SAX parsers
of parse() and accept(), which never do, so that no
code for recovering is generated for them
*/
template<bool AllowRecovery, typename SAX>
private:
template<typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse_internal(SAX* sax)
{
const std::integral_constant<bool, AllowRecovery> allow_recovery{};
// stack to remember the hierarchy of structured values we are parsing
// true = array; false = object
std::vector<bool> states;
@@ -283,18 +241,12 @@ class parser
break;
}
// remember we are now inside an object
states.push_back(false);
// parse key (the steps of parse_key(), which are
// repeated here and below for speed)
// parse key
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
if (!continue_after(key_error(sax, allow_recovery, false), skip_to_state_evaluation))
{
return false;
}
continue;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
@@ -304,13 +256,14 @@ class parser
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
if (!continue_after(key_error(sax, allow_recovery, true), skip_to_state_evaluation))
{
return false;
}
continue;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr));
}
// remember we are now inside an object
states.push_back(false);
// parse values
get_token();
continue;
@@ -346,11 +299,9 @@ class parser
if (JSON_HEDLEY_UNLIKELY(!std::isfinite(res)))
{
if (!overflow_error(sax, res, allow_recovery))
{
return false;
}
break;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
out_of_range::create(406, concat("number overflow parsing '", m_lexer.get_token_string(), '\''), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->number_float(res, m_lexer.get_string())))
@@ -418,63 +369,23 @@ class parser
case token_type::parse_error:
{
// using "uninitialized" to avoid an "expected" message
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::uninitialized, "value"), nullptr), allow_recovery))
{
return false;
}
// recover: keep what can be read of the token
recover_token();
if (last_token != token_type::uninitialized)
{
// a string or a number
continue;
}
if (states.empty())
{
// look for the value after the garbage
if (!skip_to_value())
{
return false;
}
continue;
}
// nothing could be read
if (JSON_HEDLEY_UNLIKELY(!sax->null()))
{
return false;
}
break;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::uninitialized, "value"), nullptr));
}
case token_type::end_of_input:
{
if (JSON_HEDLEY_UNLIKELY(m_lexer.get_position().chars_read_total == 1))
{
// there is nothing to recover
static_cast<void>(report_error(sax, parse_error::create(101, m_lexer.get_position(),
"attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr), allow_recovery));
return false;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(),
"attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
}
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr), allow_recovery))
{
return false;
}
// recover: the input ends where a value is missing
if (states.empty())
{
// there is no value
return false;
}
if (!recover_missing_value(sax, states))
{
return false;
}
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
continue;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr));
}
case token_type::uninitialized:
case token_type::end_array:
@@ -484,35 +395,9 @@ class parser
case token_type::literal_or_value:
default: // the last token was unexpected
{
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr), allow_recovery))
{
return false;
}
// recover
if (states.empty())
{
// look for the value after the garbage
if (!skip_to_value())
{
return false;
}
continue;
}
if (last_token == token_type::name_separator)
{
// a stray ':'; the value may follow
get_token();
continue;
}
if (!recover_missing_value(sax, states))
{
return false;
}
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
continue;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr));
}
}
}
@@ -562,30 +447,9 @@ class parser
continue;
}
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_array, "array"), nullptr), allow_recovery))
{
return false;
}
// recover
if (last_token == token_type::end_of_input)
{
// the input ends inside the array
return close_containers(sax, states);
}
if (last_token == token_type::end_object)
{
// a wrong closing bracket closes the innermost container
if (JSON_HEDLEY_UNLIKELY(!sax->end_array()))
{
return false;
}
states.pop_back();
skip_to_state_evaluation = true;
}
// otherwise, a missing ',' (or a stray ':', which value
// parsing drops): the next value begins here
continue;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_array, "array"), nullptr));
}
// states.back() is false -> object
@@ -602,12 +466,11 @@ class parser
// parse key
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
if (!continue_after(key_error(sax, allow_recovery, false), skip_to_state_evaluation))
{
return false;
}
continue;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
return false;
@@ -616,11 +479,9 @@ class parser
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
if (!continue_after(key_error(sax, allow_recovery, true), skip_to_state_evaluation))
{
return false;
}
continue;
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr));
}
// parse values
@@ -647,479 +508,12 @@ class parser
continue;
}
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_object, "object"), nullptr), allow_recovery))
{
return false;
}
// recover
if (last_token == token_type::end_of_input)
{
// the input ends inside the object
return close_containers(sax, states);
}
if (last_token == token_type::end_array)
{
// a wrong closing bracket closes the innermost container
if (JSON_HEDLEY_UNLIKELY(!sax->end_object()))
{
return false;
}
states.pop_back();
skip_to_state_evaluation = true;
continue;
}
if (!continue_after(recover_member(sax, allow_recovery), skip_to_state_evaluation))
{
return false;
}
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_object, "object"), nullptr));
}
}
/*!
@brief continue sax_parse_internal() after a recovery
@return whether to continue parsing
*/
bool continue_after(const next_step step, bool& skip_to_state_evaluation)
{
if (step == next_step::evaluate_state)
{
// the state evaluation reads the token again
m_lexer.unget_token();
skip_to_state_evaluation = true;
}
return step != next_step::stop;
}
/// the parser for parse() and accept() never recovers: stop parsing
static std::false_type continue_after(std::false_type /*step*/, bool& /*skip_to_state_evaluation*/) noexcept
{
return {};
}
/*!
@brief parse an object key and the name separator (:) after it
last_token is the token where the key is expected. sax_parse_internal()
repeats these steps rather than calling this function, which is used
when recovering from an error.
@return next_step::parse_value if the value follows, with last_token its
first token; next_step::evaluate_state if the object's state is
to be evaluated after recovering from an error; next_step::stop
to stop parsing
*/
template<typename SAX>
next_step parse_key(SAX* sax)
{
const std::true_type allow_recovery{};
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
return key_error(sax, allow_recovery, false);
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
return next_step::stop;
}
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
return key_error(sax, allow_recovery, true);
}
// the value begins with the next token
get_token();
return next_step::parse_value;
}
/*!
@brief report a number that is too large for number_float_t, and recover
from the error by passing the value on; the SAX parser gets the
number's text as well
This is a separate function, as reading other numbers is measurably
slower if the error is handled where they are read.
@param[in] sax the SAX parser
@param[in] value the value that is not finite
@return whether to continue parsing
*/
template<typename SAX, typename AllowRecovery>
bool overflow_error(SAX* sax, const number_float_t value, AllowRecovery allow_recovery)
{
if (!report_error(sax, out_of_range::create(406, concat("number overflow parsing '", m_lexer.get_token_string(), '\''), nullptr), allow_recovery))
{
return false;
}
return sax->number_float(value, m_lexer.get_string());
}
/*!
@brief report a missing key, or a missing name separator (:) after the
key; the parser for parse() and accept() never recovers
@param[in] key_read whether the key was read, so that the name separator
is missing
@return std::false_type, see report_error()
*/
template<typename SAX>
std::false_type key_error(SAX* sax, std::false_type allow_recovery, const bool key_read)
{
return report_error(sax, parse_error::create(101, m_lexer.get_position(), key_read
? exception_message(token_type::name_separator, "object separator")
: exception_message(token_type::value_string, "object key"), nullptr), allow_recovery);
}
/*!
@brief report a missing key, or a missing name separator (:) after the
key, and recover from it
@param[in] key_read whether the key was read, so that the name separator
is missing
*/
template<typename SAX>
next_step key_error(SAX* sax, std::true_type allow_recovery, const bool key_read)
{
if (!key_read)
{
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr), allow_recovery))
{
return next_step::stop;
}
return recover_key(sax);
}
if (!report_error(sax, parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr), allow_recovery))
{
return next_step::stop;
}
return recover_name_separator(sax);
}
/////////////////////
// error recovery
/////////////////////
/*
The functions below repair an error after the SAX parser's parse_error()
returned true (see #3989). Each mistake is repaired by the smallest local
edit: a missing ',' or ':' is inserted, a stray token is removed, what can
be read of an invalid string or number is kept (see
lexer::recover_token()), a missing value becomes null, a wrong closing
bracket closes the innermost container, and the end of the input closes
all of them. The events stay balanced, and every key() is followed by
exactly one value.
A repair hands a token to the state evaluation, by returning it to the
lexer (lexer::unget_token()) so that the state evaluation reads it again,
only if it is ',', ']', '}', or the end of the input. The state evaluation
hands a token to value or key parsing only if it is none of them, so a
token is never handed back and forth. Every other step reads a token or
closes a container, so parsing always ends.
*/
/*!
@brief report an error to the SAX parser; the parser for parse() and
accept() never recovers
@return std::false_type rather than false: its value is known where the
function is called even if the call is not inlined, so the code
for recovering is not generated
*/
template<typename SAX, typename Exception>
std::false_type report_error(SAX* sax, const Exception& ex, std::false_type /*allow_recovery*/)
{
error_reported = true;
static_cast<void>(sax->parse_error(m_lexer.get_position(), m_lexer.get_token_string(), ex));
return {};
}
/*!
@brief report an error to the SAX parser
@return whether to recover from the error
*/
template<typename SAX, typename Exception>
bool report_error(SAX* sax, const Exception& ex, std::true_type /*allow_recovery*/)
{
const std::size_t position = m_lexer.get_position().chars_read_total;
if (error_reported && position == last_error_position && last_token == last_error_token)
{
// a repair handed on the token of the error it repaired; the
// token was reported already, and the SAX parser asked to recover
return true;
}
error_reported = true;
last_error_position = position;
last_error_token = last_token;
if (!sax->parse_error(m_lexer.get_position(), m_lexer.get_token_string(), ex))
{
return false;
}
// the token string of the next error begins here
m_lexer.restart_token_string();
return true;
}
/*!
@brief keep what can be read of the token that the lexer rejected
The error was reported for the rejected token, so it is not reported again
for the token it is repaired to (see lexer::recover_token()).
*/
token_type recover_token()
{
last_token = m_lexer.recover_token();
last_error_position = m_lexer.get_position().chars_read_total;
last_error_token = last_token;
return last_token;
}
/// pass the end events of all open containers
template<typename SAX>
bool close_containers(SAX* sax, std::vector<bool>& states)
{
while (!states.empty())
{
const bool is_array = states.back();
states.pop_back();
if (JSON_HEDLEY_UNLIKELY(is_array ? !sax->end_array() : !sax->end_object()))
{
return false;
}
}
return true;
}
/*!
@brief read tokens until one begins a value, skipping everything before
the top-level value
@return whether a value begins with last_token
*/
bool skip_to_value()
{
while (true)
{
switch (get_token())
{
case token_type::begin_array:
case token_type::begin_object:
case token_type::literal_false:
case token_type::literal_null:
case token_type::literal_true:
case token_type::value_float:
case token_type::value_integer:
case token_type::value_string:
case token_type::value_unsigned:
return true;
case token_type::end_of_input:
return false;
case token_type::parse_error:
recover_token();
if (last_token != token_type::uninitialized)
{
return true;
}
break;
case token_type::uninitialized:
case token_type::end_array:
case token_type::end_object:
case token_type::name_separator:
case token_type::value_separator:
case token_type::literal_or_value:
default:
break;
}
}
}
/*!
@brief skip the rest of an object member that cannot be read
Reads tokens, beginning with last_token, until a ',', '}', or ']' that is
not inside a container that begins in the skipped tokens, or the end of
the input.
*/
void skip_member()
{
std::size_t depth = 0;
while (true)
{
switch (last_token)
{
case token_type::begin_array:
case token_type::begin_object:
++depth;
break;
case token_type::end_array:
case token_type::end_object:
if (depth == 0)
{
return;
}
--depth;
break;
case token_type::value_separator:
if (depth == 0)
{
return;
}
break;
case token_type::end_of_input:
return;
case token_type::parse_error:
recover_token();
break;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::name_separator:
case token_type::literal_or_value:
default:
break;
}
get_token();
}
}
/*!
@brief pass a value where it is missing
last_token is ',', ']', '}', or the end of the input, where a value was
expected. In an object, the key gets null; in an array, a ',' where a
value is missing stands for null (as in JavaScript), while an array that
ends there just ends.
*/
template<typename SAX>
bool recover_missing_value(SAX* sax, const std::vector<bool>& states)
{
JSON_ASSERT(!states.empty());
if (!states.back() || last_token == token_type::value_separator)
{
return sax->null();
}
return true;
}
/// recover from a missing key; last_token is where it was expected
template<typename SAX>
next_step recover_key(SAX* sax)
{
switch (last_token)
{
case token_type::value_separator:
case token_type::end_object:
case token_type::end_array:
case token_type::end_of_input:
// no member: the object's state handles the token
return next_step::evaluate_state;
case token_type::parse_error:
recover_token();
if (last_token == token_type::value_string)
{
// a key that could be repaired
return parse_key(sax);
}
skip_member();
return next_step::evaluate_state;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::begin_array:
case token_type::begin_object:
case token_type::name_separator:
case token_type::literal_or_value:
default:
// a member without a key
skip_member();
return next_step::evaluate_state;
}
}
/// recover from a missing name separator (:) after the key; last_token
/// is where it was expected
template<typename SAX>
next_step recover_name_separator(SAX* sax)
{
switch (last_token)
{
case token_type::value_separator:
case token_type::end_object:
case token_type::end_array:
case token_type::end_of_input:
// the value is missing as well
return sax->null() ? next_step::evaluate_state : next_step::stop;
case token_type::uninitialized:
case token_type::literal_true:
case token_type::literal_false:
case token_type::literal_null:
case token_type::value_string:
case token_type::value_unsigned:
case token_type::value_integer:
case token_type::value_float:
case token_type::begin_array:
case token_type::begin_object:
case token_type::name_separator:
case token_type::parse_error:
case token_type::literal_or_value:
default:
// a missing ':'; the value begins here
return next_step::parse_value;
}
}
/// recover from a token after an object member that is neither ',' nor
/// '}' (nor ']' or the end of the input, which the caller handles)
template<typename SAX>
next_step recover_member(SAX* sax, std::true_type /*allow_recovery*/)
{
if (last_token == token_type::parse_error)
{
recover_token();
}
if (last_token == token_type::value_string)
{
// a missing ','; the next key begins here
return parse_key(sax);
}
skip_member();
return next_step::evaluate_state;
}
/// the parser for parse() and accept() never recovers (and does not come
/// here, as report_error() returned false)
template<typename SAX>
std::false_type recover_member(SAX* /*sax*/, std::false_type /*allow_recovery*/) const noexcept
{
return {};
}
/// get next token from lexer
token_type get_token()
{
@@ -1166,12 +560,6 @@ class parser
const bool allow_exceptions = true;
/// whether trailing commas in objects and arrays should be ignored (true) or signaled as errors (false)
const bool ignore_trailing_commas = false;
/// whether an error was reported to the SAX parser
bool error_reported = false;
/// the position of the last reported error
std::size_t last_error_position = 0;
/// the token of the last reported error
token_type last_error_token = token_type::uninitialized;
};
} // namespace detail
@@ -60,9 +60,11 @@ class iter_impl // NOLINT(cppcoreguidelines-special-member-functions,hicpp-speci
static_assert(is_basic_json<typename std::remove_const<BasicJsonType>::type>::value,
"iter_impl only accepts (const) basic_json");
// superficial check for the LegacyBidirectionalIterator named requirement
static_assert(std::is_base_of<std::bidirectional_iterator_tag, std::bidirectional_iterator_tag>::value
&& std::is_base_of<std::bidirectional_iterator_tag, typename std::iterator_traits<typename array_t::iterator>::iterator_category>::value,
"basic_json iterator assumes array and object type iterators satisfy the LegacyBidirectionalIterator named requirement.");
// note: only array_t::iterator is checked here; object_t::iterator may be
// a forward-only iterator as long as reverse iteration and operator--
// are never used on it
static_assert(std::is_base_of<std::bidirectional_iterator_tag, typename std::iterator_traits<typename array_t::iterator>::iterator_category>::value,
"basic_json iterator assumes array type iterators satisfy the LegacyBidirectionalIterator named requirement.");
public:
/// The std::iterator class template (used as a base class to provide typedefs) is deprecated in C++17.
+95 -121
View File
@@ -26,6 +26,7 @@
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/string_escape.hpp>
#include <nlohmann/detail/string_utils.hpp>
#include <nlohmann/detail/value_t.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
@@ -116,7 +117,7 @@ class json_pointer
/// @sa https://json.nlohmann.me/api/json_pointer/operator_slasheq/
json_pointer& operator/=(std::size_t array_idx)
{
return *this /= std::to_string(array_idx);
return *this /= detail::to_string<string_t>(array_idx);
}
/// @brief create a new JSON pointer by appending the right JSON pointer at the end of the left JSON pointer
@@ -239,6 +240,72 @@ class json_pointer
}
private:
/*!
@brief result of @ref parse_array_index
@ref array_index maps each value to the corresponding parse_error/out_of_range
exception; @ref contains and @ref get_checked_or_null, which must not throw for
an out-of-range or unrepresentable index, switch on it directly instead.
*/
enum class array_index_status
{
ok, ///< @a s is a valid, representable array index
leading_zero, ///< @a s begins with '0' but has more than one character
not_a_number, ///< @a s does not begin with a digit
unresolved, ///< @a s could not be converted to an integer
exceeds_size_type ///< @a s converts to an integer that exceeds size_type
};
/*!
@param[in] s reference token to be converted into an array index
@param[out] idx the integer representation of @a s if @ref array_index_status::ok
is returned; left unchanged otherwise
@return whether @a s is a valid array index, and if not, why
@note this function never throws; @ref array_index and the callers that must not
throw (@ref contains, @ref get_checked_or_null) build on it instead of each
re-implementing the RFC 6901 digit rules and the @a size_type range check
*/
template<typename BasicJsonType>
static array_index_status parse_array_index(const string_t& s, typename BasicJsonType::size_type& idx) noexcept
{
using size_type = typename BasicJsonType::size_type;
// error condition (cf. RFC 6901, Sect. 4)
if (JSON_HEDLEY_UNLIKELY(s.size() > 1 && s[0] == '0'))
{
return array_index_status::leading_zero;
}
// error condition (cf. RFC 6901, Sect. 4)
if (JSON_HEDLEY_UNLIKELY(s.size() > 1 && !(s[0] >= '1' && s[0] <= '9')))
{
return array_index_status::not_a_number;
}
const char* p = s.data();
char* p_end = nullptr; // NOLINT(misc-const-correctness)
errno = 0; // strtoull doesn't reset errno
const unsigned long long res = std::strtoull(p, &p_end, 10); // NOLINT(runtime/int)
if (p == p_end // invalid input or empty string
|| errno == ERANGE // out of range
|| JSON_HEDLEY_UNLIKELY(static_cast<std::size_t>(p_end - p) != s.size())) // incomplete read
{
return array_index_status::unresolved;
}
// only triggered on special platforms (like 32bit), see also
// https://github.com/nlohmann/json/pull/2203
if (res >= static_cast<unsigned long long>((std::numeric_limits<size_type>::max)())) // NOLINT(runtime/int)
{
return array_index_status::exceeds_size_type; // LCOV_EXCL_LINE
}
idx = static_cast<size_type>(res);
return array_index_status::ok;
}
/*!
@param[in] s reference token to be converted into an array index
@@ -252,39 +319,30 @@ class json_pointer
template<typename BasicJsonType>
static typename BasicJsonType::size_type array_index(const string_t& s)
{
using size_type = typename BasicJsonType::size_type;
typename BasicJsonType::size_type idx{};
const auto status = parse_array_index<BasicJsonType>(s, idx);
// error condition (cf. RFC 6901, Sect. 4)
if (JSON_HEDLEY_UNLIKELY(s.size() > 1 && s[0] == '0'))
if (JSON_HEDLEY_UNLIKELY(status == array_index_status::leading_zero))
{
JSON_THROW(detail::parse_error::create(106, 0, detail::concat("array index '", s, "' must not begin with '0'"), nullptr));
}
// error condition (cf. RFC 6901, Sect. 4)
if (JSON_HEDLEY_UNLIKELY(s.size() > 1 && !(s[0] >= '1' && s[0] <= '9')))
if (JSON_HEDLEY_UNLIKELY(status == array_index_status::not_a_number))
{
JSON_THROW(detail::parse_error::create(109, 0, detail::concat("array index '", s, "' is not a number"), nullptr));
}
const char* p = s.data();
char* p_end = nullptr; // NOLINT(misc-const-correctness)
errno = 0; // strtoull doesn't reset errno
const unsigned long long res = std::strtoull(p, &p_end, 10); // NOLINT(runtime/int)
if (p == p_end // invalid input or empty string
|| errno == ERANGE // out of range
|| JSON_HEDLEY_UNLIKELY(static_cast<std::size_t>(p_end - p) != s.size())) // incomplete read
if (JSON_HEDLEY_UNLIKELY(status == array_index_status::unresolved))
{
JSON_THROW(detail::out_of_range::create(404, detail::concat("unresolved reference token '", s, "'"), nullptr));
}
// only triggered on special platforms (like 32bit), see also
// https://github.com/nlohmann/json/pull/2203
if (res >= static_cast<unsigned long long>((std::numeric_limits<size_type>::max)())) // NOLINT(runtime/int)
if (JSON_HEDLEY_UNLIKELY(status == array_index_status::exceeds_size_type))
{
JSON_THROW(detail::out_of_range::create(410, detail::concat("array index ", s, " exceeds size_type"), nullptr)); // LCOV_EXCL_LINE
}
return static_cast<size_type>(res);
return idx;
}
JSON_PRIVATE_UNLESS_TESTED:
@@ -583,63 +641,6 @@ class json_pointer
return *ptr;
}
/*!
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if an array index was not a number
@throw out_of_range.402 if the array index '-' is used
@throw out_of_range.404 if the JSON pointer can not be resolved
*/
template<typename BasicJsonType>
const BasicJsonType& get_checked(const BasicJsonType* ptr) const
{
for (const auto& reference_token : reference_tokens)
{
switch (ptr->type())
{
case detail::value_t::object:
{
// note: at performs range check
ptr = &ptr->at(reference_token);
break;
}
case detail::value_t::array:
{
if (JSON_HEDLEY_UNLIKELY(reference_token == "-"))
{
// "-" always fails the range check
JSON_THROW(detail::out_of_range::create(402, detail::concat(
"array index '-' (", std::to_string(ptr->m_data.m_value.array->size()),
") is out of range"), ptr));
}
const auto idx = array_index<BasicJsonType>(reference_token);
// Bounds check before access to avoid exception with JSON_NOEXCEPTION
if (JSON_HEDLEY_UNLIKELY(idx >= ptr->m_data.m_value.array->size()))
{
JSON_THROW(detail::out_of_range::create(401, detail::concat(
"array index ", std::to_string(idx), " is out of range"), ptr));
}
ptr = &ptr->operator[](idx);
break;
}
case detail::value_t::null:
case detail::value_t::string:
case detail::value_t::boolean:
case detail::value_t::number_integer:
case detail::value_t::number_unsigned:
case detail::value_t::number_float:
case detail::value_t::binary:
case detail::value_t::discarded:
default:
JSON_THROW(detail::out_of_range::create(404, detail::concat("unresolved reference token '", reference_token, "'"), ptr));
}
}
return *ptr;
}
/*!
@brief return a pointer to the pointed to value, or `nullptr` if the
pointer cannot be resolved because a key is missing, an array
@@ -678,18 +679,23 @@ class json_pointer
return nullptr;
}
// may throw parse_error.106/109 for a malformed index; an
// a malformed index still throws parse_error.106/109; an
// index that is syntactically valid but cannot be
// represented (out_of_range.404/410) is treated like an
// out-of-range index below
typename BasicJsonType::size_type idx{};
JSON_TRY
switch (parse_array_index<BasicJsonType>(reference_token, idx))
{
idx = array_index<BasicJsonType>(reference_token);
}
JSON_INTERNAL_CATCH (detail::out_of_range&)
{
return nullptr;
case array_index_status::leading_zero:
JSON_THROW(detail::parse_error::create(106, 0, detail::concat("array index '", reference_token, "' must not begin with '0'"), nullptr));
case array_index_status::not_a_number:
JSON_THROW(detail::parse_error::create(109, 0, detail::concat("array index '", reference_token, "' is not a number"), nullptr));
case array_index_status::unresolved:
case array_index_status::exceeds_size_type:
return nullptr;
case array_index_status::ok:
default:
break;
}
if (JSON_HEDLEY_UNLIKELY(idx >= ptr->m_data.m_value.array->size()))
@@ -717,8 +723,8 @@ class json_pointer
}
/*!
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if an array index was not a number
@note unlike array_index(), this never throws: a malformed or unrepresentable
array index reference token is treated like a missing key (see #5395)
*/
template<typename BasicJsonType>
bool contains(const BasicJsonType* ptr) const
@@ -746,49 +752,17 @@ class json_pointer
// "-" always fails the range check
return false;
}
if (JSON_HEDLEY_UNLIKELY(reference_token.empty()))
{
// an empty reference token is not an array index; array_index()
// would throw out_of_range.404 -- contains() must not throw (see #5395)
return false;
}
if (JSON_HEDLEY_UNLIKELY(reference_token.size() == 1 && !("0" <= reference_token && reference_token <= "9")))
{
// invalid char
return false;
}
if (JSON_HEDLEY_UNLIKELY(reference_token.size() > 1))
{
if (JSON_HEDLEY_UNLIKELY(!('1' <= reference_token[0] && reference_token[0] <= '9')))
{
// the first char should be between '1' and '9'
return false;
}
for (std::size_t i = 1; i < reference_token.size(); i++)
{
if (JSON_HEDLEY_UNLIKELY(!('0' <= reference_token[i] && reference_token[i] <= '9')))
{
// other char should be between '0' and '9'
return false;
}
}
}
// the reference token consists only of digits at this point (cf. checks
// above); however, its numeric value might not be representable, in which
// case array_index() would throw out_of_range.404/410 -- contains() must
// not throw (see #5395), so such a reference token is treated as "not found"
errno = 0; // strtoull() does not reset errno on success
char* p_end = nullptr; // NOLINT(misc-const-correctness)
const unsigned long long magnitude = std::strtoull(reference_token.c_str(), &p_end, 10); // NOLINT(runtime/int)
if (JSON_HEDLEY_UNLIKELY(errno == ERANGE // the value exceeds ULLONG_MAX
|| magnitude >= static_cast<unsigned long long>((std::numeric_limits<typename BasicJsonType::size_type>::max)()))) // NOLINT(runtime/int)
// any parse failure (malformed index, or one that is syntactically
// valid but not representable as size_type) means the reference
// token cannot denote an existing array element -- contains() must
// not throw (see #5395), so it is treated as "not found"
typename BasicJsonType::size_type idx{};
if (JSON_HEDLEY_UNLIKELY(parse_array_index<BasicJsonType>(reference_token, idx) != array_index_status::ok))
{
// the array index cannot be represented as size_type
return false;
}
const auto idx = array_index<BasicJsonType>(reference_token);
if (idx >= ptr->size())
{
// index out of range
+18 -44
View File
@@ -9,7 +9,6 @@
#pragma once
#include <utility> // declval, pair
#include <nlohmann/detail/meta/detected.hpp>
#include <nlohmann/thirdparty/hedley/hedley.hpp>
// This file contains all internal macro definitions (except those affecting ABI)
@@ -140,10 +139,12 @@
// libstdc++ < 11 has incomplete C++20 ranges (issue #4440)
#elif defined(_GLIBCXX_RELEASE) && _GLIBCXX_RELEASE < 11
#define JSON_HAS_RANGES 0
// libc++ < 16 has incomplete C++20 ranges (issue #4440)
// clang < 16 with libstdc++ does not implement the ranges customization
// points libstdc++ declares, so its C++20 ranges support is incomplete (issue #5161)
#elif defined(__clang__) && !defined(__apple_build_version__) \
&& __clang_major__ < 16 && defined(__GLIBCXX__)
#define JSON_HAS_RANGES 0
// libc++ < 16 has incomplete C++20 ranges (issue #4440)
#elif defined(_LIBCPP_VERSION) && _LIBCPP_VERSION < 160000
#define JSON_HAS_RANGES 0
// nvcc CUDA 12.0/12.1 chokes on the enable_borrowed_range variable-template
@@ -158,6 +159,18 @@
#endif
#endif
// std::ranges view conversion (to_json/is_compatible_array_type_impl) additionally
// needs to be disabled on MinGW, whose std::ranges support is incomplete
// (issue #4916); this macro combines both conditions so the check and its
// reason are not duplicated at every use site.
#ifndef JSON_HAS_RANGE_VIEW_CONVERSION
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#define JSON_HAS_RANGE_VIEW_CONVERSION 1
#else
#define JSON_HAS_RANGE_VIEW_CONVERSION 0
#endif
#endif
#ifndef JSON_HAS_STD_FORMAT
#if defined(JSON_HAS_CPP_20) && defined(__cpp_lib_format)
#define JSON_HAS_STD_FORMAT 1
@@ -286,26 +299,11 @@
/*!
@brief function to wrap JSON_THROW_MACRO - there can be compilation errors about
there being no arguments to JSON_THROW that depend on template arguments
if this is not used to call JSON_THROW
*/
template<typename ExceptionType>
void templated_json_throw(ExceptionType exception)
{
JSON_THROW(exception);
/* JSON_THROW(exception) discards exception and aborts - void cast needed to supress
compilation error if compiled with -Werror and Wunused-parameter */
(void)exception;
}
/*!
@brief macro to briefly define a mapping between an enum and JSON with exception
on invalid input
@def NLOHMANN_JSON_SERIALIZE_ENUM_STRICT
@since version 3.12.0
@since version 3.13.0
*/
#define NLOHMANN_JSON_SERIALIZE_ENUM_STRICT(ENUM_TYPE, ...) \
template<typename BasicJsonType> \
@@ -321,7 +319,7 @@ void templated_json_throw(ExceptionType exception)
return ej_pair.first == e; \
}); \
if (it != std::end(m)) j = it->second; \
else templated_json_throw<nlohmann::detail::out_of_range>(nlohmann::detail::out_of_range::create(410,"enum value out of range for " #ENUM_TYPE, nullptr)); \
else ::nlohmann::detail::templated_json_throw<nlohmann::detail::out_of_range>(nlohmann::detail::out_of_range::create(410,"enum value out of range for " #ENUM_TYPE, nullptr)); \
} \
template<typename BasicJsonType> \
inline void from_json(const BasicJsonType& j, ENUM_TYPE& e) \
@@ -336,7 +334,7 @@ void templated_json_throw(ExceptionType exception)
return ej_pair.second == j; \
}); \
if (it != std::end(m)) e = it->first; \
else templated_json_throw<nlohmann::detail::out_of_range>(nlohmann::detail::out_of_range::create(410, nlohmann::detail::concat("enum value out of range for " #ENUM_TYPE ": ", j.dump(-1, ' ', false, nlohmann::detail::error_handler_t::replace)), &j)); \
else ::nlohmann::detail::templated_json_throw<nlohmann::detail::out_of_range>(nlohmann::detail::out_of_range::create(410, nlohmann::detail::concat("enum value out of range for " #ENUM_TYPE ": ", j.dump(-1, ' ', false, nlohmann::detail::error_handler_t::replace)), &j)); \
}
// Ugly macros to avoid uglier copy-paste when specializing basic_json. They
@@ -876,30 +874,6 @@ void templated_json_throw(ExceptionType exception)
\
template<typename... T> \
using result_of_##std_name = decltype(std_name(std::declval<T>()...)); \
} \
\
namespace detail2 { \
struct std_name##_tag \
{ \
}; \
\
template<typename... T> \
std_name##_tag std_name(T&&...); \
\
template<typename... T> \
using result_of_##std_name = decltype(std_name(std::declval<T>()...)); \
\
template<typename... T> \
struct would_call_std_##std_name \
{ \
static constexpr auto const value = ::nlohmann::detail:: \
is_detected_exact<std_name##_tag, result_of_##std_name, T...>::value; \
}; \
} /* namespace detail2 */ \
\
template<typename... T> \
struct would_call_std_##std_name : detail2::would_call_std_##std_name<T...> \
{ \
}
#ifndef JSON_USE_IMPLICIT_CONVERSIONS
@@ -39,6 +39,7 @@
#undef JSON_HAS_EXPERIMENTAL_FILESYSTEM
#undef JSON_HAS_THREE_WAY_COMPARISON
#undef JSON_HAS_RANGES
#undef JSON_HAS_RANGE_VIEW_CONVERSION
#undef JSON_HAS_STD_FORMAT
#undef JSON_HAS_STATIC_RTTI
#undef JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
@@ -12,6 +12,6 @@
NLOHMANN_JSON_NAMESPACE_BEGIN
NLOHMANN_CAN_CALL_STD_FUNC_IMPL(begin);
NLOHMANN_CAN_CALL_STD_FUNC_IMPL(begin)
NLOHMANN_JSON_NAMESPACE_END
@@ -12,6 +12,6 @@
NLOHMANN_JSON_NAMESPACE_BEGIN
NLOHMANN_CAN_CALL_STD_FUNC_IMPL(end);
NLOHMANN_CAN_CALL_STD_FUNC_IMPL(end)
NLOHMANN_JSON_NAMESPACE_END
@@ -62,9 +62,5 @@ using detected_or_t = typename detected_or<Default, Op, Args...>::type;
template<class Expected, template<class...> class Op, class... Args>
using is_detected_exact = std::is_same<Expected, detected_t<Op, Args...>>;
template<class To, template<class...> class Op, class... Args>
using is_detected_convertible =
std::is_convertible<detected_t<Op, Args...>, To>;
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+1 -34
View File
@@ -8,7 +8,7 @@
#pragma once
#include <cstdint> // size_t
#include <cstddef> // size_t
#include <utility> // declval
#include <string> // string
@@ -70,37 +70,6 @@ using parse_error_function_t = decltype(std::declval<T&>().parse_error(
std::declval<std::size_t>(), std::declval<const std::string&>(),
std::declval<const Exception&>()));
template<typename SAX, typename BasicJsonType>
struct is_sax
{
private:
static_assert(is_basic_json<BasicJsonType>::value,
"BasicJsonType must be of type basic_json<...>");
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t;
using string_t = typename BasicJsonType::string_t;
using binary_t = typename BasicJsonType::binary_t;
using exception_t = typename BasicJsonType::exception;
public:
static constexpr bool value =
is_detected_exact<bool, null_function_t, SAX>::value &&
is_detected_exact<bool, boolean_function_t, SAX>::value &&
is_detected_exact<bool, number_integer_function_t, SAX, number_integer_t>::value &&
is_detected_exact<bool, number_unsigned_function_t, SAX, number_unsigned_t>::value &&
is_detected_exact<bool, number_float_function_t, SAX, number_float_t, string_t>::value &&
is_detected_exact<bool, string_function_t, SAX, string_t>::value &&
is_detected_exact<bool, binary_function_t, SAX, binary_t>::value &&
is_detected_exact<bool, start_object_function_t, SAX>::value &&
is_detected_exact<bool, key_function_t, SAX, string_t>::value &&
is_detected_exact<bool, end_object_function_t, SAX>::value &&
is_detected_exact<bool, start_array_function_t, SAX>::value &&
is_detected_exact<bool, end_array_function_t, SAX>::value &&
is_detected_exact<bool, parse_error_function_t, SAX, exception_t>::value;
};
template<typename SAX, typename BasicJsonType>
struct is_sax_static_asserts
{
@@ -120,8 +89,6 @@ struct is_sax_static_asserts
"Missing/invalid function: bool null()");
static_assert(is_detected_exact<bool, boolean_function_t, SAX>::value,
"Missing/invalid function: bool boolean(bool)");
static_assert(is_detected_exact<bool, boolean_function_t, SAX>::value,
"Missing/invalid function: bool boolean(bool)");
static_assert(
is_detected_exact<bool, number_integer_function_t, SAX,
number_integer_t>::value,
-54
View File
@@ -1,54 +0,0 @@
#pragma once
#include <nlohmann/detail/macro_scope.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
#ifdef JSON_HAS_CPP_17
template<bool... Booleans>
struct cxpr_or_impl : std::integral_constant < bool, (Booleans || ...) > {};
template<bool... Booleans>
struct cxpr_and_impl : std::integral_constant < bool, (Booleans &&...) > {};
#else
template<bool... Booleans>
struct cxpr_or_impl : std::false_type {};
template<bool... Booleans>
struct cxpr_or_impl<true, Booleans...> : std::true_type {};
template<bool... Booleans>
struct cxpr_or_impl<false, Booleans...> : cxpr_or_impl<Booleans...> {};
template<bool... Booleans>
struct cxpr_and_impl : std::true_type {};
template<bool... Booleans>
struct cxpr_and_impl<true, Booleans...> : cxpr_and_impl<Booleans...> {};
template<bool... Booleans>
struct cxpr_and_impl<false, Booleans...> : std::false_type {};
#endif
template<class Boolean>
struct cxpr_not : std::integral_constant < bool, !Boolean::value > {};
template<class... Booleans>
struct cxpr_or : cxpr_or_impl<Booleans::value...> {};
template<bool... Booleans>
struct cxpr_or_c : cxpr_or_impl<Booleans...> {};
template<class... Booleans>
struct cxpr_and : cxpr_and_impl<Booleans::value...> {};
template<bool... Booleans>
struct cxpr_and_c : cxpr_and_impl<Booleans...> {};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+12 -23
View File
@@ -283,6 +283,13 @@ template<class B, class... Bn>
struct conjunction<B, Bn...>
: std::conditional<static_cast<bool>(B::value), conjunction<Bn...>, B>::type {};
// https://en.cppreference.com/w/cpp/types/disjunction
template<class...> struct disjunction : std::false_type { };
template<class B> struct disjunction<B> : B { };
template<class B, class... Bn>
struct disjunction<B, Bn...>
: std::conditional<static_cast<bool>(B::value), B, disjunction<Bn...>>::type {};
// https://en.cppreference.com/w/cpp/types/negation
template<class B> struct negation : std::integral_constant < bool, !B::value > { };
@@ -477,9 +484,7 @@ template<typename T> struct is_range_view_optional_type<std::optional<T>> : std:
template<typename T> struct is_range_view_optional_type : std::false_type {};
#endif
// std::ranges does not work properly on MinGW due to incomplete C++20 support
// see https://github.com/nlohmann/json/issues/4916
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#if JSON_HAS_RANGE_VIEW_CONVERSION
// SafeToCheck guards against types that trigger circular constraints when
// std::ranges::view<T> is evaluated on GCC 12 / libstdc++ 12:
@@ -518,7 +523,7 @@ struct is_compatible_array_type_impl <
// filter_view) can match BOTH this iterator-based specialization AND the view-based one
// below, causing ambiguity. Exclude views here so the two specializations are mutually
// exclusive: this one handles plain iterable containers, the other handles views.
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#if JSON_HAS_RANGE_VIEW_CONVERSION
&& !is_compatible_range_view<CompatibleArrayType>::value
#endif
>>
@@ -528,7 +533,7 @@ struct is_compatible_array_type_impl <
range_value_t<CompatibleArrayType>>::value;
};
#if JSON_HAS_RANGES && !defined(__MINGW32__)
#if JSON_HAS_RANGE_VIEW_CONVERSION
template<typename BasicJsonType, typename CompatibleArrayType>
struct is_compatible_array_type_impl <
BasicJsonType, CompatibleArrayType,
@@ -604,7 +609,6 @@ struct is_compatible_integer_type_impl <
std::is_integral<CompatibleNumberIntegerType>::value&&
!std::is_same<bool, CompatibleNumberIntegerType>::value >>
{
// is there an assert somewhere on overflows?
using RealLimits = std::numeric_limits<RealIntegerType>;
using CompatibleLimits = std::numeric_limits<CompatibleNumberIntegerType>;
@@ -798,20 +802,7 @@ struct has_capacity : std::integral_constant<bool, is_detected<detect_capacity,
// a naive helper to check if a type is an ordered_map (exploits the fact that
// ordered_map inherits capacity() from std::vector)
template <typename T>
struct is_ordered_map
{
using one = char;
struct two
{
char x[2]; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
};
template <typename C> static one test( decltype(&C::capacity) ) ;
template <typename C> static two test(...);
enum { value = sizeof(test<T>(nullptr)) == sizeof(char) }; // NOLINT(cppcoreguidelines-pro-type-vararg,hicpp-vararg,cppcoreguidelines-use-enum-class)
};
struct is_ordered_map : has_capacity<T> {};
// to avoid useless casts (see https://github.com/nlohmann/json/issues/2893#issuecomment-889152324)
template < typename T, typename U, enable_if_t < !std::is_same<T, U>::value, int > = 0 >
@@ -835,10 +826,8 @@ using all_signed = conjunction<std::is_signed<Types>...>;
template<typename... Types>
using all_unsigned = conjunction<std::is_unsigned<Types>...>;
// there's a disjunction trait in another PR; replace when merged
template<typename... Types>
using same_sign = std::integral_constant < bool,
all_signed<Types...>::value || all_unsigned<Types...>::value >;
using same_sign = disjunction<all_signed<Types...>, all_unsigned<Types...>>;
template<typename OfType, typename T>
using never_out_of_range = std::integral_constant < bool,
-74
View File
@@ -12,7 +12,6 @@
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint32_t
#include <string> // string, to_string
#include <utility> // move
#include <nlohmann/detail/abi_macros.hpp>
#include <nlohmann/detail/macro_scope.hpp>
@@ -134,78 +133,5 @@ inline bool is_valid_utf8(const StringType& s, const std::size_t first = 0) noex
return state == UTF8_ACCEPT;
}
/*!
@brief append U+FFFD REPLACEMENT CHARACTER, encoded in UTF-8
@param[in,out] s the string to append to
*/
template<typename StringType>
inline void append_replacement_character(StringType& s)
{
s.push_back(static_cast<typename StringType::value_type>(0xEFu));
s.push_back(static_cast<typename StringType::value_type>(0xBFu));
s.push_back(static_cast<typename StringType::value_type>(0xBDu));
}
/*!
@brief replace ill-formed UTF-8 with U+FFFD REPLACEMENT CHARACTER
Each maximal subpart of an ill-formed sequence becomes one U+FFFD, as the
Unicode Standard recommends (Section 3.9, "U+FFFD Substitution of Maximal
Subparts"), and as the parser for JSON text does when it recovers from errors.
@param[in,out] s the string to repair
@param[in] first index of the first byte to repair; the bytes before it are
assumed to be valid UTF-8 that ends on a code point boundary
*/
template<typename StringType>
inline void replace_invalid_utf8(StringType& s, const std::size_t first = 0)
{
StringType result = s;
result.resize(first);
std::uint8_t state = UTF8_ACCEPT;
std::uint32_t codepoint = 0;
// the first byte of the sequence being decoded
std::size_t sequence_start = first;
std::size_t i = first;
while (i < s.size())
{
switch (decode(state, codepoint, static_cast<std::uint8_t>(s[i])))
{
case UTF8_ACCEPT:
for (++i; sequence_start < i; ++sequence_start)
{
result.push_back(s[sequence_start]);
}
break;
case UTF8_REJECT:
append_replacement_character(result);
// the byte that made the sequence ill-formed begins the next
// one, unless it began this one
if (i == sequence_start)
{
++i;
}
state = UTF8_ACCEPT;
sequence_start = i;
break;
default: // in the middle of a sequence
++i;
break;
}
}
// a sequence that the string ends in the middle of
if (state != UTF8_ACCEPT)
{
append_replacement_character(result);
}
s = std::move(result);
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+8 -5
View File
@@ -145,7 +145,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
friend class ::nlohmann::detail::iter_impl;
template<typename BasicJsonType, typename CharType, typename OutputSinkType>
friend class ::nlohmann::detail::binary_writer;
template<typename BasicJsonType, typename InputType, typename SAX, bool AllowRecovery>
template<typename BasicJsonType, typename InputType, typename SAX>
friend class ::nlohmann::detail::binary_reader;
template<typename BasicJsonType, typename InputAdapterType>
friend class ::nlohmann::detail::json_sax_dom_parser;
@@ -5034,7 +5034,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::forward<InputType>(i));
return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(format, sax, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(format, sax, strict);
}
/// @brief generate SAX events (iterator pair, or iterator+sentinel pair for C++20 ranges support)
@@ -5051,7 +5051,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
auto ia = detail::input_adapter(std::move(first), std::move(last));
return format == input_format_t::json
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(format, sax, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(format, sax, strict);
}
/// @brief generate SAX events
@@ -5073,7 +5073,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
? parser(std::move(ia), nullptr, true, ignore_comments, ignore_trailing_commas).sax_parse(sax, strict)
// NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg)
: detail::binary_reader<basic_json, decltype(ia), SAX, true>(std::move(ia), format).sax_parse(format, sax, strict);
: detail::binary_reader<basic_json, decltype(ia), SAX>(std::move(ia), format).sax_parse(format, sax, strict);
}
#ifndef JSON_NO_IO
/// @brief deserialize from stream
@@ -5092,7 +5092,10 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/// @sa https://json.nlohmann.me/api/basic_json/operator_gtgt/
friend std::istream& operator>>(std::istream& i, basic_json& j)
{
parser(detail::input_adapter(i)).parse(false, j);
// parse into a temporary so that j is left unchanged if parsing fails
basic_json result;
parser(detail::input_adapter(i)).parse(false, result);
j = std::move(result);
return i;
}
#endif // JSON_NO_IO
File diff suppressed because it is too large Load Diff
-1
View File
@@ -47,7 +47,6 @@ inline namespace json_literals
namespace detail
{
using NLOHMANN_JSON_NAMESPACE::detail::json_sax_dom_callback_parser;
using NLOHMANN_JSON_NAMESPACE::detail::json_sax_dom_parser;
using NLOHMANN_JSON_NAMESPACE::detail::unknown_size;
} // namespace detail
-12
View File
@@ -45,10 +45,6 @@ dumps is stable under exactly the same values that break operator==.
The unit tests run the same checks on a fixed corpus (see the "BJData round-trip
invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bjdata() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -63,8 +59,6 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// value-stable comparison for the round-trip checks below; see the note
@@ -77,15 +71,11 @@ static bool is_value_stable(const json& lhs, const json& rhs)
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bjdata).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bjdata(vec1);
assert(recovered_without_errors);
try
{
@@ -119,7 +109,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -128,7 +117,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
-12
View File
@@ -19,10 +19,6 @@ It also checks that reading the data from a stream, which reads strings byte by
byte, gives the same value or error as reading it from contiguous memory, which
copies strings in bulk.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bon8() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -37,8 +33,6 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
namespace
@@ -62,9 +56,6 @@ std::string read_bon8(InputType&& input)
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bon8).errors == 0;
// contiguous and stream input must be read alike
{
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
@@ -76,7 +67,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bon8(vec1);
assert(recovered_without_errors);
try
{
@@ -98,7 +88,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -107,7 +96,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
-12
View File
@@ -15,10 +15,6 @@ array data, it performs the following steps:
- j2 = from_bson(vec)
- assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_bson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -33,22 +29,16 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::bson).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_bson(vec1);
assert(recovered_without_errors);
if (j1.is_discarded())
{
@@ -75,7 +65,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -84,7 +73,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors can occur during parsing, too
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
-12
View File
@@ -15,10 +15,6 @@ array data, it performs the following steps:
- j2 = from_cbor(vec)
- assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_cbor() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -33,22 +29,16 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::cbor).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_cbor(vec1);
assert(recovered_without_errors);
try
{
@@ -70,7 +60,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -79,7 +68,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors can occur during parsing, too
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
-14
View File
@@ -16,10 +16,6 @@ array data, it performs the following steps:
- s2 = serialize(j2)
- assert(s1 == s2)
Furthermore, it parses data with a SAX parser that recovers from every error
and checks that the events are balanced, that parsing ends, and that valid
input is parsed without errors (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -27,7 +23,6 @@ drivers.
#include <cassert>
#include <iostream>
#include <sstream>
#include <string>
#include <nlohmann/json.hpp>
// the round-trip checks below are assertions; NDEBUG would compile them away
@@ -35,20 +30,11 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
{
const auto checker = check_recovering_parse(data, size, json::input_format_t::json);
assert(checker.events <= (4 * size) + 4);
assert((checker.errors == 0) == json::accept(data, data + size));
}
try
{
// step 1: parse input
-12
View File
@@ -15,10 +15,6 @@ array data, it performs the following steps:
- j2 = from_msgpack(vec)
- assert(j1 == j2)
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_msgpack() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -33,22 +29,16 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::msgpack).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_msgpack(vec1);
assert(recovered_without_errors);
try
{
@@ -70,7 +60,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -79,7 +68,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
-12
View File
@@ -24,10 +24,6 @@ array data, it performs the following steps:
The unit tests run the same checks on a fixed corpus (see the "UBJSON round-trip
invariants" test case), so keep both in sync.
Furthermore, it reads data with a SAX parser that recovers from every error
and checks that the events are balanced, that reading ends, and that it
reports an error exactly when from_ubjson() fails (see #3989).
The provided function `LLVMFuzzerTestOneInput` can be used in different fuzzer
drivers.
*/
@@ -42,22 +38,16 @@ drivers.
#error "the fuzzer drivers must be built without NDEBUG"
#endif
#include "fuzzer-recovering_checker.hpp"
using json = nlohmann::json;
// see http://llvm.org/docs/LibFuzzer.html
extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
{
// step 0: recover from all errors, reading from memory and from a stream
const bool recovered_without_errors = check_recovering_parse(data, size, json::input_format_t::ubjson).errors == 0;
try
{
// step 1: parse input
std::vector<uint8_t> const vec1(data, data + size);
json const j1 = json::from_ubjson(vec1);
assert(recovered_without_errors);
try
{
@@ -89,7 +79,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::parse_error&)
{
// parse errors are ok, because input may be random bytes
assert(!recovered_without_errors);
}
catch (const json::type_error&)
{
@@ -98,7 +87,6 @@ extern "C" int LLVMFuzzerTestOneInput(const uint8_t* data, size_t size)
catch (const json::out_of_range&)
{
// out of range errors may happen if provided sizes are excessive
assert(!recovered_without_errors);
}
// return 0 - non-zero return values are reserved for future use
-154
View File
@@ -1,154 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cassert>
#include <cstddef>
#include <cstdint>
#include <sstream>
#include <string>
#include <vector>
#include <nlohmann/json.hpp>
namespace
{
// a SAX parser that recovers from every error and checks that the events are
// balanced and that every key is followed by exactly one value
class recovering_checker : public nlohmann::json_sax<nlohmann::json>
{
public:
bool null() override
{
return value();
}
bool boolean(bool /*val*/) override
{
return value();
}
bool number_integer(number_integer_t /*val*/) override
{
return value();
}
bool number_unsigned(number_unsigned_t /*val*/) override
{
return value();
}
bool number_float(number_float_t /*val*/, const string_t& /*s*/) override
{
return value();
}
bool string(string_t& /*val*/) override
{
return value();
}
bool binary(binary_t& /*val*/) override
{
return value();
}
bool start_object(std::size_t /*elements*/) override
{
value();
stack.push_back('o');
return true;
}
bool key(string_t& /*val*/) override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.back() = 'v';
return true;
}
bool end_object() override
{
++events;
assert(!stack.empty() && stack.back() == 'o');
stack.pop_back();
return true;
}
bool start_array(std::size_t /*elements*/) override
{
value();
stack.push_back('a');
return true;
}
bool end_array() override
{
++events;
assert(!stack.empty() && stack.back() == 'a');
stack.pop_back();
return true;
}
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const nlohmann::detail::exception& /*ex*/) override
{
++errors;
return true;
}
bool complete() const
{
return stack.empty();
}
std::size_t events = 0;
std::size_t errors = 0;
private:
bool value()
{
++events;
if (!stack.empty())
{
// an array element, or the value of a key
assert(stack.back() != 'o');
if (stack.back() == 'v')
{
stack.back() = 'o';
}
}
return true;
}
// 'a' for an array, 'o' for an object that expects a key, 'v' for an
// object that expects the value of a key
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
};
/// parses @a data with a recovering_checker from memory and from a stream,
/// checks that both see the same, that the events are balanced, and that the
/// number of errors is bounded, and returns the checker (see #3989)
inline recovering_checker check_recovering_parse(const std::uint8_t* data, const std::size_t size, const nlohmann::json::input_format_t format)
{
recovering_checker checker;
const bool ok = nlohmann::json::sax_parse(data, data + size, &checker, format);
assert(checker.complete());
assert(checker.errors <= size + 1);
assert(ok == (checker.errors == 0));
std::istringstream stream(std::string(reinterpret_cast<const char*>(data), size));
recovering_checker stream_checker;
assert(nlohmann::json::sax_parse(stream, &stream_checker, format) == ok);
assert(stream_checker.complete());
assert(stream_checker.events == checker.events);
assert(stream_checker.errors == checker.errors);
return checker;
}
} // namespace
+35 -40
View File
@@ -352,6 +352,41 @@ TEST_CASE("alternative string type")
CHECK(j2.flatten().unflatten() == j2);
}
SECTION("contains(json_pointer)")
{
// contains(json_pointer) must compile and work with a string_t that has
// no c_str() and no comparison with const char* (see #5666)
auto j = alt_json::parse(R"({"foo": ["bar", "baz"]})");
// present: object key and array indices
CHECK(j.contains(alt_json::json_pointer("/foo")));
CHECK(j.contains(alt_json::json_pointer("/foo/0")));
CHECK(j.contains(alt_json::json_pointer("/foo/1")));
// missing: absent object key and out-of-range array index
CHECK_FALSE(j.contains(alt_json::json_pointer("/bar")));
CHECK_FALSE(j.contains(alt_json::json_pointer("/foo/2")));
// "-" always fails the range check
CHECK_FALSE(j.contains(alt_json::json_pointer("/foo/-")));
// an array index must not have a leading zero
CHECK_FALSE(j.contains(alt_json::json_pointer("/foo/01")));
// a reference token that is not a number
CHECK_FALSE(j.contains(alt_json::json_pointer("/foo/bar")));
}
SECTION("operator/(std::size_t)")
{
// json_pointer::operator/=(std::size_t) must compile without string_t
// being constructible from std::string (see #5666)
auto j = alt_json::parse(R"({"foo": ["bar", "baz"]})");
CHECK(j.at(alt_json::json_pointer("/foo") / std::size_t(0)) == j["foo"][0]);
CHECK(j.at(alt_json::json_pointer("/foo") / std::size_t(1)) == j["foo"][1]);
}
SECTION("patch")
{
alt_json const patch1 = alt_json::parse(R"([{ "op": "add", "path": "/a/b", "value": [ "foo", "bar" ] }])");
@@ -384,46 +419,6 @@ TEST_CASE("alternative string type")
CHECK(j2.dump() == R"({"/foo/0":"bar","/foo/1":"baz"})");
}
SECTION("error recovery")
{
// a SAX parser that recovers from every error (see #3989)
struct recovering_parser : nlohmann::detail::json_sax_dom_parser<alt_json>
{
explicit recovering_parser(alt_json& j)
: nlohmann::detail::json_sax_dom_parser<alt_json>(j, false)
{}
// sax_parse() calls the SAX parser's own parse_error(), so hiding
// the one of the base class is what recovering takes
// NOLINTNEXTLINE(bugprone-derived-method-shadowing-base-method)
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const nlohmann::detail::exception& /*unused*/)
{
++errors;
return true;
}
std::size_t errors = 0;
};
alt_json j;
recovering_parser sax(j);
// not inside CHECK(): MSVC reads the escape in a stringized raw string
const std::string input = R"([1., "a\qb", tru, {"k" 2}])";
CHECK(!alt_json::sax_parse(input, &sax));
CHECK(sax.errors == 4);
CHECK(j.dump() == R"([1,"aqb",null,{"k":2}])");
// a UBJSON high-precision number, a CBOR key that is not a string
alt_json u;
recovering_parser ubjson_sax(u);
CHECK(!alt_json::sax_parse(std::vector<std::uint8_t> {'[', 'H', 'i', 2, '1', '.', ']'}, &ubjson_sax, alt_json::input_format_t::ubjson));
CHECK(u.dump() == "[1]");
alt_json c;
recovering_parser cbor_sax(c);
CHECK(!alt_json::sax_parse(std::vector<std::uint8_t> {0xA2, 0x01, 0x02, 0x61, 'a', 0x03}, &cbor_sax, alt_json::input_format_t::cbor));
CHECK(c.dump() == R"({"a":3})");
}
SECTION("strict enum")
{
// regression test for #5667: NLOHMANN_JSON_SERIALIZE_ENUM_STRICT's from_json
+1 -583
View File
@@ -143,13 +143,11 @@ class SaxEventLogger
{
errored = true;
events.push_back("parse_error(" + std::to_string(position) + ")");
return recover;
return false;
}
std::vector<std::string> events {}; // NOLINT(readability-redundant-member-init)
bool errored = false;
/// whether parse_error() asks the parser to recover from the error (see #3989)
bool recover = false;
};
class SaxCountdown : public nlohmann::json::json_sax_t
@@ -2939,583 +2937,3 @@ TEST_CASE("diagnostic positions: value lifetime, input adapters, and SAX")
}
}
#endif
namespace
{
/// builds a value like json::parse(), but asks the parser to recover from
/// errors (see #3989), and checks that the events it receives are balanced
class RecoveringDomParser
{
public:
explicit RecoveringDomParser(json& j, std::size_t max_errors_ = static_cast<std::size_t>(-1))
: dom(j, false)
, max_errors(max_errors_)
{}
bool null()
{
value();
return dom.null();
}
bool boolean(bool val)
{
value();
return dom.boolean(val);
}
bool number_integer(json::number_integer_t val)
{
value();
return dom.number_integer(val);
}
bool number_unsigned(json::number_unsigned_t val)
{
value();
return dom.number_unsigned(val);
}
bool number_float(json::number_float_t val, const std::string& s)
{
value();
return dom.number_float(val, s);
}
bool string(std::string& val)
{
value();
return dom.string(val);
}
bool binary(json::binary_t& val)
{
value();
return dom.binary(val);
}
bool start_object(std::size_t elements)
{
value();
stack.push_back('o');
return dom.start_object(elements);
}
bool key(std::string& val)
{
++events;
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.back() = 'v';
return dom.key(val);
}
bool end_object()
{
++events;
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_object();
}
bool start_array(std::size_t elements)
{
value();
stack.push_back('a');
return dom.start_array(elements);
}
bool end_array()
{
++events;
if (stack.empty() || stack.back() != 'a')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_array();
}
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& ex)
{
errors.emplace_back(ex.what());
return errors.size() < max_errors;
}
/// whether the events were balanced and every key was followed by a value
bool balanced() const
{
return well_formed && stack.empty();
}
/// builds the value
nlohmann::detail::json_sax_dom_parser<json> dom;
std::vector<std::string> errors {}; // NOLINT(readability-redundant-member-init)
std::size_t events = 0;
/// the open containers: 'a' for an array, 'o' for an object that expects
/// a key, 'v' for an object that expects the value of a key
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
bool well_formed = true;
std::size_t max_errors;
private:
/// a value is passed: it is an array element, or the value of a key
void value()
{
++events;
if (!stack.empty())
{
if (stack.back() == 'v')
{
stack.back() = 'o';
}
else if (stack.back() == 'o')
{
// a value without a key
well_formed = false;
}
}
}
};
struct RecoveryResult
{
json value;
std::vector<std::string> errors;
std::size_t events;
bool ok;
bool balanced;
};
template<typename InputType>
RecoveryResult parse_recovering(InputType&& input, const bool strict = true,
const bool ignore_comments = false, const bool ignore_trailing_commas = false)
{
json j;
RecoveringDomParser sax(j);
const bool ok = json::sax_parse(std::forward<InputType>(input), &sax, json::input_format_t::json,
strict, ignore_comments, ignore_trailing_commas);
return {j, sax.errors, sax.events, ok, sax.balanced()};
}
/// stops after a number of events, but recovers from errors
class RecoveringCountdown : public SaxCountdown
{
public:
using SaxCountdown::SaxCountdown;
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const json::exception& /*ex*/) override
{
return true;
}
};
/// a repaired input: the value it is repaired to, and the number of errors
struct Repair
{
const char* input;
const char* expected;
std::size_t errors;
};
} // namespace
TEST_CASE("parser error recovery (#3989)")
{
SECTION("repairs")
{
const std::vector<Repair> repairs =
{
// a missing separator is inserted
{"[1 2]", "[1,2]", 1},
{R"({"a":1 "b":2})", R"({"a":1,"b":2})", 1},
{R"({"a" 1})", R"({"a":1})", 1},
{"[1 tru 2]", "[1,null,2]", 2},
{R"({"a" "b": 1})", R"({"a":"b"})", 2},
// a missing value is null in an object; in an array, a ',' stands
// for null, while an array that ends there just ends
{R"({"a":})", R"({"a":null})", 1},
{R"({"a"})", R"({"a":null})", 1},
{R"({"a","b":1})", R"({"a":null,"b":1})", 1},
{"[1,,2]", "[1,null,2]", 1},
{"[,1]", "[null,1]", 1},
{"[1,]", "[1]", 1},
{"[1,2,3,]", "[1,2,3]", 1},
{R"({"a":1,})", R"({"a":1})", 1},
// a broken string keeps what can be read
{R"(["a\qb"])", R"(["aqb"])", 1},
{R"({"na\me":1})", R"({"name":1})", 1},
{"[\"\xFF\"]", R"(["\uFFFD"])", 1},
{"[\"a\xC3(\"]", R"(["a\uFFFD("])", 1},
{"[\"\xE2\x82\"]", R"(["\uFFFD"])", 1},
{"[\"\xC3\\\\\", 1]", R"(["\uFFFD\\",1])", 1},
{R"(["\u12"])", R"(["\uFFFD"])", 1},
{R"(["\u12G4"])", R"(["\uFFFDG4"])", 1},
{R"(["\uDC00x"])", R"(["\uFFFDx"])", 1},
{R"(["\uD800x"])", R"(["\uFFFDx"])", 1},
{R"(["\uD800\u0041"])", R"(["\uFFFDA"])", 1},
{R"(["\uD800\uD800\uDC00"])", R"(["\uFFFD\uD800\uDC00"])", 1},
{R"(["\uD800\uD800\uD800x"])", R"(["\uFFFD\uFFFD\uFFFDx"])", 1},
{
R"(["\uD800\"x", 1])", R"(["\uFFFD\"x",1])", 1
},
{R"(["\uD800\q"])", R"(["\uFFFDq"])", 1},
{"[\"a\tb\"]", R"(["a\tb"])", 1},
{R"(["a\qb\u0041\x"])", R"(["aqbAx"])", 1},
// a broken number keeps its longest valid prefix
{"[1.]", "[1]", 1},
{"[-2.]", "[-2]", 1},
{"[1.5e]", "[1.5]", 1},
{"[1e+]", "[1]", 1},
{"[1.x2, 3]", "[1,3]", 1},
// what cannot be read at all is null
{"[1,NaN,3]", "[1,null,3]", 1},
{"[tru]", "[null]", 1},
{"[-]", "[null]", 1},
{R"({"a":Infinity})", R"({"a":null})", 1},
// a stray token is dropped
{"[:1]", "[1]", 1},
{R"(["a":1])", R"(["a",1])", 1},
{R"({"a"::1})", R"({"a":1})", 1},
// a member that cannot be read is skipped
{R"({1:2,"b":3})", R"({"b":3})", 1},
{R"({"a":1 2})", R"({"a":1})", 1},
{R"({,"a":1})", R"({"a":1})", 1},
{R"({"a":1,,"b":2})", R"({"a":1,"b":2})", 1},
{"{a:1}", "{}", 1},
{R"({"a":1 [1,{"b":2}], "c":3})", R"({"a":1,"c":3})", 1},
{R"([{1}, "a"])", R"([{},"a"])", 1},
// a wrong closing bracket closes the innermost container
{R"({"a":[1,2}, "b":3})", R"({"a":[1,2],"b":3})", 1},
{R"([{"a":1], 2])", R"([{"a":1},2])", 1},
{"{]", "{}", 1},
{"[}", "[]", 1},
// the end of the input closes all containers
{R"({"a":[1,2)", R"({"a":[1,2]})", 1},
{"[", "[]", 1},
{"{", "{}", 1},
{R"({"a")", R"({"a":null})", 1},
{R"({"a":)", R"({"a":null})", 1},
{"[1,", "[1]", 1},
{"[[[1", "[[[1]]]", 1},
{
R"(["abc)", R"(["abc"])", 2
},
{"[1,tr", "[1,null]", 2},
{"\"abc", "\"abc\"", 1},
{"[\"ab\ncd\"]", R"(["ab",null,"]"])", 4},
// what comes before the top-level value is skipped
{")]}'\n{\"a\":1}", R"({"a":1})", 1},
{R"(data: {"a":1})", R"({"a":1})", 1},
{"\xEF\xBB[1]", "[1]", 1},
// what comes after it is an error that ends parsing
{R"({"a":1}})", R"({"a":1})", 1},
{"[1}]", "[1]", 2},
{"[1] [2]", "[1]", 1},
};
for (const auto& repair : repairs)
{
CAPTURE(repair.input);
const auto result = parse_recovering(std::string(repair.input));
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.value == json::parse(repair.expected));
CHECK(result.errors.size() == repair.errors);
}
}
SECTION("number overflow")
{
const auto result = parse_recovering(std::string("[1e999,-1e999]"));
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.errors.size() == 2);
CHECK(result.errors[0] == "[json.exception.out_of_range.406] number overflow parsing '1e999'");
REQUIRE(result.value.size() == 2);
CHECK(result.value[0].is_number_float());
CHECK(result.value[0].get<double>() == std::numeric_limits<double>::infinity());
CHECK(result.value[1].get<double>() == -std::numeric_limits<double>::infinity());
// the SAX parser gets the number's text
SaxEventLogger logger;
logger.recover = true;
CHECK(!json::sax_parse("1e999", &logger));
CHECK(logger.events == std::vector<std::string>({"parse_error(5)", "number_float(1e999)"}));
}
SECTION("nothing to recover")
{
for (const std::string s :
{
"", " ", "]", "tru", "NaN", ",:", "/* comment"
})
{
CAPTURE(s);
const auto result = parse_recovering(s, true, true);
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.events == 0);
CHECK(result.value == nullptr);
CHECK(result.errors.size() == 1);
}
}
SECTION("error messages")
{
// the first error is reported as without recovery
for (const std::string s :
{
"[1 2]", R"({"a":1 "b":2})", R"({"a" 1})", R"({"a":})", "[1,]", "[1.]",
R"(["a\qb"])", "[1e999]", "{1:2}", R"({"a":[1,2}})", "[1,", "[1] [2]", "{a:1}"
})
{
CAPTURE(s);
const auto result = parse_recovering(s);
REQUIRE(!result.errors.empty());
json _;
CHECK_THROWS_WITH_STD_STR(_ = json::parse(s), result.errors.front());
}
// the token of an error begins where the previous error was
const auto result = parse_recovering(std::string("[tru, fals, nul]"));
CHECK(result.errors == std::vector<std::string>(
{
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: '[tru,'",
"[json.exception.parse_error.101] parse error at line 1, column 11: syntax error while parsing value - invalid literal; last read: ', fals,'",
"[json.exception.parse_error.101] parse error at line 1, column 16: syntax error while parsing value - invalid literal; last read: ', nul]'"
}));
CHECK(result.value == json::parse("[null,null,null]"));
}
SECTION("events")
{
// see #4522
SaxEventLogger logger;
logger.recover = true;
CHECK(!json::sax_parse(R"([{1}, "a"])", &logger));
CHECK(logger.events == std::vector<std::string>(
{
"start_array()", "start_object()", "parse_error(3)", "end_object()", "string(a)", "end_array()"
}));
}
SECTION("options")
{
SECTION("strict")
{
const auto result = parse_recovering(std::string("[1 2] [3]"), false);
CHECK(!result.ok);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
}
SECTION("ignore_trailing_commas")
{
for (const std::string s :
{
"[1,]", R"({"a":1,})", "[[1,],]"
})
{
CAPTURE(s);
const auto result = parse_recovering(s, true, false, true);
CHECK(result.ok);
CHECK(result.errors.empty());
}
auto result = parse_recovering(std::string("[1,,]"), true, false, true);
CHECK(result.value == json::parse("[1,null]"));
CHECK(result.errors.size() == 1);
result = parse_recovering(std::string(R"({"a":1,,})"), true, false, true);
CHECK(result.value == json::parse(R"({"a":1})"));
CHECK(result.errors.size() == 1);
}
SECTION("ignore_comments")
{
auto result = parse_recovering(std::string("[1 /* one */ 2]"), true, true);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
// a comment that is not closed runs to the end of the input, which
// is not reported again
result = parse_recovering(std::string("[1, 2 /* unterminated"), true, true);
CHECK(result.balanced);
CHECK(result.value == json::parse("[1,2]"));
CHECK(result.errors.size() == 1);
// a '/' that does not begin a comment is garbage
result = parse_recovering(std::string("[1, /x, 2]"), true, true);
CHECK(result.balanced);
CHECK(result.value == json::parse("[1,null,2]"));
CHECK(result.errors.size() == 1);
}
}
SECTION("null bytes")
{
// a null byte ends the input, unless JSON_STRICT_NUL_HANDLING is set
const auto result = parse_recovering(std::string("[1,\0x", 5));
CHECK(result.balanced);
CHECK(!result.ok);
#ifdef JSON_TEST_STRICT_NUL_HANDLING_ENABLED
CHECK(result.value == json::parse("[1,null]"));
#else
CHECK(result.value == json::parse("[1]"));
CHECK(result.errors.size() == 1);
#endif
const auto in_string = parse_recovering(std::string("[\"a\0b\"]", 7));
CHECK(in_string.balanced);
#ifdef JSON_TEST_STRICT_NUL_HANDLING_ENABLED
CHECK(in_string.value == json::array({std::string("a\0b", 3)}));
#else
CHECK(in_string.value == json::parse(R"(["a"])"));
#endif
}
SECTION("the SAX parser stops recovering")
{
json j;
RecoveringDomParser sax(j, 2);
CHECK(!json::sax_parse("[1 2 3 4 5]", &sax));
CHECK(sax.errors.size() == 2);
// an error at a delimiter that an invalid token consumed is reported
// to the SAX parser, too
json j2;
RecoveringDomParser sax2(j2, 2);
CHECK(!json::sax_parse("[tru}, 1]", &sax2));
CHECK(sax2.errors.size() == 2);
}
SECTION("an event stops parsing during a repair")
{
// start_object() and key() are passed, then null() for the missing
// value returns false
RecoveringCountdown countdown(2);
CHECK(!json::sax_parse(R"({"a":})", &countdown));
// the end of the input: end_array() for the second array returns false
RecoveringCountdown countdown2(4);
CHECK(!json::sax_parse("[[1", &countdown2));
}
SECTION("input adapters")
{
// the lexer reads contiguous and streaming input differently, and it
// puts back a character that ended an invalid token
for (const std::string s :
{
"[1 2]", "[tru}, 1]", R"({"a" "b\q", "c":[1.x, 2}})", "[\"\xFF\xC3(\", -, 1e+]", "{a:1,\"b\":2", ")]}' [1]"
})
{
CAPTURE(s);
const auto reference = parse_recovering(s);
CHECK(reference.balanced);
const auto from_c_string = parse_recovering(s.c_str());
CHECK(from_c_string.value == reference.value);
CHECK(from_c_string.errors == reference.errors);
const std::list<char> l(s.begin(), s.end());
json j;
RecoveringDomParser sax(j);
CHECK(!json::sax_parse(l.begin(), l.end(), &sax));
CHECK(j == reference.value);
CHECK(sax.errors == reference.errors);
std::istringstream ss(s);
const auto from_stream = parse_recovering(ss);
CHECK(from_stream.value == reference.value);
CHECK(from_stream.errors == reference.errors);
}
}
SECTION("long runs of errors")
{
// no error may copy all the input read before it
const auto closing = parse_recovering("[" + std::string(100000, '}'));
CHECK(closing.balanced);
CHECK(closing.value == json::array());
const auto garbage = parse_recovering("[" + std::string(100000, 'x') + "]");
CHECK(garbage.balanced);
CHECK(garbage.errors.size() == 1);
const auto commas = parse_recovering("{" + std::string(100000, ',') + "}");
CHECK(commas.balanced);
CHECK(commas.value == json::object());
}
SECTION("mutations of valid input")
{
// whatever the input, the events are balanced, every error is reported
// at most once, and valid input is parsed as usual
const std::vector<std::string> documents =
{
R"({"name": "value", "list": [1, -2.5, true, null, {"x": [[]]}], "e": "\u00e9"})",
R"([{"a": [1, 2, {"b": "c"}]}, [], {}, "\ud83d\ude00", 1e10])",
"{\"\xC3\xA9\": \"\xF0\x9F\x98\x80\"}",
R"( {"k" : [ "v" , 0 ] } )",
};
// each character that can be inserted, including a null byte
const std::string insertions("[]{},:\"x\\\0\xFF", 11);
std::vector<std::string> inputs;
for (const auto& doc : documents)
{
for (std::size_t i = 0; i <= doc.size(); ++i)
{
inputs.push_back(doc.substr(0, i));
if (i < doc.size())
{
inputs.push_back(doc.substr(0, i) + doc.substr(i + 1));
}
for (const char c : insertions)
{
inputs.push_back(doc.substr(0, i) + c + doc.substr(i));
}
}
}
for (const auto& s : inputs)
{
CAPTURE(s);
const auto result = parse_recovering(s);
CHECK(result.balanced);
CHECK(result.errors.size() <= s.size() + 1);
CHECK(result.events <= (4 * s.size()) + 4);
if (json::accept(s))
{
CHECK(result.ok);
CHECK(result.errors.empty());
CHECK(result.value == json::parse(s));
}
else
{
CHECK(!result.ok);
CHECK(!result.errors.empty());
}
}
}
}
+59
View File
@@ -430,6 +430,37 @@ TEST_CASE("value conversion")
CHECK(std::equal(std::begin(nbs[0][0][0]), std::end(nbs[1][1][1]), std::begin(nbs2[0][0][0])));
}
SECTION("built-in arrays: 5D")
{
// NOLINTBEGIN(misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
const int nbs[1][1][1][2][2] = {{{{{0, 1}, {2, 3}}}}};
int nbs2[1][1][1][2][2] = {{{{{0, 0}, {0, 0}}}}};
// NOLINTEND(misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
const json j2 = nbs;
j2.get_to(nbs2);
CHECK(std::equal(std::begin(nbs[0][0][0][0]), std::end(nbs[0][0][0][1]), std::begin(nbs2[0][0][0][0])));
}
SECTION("built-in arrays: mismatched shape")
{
// NOLINTBEGIN(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
int nbs2[2][3] = {{0, 0, 0}, {0, 0, 0}};
// NOLINTEND(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
SECTION("not an array")
{
const json j2 = 42;
CHECK_THROWS_WITH_AS(j2.get_to(nbs2), "[json.exception.type_error.304] cannot use at() with number", json::type_error&);
}
SECTION("too few elements")
{
const json j2 = {{0, 1, 2}};
CHECK_THROWS_WITH_AS(j2.get_to(nbs2), "[json.exception.out_of_range.401] array index 1 is out of range", json::out_of_range&);
}
}
SECTION("std::deque<json>")
{
std::deque<json> a{"previous", "value"};
@@ -1726,6 +1757,34 @@ NLOHMANN_JSON_SERIALIZE_ENUM_STRICT(StrictTaskState,
{STRICT_TS_COMPLETED, "completed"},
})
// regression test for #5708 item 2: NLOHMANN_JSON_SERIALIZE_ENUM_STRICT must not rely on
// unqualified lookup of a helper name that a user's own namespace may also declare
namespace ns_with_colliding_name
{
// NOLINTNEXTLINE(misc-use-internal-linkage) - used to shadow the library's internal helper name
inline void templated_json_throw(int /*unused*/) {}
enum class colliding_enum { a, b };
// NOLINTNEXTLINE(misc-use-internal-linkage,misc-const-correctness,cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays) - false positive
NLOHMANN_JSON_SERIALIZE_ENUM_STRICT(colliding_enum,
{
{colliding_enum::a, "a"},
{colliding_enum::b, "b"}
})
} // namespace ns_with_colliding_name
TEST_CASE("NLOHMANN_JSON_SERIALIZE_ENUM_STRICT in a namespace with a colliding name")
{
using ns_with_colliding_name::colliding_enum;
CHECK(json(colliding_enum::a) == "a");
CHECK(colliding_enum::b == json("b"));
json _;
CHECK_THROWS_WITH_AS(_ = json("nope").get<colliding_enum>(), "[json.exception.out_of_range.410] enum value out of range for colliding_enum: \"nope\"", json::out_of_range&);
}
TEST_CASE("Strict JSON to enum mapping")
{
SECTION("enum class")
+16
View File
@@ -19,6 +19,7 @@ using nlohmann::json;
#include <map>
#include <unordered_map>
#include <sstream>
TEST_CASE("Better diagnostics")
{
@@ -492,6 +493,21 @@ TEST_CASE("Regression tests for extended diagnostics")
CHECK(copy == j);
}
}
SECTION("Regression test for issue #5652 - operator>> leaves a partial value in its target on a parse error")
{
json j = "old value";
std::istringstream is("[1, x");
CHECK_THROWS_WITH_AS(is >> j, "[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: '1, x'", json::parse_error);
// j must be left unchanged, as json::parse() guarantees for its result
CHECK(j == "old value");
// copying j must not trigger assert_invariant(): a failed parse must
// not leave array/object elements without a parent pointer
json const copy = j; // NOLINT(performance-unnecessary-copy-initialization)
CHECK(copy == j);
}
}
TEST_CASE("Better diagnostics past the descent bound of update() and merge_patch()")
+10
View File
@@ -516,6 +516,16 @@ TEST_CASE_TEMPLATE("element access 2", Json, nlohmann::json, nlohmann::ordered_j
CHECK(j_array.value("/-"_json_pointer, 42) == 42);
CHECK(j_array_const.value("/-"_json_pointer, 42) == 42);
// Test an index that is syntactically valid but exceeds size_type, and one
// with a trailing non-digit: both must yield the default value rather than
// throw, even with exceptions disabled (regression test for #5708 item 4 /
// #5672: get_checked_or_null() no longer relies on a JSON_TRY/
// JSON_INTERNAL_CATCH around array_index() to turn these into "not found")
CHECK(j_array.value("/18446744073709551615"_json_pointer, 42) == 42);
CHECK(j_array_const.value("/18446744073709551615"_json_pointer, 42) == 42);
CHECK(j_array.value("/1a"_json_pointer, 42) == 42);
CHECK(j_array_const.value("/1a"_json_pointer, 42) == 42);
#if !defined(JSON_NOEXCEPTION)
// Test malformed index (non-numeric) throws parse_error
CHECK_THROWS_WITH_AS(j_array.value("/foo"_json_pointer, 1), "[json.exception.parse_error.109] parse error: array index 'foo' is not a number", typename Json::parse_error&);
-581
View File
@@ -870,585 +870,4 @@ TEST_CASE("regression test - excessive binary container size honors allow_except
CHECK(json::from_cbor(std::vector<std::uint8_t> {0x9b, 0, 0, 0, 0, 0, 0, 0, 0x02}, true, false).is_discarded());
}
namespace
{
/// builds a value from SAX events, asks the parser to recover from its first
/// 100 errors, and checks that the events are balanced (see #3989)
class RecoveringParser
{
public:
explicit RecoveringParser(json& j)
: dom(j, false)
{}
bool null()
{
value();
return dom.null();
}
bool boolean(bool val)
{
value();
return dom.boolean(val);
}
bool number_integer(json::number_integer_t val)
{
value();
return dom.number_integer(val);
}
bool number_unsigned(json::number_unsigned_t val)
{
value();
return dom.number_unsigned(val);
}
bool number_float(json::number_float_t val, const std::string& s)
{
value();
return dom.number_float(val, s);
}
bool string(std::string& val)
{
value();
return dom.string(val);
}
bool binary(json::binary_t& val)
{
value();
return dom.binary(val);
}
bool start_object(std::size_t elements)
{
value();
stack.push_back('o');
return dom.start_object(elements);
}
bool key(std::string& val)
{
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.back() = 'v';
return dom.key(val);
}
bool end_object()
{
if (stack.empty() || stack.back() != 'o')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_object();
}
bool start_array(std::size_t elements)
{
value();
stack.push_back('a');
return dom.start_array(elements);
}
bool end_array()
{
if (stack.empty() || stack.back() != 'a')
{
well_formed = false;
return false;
}
stack.pop_back();
return dom.end_array();
}
bool parse_error(std::size_t /*unused*/, const std::string& /*unused*/, const json::exception& ex)
{
messages.emplace_back(ex.what());
// a limit, so that a reader that does not stop fails the test
// instead of making it hang
return ++errors < 100;
}
/// whether the events were balanced and every key was followed by a value
bool balanced() const
{
return well_formed && stack.empty();
}
/// builds the value
nlohmann::detail::json_sax_dom_parser<json> dom;
std::size_t errors = 0;
std::vector<std::string> messages {}; // NOLINT(readability-redundant-member-init)
std::vector<char> stack {}; // NOLINT(readability-redundant-member-init)
bool well_formed = true;
private:
void value()
{
if (!stack.empty())
{
if (stack.back() == 'v')
{
stack.back() = 'o';
}
else if (stack.back() == 'o')
{
well_formed = false;
}
}
}
};
struct BinaryParseResult
{
json value;
std::size_t errors;
std::vector<std::string> messages;
bool ok;
bool balanced;
};
BinaryParseResult parse_binary_recovering(const std::vector<std::uint8_t>& input, const json::input_format_t format)
{
json j;
RecoveringParser sax(j);
const bool ok = json::sax_parse(input, &sax, format);
return {j, sax.errors, sax.messages, ok, sax.balanced()};
}
#if !defined(JSON_NOEXCEPTION)
/// the message of the exception that reading @a input into a JSON value
/// throws, or an empty string if reading succeeds
std::string binary_error_message(const std::vector<std::uint8_t>& input, const json::input_format_t format)
{
try
{
json _;
switch (format)
{
case json::input_format_t::cbor:
_ = json::from_cbor(input);
break;
case json::input_format_t::msgpack:
_ = json::from_msgpack(input);
break;
case json::input_format_t::ubjson:
_ = json::from_ubjson(input);
break;
case json::input_format_t::bjdata:
_ = json::from_bjdata(input);
break;
case json::input_format_t::bson:
_ = json::from_bson(input);
break;
case json::input_format_t::bon8:
_ = json::from_bon8(input);
break;
case json::input_format_t::json:
default:
break;
}
}
catch (const json::exception& e)
{
return e.what();
}
return "";
}
#endif
/// a BSON element: its type, its name, and its value
std::vector<std::uint8_t> bson_element(const std::uint8_t type, const std::string& name, const std::vector<std::uint8_t>& value)
{
std::vector<std::uint8_t> result = {type};
result.insert(result.end(), name.begin(), name.end());
result.push_back(0x00);
result.insert(result.end(), value.begin(), value.end());
return result;
}
/// a BSON document of the given elements; @a size_offset is added to the
/// size it declares
std::vector<std::uint8_t> bson_document(const std::vector<std::vector<std::uint8_t>>& elements, const int size_offset = 0)
{
std::vector<std::uint8_t> body;
for (const auto& element : elements)
{
body.insert(body.end(), element.begin(), element.end());
}
const auto size = static_cast<std::uint32_t>(static_cast<int>(body.size()) + 5 + size_offset);
std::vector<std::uint8_t> result = {static_cast<std::uint8_t>(size & 0xFFu), static_cast<std::uint8_t>((size >> 8u) & 0xFFu),
static_cast<std::uint8_t>((size >> 16u) & 0xFFu), static_cast<std::uint8_t>((size >> 24u) & 0xFFu)
};
result.insert(result.end(), body.begin(), body.end());
result.push_back(0x00);
return result;
}
/// a BSON int32 value
std::vector<std::uint8_t> bson_int32(const std::int32_t value)
{
const auto u = static_cast<std::uint32_t>(value);
return {static_cast<std::uint8_t>(u & 0xFFu), static_cast<std::uint8_t>((u >> 8u) & 0xFFu),
static_cast<std::uint8_t>((u >> 16u) & 0xFFu), static_cast<std::uint8_t>((u >> 24u) & 0xFFu)};
}
/// a BSON string value, whose length is @a length_offset off
std::vector<std::uint8_t> bson_string(const std::string& value, const std::int32_t length_offset = 0)
{
auto result = bson_int32(static_cast<std::int32_t>(value.size() + 1) + length_offset);
result.insert(result.end(), value.begin(), value.end());
result.push_back(0x00);
return result;
}
/// @a count bytes of value 0xAB
std::vector<std::uint8_t> bytes(const std::size_t count)
{
return std::vector<std::uint8_t>(count, 0xAB);
}
template<typename... Parts>
std::vector<std::uint8_t> concatenated(const std::vector<std::uint8_t>& first, const Parts& ... rest)
{
std::vector<std::uint8_t> result = first;
for (const auto& part : std::initializer_list<std::vector<std::uint8_t>> {rest...})
{
result.insert(result.end(), part.begin(), part.end());
}
return result;
}
/// U+FFFD REPLACEMENT CHARACTER
std::string replacement_character()
{
return "\xEF\xBF\xBD";
}
} // namespace
TEST_CASE("regression test - #3989 SAX parse_error() returning true")
{
SECTION("binary formats complete what was read before the input ends")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}};
const std::vector<std::pair<json::input_format_t, std::vector<std::uint8_t>>> encodings =
{
{json::input_format_t::cbor, json::to_cbor(j)},
{json::input_format_t::msgpack, json::to_msgpack(j)},
{json::input_format_t::ubjson, json::to_ubjson(j)},
{json::input_format_t::ubjson, json::to_ubjson(j, true, true)},
{json::input_format_t::bjdata, json::to_bjdata(j)},
{json::input_format_t::bjdata, json::to_bjdata(j, true, true)},
{json::input_format_t::bson, json::to_bson(j)},
{json::input_format_t::bon8, json::to_bon8(j)},
};
for (const auto& encoding : encodings)
{
const auto format = encoding.first;
const auto& bytes = encoding.second;
CAPTURE(format);
// every prefix is truncated input
for (std::size_t length = 0; length < bytes.size(); ++length)
{
CAPTURE(length);
const auto result = parse_binary_recovering(std::vector<std::uint8_t>(bytes.begin(), bytes.begin() + static_cast<std::ptrdiff_t>(length)), format);
CHECK(!result.ok);
CHECK(result.errors == 1);
CHECK(result.balanced);
}
// the complete input is read as usual (binary values do not
// round-trip through every format, so compare with a plain parse)
json expected;
nlohmann::detail::json_sax_dom_parser<json> dom(expected);
CHECK(json::sax_parse(bytes, &dom, format));
const auto complete = parse_binary_recovering(bytes, format);
CHECK(complete.ok);
CHECK(complete.errors == 0);
CHECK(complete.value == expected);
// a byte after the value
auto trailing_bytes = bytes;
trailing_bytes.push_back(0x01);
const auto trailing = parse_binary_recovering(trailing_bytes, format);
CHECK(!trailing.ok);
CHECK(trailing.errors == 1);
CHECK(trailing.value == expected);
}
}
SECTION("containers without an end")
{
// these made the readers loop, or read on, after the error
const auto cbor_array = parse_binary_recovering({0x9F}, json::input_format_t::cbor);
CHECK(cbor_array.errors == 1);
CHECK(cbor_array.value == json::array());
const auto cbor_map = parse_binary_recovering({0xBF, 0x61, 'a'}, json::input_format_t::cbor);
CHECK(cbor_map.errors == 1);
CHECK(cbor_map.value == json({{"a", nullptr}}));
const auto msgpack_array = parse_binary_recovering({0xDD, 0xFF, 0xFF, 0xFF, 0xFF}, json::input_format_t::msgpack);
CHECK(msgpack_array.errors == 1);
CHECK(msgpack_array.value == json::array());
const auto msgpack_map = parse_binary_recovering({0x81, 0xA1, 'a', 0x92, 0x01}, json::input_format_t::msgpack);
CHECK(msgpack_map.errors == 1);
CHECK(msgpack_map.value == json({{"a", {1}}}));
}
SECTION("BJData ndarray")
{
// a 2x3 int8 array with two of its six elements; the annotated array
// format opens an object and two arrays of its own
const auto result = parse_binary_recovering({'[', '$', 'i', '#', '[', '$', 'i', '#', 'i', 2, 2, 3, 1, 2}, json::input_format_t::bjdata);
CHECK(result.errors == 1);
CHECK(result.balanced);
CHECK(result.value == json({{"_ArrayType_", "int8"}, {"_ArraySize_", {2, 3}}, {"_ArrayData_", {1, 2}}}));
}
SECTION("binary formats repair items whose end is known")
{
struct Repair
{
json::input_format_t format;
std::vector<std::uint8_t> input;
json expected;
std::size_t errors;
};
const std::vector<Repair> repairs =
{
// CBOR: tags are ignored (here tag 1 and the self-describe tag 55799)
{json::input_format_t::cbor, {0x82, 0xC1, 0x05, 0xD9, 0xD9, 0xF7, 0x06}, {5, 6}, 2},
// CBOR: undefined and other simple values become null
{json::input_format_t::cbor, {0x84, 0xF7, 0xE0, 0xF8, 0x20, 0x01}, {nullptr, nullptr, nullptr, 1}, 3},
// CBOR: ill-formed UTF-8 becomes U+FFFD, also in keys
{json::input_format_t::cbor, {0xA1, 0x61, 0xFF, 0x62, 0xC3, 0x28}, {{replacement_character(), replacement_character() + "("}}, 2},
// CBOR: members whose key is not a string are skipped, whatever their key and value
{json::input_format_t::cbor, {0xA4, 0x01, 0x02, 0x82, 0x01, 0x02, 0xA1, 0x61, 'x', 0x9F, 0xFF, 0xC1, 0x01, 0x5F, 0x41, 0x00, 0xFF, 0x61, 'a', 0x03}, {{"a", 3}}, 3},
{json::input_format_t::cbor, {0xBF, 0xF5, 0xBF, 0x61, 'x', 0x7F, 0x61, 'y', 0xFF, 0xFF, 0x61, 'a', 0x03, 0xFF}, {{"a", 3}}, 1},
// MessagePack: members whose key is not a string are skipped
{json::input_format_t::msgpack, {0x84, 0x01, 0x02, 0x81, 0xA1, 'x', 0x01, 0x92, 0x01, 0x02, 0xD4, 0x01, 0x02, 0xC0, 0xA1, 'a', 0x04}, {{"a", 4}}, 3},
// MessagePack: ill-formed UTF-8 becomes U+FFFD
{json::input_format_t::msgpack, {0x92, 0xA2, 0xC3, 0x28, 0xA3, 0xE2, 0x82, 'x'}, {replacement_character() + "(", replacement_character() + "x"}, 2},
// UBJSON: a char that is not ASCII becomes U+FFFD
{json::input_format_t::ubjson, {'[', 'C', 0x80, 'C', 'A', ']'}, {replacement_character(), "A"}, 1},
// UBJSON: the longest beginning of a high-precision number is kept
{json::input_format_t::ubjson, {'[', 'H', 'i', 5, '1', '2', 'a', 'b', 'c', 'H', 'i', 2, '1', '.', 'H', 'i', 3, 'a', 'b', 'c', 'H', 'i', 3, '4', '.', '5', ']'}, {12, 1, nullptr, 4.5}, 3},
// BJData, too
{json::input_format_t::bjdata, {'[', 'C', 0xFF, 'H', 'i', 2, '-', '1', 'H', 'i', 2, '-', 'x', ']'}, {replacement_character(), -1, nullptr}, 2},
// BON8: members whose key is not a string are skipped
{json::input_format_t::bon8, {0x89, 0x91, 0x92, 0xC9, 0x40, 0x82, 0x91, 0x92, 0x61, 0x93}, {{"a", 3}}, 2},
{json::input_format_t::bon8, {0x8B, 0x91, 0x85, 0x91, 0xFE, 0xFA, 0x8B, 'x', 0x91, 0xFE, 0x61, 0x93, 0xFE}, {{"a", 3}}, 2},
// BSON: elements of types the library does not read become null
{
json::input_format_t::bson, bson_document(
{
bson_element(0x07, "_id", bytes(12)), // ObjectId
bson_element(0x09, "date", bytes(8)), // UTC datetime
bson_element(0x13, "decimal", bytes(16)), // 128-bit decimal
bson_element(0x0B, "regex", {'a', '+', 0, 'i', 0}), // regular expression
bson_element(0x0D, "code", bson_string("f()")), // JavaScript code
bson_element(0x0E, "symbol", bson_string("s")), // symbol
bson_element(0x0C, "pointer", concatenated(bson_string("c"), bytes(12))), // DBPointer
bson_element(0x0F, "scope", concatenated(bson_int32(15), bson_string("g"), bson_document({}))), // code with scope
bson_element(0x06, "undefined", {}), // undefined
bson_element(0xFF, "min", {}), // min key
bson_element(0x7F, "max", {}), // max key
bson_element(0x10, "z", bson_int32(7)),
}),
{{"_id", nullptr}, {"date", nullptr}, {"decimal", nullptr}, {"regex", nullptr}, {"code", nullptr}, {"symbol", nullptr}, {"pointer", nullptr}, {"scope", nullptr}, {"undefined", nullptr}, {"min", nullptr}, {"max", nullptr}, {"z", 7}},
11
},
// BSON: an element of an unknown type becomes null, and the rest of its document is skipped
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x04, "array", bson_document({bson_element(0x10, "0", bson_int32(1)), bson_element(0x42, "1", bytes(3))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"a", 1}, {"x", nullptr}}}, {"array", {1, nullptr}}, {"after", 3}},
2
},
// BSON: so does a string or byte array whose length cannot be right
{
json::input_format_t::bson, bson_document(
{
bson_element(0x03, "inner", bson_document({bson_element(0x02, "s", bson_string("abc", -10)), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x03, "bin", bson_document({bson_element(0x05, "b", concatenated(bson_int32(-1), bytes(1))), bson_element(0x10, "b", bson_int32(2))})),
bson_element(0x10, "after", bson_int32(3)),
}),
{{"inner", {{"s", nullptr}}}, {"bin", {{"b", nullptr}}}, {"after", 3}},
2
},
// BSON: a string without its terminator, and a document whose size does not match, are kept
{
json::input_format_t::bson, bson_document(
{
bson_element(0x02, "s", {2, 0, 0, 0, 'a', 'X'}),
bson_element(0x03, "inner", bson_document({bson_element(0x10, "a", bson_int32(1))}, 1)),
}),
{{"s", "a"}, {"inner", {{"a", 1}}}},
2
},
};
for (const auto& repair : repairs)
{
CAPTURE(repair.format);
CAPTURE(repair.input);
const auto result = parse_binary_recovering(repair.input, repair.format);
CHECK(!result.ok);
CHECK(result.balanced);
CHECK(result.errors == repair.errors);
CHECK(result.value == repair.expected);
REQUIRE(!result.messages.empty());
#if !defined(JSON_NOEXCEPTION)
// the first error is the one reported without recovering; under
// JSON_NOEXCEPTION, reading without recovering aborts instead of
// throwing, so there is no message to compare with
CHECK(result.messages.front() == binary_error_message(repair.input, repair.format));
#endif
}
}
SECTION("binary formats repair numbers that are out of range")
{
// CBOR: a negative integer below the range of number_integer_t
const auto cbor = parse_binary_recovering({0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value.is_number_float());
CHECK(cbor.value.get<double>() == -18446744073709551616.0);
// UBJSON: a high-precision number too large for number_float_t
const auto ubjson = parse_binary_recovering({'H', 'i', 5, '1', 'e', '9', '9', '9'}, json::input_format_t::ubjson);
CHECK(ubjson.errors == 1);
CHECK(ubjson.value.is_number_float());
CHECK(std::isinf(ubjson.value.get<double>()));
}
SECTION("binary formats stop where the end of an item is not known")
{
// a byte that begins no item
const auto cbor = parse_binary_recovering({0x82, 0x01, 0x1C, 0x02}, json::input_format_t::cbor);
CHECK(cbor.errors == 1);
CHECK(cbor.value == json({1}));
// a key that is no item: the unused MessagePack byte, a CBOR break
// in a map of known size, and the end of a BON8 container
const auto msgpack = parse_binary_recovering({0x82, 0xA1, 'a', 0x01, 0xC1, 0x02}, json::input_format_t::msgpack);
CHECK(msgpack.errors == 1);
CHECK(msgpack.value == json({{"a", 1}}));
const auto cbor_break = parse_binary_recovering({0xA2, 0x61, 'a', 0x01, 0xFF, 0x02}, json::input_format_t::cbor);
CHECK(cbor_break.errors == 1);
CHECK(cbor_break.value == json({{"a", 1}}));
const auto bon8 = parse_binary_recovering({0x88, 0x61, 0x91, 0xFE}, json::input_format_t::bon8);
CHECK(bon8.errors == 1);
CHECK(bon8.value == json({{"a", 1}}));
// a skipped member that the input ends in
const auto truncated = parse_binary_recovering({0xA2, 0x01, 0x82, 0x01}, json::input_format_t::cbor);
CHECK(truncated.errors == 2);
CHECK(truncated.balanced);
CHECK(truncated.value == json::object());
// a BSON element of an unknown type in a document whose size cannot be right
const auto bson = parse_binary_recovering(bson_document({bson_element(0x10, "a", bson_int32(1)), bson_element(0x42, "x", bytes(3))}, -10), json::input_format_t::bson);
CHECK(bson.errors == 1);
CHECK(bson.value == json({{"a", 1}, {"x", nullptr}}));
}
SECTION("changed bytes in binary input")
{
const json j = {{"a", {1, -2, {{"b", "c"}}, json::array()}}, {"d", {{"e", nullptr}, {"f", true}}}, {"g", 1.5}, {"h", json::binary({1, 2, 3})}, {"i", "\xC3\xA4"}};
const std::vector<std::pair<json::input_format_t, std::vector<std::uint8_t>>> encodings =
{
{json::input_format_t::cbor, json::to_cbor(j)},
{json::input_format_t::msgpack, json::to_msgpack(j)},
{json::input_format_t::ubjson, json::to_ubjson(j)},
{json::input_format_t::ubjson, json::to_ubjson(j, true, true)},
{json::input_format_t::bjdata, json::to_bjdata(j)},
{json::input_format_t::bjdata, json::to_bjdata(j, true, true)},
{json::input_format_t::bson, json::to_bson(j)},
{json::input_format_t::bon8, json::to_bon8(j)},
};
const std::vector<std::uint8_t> replacements = {0x00, 0x01, 0x7F, 0x80, 0xC1, 0xD9, 0xE0, 0xF7, 0xFE, 0xFF};
for (const auto& encoding : encodings)
{
const auto format = encoding.first;
const auto& original = encoding.second;
CAPTURE(format);
std::vector<std::vector<std::uint8_t>> inputs;
for (std::size_t position = 0; position < original.size(); ++position)
{
for (const auto replacement : replacements)
{
auto changed = original;
changed[position] = replacement;
inputs.push_back(changed);
}
auto removed = original;
removed.erase(removed.begin() + static_cast<std::ptrdiff_t>(position));
inputs.push_back(removed);
}
for (const auto& input : inputs)
{
CAPTURE(input);
const auto result = parse_binary_recovering(input, format);
CHECK(result.balanced);
CHECK(result.errors <= input.size() + 1);
#if !defined(JSON_NOEXCEPTION)
// an error is reported exactly if reading into a JSON value
// fails, and the first one is the same (under JSON_NOEXCEPTION,
// that reading aborts instead of throwing)
const auto message = binary_error_message(input, format);
CHECK(result.ok == message.empty());
if (!result.ok && result.errors < 100)
{
CHECK(result.messages.front() == message);
}
#endif
}
}
}
SECTION("JSON text")
{
// the parser stopped, but reported success
json j;
RecoveringParser sax(j);
CHECK(!json::sax_parse("[1,2,3,]", &sax));
CHECK(sax.errors == 1);
CHECK(j == json({1, 2, 3}));
}
SECTION("the SAX parsers of the library stop")
{
json _;
CHECK(json::from_cbor(std::vector<std::uint8_t> {0x9F}, true, false).is_discarded());
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<std::uint8_t> {0x9F}), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing CBOR value: unexpected end of input", json::parse_error&);
CHECK(json::parse("[1,2,3,]", nullptr, false).is_discarded());
CHECK(!json::accept("[1,2,3,]"));
}
}
DOCTEST_CLANG_SUPPRESS_WARNING_POP
+12
View File
@@ -10,10 +10,14 @@
#if JSON_TEST_USING_MULTIPLE_HEADERS
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/ordered_map.hpp>
#else
#include <nlohmann/json.hpp>
#endif
#include <map>
#include <string>
TEST_CASE("type traits")
{
SECTION("is_c_string")
@@ -83,4 +87,12 @@ TEST_CASE("type traits")
}
}
}
SECTION("is_ordered_map")
{
using nlohmann::detail::is_ordered_map;
CHECK(is_ordered_map<nlohmann::ordered_map<std::string, int>>::value);
CHECK_FALSE(is_ordered_map<std::map<std::string, int>>::value);
}
}