Files
json/tests/src/unit-class_parser.cpp
T
Niels Lohmann a0b71e2720 Deduplicate basic_json internals; make insert(pos, json&&) move (#5727)
* Remove unused private aliases from basic_json

The private aliases primitive_iterator_t, internal_iterator and
output_adapter_t are not used anywhere: iter_impl, binary_writer and
the tests refer to the detail:: names directly. As the aliases are
private, no user or derived class can depend on them. The
internal_iterator.hpp include stays because iter_impl needs it.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix meta()'s dead, syntactically invalid HP aCC branch

The HP aCC branch of basic_json::meta() was missing a semicolon and
has therefore never compiled; adding only a semicolon would also make
it throw type_error.305, since it assigned a plain string to
result["compiler"] and then indexed into it like the other branches
do into an object. Make the branch consistent with the others by
assigning an object with "family" and "version" keys, narrow the
condition to __HP_aCC (a C compiler cannot build this header-only
library), and fix meta.md, which documented the old (impossible)
plain-string behavior.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the noexcept null constructor delegating to a throwing one

basic_json(std::nullptr_t) delegated to basic_json(value_t), whose
underlying json_value(value_t) constructor allocates for other types
and can therefore throw, which is why the noexcept had a
NOLINT(bugprone-exception-escape). The delegated-to constructor also
called assert_invariant() a second time. The default member
initializers of data already produce the same null state (a
value-initialized, i.e. zeroed, union with object == nullptr), so the
delegation and its NOLINT can simply be dropped.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove four cppcheck accessForwarded suppressions in the move constructor

basic_json(basic_json&&) built its base subobject with
std::forward<json_base_class_t>(other), so cppcheck saw the whole of
other as forwarded and flagged every subsequent access to it as
accessForwarded, three of them still marked "TODO check". Only the
base subobject is actually moved from; cast explicitly to the base
type instead, the way ordered_map already does, so cppcheck can tell
the two are unrelated. Behavior is unchanged: for a non-reference T,
std::forward<T>(x) is defined as static_cast<T&&>(x).

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove stale cppcheck suppressions and name the local parser in parse()

Running the pinned cppcheck (ci_cppcheck's invocation) without
--inline-suppr across all configurations reports no syntaxError, no
ignoredReturnValue and no assertWithSideEffect, so the corresponding
suppressions in json_fwd.hpp, string_concat.hpp and
assert_invariant() no longer match anything (json_fwd.hpp's is kept,
since downstream users who run an older cppcheck against it could
still hit the warning it once silenced).

The three basic_json::parse() overloads still trigger a false-positive
accessMoved/accessForwarded because they build a temporary parser and
call .parse() on it in the same expression; giving that parser a name
makes the warning go away without changing behavior, and removes the
last of the inline suppressions on these functions.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the string/binary cleanup in the two erase() overloads

erase(pos) and erase(first, last) each carried a byte-identical
14-line block that destroys and deallocates a string or binary
value before resetting the type to null. That reimplements the
string/binary cases of json_value::destroy(), so any future change to
how those values are freed would have to be made in three places
instead of one. Both overloads now just call destroy() and reset the
union; for the other primitive types (boolean, numbers) destroy() is
a no-op, so behavior is unchanged.

Also fix erase(first, last)'s error-path branch hint, which used
JSON_HEDLEY_LIKELY where erase(pos), the iterator-range constructor,
and every other error path in the class use JSON_HEDLEY_UNLIKELY.
This only affects code layout, not semantics.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale and copy-pasted comments in basic_json

Several comments no longer match the code: the class invariant and
assert_invariant()'s doc still named the members m_value/m_type
(now m_data.m_value/m_data.m_type) and did not mention the binary
invariant that assert_invariant() already checks; the json_value note
and the get<PointerType>() @tparam list omitted binary_t even though
binary is a variable-length, pointer-stored type like the others; the
key-based value() overload's brief said "via JSON Pointer", which is
the other overload; and swap(binary_t&)/swap(binary_t::container_type&)
both carried "swap only works for strings", copied from swap(string_t&).

Comment-only change; behavior, the public API and the ABI are
unchanged. The private get_impl() doxygen and the emplace() comments
that border #5585's hunk are intentionally left alone.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the 16 copied from_cbor/msgpack/ubjson/bjdata/bon8/bson bodies

Each of the 16 binary deserialization overloads (from_cbor,
from_msgpack, from_ubjson, from_bjdata, from_bon8, from_bson, each in
an InputType&& and an iterator/sentinel version, plus the deprecated
span overloads of from_cbor/from_msgpack/from_ubjson/from_bson) had
the same body, differing only in the input_format_t value. Every copy
built a temporary binary_reader from std::move(ia) and called
sax_parse on it in the same expression, which also produced a
false-positive cppcheck accessMoved on all 16 lines and needed a
NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) on
the four span overloads.

Add a private from_binary_impl() helper that builds the reader as a
named local instead, and make each of the 16 overloads a one-line
forward to it. All public signatures, default arguments,
JSON_HEDLEY_WARN_UNUSED_RESULT and JSON_HEDLEY_DEPRECATED_FOR
attributes are unchanged, tag_handler keeps defaulting to
cbor_tag_handler_t::error for the non-CBOR formats (matching
binary_reader::sax_parse's own default), and the helper is placed in
the existing private section before the binary section banner rather
than between the from_* overloads, so from_binary_impl() itself does
not collide with #5688's insertion point. Collapsing the
from_bjdata/from_bon8 bodies into one-line forwards does rewrite the
"return result; }" context lines that #5688 inserts its two
deprecated overloads after, so that PR will need a small manual
rebase (reinserting its overloads after the new one-line bodies)
rather than applying cleanly.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make insert(pos, basic_json&&) move its argument instead of copying it

insert(const_iterator pos, basic_json&& val) delegated to
insert(pos, val), but val is a named rvalue reference, so inside the
function it is an lvalue: the call always resolved to
insert(const_iterator, const basic_json&) and deep-copied the value.
This has been the case since the overload was introduced, in every
release. push_back(basic_json&&), by contrast, already moves.

Give the rvalue overload its own body with the same two checks
(type_error.309, invalid_iterator.202), then move the argument into a
local before inserting it. Moving into a local first, rather than
inserting std::move(val) directly, keeps this safe even when val
aliases an element of the same array (e.g.
arr.insert(arr.begin(), std::move(arr[1]))), since
std::vector::insert(pos, T&&) is not guaranteed to handle an argument
that aliases one of its own elements.

This is a deliberate, small behavior change: the moved-from argument
now ends up null afterwards, the same as after push_back(&&), instead
of keeping its old value unchanged. No signature changes, so the
public API and ABI are unaffected. Add unit-modifiers coverage for
the moved-from state and for self-aliasing insertion, both with and
without reallocation of the underlying array.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate object key lookup and checked at() access

at()/find()/count()/contains()/const operator[]/erase_internal() each
repeated the raw object lookup (m_value.object->find(key)), and the
six at() overloads additionally repeated the type_error.304 check and
out_of_range.401/403 throw. Route them all through two new private
helpers, object_lookup()/object_at() (plus array_at() for the index
overloads of at()), templated on the constness of the receiver so one
body serves both the const and non-const overload. count() is left
untouched, since it already goes through object_t::count() rather
than a second find().

The at(KeyType&&) overloads used to forward the same key twice: once
into object->find() and again, on the not-found path, into the
string_t() conversion for the exception message. object_at() now
forwards it only into the lookup and reuses the (unmoved) key for the
message. clang-tidy 22 (Docker silkeh/clang:22) still flags that reuse
under bugprone-use-after-move/hicpp-invalid-access-moved even with the
single forward, since it cannot see that object_t::find() (a plain
std::map or ordered_map) never actually moves from its argument; add a
NOLINTNEXTLINE with that reasoning rather than avoid the pattern.

No signature, exception id/message, or set_parent() behavior changes.

Overlaps #5689, #5705, #5606, #5687 and #5585, which touch the same
hunks; whichever of this commit and those PRs lands second will need
a small rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the lookup/default/throw body of value()

The six non-deprecated value() overloads each held a full copy of the
same body: the four key-based overloads looked up the key and either
returned the found element converted to the requested type or the
default value (throwing type_error.306 if this is not an object), and
the two json_pointer overloads did the same via
ptr.get_checked_or_null(), throwing type_error.306 unless
is_structured(). Replace the duplicated bodies with two private
helpers, value_member() and value_pointee(), that return a
const basic_json* (null when not found) and do the type check/throw
once each. Every value() overload now just picks between the found
pointer's get<T>() and the default.

Same signatures, template parameters, SFINAE conditions, exception id,
message and this context on every overload.

Overlaps #5689 (routes find() through lookup_key()) and #5705 (adds a
deleted integral-key value() next to these overloads); whichever of
this commit and those PRs lands second will need a small rebase.

#5724 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the null-to-container conversion into convert_null_to()

Nine sites wrote out the same "turn a null value into an empty array
or object" logic with two different idioms: operator[](size_type),
operator[](key_type), operator[](KeyType&&) and update() set m_type
then assigned m_value.array/object directly via create<T>(), while the
three push_back() overloads, emplace_back() and emplace() set m_type
then assigned m_value = value_t::array/object (going through
json_value's converting constructor and a temporary). Both idioms end
up calling create<T>() and produce the same state, just via a
different path; both also share a latent exception-safety bug, since
m_type is written before the (possibly throwing) allocation, so a
throwing allocator leaves m_type == array/object with a null pointer
behind it, violating the class invariant and crashing on the next
access to, or destruction of, the value.

Add a private convert_null_to(value_t) helper and call it from all
nine sites. Unlike the idioms it replaces, it allocates the container
first and only then writes m_type, so a throwing allocation leaves the
value as a valid null instead of a mistyped, half-constructed one;
verified with a throwing allocator (see unit-allocator.cpp's
bad_allocator) that j["x"] = ... on a null j now stays null, and no
longer trips assert_invariant()/crashes, when create<object_t>()
throws. Same allocator usage and assert_invariant() call as before,
otherwise.

Overlaps #5585, which reorders these same nine blocks for exception
safety; whichever of this commit and that PR lands second will need a
small rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the linear key search in ordered_map

emplace, at, erase(key), count and find each repeated the same
"for (auto it = begin(); it != end(); ++it) if (m_compare(it->first,
key)) ..." loop (15 copies across their key_type and transparent
KeyType&& overloads), and both erase(key) overloads additionally
repeated the exception-sensitive in-place reconstruction (destroy,
placement-new, pop_back) used to remove an element while keeping the
const Key non-movable.

Add two private helpers: find_impl(Self&, KeyType&&), a static member
template that runs the search once for either constness of the
receiver, and erase_at(iterator), which keeps the existing
pop_back-based reconstruction instead of switching to erase()/resize()
(which would add a DefaultInsertable requirement). Route find, at,
count, emplace, insert(const value_type&) and both erase(key)
overloads through them.

Same signatures, is_usable_as_key_type constraints and exception
messages/types.

Overlaps #5609 and #5685, which both rewrite emplace (#5609 also
touches insert and adds private members at the end of the class);
whichever of this commit and those PRs lands second will need a
rebase.

#5724 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Re-enable bugprone-use-after-move/hicpp-invalid-access-moved

These two checks (and portability-template-virtual-member-function)
were disabled in #4489 (November 2024) "only removed to get the CI
going". portability-template-virtual-member-function is a separate,
still-open cleanup (#5725 item 3 on its own branch) and stays
disabled here; this commit only re-enables the move/forward checks
and cleans up what they flag on this branch.

The move constructor (json.hpp) already casts to the base type
instead of forwarding the whole object (#5724 item 9), so it no
longer trips either check. The at(KeyType&&) double-forward this
check used to flag was reduced to a single forward with the
now-unforwarded reuse annotated by a NOLINTNEXTLINE in #5724 item 3's
object_at() helper (clang-tidy 22 still flags that reuse even after a
single forward; see that commit's message). What is left here:

- from_json_inplace_array_impl(), from_json_tuple_impl_base() and the
  std::pair overload of from_json_tuple_impl() forwarded j into every
  j.at(...) call in a pack expansion or a pair of calls. at() has no
  ref-qualified overloads, so the forward was a no-op; call j.at(...)
  directly.
- container_input_adapter_factory::create() forwards container twice
  on purpose, into begin() and end(), so both see the same value
  category and produce matching iterator types. Annotate it with
  NOLINTNEXTLINE and a comment instead of changing it.
- unit-class_parser.cpp's "move constructor resets the moved-from
  value to npos" test still pointed at the pre-static_cast move
  constructor by line number and mentioned the cppcheck-suppress
  annotation that #5724 item 9 already removed; update the comment.

No behavior change anywhere in include/. Verified with clang-tidy 22.1.8
(Docker silkeh/clang:22, --platform linux/amd64) against a TU including
json.hpp with the repo's .clang-tidy: bugprone-use-after-move and
hicpp-invalid-access-moved report nothing unsuppressed.

Overlaps #5737 (open PR for the rest of #5725 item 3: the
from_json.hpp/input_adapters.hpp cleanup above, and
portability-template-virtual-member-function), which currently keeps
both checks disabled pending this move-constructor change; whichever
of this commit and that PR lands second will need a small rebase of
.clang-tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make convert_null_to() take the container type as a template argument

Passing array_t or object_t instead of a value_t makes an invalid target
a compile error instead of a runtime assertion, and removes the branch.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:57:41 +02:00

2977 lines
140 KiB
C++
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
// capture whether JSON_STRICT_NUL_HANDLING was enabled on the command line
// (e.g. -DJSON_STRICT_NUL_HANDLING=1) *before* including json.hpp, since the
// library #undefs JSON_STRICT_NUL_HANDLING itself once the header has been
// fully processed unless JSON_TEST_KEEP_MACROS is defined (see
// include/nlohmann/detail/macro_unscope.hpp)
#if defined(JSON_STRICT_NUL_HANDLING) && (JSON_STRICT_NUL_HANDLING == 1)
#define JSON_TEST_STRICT_NUL_HANDLING_ENABLED 1
#endif
#define JSON_TEST_STRINGIZE_EX(x) #x
#define JSON_TEST_STRINGIZE(x) JSON_TEST_STRINGIZE_EX(x)
#define JSON_TESTS_PRIVATE
#include <nlohmann/json.hpp>
using nlohmann::json;
#ifdef JSON_TEST_NO_GLOBAL_UDLS
using namespace nlohmann::literals; // NOLINT(google-build-using-namespace)
#endif
#include <valarray>
#include <algorithm>
#include <cstdio>
#include <fstream>
#include <list>
#include <sstream>
#include <string>
#include <utility>
#include <vector>
#include "test_utils.hpp"
namespace
{
class SaxEventLogger
{
public:
bool null()
{
events.emplace_back("null()");
return true;
}
bool boolean(bool val)
{
events.emplace_back(val ? "boolean(true)" : "boolean(false)");
return true;
}
bool number_integer(json::number_integer_t val)
{
events.push_back("number_integer(" + std::to_string(val) + ")");
return true;
}
bool number_unsigned(json::number_unsigned_t val)
{
events.push_back("number_unsigned(" + std::to_string(val) + ")");
return true;
}
bool number_float(json::number_float_t /*unused*/, const std::string& s)
{
events.push_back("number_float(" + s + ")");
return true;
}
bool string(std::string& val)
{
events.push_back("string(" + val + ")");
return true;
}
bool binary(json::binary_t& val)
{
std::string binary_contents = "binary(";
std::string comma_space;
for (auto b : val)
{
binary_contents.append(comma_space);
binary_contents.append(std::to_string(static_cast<int>(b)));
comma_space = ", ";
}
binary_contents.append(")");
events.push_back(binary_contents);
return true;
}
bool start_object(std::size_t elements)
{
if (elements == (std::numeric_limits<std::size_t>::max)())
{
events.emplace_back("start_object()");
}
else
{
events.push_back("start_object(" + std::to_string(elements) + ")");
}
return true;
}
bool key(std::string& val)
{
events.push_back("key(" + val + ")");
return true;
}
bool end_object()
{
events.emplace_back("end_object()");
return true;
}
bool start_array(std::size_t elements)
{
if (elements == (std::numeric_limits<std::size_t>::max)())
{
events.emplace_back("start_array()");
}
else
{
events.push_back("start_array(" + std::to_string(elements) + ")");
}
return true;
}
bool end_array()
{
events.emplace_back("end_array()");
return true;
}
bool parse_error(std::size_t position, const std::string& /*unused*/, const json::exception& /*unused*/)
{
errored = true;
events.push_back("parse_error(" + std::to_string(position) + ")");
return false;
}
std::vector<std::string> events {}; // NOLINT(readability-redundant-member-init)
bool errored = false;
};
class SaxCountdown : public nlohmann::json::json_sax_t
{
public:
explicit SaxCountdown(const int count) : events_left(count)
{}
bool null() override
{
return events_left-- > 0;
}
bool boolean(bool /*val*/) override
{
return events_left-- > 0;
}
bool number_integer(json::number_integer_t /*val*/) override
{
return events_left-- > 0;
}
bool number_unsigned(json::number_unsigned_t /*val*/) override
{
return events_left-- > 0;
}
bool number_float(json::number_float_t /*val*/, const std::string& /*s*/) override
{
return events_left-- > 0;
}
bool string(std::string& /*val*/) override
{
return events_left-- > 0;
}
bool binary(json::binary_t& /*val*/) override
{
return events_left-- > 0;
}
bool start_object(std::size_t /*elements*/) override
{
return events_left-- > 0;
}
bool key(std::string& /*val*/) override
{
return events_left-- > 0;
}
bool end_object() override
{
return events_left-- > 0;
}
bool start_array(std::size_t /*elements*/) override
{
return events_left-- > 0;
}
bool end_array() override
{
return events_left-- > 0;
}
bool parse_error(std::size_t /*position*/, const std::string& /*last_token*/, const json::exception& /*ex*/) override
{
return false;
}
private:
int events_left = 0;
};
json parser_helper(const std::string& s);
bool accept_helper(const std::string& s);
void comments_helper(const std::string& s);
void trailing_comma_helper(const std::string& s);
json parser_helper(const std::string& s)
{
json j;
json::parser(nlohmann::detail::input_adapter(s)).parse(true, j);
// if this line was reached, no exception occurred
// -> check if result is the same without exceptions
json j_nothrow;
CHECK_NOTHROW(json::parser(nlohmann::detail::input_adapter(s), nullptr, false).parse(true, j_nothrow));
CHECK(j_nothrow == j);
json j_sax;
nlohmann::detail::json_sax_dom_parser<json, nlohmann::detail::string_input_adapter_type> sdp(j_sax);
json::sax_parse(s, &sdp);
CHECK(j_sax == j);
comments_helper(s);
trailing_comma_helper(s);
return j;
}
bool accept_helper(const std::string& s)
{
CAPTURE(s)
// 1. parse s without exceptions
json j;
CHECK_NOTHROW(json::parser(nlohmann::detail::input_adapter(s), nullptr, false).parse(true, j));
const bool ok_noexcept = !j.is_discarded();
// 2. accept s
const bool ok_accept = json::parser(nlohmann::detail::input_adapter(s)).accept(true);
// 3. check if both approaches come to the same result
CHECK(ok_noexcept == ok_accept);
// 4. parse with SAX (compare with relaxed accept result)
SaxEventLogger el;
CHECK_NOTHROW(json::sax_parse(s, &el, json::input_format_t::json, false));
CHECK(json::parser(nlohmann::detail::input_adapter(s)).accept(false) == !el.errored);
// 5. parse with simple callback
json::parser_callback_t const cb = [](int /*unused*/, json::parse_event_t /*unused*/, json& /*unused*/) noexcept
{
return true;
};
json const j_cb = json::parse(s, cb, false);
const bool ok_noexcept_cb = !j_cb.is_discarded();
// 6. check if this approach came to the same result
CHECK(ok_noexcept == ok_noexcept_cb);
// 7. check if comments or trailing commas are properly ignored
if (ok_accept)
{
comments_helper(s);
trailing_comma_helper(s);
}
// 8. return result
return ok_accept;
}
void comments_helper(const std::string& s)
{
json _;
// parse/accept with default parser
CHECK_NOTHROW(_ = json::parse(s));
CHECK(json::accept(s));
// parse/accept while skipping comments
CHECK_NOTHROW(_ = json::parse(s, nullptr, false, true));
CHECK(json::accept(s, true));
std::vector<std::string> json_with_comments;
// start with a comment
json_with_comments.push_back(std::string("// this is a comment\n") + s);
json_with_comments.push_back(std::string("/* this is a comment */") + s);
// end with a comment
json_with_comments.push_back(s + "// this is a comment");
json_with_comments.push_back(s + "/* this is a comment */");
// check all strings
for (const auto& json_with_comment : json_with_comments)
{
CAPTURE(json_with_comment)
CHECK_THROWS_AS(_ = json::parse(json_with_comment), json::parse_error);
CHECK(!json::accept(json_with_comment));
CHECK_NOTHROW(_ = json::parse(json_with_comment, nullptr, true, true));
CHECK(json::accept(json_with_comment, true));
}
}
void trailing_comma_helper(const std::string& s)
{
json _;
// parse/accept with default parser
CHECK_NOTHROW(_ = json::parse(s));
CHECK(json::accept(s));
// parse/accept while allowing trailing commas
CHECK_NOTHROW(_ = json::parse(s, nullptr, false, false, true));
CHECK(json::accept(s, false, true));
// note: [,] and {,} are not allowed
if (s.size() > 1 && (s.back() == ']' || s.back() == '}') && !_.empty())
{
std::vector<std::string> json_with_trailing_commas;
json_with_trailing_commas.push_back(s.substr(0, s.size() - 1) + " ," + s.back());
json_with_trailing_commas.push_back(s.substr(0, s.size() - 1) + "," + s.back());
json_with_trailing_commas.push_back(s.substr(0, s.size() - 1) + ", " + s.back());
for (const auto& json_with_trailing_comma : json_with_trailing_commas)
{
CAPTURE(json_with_trailing_comma)
CHECK_THROWS_AS(_ = json::parse(json_with_trailing_comma), json::parse_error);
CHECK(!json::accept(json_with_trailing_comma));
CHECK_NOTHROW(_ = json::parse(json_with_trailing_comma, nullptr, true, false, true));
CHECK(json::accept(json_with_trailing_comma, false, true));
}
}
}
#if JSON_DIAGNOSTIC_POSITIONS
/**
* Validates that the generated JSON object is the same as expected
* Validates that the start position and end position match the start and end of the string
*
* This check assumes that there is no whitespace around the json object in the original string.
*/
void validate_generated_json_and_start_end_pos_helper(const std::string& original_string, const json& j, const json& check)
{
CHECK(j == check);
CHECK(j.start_pos() == 0);
CHECK(j.end_pos() == original_string.size());
}
/**
* Parses the root object from the given root string and validates that the start and end positions for the nested object are correct.
*
* This checks that whitespace around the nested object is included in the start and end positions of the root object.
*/
void validate_start_end_pos_for_nested_obj_helper(const std::string& nested_type_json_str, const std::string& root_type_json_str, const json& expected_json, const json::parser_callback_t& cb = nullptr)
{
json j;
// 1. If callback is provided, use callback version of parse()
if (cb)
{
j = json::parse(root_type_json_str, cb);
}
else
{
j = json::parse(root_type_json_str);
}
// 2. Check if the generated JSON is as expected
// Assumptions: The root_type_json_str does not have any whitespace around the json object
validate_generated_json_and_start_end_pos_helper(root_type_json_str, j, expected_json);
// 3. Get the nested object
const auto& nested = j["nested"];
// 4. Check if the start and end positions are generated correctly for nested objects and arrays
CHECK(nested_type_json_str == root_type_json_str.substr(nested.start_pos(), nested.end_pos() - nested.start_pos()));
}
#endif
} // namespace
TEST_CASE("parser class")
{
SECTION("parse")
{
SECTION("null")
{
CHECK(parser_helper("null") == json(nullptr));
}
SECTION("true")
{
CHECK(parser_helper("true") == json(true));
}
SECTION("false")
{
CHECK(parser_helper("false") == json(false));
}
SECTION("array")
{
SECTION("empty array")
{
CHECK(parser_helper("[]") == json(json::value_t::array));
CHECK(parser_helper("[ ]") == json(json::value_t::array));
}
SECTION("nonempty array")
{
CHECK(parser_helper("[true, false, null]") == json({true, false, nullptr}));
}
}
SECTION("object")
{
SECTION("empty object")
{
CHECK(parser_helper("{}") == json(json::value_t::object));
CHECK(parser_helper("{ }") == json(json::value_t::object));
}
SECTION("nonempty object")
{
CHECK(parser_helper("{\"\": true, \"one\": 1, \"two\": null}") == json({{"", true}, {"one", 1}, {"two", nullptr}}));
}
}
SECTION("string")
{
// empty string
CHECK(parser_helper("\"\"") == json(json::value_t::string));
SECTION("errors")
{
// error: tab in string
CHECK_THROWS_WITH_AS(parser_helper("\"\t\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0009 (HT) must be escaped to \\u0009 or \\t; last read: '\"<U+0009>'", json::parse_error&);
// error: newline in string
CHECK_THROWS_WITH_AS(parser_helper("\"\n\""), "[json.exception.parse_error.101] parse error at line 2, column 0: syntax error while parsing value - invalid string: control character U+000A (LF) must be escaped to \\u000A or \\n; last read: '\"<U+000A>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\r\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+000D (CR) must be escaped to \\u000D or \\r; last read: '\"<U+000D>'", json::parse_error&);
// error: backspace in string
CHECK_THROWS_WITH_AS(parser_helper("\"\b\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0008 (BS) must be escaped to \\u0008 or \\b; last read: '\"<U+0008>'", json::parse_error&);
// improve code coverage
CHECK_THROWS_AS(parser_helper("\uFF01"), json::parse_error&);
CHECK_THROWS_AS(parser_helper("[-4:1,]"), json::parse_error&);
// unescaped control characters
CHECK_THROWS_WITH_AS(parser_helper("\"\x00\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: missing closing quote; last read: '\"'", json::parse_error&); // NOLINT(bugprone-string-literal-with-embedded-nul)
CHECK_THROWS_WITH_AS(parser_helper("\"\x01\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0001 (SOH) must be escaped to \\u0001; last read: '\"<U+0001>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x02\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0002 (STX) must be escaped to \\u0002; last read: '\"<U+0002>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x03\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0003 (ETX) must be escaped to \\u0003; last read: '\"<U+0003>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x04\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0004 (EOT) must be escaped to \\u0004; last read: '\"<U+0004>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x05\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0005 (ENQ) must be escaped to \\u0005; last read: '\"<U+0005>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x06\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0006 (ACK) must be escaped to \\u0006; last read: '\"<U+0006>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x07\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0007 (BEL) must be escaped to \\u0007; last read: '\"<U+0007>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x08\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0008 (BS) must be escaped to \\u0008 or \\b; last read: '\"<U+0008>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x09\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0009 (HT) must be escaped to \\u0009 or \\t; last read: '\"<U+0009>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x0a\""), "[json.exception.parse_error.101] parse error at line 2, column 0: syntax error while parsing value - invalid string: control character U+000A (LF) must be escaped to \\u000A or \\n; last read: '\"<U+000A>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x0b\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+000B (VT) must be escaped to \\u000B; last read: '\"<U+000B>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x0c\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+000C (FF) must be escaped to \\u000C or \\f; last read: '\"<U+000C>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x0d\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+000D (CR) must be escaped to \\u000D or \\r; last read: '\"<U+000D>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x0e\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+000E (SO) must be escaped to \\u000E; last read: '\"<U+000E>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x0f\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+000F (SI) must be escaped to \\u000F; last read: '\"<U+000F>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x10\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0010 (DLE) must be escaped to \\u0010; last read: '\"<U+0010>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x11\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0011 (DC1) must be escaped to \\u0011; last read: '\"<U+0011>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x12\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0012 (DC2) must be escaped to \\u0012; last read: '\"<U+0012>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x13\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0013 (DC3) must be escaped to \\u0013; last read: '\"<U+0013>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x14\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0014 (DC4) must be escaped to \\u0014; last read: '\"<U+0014>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x15\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0015 (NAK) must be escaped to \\u0015; last read: '\"<U+0015>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x16\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0016 (SYN) must be escaped to \\u0016; last read: '\"<U+0016>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x17\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0017 (ETB) must be escaped to \\u0017; last read: '\"<U+0017>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x18\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0018 (CAN) must be escaped to \\u0018; last read: '\"<U+0018>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x19\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0019 (EM) must be escaped to \\u0019; last read: '\"<U+0019>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x1a\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+001A (SUB) must be escaped to \\u001A; last read: '\"<U+001A>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x1b\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+001B (ESC) must be escaped to \\u001B; last read: '\"<U+001B>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x1c\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+001C (FS) must be escaped to \\u001C; last read: '\"<U+001C>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x1d\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+001D (GS) must be escaped to \\u001D; last read: '\"<U+001D>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x1e\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+001E (RS) must be escaped to \\u001E; last read: '\"<U+001E>'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\x1f\""), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+001F (US) must be escaped to \\u001F; last read: '\"<U+001F>'", json::parse_error&);
SECTION("additional test for null byte")
{
// The test above for the null byte is wrong, because passing
// a string to the parser only reads int until it encounters
// a null byte. This test inserts the null byte later on and
// uses an iterator range.
std::string s = "\"1\"";
s[1] = '\0';
json _;
CHECK_THROWS_WITH_AS(_ = json::parse(s.begin(), s.end()), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: control character U+0000 (NUL) must be escaped to \\u0000; last read: '\"<U+0000>'", json::parse_error&);
}
}
SECTION("escaped")
{
// quotation mark "\""
auto r1 = R"("\"")"_json;
CHECK(parser_helper("\"\\\"\"") == r1);
// reverse solidus "\\"
auto r2 = R"("\\")"_json;
CHECK(parser_helper("\"\\\\\"") == r2);
// solidus
CHECK(parser_helper("\"\\/\"") == R"("/")"_json);
// backspace
CHECK(parser_helper("\"\\b\"") == json("\b"));
// formfeed
CHECK(parser_helper("\"\\f\"") == json("\f"));
// newline
CHECK(parser_helper("\"\\n\"") == json("\n"));
// carriage return
CHECK(parser_helper("\"\\r\"") == json("\r"));
// horizontal tab
CHECK(parser_helper("\"\\t\"") == json("\t"));
CHECK(parser_helper("\"\\u0001\"").get<json::string_t>() == "\x01");
CHECK(parser_helper("\"\\u000a\"").get<json::string_t>() == "\n");
CHECK(parser_helper("\"\\u00b0\"").get<json::string_t>() == "°");
CHECK(parser_helper("\"\\u0c00\"").get<json::string_t>() == "ఀ");
CHECK(parser_helper("\"\\ud000\"").get<json::string_t>() == "퀀");
CHECK(parser_helper("\"\\u000E\"").get<json::string_t>() == "\x0E");
CHECK(parser_helper("\"\\u00F0\"").get<json::string_t>() == "ð");
CHECK(parser_helper("\"\\u0100\"").get<json::string_t>() == "Ā");
CHECK(parser_helper("\"\\u2000\"").get<json::string_t>() == " ");
CHECK(parser_helper("\"\\uFFFF\"").get<json::string_t>() == "￿");
CHECK(parser_helper("\"\\u20AC\"").get<json::string_t>() == "€");
CHECK(parser_helper("\"€\"").get<json::string_t>() == "€");
CHECK(parser_helper("\"🎈\"").get<json::string_t>() == "🎈");
CHECK(parser_helper("\"\\ud80c\\udc60\"").get<json::string_t>() == "\xf0\x93\x81\xa0");
CHECK(parser_helper("\"\\ud83c\\udf1e\"").get<json::string_t>() == "🌞");
}
}
SECTION("NUL byte handling (issue #5530, JSON_STRICT_NUL_HANDLING)")
{
// by default, a NUL byte anywhere in the input (not inside a quoted
// string, which is covered above) is silently treated the same as
// real end of input; JSON_STRICT_NUL_HANDLING (off by default, see
// docs/mkdocs/docs/api/macros/json_strict_nul_handling.md) makes a
// NUL byte an error like any other unexpected byte instead.
//
// The two sections below are mutually exclusive: this whole test
// binary is compiled once, with JSON_STRICT_NUL_HANDLING either
// left at its default or forced to 1 (e.g. by the dedicated
// ci_test_strict_nul_handling CI target), so only the section
// matching the actual, compiled-in behavior can pass.
SECTION("the macro is part of the ABI tag")
{
const std::string ns = JSON_TEST_STRINGIZE(NLOHMANN_JSON_NAMESPACE);
#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
CHECK(ns.find("_snul") != std::string::npos);
#else
CHECK(ns.find("_snul") == std::string::npos);
#endif
}
#if !defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
SECTION("default behavior (macro not enabled)")
{
// a NUL byte after a complete value silently truncates the input
std::string s = "123";
s.push_back('\0');
s += "4";
CHECK(json::parse(s) == json(123));
CHECK(json::accept(s));
// parsing from a string literal is unaffected either way
CHECK(json::parse("123") == json(123));
// a NUL byte that ends a // comment ends the input just
// like a NUL byte anywhere else (issue #5659); before the
// fix, the NUL was consumed as part of the comment, and
// scanning continued with whatever followed it
{
// same as "//c" alone (real end of input after the
// comment), rather than continuing with "[1]"
std::string s1 = "//c";
s1.push_back('\0');
s1 += "[1]";
json _; // NOLINT(readability-identifier-naming)
CHECK_THROWS_WITH_AS(_ = json::parse(s1, nullptr, true, true),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - unexpected end of input; expected '[', '{', or a literal",
json::parse_error&);
CHECK_FALSE(json::accept(s1, true, true));
}
{
// same as "[1, //c" alone, rather than continuing with " 2]"
std::string s2 = "[1, //c";
s2.push_back('\0');
s2 += " 2]";
json _; // NOLINT(readability-identifier-naming)
CHECK_THROWS_WITH_AS(_ = json::parse(s2, nullptr, true, true),
"[json.exception.parse_error.101] parse error at line 1, column 8: syntax error while parsing value - unexpected end of input; expected '[', '{', or a literal",
json::parse_error&);
CHECK_FALSE(json::accept(s2, true, true));
}
{
// same as "1 //c" alone: the comment (and the NUL that
// ends it) is ignored, and "x" is never reached
std::string s3 = "1 //c";
s3.push_back('\0');
s3 += "x";
CHECK(json::parse(s3, nullptr, true, true) == json(1));
CHECK(json::accept(s3, true, true));
}
}
#endif
#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
SECTION("opt-in strict behavior (JSON_STRICT_NUL_HANDLING == 1)")
{
// a NUL byte after a complete value is now a parse error,
// instead of silently truncating the input
{
std::string s = "123";
s.push_back('\0');
json _; // NOLINT(readability-identifier-naming)
CHECK_THROWS_WITH_AS(_ = json::parse(s),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: '123<U+0000>'; expected end of input",
json::parse_error&);
CHECK_FALSE(json::accept(s));
}
// a NUL byte where a value is expected is now a parse error,
// instead of being treated the same as an empty input
{
const std::string s(1, '\0');
json _; // NOLINT(readability-identifier-naming)
CHECK_THROWS_WITH_AS(_ = json::parse(s),
"[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - invalid literal; last read: '<U+0000>'",
json::parse_error&);
CHECK_FALSE(json::accept(s));
}
// a NUL byte inside a // comment no longer stops the comment
// scan early; scanning continues correctly past it
{
std::string s = "1 // a";
s.push_back('\0');
s += "b\n";
CHECK(json::parse(s, nullptr, true, true) == json(1));
CHECK(json::accept(s, true, true));
}
// a NUL byte inside a /* */ comment no longer stops the
// comment scan early either
{
std::string s = "1 /* a";
s.push_back('\0');
s += "b */ ";
CHECK(json::parse(s, nullptr, true, true) == json(1));
CHECK(json::accept(s, true, true));
}
// regression guard: parsing from a string literal (which
// carries a compiler-appended trailing '\0') still works,
// even though a NUL byte is now rejected everywhere else
CHECK(json::parse("123") == json(123));
// regression test for issue #5658: the same holds for wide,
// UTF-16, UTF-32, and (C++20) UTF-8 string literals, whose
// compiler-appended trailing '\0' is not of type `char`
CHECK(json::parse(L"[1]") == json({1}));
CHECK(json::accept(L"[1]"));
CHECK(json::parse(u"[1]") == json({1}));
CHECK(json::accept(u"[1]"));
CHECK(json::parse(U"[1]") == json({1}));
CHECK(json::accept(U"[1]"));
#if defined(__cpp_char8_t)
CHECK(json::parse(u8"[1]") == json({1}));
CHECK(json::accept(u8"[1]"));
#endif
// a NUL byte inside such a literal, as opposed to the single
// compiler-appended trailing one, is still rejected
{
json _; // NOLINT(readability-identifier-naming)
CHECK_THROWS_WITH_AS(_ = json::parse(L"[1\0]"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing array - invalid literal; last read: '1<U+0000>'; expected ']'",
json::parse_error&);
CHECK_FALSE(json::accept(L"[1\0]"));
}
}
#endif
}
SECTION("number")
{
SECTION("integers")
{
SECTION("without exponent")
{
CHECK(parser_helper("-128") == json(-128));
CHECK(parser_helper("-0") == json(-0));
CHECK(parser_helper("0") == json(0));
CHECK(parser_helper("128") == json(128));
}
SECTION("with exponent")
{
CHECK(parser_helper("0e1") == json(0e1));
CHECK(parser_helper("0E1") == json(0e1));
CHECK(parser_helper("10000E-4") == json(10000e-4));
CHECK(parser_helper("10000E-3") == json(10000e-3));
CHECK(parser_helper("10000E-2") == json(10000e-2));
CHECK(parser_helper("10000E-1") == json(10000e-1));
CHECK(parser_helper("10000E0") == json(10000e0));
CHECK(parser_helper("10000E1") == json(10000e1));
CHECK(parser_helper("10000E2") == json(10000e2));
CHECK(parser_helper("10000E3") == json(10000e3));
CHECK(parser_helper("10000E4") == json(10000e4));
CHECK(parser_helper("10000e-4") == json(10000e-4));
CHECK(parser_helper("10000e-3") == json(10000e-3));
CHECK(parser_helper("10000e-2") == json(10000e-2));
CHECK(parser_helper("10000e-1") == json(10000e-1));
CHECK(parser_helper("10000e0") == json(10000e0));
CHECK(parser_helper("10000e1") == json(10000e1));
CHECK(parser_helper("10000e2") == json(10000e2));
CHECK(parser_helper("10000e3") == json(10000e3));
CHECK(parser_helper("10000e4") == json(10000e4));
CHECK(parser_helper("-0e1") == json(-0e1));
CHECK(parser_helper("-0E1") == json(-0e1));
CHECK(parser_helper("-0E123") == json(-0e123));
// numbers after exponent
CHECK(parser_helper("10E0") == json(10e0));
CHECK(parser_helper("10E1") == json(10e1));
CHECK(parser_helper("10E2") == json(10e2));
CHECK(parser_helper("10E3") == json(10e3));
CHECK(parser_helper("10E4") == json(10e4));
CHECK(parser_helper("10E5") == json(10e5));
CHECK(parser_helper("10E6") == json(10e6));
CHECK(parser_helper("10E7") == json(10e7));
CHECK(parser_helper("10E8") == json(10e8));
CHECK(parser_helper("10E9") == json(10e9));
CHECK(parser_helper("10E+0") == json(10e0));
CHECK(parser_helper("10E+1") == json(10e1));
CHECK(parser_helper("10E+2") == json(10e2));
CHECK(parser_helper("10E+3") == json(10e3));
CHECK(parser_helper("10E+4") == json(10e4));
CHECK(parser_helper("10E+5") == json(10e5));
CHECK(parser_helper("10E+6") == json(10e6));
CHECK(parser_helper("10E+7") == json(10e7));
CHECK(parser_helper("10E+8") == json(10e8));
CHECK(parser_helper("10E+9") == json(10e9));
CHECK(parser_helper("10E-1") == json(10e-1));
CHECK(parser_helper("10E-2") == json(10e-2));
CHECK(parser_helper("10E-3") == json(10e-3));
CHECK(parser_helper("10E-4") == json(10e-4));
CHECK(parser_helper("10E-5") == json(10e-5));
CHECK(parser_helper("10E-6") == json(10e-6));
CHECK(parser_helper("10E-7") == json(10e-7));
CHECK(parser_helper("10E-8") == json(10e-8));
CHECK(parser_helper("10E-9") == json(10e-9));
}
SECTION("edge cases")
{
// From RFC8259, Section 6:
// Note that when such software is used, numbers that are
// integers and are in the range [-(2**53)+1, (2**53)-1]
// are interoperable in the sense that implementations will
// agree exactly on their numeric values.
// -(2**53)+1
CHECK(parser_helper("-9007199254740991").get<int64_t>() == -9007199254740991);
// (2**53)-1
CHECK(parser_helper("9007199254740991").get<int64_t>() == 9007199254740991);
}
SECTION("over the edge cases") // issue #178 - Integer conversion to unsigned (incorrect handling of 64-bit integers)
{
// While RFC8259, Section 6 specifies a preference for support
// for ranges in range of IEEE 754-2008 binary64 (double precision)
// this does not accommodate 64-bit integers without loss of accuracy.
// As 64-bit integers are now widely used in software, it is desirable
// to expand support to the full 64 bit (signed and unsigned) range
// i.e. -(2**63) -> (2**64)-1.
// -(2**63) ** Note: compilers see negative literals as negated positive numbers (hence the -1))
CHECK(parser_helper("-9223372036854775808").get<int64_t>() == -9223372036854775807 - 1);
// (2**63)-1
CHECK(parser_helper("9223372036854775807").get<int64_t>() == 9223372036854775807);
// (2**64)-1
CHECK(parser_helper("18446744073709551615").get<uint64_t>() == 18446744073709551615u);
}
}
SECTION("floating-point")
{
SECTION("without exponent")
{
CHECK(parser_helper("-128.5") == json(-128.5));
CHECK(parser_helper("0.999") == json(0.999));
CHECK(parser_helper("128.5") == json(128.5));
CHECK(parser_helper("-0.0") == json(-0.0));
}
SECTION("with exponent")
{
CHECK(parser_helper("-128.5E3") == json(-128.5E3));
CHECK(parser_helper("-128.5E-3") == json(-128.5E-3));
CHECK(parser_helper("-0.0e1") == json(-0.0e1));
CHECK(parser_helper("-0.0E1") == json(-0.0e1));
}
}
SECTION("overflow")
{
// overflows during parsing yield an exception
// empty() is nodiscard; the exception is thrown by parser_helper() itself, before empty() would run
CHECK_THROWS_WITH_AS(utils::ignore_return_value(parser_helper("1.18973e+4932").empty()), "[json.exception.out_of_range.406] number overflow parsing '1.18973e+4932'", json::out_of_range&);
}
SECTION("invalid numbers")
{
// numbers must not begin with "+"
CHECK_THROWS_AS(parser_helper("+1"), json::parse_error&);
CHECK_THROWS_AS(parser_helper("+0"), json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("01"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - unexpected number literal; expected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-01"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - unexpected number literal; expected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("--1"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid number; expected digit after '-'; last read: '--'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1."),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected digit after '.'; last read: '1.'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1E"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '1E'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1E-"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid number; expected digit after exponent sign; last read: '1E-'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1.E1"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected digit after '.'; last read: '1.E'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-1E"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '-1E'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0E#"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '-0E#'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0E-#"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid number; expected digit after exponent sign; last read: '-0E-#'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0#"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid literal; last read: '-0#'; expected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0.0:"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - unexpected ':'; expected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0.0Z"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: '-0.0Z'; expected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0E123:"),
"[json.exception.parse_error.101] parse error at line 1, column 7: syntax error while parsing value - unexpected ':'; expected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0e0-:"),
"[json.exception.parse_error.101] parse error at line 1, column 6: syntax error while parsing value - invalid number; expected digit after '-'; last read: '-:'; expected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0e-:"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid number; expected digit after exponent sign; last read: '-0e-:'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0f"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: '-0f'; expected end of input", json::parse_error&);
}
}
}
SECTION("accept")
{
SECTION("null")
{
CHECK(accept_helper("null"));
}
SECTION("true")
{
CHECK(accept_helper("true"));
}
SECTION("false")
{
CHECK(accept_helper("false"));
}
SECTION("array")
{
SECTION("empty array")
{
CHECK(accept_helper("[]"));
CHECK(accept_helper("[ ]"));
}
SECTION("nonempty array")
{
CHECK(accept_helper("[true, false, null]"));
}
}
SECTION("object")
{
SECTION("empty object")
{
CHECK(accept_helper("{}"));
CHECK(accept_helper("{ }"));
}
SECTION("nonempty object")
{
CHECK(accept_helper("{\"\": true, \"one\": 1, \"two\": null}"));
}
}
SECTION("string")
{
// empty string
CHECK(accept_helper("\"\""));
SECTION("errors")
{
// error: tab in string
CHECK(accept_helper("\"\t\"") == false);
// error: newline in string
CHECK(accept_helper("\"\n\"") == false);
CHECK(accept_helper("\"\r\"") == false);
// error: backspace in string
CHECK(accept_helper("\"\b\"") == false);
// improve code coverage
CHECK(accept_helper("\uFF01") == false);
CHECK(accept_helper("[-4:1,]") == false);
// unescaped control characters
CHECK(accept_helper("\"\x00\"") == false); // NOLINT(bugprone-string-literal-with-embedded-nul)
CHECK(accept_helper("\"\x01\"") == false);
CHECK(accept_helper("\"\x02\"") == false);
CHECK(accept_helper("\"\x03\"") == false);
CHECK(accept_helper("\"\x04\"") == false);
CHECK(accept_helper("\"\x05\"") == false);
CHECK(accept_helper("\"\x06\"") == false);
CHECK(accept_helper("\"\x07\"") == false);
CHECK(accept_helper("\"\x08\"") == false);
CHECK(accept_helper("\"\x09\"") == false);
CHECK(accept_helper("\"\x0a\"") == false);
CHECK(accept_helper("\"\x0b\"") == false);
CHECK(accept_helper("\"\x0c\"") == false);
CHECK(accept_helper("\"\x0d\"") == false);
CHECK(accept_helper("\"\x0e\"") == false);
CHECK(accept_helper("\"\x0f\"") == false);
CHECK(accept_helper("\"\x10\"") == false);
CHECK(accept_helper("\"\x11\"") == false);
CHECK(accept_helper("\"\x12\"") == false);
CHECK(accept_helper("\"\x13\"") == false);
CHECK(accept_helper("\"\x14\"") == false);
CHECK(accept_helper("\"\x15\"") == false);
CHECK(accept_helper("\"\x16\"") == false);
CHECK(accept_helper("\"\x17\"") == false);
CHECK(accept_helper("\"\x18\"") == false);
CHECK(accept_helper("\"\x19\"") == false);
CHECK(accept_helper("\"\x1a\"") == false);
CHECK(accept_helper("\"\x1b\"") == false);
CHECK(accept_helper("\"\x1c\"") == false);
CHECK(accept_helper("\"\x1d\"") == false);
CHECK(accept_helper("\"\x1e\"") == false);
CHECK(accept_helper("\"\x1f\"") == false);
}
SECTION("escaped")
{
// quotation mark "\""
auto r1 = R"("\"")"_json;
CHECK(accept_helper("\"\\\"\""));
// reverse solidus "\\"
auto r2 = R"("\\")"_json;
CHECK(accept_helper("\"\\\\\""));
// solidus
CHECK(accept_helper("\"\\/\""));
// backspace
CHECK(accept_helper("\"\\b\""));
// formfeed
CHECK(accept_helper("\"\\f\""));
// newline
CHECK(accept_helper("\"\\n\""));
// carriage return
CHECK(accept_helper("\"\\r\""));
// horizontal tab
CHECK(accept_helper("\"\\t\""));
CHECK(accept_helper("\"\\u0001\""));
CHECK(accept_helper("\"\\u000a\""));
CHECK(accept_helper("\"\\u00b0\""));
CHECK(accept_helper("\"\\u0c00\""));
CHECK(accept_helper("\"\\ud000\""));
CHECK(accept_helper("\"\\u000E\""));
CHECK(accept_helper("\"\\u00F0\""));
CHECK(accept_helper("\"\\u0100\""));
CHECK(accept_helper("\"\\u2000\""));
CHECK(accept_helper("\"\\uFFFF\""));
CHECK(accept_helper("\"\\u20AC\""));
CHECK(accept_helper("\"€\""));
CHECK(accept_helper("\"🎈\""));
CHECK(accept_helper("\"\\ud80c\\udc60\""));
CHECK(accept_helper("\"\\ud83c\\udf1e\""));
}
}
SECTION("number")
{
SECTION("integers")
{
SECTION("without exponent")
{
CHECK(accept_helper("-128"));
CHECK(accept_helper("-0"));
CHECK(accept_helper("0"));
CHECK(accept_helper("128"));
}
SECTION("with exponent")
{
CHECK(accept_helper("0e1"));
CHECK(accept_helper("0E1"));
CHECK(accept_helper("10000E-4"));
CHECK(accept_helper("10000E-3"));
CHECK(accept_helper("10000E-2"));
CHECK(accept_helper("10000E-1"));
CHECK(accept_helper("10000E0"));
CHECK(accept_helper("10000E1"));
CHECK(accept_helper("10000E2"));
CHECK(accept_helper("10000E3"));
CHECK(accept_helper("10000E4"));
CHECK(accept_helper("10000e-4"));
CHECK(accept_helper("10000e-3"));
CHECK(accept_helper("10000e-2"));
CHECK(accept_helper("10000e-1"));
CHECK(accept_helper("10000e0"));
CHECK(accept_helper("10000e1"));
CHECK(accept_helper("10000e2"));
CHECK(accept_helper("10000e3"));
CHECK(accept_helper("10000e4"));
CHECK(accept_helper("-0e1"));
CHECK(accept_helper("-0E1"));
CHECK(accept_helper("-0E123"));
}
SECTION("edge cases")
{
// From RFC8259, Section 6:
// Note that when such software is used, numbers that are
// integers and are in the range [-(2**53)+1, (2**53)-1]
// are interoperable in the sense that implementations will
// agree exactly on their numeric values.
// -(2**53)+1
CHECK(accept_helper("-9007199254740991"));
// (2**53)-1
CHECK(accept_helper("9007199254740991"));
}
SECTION("over the edge cases") // issue #178 - Integer conversion to unsigned (incorrect handling of 64-bit integers)
{
// While RFC8259, Section 6 specifies a preference for support
// for ranges in range of IEEE 754-2008 binary64 (double precision)
// this does not accommodate 64 bit integers without loss of accuracy.
// As 64 bit integers are now widely used in software, it is desirable
// to expand support to the full 64 bit (signed and unsigned) range
// i.e. -(2**63) -> (2**64)-1.
// -(2**63) ** Note: compilers see negative literals as negated positive numbers (hence the -1))
CHECK(accept_helper("-9223372036854775808"));
// (2**63)-1
CHECK(accept_helper("9223372036854775807"));
// (2**64)-1
CHECK(accept_helper("18446744073709551615"));
}
}
SECTION("floating-point")
{
SECTION("without exponent")
{
CHECK(accept_helper("-128.5"));
CHECK(accept_helper("0.999"));
CHECK(accept_helper("128.5"));
CHECK(accept_helper("-0.0"));
}
SECTION("with exponent")
{
CHECK(accept_helper("-128.5E3"));
CHECK(accept_helper("-128.5E-3"));
CHECK(accept_helper("-0.0e1"));
CHECK(accept_helper("-0.0E1"));
}
}
SECTION("overflow")
{
// overflows during parsing
CHECK(!accept_helper("1.18973e+4932"));
}
SECTION("invalid numbers")
{
CHECK(accept_helper("01") == false);
CHECK(accept_helper("--1") == false);
CHECK(accept_helper("1.") == false);
CHECK(accept_helper("1E") == false);
CHECK(accept_helper("1E-") == false);
CHECK(accept_helper("1.E1") == false);
CHECK(accept_helper("-1E") == false);
CHECK(accept_helper("-0E#") == false);
CHECK(accept_helper("-0E-#") == false);
CHECK(accept_helper("-0#") == false);
CHECK(accept_helper("-0.0:") == false);
CHECK(accept_helper("-0.0Z") == false);
CHECK(accept_helper("-0E123:") == false);
CHECK(accept_helper("-0e0-:") == false);
CHECK(accept_helper("-0e-:") == false);
CHECK(accept_helper("-0f") == false);
// numbers must not begin with "+"
CHECK(accept_helper("+1") == false);
CHECK(accept_helper("+0") == false);
}
SECTION("issue #5411 - skip conversion when accept() does not need the numeric value")
{
// lexer::scan_number() may skip strtoull()/strtoll() for
// value_unsigned/value_integer tokens when the caller (e.g.
// json::accept()) does not need the converted value, as long
// as the digit count alone guarantees no 64-bit overflow (see
// the "safe_digit_count" fast path in scan_number()). This
// differential test checks that json::accept() (which enables
// the fast path) and json::parse() (which never does) always
// agree, over a corpus that exercises both the fast path
// (<=18 digits) and the untouched, exact fallback path (>=19
// digits) -- including reclassification of huge digit-only
// integers to a (possibly non-finite) floating-point value.
const std::vector<std::pair<std::string, bool>> cases =
{
// normal small/large integers, both signs
{"0", true}, {"1", true}, {"-1", true}, {"42", true}, {"-42", true},
{"123456789", true}, {"-123456789", true},
// digit-count boundary around the 18-digit safe cutoff (both signs)
{std::string(17, '9'), true},
{std::string(18, '9'), true},
{std::string(19, '9'), true},
{std::string(20, '9'), true},
{"-" + std::string(17, '9'), true},
{"-" + std::string(18, '9'), true},
{"-" + std::string(19, '9'), true},
{"-" + std::string(20, '9'), true},
// 64-bit boundaries
{"9223372036854775807", true}, // INT64_MAX
{"-9223372036854775808", true}, // INT64_MIN
{"18446744073709551615", true}, // UINT64_MAX
{"18446744073709551616", true}, // UINT64_MAX + 1 (overflows uint64_t, finite double)
// the 28-digit example from the issue: overflows uint64_t
// but is finite as a double, so the scanner reclassifies
// it to value_float and it is accepted
{"9999999999999999999999999999", true},
// huge digit-only integers that overflow even a double -> rejected
{std::string(309, '9'), false},
{std::string(400, '9'), false},
{"1" + std::string(400, '0'), false},
// 1e999 / 1e400 style overflow -> rejected
{"1e999", false},
{"1e400", false},
{"-1e999", false},
{"1E999", false},
// values straddling DBL_MAX
{"1.7976931348623157e308", true}, // <= DBL_MAX, finite
{"1.7976931348623159e308", false}, // > DBL_MAX, overflows to inf
// a mix of other valid/invalid numeric syntax
{"3.14159", true},
{"-0.0", true},
{"1.0e10", true},
{"01", false},
{"-", false},
{"1.", false},
{"1e", false},
{"+1", false},
};
for (const auto& c : cases)
{
const std::string& number = c.first;
const bool expected = c.second;
CAPTURE(number)
CAPTURE(expected)
// accept() takes the fast path (skips conversion when possible)
CHECK(json::accept(number) == expected);
// parse() always performs the full conversion; it must agree
json j;
CHECK_NOTHROW(json::parser(nlohmann::detail::input_adapter(number), nullptr, false).parse(true, j));
CHECK(!j.is_discarded() == expected);
// wrap in an array so get_token() is exercised beyond the
// very first (constructor-time) scan as well
std::string wrapped = "[";
wrapped += number;
wrapped += ",";
wrapped += number;
wrapped += "]";
CHECK(json::accept(wrapped) == expected);
}
}
}
}
SECTION("parse errors")
{
// unexpected end of number
CHECK_THROWS_WITH_AS(parser_helper("0."),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected digit after '.'; last read: '0.'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid number; expected digit after '-'; last read: '-'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("--"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid number; expected digit after '-'; last read: '--'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-0."),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid number; expected digit after '.'; last read: '-0.'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-."),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid number; expected digit after '-'; last read: '-.'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("-:"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid number; expected digit after '-'; last read: '-:'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("0.:"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected digit after '.'; last read: '0.:'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("e."),
"[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - invalid literal; last read: 'e'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1e."),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '1e.'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1e/"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '1e/'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1e:"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '1e:'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1E."),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '1E.'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1E/"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '1E/'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("1E:"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid number; expected '+', '-', or digit after exponent; last read: '1E:'", json::parse_error&);
// unexpected end of null
CHECK_THROWS_WITH_AS(parser_helper("n"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid literal; last read: 'n'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("nu"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid literal; last read: 'nu'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("nul"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: 'nul'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("nulk"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: 'nulk'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("nulm"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: 'nulm'", json::parse_error&);
// unexpected end of true
CHECK_THROWS_WITH_AS(parser_helper("t"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid literal; last read: 't'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("tr"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid literal; last read: 'tr'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("tru"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: 'tru'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("trud"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: 'trud'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("truf"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: 'truf'", json::parse_error&);
// unexpected end of false
CHECK_THROWS_WITH_AS(parser_helper("f"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid literal; last read: 'f'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("fa"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid literal; last read: 'fa'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("fal"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid literal; last read: 'fal'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("fals"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: 'fals'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("falsd"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: 'falsd'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("falsf"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid literal; last read: 'falsf'", json::parse_error&);
// missing/unexpected end of array
CHECK_THROWS_WITH_AS(parser_helper("["),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - unexpected end of input; expected '[', '{', or a literal", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("[1"),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing array - unexpected end of input; expected ']'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("[1,"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - unexpected end of input; expected '[', '{', or a literal", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("[1,]"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - unexpected ']'; expected '[', '{', or a literal", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("]"),
"[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - unexpected ']'; expected '[', '{', or a literal", json::parse_error&);
// missing/unexpected end of object
CHECK_THROWS_WITH_AS(parser_helper("{"),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing object key - unexpected end of input; expected string literal", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("{\"foo\""),
"[json.exception.parse_error.101] parse error at line 1, column 7: syntax error while parsing object separator - unexpected end of input; expected ':'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("{\"foo\":"),
"[json.exception.parse_error.101] parse error at line 1, column 8: syntax error while parsing value - unexpected end of input; expected '[', '{', or a literal", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("{\"foo\":}"),
"[json.exception.parse_error.101] parse error at line 1, column 8: syntax error while parsing value - unexpected '}'; expected '[', '{', or a literal", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("{\"foo\":1,}"),
"[json.exception.parse_error.101] parse error at line 1, column 10: syntax error while parsing object key - unexpected '}'; expected string literal", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("}"),
"[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - unexpected '}'; expected '[', '{', or a literal", json::parse_error&);
// missing/unexpected end of string
CHECK_THROWS_WITH_AS(parser_helper("\""),
"[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: missing closing quote; last read: '\"'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\\""),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid string: missing closing quote; last read: '\"\\\"'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u\""),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u\"'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u0\""),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u0\"'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u01\""),
"[json.exception.parse_error.101] parse error at line 1, column 6: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u01\"'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u012\""),
"[json.exception.parse_error.101] parse error at line 1, column 7: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u012\"'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u"),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u0"),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u0'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u01"),
"[json.exception.parse_error.101] parse error at line 1, column 6: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u01'", json::parse_error&);
CHECK_THROWS_WITH_AS(parser_helper("\"\\u012"),
"[json.exception.parse_error.101] parse error at line 1, column 7: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '\"\\u012'", json::parse_error&);
// invalid escapes
for (int c = 1; c < 128; ++c)
{
auto s = std::string("\"\\") + std::string(1, static_cast<char>(c)) + "\"";
switch (c)
{
// valid escapes
case ('"'):
case ('\\'):
case ('/'):
case ('b'):
case ('f'):
case ('n'):
case ('r'):
case ('t'):
{
CHECK_NOTHROW(parser_helper(s));
break;
}
// \u must be followed with four numbers, so we skip it here
case ('u'):
{
break;
}
// any other combination of backslash and character is invalid
default:
{
CHECK_THROWS_AS(parser_helper(s), json::parse_error&);
// only check error message if c is not a control character
if (c > 0x1f)
{
CHECK_THROWS_WITH_STD_STR(parser_helper(s),
"[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid string: forbidden character after backslash; last read: '\"\\" + std::string(1, static_cast<char>(c)) + "'");
}
break;
}
}
}
// invalid \uxxxx escapes
{
// check whether character is a valid hex character
const auto valid = [](int c)
{
switch (c)
{
case ('0'):
case ('1'):
case ('2'):
case ('3'):
case ('4'):
case ('5'):
case ('6'):
case ('7'):
case ('8'):
case ('9'):
case ('a'):
case ('b'):
case ('c'):
case ('d'):
case ('e'):
case ('f'):
case ('A'):
case ('B'):
case ('C'):
case ('D'):
case ('E'):
case ('F'):
{
return true;
}
default:
{
return false;
}
}
};
for (int c = 1; c < 128; ++c)
{
std::string const s = "\"\\u";
// create a string with the iterated character at each position
auto s1 = s + "000" + std::string(1, static_cast<char>(c)) + "\"";
auto s2 = s + "00" + std::string(1, static_cast<char>(c)) + "0\"";
auto s3 = s + "0" + std::string(1, static_cast<char>(c)) + "00\"";
auto s4 = s + std::string(1, static_cast<char>(c)) + "000\"";
if (valid(c))
{
CAPTURE(s1)
CHECK_NOTHROW(parser_helper(s1));
CAPTURE(s2)
CHECK_NOTHROW(parser_helper(s2));
CAPTURE(s3)
CHECK_NOTHROW(parser_helper(s3));
CAPTURE(s4)
CHECK_NOTHROW(parser_helper(s4));
}
else
{
CAPTURE(s1)
CHECK_THROWS_AS(parser_helper(s1), json::parse_error&);
// only check error message if c is not a control character
if (c > 0x1f)
{
CHECK_THROWS_WITH_STD_STR(parser_helper(s1),
"[json.exception.parse_error.101] parse error at line 1, column 7: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '" + s1.substr(0, 7) + "'");
}
CAPTURE(s2)
CHECK_THROWS_AS(parser_helper(s2), json::parse_error&);
// only check error message if c is not a control character
if (c > 0x1f)
{
CHECK_THROWS_WITH_STD_STR(parser_helper(s2),
"[json.exception.parse_error.101] parse error at line 1, column 6: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '" + s2.substr(0, 6) + "'");
}
CAPTURE(s3)
CHECK_THROWS_AS(parser_helper(s3), json::parse_error&);
// only check error message if c is not a control character
if (c > 0x1f)
{
CHECK_THROWS_WITH_STD_STR(parser_helper(s3),
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '" + s3.substr(0, 5) + "'");
}
CAPTURE(s4)
CHECK_THROWS_AS(parser_helper(s4), json::parse_error&);
// only check error message if c is not a control character
if (c > 0x1f)
{
CHECK_THROWS_WITH_STD_STR(parser_helper(s4),
"[json.exception.parse_error.101] parse error at line 1, column 4: syntax error while parsing value - invalid string: '\\u' must be followed by 4 hex digits; last read: '" + s4.substr(0, 4) + "'");
}
}
}
}
json _;
// missing part of a surrogate pair
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD80C\""), "[json.exception.parse_error.101] parse error at line 1, column 8: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD80C\"'", json::parse_error&);
// invalid surrogate pair
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD80C\\uD80C\""),
"[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD80C\\uD80C'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD80C\\u0000\""),
"[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD80C\\u0000'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("\"\\uD80C\\uFFFF\""),
"[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing value - invalid string: surrogate U+D800..U+DBFF must be followed by U+DC00..U+DFFF; last read: '\"\\uD80C\\uFFFF'", json::parse_error&);
}
SECTION("parse errors (accept)")
{
// unexpected end of number
CHECK(accept_helper("0.") == false);
CHECK(accept_helper("-") == false);
CHECK(accept_helper("--") == false);
CHECK(accept_helper("-0.") == false);
CHECK(accept_helper("-.") == false);
CHECK(accept_helper("-:") == false);
CHECK(accept_helper("0.:") == false);
CHECK(accept_helper("e.") == false);
CHECK(accept_helper("1e.") == false);
CHECK(accept_helper("1e/") == false);
CHECK(accept_helper("1e:") == false);
CHECK(accept_helper("1E.") == false);
CHECK(accept_helper("1E/") == false);
CHECK(accept_helper("1E:") == false);
// unexpected end of null
CHECK(accept_helper("n") == false);
CHECK(accept_helper("nu") == false);
CHECK(accept_helper("nul") == false);
// unexpected end of true
CHECK(accept_helper("t") == false);
CHECK(accept_helper("tr") == false);
CHECK(accept_helper("tru") == false);
// unexpected end of false
CHECK(accept_helper("f") == false);
CHECK(accept_helper("fa") == false);
CHECK(accept_helper("fal") == false);
CHECK(accept_helper("fals") == false);
// missing/unexpected end of array
CHECK(accept_helper("[") == false);
CHECK(accept_helper("[1") == false);
CHECK(accept_helper("[1,") == false);
CHECK(accept_helper("[1,]") == false);
CHECK(accept_helper("]") == false);
// missing/unexpected end of object
CHECK(accept_helper("{") == false);
CHECK(accept_helper("{\"foo\"") == false);
CHECK(accept_helper("{\"foo\":") == false);
CHECK(accept_helper("{\"foo\":}") == false);
CHECK(accept_helper("{\"foo\":1,}") == false);
CHECK(accept_helper("}") == false);
// missing/unexpected end of string
CHECK(accept_helper("\"") == false);
CHECK(accept_helper("\"\\\"") == false);
CHECK(accept_helper("\"\\u\"") == false);
CHECK(accept_helper("\"\\u0\"") == false);
CHECK(accept_helper("\"\\u01\"") == false);
CHECK(accept_helper("\"\\u012\"") == false);
CHECK(accept_helper("\"\\u") == false);
CHECK(accept_helper("\"\\u0") == false);
CHECK(accept_helper("\"\\u01") == false);
CHECK(accept_helper("\"\\u012") == false);
// unget of newline
CHECK(parser_helper("\n123\n") == 123);
// invalid escapes
for (int c = 1; c < 128; ++c)
{
auto s = std::string("\"\\") + std::string(1, static_cast<char>(c)) + "\"";
switch (c)
{
// valid escapes
case ('"'):
case ('\\'):
case ('/'):
case ('b'):
case ('f'):
case ('n'):
case ('r'):
case ('t'):
{
CHECK(json::parser(nlohmann::detail::input_adapter(s)).accept());
break;
}
// \u must be followed with four numbers, so we skip it here
case ('u'):
{
break;
}
// any other combination of backslash and character is invalid
default:
{
CHECK(json::parser(nlohmann::detail::input_adapter(s)).accept() == false);
break;
}
}
}
// invalid \uxxxx escapes
{
// check whether character is a valid hex character
const auto valid = [](int c)
{
switch (c)
{
case ('0'):
case ('1'):
case ('2'):
case ('3'):
case ('4'):
case ('5'):
case ('6'):
case ('7'):
case ('8'):
case ('9'):
case ('a'):
case ('b'):
case ('c'):
case ('d'):
case ('e'):
case ('f'):
case ('A'):
case ('B'):
case ('C'):
case ('D'):
case ('E'):
case ('F'):
{
return true;
}
default:
{
return false;
}
}
};
for (int c = 1; c < 128; ++c)
{
std::string const s = "\"\\u";
// create a string with the iterated character at each position
const auto s1 = s + "000" + std::string(1, static_cast<char>(c)) + "\"";
const auto s2 = s + "00" + std::string(1, static_cast<char>(c)) + "0\"";
const auto s3 = s + "0" + std::string(1, static_cast<char>(c)) + "00\"";
const auto s4 = s + std::string(1, static_cast<char>(c)) + "000\"";
if (valid(c))
{
CAPTURE(s1)
CHECK(json::parser(nlohmann::detail::input_adapter(s1)).accept());
CAPTURE(s2)
CHECK(json::parser(nlohmann::detail::input_adapter(s2)).accept());
CAPTURE(s3)
CHECK(json::parser(nlohmann::detail::input_adapter(s3)).accept());
CAPTURE(s4)
CHECK(json::parser(nlohmann::detail::input_adapter(s4)).accept());
}
else
{
CAPTURE(s1)
CHECK(json::parser(nlohmann::detail::input_adapter(s1)).accept() == false);
CAPTURE(s2)
CHECK(json::parser(nlohmann::detail::input_adapter(s2)).accept() == false);
CAPTURE(s3)
CHECK(json::parser(nlohmann::detail::input_adapter(s3)).accept() == false);
CAPTURE(s4)
CHECK(json::parser(nlohmann::detail::input_adapter(s4)).accept() == false);
}
}
}
// missing part of a surrogate pair
CHECK(accept_helper("\"\\uD80C\"") == false);
// invalid surrogate pair
CHECK(accept_helper("\"\\uD80C\\uD80C\"") == false);
CHECK(accept_helper("\"\\uD80C\\u0000\"") == false);
CHECK(accept_helper("\"\\uD80C\\uFFFF\"") == false);
}
#if !defined(JSON_NOEXCEPTION)
SECTION("issue #5412 - whitespace skipping bookkeeping (compact vs. pretty-printed)")
{
// lexer::skip_whitespace() reads its first character with get() (to
// honor a possibly pending unget() from the previous token) and every
// further whitespace character with get_ignoring_pending_unget() (a
// get() variant that skips the then-always-false next_unget check).
// This must not change the reported byte offset, line, or column of
// a syntax error, even when a long run of whitespace containing
// multiple newlines is skipped beforehand (as with pretty-printed
// input). The expected values below were captured from the
// unmodified do-while(get()) loop, so any regression that miscounts
// characters or newlines while skipping whitespace changes them.
const auto check_error = [](const std::string & input, std::size_t expected_byte,
const std::string & expected_what)
{
CAPTURE(input)
try
{
json _ = json::parse(input);
FAIL_CHECK("expected a parse_error, but parsing succeeded");
}
catch (const json::parse_error& e)
{
CHECK(e.byte == expected_byte);
CHECK(std::string(e.what()) == expected_what);
}
};
// a nested document, serialized both compactly and pretty-printed
// (dump(4)), each truncated right before the final closing '}' so
// that the parser hits EOF after skipping all of the (in the
// pretty-printed case, substantial) indentation whitespace
const json doc =
{
{"a", 1},
{"b", json::array({true, false, nullptr, "x"})},
{"c", json::object({{"d", 3.14}, {"e", json::array({1, 2, 3})}})}
};
const std::string compact = doc.dump();
const std::string pretty = doc.dump(4);
check_error(compact.substr(0, compact.size() - 1), 60,
"[json.exception.parse_error.101] parse error at line 1, column 60: syntax error while parsing object - unexpected end of input; expected '}'");
check_error(pretty.substr(0, pretty.size() - 1), 193,
"[json.exception.parse_error.101] parse error at line 17, column 1: syntax error while parsing object - unexpected end of input; expected '}'");
// an invalid token appearing after several indented, multi-line
// whitespace runs vs. the same document without any of that
// whitespace
check_error(R"({
"a": 1,
"b": [
true,
false
],
"c": @
})", 70,
"[json.exception.parse_error.101] parse error at line 7, column 10: syntax error while parsing value - invalid literal; last read: '\"c\": @'");
check_error(R"({"a":1,"b":[true,false],"c":@})", 29,
"[json.exception.parse_error.101] parse error at line 1, column 29: syntax error while parsing value - invalid literal; last read: '\"c\":@'");
}
#endif
SECTION("tests found by mutate++")
{
// test case to make sure no comma precedes the first key
CHECK_THROWS_WITH_AS(parser_helper("{,\"key\": false}"), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing object key - unexpected ','; expected string literal", json::parse_error&);
// test case to make sure an object is properly closed
CHECK_THROWS_WITH_AS(parser_helper("[{\"key\": false true]"), "[json.exception.parse_error.101] parse error at line 1, column 19: syntax error while parsing object - unexpected true literal; expected '}'", json::parse_error&);
// test case to make sure the callback is properly evaluated after reading a key
{
json::parser_callback_t const cb = [](int /*unused*/, json::parse_event_t event, json& /*unused*/) noexcept
{
return event != json::parse_event_t::key;
};
const json x = json::parse("{\"key\": false}", cb);
CHECK(x == json::object());
}
}
SECTION("callback function")
{
const auto* s_object = R"(
{
"foo": 2,
"bar": {
"baz": 1
}
}
)";
const auto* s_array = R"(
[1,2,[3,4,5],4,5]
)";
const auto* structured_array = R"(
[
1,
{
"foo": "bar"
},
{
"qux": "baz"
}
]
)";
const auto* structured_object = R"(
{
"foo": [1, 2],
"bar": 3
}
)";
SECTION("filter nothing")
{
const json j_object = json::parse(s_object, [](int /*unused*/, json::parse_event_t /*unused*/, const json& /*unused*/) noexcept
{
return true;
});
CHECK (j_object == json({{"foo", 2}, {"bar", {{"baz", 1}}}}));
const json j_array = json::parse(s_array, [](int /*unused*/, json::parse_event_t /*unused*/, const json& /*unused*/) noexcept
{
return true;
});
CHECK (j_array == json({1, 2, {3, 4, 5}, 4, 5}));
}
SECTION("filter everything")
{
json const j_object = json::parse(s_object, [](int /*unused*/, json::parse_event_t /*unused*/, const json& /*unused*/) noexcept
{
return false;
});
// the top-level object will be discarded, leaving a null
CHECK (j_object.is_null());
json const j_array = json::parse(s_array, [](int /*unused*/, json::parse_event_t /*unused*/, const json& /*unused*/) noexcept
{
return false;
});
// the top-level array will be discarded, leaving a null
CHECK (j_array.is_null());
}
SECTION("filter specific element")
{
const json j_object = json::parse(s_object, [](int /*unused*/, json::parse_event_t event, const json & j) noexcept
{
// filter all number(2) elements
return event != json::parse_event_t::value || j != json(2);
});
CHECK (j_object == json({{"bar", {{"baz", 1}}}}));
const json j_array = json::parse(s_array, [](int /*unused*/, json::parse_event_t event, const json & j) noexcept
{
return event != json::parse_event_t::value || j != json(2);
});
CHECK (j_array == json({1, {3, 4, 5}, 4, 5}));
}
SECTION("filter object in array")
{
const json j_filtered1 = json::parse(structured_array, [](int /*unused*/, json::parse_event_t e, const json & parsed)
{
return !(e == json::parse_event_t::object_end && parsed.contains("foo"));
});
// the specified object will be discarded, and removed.
CHECK (j_filtered1.size() == 2);
CHECK (j_filtered1 == json({1, {{"qux", "baz"}}}));
const json j_filtered2 = json::parse(structured_array, [](int /*unused*/, json::parse_event_t e, const json& /*parsed*/) noexcept
{
return e != json::parse_event_t::object_end;
});
// removed all objects in array.
CHECK (j_filtered2.size() == 1);
CHECK (j_filtered2 == json({1}));
}
SECTION("filter array in object")
{
// the array is discarded once it is already stored under its key
const json j_filtered1 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json& /*parsed*/) noexcept
{
return e != json::parse_event_t::array_end;
});
CHECK (j_filtered1 == json({{"bar", 3}}));
// the array is discarded before it is stored, leaving the
// placeholder the key event wrote
const json j_filtered2 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json& /*parsed*/) noexcept
{
return e != json::parse_event_t::array_start;
});
CHECK (j_filtered2 == json({{"bar", 3}}));
}
SECTION("filter value in object")
{
// the value is discarded after its key was kept, leaving the
// placeholder the key event wrote
const json j_filtered1 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json & parsed) noexcept
{
return !(e == json::parse_event_t::value && parsed == json(3));
});
CHECK (j_filtered1 == json({{"foo", {1, 2}}}));
// the same value is discarded together with its key, so no
// placeholder was stored for it
const json j_filtered2 = json::parse(structured_object, [](int /*unused*/, json::parse_event_t e, const json & parsed) noexcept
{
return !((e == json::parse_event_t::key && parsed == json("bar")) ||
(e == json::parse_event_t::value && parsed == json(3)));
});
CHECK (j_filtered2 == json({{"foo", {1, 2}}}));
}
SECTION("filter many members of one container")
{
// Rejecting a value makes the parser remove the placeholder its key
// event stored. Locating that placeholder used to be a scan of the
// whole parent, which made filtering a large container quadratic:
// 128k members took ~25 s. These cases keep many members alive
// while discarding many others, so the removal cost is the whole
// point; they run in milliseconds when the placeholder is erased
// directly.
constexpr int count = 20000;
std::string s = "{";
for (int i = 0; i < count; ++i)
{
// "a<i>" is kept, "z<i>" is discarded
s += "\"a" + std::to_string(i) + "\":" + std::to_string(i) + ",";
s += "\"z" + std::to_string(i) + "\":-1,";
}
s.back() = '}';
const json j_values = json::parse(s, [](int /*unused*/, json::parse_event_t e, const json & parsed) noexcept
{
return !(e == json::parse_event_t::value && parsed == json(-1));
});
CHECK(j_values.size() == count);
CHECK(j_values.at("a0") == json(0));
CHECK(j_values.at("a" + std::to_string(count - 1)) == json(count - 1));
CHECK_FALSE(j_values.contains("z0"));
CHECK_FALSE(j_values.contains("z" + std::to_string(count - 1)));
// the same, but discarding whole containers rather than values,
// which takes the end_object()/end_array() removal path
std::string s_nested = "{";
for (int i = 0; i < count; ++i)
{
s_nested += "\"a" + std::to_string(i) + "\":" + std::to_string(i) + ",";
s_nested += "\"z" + std::to_string(i) + "\":[1,2],";
}
s_nested.back() = '}';
const json j_arrays = json::parse(s_nested, [](int /*unused*/, json::parse_event_t e, const json& /*unused*/) noexcept
{
return e != json::parse_event_t::array_end;
});
CHECK(j_arrays.size() == count);
CHECK(j_arrays.at("a0") == json(0));
CHECK_FALSE(j_arrays.contains("z0"));
CHECK_FALSE(j_arrays.contains("z" + std::to_string(count - 1)));
}
SECTION("filter specific events")
{
SECTION("first closing event")
{
{
const json j_object = json::parse(s_object, [](int /*unused*/, json::parse_event_t e, const json& /*unused*/) noexcept
{
static bool first = true;
if (e == json::parse_event_t::object_end && first)
{
first = false;
return false;
}
return true;
});
// the first completed object will be discarded
CHECK (j_object == json({{"foo", 2}}));
}
{
const json j_array = json::parse(s_array, [](int /*unused*/, json::parse_event_t e, const json& /*unused*/) noexcept
{
static bool first = true;
if (e == json::parse_event_t::array_end && first)
{
first = false;
return false;
}
return true;
});
// the first completed array will be discarded
CHECK (j_array == json({1, 2, 4, 5}));
}
}
}
SECTION("no callback for the content of a discarded container (#5643)")
{
// discarding a container at its start event must also hide
// everything inside it from the callback: none of the nested
// keys, values, or nested containers' own start/end events may
// be reported
std::vector<std::string> log;
bool first = true;
const json j = json::parse(R"({"skip": {"k1": 1, "k2": [2, {"k3": 3}]}, "keep": 1})",
[&](int depth, json::parse_event_t event, json & parsed)
{
static const char* const names[] = {"object_start", "object_end", "array_start", "array_end", "key", "value"}; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
log.push_back(std::to_string(depth) + " " + names[static_cast<int>(event)] + " " + parsed.dump());
if (depth == 1 && event == json::parse_event_t::object_start && first)
{
// discard "skip" right at its object_start event
first = false;
return false;
}
return true;
});
CHECK(log == std::vector<std::string>
{
"0 object_start <discarded>",
"1 key \"skip\"",
"1 object_start <discarded>",
"1 key \"keep\"",
"1 value 1",
"0 object_end {\"keep\":1}"
});
CHECK(j == json({{"keep", 1}}));
}
SECTION("callback still called inside a container whose key was rejected (#5643)")
{
// rejecting a key does not discard its value's container at the
// container's own start event, so the callback is still called
// for that container's content; only storing the container
// under the rejected key is skipped
// (documented for parser_callback_t: "the callback is still
// called for the associated value, but its return value has no
// further effect")
const auto record = [](std::vector<std::string>& log, int depth, json::parse_event_t event, const json & parsed)
{
static const char* const names[] = {"object_start", "object_end", "array_start", "array_end", "key", "value"}; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
log.push_back(std::to_string(depth) + " " + names[static_cast<int>(event)] + " " + parsed.dump());
};
std::vector<std::string> log_object;
const json j_object = json::parse(R"({"skip": {"k1": 1}, "keep": 2})",
[&](int depth, json::parse_event_t event, json & parsed)
{
record(log_object, depth, event, parsed);
return !(event == json::parse_event_t::key && parsed == json("skip"));
});
CHECK(log_object == std::vector<std::string>
{
"0 object_start <discarded>",
"1 key \"skip\"",
"1 object_start <discarded>",
"2 key \"k1\"",
"2 value 1",
"1 key \"keep\"",
"1 value 2",
"0 object_end {\"keep\":2}"
});
CHECK(j_object == json({{"keep", 2}}));
// same for a rejected key whose value is an array rather than an object
std::vector<std::string> log_array;
const json j_array = json::parse(R"({"skip": [1, {"k1": 2}], "keep": 2})",
[&](int depth, json::parse_event_t event, json & parsed)
{
record(log_array, depth, event, parsed);
return !(event == json::parse_event_t::key && parsed == json("skip"));
});
CHECK(log_array == std::vector<std::string>
{
"0 object_start <discarded>",
"1 key \"skip\"",
"1 array_start <discarded>",
"2 value 1",
"2 object_start <discarded>",
"3 key \"k1\"",
"3 value 2",
"1 key \"keep\"",
"1 value 2",
"0 object_end {\"keep\":2}"
});
CHECK(j_array == json({{"keep", 2}}));
}
SECTION("special cases")
{
// the following test cases cover the situation in which an empty
// object and array is discarded only after the closing character
// has been read
const json j_empty_object = json::parse("{}", [](int /*unused*/, json::parse_event_t e, const json& /*unused*/) noexcept
{
return e != json::parse_event_t::object_end;
});
CHECK(j_empty_object == json());
const json j_empty_array = json::parse("[]", [](int /*unused*/, json::parse_event_t e, const json& /*unused*/) noexcept
{
return e != json::parse_event_t::array_end;
});
CHECK(j_empty_array == json());
}
}
SECTION("constructing from contiguous containers")
{
SECTION("from std::vector")
{
std::vector<uint8_t> v = {'t', 'r', 'u', 'e'};
json j;
json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j);
CHECK(j == json(true));
}
SECTION("from std::array")
{
// NOTE: this array is sized to exactly the length of "true" (unlike
// the trailing-NUL-tolerant default behavior elsewhere in this file,
// see the "NUL byte handling" section above); a size of 5 here would
// leave a value-initialized trailing 0x00 element that is only
// silently accepted as end-of-input by default and would fail under
// JSON_STRICT_NUL_HANDLING
std::array<uint8_t, 4> v { {'t', 'r', 'u', 'e'} };
json j;
json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j);
CHECK(j == json(true));
}
SECTION("from array")
{
uint8_t v[] = {'t', 'r', 'u', 'e'}; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
json j;
json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j);
CHECK(j == json(true));
}
SECTION("from char literal")
{
CHECK(parser_helper("true") == json(true));
}
SECTION("from std::string")
{
std::string v = {'t', 'r', 'u', 'e'};
json j;
json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j);
CHECK(j == json(true));
}
SECTION("from std::initializer_list")
{
std::initializer_list<uint8_t> const v = {'t', 'r', 'u', 'e'};
json j;
json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j);
CHECK(j == json(true));
}
SECTION("from std::valarray")
{
std::valarray<uint8_t> v = {'t', 'r', 'u', 'e'};
json j;
json::parser(nlohmann::detail::input_adapter(std::begin(v), std::end(v))).parse(true, j);
CHECK(j == json(true));
}
}
SECTION("improve test coverage")
{
SECTION("parser with callback")
{
json::parser_callback_t const cb = [](int /*unused*/, json::parse_event_t /*unused*/, json& /*unused*/) noexcept
{
return true;
};
CHECK(json::parse("{\"foo\": true:", cb, false).is_discarded());
json _;
CHECK_THROWS_WITH_AS(_ = json::parse("{\"foo\": true:", cb), "[json.exception.parse_error.101] parse error at line 1, column 13: syntax error while parsing object - unexpected ':'; expected '}'", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::parse("1.18973e+4932", cb), "[json.exception.out_of_range.406] number overflow parsing '1.18973e+4932'", json::out_of_range&);
}
SECTION("SAX parser")
{
SECTION("} without value")
{
SaxCountdown s(1);
CHECK(json::sax_parse("{}", &s) == false);
}
SECTION("} with value")
{
SaxCountdown s(3);
CHECK(json::sax_parse("{\"k1\": true}", &s) == false);
}
SECTION("second key")
{
SaxCountdown s(3);
CHECK(json::sax_parse("{\"k1\": true, \"k2\": false}", &s) == false);
}
SECTION("] without value")
{
SaxCountdown s(1);
CHECK(json::sax_parse("[]", &s) == false);
}
SECTION("] with value")
{
SaxCountdown s(2);
CHECK(json::sax_parse("[1]", &s) == false);
}
SECTION("float")
{
SaxCountdown s(0);
CHECK(json::sax_parse("3.14", &s) == false);
}
SECTION("false")
{
SaxCountdown s(0);
CHECK(json::sax_parse("false", &s) == false);
}
SECTION("null")
{
SaxCountdown s(0);
CHECK(json::sax_parse("null", &s) == false);
}
SECTION("true")
{
SaxCountdown s(0);
CHECK(json::sax_parse("true", &s) == false);
}
SECTION("unsigned")
{
SaxCountdown s(0);
CHECK(json::sax_parse("12", &s) == false);
}
SECTION("integer")
{
SaxCountdown s(0);
CHECK(json::sax_parse("-12", &s) == false);
}
SECTION("string")
{
SaxCountdown s(0);
CHECK(json::sax_parse("\"foo\"", &s) == false);
}
}
}
SECTION("error messages for comments")
{
json _;
CHECK_THROWS_WITH_AS(_ = json::parse("/a", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid comment; expecting '/' or '*' after '/'; last read: '/a'", json::parse_error);
// "/*" is a string literal, so it carries a compiler-appended trailing
// '\0'; by default that NUL is read like any other byte and shows up
// in "last read", but JSON_STRICT_NUL_HANDLING trims exactly that one
// trailing byte from a char array (see
// docs/mkdocs/docs/api/macros/json_strict_nul_handling.md), so it no
// longer appears in the message in that state
#if defined(JSON_TEST_STRICT_NUL_HANDLING_ENABLED)
CHECK_THROWS_WITH_AS(_ = json::parse("/*", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid comment; missing closing '*/'; last read: '/*'", json::parse_error);
#else
CHECK_THROWS_WITH_AS(_ = json::parse("/*", nullptr, true, true), "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid comment; missing closing '*/'; last read: '/*<U+0000>'", json::parse_error);
#endif
}
#if JSON_DIAGNOSTIC_POSITIONS
// Macro for all test cases for start_pos and end_pos
#define SETUP_TESTCASES() \
SECTION("with callback") \
{ \
SECTION("filter nothing") \
{ \
json::parser_callback_t const cb = [](int /*unused*/, json::parse_event_t /*unused*/, json& /*unused*/) noexcept \
{ \
return true; \
}; \
validate_start_end_pos_for_nested_obj_helper(nested_type_json_str, root_type_json_str, expected, cb); \
} \
SECTION("filter element") \
{ \
json::parser_callback_t const cb = [](int /*unused*/, json::parse_event_t event, json& j) noexcept \
{ \
return (event != json::parse_event_t::key && event != json::parse_event_t::value) || j != json("a"); \
}; \
validate_start_end_pos_for_nested_obj_helper(nested_type_json_str, root_type_json_str, filteredExpected, cb); \
} \
} \
SECTION("without callback") \
{ \
validate_start_end_pos_for_nested_obj_helper(nested_type_json_str, root_type_json_str, expected); \
}
SECTION("retrieve start position and end position")
{
SECTION("for object")
{
// Create an object with spaces to test the start and end positions. Spaces will not be included in the
// JSON object, however, the start and end positions should include the spaces from the input JSON string.
const std::string nested_type_json_str = R"({ "a": 1,"b" : "test1"})";
const std::string root_type_json_str = R"({ "nested": )" + nested_type_json_str + R"(, "anotherValue": "test2"})";
auto expected = json({{"nested", {{"a", 1}, {"b", "test1"}}}, {"anotherValue", "test2"}});
auto filteredExpected = expected;
filteredExpected["nested"].erase("a");
SETUP_TESTCASES()
}
SECTION("for array")
{
const std::string nested_type_json_str = R"(["a", "test", 45])";
const std::string root_type_json_str = R"({ "nested": )" + nested_type_json_str + R"(, "anotherValue": "test" })";
auto expected = json({{"nested", {"a", "test", 45}}, {"anotherValue", "test"}});
auto filteredExpected = expected;
filteredExpected["nested"] = json({"test", 45});
SETUP_TESTCASES()
}
SECTION("for array with objects")
{
const std::string nested_type_json_str = R"([{"a": 1, "b": "test"}, {"c": 2, "d": "test2"}])";
const std::string root_type_json_str = R"({ "nested": )" + nested_type_json_str + R"(, "anotherValue": "test" })";
auto expected = json({{"nested", {{{"a", 1}, {"b", "test"}}, {{"c", 2}, {"d", "test2"}}}}, {"anotherValue", "test"}});
auto filteredExpected = expected;
filteredExpected["nested"][0].erase("a");
SETUP_TESTCASES()
auto j = json::parse(root_type_json_str);
auto nested_array = j["nested"];
const auto& nested_obj = nested_array[0];
CHECK(nested_type_json_str.substr(1, 21) == root_type_json_str.substr(nested_obj.start_pos(), nested_obj.end_pos() - nested_obj.start_pos()));
CHECK(nested_type_json_str.substr(24, 22) == root_type_json_str.substr(nested_array[1].start_pos(), nested_array[1].end_pos() - nested_array[1].start_pos()));
}
SECTION("for two levels of nesting objects")
{
const std::string nested_type_json_str = R"({"nested2": {"b": "test"}})";
const std::string root_type_json_str = R"({ "a": 2, "nested": )" + nested_type_json_str + R"(, "anotherValue": "test" })";
auto expected = json({{"a", 2}, {"nested", {{"nested2", {{"b", "test"}}}}}, {"anotherValue", "test"}});
auto filteredExpected = expected;
filteredExpected.erase("a");
SETUP_TESTCASES()
auto j = json::parse(root_type_json_str);
auto nested_obj = j["nested"]["nested2"];
CHECK(nested_type_json_str.substr(12, 13) == root_type_json_str.substr(nested_obj.start_pos(), nested_obj.end_pos() - nested_obj.start_pos()));
}
SECTION("for simple types")
{
SECTION("no nested")
{
SECTION("with callback")
{
json::parser_callback_t const cb = [](int /*unused*/, json::parse_event_t /*unused*/, json& /*unused*/) noexcept
{
return true;
};
// 1. string type
std::string json_str = R"("test")";
auto j = json::parse(json_str, cb);
validate_generated_json_and_start_end_pos_helper(json_str, j, "test");
// 2. number type
json_str = R"(1)";
j = json::parse(json_str, cb);
validate_generated_json_and_start_end_pos_helper(json_str, j, 1);
// 3. boolean type
json_str = R"(true)";
j = json::parse(json_str, cb);
validate_generated_json_and_start_end_pos_helper(json_str, j, true);
// 4. null type
json_str = R"(null)";
j = json::parse(json_str, cb);
validate_generated_json_and_start_end_pos_helper(json_str, j, nullptr);
}
SECTION("without callback")
{
// 1. string type
std::string json_str = R"("test")";
auto j = json::parse(json_str);
validate_generated_json_and_start_end_pos_helper(json_str, j, "test");
// 2. number type
json_str = R"(1)";
j = json::parse(json_str);
validate_generated_json_and_start_end_pos_helper(json_str, j, 1);
json_str = R"(1.001239923)";
j = json::parse(json_str);
validate_generated_json_and_start_end_pos_helper(json_str, j, 1.001239923);
json_str = R"(1.123812389000000)";
j = json::parse(json_str);
validate_generated_json_and_start_end_pos_helper(json_str, j, 1.123812389);
// 3. boolean type
json_str = R"(true)";
j = json::parse(json_str);
validate_generated_json_and_start_end_pos_helper(json_str, j, true);
json_str = R"(false)";
j = json::parse(json_str);
validate_generated_json_and_start_end_pos_helper(json_str, j, false);
// 4. null type
json_str = R"(null)";
j = json::parse(json_str);
validate_generated_json_and_start_end_pos_helper(json_str, j, nullptr);
}
}
SECTION("string type")
{
const std::string nested_type_json_str = R"("test")";
const std::string root_type_json_str = R"({ "a": 1, "nested": )" + nested_type_json_str + R"(, "anotherValue": "test" })";
auto expected = json({{"nested", "test"}, {"anotherValue", "test"}, {"a", 1}});
auto filteredExpected = expected;
filteredExpected.erase("a");
SETUP_TESTCASES()
}
SECTION("number type")
{
const std::string nested_type_json_str = R"(2)";
const std::string root_type_json_str = R"({ "a": 1, "nested": )" + nested_type_json_str + R"(, "anotherValue": "test" })";
auto expected = json({{"nested", 2}, {"anotherValue", "test"}, {"a", 1}});
auto filteredExpected = expected;
filteredExpected.erase("a");
SETUP_TESTCASES()
}
SECTION("boolean type")
{
const std::string nested_type_json_str = R"(true)";
const std::string root_type_json_str = R"({ "a": 1, "nested": )" + nested_type_json_str + R"(, "anotherValue": "test" })";
auto expected = json({{"nested", true}, {"anotherValue", "test"}, {"a", 1}});
auto filteredExpected = expected;
filteredExpected.erase("a");
SETUP_TESTCASES()
}
SECTION("null type")
{
const std::string nested_type_json_str = R"(null)";
const std::string root_type_json_str = R"({ "a": 1, "nested": )" + nested_type_json_str + R"(, "anotherValue": "test" })";
auto expected = json({{"nested", nullptr}, {"anotherValue", "test"}, {"a", 1}});
auto filteredExpected = expected;
filteredExpected.erase("a");
SETUP_TESTCASES()
}
}
SECTION("with leading whitespace and newlines around root JSON")
{
const std::string initial_whitespace = R"(
)";
const std::string nested_type_json_str = R"({
"a": 1,
"nested": {
"b": "test"
},
"anotherValue": "test"
})";
const std::string end_whitespace = R"(
)";
const std::string root_type_json_str = initial_whitespace + nested_type_json_str + end_whitespace;
auto expected = json({{"a", 1}, {"nested", {{"b", "test"}}}, {"anotherValue", "test"}});
auto j = json::parse(root_type_json_str);
// 2. Check if the generated JSON is as expected
CHECK(j == expected);
// 3. Check if the start and end positions do not include the surrounding whitespace
CHECK(j.start_pos() == initial_whitespace.size());
CHECK(j.end_pos() == root_type_json_str.size() - end_whitespace.size());
}
}
#undef SETUP_TESTCASES
#endif
}
// this test relies on parse errors being thrown, so it is skipped when
// exceptions are disabled (json::parse aborts instead of throwing there)
#if !defined(JSON_NOEXCEPTION)
namespace
{
// Return the exception message from parsing @a input, or a "<no error ...>"
// sentinel if the parse unexpectedly succeeds. json::parse is nodiscard, so the
// result is consumed (via size()) to keep -Wunused-result / -Werror happy.
template<typename InputType>
std::string parse_error_message(InputType&& input)
{
try
{
const json j = json::parse(std::forward<InputType>(input));
return "<no error, size " + std::to_string(j.size()) + ">";
}
catch (const json::exception& e)
{
return e.what();
}
}
template<typename IteratorType>
std::string parse_error_message_range(IteratorType first, IteratorType last)
{
try
{
const json j = json::parse(first, last);
return "<no error, size " + std::to_string(j.size()) + ">";
}
catch (const json::exception& e)
{
return e.what();
}
}
} // namespace
TEST_CASE("last-read diagnostics are identical across input adapters")
{
// The lexer reconstructs the "last read" token lazily for seekable adapters
// (contiguous byte input) and copies it eagerly for streaming adapters.
// Both strategies must yield byte-for-byte identical error messages.
// a selection of malformed inputs that exercise different token kinds,
// whitespace/structural accumulation, number overflow, and control-char
// escaping in the reconstructed "last read" token
const std::vector<std::string> inputs =
{
"[1,2,x]",
" \n @",
"{\"a\": }",
"1.18973e+4932",
"\"\t\"",
"tru",
"[1 2]",
"\xEF\xBB\xBF nul",
};
for (const auto& s : inputs)
{
CAPTURE(s)
// reference: contiguous std::string -> seekable (lazy) path
const std::string reference = parse_error_message(s);
// every input is malformed, so parsing must fail (error messages start
// with '['; the success sentinel returned above starts with '<')
CHECK(reference.front() == '[');
// const char* -> also seekable
CHECK(parse_error_message(s.c_str()) == reference);
// std::vector<char> iterators -> seekable (random-access)
{
const std::vector<char> v(s.begin(), s.end());
CHECK(parse_error_message_range(v.begin(), v.end()) == reference);
}
// std::list iterators -> non-seekable (bidirectional) eager path
{
const std::list<char> l(s.begin(), s.end());
CHECK(parse_error_message_range(l.begin(), l.end()) == reference);
}
// std::istringstream -> non-seekable streaming eager path
{
std::istringstream ss(s);
CHECK(parse_error_message(ss) == reference);
}
// wide strings -> wide_string_input_adapter eager path; only comparable
// for ASCII input, as non-ASCII bytes are transcoded to different UTF-8
const bool is_ascii = std::all_of(s.begin(), s.end(), [](char c)
{
return static_cast<unsigned char>(c) < 0x80;
});
if (is_ascii)
{
const std::u16string w16(s.begin(), s.end());
CHECK(parse_error_message(w16) == reference);
const std::u32string w32(s.begin(), s.end());
CHECK(parse_error_message(w32) == reference);
}
}
}
#endif // !defined(JSON_NOEXCEPTION)
// this test characterizes the current (documented-by-example, not otherwise
// specified) behavior of JSON_DIAGNOSTIC_POSITIONS positions with respect to
// value lifetime (copy/move/swap/mutation), the various input adapters, and
// user-driven SAX usage. It is regression protection, not a behavior
// specification: if any of these checks fail after a change to json.hpp,
// that change deliberately altered observable behavior and the test (and
// this comment) should be updated accordingly, rather than "fixed" blindly.
#if JSON_DIAGNOSTIC_POSITIONS
TEST_CASE("diagnostic positions: value lifetime, input adapters, and SAX")
{
SECTION("value lifetime")
{
SECTION("copy constructor copies positions, recursively")
{
// basic_json(const basic_json&) (json.hpp, around line 1192) copies
// start_position/end_position for the value itself; nested values
// are copied via their own copy constructor (through the copied
// object/array container), so positions are preserved throughout
// the whole tree.
const std::string s = R"({"a":1,"b":[1,2,3]})";
const json a = json::parse(s);
const json b = a; // NOLINT(performance-unnecessary-copy-initialization)
CHECK(b.start_pos() == a.start_pos());
CHECK(b.end_pos() == a.end_pos());
CHECK(b["b"].start_pos() == a["b"].start_pos());
CHECK(b["b"].end_pos() == a["b"].end_pos());
CHECK(b["b"][0].start_pos() == a["b"][0].start_pos());
CHECK(b["b"][0].end_pos() == a["b"][0].end_pos());
// sanity: the positions are meaningful (not all npos)
CHECK(b.start_pos() == 0);
CHECK(b.end_pos() == s.size());
}
SECTION("move constructor resets the moved-from value to npos")
{
// basic_json(basic_json&&) copies
// other's start_position/end_position into *this and then resets
// other's to npos. Only the top-level moved-from value is
// affected; its (moved-away) children are gone along with it.
const std::string s = R"({"a":1,"b":[1,2,3]})";
json a = json::parse(s);
const auto a_start = a.start_pos();
const auto a_end = a.end_pos();
const auto nested_start = a["b"].start_pos();
const auto nested_end = a["b"].end_pos();
const json b(std::move(a));
// the destination retains the original positions, recursively
CHECK(b.start_pos() == a_start);
CHECK(b.end_pos() == a_end);
CHECK(b["b"].start_pos() == nested_start);
CHECK(b["b"].end_pos() == nested_end);
// the moved-from value is reset to a null and reports npos
CHECK(a.is_null()); // NOLINT(bugprone-use-after-move,clang-analyzer-cplusplus.Move)
CHECK(a.start_pos() == std::string::npos); // NOLINT(bugprone-use-after-move,clang-analyzer-cplusplus.Move)
CHECK(a.end_pos() == std::string::npos); // NOLINT(bugprone-use-after-move,clang-analyzer-cplusplus.Move)
}
SECTION("swap() exchanges positions along with values")
{
// basic_json::swap() (json.hpp, around line 3626, and the friend
// swap() that forwards to it) swaps start_position/end_position
// together with m_data.m_type and m_data.m_value, so after
// swap(a, b) each variable's position describes its own new
// content, consistent with copy-assignment's
// operator=(basic_json) (json.hpp, around line 1291), which also
// swaps positions as part of its copy-and-swap implementation.
json a = json::parse(R"({"a":1})");
json b = json::parse(R"([1,2,3,4,5])");
const auto a_start = a.start_pos();
const auto a_end = a.end_pos();
const auto b_start = b.start_pos();
const auto b_end = b.end_pos();
// both start at 0 (root values start right away), but their
// lengths (and thus end positions) differ, which is enough to
// tell after the swap whether positions actually moved with
// the values
CHECK(a_end != b_end);
using std::swap;
swap(a, b);
// values were exchanged as expected ...
CHECK(a == json::parse(R"([1,2,3,4,5])"));
CHECK(b == json::parse(R"({"a":1})"));
// ... and so were positions: each variable now carries the
// other's original position, describing its own new content
CHECK(a.start_pos() == b_start);
CHECK(a.end_pos() == b_end);
CHECK(b.start_pos() == a_start);
CHECK(b.end_pos() == a_end);
// member swap() behaves the same as the free function
json c = json::parse(R"({"a":1})");
json d = json::parse(R"([1,2,3,4,5])");
const auto c_start = c.start_pos();
const auto c_end = c.end_pos();
const auto d_start = d.start_pos();
const auto d_end = d.end_pos();
c.swap(d);
CHECK(c.start_pos() == d_start);
CHECK(c.end_pos() == d_end);
CHECK(d.start_pos() == c_start);
CHECK(d.end_pos() == c_end);
}
SECTION("mutating a parsed document leaves positions of unrelated values untouched")
{
// Positions are recorded once, during parsing, and are not
// recomputed on mutation. As a consequence, after a mutation the
// parent's own recorded span may no longer describe its current
// (serialized) content -- it still describes what was originally
// parsed. This is characterized here as current behavior, not
// asserted to be desirable or specified.
SECTION("operator[] adding a new object key")
{
const std::string s = R"({"a":1})";
json j = json::parse(s);
const auto root_start = j.start_pos();
const auto root_end = j.end_pos();
const auto a_start = j["a"].start_pos();
const auto a_end = j["a"].end_pos();
j["c"] = 42;
// the newly-added value was never parsed, so it has no position
CHECK(j["c"].start_pos() == std::string::npos);
CHECK(j["c"].end_pos() == std::string::npos);
// the existing sibling's position is unaffected
CHECK(j["a"].start_pos() == a_start);
CHECK(j["a"].end_pos() == a_end);
// the parent's own recorded span is left as-is (now stale:
// it still reflects the original, shorter `{"a":1}` string)
CHECK(j.start_pos() == root_start);
CHECK(j.end_pos() == root_end);
}
SECTION("push_back on a parsed array")
{
const std::string s = R"([1,2,3])";
json j = json::parse(s);
const auto root_start = j.start_pos();
const auto root_end = j.end_pos();
const auto first_start = j[0].start_pos();
j.push_back(4);
CHECK(j.back().start_pos() == std::string::npos);
CHECK(j.back().end_pos() == std::string::npos);
CHECK(j[0].start_pos() == first_start);
CHECK(j.start_pos() == root_start);
CHECK(j.end_pos() == root_end);
}
SECTION("erase on a parsed array shifts elements but keeps their own positions")
{
const std::string s = R"([1,2,3])";
json j = json::parse(s);
const auto second_start = j[1].start_pos();
const auto third_start = j[2].start_pos();
const auto root_start = j.start_pos();
const auto root_end = j.end_pos();
j.erase(0);
// remaining elements moved down an index, but each one still
// reports the position it had *before* the erase (i.e. its
// position in the original source string, not a
// recalculated one)
CHECK(j[0].start_pos() == second_start);
CHECK(j[1].start_pos() == third_start);
// the parent's own recorded span is again left as-is
CHECK(j.start_pos() == root_start);
CHECK(j.end_pos() == root_end);
}
}
}
SECTION("input adapters")
{
SECTION("wide string input: positions count transcoded UTF-8 bytes, not wide characters")
{
// 'é' (U+00E9) is a single code unit in a wchar_t/UTF-16 string, but
// transcodes to 2 bytes in UTF-8; the lexer only ever sees the
// transcoded UTF-8 byte stream, so reported positions are byte
// offsets into that UTF-8 stream, not indices into the original
// std::wstring.
// é (rather than a literal 'é' byte sequence in this source
// file) so the wide-string literal's meaning does not depend on
// the compiler's assumed source character set (MSVC, without
// /utf-8, would otherwise decode the raw UTF-8 bytes using the
// system code page instead of as UTF-8)
const std::wstring ws = L"{\"a\":\"\u00e9\u00e9\"}";
CHECK(ws.size() == 10); // 10 wide characters
const json j = json::parse(ws);
CHECK(j.start_pos() == 0);
// the transcoded UTF-8 form is 2 bytes longer than the wide string,
// because each of the two 'é' characters becomes 2 UTF-8 bytes
CHECK(j.end_pos() == 12);
CHECK(j.end_pos() != ws.size());
const json& a = j["a"];
CHECK(a.start_pos() == 5);
CHECK(a.end_pos() == 11);
}
SECTION("BOM-prefixed input: start_pos() reflects the skipped 3-byte BOM")
{
const std::string s = "\xEF\xBB\xBF{\"a\":1}";
const json j = json::parse(s);
// the lexer silently skips the BOM before parsing the value, so
// the root value's recorded span starts right after it
CHECK(j.start_pos() == 3);
CHECK(j.end_pos() == s.size());
}
SECTION("std::istringstream: positions are consistent, not npos")
{
const std::string s = R"({"a":1,"b":2})";
std::istringstream ss(s);
const json j = json::parse(ss);
CHECK(j.start_pos() == 0);
CHECK(j.end_pos() == s.size());
CHECK(j["a"].start_pos() == 5);
}
SECTION("std::ifstream: positions are consistent, not npos")
{
const std::string s = R"({"a":1,"b":2})";
{
std::ofstream file("unit-class_parser_diagnostic_positions.tmp");
file << s;
}
{
std::ifstream f("unit-class_parser_diagnostic_positions.tmp");
const json j = json::parse(f);
CHECK(j.start_pos() == 0);
CHECK(j.end_pos() == s.size());
CHECK(j["a"].start_pos() == 5);
}
static_cast<void>(std::remove("unit-class_parser_diagnostic_positions.tmp"));
}
SECTION("iterator-pair input: positions are consistent, not npos")
{
const std::string s = R"({"a":1,"b":2})";
const json j = json::parse(s.begin(), s.end());
CHECK(j.start_pos() == 0);
CHECK(j.end_pos() == s.size());
CHECK(j["a"].start_pos() == 5);
}
SECTION("binary formats have no text positions")
{
// binary formats (BJData, BON8, BSON, CBOR, MessagePack, UBJSON) are
// parsed via detail::binary_reader, which never sets
// start_position/end_position on the values it produces (they
// have no notion of a text offset), so every value's position
// stays at its default of npos.
const json src = json::parse(R"({"a":1,"b":[1,2]})");
const json from_cbor = json::from_cbor(json::to_cbor(src));
CHECK(from_cbor.start_pos() == std::string::npos);
CHECK(from_cbor.end_pos() == std::string::npos);
CHECK(from_cbor["a"].start_pos() == std::string::npos);
CHECK(from_cbor["b"][0].start_pos() == std::string::npos);
const json from_msgpack = json::from_msgpack(json::to_msgpack(src));
CHECK(from_msgpack.start_pos() == std::string::npos);
CHECK(from_msgpack.end_pos() == std::string::npos);
const json from_bon8 = json::from_bon8(json::to_bon8(src));
CHECK(from_bon8.start_pos() == std::string::npos);
CHECK(from_bon8.end_pos() == std::string::npos);
const json from_ubjson = json::from_ubjson(json::to_ubjson(src));
CHECK(from_ubjson.start_pos() == std::string::npos);
CHECK(from_ubjson.end_pos() == std::string::npos);
const json from_bson_val = json::from_bson(json::to_bson(src));
CHECK(from_bson_val.start_pos() == std::string::npos);
CHECK(from_bson_val.end_pos() == std::string::npos);
}
}
SECTION("user-driven SAX consumers with no lexer report npos")
{
// json::parse() internally wires up its json_sax_dom_parser with a
// pointer to its own lexer (see parser.hpp), which is how positions
// get set at all. A user who constructs a json_sax_dom_parser
// directly (e.g. to drive it via json::sax_parse()) and does not
// supply a lexer pointer gets a consumer with m_lexer_ref == nullptr;
// every "if (m_lexer_ref)" guard in json_sax.hpp is then skipped, so
// every value it produces keeps its default, unset position (npos).
// This was previously true but silently unasserted (operator==
// ignores positions), see #5420.
json result;
nlohmann::detail::json_sax_dom_parser<json, nlohmann::detail::string_input_adapter_type> sdp(result);
const std::string s = R"({"a":1,"b":[1,2,3]})";
CHECK(json::sax_parse(s, &sdp));
CHECK(result.start_pos() == std::string::npos);
CHECK(result.end_pos() == std::string::npos);
CHECK(result["a"].start_pos() == std::string::npos);
CHECK(result["a"].end_pos() == std::string::npos);
CHECK(result["b"][0].start_pos() == std::string::npos);
CHECK(result["b"][0].end_pos() == std::string::npos);
}
}
#endif