Files
json/include/nlohmann/detail/input/input_adapters.hpp
T
Niels Lohmann a0b71e2720 Deduplicate basic_json internals; make insert(pos, json&&) move (#5727)
* Remove unused private aliases from basic_json

The private aliases primitive_iterator_t, internal_iterator and
output_adapter_t are not used anywhere: iter_impl, binary_writer and
the tests refer to the detail:: names directly. As the aliases are
private, no user or derived class can depend on them. The
internal_iterator.hpp include stays because iter_impl needs it.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix meta()'s dead, syntactically invalid HP aCC branch

The HP aCC branch of basic_json::meta() was missing a semicolon and
has therefore never compiled; adding only a semicolon would also make
it throw type_error.305, since it assigned a plain string to
result["compiler"] and then indexed into it like the other branches
do into an object. Make the branch consistent with the others by
assigning an object with "family" and "version" keys, narrow the
condition to __HP_aCC (a C compiler cannot build this header-only
library), and fix meta.md, which documented the old (impossible)
plain-string behavior.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Stop the noexcept null constructor delegating to a throwing one

basic_json(std::nullptr_t) delegated to basic_json(value_t), whose
underlying json_value(value_t) constructor allocates for other types
and can therefore throw, which is why the noexcept had a
NOLINT(bugprone-exception-escape). The delegated-to constructor also
called assert_invariant() a second time. The default member
initializers of data already produce the same null state (a
value-initialized, i.e. zeroed, union with object == nullptr), so the
delegation and its NOLINT can simply be dropped.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove four cppcheck accessForwarded suppressions in the move constructor

basic_json(basic_json&&) built its base subobject with
std::forward<json_base_class_t>(other), so cppcheck saw the whole of
other as forwarded and flagged every subsequent access to it as
accessForwarded, three of them still marked "TODO check". Only the
base subobject is actually moved from; cast explicitly to the base
type instead, the way ordered_map already does, so cppcheck can tell
the two are unrelated. Behavior is unchanged: for a non-reference T,
std::forward<T>(x) is defined as static_cast<T&&>(x).

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Remove stale cppcheck suppressions and name the local parser in parse()

Running the pinned cppcheck (ci_cppcheck's invocation) without
--inline-suppr across all configurations reports no syntaxError, no
ignoredReturnValue and no assertWithSideEffect, so the corresponding
suppressions in json_fwd.hpp, string_concat.hpp and
assert_invariant() no longer match anything (json_fwd.hpp's is kept,
since downstream users who run an older cppcheck against it could
still hit the warning it once silenced).

The three basic_json::parse() overloads still trigger a false-positive
accessMoved/accessForwarded because they build a temporary parser and
call .parse() on it in the same expression; giving that parser a name
makes the warning go away without changing behavior, and removes the
last of the inline suppressions on these functions.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the string/binary cleanup in the two erase() overloads

erase(pos) and erase(first, last) each carried a byte-identical
14-line block that destroys and deallocates a string or binary
value before resetting the type to null. That reimplements the
string/binary cases of json_value::destroy(), so any future change to
how those values are freed would have to be made in three places
instead of one. Both overloads now just call destroy() and reset the
union; for the other primitive types (boolean, numbers) destroy() is
a no-op, so behavior is unchanged.

Also fix erase(first, last)'s error-path branch hint, which used
JSON_HEDLEY_LIKELY where erase(pos), the iterator-range constructor,
and every other error path in the class use JSON_HEDLEY_UNLIKELY.
This only affects code layout, not semantics.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale and copy-pasted comments in basic_json

Several comments no longer match the code: the class invariant and
assert_invariant()'s doc still named the members m_value/m_type
(now m_data.m_value/m_data.m_type) and did not mention the binary
invariant that assert_invariant() already checks; the json_value note
and the get<PointerType>() @tparam list omitted binary_t even though
binary is a variable-length, pointer-stored type like the others; the
key-based value() overload's brief said "via JSON Pointer", which is
the other overload; and swap(binary_t&)/swap(binary_t::container_type&)
both carried "swap only works for strings", copied from swap(string_t&).

Comment-only change; behavior, the public API and the ABI are
unchanged. The private get_impl() doxygen and the emplace() comments
that border #5585's hunk are intentionally left alone.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the 16 copied from_cbor/msgpack/ubjson/bjdata/bon8/bson bodies

Each of the 16 binary deserialization overloads (from_cbor,
from_msgpack, from_ubjson, from_bjdata, from_bon8, from_bson, each in
an InputType&& and an iterator/sentinel version, plus the deprecated
span overloads of from_cbor/from_msgpack/from_ubjson/from_bson) had
the same body, differing only in the input_format_t value. Every copy
built a temporary binary_reader from std::move(ia) and called
sax_parse on it in the same expression, which also produced a
false-positive cppcheck accessMoved on all 16 lines and needed a
NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) on
the four span overloads.

Add a private from_binary_impl() helper that builds the reader as a
named local instead, and make each of the 16 overloads a one-line
forward to it. All public signatures, default arguments,
JSON_HEDLEY_WARN_UNUSED_RESULT and JSON_HEDLEY_DEPRECATED_FOR
attributes are unchanged, tag_handler keeps defaulting to
cbor_tag_handler_t::error for the non-CBOR formats (matching
binary_reader::sax_parse's own default), and the helper is placed in
the existing private section before the binary section banner rather
than between the from_* overloads, so from_binary_impl() itself does
not collide with #5688's insertion point. Collapsing the
from_bjdata/from_bon8 bodies into one-line forwards does rewrite the
"return result; }" context lines that #5688 inserts its two
deprecated overloads after, so that PR will need a small manual
rebase (reinserting its overloads after the new one-line bodies)
rather than applying cleanly.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make insert(pos, basic_json&&) move its argument instead of copying it

insert(const_iterator pos, basic_json&& val) delegated to
insert(pos, val), but val is a named rvalue reference, so inside the
function it is an lvalue: the call always resolved to
insert(const_iterator, const basic_json&) and deep-copied the value.
This has been the case since the overload was introduced, in every
release. push_back(basic_json&&), by contrast, already moves.

Give the rvalue overload its own body with the same two checks
(type_error.309, invalid_iterator.202), then move the argument into a
local before inserting it. Moving into a local first, rather than
inserting std::move(val) directly, keeps this safe even when val
aliases an element of the same array (e.g.
arr.insert(arr.begin(), std::move(arr[1]))), since
std::vector::insert(pos, T&&) is not guaranteed to handle an argument
that aliases one of its own elements.

This is a deliberate, small behavior change: the moved-from argument
now ends up null afterwards, the same as after push_back(&&), instead
of keeping its old value unchanged. No signature changes, so the
public API and ABI are unaffected. Add unit-modifiers coverage for
the moved-from state and for self-aliasing insertion, both with and
without reallocation of the underlying array.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate object key lookup and checked at() access

at()/find()/count()/contains()/const operator[]/erase_internal() each
repeated the raw object lookup (m_value.object->find(key)), and the
six at() overloads additionally repeated the type_error.304 check and
out_of_range.401/403 throw. Route them all through two new private
helpers, object_lookup()/object_at() (plus array_at() for the index
overloads of at()), templated on the constness of the receiver so one
body serves both the const and non-const overload. count() is left
untouched, since it already goes through object_t::count() rather
than a second find().

The at(KeyType&&) overloads used to forward the same key twice: once
into object->find() and again, on the not-found path, into the
string_t() conversion for the exception message. object_at() now
forwards it only into the lookup and reuses the (unmoved) key for the
message. clang-tidy 22 (Docker silkeh/clang:22) still flags that reuse
under bugprone-use-after-move/hicpp-invalid-access-moved even with the
single forward, since it cannot see that object_t::find() (a plain
std::map or ordered_map) never actually moves from its argument; add a
NOLINTNEXTLINE with that reasoning rather than avoid the pattern.

No signature, exception id/message, or set_parent() behavior changes.

Overlaps #5689, #5705, #5606, #5687 and #5585, which touch the same
hunks; whichever of this commit and those PRs lands second will need
a small rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the lookup/default/throw body of value()

The six non-deprecated value() overloads each held a full copy of the
same body: the four key-based overloads looked up the key and either
returned the found element converted to the requested type or the
default value (throwing type_error.306 if this is not an object), and
the two json_pointer overloads did the same via
ptr.get_checked_or_null(), throwing type_error.306 unless
is_structured(). Replace the duplicated bodies with two private
helpers, value_member() and value_pointee(), that return a
const basic_json* (null when not found) and do the type check/throw
once each. Every value() overload now just picks between the found
pointer's get<T>() and the default.

Same signatures, template parameters, SFINAE conditions, exception id,
message and this context on every overload.

Overlaps #5689 (routes find() through lookup_key()) and #5705 (adds a
deleted integral-key value() next to these overloads); whichever of
this commit and those PRs lands second will need a small rebase.

#5724 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the null-to-container conversion into convert_null_to()

Nine sites wrote out the same "turn a null value into an empty array
or object" logic with two different idioms: operator[](size_type),
operator[](key_type), operator[](KeyType&&) and update() set m_type
then assigned m_value.array/object directly via create<T>(), while the
three push_back() overloads, emplace_back() and emplace() set m_type
then assigned m_value = value_t::array/object (going through
json_value's converting constructor and a temporary). Both idioms end
up calling create<T>() and produce the same state, just via a
different path; both also share a latent exception-safety bug, since
m_type is written before the (possibly throwing) allocation, so a
throwing allocator leaves m_type == array/object with a null pointer
behind it, violating the class invariant and crashing on the next
access to, or destruction of, the value.

Add a private convert_null_to(value_t) helper and call it from all
nine sites. Unlike the idioms it replaces, it allocates the container
first and only then writes m_type, so a throwing allocation leaves the
value as a valid null instead of a mistyped, half-constructed one;
verified with a throwing allocator (see unit-allocator.cpp's
bad_allocator) that j["x"] = ... on a null j now stays null, and no
longer trips assert_invariant()/crashes, when create<object_t>()
throws. Same allocator usage and assert_invariant() call as before,
otherwise.

Overlaps #5585, which reorders these same nine blocks for exception
safety; whichever of this commit and that PR lands second will need a
small rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the linear key search in ordered_map

emplace, at, erase(key), count and find each repeated the same
"for (auto it = begin(); it != end(); ++it) if (m_compare(it->first,
key)) ..." loop (15 copies across their key_type and transparent
KeyType&& overloads), and both erase(key) overloads additionally
repeated the exception-sensitive in-place reconstruction (destroy,
placement-new, pop_back) used to remove an element while keeping the
const Key non-movable.

Add two private helpers: find_impl(Self&, KeyType&&), a static member
template that runs the search once for either constness of the
receiver, and erase_at(iterator), which keeps the existing
pop_back-based reconstruction instead of switching to erase()/resize()
(which would add a DefaultInsertable requirement). Route find, at,
count, emplace, insert(const value_type&) and both erase(key)
overloads through them.

Same signatures, is_usable_as_key_type constraints and exception
messages/types.

Overlaps #5609 and #5685, which both rewrite emplace (#5609 also
touches insert and adds private members at the end of the class);
whichever of this commit and those PRs lands second will need a
rebase.

#5724 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Re-enable bugprone-use-after-move/hicpp-invalid-access-moved

These two checks (and portability-template-virtual-member-function)
were disabled in #4489 (November 2024) "only removed to get the CI
going". portability-template-virtual-member-function is a separate,
still-open cleanup (#5725 item 3 on its own branch) and stays
disabled here; this commit only re-enables the move/forward checks
and cleans up what they flag on this branch.

The move constructor (json.hpp) already casts to the base type
instead of forwarding the whole object (#5724 item 9), so it no
longer trips either check. The at(KeyType&&) double-forward this
check used to flag was reduced to a single forward with the
now-unforwarded reuse annotated by a NOLINTNEXTLINE in #5724 item 3's
object_at() helper (clang-tidy 22 still flags that reuse even after a
single forward; see that commit's message). What is left here:

- from_json_inplace_array_impl(), from_json_tuple_impl_base() and the
  std::pair overload of from_json_tuple_impl() forwarded j into every
  j.at(...) call in a pack expansion or a pair of calls. at() has no
  ref-qualified overloads, so the forward was a no-op; call j.at(...)
  directly.
- container_input_adapter_factory::create() forwards container twice
  on purpose, into begin() and end(), so both see the same value
  category and produce matching iterator types. Annotate it with
  NOLINTNEXTLINE and a comment instead of changing it.
- unit-class_parser.cpp's "move constructor resets the moved-from
  value to npos" test still pointed at the pre-static_cast move
  constructor by line number and mentioned the cppcheck-suppress
  annotation that #5724 item 9 already removed; update the comment.

No behavior change anywhere in include/. Verified with clang-tidy 22.1.8
(Docker silkeh/clang:22, --platform linux/amd64) against a TU including
json.hpp with the repo's .clang-tidy: bugprone-use-after-move and
hicpp-invalid-access-moved report nothing unsuppressed.

Overlaps #5737 (open PR for the rest of #5725 item 3: the
from_json.hpp/input_adapters.hpp cleanup above, and
portability-template-virtual-member-function), which currently keeps
both checks disabled pending this move-constructor change; whichever
of this commit and that PR lands second will need a small rebase of
.clang-tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Make convert_null_to() take the container type as a template argument

Passing array_t or object_t instead of a value_t makes an invalid target
a compile error instead of a runtime assertion, and removes the branch.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-04 11:57:41 +02:00

898 lines
36 KiB
C++

// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <algorithm> // min
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint32_t
#include <cstring> // strlen
#include <iterator> // begin, end, iterator_traits, random_access_iterator_tag, distance, next
#include <streambuf> // streambuf
#include <string> // string, char_traits
#include <type_traits> // enable_if, is_base_of, is_pointer, is_integral, remove_pointer
#include <utility> // pair, declval
#ifndef JSON_NO_IO
#include <cstdio> // FILE *
#include <istream> // istream
#endif // JSON_NO_IO
#include <nlohmann/detail/exceptions.hpp>
#include <nlohmann/detail/iterators/iterator_traits.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/string_utils.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/// the supported input formats
enum class input_format_t { json, cbor, msgpack, ubjson, bson, bjdata, bon8 };
////////////////////
// input adapters //
////////////////////
#ifndef JSON_NO_IO
/*!
Input adapter for stdio file access. This adapter read only 1 byte and do not use any
buffer. This adapter is a very low level adapter.
*/
class file_input_adapter
{
public:
using char_type = char;
JSON_HEDLEY_NON_NULL(2)
explicit file_input_adapter(std::FILE* f) noexcept
: m_file(f)
{
JSON_ASSERT(m_file != nullptr);
}
// make class move-only
file_input_adapter(const file_input_adapter&) = delete;
file_input_adapter(file_input_adapter&&) noexcept = default;
file_input_adapter& operator=(const file_input_adapter&) = delete;
file_input_adapter& operator=(file_input_adapter&&) = delete;
~file_input_adapter() = default;
std::char_traits<char>::int_type get_character() noexcept
{
return std::fgetc(m_file);
}
// returns the number of characters successfully read
template<class T>
std::size_t get_elements(T* dest, std::size_t count = 1)
{
return fread(dest, 1, sizeof(T) * count, m_file);
}
private:
/// the file pointer to read from
std::FILE* m_file;
};
/*!
Input adapter for a (caching) istream. Does not skip a UTF Byte Order Mark
itself; that is done by the lexer's skip_bom(). Does not support changing
the underlying std::streambuf
in mid-input. Maintains underlying std::istream and std::streambuf to support
subsequent use of standard std::istream operations to process any input
characters following those used in parsing the JSON input. Clears the
std::istream flags; any input errors (e.g., EOF) will be detected by the first
subsequent call for input from the std::istream.
*/
class input_stream_adapter
{
public:
using char_type = char;
~input_stream_adapter()
{
// clear stream flags; we use underlying streambuf I/O, do not
// maintain ifstream flags, except eof
if (is != nullptr)
{
#if JSON_PRECISE_STREAM_POSITION
// consume the character last returned by get_character() unless it
// was given back with release_lookahead()
commit_lookahead();
#endif
// only call clear() if there is something to clear: it throws
// std::ios_base::failure if the stream has exceptions() enabled
// for a state bit that remains set, and a destructor must not throw
if ((is->rdstate() & ~std::ios::eofbit) != 0)
{
is->clear(is->rdstate() & std::ios::eofbit);
}
}
}
explicit input_stream_adapter(std::istream& i)
: is(&i), sb(i.rdbuf())
{}
// deleted because of pointer members
input_stream_adapter(const input_stream_adapter&) = delete;
input_stream_adapter& operator=(input_stream_adapter&) = delete;
input_stream_adapter& operator=(input_stream_adapter&&) = delete;
#if JSON_PRECISE_STREAM_POSITION
input_stream_adapter(input_stream_adapter&& rhs) noexcept
: is(rhs.is), sb(rhs.sb), lookahead(rhs.lookahead)
{
rhs.is = nullptr;
rhs.sb = nullptr;
rhs.lookahead = false;
}
// Whether the character last returned by get_character() can be given back
// to the input with release_lookahead().
static constexpr bool supports_lookahead = true;
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
// end up as the same value, e.g., 0xFFFFFFFF.
//
// The character is peeked rather than consumed: it is only stepped over
// once the next character is requested, or when the adapter is destroyed.
// Until then, release_lookahead() can leave it in the input.
std::char_traits<char>::int_type get_character()
{
if (lookahead)
{
// step over the character returned by the previous call
sb->sbumpc();
}
auto res = sb->sgetc();
// set eof manually, as we don't use the istream interface.
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
{
// there is nothing to step over next time
lookahead = false;
is->clear(is->rdstate() | std::ios::eofbit);
}
else
{
lookahead = true;
}
return res;
}
// Leave the character last returned by get_character() in the input, so
// that the next read from the stream - by this adapter or by the caller
// once parsing is done - sees it again. Unlike putting a consumed
// character back, this cannot fail.
void release_lookahead() noexcept
{
lookahead = false;
}
#else
input_stream_adapter(input_stream_adapter&& rhs) noexcept
: is(rhs.is), sb(rhs.sb)
{
rhs.is = nullptr;
rhs.sb = nullptr;
}
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
// end up as the same value, e.g., 0xFFFFFFFF.
//
// The character is consumed, so the character that terminates a number
// stays consumed after parsing; see JSON_PRECISE_STREAM_POSITION.
std::char_traits<char>::int_type get_character()
{
auto res = sb->sbumpc();
// set eof manually, as we don't use the istream interface.
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
{
is->clear(is->rdstate() | std::ios::eofbit);
}
return res;
}
#endif
template<class T>
std::size_t get_elements(T* dest, std::size_t count = 1)
{
#if JSON_PRECISE_STREAM_POSITION
commit_lookahead();
#endif
auto res = static_cast<std::size_t>(sb->sgetn(reinterpret_cast<char*>(dest), static_cast<std::streamsize>(count * sizeof(T))));
if (JSON_HEDLEY_UNLIKELY(res < count * sizeof(T)))
{
is->clear(is->rdstate() | std::ios::eofbit);
}
return res;
}
private:
#if JSON_PRECISE_STREAM_POSITION
// Step over the character last returned by get_character(). The character
// has already been peeked successfully, so for every streambuf with a get
// area this is a pointer increment that cannot fail.
void commit_lookahead()
{
if (lookahead)
{
lookahead = false;
sb->sbumpc();
}
}
#endif
/// the associated input stream
std::istream* is = nullptr;
std::streambuf* sb = nullptr;
#if JSON_PRECISE_STREAM_POSITION
/// whether get_character() peeked a character that is not consumed yet
bool lookahead = false;
#endif
};
#endif // JSON_NO_IO
// General-purpose iterator-based adapter. It might not be as fast as
// theoretically possible for some containers, but it is extremely versatile.
// SentinelType defaults to IteratorType for backward compatibility, but may be
// a different type, e.g. a C++20 sentinel such as std::default_sentinel_t when
// IteratorType is a std::counted_iterator.
template<typename IteratorType, typename SentinelType = IteratorType>
class iterator_input_adapter
{
// Whether the number of elements between two positions can be computed in
// O(1): either the iterator and the sentinel have the same type (plain
// std::distance) or, in C++20, the sentinel is a sized sentinel for the
// iterator (std::ranges::distance), e.g. std::default_sentinel_t paired
// with std::counted_iterator.
//
// JSON_HAS_RANGES gates the C++20 branch: on standard libraries with an
// incomplete <ranges> (libstdc++ < 11, see #4440) evaluating
// std::contiguous_iterator on a std::counted_iterator is a hard error
// instead of yielding false, and these traits are instantiated for every
// adapter. Such toolchains fall back to the pointer-only test and simply
// use the byte-at-a-time scanner.
static constexpr bool sentinel_is_sized =
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
std::is_same<IteratorType, SentinelType>::value || std::sized_sentinel_for<SentinelType, IteratorType>;
#else
std::is_same<IteratorType, SentinelType>::value;
#endif
public:
using char_type = typename std::iterator_traits<IteratorType>::value_type;
// Whether the lexer may reconstruct already-consumed input on demand (for
// diagnostics) instead of copying every scanned character eagerly. This is
// only sound for multi-pass, randomly-addressable byte input: the iterator
// must be random-access (so the consumed prefix can be revisited in O(1))
// and each element must map 1:1 to an input byte (wide inputs are wrapped
// in wide_string_input_adapter, which does not expose this).
static constexpr bool supports_seek =
std::is_same<typename std::iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value
&& sentinel_is_sized
&& sizeof(char_type) == 1;
iterator_input_adapter(IteratorType first, SentinelType last)
: begin(first), current(std::move(first)), end(std::move(last))
{}
typename char_traits<char_type>::int_type get_character()
{
if (JSON_HEDLEY_LIKELY(current != end))
{
auto result = char_traits<char_type>::to_int_type(*current);
std::advance(current, 1);
return result;
}
return char_traits<char_type>::eof();
}
// number of characters consumed from the input so far
std::size_t get_consumed_count() const
{
return static_cast<std::size_t>(std::distance(begin, current));
}
// append the already-consumed characters in the half-open range
// [first_index, last_index) to @a out; only valid when supports_seek
template<typename ContainerType>
void copy_consumed_range(std::size_t first_index, std::size_t last_index, ContainerType& out) const
{
const auto from = std::next(begin, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(first_index));
const auto to = std::next(begin, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(last_index));
out.insert(out.end(), from, to);
}
// Copy up to count * sizeof(T) bytes into dest, returning the number of
// bytes actually read. For contiguous iterators (e.g. pointers) this is a
// single std::memcpy; for general iterators we fall back to processing the
// range one-by-one.
template<class T>
std::size_t get_elements(T* dest, std::size_t count = 1)
{
return get_elements_impl(dest, count, std::integral_constant<bool, iterator_is_contiguous> {});
}
private:
// whether IteratorType refers to a contiguous range and therefore supports
// a std::memcpy fast path (pointers always do; in C++20 we can also detect
// library iterators such as those of std::vector and std::string). The
// available element count must also be computable in O(1), hence
// sentinel_is_sized.
static constexpr bool iterator_is_contiguous = sentinel_is_sized &&
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
(std::contiguous_iterator<IteratorType> || std::is_pointer<IteratorType>::value);
#else
std::is_pointer<IteratorType>::value;
#endif
// number of unread elements in [current, end)
std::size_t remaining_count() const
{
#if JSON_HAS_RANGES && defined(__cpp_lib_concepts) && defined(JSON_HAS_CPP_20)
// std::ranges::distance also supports sized sentinels of a different
// type (e.g. std::counted_iterator + std::default_sentinel_t)
return static_cast<std::size_t>(std::ranges::distance(current, end));
#else
return static_cast<std::size_t>(std::distance(current, end));
#endif
}
public:
// Whether the remaining input is a single contiguous block of 1-byte
// elements that the lexer can inspect directly (used for the SWAR string
// fast path).
static constexpr bool supports_bulk_scan =
iterator_is_contiguous && sizeof(char_type) == 1;
// Pointer to the next unread element; only valid when bulk_remaining() > 0.
const char_type* bulk_data() const
{
return &*current;
}
// Number of unread elements available as one contiguous block.
std::size_t bulk_remaining() const
{
return remaining_count();
}
// Consume @a n elements previously inspected via bulk_data().
void bulk_skip(std::size_t n)
{
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(n));
}
private:
// contiguous fast path: bulk copy the remaining range with std::memcpy
template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
{
const std::size_t wanted = count * sizeof(T);
const std::size_t available = remaining_count() * sizeof(char_type);
const std::size_t copied = (std::min)(wanted, available);
if (JSON_HEDLEY_LIKELY(copied != 0))
{
// the copy must stay within both buffers: the caller-provided
// destination holds `wanted` bytes and the remaining input range
// holds `available` bytes, and `copied` is the minimum of the two
JSON_ASSERT(copied <= wanted); // does not overrun the destination
JSON_ASSERT(copied <= available); // does not read past the input end
// &*current yields the raw address for both raw pointers and
// non-pointer contiguous iterators (e.g. std::vector's iterator)
std::memcpy(dest, &*current, copied);
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(copied / sizeof(char_type)));
}
return copied;
}
// general fallback: copy the range one element at a time
template<class T>
std::size_t get_elements_impl(T* dest, std::size_t count, std::false_type /*contiguous*/)
{
auto* ptr = reinterpret_cast<char*>(dest);
for (std::size_t read_index = 0; read_index < count * sizeof(T); ++read_index)
{
if (JSON_HEDLEY_LIKELY(current != end))
{
ptr[read_index] = static_cast<char>(*current);
std::advance(current, 1);
}
else
{
return read_index;
}
}
return count * sizeof(T);
}
IteratorType begin;
IteratorType current;
SentinelType end;
template<typename BaseInputAdapter, size_t T>
friend struct wide_string_input_helper;
bool empty() const
{
return current == end;
}
};
template<typename BaseInputAdapter, size_t T>
struct wide_string_input_helper;
template<typename BaseInputAdapter>
struct wide_string_input_helper<BaseInputAdapter, 4>
{
// UTF-32
static void fill_buffer(BaseInputAdapter& input,
std::array<std::char_traits<char>::int_type, 4>& utf8_bytes,
size_t& utf8_bytes_index,
size_t& utf8_bytes_filled)
{
utf8_bytes_index = 0;
if (JSON_HEDLEY_UNLIKELY(input.empty()))
{
utf8_bytes[0] = std::char_traits<char>::eof();
utf8_bytes_filled = 1;
}
else
{
// get the current character
const auto wc = input.get_character();
if (wc <= 0x10FFFF)
{
// UTF-32 to UTF-8 encoding
utf8_bytes_filled = 0;
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
{
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
});
}
else
{
// A code point above U+10FFFF has no UTF-8 encoding. Passing the
// unit through would narrow it to int, where 0xFFFFFFFF becomes
// char_traits<char>::eof() and would end the input silently, so
// emit a byte that is never valid UTF-8 and let the decoder
// reject it.
utf8_bytes[0] = 0xFF;
utf8_bytes_filled = 1;
}
}
}
};
template<typename BaseInputAdapter>
struct wide_string_input_helper<BaseInputAdapter, 2>
{
// UTF-16
static void fill_buffer(BaseInputAdapter& input,
std::array<std::char_traits<char>::int_type, 4>& utf8_bytes,
size_t& utf8_bytes_index,
size_t& utf8_bytes_filled)
{
utf8_bytes_index = 0;
if (JSON_HEDLEY_UNLIKELY(input.empty()))
{
utf8_bytes[0] = std::char_traits<char>::eof();
utf8_bytes_filled = 1;
}
else
{
// get the current character
const auto wc = input.get_character();
if (0xD800 > wc || wc >= 0xE000)
{
// a UTF-16 code unit outside the surrogate range is a valid
// code point (at most U+FFFF) on its own
utf8_bytes_filled = 0;
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
{
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
});
}
else
{
// A supplementary code point is a high surrogate (0xD800..0xDBFF)
// followed by a low surrogate (0xDC00..0xDFFF). A lone low
// surrogate, a high surrogate at the end of the input, or a high
// surrogate followed by any other unit is malformed UTF-16. In
// that case the offending unit is passed through unchanged so the
// UTF-8 decoder rejects it, matching how \uXXXX surrogate escapes
// are handled in the lexer.
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{
const auto wc2 = static_cast<unsigned int>(input.get_character());
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
{
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes_filled = 0;
encode_utf8(charcode, [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
{
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
});
valid_pair = true;
}
}
if (!valid_pair)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
}
}
}
}
};
// Wraps another input adapter to convert wide character types into individual bytes.
template<typename BaseInputAdapter, typename WideCharType>
class wide_string_input_adapter
{
public:
using char_type = char;
wide_string_input_adapter(BaseInputAdapter base)
: base_adapter(base) {}
typename std::char_traits<char>::int_type get_character() noexcept
{
// check if the buffer needs to be filled
if (utf8_bytes_index == utf8_bytes_filled)
{
fill_buffer<sizeof(WideCharType)>();
JSON_ASSERT(utf8_bytes_filled > 0);
JSON_ASSERT(utf8_bytes_index == 0);
}
// use buffer
JSON_ASSERT(utf8_bytes_filled > 0);
JSON_ASSERT(utf8_bytes_index < utf8_bytes_filled);
return utf8_bytes[utf8_bytes_index++];
}
// parsing binary with wchar doesn't make sense, but since the parsing mode can be runtime, we need something here
template<class T>
JSON_HEDLEY_NO_RETURN std::size_t get_elements(T* /*dest*/, std::size_t /*count*/ = 1)
{
JSON_THROW(parse_error::create(112, 1, "wide string type cannot be interpreted as binary data", nullptr));
}
private:
BaseInputAdapter base_adapter;
template<size_t T>
void fill_buffer()
{
wide_string_input_helper<BaseInputAdapter, T>::fill_buffer(base_adapter, utf8_bytes, utf8_bytes_index, utf8_bytes_filled);
}
/// a buffer for UTF-8 bytes
std::array<std::char_traits<char>::int_type, 4> utf8_bytes = {{0, 0, 0, 0}};
/// index to the utf8_codes array for the next valid byte
std::size_t utf8_bytes_index = 0;
/// number of valid bytes in the utf8_codes array
std::size_t utf8_bytes_filled = 0;
};
template<typename IteratorType, typename SentinelType = IteratorType, typename Enable = void>
struct iterator_input_adapter_factory
{
using iterator_type = IteratorType;
using sentinel_type = SentinelType;
using char_type = typename std::iterator_traits<iterator_type>::value_type;
using adapter_type = iterator_input_adapter<iterator_type, sentinel_type>;
static adapter_type create(IteratorType first, SentinelType last)
{
return adapter_type(std::move(first), std::move(last));
}
};
// Detection: whether IteratorType and SentinelType can be compared with !=
template<typename IteratorType, typename SentinelType, typename = void>
struct can_compare_ne_impl : std::false_type {};
template<typename IteratorType, typename SentinelType>
struct can_compare_ne_impl < IteratorType, SentinelType,
void_t < decltype(std::declval<IteratorType>() != std::declval<SentinelType>()) >>
: std::true_type {};
// Workaround for reversed operator order
template<typename IteratorType, typename SentinelType, typename = void>
struct can_compare_ne_reversed : std::false_type {};
template<typename IteratorType, typename SentinelType>
struct can_compare_ne_reversed < IteratorType, SentinelType,
void_t < decltype(std::declval<SentinelType>() != std::declval<IteratorType>()) >>
: std::true_type {};
template<typename IteratorType, typename SentinelType>
struct can_compare_ne_either_order : std::integral_constant < bool,
can_compare_ne_impl<IteratorType, SentinelType>::value ||
can_compare_ne_reversed<IteratorType, SentinelType>::value > {};
// std::nullptr_t is excluded explicitly: a literal `nullptr` passed as a
// trailing default argument (e.g. parse(s, nullptr, ...)) must never be
// mistaken for a sentinel, and some compilers (e.g. GCC 4.8) unreliably
// SFINAE the `operator!=` detection above for std::nullptr_t against
// container/string types, which would otherwise make such calls ambiguous
// with the compatible-input overload.
template<typename IteratorType, typename SentinelType>
struct can_compare_ne : std::integral_constant < bool,
!std::is_same<SentinelType, std::nullptr_t>::value &&
can_compare_ne_either_order<IteratorType, SentinelType>::value > {};
template<typename T>
struct is_iterator_of_multibyte
{
using value_type = typename std::iterator_traits<T>::value_type;
enum // NOLINT(cppcoreguidelines-use-enum-class)
{
value = sizeof(value_type) > 1
};
};
template<typename IteratorType, typename SentinelType>
struct iterator_input_adapter_factory<IteratorType, SentinelType, enable_if_t<is_iterator_of_multibyte<IteratorType>::value>>
{
using iterator_type = IteratorType;
using sentinel_type = SentinelType;
using char_type = typename std::iterator_traits<iterator_type>::value_type;
using base_adapter_type = iterator_input_adapter<iterator_type, sentinel_type>;
using adapter_type = wide_string_input_adapter<base_adapter_type, char_type>;
static adapter_type create(IteratorType first, SentinelType last)
{
return adapter_type(base_adapter_type(std::move(first), std::move(last)));
}
};
// General purpose iterator-based input (iterator+sentinel pair; SentinelType
// defaults to IteratorType for the common same-type case, but may differ for
// C++20 ranges-style iterator+sentinel pairs). Only enable for types that can
// be compared with !=.
template < typename IteratorType, typename SentinelType = IteratorType,
typename = typename std::enable_if <
can_compare_ne<IteratorType, SentinelType>::value >::type >
typename iterator_input_adapter_factory<IteratorType, SentinelType>::adapter_type input_adapter(IteratorType first, SentinelType last)
{
using factory_type = iterator_input_adapter_factory<IteratorType, SentinelType>;
return factory_type::create(first, last);
}
// The element type a container's data() points at, cv-qualifiers removed.
// Ill-formed - and therefore SFINAE-friendly - for types without data().
template<typename ContainerType>
using container_data_t = typename std::remove_cv<typename std::remove_pointer <
decltype(std::declval<const ContainerType&>().data()) >::type >::type;
// The container's own element type, cv-qualifiers removed. It is looked up on
// the bare type so it is also found when ContainerType is deduced as a
// reference by the forwarding-reference overload below.
template<typename ContainerType>
using container_value_t = typename std::remove_cv <
typename std::remove_cv<typename std::remove_reference<ContainerType>::type>::type::value_type >::type;
// Detect a container that stores its elements contiguously as single bytes
// (std::string, std::vector<char/unsigned char>, std::array<char, N>,
// std::string_view, ...). Such inputs are wrapped in a pointer-based adapter so
// they benefit from the contiguous fast paths (bulk string scanning, memcpy for
// binary formats) in every C++ standard - not only in C++20, where the standard
// library iterators model std::contiguous_iterator and are detected directly.
//
// data() and size() on their own would be duck typing: they say nothing about
// size() counting the units data() points at, and reading [data(), data() +
// size()) as bytes would be wrong for a type where it does not. Requiring the
// container's own value_type to be that same single-byte element ties the two
// together; every contiguous standard container satisfies it. Anything else
// keeps the iterator-based adapter, which is always correct - only slower.
template<typename ContainerType, typename = void>
struct is_contiguous_byte_container : std::false_type {};
template<typename ContainerType>
struct is_contiguous_byte_container < ContainerType, void_t <
container_data_t<ContainerType>,
container_value_t<ContainerType>,
decltype(std::declval<const ContainerType&>().size()) >>
: std::integral_constant < bool,
std::is_pointer<decltype(std::declval<const ContainerType&>().data())>::value&&
std::is_integral<container_data_t<ContainerType>>::value&&
sizeof(container_data_t<ContainerType>) == 1 &&
std::is_same<container_data_t<ContainerType>, container_value_t<ContainerType>>::value > {};
// Convenience shorthand from container to iterator
// Enables ADL on begin(container) and end(container)
// Encloses the using declarations in namespace for not to leak them to outside scope
namespace container_input_adapter_factory_impl
{
using std::begin;
using std::end;
template<typename ContainerType, typename Enable = void>
struct container_input_adapter_factory {};
template<typename ContainerType>
struct container_input_adapter_factory< ContainerType,
void_t<decltype(begin(std::declval<ContainerType>()), end(std::declval<ContainerType>()))>>
{
using adapter_type = decltype(input_adapter(begin(std::declval<ContainerType>()), end(std::declval<ContainerType>())));
static adapter_type create(ContainerType&& container)
{
// container is forwarded twice on purpose: the resulting begin/end
// iterator types must match adapter_type, computed the same way
// NOLINTNEXTLINE(bugprone-use-after-move,hicpp-invalid-access-moved)
return input_adapter(begin(std::forward<ContainerType>(container)), end(std::forward<ContainerType>(container)));
}
};
} // namespace container_input_adapter_factory_impl
// General container path (iterator-based). Contiguous single-byte containers
// are excluded here and routed through the pointer-based overload below.
template < typename ContainerType,
enable_if_t < !is_contiguous_byte_container<ContainerType>::value, int > = 0 >
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType && container)
{
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
}
// Contiguous single-byte containers (std::string, std::vector<char>, ...) are
// wrapped in a pointer-based adapter so the contiguous fast paths apply in every
// standard. The pointer keeps the container's own element type (const char* for
// std::string, const std::uint8_t* for std::vector<std::uint8_t>, ...), so the
// resulting char_type - and therefore the parsing behavior - is byte-for-byte
// identical to the iterator-based path; only the raw pointer additionally
// enables the bulk fast paths. The container outlives the adapter for the whole
// parse (temporaries live until the end of the full expression), exactly as the
// iterators it replaces did.
template < typename ContainerType,
enable_if_t < is_contiguous_byte_container<ContainerType>::value, int > = 0 >
auto input_adapter(const ContainerType& container)
-> decltype(input_adapter(container.data(), container.data() + container.size()))
{
return input_adapter(container.data(), container.data() + container.size());
}
// specialization for std::string
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
#ifndef JSON_NO_IO
// Special cases with fast paths
inline file_input_adapter input_adapter(std::FILE* file)
{
if (file == nullptr)
{
JSON_THROW(parse_error::create(101, 0, "attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
}
return file_input_adapter(file);
}
inline input_stream_adapter input_adapter(std::istream& stream)
{
if (stream.rdbuf() == nullptr)
{
JSON_THROW(parse_error::create(101, 0, "attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
}
return input_stream_adapter(stream);
}
inline input_stream_adapter input_adapter(std::istream&& stream)
{
return input_adapter(stream);
}
#endif // JSON_NO_IO
using contiguous_bytes_input_adapter = decltype(input_adapter(std::declval<const char*>(), std::declval<const char*>()));
// Null-delimited strings, and the like.
template < typename CharT,
typename std::enable_if <
std::is_pointer<CharT>::value&&
!std::is_array<CharT>::value&&
std::is_integral<typename std::remove_pointer<CharT>::type>::value&&
sizeof(typename std::remove_pointer<CharT>::type) == 1,
int >::type = 0 >
contiguous_bytes_input_adapter input_adapter(CharT b)
{
if (b == nullptr)
{
JSON_THROW(parse_error::create(101, 0, "attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
}
auto length = std::strlen(reinterpret_cast<const char*>(b));
const auto* ptr = reinterpret_cast<const char*>(b);
return input_adapter(ptr, ptr + length); // cppcheck-suppress[nullPointerArithmeticRedundantCheck]
}
template<typename T, std::size_t N>
auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
{
#if JSON_STRICT_NUL_HANDLING
// A text-literal array from string-literal initialization (e.g.
// json::parse("123") or json::parse(L"123")) carries a trailing '\0'
// contributed by the compiler, not by the source text; drop exactly that
// one byte so it is not mistaken for real trailing data. This covers all
// character types that string literals can use: char, wchar_t, char16_t,
// char32_t, and (C++20) char8_t. Every other element type (unsigned char,
// std::uint8_t, ...) keeps the full extent unconditionally, since a
// trailing zero byte there is genuine data (e.g. CBOR/MessagePack). This
// intentionally does not strlen()-scan the array (as the pointer overload
// above does for a null-delimited string): for an array that is not
// NUL-terminated within its bounds, that would read past the end of the
// array.
using char_t = typename std::remove_cv<T>::type;
constexpr bool is_text_literal_type = std::is_same<char_t, char>::value
|| std::is_same<char_t, wchar_t>::value
|| std::is_same<char_t, char16_t>::value
|| std::is_same<char_t, char32_t>::value
#if defined(__cpp_char8_t)
|| std::is_same<char_t, char8_t>::value
#endif
;
if (is_text_literal_type && N > 0 && array[N - 1] == 0)
{
return input_adapter(array, array + N - 1);
}
#endif
return input_adapter(array, array + N);
}
// This class only handles inputs that construct a contiguous_bytes_input_adapter
// (e.g. span_input_adapter). It's required so that expressions like {ptr, len}
// can be implicitly cast to the correct adapter.
class span_input_adapter
{
public:
template < typename CharT,
typename std::enable_if <
std::is_pointer<CharT>::value&&
std::is_integral<typename std::remove_pointer<CharT>::type>::value&&
sizeof(typename std::remove_pointer<CharT>::type) == 1,
int >::type = 0 >
span_input_adapter(CharT b, std::size_t l)
: ia(reinterpret_cast<const char*>(b), reinterpret_cast<const char*>(b) + l) {}
template<class IteratorType,
typename std::enable_if<
std::is_same<typename iterator_traits<IteratorType>::iterator_category, std::random_access_iterator_tag>::value,
int>::type = 0>
span_input_adapter(IteratorType first, IteratorType last)
: ia(input_adapter(first, last)) {}
contiguous_bytes_input_adapter&& get()
{
return std::move(ia); // NOLINT(hicpp-move-const-arg,performance-move-const-arg)
}
private:
contiguous_bytes_input_adapter ia;
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END