Compare commits

..
Author SHA1 Message Date
Niels Lohmann 3a04f78507 Make convert_null_to() take the container type as a template argument
Passing array_t or object_t instead of a value_t makes an invalid target
a compile error instead of a runtime assertion, and removes the branch.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 21:31:58 +02:00
Niels Lohmann 09d41a894b Re-enable bugprone-use-after-move/hicpp-invalid-access-moved
These two checks (and portability-template-virtual-member-function)
were disabled in #4489 (November 2024) "only removed to get the CI
going". portability-template-virtual-member-function is a separate,
still-open cleanup (#5725 item 3 on its own branch) and stays
disabled here; this commit only re-enables the move/forward checks
and cleans up what they flag on this branch.

The move constructor (json.hpp) already casts to the base type
instead of forwarding the whole object (#5724 item 9), so it no
longer trips either check. The at(KeyType&&) double-forward this
check used to flag was reduced to a single forward with the
now-unforwarded reuse annotated by a NOLINTNEXTLINE in #5724 item 3's
object_at() helper (clang-tidy 22 still flags that reuse even after a
single forward; see that commit's message). What is left here:

- from_json_inplace_array_impl(), from_json_tuple_impl_base() and the
  std::pair overload of from_json_tuple_impl() forwarded j into every
  j.at(...) call in a pack expansion or a pair of calls. at() has no
  ref-qualified overloads, so the forward was a no-op; call j.at(...)
  directly.
- container_input_adapter_factory::create() forwards container twice
  on purpose, into begin() and end(), so both see the same value
  category and produce matching iterator types. Annotate it with
  NOLINTNEXTLINE and a comment instead of changing it.
- unit-class_parser.cpp's "move constructor resets the moved-from
  value to npos" test still pointed at the pre-static_cast move
  constructor by line number and mentioned the cppcheck-suppress
  annotation that #5724 item 9 already removed; update the comment.

No behavior change anywhere in include/. Verified with clang-tidy 22.1.8
(Docker silkeh/clang:22, --platform linux/amd64) against a TU including
json.hpp with the repo's .clang-tidy: bugprone-use-after-move and
hicpp-invalid-access-moved report nothing unsuppressed.

Overlaps #5737 (open PR for the rest of #5725 item 3: the
from_json.hpp/input_adapters.hpp cleanup above, and
portability-template-virtual-member-function), which currently keeps
both checks disabled pending this move-constructor change; whichever
of this commit and that PR lands second will need a small rebase of
.clang-tidy.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 19:05:07 +02:00
Niels Lohmann 3e771c2ad7 Deduplicate the linear key search in ordered_map
emplace, at, erase(key), count and find each repeated the same
"for (auto it = begin(); it != end(); ++it) if (m_compare(it->first,
key)) ..." loop (15 copies across their key_type and transparent
KeyType&& overloads), and both erase(key) overloads additionally
repeated the exception-sensitive in-place reconstruction (destroy,
placement-new, pop_back) used to remove an element while keeping the
const Key non-movable.

Add two private helpers: find_impl(Self&, KeyType&&), a static member
template that runs the search once for either constness of the
receiver, and erase_at(iterator), which keeps the existing
pop_back-based reconstruction instead of switching to erase()/resize()
(which would add a DefaultInsertable requirement). Route find, at,
count, emplace, insert(const value_type&) and both erase(key)
overloads through them.

Same signatures, is_usable_as_key_type constraints and exception
messages/types.

Overlaps #5609 and #5685, which both rewrite emplace (#5609 also
touches insert and adds private members at the end of the class);
whichever of this commit and those PRs lands second will need a
rebase.

#5724 item 6

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 19:05:05 +02:00
Niels Lohmann 51233dc17d Deduplicate the null-to-container conversion into convert_null_to()
Nine sites wrote out the same "turn a null value into an empty array
or object" logic with two different idioms: operator[](size_type),
operator[](key_type), operator[](KeyType&&) and update() set m_type
then assigned m_value.array/object directly via create<T>(), while the
three push_back() overloads, emplace_back() and emplace() set m_type
then assigned m_value = value_t::array/object (going through
json_value's converting constructor and a temporary). Both idioms end
up calling create<T>() and produce the same state, just via a
different path; both also share a latent exception-safety bug, since
m_type is written before the (possibly throwing) allocation, so a
throwing allocator leaves m_type == array/object with a null pointer
behind it, violating the class invariant and crashing on the next
access to, or destruction of, the value.

Add a private convert_null_to(value_t) helper and call it from all
nine sites. Unlike the idioms it replaces, it allocates the container
first and only then writes m_type, so a throwing allocation leaves the
value as a valid null instead of a mistyped, half-constructed one;
verified with a throwing allocator (see unit-allocator.cpp's
bad_allocator) that j["x"] = ... on a null j now stays null, and no
longer trips assert_invariant()/crashes, when create<object_t>()
throws. Same allocator usage and assert_invariant() call as before,
otherwise.

Overlaps #5585, which reorders these same nine blocks for exception
safety; whichever of this commit and that PR lands second will need a
small rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 19:05:05 +02:00
Niels Lohmann 03d53596f9 Deduplicate the lookup/default/throw body of value()
The six non-deprecated value() overloads each held a full copy of the
same body: the four key-based overloads looked up the key and either
returned the found element converted to the requested type or the
default value (throwing type_error.306 if this is not an object), and
the two json_pointer overloads did the same via
ptr.get_checked_or_null(), throwing type_error.306 unless
is_structured(). Replace the duplicated bodies with two private
helpers, value_member() and value_pointee(), that return a
const basic_json* (null when not found) and do the type check/throw
once each. Every value() overload now just picks between the found
pointer's get<T>() and the default.

Same signatures, template parameters, SFINAE conditions, exception id,
message and this context on every overload.

Overlaps #5689 (routes find() through lookup_key()) and #5705 (adds a
deleted integral-key value() next to these overloads); whichever of
this commit and those PRs lands second will need a small rebase.

#5724 item 4

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 19:05:03 +02:00
Niels Lohmann a0e01e3b9c Deduplicate object key lookup and checked at() access
at()/find()/count()/contains()/const operator[]/erase_internal() each
repeated the raw object lookup (m_value.object->find(key)), and the
six at() overloads additionally repeated the type_error.304 check and
out_of_range.401/403 throw. Route them all through two new private
helpers, object_lookup()/object_at() (plus array_at() for the index
overloads of at()), templated on the constness of the receiver so one
body serves both the const and non-const overload. count() is left
untouched, since it already goes through object_t::count() rather
than a second find().

The at(KeyType&&) overloads used to forward the same key twice: once
into object->find() and again, on the not-found path, into the
string_t() conversion for the exception message. object_at() now
forwards it only into the lookup and reuses the (unmoved) key for the
message. clang-tidy 22 (Docker silkeh/clang:22) still flags that reuse
under bugprone-use-after-move/hicpp-invalid-access-moved even with the
single forward, since it cannot see that object_t::find() (a plain
std::map or ordered_map) never actually moves from its argument; add a
NOLINTNEXTLINE with that reasoning rather than avoid the pattern.

No signature, exception id/message, or set_parent() behavior changes.

Overlaps #5689, #5705, #5606, #5687 and #5585, which touch the same
hunks; whichever of this commit and those PRs lands second will need
a small rebase.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 19:05:02 +02:00
Niels Lohmann 0c0d5edda3 Make insert(pos, basic_json&&) move its argument instead of copying it
insert(const_iterator pos, basic_json&& val) delegated to
insert(pos, val), but val is a named rvalue reference, so inside the
function it is an lvalue: the call always resolved to
insert(const_iterator, const basic_json&) and deep-copied the value.
This has been the case since the overload was introduced, in every
release. push_back(basic_json&&), by contrast, already moves.

Give the rvalue overload its own body with the same two checks
(type_error.309, invalid_iterator.202), then move the argument into a
local before inserting it. Moving into a local first, rather than
inserting std::move(val) directly, keeps this safe even when val
aliases an element of the same array (e.g.
arr.insert(arr.begin(), std::move(arr[1]))), since
std::vector::insert(pos, T&&) is not guaranteed to handle an argument
that aliases one of its own elements.

This is a deliberate, small behavior change: the moved-from argument
now ends up null afterwards, the same as after push_back(&&), instead
of keeping its old value unchanged. No signature changes, so the
public API and ABI are unaffected. Add unit-modifiers coverage for
the moved-from state and for self-aliasing insertion, both with and
without reallocation of the underlying array.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 11:11:16 +02:00
Niels Lohmann ff2ec8116d Deduplicate the 16 copied from_cbor/msgpack/ubjson/bjdata/bon8/bson bodies
Each of the 16 binary deserialization overloads (from_cbor,
from_msgpack, from_ubjson, from_bjdata, from_bon8, from_bson, each in
an InputType&& and an iterator/sentinel version, plus the deprecated
span overloads of from_cbor/from_msgpack/from_ubjson/from_bson) had
the same body, differing only in the input_format_t value. Every copy
built a temporary binary_reader from std::move(ia) and called
sax_parse on it in the same expression, which also produced a
false-positive cppcheck accessMoved on all 16 lines and needed a
NOLINTNEXTLINE(hicpp-move-const-arg,performance-move-const-arg) on
the four span overloads.

Add a private from_binary_impl() helper that builds the reader as a
named local instead, and make each of the 16 overloads a one-line
forward to it. All public signatures, default arguments,
JSON_HEDLEY_WARN_UNUSED_RESULT and JSON_HEDLEY_DEPRECATED_FOR
attributes are unchanged, tag_handler keeps defaulting to
cbor_tag_handler_t::error for the non-CBOR formats (matching
binary_reader::sax_parse's own default), and the helper is placed in
the existing private section before the binary section banner rather
than between the from_* overloads, so from_binary_impl() itself does
not collide with #5688's insertion point. Collapsing the
from_bjdata/from_bon8 bodies into one-line forwards does rewrite the
"return result; }" context lines that #5688 inserts its two
deprecated overloads after, so that PR will need a small manual
rebase (reinserting its overloads after the new one-line bodies)
rather than applying cleanly.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 11:11:15 +02:00
Niels Lohmann b591304d04 Fix stale and copy-pasted comments in basic_json
Several comments no longer match the code: the class invariant and
assert_invariant()'s doc still named the members m_value/m_type
(now m_data.m_value/m_data.m_type) and did not mention the binary
invariant that assert_invariant() already checks; the json_value note
and the get<PointerType>() @tparam list omitted binary_t even though
binary is a variable-length, pointer-stored type like the others; the
key-based value() overload's brief said "via JSON Pointer", which is
the other overload; and swap(binary_t&)/swap(binary_t::container_type&)
both carried "swap only works for strings", copied from swap(string_t&).

Comment-only change; behavior, the public API and the ABI are
unchanged. The private get_impl() doxygen and the emplace() comments
that border #5585's hunk are intentionally left alone.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:42:29 +02:00
Niels Lohmann c2553bac5a Deduplicate the string/binary cleanup in the two erase() overloads
erase(pos) and erase(first, last) each carried a byte-identical
14-line block that destroys and deallocates a string or binary
value before resetting the type to null. That reimplements the
string/binary cases of json_value::destroy(), so any future change to
how those values are freed would have to be made in three places
instead of one. Both overloads now just call destroy() and reset the
union; for the other primitive types (boolean, numbers) destroy() is
a no-op, so behavior is unchanged.

Also fix erase(first, last)'s error-path branch hint, which used
JSON_HEDLEY_LIKELY where erase(pos), the iterator-range constructor,
and every other error path in the class use JSON_HEDLEY_UNLIKELY.
This only affects code layout, not semantics.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:40:04 +02:00
Niels Lohmann 1274e0f250 Remove stale cppcheck suppressions and name the local parser in parse()
Running the pinned cppcheck (ci_cppcheck's invocation) without
--inline-suppr across all configurations reports no syntaxError, no
ignoredReturnValue and no assertWithSideEffect, so the corresponding
suppressions in json_fwd.hpp, string_concat.hpp and
assert_invariant() no longer match anything (json_fwd.hpp's is kept,
since downstream users who run an older cppcheck against it could
still hit the warning it once silenced).

The three basic_json::parse() overloads still trigger a false-positive
accessMoved/accessForwarded because they build a temporary parser and
call .parse() on it in the same expression; giving that parser a name
makes the warning go away without changing behavior, and removes the
last of the inline suppressions on these functions.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:37:48 +02:00
Niels Lohmann ffd3a3722b Remove four cppcheck accessForwarded suppressions in the move constructor
basic_json(basic_json&&) built its base subobject with
std::forward<json_base_class_t>(other), so cppcheck saw the whole of
other as forwarded and flagged every subsequent access to it as
accessForwarded, three of them still marked "TODO check". Only the
base subobject is actually moved from; cast explicitly to the base
type instead, the way ordered_map already does, so cppcheck can tell
the two are unrelated. Behavior is unchanged: for a non-reference T,
std::forward<T>(x) is defined as static_cast<T&&>(x).

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:36:41 +02:00
Niels Lohmann f5e4b9b8cf Stop the noexcept null constructor delegating to a throwing one
basic_json(std::nullptr_t) delegated to basic_json(value_t), whose
underlying json_value(value_t) constructor allocates for other types
and can therefore throw, which is why the noexcept had a
NOLINT(bugprone-exception-escape). The delegated-to constructor also
called assert_invariant() a second time. The default member
initializers of data already produce the same null state (a
value-initialized, i.e. zeroed, union with object == nullptr), so the
delegation and its NOLINT can simply be dropped.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:34:05 +02:00
Niels Lohmann e5c3e81279 Fix meta()'s dead, syntactically invalid HP aCC branch
The HP aCC branch of basic_json::meta() was missing a semicolon and
has therefore never compiled; adding only a semicolon would also make
it throw type_error.305, since it assigned a plain string to
result["compiler"] and then indexed into it like the other branches
do into an object. Make the branch consistent with the others by
assigning an object with "family" and "version" keys, narrow the
condition to __HP_aCC (a C compiler cannot build this header-only
library), and fix meta.md, which documented the old (impossible)
plain-string behavior.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 10:32:07 +02:00
Niels Lohmann 1d98396a19 Remove unused private aliases from basic_json
The private aliases primitive_iterator_t, internal_iterator and
output_adapter_t are not used anywhere: iter_impl, binary_writer and
the tests refer to the detail:: names directly. As the aliases are
private, no user or derived class can depend on them. The
internal_iterator.hpp include stays because iter_impl needs it.

Part of #5724

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 09:47:00 +02:00
15 changed files with 1330 additions and 1599 deletions
+1 -3
View File
@@ -1,11 +1,9 @@
# TODO: The first three checks are only removed to get the CI going. They have to be addressed at some point.
# TODO: portability-template-virtual-member-function is only removed to get the CI going. It has to be addressed at some point.
# TODO: portability-avoid-pragma-once: should be fixed eventually
Checks: '*,
-portability-template-virtual-member-function,
-bugprone-use-after-move,
-hicpp-invalid-access-moved,
-altera-id-dependent-backward-branch,
-altera-struct-pack-align,
+1 -1
View File
@@ -13,7 +13,7 @@ JSON object holding version information
| key | description |
|-------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `compiler` | Information on the used compiler. It is an object with the following keys: `c++` (the used C++ standard), `family` (the compiler family; possible values are `clang`, `icc`, `gcc`, `ilecpp`, `msvc`, `pgcpp`, `sunpro`, and `unknown`), and `version` (the compiler version). On HP aCC compilers, `compiler` is instead the plain string `hp`. |
| `compiler` | Information on the used compiler. It is an object with the following keys: `c++` (the used C++ standard), `family` (the compiler family; possible values are `clang`, `icc`, `gcc`, `hp`, `ilecpp`, `msvc`, `pgcpp`, `sunpro`, and `unknown`), and `version` (the compiler version). |
| `copyright` | The copyright line for the library as string. |
| `name` | The name of the library as string. |
| `platform` | The used platform as string. Possible values are `win32`, `linux`, `apple`, `unix`, and `unknown`. |
@@ -353,7 +353,7 @@ template < typename BasicJsonType, typename T, std::size_t... Idx >
std::array<T, sizeof...(Idx)> from_json_inplace_array_impl(BasicJsonType&& j,
identity_tag<std::array<T, sizeof...(Idx)>> /*unused*/, index_sequence<Idx...> /*unused*/)
{
return { { std::forward<BasicJsonType>(j).at(Idx).template get<T>()... } };
return { { j.at(Idx).template get<T>()... } };
}
template < typename BasicJsonType, typename T, std::size_t N >
@@ -502,7 +502,7 @@ using tuple_type = std::tuple < decltype(from_json_tuple_get_impl(std::declval<B
template<std::size_t PTagValue, typename... Args, typename BasicJsonType, std::size_t... Idx>
tuple_type<PTagValue, BasicJsonType, Args...> from_json_tuple_impl_base(BasicJsonType&& j, index_sequence<Idx...> /*unused*/)
{
return tuple_type<PTagValue, BasicJsonType, Args...>(from_json_tuple_get_impl(std::forward<BasicJsonType>(j).at(Idx), detail::identity_tag<Args> {}, detail::priority_tag<PTagValue> {})...);
return tuple_type<PTagValue, BasicJsonType, Args...>(from_json_tuple_get_impl(j.at(Idx), detail::identity_tag<Args> {}, detail::priority_tag<PTagValue> {})...);
}
template<std::size_t PTagValue, typename BasicJsonType>
@@ -514,8 +514,8 @@ std::tuple<> from_json_tuple_impl_base(BasicJsonType& /*unused*/, index_sequence
template < typename BasicJsonType, class A1, class A2 >
std::pair<A1, A2> from_json_tuple_impl(BasicJsonType&& j, identity_tag<std::pair<A1, A2>> /*unused*/, priority_tag<0> /*unused*/)
{
return {std::forward<BasicJsonType>(j).at(0).template get<A1>(),
std::forward<BasicJsonType>(j).at(1).template get<A2>()};
return {j.at(0).template get<A1>(),
j.at(1).template get<A2>()};
}
template<typename BasicJsonType, typename A1, typename A2>
@@ -3344,8 +3344,8 @@ class binary_reader
return enter_object(detail::unknown_size());
}
// Note, UBJSON has no binary type of its own; BJData, which shares this
// reader, decodes optimized 'B' arrays as binary in get_ubjson_array().
// Note, no reader for UBJSON binary types is implemented because they do
// not exist
bool get_ubjson_high_precision_number()
{
@@ -4052,7 +4052,7 @@ class binary_reader
#endif
}
/*!
/*
@brief read a number from the input
@tparam NumberType the type of the number
@@ -4062,10 +4062,10 @@ class binary_reader
@return whether conversion completed
@note This function needs to respect the system's endianness, because
bytes in CBOR, MessagePack, UBJSON, and BON8 are stored in network
order (big endian) and therefore need reordering on little endian
systems. On the other hand, BSON and BJData use little endian and
should reorder on big endian systems.
bytes in CBOR, MessagePack, and UBJSON are stored in network order
(big endian) and therefore need reordering on little endian systems.
On the other hand, BSON and BJData use little endian and should reorder
on big endian systems.
*/
template<typename NumberType, bool InputIsLittleEndian = false>
bool get_number(const input_format_t format, NumberType& result)
@@ -8,12 +8,12 @@
#pragma once
#include <algorithm> // min
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint32_t
#include <cstring> // strlen
#include <iterator> // begin, end, iterator_traits, random_access_iterator_tag, distance, next
#include <memory> // shared_ptr, make_shared, addressof
#include <numeric> // accumulate
#include <streambuf> // streambuf
#include <string> // string, char_traits
#include <type_traits> // enable_if, is_base_of, is_pointer, is_integral, remove_pointer
@@ -28,7 +28,6 @@
#include <nlohmann/detail/iterators/iterator_traits.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/string_utils.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -83,9 +82,8 @@ class file_input_adapter
};
/*!
Input adapter for a (caching) istream. Does not skip a UTF Byte Order Mark
itself; that is done by the lexer's skip_bom(). Does not support changing
the underlying std::streambuf
Input adapter for a (caching) istream. Ignores a UFT Byte Order Mark at
beginning of input. Does not support changing the underlying std::streambuf
in mid-input. Maintains underlying std::istream and std::streambuf to support
subsequent use of standard std::istream operations to process any input
characters following those used in parsing the JSON input. Clears the
@@ -450,14 +448,32 @@ struct wide_string_input_helper<BaseInputAdapter, 4>
// get the current character
const auto wc = input.get_character();
if (wc <= 0x10FFFF)
// UTF-32 to UTF-8 encoding
if (wc < 0x80)
{
// UTF-32 to UTF-8 encoding
utf8_bytes_filled = 0;
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
{
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
});
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
}
else if (wc <= 0x7FF)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xC0u | ((static_cast<unsigned int>(wc) >> 6u) & 0x1Fu));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
utf8_bytes_filled = 2;
}
else if (wc <= 0xFFFF)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xE0u | ((static_cast<unsigned int>(wc) >> 12u) & 0x0Fu));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
utf8_bytes_filled = 3;
}
else if (wc <= 0x10FFFF)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | ((static_cast<unsigned int>(wc) >> 18u) & 0x07u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
utf8_bytes_filled = 4;
}
else
{
@@ -494,15 +510,24 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
// get the current character
const auto wc = input.get_character();
if (0xD800 > wc || wc >= 0xE000)
// UTF-16 to UTF-8 encoding
if (wc < 0x80)
{
// a UTF-16 code unit outside the surrogate range is a valid
// code point (at most U+FFFF) on its own
utf8_bytes_filled = 0;
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
{
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
});
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
}
else if (wc <= 0x7FF)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xC0u | ((static_cast<unsigned int>(wc) >> 6u)));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
utf8_bytes_filled = 2;
}
else if (0xD800 > wc || wc >= 0xE000)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xE0u | ((static_cast<unsigned int>(wc) >> 12u)));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
utf8_bytes_filled = 3;
}
else
{
@@ -520,11 +545,11 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
{
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes_filled = 0;
encode_utf8(charcode, [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
{
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
});
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
valid_pair = true;
}
}
@@ -738,6 +763,7 @@ struct container_input_adapter_factory< ContainerType,
static adapter_type create(ContainerType&& container)
{
// NOLINTNEXTLINE(bugprone-use-after-move,hicpp-invalid-access-moved) forwarded twice on purpose, so begin() and end() see the same value category and yield matching iterator types
return input_adapter(begin(std::forward<ContainerType>(container)), end(std::forward<ContainerType>(container)));
}
};
@@ -837,9 +863,9 @@ auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) /
return input_adapter(array, array + N);
}
// This class only handles inputs that construct a contiguous_bytes_input_adapter
// (e.g. span_input_adapter). It's required so that expressions like {ptr, len}
// can be implicitly cast to the correct adapter.
// This class only handles inputs of input_buffer_adapter type.
// It's required so that expressions like {ptr, len} can be implicitly cast
// to the correct adapter.
class span_input_adapter
{
public:
+142 -89
View File
@@ -10,7 +10,6 @@
#include <algorithm> // find_if, min
#include <cstddef>
#include <limits> // numeric_limits
#include <string> // string
#include <type_traits> // enable_if_t
#include <utility> // move, pair
@@ -176,88 +175,6 @@ template<typename ArrayType>
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
{}
#if JSON_DIAGNOSTIC_POSITIONS
/*!
@brief set the diagnostic positions of a value the DOM SAX parsers just stored
Shared by json_sax_dom_parser and json_sax_dom_callback_parser. basic_json
befriends this struct, as the position members are private.
*/
struct diagnostic_positions
{
/*!
@param[in,out] v the value that was just parsed
@param[in] lexer the lexer that read it, or nullptr to leave @a v alone
*/
template<typename BasicJsonType, typename LexerType>
static void set_from_lexer(BasicJsonType& v, LexerType* lexer)
{
if (lexer)
{
// Lexer has read past the current field value, so set the end position to the current position.
// The start position will be set below based on the length of the string representation
// of the value.
v.end_position = lexer->get_position();
switch (v.type())
{
case value_t::boolean:
{
// 4 and 5 are the string length of "true" and "false"
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
break;
}
case value_t::null:
{
// 4 is the string length of "null"
v.start_position = v.end_position - 4;
break;
}
case value_t::string:
{
// escape sequences make the token longer than the value it
// parses to, so the start position cannot be derived from
// the value; use the offset the lexer recorded instead
v.start_position = lexer->get_token_start_position();
break;
}
case value_t::discarded:
{
// an object or array the callback of
// json_sax_dom_callback_parser rejected has no position
v.end_position = std::string::npos;
v.start_position = v.end_position;
break;
}
case value_t::binary:
case value_t::number_integer:
case value_t::number_unsigned:
case value_t::number_float:
{
v.start_position = v.end_position - lexer->get_string().size();
break;
}
case value_t::object:
case value_t::array:
{
// object and array are handled in start_object() and start_array() handlers
// skip setting the values here.
break;
}
default: // LCOV_EXCL_LINE
// Handle all possible types discretely, default handler should never be reached.
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
}
}
};
#endif
/*!
@brief SAX implementation to create a JSON value from SAX events
@@ -459,6 +376,76 @@ class json_sax_dom_parser
private:
#if JSON_DIAGNOSTIC_POSITIONS
void handle_diagnostic_positions_for_json_value(BasicJsonType& v)
{
if (m_lexer_ref)
{
// Lexer has read past the current field value, so set the end position to the current position.
// The start position will be set below based on the length of the string representation
// of the value.
v.end_position = m_lexer_ref->get_position();
switch (v.type())
{
case value_t::boolean:
{
// 4 and 5 are the string length of "true" and "false"
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
break;
}
case value_t::null:
{
// 4 is the string length of "null"
v.start_position = v.end_position - 4;
break;
}
case value_t::string:
{
// escape sequences make the token longer than the value it
// parses to, so the start position cannot be derived from
// the value; use the offset the lexer recorded instead
v.start_position = m_lexer_ref->get_token_start_position();
break;
}
// As we handle the start and end positions for values created during parsing,
// we do not expect the following value type to be called. Regardless, set the positions
// in case this is created manually or through a different constructor. Exclude from lcov
// since the exact condition of this switch is esoteric.
// LCOV_EXCL_START
case value_t::discarded:
{
v.end_position = std::string::npos;
v.start_position = v.end_position;
break;
}
// LCOV_EXCL_STOP
case value_t::binary:
case value_t::number_integer:
case value_t::number_unsigned:
case value_t::number_float:
{
v.start_position = v.end_position - m_lexer_ref->get_string().size();
break;
}
case value_t::object:
case value_t::array:
{
// object and array are handled in start_object() and start_array() handlers
// skip setting the values here.
break;
}
default: // LCOV_EXCL_LINE
// Handle all possible types discretely, default handler should never be reached.
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert,-warnings-as-errors) LCOV_EXCL_LINE
}
}
}
#endif
/*!
@invariant If the ref stack is empty, then the passed value will be the new
root.
@@ -474,7 +461,7 @@ class json_sax_dom_parser
root = BasicJsonType(std::forward<Value>(v));
#if JSON_DIAGNOSTIC_POSITIONS
diagnostic_positions::set_from_lexer(root, m_lexer_ref);
handle_diagnostic_positions_for_json_value(root);
#endif
return &root;
@@ -487,7 +474,7 @@ class json_sax_dom_parser
ref_stack.back()->m_data.m_value.array->emplace_back(std::forward<Value>(v));
#if JSON_DIAGNOSTIC_POSITIONS
diagnostic_positions::set_from_lexer(ref_stack.back()->m_data.m_value.array->back(), m_lexer_ref);
handle_diagnostic_positions_for_json_value(ref_stack.back()->m_data.m_value.array->back());
#endif
return &(ref_stack.back()->m_data.m_value.array->back());
@@ -498,7 +485,7 @@ class json_sax_dom_parser
*object_element = BasicJsonType(std::forward<Value>(v));
#if JSON_DIAGNOSTIC_POSITIONS
diagnostic_positions::set_from_lexer(*object_element, m_lexer_ref);
handle_diagnostic_positions_for_json_value(*object_element);
#endif
return object_element;
@@ -675,7 +662,7 @@ class json_sax_dom_callback_parser
#if JSON_DIAGNOSTIC_POSITIONS
// Set start/end positions for discarded object.
diagnostic_positions::set_from_lexer(*ref_stack.back(), m_lexer_ref);
handle_diagnostic_positions_for_json_value(*ref_stack.back());
#endif
}
}
@@ -791,7 +778,7 @@ class json_sax_dom_callback_parser
#if JSON_DIAGNOSTIC_POSITIONS
// Set start/end positions for discarded array.
diagnostic_positions::set_from_lexer(*ref_stack.back(), m_lexer_ref);
handle_diagnostic_positions_for_json_value(*ref_stack.back());
#endif
}
}
@@ -844,6 +831,72 @@ class json_sax_dom_callback_parser
private:
#if JSON_DIAGNOSTIC_POSITIONS
void handle_diagnostic_positions_for_json_value(BasicJsonType& v)
{
if (m_lexer_ref)
{
// Lexer has read past the current field value, so set the end position to the current position.
// The start position will be set below based on the length of the string representation
// of the value.
v.end_position = m_lexer_ref->get_position();
switch (v.type())
{
case value_t::boolean:
{
// 4 and 5 are the string length of "true" and "false"
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
break;
}
case value_t::null:
{
// 4 is the string length of "null"
v.start_position = v.end_position - 4;
break;
}
case value_t::string:
{
// escape sequences make the token longer than the value it
// parses to, so the start position cannot be derived from
// the value; use the offset the lexer recorded instead
v.start_position = m_lexer_ref->get_token_start_position();
break;
}
case value_t::discarded:
{
v.end_position = std::string::npos;
v.start_position = v.end_position;
break;
}
case value_t::binary:
case value_t::number_integer:
case value_t::number_unsigned:
case value_t::number_float:
{
v.start_position = v.end_position - m_lexer_ref->get_string().size();
break;
}
case value_t::object:
case value_t::array:
{
// object and array are handled in start_object() and start_array() handlers
// skip setting the values here.
break;
}
default: // LCOV_EXCL_LINE
// Handle all possible types discretely, default handler should never be reached.
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert,-warnings-as-errors) LCOV_EXCL_LINE
}
}
}
#endif
/// if there is a pending duplicate-key stash entry for this exact slot,
/// remove it from the stash; if restore_value is true, the stashed
/// previous value is moved back into the slot first (use this when the
@@ -965,7 +1018,7 @@ class json_sax_dom_callback_parser
auto value = BasicJsonType(std::forward<Value>(v));
#if JSON_DIAGNOSTIC_POSITIONS
diagnostic_positions::set_from_lexer(value, m_lexer_ref);
handle_diagnostic_positions_for_json_value(value);
#endif
// check callback
+68 -33
View File
@@ -11,7 +11,6 @@
#include <array> // array
#include <clocale> // localeconv
#include <cstddef> // size_t
#include <cstdint> // uint32_t
#include <cstdio> // snprintf
#include <cstdlib> // strtof, strtod, strtold, strtoll, strtoull
#include <initializer_list> // initializer_list
@@ -25,7 +24,6 @@
#include <nlohmann/detail/input/string_scan.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/type_traits.hpp>
#include <nlohmann/detail/string_utils.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
@@ -502,10 +500,32 @@ class lexer : public lexer_base<BasicJsonType>
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
// translate codepoint into bytes
encode_utf8(static_cast<std::uint32_t>(codepoint), [this](std::uint32_t byte)
if (codepoint < 0x80)
{
add(static_cast<char_int_type>(byte));
});
// 1-byte characters: 0xxxxxxx (ASCII)
add(static_cast<char_int_type>(codepoint));
}
else if (codepoint <= 0x7FF)
{
// 2-byte characters: 110xxxxx 10xxxxxx
add(static_cast<char_int_type>(0xC0u | (static_cast<unsigned int>(codepoint) >> 6u)));
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
}
else if (codepoint <= 0xFFFF)
{
// 3-byte characters: 1110xxxx 10xxxxxx 10xxxxxx
add(static_cast<char_int_type>(0xE0u | (static_cast<unsigned int>(codepoint) >> 12u)));
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
}
else
{
// 4-byte characters: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
add(static_cast<char_int_type>(0xF0u | (static_cast<unsigned int>(codepoint) >> 18u)));
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 12u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 6u) & 0x3Fu)));
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
}
break;
}
@@ -1474,30 +1494,45 @@ scan_number_done:
*/
token_type convert_number(token_type number_type, std::size_t mantissa_end)
{
// accept() only needs to know whether the input is valid, so it sets
// discard_number_values (see json.hpp), and an integer token whose
// digit count shows that it fits is reported without calling
// convert_integer(). A number with up to 18 digits always fits into
// both std::uint64_t and std::int64_t (18 nines is about 1e18, below
// INT64_MAX, which is about 9.2e18). Longer tokens take the exact path
// below, including the fallback to floating point when the value does
// not fit.
// If the caller does not need the converted value (only whether the
// input is syntactically valid; see json_sax_acceptor/accept()), an
// unsigned/integer token can be reported without calling
// strtoull()/strtoll() at all, *provided* we can already tell from
// the digit count alone that the conversion cannot overflow 64 bits.
// Such tokens are always finite and are accepted unconditionally by
// the parser regardless of their actual value (parser::sax_parse_internal()
// never checks finiteness for value_unsigned/value_integer), so the
// classification below is all that is needed.
//
// With a narrower number_unsigned_t/number_integer_t (e.g.
// std::uint32_t), the exact path would reclassify some of these tokens
// as (finite) floats, while this check reports integers. That does not
// change the result of accept(): it always parses through
// json_sax_acceptor, whose number callbacks discard their argument and
// return true, and the parser rejects neither integers nor finite
// floats. value_unsigned/value_integer are left unset here, so a caller
// that reads the converted value must not set discard_number_values.
// A decimal number with up to 18 digits is always representable in
// both std::uint64_t and std::int64_t (18 nines is ~1e18, well below
// both UINT64_MAX ~1.8e19 and INT64_MAX ~9.2e18), so strtoull()/strtoll()
// could not have set errno to ERANGE for it. Numbers with more digits
// (rare in practice) fall through to the exact code below, unchanged,
// so their handling -- including reclassification to value_float when
// the value overflows 64 bits, and rejection when it is not even
// finite as a double -- is bit-for-bit identical to before this
// optimization.
//
// On contiguous input, scan_number_bulk_contiguous() converts integer
// tokens itself and does not pass them to this function, unless
// JSON_DIAGNOSTIC_POSITIONS is enabled. This check is therefore only
// reached for input without bulk access (e.g. streams), with
// JSON_DIAGNOSTIC_POSITIONS, or when scan_number_bulk_contiguous()
// falls back to scan_number().
// Note this reasons about std::uint64_t/std::int64_t, not about
// number_unsigned_t/number_integer_t (BasicJsonType's own, possibly
// narrower, template parameters -- e.g. std::uint32_t). That is fine
// *only* because discard_number_values is exclusively set by
// accept() (see json.hpp), and accept() always parses through the
// library's own json_sax_acceptor -- never a user-supplied SAX
// consumer -- whose number_unsigned()/number_integer()/number_float()
// callbacks unconditionally discard their argument and return true.
// So for every caller that can reach this branch, neither the token
// classification below nor the eventual (possibly narrowed, and on
// this fast path left stale/unset) value_unsigned/value_integer is
// ever consulted -- an unsigned/integer token is accepted outright,
// and even a >18-digit token that this fast path deliberately falls
// through for is, once reclassified to value_float, still finite
// (and thus accepted) for any digit count that fits in number_unsigned_t
// or number_integer_t regardless of that type's width. If this
// function is ever taught to run with discard_number_values true for
// a caller that *does* read the converted value, this reasoning (and
// the fast path below) would need to be revisited.
if (discard_number_values)
{
constexpr std::size_t safe_digit_count = 18;
@@ -1988,7 +2023,7 @@ scan_number_done:
return value_float;
}
/// return current string value
/// return current string value (implicitly resets the token; useful only once)
string_t& get_string()
{
// a number token holds '.' regardless of the locale (#4084)
@@ -2290,11 +2325,11 @@ scan_number_done:
/// the position of the decimal point in token_buffer
std::size_t decimal_point_position = std::string::npos;
/// whether the caller only needs the token types and never looks at the
/// converted numeric values; set only by accept(), which parses through
/// json_sax_acceptor. When set, convert_number() skips converting integer
/// tokens whose digit count guarantees that they fit into 64 bits (see
/// there)
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
/// token classification and never looks at the converted numeric value;
/// when set, scan_number() may skip strtoull()/strtoll() for
/// value_unsigned/value_integer tokens whose digit count guarantees they
/// fit into 64 bits (see scan_number())
const bool discard_number_values = false;
};
+43 -50
View File
@@ -54,8 +54,7 @@ using parser_callback_t =
/*!
@brief syntax analysis
This class implements an iterative parser that keeps the open containers on
an explicit stack and reports what it reads as SAX events.
This class implements a recursive descent parser.
*/
template<typename BasicJsonType, typename InputAdapterType>
class parser
@@ -99,9 +98,28 @@ class parser
if (callback)
{
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
sax_parse_internal(&sdp);
if (strict)
{
// in strict mode, input must be completely read
if (get_token() != token_type::end_of_input)
{
sdp.parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(),
exception_message(token_type::end_of_input, "value"), nullptr));
}
}
else
{
// the caller keeps using the input: position it right after
// the value by leaving the character that terminated it
m_lexer.release_lookahead();
}
// in case of an error, return a discarded value
if (!parse_dom(sdp, strict))
if (sdp.is_errored())
{
result = value_t::discarded;
return;
@@ -117,9 +135,26 @@ class parser
else
{
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
sax_parse_internal(&sdp);
if (strict)
{
// in strict mode, input must be completely read
if (get_token() != token_type::end_of_input)
{
sdp.parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr));
}
}
else
{
// see above
m_lexer.release_lookahead();
}
// in case of an error, return a discarded value
if (!parse_dom(sdp, strict))
if (sdp.is_errored())
{
result = value_t::discarded;
return;
@@ -172,46 +207,6 @@ class parser
}
private:
/*!
@brief run a DOM SAX parser to completion and position the lexer
Shared by both branches of @ref parse(): builds no SAX parser itself,
but drives an already-constructed @a json_sax_dom_parser or
@ref json_sax_dom_callback_parser through @ref sax_parse_internal(),
then applies the strict-EOF check (reporting parse_error.101 through
@a sdp on failure) or, in non-strict mode, releases the lookahead so
the caller can keep reading the input right after the parsed value.
@param[in,out] sdp the DOM SAX parser to run
@param[in] strict whether to expect the last token to be EOF
@return whether @a sdp did not report an error
*/
template<typename DomSax>
bool parse_dom(DomSax& sdp, const bool strict)
{
sax_parse_internal(&sdp);
if (strict)
{
// in strict mode, input must be completely read
if (get_token() != token_type::end_of_input)
{
sdp.parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(),
exception_message(token_type::end_of_input, "value"), nullptr));
}
}
else
{
// the caller keeps using the input: position it right after
// the value by leaving the character that terminated it
m_lexer.release_lookahead();
}
return !sdp.is_errored();
}
template<typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse_internal(SAX* sax)
@@ -444,9 +439,8 @@ class parser
// We are done with this array. Before we can parse a
// new value, we need to evaluate the new state first.
// By setting skip_to_state_evaluation to true, the next
// iteration skips parsing a value and evaluates the
// enclosing state directly.
// By setting skip_to_state_evaluation to false, we
// are effectively jumping to the beginning of this if.
JSON_ASSERT(!states.empty());
states.pop_back();
skip_to_state_evaluation = true;
@@ -506,9 +500,8 @@ class parser
// We are done with this object. Before we can parse a
// new value, we need to evaluate the new state first.
// By setting skip_to_state_evaluation to true, the next
// iteration skips parsing a value and evaluates the
// enclosing state directly.
// By setting skip_to_state_evaluation to false, we
// are effectively jumping to the beginning of this if.
JSON_ASSERT(!states.empty());
states.pop_back();
skip_to_state_evaluation = true;
@@ -39,7 +39,6 @@ inline std::size_t concat_length(const char /*c*/, const Args& ... rest)
template<typename... Args>
inline std::size_t concat_length(const char* cstr, const Args& ... rest)
{
// cppcheck-suppress ignoredReturnValue
return ::strlen(cstr) + concat_length(rest...);
}
+5 -72
View File
@@ -36,61 +36,6 @@ StringType to_string(std::size_t value)
return result;
}
///////////////////
// UTF-8 encoding //
///////////////////
/*!
@brief encode a Unicode code point as UTF-8
Used to turn a decoded code point back into bytes: by the wide-string input
adapters in input_adapters.hpp (one code point per UTF-32 unit, per UTF-16
unit outside the surrogate range, and per valid UTF-16 surrogate pair), and
by the lexer's `\uXXXX`/`\uXXXX\uYYYY` handling in lexer.hpp. Passing a
code point above U+10FFFF, or one in the surrogate range U+D800..U+DFFF, is
undefined behavior; callers are expected to have rejected those already
(the wide-string adapters pass malformed units through unencoded instead of
calling this function, and the lexer rejects unpaired surrogates before
reaching it).
@tparam Out a callable invoked with one byte (as std::uint32_t, 0x00..0xFF)
at a time, most significant byte first
@param[in] cp the code point to encode (at most U+10FFFF)
@param[in] out called once for each byte of the UTF-8 encoding of @a cp
*/
template<typename Out>
void encode_utf8(std::uint32_t cp, Out&& out)
{
JSON_ASSERT(cp <= 0x10FFFF);
if (cp < 0x80)
{
// 1-byte characters: 0xxxxxxx (ASCII)
out(cp);
}
else if (cp <= 0x7FF)
{
// 2-byte characters: 110xxxxx 10xxxxxx
out(0xC0u | (cp >> 6u));
out(0x80u | (cp & 0x3Fu));
}
else if (cp <= 0xFFFF)
{
// 3-byte characters: 1110xxxx 10xxxxxx 10xxxxxx
out(0xE0u | (cp >> 12u));
out(0x80u | ((cp >> 6u) & 0x3Fu));
out(0x80u | (cp & 0x3Fu));
}
else
{
// 4-byte characters: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
out(0xF0u | (cp >> 18u));
out(0x80u | ((cp >> 12u) & 0x3Fu));
out(0x80u | ((cp >> 6u) & 0x3Fu));
out(0x80u | (cp & 0x3Fu));
}
}
///////////////////
// UTF-8 decoding //
///////////////////
@@ -106,23 +51,11 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
written by Björn Hoehrmann. See
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
places, which differ in speed, diagnostics, and how they read the input:
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
text strings at decode time).
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
diagnostic for each kind of error.
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
bulk string scan, the bulk path of the BON8 reader, and the BON8 writer.
They must accept exactly what the lexer's switch accepts.
- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk
access, and the bytes the bulk path leaves to it.
All four must accept the same set of sequences, so a change to one needs a
matching change to the others.
This decoder is the single source of truth for UTF-8 validation in this
library: it is used both by the serializer (to escape and, in strict mode,
reject ill-formed UTF-8 when dumping a string) and by the binary readers
(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at
decode time; see @ref is_valid_utf8 below).
@param[in,out] state the current decoder state
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
+249 -394
View File
File diff suppressed because it is too large Load Diff
+68 -115
View File
@@ -70,14 +70,41 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
return *this;
}
private:
/// @brief find the entry for @a key, for either constness of @a self
/// @note the single place that performs the linear key search
template<typename Self, typename KeyType>
static auto find_impl(Self& self, KeyType&& key) -> decltype(self.begin())
{
for (auto it = self.begin(); it != self.end(); ++it)
{
if (self.m_compare(it->first, key))
{
return it;
}
}
return self.end();
}
/// @brief remove the entry @a it points to, preserving order
/// @note keys are not movable, so the tail is destroyed and re-constructed in place
void erase_at(iterator it)
{
for (auto next = it; ++next != this->end(); ++it)
{
it->~value_type(); // Destroy but keep allocation
new (&*it) value_type{std::move(*next)};
}
Container::pop_back();
}
public:
std::pair<iterator, bool> emplace(const key_type& key, T&& t)
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it != this->end())
{
if (m_compare(it->first, key))
{
return {it, false};
}
return {it, false};
}
Container::emplace_back(key, std::forward<T>(t));
return {std::prev(this->end()), true};
@@ -87,12 +114,10 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
detail::is_usable_as_key_type<key_compare, key_type, KeyType>::value, int> = 0>
std::pair<iterator, bool> emplace(KeyType && key, T && t)
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it != this->end())
{
if (m_compare(it->first, key))
{
return {it, false};
}
return {it, false};
}
Container::emplace_back(std::forward<KeyType>(key), std::forward<T>(t));
return {std::prev(this->end()), true};
@@ -124,75 +149,55 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
T& at(const key_type& key)
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it == this->end())
{
if (m_compare(it->first, key))
{
return it->second;
}
JSON_THROW(std::out_of_range("key not found"));
}
JSON_THROW(std::out_of_range("key not found"));
return it->second;
}
template<class KeyType, detail::enable_if_t<
detail::is_usable_as_key_type<key_compare, key_type, KeyType>::value, int> = 0>
T & at(KeyType && key) // NOLINT(cppcoreguidelines-missing-std-forward)
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it == this->end())
{
if (m_compare(it->first, key))
{
return it->second;
}
JSON_THROW(std::out_of_range("key not found"));
}
JSON_THROW(std::out_of_range("key not found"));
return it->second;
}
const T& at(const key_type& key) const
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it == this->end())
{
if (m_compare(it->first, key))
{
return it->second;
}
JSON_THROW(std::out_of_range("key not found"));
}
JSON_THROW(std::out_of_range("key not found"));
return it->second;
}
template<class KeyType, detail::enable_if_t<
detail::is_usable_as_key_type<key_compare, key_type, KeyType>::value, int> = 0>
const T & at(KeyType && key) const // NOLINT(cppcoreguidelines-missing-std-forward)
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it == this->end())
{
if (m_compare(it->first, key))
{
return it->second;
}
JSON_THROW(std::out_of_range("key not found"));
}
JSON_THROW(std::out_of_range("key not found"));
return it->second;
}
size_type erase(const key_type& key)
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it != this->end())
{
if (m_compare(it->first, key))
{
// Since we cannot move const Keys, re-construct them in place
for (auto next = it; ++next != this->end(); ++it)
{
it->~value_type(); // Destroy but keep allocation
new (&*it) value_type{std::move(*next)};
}
Container::pop_back();
return 1;
}
erase_at(it);
return 1;
}
return 0;
}
@@ -201,19 +206,11 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
detail::is_usable_as_key_type<key_compare, key_type, KeyType>::value, int> = 0>
size_type erase(KeyType && key) // NOLINT(cppcoreguidelines-missing-std-forward)
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, key);
if (it != this->end())
{
if (m_compare(it->first, key))
{
// Since we cannot move const Keys, re-construct them in place
for (auto next = it; ++next != this->end(); ++it)
{
it->~value_type(); // Destroy but keep allocation
new (&*it) value_type{std::move(*next)};
}
Container::pop_back();
return 1;
}
erase_at(it);
return 1;
}
return 0;
}
@@ -278,80 +275,38 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
size_type count(const key_type& key) const
{
for (auto it = this->begin(); it != this->end(); ++it)
{
if (m_compare(it->first, key))
{
return 1;
}
}
return 0;
return find_impl(*this, key) != this->end() ? 1 : 0;
}
template<class KeyType, detail::enable_if_t<
detail::is_usable_as_key_type<key_compare, key_type, KeyType>::value, int> = 0>
size_type count(KeyType && key) const // NOLINT(cppcoreguidelines-missing-std-forward)
{
for (auto it = this->begin(); it != this->end(); ++it)
{
if (m_compare(it->first, key))
{
return 1;
}
}
return 0;
return find_impl(*this, key) != this->end() ? 1 : 0;
}
iterator find(const key_type& key)
{
for (auto it = this->begin(); it != this->end(); ++it)
{
if (m_compare(it->first, key))
{
return it;
}
}
return Container::end();
return find_impl(*this, key);
}
template<class KeyType, detail::enable_if_t<
detail::is_usable_as_key_type<key_compare, key_type, KeyType>::value, int> = 0>
iterator find(KeyType && key) // NOLINT(cppcoreguidelines-missing-std-forward)
{
for (auto it = this->begin(); it != this->end(); ++it)
{
if (m_compare(it->first, key))
{
return it;
}
}
return Container::end();
return find_impl(*this, key);
}
const_iterator find(const key_type& key) const
{
for (auto it = this->begin(); it != this->end(); ++it)
{
if (m_compare(it->first, key))
{
return it;
}
}
return Container::end();
return find_impl(*this, key);
}
template<class KeyType, detail::enable_if_t<
detail::is_usable_as_key_type<key_compare, key_type, KeyType>::value, int> = 0>
const_iterator find(KeyType && key) const // NOLINT(cppcoreguidelines-missing-std-forward)
{
for (auto it = this->begin(); it != this->end(); ++it)
{
if (m_compare(it->first, key))
{
return it;
}
}
return Container::end();
return find_impl(*this, key);
}
std::pair<iterator, bool> insert( value_type&& value )
@@ -361,12 +316,10 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
std::pair<iterator, bool> insert( const value_type& value )
{
for (auto it = this->begin(); it != this->end(); ++it)
const auto it = find_impl(*this, value.first);
if (it != this->end())
{
if (m_compare(it->first, value.first))
{
return {it, false};
}
return {it, false};
}
Container::push_back(value);
return {--this->end(), true};
File diff suppressed because it is too large Load Diff
+3 -5
View File
@@ -2525,12 +2525,10 @@ TEST_CASE("diagnostic positions: value lifetime, input adapters, and SAX")
SECTION("move constructor resets the moved-from value to npos")
{
// basic_json(basic_json&&) (json.hpp, around line 1265) copies
// basic_json(basic_json&&) (json.hpp, around line 1944) copies
// other's start_position/end_position into *this and then resets
// other's to npos (see the cppcheck-suppress[accessForwarded]
// annotation there, which flags this reset as worth a second
// look). Only the top-level moved-from value is affected; its
// (moved-away) children are gone along with it.
// other's to npos. Only the top-level moved-from value is
// affected; its (moved-away) children are gone along with it.
const std::string s = R"({"a":1,"b":[1,2,3]})";
json a = json::parse(s);
const auto a_start = a.start_pos();
+43
View File
@@ -618,6 +618,49 @@ TEST_CASE("modifiers")
}
}
SECTION("rvalue at position moves rather than copies")
{
// regression test: insert(pos, basic_json&&) used to forward to
// insert(pos, const basic_json&) because the named rvalue
// reference parameter is itself an lvalue, so it always
// deep-copied its argument instead of moving it
json j_big = std::string(1000, 'x');
const auto* const original_buffer = j_big.get_ref<const std::string&>().data();
auto it = j_array.insert(j_array.begin(), std::move(j_big));
CHECK(j_array.size() == 5);
CHECK(*it == json(std::string(1000, 'x')));
CHECK((*it).get_ref<const std::string&>().data() == original_buffer);
// the moved-from value is null, the same as after push_back(&&)
CHECK(j_big.is_null()); // NOLINT(bugprone-use-after-move,hicpp-invalid-access-moved)
}
SECTION("self-aliasing insertion")
{
SECTION("without reallocation")
{
json j_self = {1, 2, 3, 4};
j_self.get_ref<json::array_t&>().reserve(j_self.size() + 1);
auto it = j_self.insert(j_self.begin(), std::move(j_self[1]));
CHECK(j_self.size() == 5);
CHECK(*it == json(2));
CHECK(j_self == json({2, 1, nullptr, 3, 4}));
}
SECTION("with reallocation")
{
json j_self = {1, 2, 3, 4};
j_self.get_ref<json::array_t&>().shrink_to_fit();
auto it = j_self.insert(j_self.begin(), std::move(j_self[1]));
CHECK(j_self.size() == 5);
CHECK(*it == json(2));
CHECK(j_self == json({2, 1, nullptr, 3, 4}));
}
}
SECTION("copies at position")
{
SECTION("insert before begin()")