Compare commits

...
Author SHA1 Message Date
Niels Lohmann e96e2982a5 Fix CI: useless cast in the growth of the view's node index
GCC -Werror=useless-cast (ci_test_gcc on Linux x86-64) rejected
static_cast<std::size_t>(guess + (guess / 4) + 64): the sum is a
std::uint64_t prvalue, the same type as std::size_t there, while the cast
is needed where std::size_t is 32 bits wide. Cast a named variable
instead, which GCC does not report. The build stopped at an earlier error
before, so the previous CI run did not show this one.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 0a64e6c99b Fix CI: MSVC C4127 in the view builder and the single-header test build
- msvc (Win32, /W4 /WX) reported C4127 (conditional expression is
  constant) for `TrailingCommas && cur() == ']'` and the like when the
  option is off. Route the template arguments through a static enabled()
  function, as json.hpp's nesting_depth_exhausted() does.
- ci_test_single_header compiled unit-json_view_builder.cpp against
  single_include/, which does not contain the internal
  nlohmann/detail/view headers. Build that test only with the multiple
  headers.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann d8b8c2498f Test the C++11 stand-in for std::string_view of json_view
Six members of string_ref (length, begin, end, operator[], operator!=,
and operator<<) were not reached before C++17, where string_ref is
std::string_view.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 2fb66ba859 Remove the eight-digit case of the view's digit parser
Since whole blocks of eight digits are read directly, parse_upto8() only
gets fewer than eight digits; its eight-digit case was dead code.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann ab49a7b1bf Address the cpplint findings of the view's parser
The exponent of the overflow check is an std::int64_t instead of a long
(runtime/int).

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 154f240022 Read whole blocks of eight digits of json_view directly
A block of eight digits of a number token lies inside the input (the
digits were counted while scanning, or are recorded in the digit
layout), so parse_upto19() reads it without the bounds check of the
last, partial block. Traversing canada.json: -11% instructions, -6%
cycles.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 8d3280de1e Classify a NUL inside a string of json_view as a control character
json::parse ends the input at a NUL only between values (where it does
at all); inside a string, a NUL is a control character that must be
escaped. The view reported it as a missing closing quote. (Only the
error code differed: the exception comes from the library parser.)

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 118af5015e Test the overflow check of json_view at the largest double
Numbers of more than 19 digits at the boundary of the largest double are
decided by the exact comparison with the midpoint in the overflow check.
The check uses the floating-point type of the document: with float, the
view rejects what parse() rejects (1e39, 3.4028236e38, the midpoint between
the largest float and 2^128), and double documents are not affected.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 9b0b4ff26d Address the clang-tidy findings of the view's parser
- tables as std::array; the frames of the first 64 levels stay a C array
  (not initialized on purpose, NOLINT)
- \u escapes are decoded with the library's hex_codepoint() instead of a
  second table
- the parse failure is private, with an accessor; the special member
  functions of the builder are all declared
- no nested conditional operators; explicit parentheses; a repeated
  branch body merged; auto for casts
- the test's C arrays, fixed seed, and escaped literals are marked, as in
  the other tests

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 694bfd1b4e Keep the arguments of the view's throw helpers used without exceptions
With JSON_NOEXCEPTION, NLOHMANN_VIEW_THROW(e) was std::abort() alone, so
the parameters of the functions that build the exceptions were unused, a
warning that the builds with -Werror turn into an error. The exception is
now evaluated before std::abort(); the program ends anyway.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann ad82821c89 Add the one-pass parser of json_view (internal)
The builder parses JSON text in one pass into the node index: strings and
numbers stay in the source (escaped strings are decoded into an arena),
integers are converted while their digits are in cache, and floats keep
their digit layout for a later conversion. It accepts exactly what
json::parse accepts, for every combination of comments and trailing commas,
with and without a terminating NUL, and with JSON_STRICT_NUL_HANDLING.

Parse state lives in a local cursor whose address never escapes, so that it
stays in registers; out-of-line helpers (errors, regrowth, escapes,
comments) are members of the builder and get the positions they need. The
value dispatch is expanded once for array elements and once for member
values. Literals are compared with memcmp and words read in a fixed byte
order, so nothing depends on the platform's byte order. Error messages come
with the public classes.

Tests (unit-json_view_builder.cpp): accept/reject and values against
json::parse for handwritten, generated, and damaged documents under all
option combinations, from std::string and from exact-size buffers (no read
past the input under AddressSanitizer), deep nesting up to 100,000 levels,
NUL/BOM/whitespace cases, and the test-suite files.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann 1bf0b1b6c2 Add the node index and scanning primitives of json_view (internal)
Internal parts of the zero-copy view (#5295), under detail/view and not
included by json.hpp, so that users of json.hpp compile nothing of it:

- macro_scope.hpp/macro_unscope.hpp: the few macros the view needs, under
  its own prefix (json.hpp undefines its own at its end); the throw macro
  honors JSON_NOEXCEPTION and JSON_THROW_USER like JSON_THROW
- string_ref.hpp: std::string_view from C++17 on, else a small stand-in
- node.hpp: the 16-byte node of the index; its kinds are value_t values
  (checked by a static_assert)
- document_data.hpp: the storage of a parsed document (node array, decode
  arena, owned input)
- scan.hpp: string and digit scanning with unrolled checks at fixed offsets
  (after yyjson) and the library's SWAR and UTF-8 checks, independent of the
  byte order

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann b88e5f9107 Keep the behavior-changing configuration readable after json.hpp
json.hpp undefines JSON_STRICT_NUL_HANDLING and
JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON at its end. Code that builds on
the library after it, such as the planned json_view.hpp, reads them from
detail::abi_config instead. The constants live in the ABI namespace, which
already encodes both settings, so they always match the basic_json in use.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:13 +02:00
Niels Lohmann d23803fd32 Decode \u escapes with a table in the lexer
get_codepoint() read the four hex digits of a \u escape with four calls
to get(), each classified by a chain of range comparisons. For contiguous
input, get_codepoint_bulk() now decodes them with one lookup per byte
(hex_codepoint() in string_scan.hpp, after yyjson's read_hex_u16): a
256-entry table maps a byte to its value, or 0xFF for anything else, and
an invalid digit shows in the OR of the four values. It then skips the
four bytes and updates the position counters as four get() calls would.
If a digit is invalid or fewer than four bytes are left, it changes
nothing and the existing loop runs, so errors are reported with the same
message and position as before.

json::parse, best of 5 runs in separate processes (M1 Max): the escaped
twitter.json (every non-ASCII character as \u) -13.6%, all other files
within 0.3%.

Tests compare the contiguous and the streaming path (value or exception
message) for valid escapes, surrogate pairs, truncated and invalid digits
at every position, and 3,000 seeded random escapes.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-30 20:19:12 +02:00
16 changed files with 2495 additions and 2 deletions
+8
View File
@@ -20,6 +20,7 @@ cc_library(
hdrs = [
"include/nlohmann/adl_serializer.hpp",
"include/nlohmann/byte_container_with_subtype.hpp",
"include/nlohmann/detail/abi_config.hpp",
"include/nlohmann/detail/abi_macros.hpp",
"include/nlohmann/detail/bit_ops.hpp",
"include/nlohmann/detail/conversions/from_json.hpp",
@@ -65,6 +66,13 @@ cc_library(
"include/nlohmann/detail/string_escape.hpp",
"include/nlohmann/detail/string_utils.hpp",
"include/nlohmann/detail/value_t.hpp",
"include/nlohmann/detail/view/builder.hpp",
"include/nlohmann/detail/view/document_data.hpp",
"include/nlohmann/detail/view/macro_scope.hpp",
"include/nlohmann/detail/view/macro_unscope.hpp",
"include/nlohmann/detail/view/node.hpp",
"include/nlohmann/detail/view/scan.hpp",
"include/nlohmann/detail/view/string_ref.hpp",
"include/nlohmann/json.hpp",
"include/nlohmann/json_fwd.hpp",
"include/nlohmann/json_literals.hpp",
+34
View File
@@ -0,0 +1,34 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief the configuration macros that change the library's behavior
json.hpp undefines these macros at its end (see macro_unscope.hpp), so code
that builds on the library after it (json_view.hpp) reads them here. Like the
macros, they are part of the ABI namespace, so they always match the
basic_json they are used with.
*/
struct abi_config
{
/// JSON_STRICT_NUL_HANDLING: a null byte is an error, not the end of input
static constexpr bool strict_nul_handling = JSON_STRICT_NUL_HANDLING != 0;
/// JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
static constexpr bool legacy_discarded_value_comparison = JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON != 0;
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+46
View File
@@ -219,6 +219,44 @@ class lexer : public lexer_base<BasicJsonType>
// scan functions
/////////////////////
/// contiguous input: try to decode the 4 hex digits following `\u`
/// directly from the input buffer via hex_codepoint(), instead of 4 calls
/// to get(). On success, advances the adapter and the position counters
/// exactly as those 4 get() calls would (a hex digit is never '\n', so
/// only the flat counters move) and leaves @a current holding the last of
/// the 4 digits, just as the last such get() would; the codepoint is
/// written to @a out. Makes no state change and returns false - for a
/// pending unget, fewer than 4 remaining bytes, or any of the 4 bytes not
/// being a hex digit - so the caller falls back unchanged to the
/// per-character loop, which then reports the same diagnostic (stopping
/// at the first invalid digit) as before this optimization.
bool get_codepoint_bulk(std::true_type /*bulk*/, int& out)
{
if (next_unget || ia.bulk_remaining() < 4)
{
return false;
}
const char_type* const raw = ia.bulk_data();
const int codepoint = hex_codepoint(reinterpret_cast<const unsigned char*>(raw));
if (codepoint < 0)
{
return false;
}
ia.bulk_skip(4);
// a hex digit is never a newline, so only the flat counters advance
position.chars_read_total += 4;
position.chars_read_current_line += 4;
current = char_traits<char_type>::to_int_type(raw[3]);
out = codepoint;
return true;
}
/// streaming input: no bulk fast path
bool get_codepoint_bulk(std::false_type /*bulk*/, int& /*out*/) const noexcept
{
return false;
}
/*!
@brief get codepoint from 4 hex characters following `\u`
@@ -238,6 +276,14 @@ class lexer : public lexer_base<BasicJsonType>
{
// this function only makes sense after reading `\u`
JSON_ASSERT(current == 'u');
// contiguous input: decode all 4 hex digits directly from the buffer
int fast_codepoint = 0;
if (get_codepoint_bulk(std::integral_constant<bool, bulk_scan> {}, fast_codepoint))
{
return fast_codepoint;
}
int codepoint = 0;
const auto factors = { 12u, 8u, 4u, 0u };
+47 -1
View File
@@ -8,8 +8,9 @@
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint64_t
#include <cstdint> // uint64_t, uint8_t
#include <cstring> // memcpy
#include <nlohmann/detail/bit_ops.hpp>
@@ -315,5 +316,50 @@ inline std::size_t string_bulk_run(const unsigned char* data, std::size_t n) noe
return scalar_string_bulk_run(data, n);
}
// Decode the 4 hex digits at [data, data+4) - the digits following a `\u`
// escape - into a codepoint 0x0000..0xFFFF via one table lookup per byte
// (after yyjson's read_hex_u16), or return -1 if any of the 4 bytes is not a
// hex digit ('0'..'9', 'A'..'F', 'a'..'f'). The caller must already have
// checked that 4 bytes are available; used by lexer::get_codepoint()'s
// contiguous fast path. On -1 it falls back to the byte-at-a-time loop, which
// stops at the first invalid digit, so the reported error and position are
// unaffected by this fast path.
inline int hex_codepoint(const unsigned char* data) noexcept
{
static const std::array<std::uint8_t, 256> hex_digit_table = // NOLINT(cppcoreguidelines-avoid-non-const-global-variables)
{
{
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 00..0F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 10..1F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 20..2F
0x00, 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08, 0x09, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 30..3F ('0'..'9')
0xFF, 0x0A, 0x0B, 0x0C, 0x0D, 0x0E, 0x0F, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 40..4F ('A'..'F')
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 50..5F
0xFF, 0x0A, 0x0B, 0x0C, 0x0D, 0x0E, 0x0F, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 60..6F ('a'..'f')
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 70..7F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 80..8F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 90..9F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // A0..AF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // B0..BF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // C0..CF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // D0..DF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // E0..EF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF // F0..FF
}
};
const std::uint8_t d0 = hex_digit_table[data[0]];
const std::uint8_t d1 = hex_digit_table[data[1]];
const std::uint8_t d2 = hex_digit_table[data[2]];
const std::uint8_t d3 = hex_digit_table[data[3]];
// every valid digit is <= 0xF; the combined OR only exceeds it if at
// least one of the four bytes was not a hex digit (looked up as 0xFF)
if ((d0 | d1 | d2 | d3) > 0x0F)
{
return -1;
}
return (d0 << 12) | (d1 << 8) | (d2 << 4) | d3;
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,130 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstring> // memcpy
#include <new> // operator new, placement new
#include <string> // string
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#include <nlohmann/detail/view/node.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
/// storage of a parsed document; heap-allocated (header and an initial node
/// array in one block) so that views survive moves of the owning document
struct document_data
{
const char* src = nullptr;
std::size_t size = 0;
node* tape = nullptr;
std::size_t tape_size = 0;
std::size_t tape_cap = 0;
node* inline_tape = nullptr; ///< node array allocated together with this header
std::size_t inline_cap = 0;
std::string arena{}; ///< decoded strings that contained escapes // NOLINT(readability-redundant-member-init)
std::string owned{}; ///< owned copy of the input, if any // NOLINT(readability-redundant-member-init)
std::array<const char*, 4> base = {{nullptr, nullptr, nullptr, nullptr}}; ///< string bases: source, arena (indexed by flags & node_flags::storage)
bool discarded = true;
/// one allocation for the header and room for `nodes` nodes; large
/// documents get a separate node array instead (so it can be trimmed)
static document_data* create(std::size_t nodes)
{
nodes = nodes <= 256 ? nodes : 0;
void* mem = ::operator new (sizeof(document_data) + (nodes * sizeof(node)));
auto* d = new (mem) document_data(); // NOLINT(cppcoreguidelines-owning-memory): owned by the returned pointer, freed by deleter
// (aligned: sizeof is a multiple of the alignment; through void*, as GCC's -Wcast-align wants)
d->inline_tape = static_cast<node*>(static_cast<void*>(static_cast<char*>(mem) + sizeof(document_data))); // NOLINT(bugprone-casting-through-void)
d->inline_cap = nodes;
d->tape = d->inline_tape;
d->tape_cap = nodes;
return d;
}
struct deleter
{
void operator()(document_data* d) const noexcept
{
d->~document_data();
::operator delete (d);
}
};
document_data() noexcept = default;
document_data(const document_data&) = delete;
document_data(document_data&&) = delete;
document_data& operator=(const document_data&) = delete;
document_data& operator=(document_data&&) = delete;
~document_data()
{
release();
}
void release() noexcept
{
if (tape != inline_tape)
{
::operator delete (tape);
}
tape = inline_tape;
tape_cap = inline_cap;
}
/// make room for n nodes; keeps the first tape_size nodes
void reserve(std::size_t n)
{
if (n <= tape_cap)
{
return;
}
node* fresh = static_cast<node*>(::operator new (n * sizeof(node)));
if (tape_size != 0)
{
std::memcpy(fresh, tape, tape_size * sizeof(node));
}
release();
tape = fresh;
tape_cap = n;
}
const char* str(const node& n) const noexcept
{
return base[n.flags & node_flags::storage] + n.off;
}
/// the node after n's subtree (containers span `next` nodes, scalars one)
static NLOHMANN_VIEW_ALWAYS_INLINE const node* after(const node* n) noexcept
{
return n + (is_container(*n) ? n->next : 1u);
}
/// first element (array) or first key (object) of a container
static NLOHMANN_VIEW_ALWAYS_INLINE const node* first_child(const node* n) noexcept
{
return n + 1;
}
/// end of the elements of a container
static NLOHMANN_VIEW_ALWAYS_INLINE const node* child_end(const node* n) noexcept
{
return n + n->next;
}
};
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -0,0 +1,63 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
// Macros of json_view.hpp and its detail headers. json.hpp undefines its own
// macros at its end (macro_unscope.hpp), so the view defines the few it needs
// under its own prefix; json_view.hpp undefines them all at its end
// (detail/view/macro_unscope.hpp). Configuration that json.hpp undefines is
// read from detail::abi_config instead.
#if (defined(__cplusplus) && __cplusplus >= 201703L) || (defined(_MSVC_LANG) && _MSVC_LANG >= 201703L)
#define NLOHMANN_VIEW_HAS_CPP_17 1
#else
#define NLOHMANN_VIEW_HAS_CPP_17 0
#endif
#if defined(__GNUC__) || defined(__clang__)
#define NLOHMANN_VIEW_LIKELY(x) __builtin_expect(!!(x), 1)
#define NLOHMANN_VIEW_UNLIKELY(x) __builtin_expect(!!(x), 0)
#define NLOHMANN_VIEW_ALWAYS_INLINE inline __attribute__((always_inline))
#define NLOHMANN_VIEW_NOINLINE __attribute__((noinline))
#elif defined(_MSC_VER)
#define NLOHMANN_VIEW_LIKELY(x) (x)
#define NLOHMANN_VIEW_UNLIKELY(x) (x)
#define NLOHMANN_VIEW_ALWAYS_INLINE __forceinline
#define NLOHMANN_VIEW_NOINLINE __declspec(noinline)
#else
#define NLOHMANN_VIEW_LIKELY(x) (x)
#define NLOHMANN_VIEW_UNLIKELY(x) (x)
#define NLOHMANN_VIEW_ALWAYS_INLINE inline
#define NLOHMANN_VIEW_NOINLINE
#endif
// exceptions as in json.hpp (JSON_NOEXCEPTION, JSON_THROW_USER)
#if (defined(__cpp_exceptions) || defined(__EXCEPTIONS) || defined(_CPPUNWIND)) && !defined(JSON_NOEXCEPTION)
#define NLOHMANN_VIEW_THROW(exception) throw exception
#else
#include <cstdlib>
// (the exception is built first, so that the arguments of the throwing
// helpers count as used; the program ends anyway)
#define NLOHMANN_VIEW_THROW(exception) (static_cast<void>(exception), std::abort())
#endif
#if defined(JSON_THROW_USER)
#undef NLOHMANN_VIEW_THROW
#define NLOHMANN_VIEW_THROW JSON_THROW_USER
#endif
// the parser stores a node's first word at once where the layout of `node` is
// known to be little-endian (MSVC targets are); elsewhere field by field
#if (defined(__BYTE_ORDER__) && defined(__ORDER_LITTLE_ENDIAN__) && __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__) || defined(_MSC_VER)
#define NLOHMANN_VIEW_LITTLE_ENDIAN 1
#else
#define NLOHMANN_VIEW_LITTLE_ENDIAN 0
#endif
/// sixteen checks at fixed offsets 0..15
#define NLOHMANN_VIEW_REPEAT16(X) X(0) X(1) X(2) X(3) X(4) X(5) X(6) X(7) X(8) X(9) X(10) X(11) X(12) X(13) X(14) X(15)
@@ -0,0 +1,20 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
// undefine the macros of detail/view/macro_scope.hpp (at the end of json_view.hpp)
#undef NLOHMANN_VIEW_HAS_CPP_17
#undef NLOHMANN_VIEW_LIKELY
#undef NLOHMANN_VIEW_UNLIKELY
#undef NLOHMANN_VIEW_ALWAYS_INLINE
#undef NLOHMANN_VIEW_NOINLINE
#undef NLOHMANN_VIEW_THROW
#undef NLOHMANN_VIEW_LITTLE_ENDIAN
#undef NLOHMANN_VIEW_REPEAT16
+97
View File
@@ -0,0 +1,97 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint16_t, uint32_t, uint64_t
#include <cstring> // memcpy
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
// the node kinds are value_t values; the tests of is_container() and of the
// number kinds depend on this numbering
static_assert(static_cast<std::uint8_t>(value_t::null) == 0 && static_cast<std::uint8_t>(value_t::object) == 1
&& static_cast<std::uint8_t>(value_t::array) == 2 && static_cast<std::uint8_t>(value_t::string) == 3
&& static_cast<std::uint8_t>(value_t::boolean) == 4 && static_cast<std::uint8_t>(value_t::number_integer) == 5
&& static_cast<std::uint8_t>(value_t::number_unsigned) == 6 && static_cast<std::uint8_t>(value_t::number_float) == 7,
"the node format depends on the numbering of value_t");
/// node flags
struct node_flags
{
static constexpr std::uint8_t escaped = 1; ///< string payload lives in the decode arena, not the source
static constexpr std::uint8_t storage = 3; ///< mask: where a string or number token lives (index into document_data::base)
static constexpr std::uint8_t is_true = 4; ///< boolean value
};
/// One entry of the flat index, in document order. An object's members are
/// stored as key node followed by the value's subtree. Integers keep their
/// converted 64-bit value in the len/next bytes (the node after a scalar is
/// always the next one, and the token length follows from `extra`).
struct node
{
std::uint8_t kind; ///< value_t
std::uint8_t flags; ///< node_flags
std::uint16_t extra; ///< numbers: integer digits (low byte) and fraction digits (high byte), 255 = "many"; otherwise 0
std::uint32_t off; ///< source offset (string content, number token, literal, bracket); arena offset if node_flags::escaped
std::uint32_t len; ///< string: decoded bytes; float: token bytes; array/object: element count
std::uint32_t next; ///< array/object: number of nodes of the subtree (its extent in the enclosing sequence)
};
static_assert(sizeof(node) == 16, "node must stay 16 bytes");
NLOHMANN_VIEW_ALWAYS_INLINE bool is_container(const node& n) noexcept
{
return static_cast<unsigned>(n.kind) - 1u <= 1u;
}
/// the converted value of an integer node (stored in len/next)
NLOHMANN_VIEW_ALWAYS_INLINE std::uint64_t integer_bits(const node& n) noexcept
{
std::uint64_t v = 0;
std::memcpy(&v, reinterpret_cast<const unsigned char*>(&n) + 8, 8); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
return v;
}
NLOHMANN_VIEW_ALWAYS_INLINE void set_integer_bits(node& n, std::uint64_t v) noexcept
{
std::memcpy(reinterpret_cast<unsigned char*>(&n) + 8, &v, 8); // NOLINT(cppcoreguidelines-pro-type-reinterpret-cast)
}
/// token length of a number node
NLOHMANN_VIEW_ALWAYS_INLINE std::uint32_t number_length(const node& n) noexcept
{
return n.kind == static_cast<std::uint8_t>(value_t::number_float) ? n.len
: (n.extra & 0xFFu) + (n.kind == static_cast<std::uint8_t>(value_t::number_integer) ? 1u : 0u);
}
/// estimated number of nodes for an input of `size` bytes (one node per ~12
/// bytes covers typical documents without regrowth)
inline std::size_t estimate_nodes(std::size_t size) noexcept
{
return (size / 12) + 16;
}
/// estimated number of nodes for the input [src, src + size): pretty-printed
/// input (whitespace after the first byte) needs about a node per 12 bytes,
/// minified input up to one per 4 (yyjson tells the two apart the same way)
inline std::size_t estimate_nodes(const char* src, std::size_t size) noexcept
{
return size >= 2 && (src[1] == ' ' || src[1] == '\n' || src[1] == '\r' || src[1] == '\t') ? estimate_nodes(size) : (size / 4) + 16;
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+189
View File
@@ -0,0 +1,189 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-FileCopyrightText: 2020 YaoYuan <https://github.com/ibireme/yyjson>
// SPDX-License-Identifier: MIT
#pragma once
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint8_t, uint16_t, uint64_t
#include <cstring> // memcpy
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
// Scanning primitives of the view's parser. The unrolled checks at fixed
// offsets follow yyjson (https://github.com/ibireme/yyjson, MIT license): the
// loads do not depend on each other, so the CPU can run ahead. Words are read
// with read_eight_bytes(), so nothing here depends on the byte order.
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
/// 1 for bytes that may appear verbatim in a string: 0x20..0x7F except '"' and '\\'
inline const std::uint8_t* string_plain() noexcept
{
static const std::array<std::uint8_t, 256> table =
{
{
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, // 0x00..0x1F
1, 1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, // 0x20..0x3F ('"')
1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1, 1, 1, // 0x40..0x5F ('\\')
1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, // 0x60..0x7F
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, // 0x80..0x9F
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, // 0xA0..0xBF
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, // 0xC0..0xDF
0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, // 0xE0..0xFF
}
};
return table.data();
}
NLOHMANN_VIEW_ALWAYS_INLINE bool is_digit(unsigned char c) noexcept
{
return static_cast<unsigned char>(c - '0') <= 9;
}
/// two bytes as they are in memory (only compared with byte-symmetric patterns)
NLOHMANN_VIEW_ALWAYS_INLINE std::uint16_t load16(const unsigned char* p) noexcept
{
std::uint16_t w = 0;
std::memcpy(&w, p, 2);
return w;
}
/// Advance over plain string bytes and well-formed UTF-8. Stops at a quote,
/// a backslash, a control character, ill-formed UTF-8, or the end. The first
/// 16 bytes are checked one by one, so that the position advances by
/// constants in predicted branches (most strings are short); longer runs
/// continue eight bytes at a time.
NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* scan_string_run(const unsigned char* p, const unsigned char* e) noexcept
{
const std::uint8_t* plain = string_plain();
for (;;)
{
if (e - p >= 16)
{
#define NLOHMANN_VIEW_STEP(i) if (NLOHMANN_VIEW_LIKELY(plain[p[i]] != 0)) {} else { p += (i); goto stop; }
NLOHMANN_VIEW_REPEAT16(NLOHMANN_VIEW_STEP)
#undef NLOHMANN_VIEW_STEP
p += 16;
while (e - p >= 8)
{
const std::uint64_t special = swar_string_special(read_eight_bytes(p));
if (special != 0)
{
p += count_trailing_zeros(special) / 8;
goto stop;
}
p += 8;
}
continue;
}
while (p != e && plain[*p] != 0)
{
++p;
}
if (p == e)
{
return p;
}
stop:
if (*p < 0x80)
{
return p; // quote, backslash, or control character
}
// non-ASCII: a run of well-formed sequences (the library's check, so
// that exactly what json::parse accepts is accepted)
do
{
const std::size_t n = validate_one_utf8(p, static_cast<std::size_t>(e - p));
if (n == 0)
{
return p;
}
p += n;
}
while (p != e && *p >= 0x80);
}
}
/// advance over ASCII digits
NLOHMANN_VIEW_ALWAYS_INLINE const unsigned char* skip_digits(const unsigned char* p, const unsigned char* e) noexcept
{
while (e - p >= 16)
{
#define NLOHMANN_VIEW_STEP(i) if (NLOHMANN_VIEW_LIKELY(is_digit(p[i]))) {} else { return p + (i); }
NLOHMANN_VIEW_REPEAT16(NLOHMANN_VIEW_STEP)
#undef NLOHMANN_VIEW_STEP
p += 16;
}
while (p != e && is_digit(*p))
{
++p;
}
return p;
}
/// powers of ten up to 10^19 as integers
inline std::uint64_t int_pow10(unsigned k) noexcept
{
static const std::array<std::uint64_t, 20> table =
{
{
1u, 10u, 100u, 1000u, 10000u, 100000u, 1000000u, 10000000u, 100000000u, 1000000000u,
10000000000u, 100000000000u, 1000000000000u, 10000000000000u, 100000000000000u, 1000000000000000u,
10000000000000000u, 100000000000000000u, 1000000000000000000u, 10000000000000000000u
}
};
return table[k];
}
/// value of 0 < k < 8 digits at p in one step if [p, p + 8) lies below
/// limit, else one digit at a time (whole blocks of eight digits are read by
/// parse_upto19() directly)
NLOHMANN_VIEW_ALWAYS_INLINE std::uint64_t parse_upto8(const unsigned char* p, unsigned k, const unsigned char* limit) noexcept
{
if (NLOHMANN_VIEW_LIKELY(limit - p >= 8))
{
// move the k digits to the top and pad the vacated low bytes with '0'
const unsigned shift = 8 * (8 - k);
return parse_eight_digits((read_eight_bytes(p) << shift) | (0x3030303030303030u >> (8 * k)));
}
std::uint64_t v = 0;
for (unsigned i = 0; i < k; ++i)
{
v = (v * 10) + static_cast<std::uint64_t>(p[i] - '0');
}
return v;
}
/// value of k <= 19 digits at p
NLOHMANN_VIEW_ALWAYS_INLINE std::uint64_t parse_upto19(const unsigned char* p, unsigned k, const unsigned char* limit) noexcept
{
std::uint64_t w = 0;
while (k >= 8)
{
// (eight digits of the token: they lie below limit)
w = (w * 100000000u) + parse_eight_digits(read_eight_bytes(p));
p += 8;
k -= 8;
}
if (k != 0)
{
w = (w * int_pow10(k)) + parse_upto8(p, k, limit);
}
return w;
}
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+111
View File
@@ -0,0 +1,111 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <algorithm> // min
#include <cstddef> // size_t
#include <cstring> // memcmp, strlen
#include <string> // basic_string
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/macro_scope.hpp>
#if NLOHMANN_VIEW_HAS_CPP_17
#include <string_view> // string_view
#endif
#ifndef JSON_NO_IO
#include <ostream> // ostream
#endif
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
namespace view
{
#if NLOHMANN_VIEW_HAS_CPP_17
using string_ref = std::string_view;
#else
/// minimal C++11 stand-in for std::string_view
class string_ref
{
public:
using size_type = std::size_t;
using const_iterator = const char*;
string_ref() noexcept = default;
string_ref(const char* s) : m_data(s), m_size(std::strlen(s)) {} // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
string_ref(const char* s, std::size_t n) noexcept : m_data(s), m_size(n) {}
template<typename Traits, typename Alloc>
string_ref(const std::basic_string<char, Traits, Alloc>& s) noexcept : m_data(s.data()), m_size(s.size()) {} // NOLINT(google-explicit-constructor,hicpp-explicit-conversions)
const char* data() const noexcept
{
return m_data;
}
std::size_t size() const noexcept
{
return m_size;
}
std::size_t length() const noexcept
{
return m_size;
}
bool empty() const noexcept
{
return m_size == 0;
}
const char* begin() const noexcept
{
return m_data;
}
const char* end() const noexcept
{
return m_data + m_size;
}
char operator[](std::size_t i) const noexcept
{
return m_data[i];
}
template<typename Traits, typename Alloc>
explicit operator std::basic_string<char, Traits, Alloc>() const
{
return std::basic_string<char, Traits, Alloc>(m_data, m_size);
}
friend bool operator==(string_ref a, string_ref b) noexcept
{
return a.m_size == b.m_size && (a.m_size == 0 || std::memcmp(a.m_data, b.m_data, a.m_size) == 0);
}
friend bool operator!=(string_ref a, string_ref b) noexcept
{
return !(a == b);
}
friend bool operator<(string_ref a, string_ref b) noexcept
{
const int c = std::memcmp(a.m_data, b.m_data, (std::min)(a.m_size, b.m_size));
return c != 0 ? c < 0 : a.m_size < b.m_size;
}
#ifndef JSON_NO_IO
friend std::ostream& operator<<(std::ostream& o, string_ref s)
{
return o.write(s.m_data, static_cast<std::streamsize>(s.m_size));
}
#endif
private:
const char* m_data = "";
std::size_t m_size = 0;
};
#endif
} // namespace view
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
+1
View File
@@ -45,6 +45,7 @@
#include <nlohmann/adl_serializer.hpp>
#include <nlohmann/byte_container_with_subtype.hpp>
#include <nlohmann/detail/abi_config.hpp>
#include <nlohmann/detail/conversions/from_json.hpp>
#include <nlohmann/detail/conversions/to_json.hpp>
#include <nlohmann/detail/exceptions.hpp>
+130 -1
View File
@@ -7236,6 +7236,43 @@ class byte_container_with_subtype : public BinaryType
NLOHMANN_JSON_NAMESPACE_END
// #include <nlohmann/detail/abi_config.hpp>
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
// #include <nlohmann/detail/abi_macros.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
/*!
@brief the configuration macros that change the library's behavior
json.hpp undefines these macros at its end (see macro_unscope.hpp), so code
that builds on the library after it (json_view.hpp) reads them here. Like the
macros, they are part of the ABI namespace, so they always match the
basic_json they are used with.
*/
struct abi_config
{
/// JSON_STRICT_NUL_HANDLING: a null byte is an error, not the end of input
static constexpr bool strict_nul_handling = JSON_STRICT_NUL_HANDLING != 0;
/// JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON
static constexpr bool legacy_discarded_value_comparison = JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON != 0;
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
// #include <nlohmann/detail/conversions/from_json.hpp>
// #include <nlohmann/detail/conversions/to_json.hpp>
@@ -10047,8 +10084,9 @@ NLOHMANN_JSON_NAMESPACE_END
#include <array> // array
#include <cstddef> // size_t
#include <cstdint> // uint64_t
#include <cstdint> // uint64_t, uint8_t
#include <cstring> // memcpy
// #include <nlohmann/detail/bit_ops.hpp>
@@ -10356,6 +10394,51 @@ inline std::size_t string_bulk_run(const unsigned char* data, std::size_t n) noe
return scalar_string_bulk_run(data, n);
}
// Decode the 4 hex digits at [data, data+4) - the digits following a `\u`
// escape - into a codepoint 0x0000..0xFFFF via one table lookup per byte
// (after yyjson's read_hex_u16), or return -1 if any of the 4 bytes is not a
// hex digit ('0'..'9', 'A'..'F', 'a'..'f'). The caller must already have
// checked that 4 bytes are available; used by lexer::get_codepoint()'s
// contiguous fast path. On -1 it falls back to the byte-at-a-time loop, which
// stops at the first invalid digit, so the reported error and position are
// unaffected by this fast path.
inline int hex_codepoint(const unsigned char* data) noexcept
{
static const std::array<std::uint8_t, 256> hex_digit_table = // NOLINT(cppcoreguidelines-avoid-non-const-global-variables)
{
{
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 00..0F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 10..1F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 20..2F
0x00, 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07, 0x08, 0x09, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 30..3F ('0'..'9')
0xFF, 0x0A, 0x0B, 0x0C, 0x0D, 0x0E, 0x0F, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 40..4F ('A'..'F')
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 50..5F
0xFF, 0x0A, 0x0B, 0x0C, 0x0D, 0x0E, 0x0F, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 60..6F ('a'..'f')
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 70..7F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 80..8F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // 90..9F
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // A0..AF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // B0..BF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // C0..CF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // D0..DF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, // E0..EF
0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF // F0..FF
}
};
const std::uint8_t d0 = hex_digit_table[data[0]];
const std::uint8_t d1 = hex_digit_table[data[1]];
const std::uint8_t d2 = hex_digit_table[data[2]];
const std::uint8_t d3 = hex_digit_table[data[3]];
// every valid digit is <= 0xF; the combined OR only exceeds it if at
// least one of the four bytes was not a hex digit (looked up as 0xFF)
if ((d0 | d1 | d2 | d3) > 0x0F)
{
return -1;
}
return (d0 << 12) | (d1 << 8) | (d2 << 4) | d3;
}
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -10560,6 +10643,44 @@ class lexer : public lexer_base<BasicJsonType>
// scan functions
/////////////////////
/// contiguous input: try to decode the 4 hex digits following `\u`
/// directly from the input buffer via hex_codepoint(), instead of 4 calls
/// to get(). On success, advances the adapter and the position counters
/// exactly as those 4 get() calls would (a hex digit is never '\n', so
/// only the flat counters move) and leaves @a current holding the last of
/// the 4 digits, just as the last such get() would; the codepoint is
/// written to @a out. Makes no state change and returns false - for a
/// pending unget, fewer than 4 remaining bytes, or any of the 4 bytes not
/// being a hex digit - so the caller falls back unchanged to the
/// per-character loop, which then reports the same diagnostic (stopping
/// at the first invalid digit) as before this optimization.
bool get_codepoint_bulk(std::true_type /*bulk*/, int& out)
{
if (next_unget || ia.bulk_remaining() < 4)
{
return false;
}
const char_type* const raw = ia.bulk_data();
const int codepoint = hex_codepoint(reinterpret_cast<const unsigned char*>(raw));
if (codepoint < 0)
{
return false;
}
ia.bulk_skip(4);
// a hex digit is never a newline, so only the flat counters advance
position.chars_read_total += 4;
position.chars_read_current_line += 4;
current = char_traits<char_type>::to_int_type(raw[3]);
out = codepoint;
return true;
}
/// streaming input: no bulk fast path
bool get_codepoint_bulk(std::false_type /*bulk*/, int& /*out*/) const noexcept
{
return false;
}
/*!
@brief get codepoint from 4 hex characters following `\u`
@@ -10579,6 +10700,14 @@ class lexer : public lexer_base<BasicJsonType>
{
// this function only makes sense after reading `\u`
JSON_ASSERT(current == 'u');
// contiguous input: decode all 4 hex digits directly from the buffer
int fast_codepoint = 0;
if (get_codepoint_bulk(std::integral_constant<bool, bulk_scan> {}, fast_codepoint))
{
return fast_codepoint;
}
int codepoint = 0;
const auto factors = { 12u, 8u, 4u, 0u };
+4
View File
@@ -279,6 +279,10 @@ if(json_32bit_test_only)
elseif(NOT json_32bit_test)
list(FILTER files EXCLUDE REGEX src/unit-32bit.cpp)
endif()
if(NOT JSON_MultipleHeaders)
# the internal headers of json_view are not part of a single header yet
list(FILTER files EXCLUDE REGEX src/unit-json_view_builder.cpp)
endif()
foreach(file ${files})
json_test_add_test_for(${file} MAIN test_main CXX_STANDARDS ${test_cxx_standards} ${test_force})
+179
View File
@@ -18,6 +18,7 @@ using nlohmann::json;
#include <cstdlib> // strtod
#include <cstring> // memcpy
#include <map> // map
#include <random> // mt19937
#include <sstream> // stringstream
#include <string> // string
#include <utility> // pair
@@ -664,6 +665,184 @@ TEST_CASE("lexer string fast path")
}
}
TEST_CASE("lexer escape fast path")
{
// json::accept() never throws, so this section stays covered without
// exceptions; it pins which of the cases below are valid/invalid and
// checks the contiguous and streaming paths agree on that classification.
SECTION("accept() parity")
{
const std::vector<std::pair<std::string, bool>> cases =
{
{"\\u0041", true}, {"\\u00e4", true}, {"\\u00E4", true},
{"\\uD83D\\uDE00", true},
{"\\u12", false}, {"\\u12G4", false}, {"\\uXYZW", false},
{"\\uD800", false}, {"\\uD800A", false}, {"\\uD800\\u0041", false},
{"\\uDC00", false}, {"\\u", false}
};
for (const auto& c : cases)
{
for (const std::size_t offset :
{
std::size_t{0}, std::size_t{9}
})
{
const std::string doc = "[\"" + std::string(offset, 'a') + c.first + "\"]";
CAPTURE(doc);
CHECK(json::accept(doc) == c.second);
std::stringstream ss(doc);
CHECK(json::accept(ss) == c.second);
}
}
}
#if !defined(JSON_NOEXCEPTION)
// the full outcome of parsing @a doc: the parsed value, or the exact
// error message, so a mismatch in either is caught
const auto outcome = [](const std::string & doc, bool streaming) -> std::string
{
try
{
if (streaming)
{
std::stringstream ss(doc);
const json j = json::parse(ss);
return j.dump();
}
const json j = json::parse(doc);
return j.dump();
}
catch (const json::exception& e)
{
return {e.what()};
}
};
SECTION("contiguous vs streaming parity")
{
const std::vector<std::string> escapes =
{
"\\u0041", // "A"
"\\u00e4", // "ä" (lowercase hex)
"\\u00E4", // "ä" (uppercase hex)
"\\uD83D\\uDE00", // valid surrogate pair (an emoji)
"\\u12", // truncated: only 2 hex digits before the closing quote
"\\u12G4", // invalid hex digit at the 3rd position
"\\uXYZW", // all 4 bytes invalid
"\\uD800", // lone high surrogate, string ends right after
"\\uD800A", // high surrogate not followed by another \u escape
"\\uD800\\u0041", // high surrogate followed by \u, but not a low surrogate
"\\uDC00", // lone low surrogate
"\\u", // '\u' with nothing after (closing quote right away)
};
// once at the start of the string and once past the first 8-byte SWAR
// word of the outer string_bulk_run, so the escape is reached both
// right after the opening quote and mid-run
for (const auto& escape : escapes)
{
for (const std::size_t offset :
{
std::size_t{0}, std::size_t{9}
})
{
const std::string doc = "[\"" + std::string(offset, 'a') + escape + "\"]";
CAPTURE(doc);
CHECK(outcome(doc, false) == outcome(doc, true));
}
// the escape is the last thing before end of input: no closing
// quote at all
const std::string truncated_doc = "[\"" + escape;
CAPTURE(truncated_doc);
CHECK(outcome(truncated_doc, false) == outcome(truncated_doc, true));
}
}
SECTION("truncated \\u escape at every distance from the end of input")
{
// ia.bulk_remaining() must correctly report fewer than 4 bytes for
// every possible count of trailing hex-looking bytes (0, 1, 2, or 3)
// before end of input, so the fast path declines and the byte path
// alone reports the "must be followed by 4 hex digits" error, at the
// same position, in every case
for (const std::string& tail :
{
std::string{}, std::string("1"), std::string("12"), std::string("123")
})
{
const std::string doc = "[\"\\u" + tail;
CAPTURE(doc);
CHECK(outcome(doc, false) == outcome(doc, true));
CHECK(outcome(doc, false).find("must be followed by 4 hex digits") != std::string::npos);
}
}
SECTION("invalid hex digit at every position of the 4")
{
// the fast path must decline for *any* invalid byte among the 4, not
// just the first, and the byte path must then stop at exactly that
// position - same as it always has
for (std::size_t bad_pos = 0; bad_pos < 4; ++bad_pos)
{
std::string digits = "1234";
digits[bad_pos] = 'g'; // not a hex digit
const std::string doc = "[\"\\u" + digits + "\"]";
CAPTURE(doc);
CHECK(outcome(doc, false) == outcome(doc, true));
CHECK(outcome(doc, false).find("must be followed by 4 hex digits") != std::string::npos);
}
}
SECTION("random escapes")
{
// A seeded PRNG builds the 4 bytes following `\u` from a mix of hex
// digits and non-hex bytes, at varying distances from the start of
// the string, to compare the two scanners on many more shapes than
// are practical to enumerate by hand.
std::mt19937 gen(7654321); // NOLINT(cert-msc32-c,cert-msc51-cpp)
const std::string hex_alphabet = "0123456789AaBbCcDdEeFf";
std::uniform_int_distribution<std::size_t> pick_hex(0, hex_alphabet.size() - 1);
std::uniform_int_distribution<int> pick_byte(1, 255); // never NUL
std::uniform_int_distribution<int> pick_is_hex(0, 4); // 4-in-5 chance of a hex digit
std::uniform_int_distribution<std::size_t> pick_offset(0, 12);
std::vector<std::string> mismatches;
for (int iter = 0; iter < 3000; ++iter)
{
std::string digits;
for (int i = 0; i < 4; ++i)
{
if (pick_is_hex(gen) != 0)
{
digits += hex_alphabet[pick_hex(gen)];
}
else
{
char c = static_cast<char>(pick_byte(gen));
if (c == '"' || c == '\\')
{
// keep the string well-formed apart from the escape
// itself, so any mismatch is attributable to the \u
// handling and not to an unrelated quote/escape
c = 'z';
}
digits += c;
}
}
const std::string doc = "[\"" + std::string(pick_offset(gen), 'a') + "\\u" + digits + "\"]";
if (outcome(doc, false) != outcome(doc, true))
{
mismatches.push_back(doc);
}
}
CAPTURE(mismatches);
CHECK(mismatches.empty());
}
#endif
}
namespace
{
// the index of the decimal point (or npos) and of the end of the mantissa of a
+399
View File
@@ -0,0 +1,399 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
#include <nlohmann/detail/view/builder.hpp>
#include <nlohmann/detail/view/string_ref.hpp>
using nlohmann::json;
#include <cstdint>
#include <fstream>
#include <map>
#include <memory>
#include <random>
#include <sstream>
#include <string>
#include <utility>
#include <vector>
#include <test_data.hpp>
namespace
{
using nlohmann::detail::view::document_data;
using nlohmann::detail::view::node;
// the node index of a text; a vector input has no terminating NUL, so that
// AddressSanitizer catches any read past the last byte
struct built
{
std::unique_ptr<document_data, document_data::deleter> data{}; // NOLINT(readability-redundant-member-init)
std::vector<char> copy{}; // NOLINT(readability-redundant-member-init)
bool ok = false;
nlohmann::detail::view::parse_failure failure{};
};
template<typename FloatType = double>
built build(const std::string& text, bool comments, bool trailing_commas, bool sentinel)
{
built r;
r.data.reset(document_data::create(nlohmann::detail::view::estimate_nodes(text.data(), text.size())));
const char* src = text.c_str();
if (!sentinel)
{
r.copy.assign(text.begin(), text.end());
src = r.copy.data();
}
r.ok = nlohmann::detail::view::build < FloatType, !nlohmann::detail::abi_config::strict_nul_handling > (*r.data, src, text.size(), comments, trailing_commas, sentinel, r.failure);
r.data->src = src;
r.data->base[0] = src;
r.data->base[1] = r.data->arena.data();
return r;
}
// the value of a subtree, as json::parse would build it
json value_of(const document_data& d, const node*& n)
{
const node& x = *n;
++n;
switch (static_cast<json::value_t>(x.kind))
{
case json::value_t::object:
{
json o = json::object();
const node* const end = &x + x.next;
while (n != end)
{
const std::string key(d.str(*n), n->len);
++n;
o[key] = value_of(d, n);
}
return o;
}
case json::value_t::array:
{
json a = json::array();
const node* const end = &x + x.next;
while (n != end)
{
a.push_back(value_of(d, n));
}
return a;
}
case json::value_t::string:
return std::string(d.str(x), x.len);
case json::value_t::boolean:
return (x.flags & nlohmann::detail::view::node_flags::is_true) != 0;
case json::value_t::number_integer:
return static_cast<std::int64_t>(nlohmann::detail::view::integer_bits(x));
case json::value_t::number_unsigned:
return nlohmann::detail::view::integer_bits(x);
case json::value_t::number_float:
return json::parse(std::string(d.src + x.off, x.len)).get<double>();
case json::value_t::null:
case json::value_t::binary:
case json::value_t::discarded:
default:
return nullptr;
}
}
json value_of(const built& b)
{
const node* n = b.data->tape;
json v = value_of(*b.data, n);
CHECK(n == b.data->tape + b.data->tape_size);
return v;
}
// accept/reject and the value must match json::parse, for all options and
// with and without a NUL after the text
void check_same(const std::string& text)
{
CAPTURE(text);
for (int options = 0; options < 4; ++options)
{
const bool comments = (options & 1) != 0;
const bool trailing_commas = (options & 2) != 0;
const bool accepted = json::accept(text, comments, trailing_commas);
for (const bool sentinel :
{
true, false
})
{
const built b = build(text, comments, trailing_commas, sentinel);
CHECK(b.ok == accepted);
if (b.ok && accepted)
{
CHECK(value_of(b) == json::parse(text, nullptr, true, comments, trailing_commas));
}
}
}
}
// a small deterministic generator of documents
struct generator
{
std::mt19937 rng{5295}; // NOLINT(cert-msc32-c,cert-msc51-cpp,bugprone-random-generator-seed)
int r(int n)
{
return static_cast<int>(rng() % static_cast<unsigned>(n));
}
void ws(std::string& o)
{
for (int n = r(4) == 0 ? r(12) : r(2); n > 0; --n)
{
o += " \n\t\r "[r(6)];
}
}
void str(std::string& o)
{
static const char* const pieces[] = {"a", "Z", " ", "~", "\\n", "\\\"", "\\\\", "\\/", "\\u00e9", "\\ud83d\\ude00", "\xc3\xa9", "\xe3\x81\x82", "\xf0\x9f\x98\x80", "\x7f", "\\u001f", "long enough text to leave the first 16 bytes"}; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
o += '"';
for (int n = r(3) == 0 ? r(20) : r(6); n > 0; --n)
{
o += pieces[r(16)];
}
o += '"';
}
void num(std::string& o)
{
static const char* const numbers[] = {"0", "-0", "1", "-1", "12", "123456789", "1234567890123456789", "9223372036854775807", "-9223372036854775808", // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
"9223372036854775808", "18446744073709551615", "18446744073709551616", "-9223372036854775809",
"1.5", "-2.25e-3", "1e10", "1E+2", "0.000001", "3.141592653589793238462643", "1e308", "-1e-400", "123.456e7"
};
o += numbers[r(22)];
}
void value(std::string& o, int depth)
{
ws(o);
const int k = depth > 5 ? 2 + r(6) : r(8);
if (k == 0 || k == 1)
{
const bool object = k == 0;
o += object ? '{' : '[';
for (int i = r(5); i > 0; --i)
{
ws(o);
if (object)
{
str(o);
ws(o);
o += ':';
}
value(o, depth + 1);
o += i > 1 ? "," : "";
}
ws(o);
o += object ? '}' : ']';
}
else if (k < 4)
{
str(o);
}
else if (k < 6)
{
num(o);
}
else
{
static const char* const literals[] = {"true", "false", "null"}; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
o += literals[r(3)];
}
ws(o);
}
};
} // namespace
TEST_CASE("json_view string_ref")
{
// std::string_view in C++17, a stand-in with the same members before
using nlohmann::detail::view::string_ref;
const std::string text = "abc";
const string_ref r(text);
CHECK(r.length() == 3);
CHECK(std::string(r.begin(), r.end()) == "abc");
CHECK(r[1] == 'b');
CHECK(r != string_ref("abd"));
CHECK_FALSE(r != string_ref("abcd", 3));
std::ostringstream o;
o << r;
CHECK(o.str() == "abc");
}
TEST_CASE("json_view builder")
{
SECTION("scalars and containers")
{
for (const char* text :
{
"null", "true", "false", "0", "-0", "42", "-42", "1.5", "\"\"", "\"abc\"", "[]", "{}", "[1,2,3]", "{\"a\":1,\"b\":[true,null]}", // NOLINT(modernize-raw-string-literal)
" [ 1 , 2 ] ", "{\"a\" : {\"b\" : {}}}", "[[[]]]", "\"\\u00e4\\n\\ud83d\\ude00\"", "{\"a\":1,\"a\":2}", "18446744073709551616", // NOLINT(modernize-raw-string-literal)
"-9223372036854775809", "123456789012345678901234567890", "1e400", "-1e400", "1.7976931348623157e308"
})
{
check_same(text);
}
// the midpoint between the largest double and 2^1024 rounds to
// infinity (an overflow), one less to the largest double: with more
// than 19 digits, Eisel-Lemire cannot decide these, and the overflow
// check needs the exact comparison with the midpoint
const std::string midpoint = "179769313486231580793728971405303415079934132710037826936173778980444968292764750946649017977587207096330286416692887910946555547851940402630657488671505820681908902000708383676273854845817711531764475730270069855571366959622842914819860834936475292719074168444365510704342711559699508093042880177904174497792";
const std::string below = "179769313486231580793728971405303415079934132710037826936173778980444968292764750946649017977587207096330286416692887910946555547851940402630657488671505820681908902000708383676273854845817711531764475730270069855571366959622842914819860834936475292719074168444365510704342711559699508093042880177904174497791";
check_same(midpoint);
check_same("-" + midpoint);
check_same(below);
check_same("[" + below + "," + midpoint + "]");
// the check uses the floating-point type of the document: with float,
// the view rejects what parse() rejects (out_of_range.406), and a
// double document is not affected
using float_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, float>;
CHECK_FALSE(float_json::accept("1e39"));
CHECK(float_json::accept("3.4028235e38"));
CHECK_FALSE(float_json::accept("3.4028236e38"));
for (const char* text :
{
"1e39", "-1e39", "3.4028235e38", "-3.4028235e38", "3.4028236e38", "-3.4028236e38", "3.4028234663852886e38", "1e38",
"340282356779733661637539395458142568448", "340282356779733661637539395458142568447.99", "0.00034028236e42",
"[1.5e38, 3.5e38]", "{\"a\": 1e-50, \"b\": 1e39}"
})
{
CAPTURE(text);
const bool float_accepted = float_json::accept(text);
for (const bool sentinel :
{
true, false
})
{
const built f = build<float>(text, false, false, sentinel);
CHECK(f.ok == float_accepted);
if (!f.ok)
{
CHECK(f.failure.code == nlohmann::detail::view::error_code::number_overflow);
}
CHECK(build<double>(text, false, false, sentinel).ok == json::accept(text));
}
}
}
SECTION("malformed input")
{
for (const char* text :
{
"", " ", "[", "]", "{", "}", "[1,]", "{\"a\":1,}", "[1 2]", "{\"a\" 1}", "{1:2}", "tru", "nul", "fals", "truex", "-", "01", "1.", ".5", "1e", "1e+",
"\"", "\"abc", "\"\\x\"", "\"\\u12\"", "\"\\u12G4\"", "\"\\ud800\"", "\"\\udc00\"", "\"\\ud800\\u0041\"", "\"\x01\"", "\"\xff\"", "\"\xc3\"", // NOLINT(modernize-raw-string-literal)
"\"\xe0\x80\x80\"", "\"\xed\xa0\x80\"", "[1]x", "[1] [2]", "/", "/*", "/* */ 1", "// c\n1", "1 // c", "[1,/*c*/2]", "[1,2,]"
})
{
check_same(text);
}
}
SECTION("NUL, BOM, and whitespace")
{
// a NUL inside a string is a control character, as for json::parse
// (where a NUL ends the input, it does so only between values)
for (const bool sentinel :
{
true, false
})
{
const built b = build(std::string("[\"ab\0cd\"]", 9), false, false, sentinel);
CHECK(!b.ok);
CHECK(b.failure.code == nlohmann::detail::view::error_code::string_control_character);
CHECK(b.failure.offset == 4);
}
check_same(std::string("[1]\0garbage", 11));
check_same(std::string("[1\0]", 4));
check_same(std::string("[1, // c\0\n2]", 12));
check_same(std::string("[1, /* c\0 */ 2]", 15));
check_same("\xEF\xBB\xBF[1]");
check_same("\xEF\xBB[1]");
check_same(" \t\r\n 7 \n");
for (const char* text :
{"[1]\r", "[1]\n", "[1]\r\n", "[1,\r2]", "[1,\r\n2]", "7\r", "\"x\"\r", "{\"a\":\r\n1}\r", "[\n 1,\n 2\n]", "{\n \"a\": [\n 1\n ]\n}"
})
{
check_same(text);
}
}
SECTION("deep nesting")
{
// the open containers beyond 64 levels live on the heap
for (const std::size_t depth :
{
63u, 64u, 65u, 1000u, 100000u
})
{
const std::string arrays = std::string(depth, '[') + std::string(depth, ']');
const built b = build(arrays, false, false, false);
REQUIRE(b.ok);
CHECK(b.data->tape_size == depth);
CHECK(b.data->tape[0].next == depth);
std::string objects;
for (std::size_t i = 0; i < depth; ++i)
{
objects += "{\"a\":";
}
objects += '1' + std::string(depth, '}');
const built o = build(objects, false, false, false);
REQUIRE(o.ok);
CHECK(o.data->tape_size == (2 * depth) + 1);
CHECK(!build(std::string(depth, '[') + std::string(depth - 1, ']'), false, false, false).ok);
}
}
SECTION("generated documents and damaged copies")
{
generator g;
for (int i = 0; i < 3000; ++i)
{
std::string text;
g.value(text, 0);
check_same(text);
// damage: flip one byte, or cut the text
std::string damaged = text;
const auto at = static_cast<std::size_t>(g.r(static_cast<int>(damaged.size())));
static const char replacements[] = {'x', '"', '\\', ',', ':', ']', '}', '[', '{', '1', '-', '.', 'e', '\0', '\n', '/'}; // NOLINT(cppcoreguidelines-avoid-c-arrays,hicpp-avoid-c-arrays,modernize-avoid-c-arrays)
damaged[at] = replacements[g.r(16)];
check_same(damaged);
check_same(text.substr(0, at));
}
}
SECTION("test files")
{
for (const char* name :
{
"/json.org/1.json", "/json.org/2.json", "/json.org/3.json", "/json.org/4.json", "/json.org/5.json",
"/json_testsuite/sample.json", "/nativejson-benchmark/canada.json", "/nativejson-benchmark/citm_catalog.json",
"/nativejson-benchmark/twitter.json", "/json_tests/pass1.json", "/json_tests/pass2.json", "/json_tests/pass3.json"
})
{
CAPTURE(name);
std::ifstream f(std::string(TEST_DATA_DIRECTORY) + name, std::ios::binary);
std::stringstream ss;
ss << f.rdbuf();
const std::string text = ss.str();
REQUIRE(!text.empty());
const built b = build(text, false, false, true);
REQUIRE(b.ok);
CHECK(value_of(b) == json::parse(text));
}
}
}