Files
json/include/nlohmann/detail/input/parser.hpp
T
Niels Lohmann cff0a61369 Share DOM SAX position handling; fix stale parser and lexer comments (#5731)
* Share the diagnostic-position setter of the DOM SAX parsers

json_sax_dom_parser and json_sax_dom_callback_parser each had a private
copy of handle_diagnostic_positions_for_json_value(), identical except
for comments. Move the body into one static member function,
detail::diagnostic_positions::set_from_lexer(value, lexer), which both
classes call with their lexer pointer. basic_json befriends the new
struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position
members are private.

The discarded case is reached through the callback parser, so the
LCOV_EXCL markers that only the dom parser's copy had are gone. The
NOLINT on the unreachable default case loses the stray
"-warnings-as-errors", which is not a check name.

The start-position setup in start_object()/start_array() is left alone,
as #5706 is editing the callback parser's versions.

Behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Correct the parser comments on recursion and skip_to_state_evaluation

The class documentation called the parser a recursive descent parser,
but sax_parse_internal() is a loop that keeps the open containers on an
explicit stack. The comment at the end of an array and of an object
said the flag is set to false while the code below it sets it to true.
Describe what the code does instead.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Update the discard_number_values comments to the current number path

The comments explaining the accept() shortcut in convert_number() and
the member documentation still argued in terms of strtoull()/strtoll()
and errno, which #5283 replaced with convert_integer(), and pointed at
scan_number() instead of convert_number(). They also did not say that
scan_number_bulk_contiguous() converts integers itself, so the shortcut
is only reached for input without bulk access, with
JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back.

Rewrite both comments to describe the digit-count check in front of
convert_integer(), keeping the 18-digit bound and the json_sax_acceptor
argument. The stale <cstdlib> comment is left for after #5616, which
edits that include block.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* List the UTF-8 validators instead of calling the DFA the only one

The documentation of decode() called the Hoehrmann DFA the single
source of truth for UTF-8 validation. It is used only by the serializer
and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text
strings). The lexer's scan_string() switch, validate_one_utf8() /
valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer)
and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges
on their own.

Replace the sentence with a list of the four validators, what each is
used for, and a note that they must accept the same sequences. Sharing
code between them was considered and dropped: it would save a few lines
in a validator that is entangled with BON8 pushback, and #5677 is
editing the BON8 byte path.

Comments only; behavior, the public API and the ABI are unchanged.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Fix stale doc comments and include lists in the input headers

input_adapters.hpp included <memory> and <numeric> for the removed
shared_ptr-based adapter design but used neither; it called
(std::min) without including <algorithm>. json_sax.hpp used
std::numeric_limits without including <limits>. Also corrected
comments that no longer matched the code: input_stream_adapter does
not skip the input's BOM (the lexer's skip_bom() does), the
span_input_adapter comment named the no-longer-existing
input_buffer_adapter type, lexer::get_string() does not reset the
token, binary_reader's get_number() doc opened with /* instead of
/*! (so Doxygen skipped it) and omitted BON8 from its endianness
note, and the UBJSON-binary-types note did not mention that BJData
'B' arrays are read as binary.

Left out: the lgtm suppression on lexer.hpp's scan_number() (in
#5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's
start_object/start_array docs (in draft #5267's hunk), per the
verdict's conflict list.

Part of #5712

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

* Deduplicate the strict-EOF/release_lookahead/error block in parser::parse()

json_sax_dom_callback_parser and json_sax_dom_parser branches of
parser::parse() ran the same ~25 lines after sax_parse_internal():
the strict-mode EOF check (raising parse_error.101 through the SAX
parser), release_lookahead() in non-strict mode, and mapping an
errored SAX parser to a discarded result. The two copies had already
drifted apart in formatting and in the second copy's "see above"
comment.

Add a private parse_dom(DomSax&, strict) member that runs this shared
sequence once and returns whether the SAX parser did not error; both
branches of parse() now only construct their DOM SAX parser, call
parse_dom(), and (for the callback parser) map a discarded top-level
value to null. sax_parse() is left untouched, since it only runs the
EOF check and release_lookahead() when sax_parse_internal() succeeded,
unlike parse(), which runs them unconditionally.

Behavior-preserving: same operations in the same order for both SAX
parser kinds. Verified with unit-class_parser (strict/non-strict,
callback and non-callback), unit-deserialization and
unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean
make amalgamate / make check-amalgamation diff.

Overlaps #5601, which touches the same lines.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5712 item 2

* Share the code point to UTF-8 encoding between the wide-string helpers and the lexer

The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three
times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32
code point, in the UTF-16 helper for both a BMP code unit and a valid
surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The
copies had drifted: the UTF-32 helper masked the leading bits of each
byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift
alone, even though both give the same result for a code point that is
already known to be in range.

Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder
that invokes a callable once per output byte, most significant byte
first. Use it in the three valid-code-point branches (UTF-32 code
points up to U+10FFFF, UTF-16 code units outside the surrogate range,
and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where
out forwards to add(). The UTF-16 helper's deliberate pass-through of
malformed surrogate units and the UTF-32 helper's 0xFF sentinel for
code points above U+10FFFF are untouched, since neither reaches the new
helper.

Behavior-preserving: same bytes in the same order for every valid code
point, verified with unit-class_lexer, unit-class_parser,
unit-deserialization, unit-wstring and the non-test-data parts of
unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse
microbenchmark that shows no change (about 73 ms either way, median of
3, 1M escape sequences). single_include/ regenerated with make
amalgamate; make check-amalgamation leaves a clean tree.

Overlaps #5704, which rewrites the wide_string_input_helper
specializations touched here.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>

#5712 item 6

---------

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-10-01 07:32:32 +02:00

574 lines
21 KiB
C++

// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <cmath> // isfinite
#include <cstdint> // uint8_t
#include <functional> // function
#include <string> // string
#include <utility> // move
#include <vector> // vector
#include <nlohmann/detail/exceptions.hpp>
#include <nlohmann/detail/input/input_adapters.hpp>
#include <nlohmann/detail/input/json_sax.hpp>
#include <nlohmann/detail/input/lexer.hpp>
#include <nlohmann/detail/macro_scope.hpp>
#include <nlohmann/detail/meta/is_sax.hpp>
#include <nlohmann/detail/string_concat.hpp>
#include <nlohmann/detail/value_t.hpp>
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
////////////
// parser //
////////////
enum class parse_event_t : std::uint8_t
{
/// the parser read `{` and started to process a JSON object
object_start,
/// the parser read `}` and finished processing a JSON object
object_end,
/// the parser read `[` and started to process a JSON array
array_start,
/// the parser read `]` and finished processing a JSON array
array_end,
/// the parser read a key of a value in an object
key,
/// the parser finished reading a JSON value
value
};
template<typename BasicJsonType>
using parser_callback_t =
std::function<bool(int /*depth*/, parse_event_t /*event*/, BasicJsonType& /*parsed*/)>;
/*!
@brief syntax analysis
This class implements an iterative parser that keeps the open containers on
an explicit stack and reports what it reads as SAX events.
*/
template<typename BasicJsonType, typename InputAdapterType>
class parser
{
using number_integer_t = typename BasicJsonType::number_integer_t;
using number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using number_float_t = typename BasicJsonType::number_float_t;
using string_t = typename BasicJsonType::string_t;
using lexer_t = lexer<BasicJsonType, InputAdapterType>;
using token_type = typename lexer_t::token_type;
public:
/// a parser reading from an input adapter
explicit parser(InputAdapterType&& adapter,
parser_callback_t<BasicJsonType> cb = nullptr,
const bool allow_exceptions_ = true,
const bool ignore_comments = false,
const bool ignore_trailing_commas_ = false,
const bool discard_number_values_ = false)
: callback(std::move(cb))
, m_lexer(std::move(adapter), ignore_comments, discard_number_values_)
, allow_exceptions(allow_exceptions_)
, ignore_trailing_commas(ignore_trailing_commas_)
{
// read first token
get_token();
}
/*!
@brief public parser interface
@param[in] strict whether to expect the last token to be EOF
@param[in,out] result parsed JSON value
@throw parse_error.101 in case of an unexpected token (including invalid
unicode escapes and surrogate errors, which are reported with a
detailed message)
*/
void parse(const bool strict, BasicJsonType& result)
{
if (callback)
{
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
// in case of an error, return a discarded value
if (!parse_dom(sdp, strict))
{
result = value_t::discarded;
return;
}
// set top-level value to null if it was discarded by the callback
// function
if (result.is_discarded())
{
result = nullptr;
}
}
else
{
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
// in case of an error, return a discarded value
if (!parse_dom(sdp, strict))
{
result = value_t::discarded;
return;
}
}
result.assert_invariant();
}
/*!
@brief public accept interface
@param[in] strict whether to expect the last token to be EOF
@return whether the input is a proper JSON text
*/
bool accept(const bool strict = true)
{
json_sax_acceptor<BasicJsonType> sax_acceptor;
return sax_parse(&sax_acceptor, strict);
}
template<typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse(SAX* sax, const bool strict = true)
{
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
const bool result = sax_parse_internal(sax);
if (result)
{
if (strict)
{
// strict mode: next byte must be EOF
if (get_token() != token_type::end_of_input)
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr));
}
}
else
{
// the caller keeps using the input: position it right after
// the value by leaving the character that terminated it
m_lexer.release_lookahead();
}
}
return result;
}
private:
/*!
@brief run a DOM SAX parser to completion and position the lexer
Shared by both branches of @ref parse(): builds no SAX parser itself,
but drives an already-constructed @a json_sax_dom_parser or
@ref json_sax_dom_callback_parser through @ref sax_parse_internal(),
then applies the strict-EOF check (reporting parse_error.101 through
@a sdp on failure) or, in non-strict mode, releases the lookahead so
the caller can keep reading the input right after the parsed value.
@param[in,out] sdp the DOM SAX parser to run
@param[in] strict whether to expect the last token to be EOF
@return whether @a sdp did not report an error
*/
template<typename DomSax>
bool parse_dom(DomSax& sdp, const bool strict)
{
sax_parse_internal(&sdp);
if (strict)
{
// in strict mode, input must be completely read
if (get_token() != token_type::end_of_input)
{
sdp.parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(),
exception_message(token_type::end_of_input, "value"), nullptr));
}
}
else
{
// the caller keeps using the input: position it right after
// the value by leaving the character that terminated it
m_lexer.release_lookahead();
}
return !sdp.is_errored();
}
template<typename SAX>
JSON_HEDLEY_NON_NULL(2)
bool sax_parse_internal(SAX* sax)
{
// stack to remember the hierarchy of structured values we are parsing
// true = array; false = object
std::vector<bool> states;
// value to avoid a goto (see comment where set to true)
bool skip_to_state_evaluation = false;
while (true)
{
if (!skip_to_state_evaluation)
{
// invariant: get_token() was called before each iteration
switch (last_token)
{
case token_type::begin_object:
{
if (JSON_HEDLEY_UNLIKELY(!sax->start_object(detail::unknown_size())))
{
return false;
}
// closing } -> we are done
if (get_token() == token_type::end_object)
{
if (JSON_HEDLEY_UNLIKELY(!sax->end_object()))
{
return false;
}
break;
}
// parse key
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
return false;
}
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr));
}
// remember we are now inside an object
states.push_back(false);
// parse values
get_token();
continue;
}
case token_type::begin_array:
{
if (JSON_HEDLEY_UNLIKELY(!sax->start_array(detail::unknown_size())))
{
return false;
}
// closing ] -> we are done
if (get_token() == token_type::end_array)
{
if (JSON_HEDLEY_UNLIKELY(!sax->end_array()))
{
return false;
}
break;
}
// remember we are now inside an array
states.push_back(true);
// parse values (no need to call get_token)
continue;
}
case token_type::value_float:
{
const auto res = m_lexer.get_number_float();
if (JSON_HEDLEY_UNLIKELY(!std::isfinite(res)))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
out_of_range::create(406, concat("number overflow parsing '", m_lexer.get_token_string(), '\''), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->number_float(res, m_lexer.get_string())))
{
return false;
}
break;
}
case token_type::literal_false:
{
if (JSON_HEDLEY_UNLIKELY(!sax->boolean(false)))
{
return false;
}
break;
}
case token_type::literal_null:
{
if (JSON_HEDLEY_UNLIKELY(!sax->null()))
{
return false;
}
break;
}
case token_type::literal_true:
{
if (JSON_HEDLEY_UNLIKELY(!sax->boolean(true)))
{
return false;
}
break;
}
case token_type::value_integer:
{
if (JSON_HEDLEY_UNLIKELY(!sax->number_integer(m_lexer.get_number_integer())))
{
return false;
}
break;
}
case token_type::value_string:
{
if (JSON_HEDLEY_UNLIKELY(!sax->string(m_lexer.get_string())))
{
return false;
}
break;
}
case token_type::value_unsigned:
{
if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(m_lexer.get_number_unsigned())))
{
return false;
}
break;
}
case token_type::parse_error:
{
// using "uninitialized" to avoid an "expected" message
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::uninitialized, "value"), nullptr));
}
case token_type::end_of_input:
{
if (JSON_HEDLEY_UNLIKELY(m_lexer.get_position().chars_read_total == 1))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(),
"attempting to parse an empty input; check that your input string or stream contains the expected JSON", nullptr));
}
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr));
}
case token_type::uninitialized:
case token_type::end_array:
case token_type::end_object:
case token_type::name_separator:
case token_type::value_separator:
case token_type::literal_or_value:
default: // the last token was unexpected
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::literal_or_value, "value"), nullptr));
}
}
}
else
{
skip_to_state_evaluation = false;
}
// we reached this line after we successfully parsed a value
if (states.empty())
{
// empty stack: we reached the end of the hierarchy: done
return true;
}
if (states.back()) // array
{
// comma -> next value
// or end of array (ignore_trailing_commas = true)
if (get_token() == token_type::value_separator)
{
// parse a new value
get_token();
// if ignore_trailing_commas and last_token is ], we can continue to "closing ]"
if (!(ignore_trailing_commas && last_token == token_type::end_array))
{
continue;
}
}
// closing ]
if (JSON_HEDLEY_LIKELY(last_token == token_type::end_array))
{
if (JSON_HEDLEY_UNLIKELY(!sax->end_array()))
{
return false;
}
// We are done with this array. Before we can parse a
// new value, we need to evaluate the new state first.
// By setting skip_to_state_evaluation to true, the next
// iteration skips parsing a value and evaluates the
// enclosing state directly.
JSON_ASSERT(!states.empty());
states.pop_back();
skip_to_state_evaluation = true;
continue;
}
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_array, "array"), nullptr));
}
// states.back() is false -> object
// comma -> next value
// or end of object (ignore_trailing_commas = true)
if (get_token() == token_type::value_separator)
{
get_token();
// if ignore_trailing_commas and last_token is }, we can continue to "closing }"
if (!(ignore_trailing_commas && last_token == token_type::end_object))
{
// parse key
if (JSON_HEDLEY_UNLIKELY(last_token != token_type::value_string))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::value_string, "object key"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!sax->key(m_lexer.get_string())))
{
return false;
}
// parse separator (:)
if (JSON_HEDLEY_UNLIKELY(get_token() != token_type::name_separator))
{
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::name_separator, "object separator"), nullptr));
}
// parse values
get_token();
continue;
}
}
// closing }
if (JSON_HEDLEY_LIKELY(last_token == token_type::end_object))
{
if (JSON_HEDLEY_UNLIKELY(!sax->end_object()))
{
return false;
}
// We are done with this object. Before we can parse a
// new value, we need to evaluate the new state first.
// By setting skip_to_state_evaluation to true, the next
// iteration skips parsing a value and evaluates the
// enclosing state directly.
JSON_ASSERT(!states.empty());
states.pop_back();
skip_to_state_evaluation = true;
continue;
}
return sax->parse_error(m_lexer.get_position(),
m_lexer.get_token_string(),
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_object, "object"), nullptr));
}
}
/// get next token from lexer
token_type get_token()
{
return last_token = m_lexer.scan();
}
std::string exception_message(const token_type expected, const std::string& context)
{
std::string error_msg = "syntax error ";
if (!context.empty())
{
error_msg += concat("while parsing ", context, ' ');
}
error_msg += "- ";
if (last_token == token_type::parse_error)
{
error_msg += concat(m_lexer.get_error_message(), "; last read: '",
m_lexer.get_token_string(), '\'');
}
else
{
error_msg += concat("unexpected ", lexer_t::token_type_name(last_token));
}
if (expected != token_type::uninitialized)
{
error_msg += concat("; expected ", lexer_t::token_type_name(expected));
}
return error_msg;
}
private:
/// callback function
const parser_callback_t<BasicJsonType> callback = nullptr;
/// the type of the last read token
token_type last_token = token_type::uninitialized;
/// the lexer
lexer_t m_lexer;
/// whether to throw exceptions in case of errors
const bool allow_exceptions = true;
/// whether trailing commas in objects and arrays should be ignored (true) or signaled as errors (false)
const bool ignore_trailing_commas = false;
};
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END