mirror of
https://github.com/nlohmann/json.git
synced 2026-10-03 21:20:30 +00:00
Share DOM SAX position handling; fix stale parser and lexer comments (#5731)
* Share the diagnostic-position setter of the DOM SAX parsers json_sax_dom_parser and json_sax_dom_callback_parser each had a private copy of handle_diagnostic_positions_for_json_value(), identical except for comments. Move the body into one static member function, detail::diagnostic_positions::set_from_lexer(value, lexer), which both classes call with their lexer pointer. basic_json befriends the new struct (only when JSON_DIAGNOSTIC_POSITIONS is enabled), as the position members are private. The discarded case is reached through the callback parser, so the LCOV_EXCL markers that only the dom parser's copy had are gone. The NOLINT on the unreachable default case loses the stray "-warnings-as-errors", which is not a check name. The start-position setup in start_object()/start_array() is left alone, as #5706 is editing the callback parser's versions. Behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Correct the parser comments on recursion and skip_to_state_evaluation The class documentation called the parser a recursive descent parser, but sax_parse_internal() is a loop that keeps the open containers on an explicit stack. The comment at the end of an array and of an object said the flag is set to false while the code below it sets it to true. Describe what the code does instead. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Update the discard_number_values comments to the current number path The comments explaining the accept() shortcut in convert_number() and the member documentation still argued in terms of strtoull()/strtoll() and errno, which #5283 replaced with convert_integer(), and pointed at scan_number() instead of convert_number(). They also did not say that scan_number_bulk_contiguous() converts integers itself, so the shortcut is only reached for input without bulk access, with JSON_DIAGNOSTIC_POSITIONS, or when the bulk scanner falls back. Rewrite both comments to describe the digit-count check in front of convert_integer(), keeping the 18-digit bound and the json_sax_acceptor argument. The stale <cstdlib> comment is left for after #5616, which edits that include block. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * List the UTF-8 validators instead of calling the DFA the only one The documentation of decode() called the Hoehrmann DFA the single source of truth for UTF-8 validation. It is used only by the serializer and by is_valid_utf8() (CBOR/MessagePack/BSON/UBJSON/BJData text strings). The lexer's scan_string() switch, validate_one_utf8() / valid_utf8_prefix() (bulk string scan, BON8 bulk path and BON8 writer) and the BON8 byte path in get_bon8_string() check the RFC 3629 ranges on their own. Replace the sentence with a list of the four validators, what each is used for, and a note that they must accept the same sequences. Sharing code between them was considered and dropped: it would save a few lines in a validator that is entangled with BON8 pushback, and #5677 is editing the BON8 byte path. Comments only; behavior, the public API and the ABI are unchanged. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Fix stale doc comments and include lists in the input headers input_adapters.hpp included <memory> and <numeric> for the removed shared_ptr-based adapter design but used neither; it called (std::min) without including <algorithm>. json_sax.hpp used std::numeric_limits without including <limits>. Also corrected comments that no longer matched the code: input_stream_adapter does not skip the input's BOM (the lexer's skip_bom() does), the span_input_adapter comment named the no-longer-existing input_buffer_adapter type, lexer::get_string() does not reset the token, binary_reader's get_number() doc opened with /* instead of /*! (so Doxygen skipped it) and omitted BON8 from its endianness note, and the UBJSON-binary-types note did not mention that BJData 'B' arrays are read as binary. Left out: the lgtm suppression on lexer.hpp's scan_number() (in #5616's hunk) and the "-1 if unknown" wording in json_sax.hpp's start_object/start_array docs (in draft #5267's hunk), per the verdict's conflict list. Part of #5712 Signed-off-by: Niels Lohmann <mail@nlohmann.me> * Deduplicate the strict-EOF/release_lookahead/error block in parser::parse() json_sax_dom_callback_parser and json_sax_dom_parser branches of parser::parse() ran the same ~25 lines after sax_parse_internal(): the strict-mode EOF check (raising parse_error.101 through the SAX parser), release_lookahead() in non-strict mode, and mapping an errored SAX parser to a discarded result. The two copies had already drifted apart in formatting and in the second copy's "see above" comment. Add a private parse_dom(DomSax&, strict) member that runs this shared sequence once and returns whether the SAX parser did not error; both branches of parse() now only construct their DOM SAX parser, call parse_dom(), and (for the callback parser) map a discarded top-level value to null. sax_parse() is left untouched, since it only runs the EOF check and release_lookahead() when sax_parse_internal() succeeded, unlike parse(), which runs them unconditionally. Behavior-preserving: same operations in the same order for both SAX parser kinds. Verified with unit-class_parser (strict/non-strict, callback and non-callback), unit-deserialization and unit-disabled_exceptions (JSON_NOEXCEPTION), plus a clean make amalgamate / make check-amalgamation diff. Overlaps #5601, which touches the same lines. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 2 * Share the code point to UTF-8 encoding between the wide-string helpers and the lexer The 1/2/3/4-byte UTF-8 encoding ladder was written out by hand three times: in wide_string_input_helper<..., 4>::fill_buffer() for a UTF-32 code point, in the UTF-16 helper for both a BMP code unit and a valid surrogate pair, and in the lexer's \uXXXX/\uXXXX\uYYYY handling. The copies had drifted: the UTF-32 helper masked the leading bits of each byte (& 0x1Fu, & 0x0Fu, & 0x07u) where the others relied on the shift alone, even though both give the same result for a code point that is already known to be in range. Add detail::encode_utf8(cp, out) in string_utils.hpp, a single encoder that invokes a callable once per output byte, most significant byte first. Use it in the three valid-code-point branches (UTF-32 code points up to U+10FFFF, UTF-16 code units outside the surrogate range, and valid UTF-16 surrogate pairs) and in the lexer's \u handling, where out forwards to add(). The UTF-16 helper's deliberate pass-through of malformed surrogate units and the UTF-32 helper's 0xFF sentinel for code points above U+10FFFF are untouched, since neither reaches the new helper. Behavior-preserving: same bytes in the same order for every valid code point, verified with unit-class_lexer, unit-class_parser, unit-deserialization, unit-wstring and the non-test-data parts of unit-unicode1..5 (ASan/UBSan, C++11/17/20), and an escape-heavy parse microbenchmark that shows no change (about 73 ms either way, median of 3, 1M escape sequences). single_include/ regenerated with make amalgamate; make check-amalgamation leaves a clean tree. Overlaps #5704, which rewrites the wide_string_input_helper specializations touched here. Signed-off-by: Niels Lohmann <mail@nlohmann.me> #5712 item 6 --------- Signed-off-by: Niels Lohmann <mail@nlohmann.me>
This commit is contained in:
@@ -3212,8 +3212,8 @@ class binary_reader
|
|||||||
return enter_object(detail::unknown_size());
|
return enter_object(detail::unknown_size());
|
||||||
}
|
}
|
||||||
|
|
||||||
// Note, no reader for UBJSON binary types is implemented because they do
|
// Note, UBJSON has no binary type of its own; BJData, which shares this
|
||||||
// not exist
|
// reader, decodes optimized 'B' arrays as binary in get_ubjson_array().
|
||||||
|
|
||||||
bool get_ubjson_high_precision_number()
|
bool get_ubjson_high_precision_number()
|
||||||
{
|
{
|
||||||
@@ -3919,7 +3919,7 @@ class binary_reader
|
|||||||
#endif
|
#endif
|
||||||
}
|
}
|
||||||
|
|
||||||
/*
|
/*!
|
||||||
@brief read a number from the input
|
@brief read a number from the input
|
||||||
|
|
||||||
@tparam NumberType the type of the number
|
@tparam NumberType the type of the number
|
||||||
@@ -3929,10 +3929,10 @@ class binary_reader
|
|||||||
@return whether conversion completed
|
@return whether conversion completed
|
||||||
|
|
||||||
@note This function needs to respect the system's endianness, because
|
@note This function needs to respect the system's endianness, because
|
||||||
bytes in CBOR, MessagePack, and UBJSON are stored in network order
|
bytes in CBOR, MessagePack, UBJSON, and BON8 are stored in network
|
||||||
(big endian) and therefore need reordering on little endian systems.
|
order (big endian) and therefore need reordering on little endian
|
||||||
On the other hand, BSON and BJData use little endian and should reorder
|
systems. On the other hand, BSON and BJData use little endian and
|
||||||
on big endian systems.
|
should reorder on big endian systems.
|
||||||
*/
|
*/
|
||||||
template<typename NumberType, bool InputIsLittleEndian = false>
|
template<typename NumberType, bool InputIsLittleEndian = false>
|
||||||
bool get_number(const input_format_t format, NumberType& result)
|
bool get_number(const input_format_t format, NumberType& result)
|
||||||
|
|||||||
@@ -8,12 +8,12 @@
|
|||||||
|
|
||||||
#pragma once
|
#pragma once
|
||||||
|
|
||||||
|
#include <algorithm> // min
|
||||||
#include <array> // array
|
#include <array> // array
|
||||||
#include <cstddef> // size_t
|
#include <cstddef> // size_t
|
||||||
|
#include <cstdint> // uint32_t
|
||||||
#include <cstring> // strlen
|
#include <cstring> // strlen
|
||||||
#include <iterator> // begin, end, iterator_traits, random_access_iterator_tag, distance, next
|
#include <iterator> // begin, end, iterator_traits, random_access_iterator_tag, distance, next
|
||||||
#include <memory> // shared_ptr, make_shared, addressof
|
|
||||||
#include <numeric> // accumulate
|
|
||||||
#include <streambuf> // streambuf
|
#include <streambuf> // streambuf
|
||||||
#include <string> // string, char_traits
|
#include <string> // string, char_traits
|
||||||
#include <type_traits> // enable_if, is_base_of, is_pointer, is_integral, remove_pointer
|
#include <type_traits> // enable_if, is_base_of, is_pointer, is_integral, remove_pointer
|
||||||
@@ -28,6 +28,7 @@
|
|||||||
#include <nlohmann/detail/iterators/iterator_traits.hpp>
|
#include <nlohmann/detail/iterators/iterator_traits.hpp>
|
||||||
#include <nlohmann/detail/macro_scope.hpp>
|
#include <nlohmann/detail/macro_scope.hpp>
|
||||||
#include <nlohmann/detail/meta/type_traits.hpp>
|
#include <nlohmann/detail/meta/type_traits.hpp>
|
||||||
|
#include <nlohmann/detail/string_utils.hpp>
|
||||||
|
|
||||||
NLOHMANN_JSON_NAMESPACE_BEGIN
|
NLOHMANN_JSON_NAMESPACE_BEGIN
|
||||||
namespace detail
|
namespace detail
|
||||||
@@ -82,8 +83,9 @@ class file_input_adapter
|
|||||||
};
|
};
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
Input adapter for a (caching) istream. Ignores a UFT Byte Order Mark at
|
Input adapter for a (caching) istream. Does not skip a UTF Byte Order Mark
|
||||||
beginning of input. Does not support changing the underlying std::streambuf
|
itself; that is done by the lexer's skip_bom(). Does not support changing
|
||||||
|
the underlying std::streambuf
|
||||||
in mid-input. Maintains underlying std::istream and std::streambuf to support
|
in mid-input. Maintains underlying std::istream and std::streambuf to support
|
||||||
subsequent use of standard std::istream operations to process any input
|
subsequent use of standard std::istream operations to process any input
|
||||||
characters following those used in parsing the JSON input. Clears the
|
characters following those used in parsing the JSON input. Clears the
|
||||||
@@ -454,32 +456,14 @@ struct wide_string_input_helper<BaseInputAdapter, 4>
|
|||||||
// get the current character
|
// get the current character
|
||||||
const auto wc = input.get_character();
|
const auto wc = input.get_character();
|
||||||
|
|
||||||
// UTF-32 to UTF-8 encoding
|
if (wc <= 0x10FFFF)
|
||||||
if (wc < 0x80)
|
|
||||||
{
|
{
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
|
// UTF-32 to UTF-8 encoding
|
||||||
utf8_bytes_filled = 1;
|
utf8_bytes_filled = 0;
|
||||||
}
|
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
||||||
else if (wc <= 0x7FF)
|
{
|
||||||
{
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xC0u | ((static_cast<unsigned int>(wc) >> 6u) & 0x1Fu));
|
});
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 2;
|
|
||||||
}
|
|
||||||
else if (wc <= 0xFFFF)
|
|
||||||
{
|
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xE0u | ((static_cast<unsigned int>(wc) >> 12u) & 0x0Fu));
|
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
|
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 3;
|
|
||||||
}
|
|
||||||
else if (wc <= 0x10FFFF)
|
|
||||||
{
|
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | ((static_cast<unsigned int>(wc) >> 18u) & 0x07u));
|
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 12u) & 0x3Fu));
|
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
|
|
||||||
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 4;
|
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
@@ -516,24 +500,15 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
|
|||||||
// get the current character
|
// get the current character
|
||||||
const auto wc = input.get_character();
|
const auto wc = input.get_character();
|
||||||
|
|
||||||
// UTF-16 to UTF-8 encoding
|
if (0xD800 > wc || wc >= 0xE000)
|
||||||
if (wc < 0x80)
|
|
||||||
{
|
{
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
|
// a UTF-16 code unit outside the surrogate range is a valid
|
||||||
utf8_bytes_filled = 1;
|
// code point (at most U+FFFF) on its own
|
||||||
}
|
utf8_bytes_filled = 0;
|
||||||
else if (wc <= 0x7FF)
|
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
||||||
{
|
{
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xC0u | ((static_cast<unsigned int>(wc) >> 6u)));
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
});
|
||||||
utf8_bytes_filled = 2;
|
|
||||||
}
|
|
||||||
else if (0xD800 > wc || wc >= 0xE000)
|
|
||||||
{
|
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xE0u | ((static_cast<unsigned int>(wc) >> 12u)));
|
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
|
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 3;
|
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
@@ -551,11 +526,11 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
|
|||||||
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
|
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
|
||||||
{
|
{
|
||||||
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
|
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
|
utf8_bytes_filled = 0;
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
|
encode_utf8(charcode, [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
|
{
|
||||||
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
||||||
utf8_bytes_filled = 4;
|
});
|
||||||
valid_pair = true;
|
valid_pair = true;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -884,9 +859,9 @@ auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) /
|
|||||||
return input_adapter(array, array + N);
|
return input_adapter(array, array + N);
|
||||||
}
|
}
|
||||||
|
|
||||||
// This class only handles inputs of input_buffer_adapter type.
|
// This class only handles inputs that construct a contiguous_bytes_input_adapter
|
||||||
// It's required so that expressions like {ptr, len} can be implicitly cast
|
// (e.g. span_input_adapter). It's required so that expressions like {ptr, len}
|
||||||
// to the correct adapter.
|
// can be implicitly cast to the correct adapter.
|
||||||
class span_input_adapter
|
class span_input_adapter
|
||||||
{
|
{
|
||||||
public:
|
public:
|
||||||
|
|||||||
@@ -10,6 +10,7 @@
|
|||||||
|
|
||||||
#include <algorithm> // find_if, min
|
#include <algorithm> // find_if, min
|
||||||
#include <cstddef>
|
#include <cstddef>
|
||||||
|
#include <limits> // numeric_limits
|
||||||
#include <string> // string
|
#include <string> // string
|
||||||
#include <type_traits> // enable_if_t
|
#include <type_traits> // enable_if_t
|
||||||
#include <utility> // move, pair
|
#include <utility> // move, pair
|
||||||
@@ -175,6 +176,88 @@ template<typename ArrayType>
|
|||||||
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
|
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
|
||||||
{}
|
{}
|
||||||
|
|
||||||
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
|
/*!
|
||||||
|
@brief set the diagnostic positions of a value the DOM SAX parsers just stored
|
||||||
|
|
||||||
|
Shared by json_sax_dom_parser and json_sax_dom_callback_parser. basic_json
|
||||||
|
befriends this struct, as the position members are private.
|
||||||
|
*/
|
||||||
|
struct diagnostic_positions
|
||||||
|
{
|
||||||
|
/*!
|
||||||
|
@param[in,out] v the value that was just parsed
|
||||||
|
@param[in] lexer the lexer that read it, or nullptr to leave @a v alone
|
||||||
|
*/
|
||||||
|
template<typename BasicJsonType, typename LexerType>
|
||||||
|
static void set_from_lexer(BasicJsonType& v, LexerType* lexer)
|
||||||
|
{
|
||||||
|
if (lexer)
|
||||||
|
{
|
||||||
|
// Lexer has read past the current field value, so set the end position to the current position.
|
||||||
|
// The start position will be set below based on the length of the string representation
|
||||||
|
// of the value.
|
||||||
|
v.end_position = lexer->get_position();
|
||||||
|
|
||||||
|
switch (v.type())
|
||||||
|
{
|
||||||
|
case value_t::boolean:
|
||||||
|
{
|
||||||
|
// 4 and 5 are the string length of "true" and "false"
|
||||||
|
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::null:
|
||||||
|
{
|
||||||
|
// 4 is the string length of "null"
|
||||||
|
v.start_position = v.end_position - 4;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::string:
|
||||||
|
{
|
||||||
|
// escape sequences make the token longer than the value it
|
||||||
|
// parses to, so the start position cannot be derived from
|
||||||
|
// the value; use the offset the lexer recorded instead
|
||||||
|
v.start_position = lexer->get_token_start_position();
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::discarded:
|
||||||
|
{
|
||||||
|
// an object or array the callback of
|
||||||
|
// json_sax_dom_callback_parser rejected has no position
|
||||||
|
v.end_position = std::string::npos;
|
||||||
|
v.start_position = v.end_position;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::binary:
|
||||||
|
case value_t::number_integer:
|
||||||
|
case value_t::number_unsigned:
|
||||||
|
case value_t::number_float:
|
||||||
|
{
|
||||||
|
v.start_position = v.end_position - lexer->get_string().size();
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::object:
|
||||||
|
case value_t::array:
|
||||||
|
{
|
||||||
|
// object and array are handled in start_object() and start_array() handlers
|
||||||
|
// skip setting the values here.
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
default: // LCOV_EXCL_LINE
|
||||||
|
// Handle all possible types discretely, default handler should never be reached.
|
||||||
|
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
};
|
||||||
|
#endif
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
@brief SAX implementation to create a JSON value from SAX events
|
@brief SAX implementation to create a JSON value from SAX events
|
||||||
|
|
||||||
@@ -376,76 +459,6 @@ class json_sax_dom_parser
|
|||||||
|
|
||||||
private:
|
private:
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
|
||||||
void handle_diagnostic_positions_for_json_value(BasicJsonType& v)
|
|
||||||
{
|
|
||||||
if (m_lexer_ref)
|
|
||||||
{
|
|
||||||
// Lexer has read past the current field value, so set the end position to the current position.
|
|
||||||
// The start position will be set below based on the length of the string representation
|
|
||||||
// of the value.
|
|
||||||
v.end_position = m_lexer_ref->get_position();
|
|
||||||
|
|
||||||
switch (v.type())
|
|
||||||
{
|
|
||||||
case value_t::boolean:
|
|
||||||
{
|
|
||||||
// 4 and 5 are the string length of "true" and "false"
|
|
||||||
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::null:
|
|
||||||
{
|
|
||||||
// 4 is the string length of "null"
|
|
||||||
v.start_position = v.end_position - 4;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::string:
|
|
||||||
{
|
|
||||||
// escape sequences make the token longer than the value it
|
|
||||||
// parses to, so the start position cannot be derived from
|
|
||||||
// the value; use the offset the lexer recorded instead
|
|
||||||
v.start_position = m_lexer_ref->get_token_start_position();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
// As we handle the start and end positions for values created during parsing,
|
|
||||||
// we do not expect the following value type to be called. Regardless, set the positions
|
|
||||||
// in case this is created manually or through a different constructor. Exclude from lcov
|
|
||||||
// since the exact condition of this switch is esoteric.
|
|
||||||
// LCOV_EXCL_START
|
|
||||||
case value_t::discarded:
|
|
||||||
{
|
|
||||||
v.end_position = std::string::npos;
|
|
||||||
v.start_position = v.end_position;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
// LCOV_EXCL_STOP
|
|
||||||
case value_t::binary:
|
|
||||||
case value_t::number_integer:
|
|
||||||
case value_t::number_unsigned:
|
|
||||||
case value_t::number_float:
|
|
||||||
{
|
|
||||||
v.start_position = v.end_position - m_lexer_ref->get_string().size();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
case value_t::object:
|
|
||||||
case value_t::array:
|
|
||||||
{
|
|
||||||
// object and array are handled in start_object() and start_array() handlers
|
|
||||||
// skip setting the values here.
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
default: // LCOV_EXCL_LINE
|
|
||||||
// Handle all possible types discretely, default handler should never be reached.
|
|
||||||
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert,-warnings-as-errors) LCOV_EXCL_LINE
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
#endif
|
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
@invariant If the ref stack is empty, then the passed value will be the new
|
@invariant If the ref stack is empty, then the passed value will be the new
|
||||||
root.
|
root.
|
||||||
@@ -461,7 +474,7 @@ class json_sax_dom_parser
|
|||||||
root = BasicJsonType(std::forward<Value>(v));
|
root = BasicJsonType(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(root);
|
diagnostic_positions::set_from_lexer(root, m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
return &root;
|
return &root;
|
||||||
@@ -474,7 +487,7 @@ class json_sax_dom_parser
|
|||||||
ref_stack.back()->m_data.m_value.array->emplace_back(std::forward<Value>(v));
|
ref_stack.back()->m_data.m_value.array->emplace_back(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(ref_stack.back()->m_data.m_value.array->back());
|
diagnostic_positions::set_from_lexer(ref_stack.back()->m_data.m_value.array->back(), m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
return &(ref_stack.back()->m_data.m_value.array->back());
|
return &(ref_stack.back()->m_data.m_value.array->back());
|
||||||
@@ -485,7 +498,7 @@ class json_sax_dom_parser
|
|||||||
*object_element = BasicJsonType(std::forward<Value>(v));
|
*object_element = BasicJsonType(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(*object_element);
|
diagnostic_positions::set_from_lexer(*object_element, m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
return object_element;
|
return object_element;
|
||||||
@@ -674,7 +687,7 @@ class json_sax_dom_callback_parser
|
|||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
// Set start/end positions for discarded object.
|
// Set start/end positions for discarded object.
|
||||||
handle_diagnostic_positions_for_json_value(*ref_stack.back());
|
diagnostic_positions::set_from_lexer(*ref_stack.back(), m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -790,7 +803,7 @@ class json_sax_dom_callback_parser
|
|||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
// Set start/end positions for discarded array.
|
// Set start/end positions for discarded array.
|
||||||
handle_diagnostic_positions_for_json_value(*ref_stack.back());
|
diagnostic_positions::set_from_lexer(*ref_stack.back(), m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -843,72 +856,6 @@ class json_sax_dom_callback_parser
|
|||||||
|
|
||||||
private:
|
private:
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
|
||||||
void handle_diagnostic_positions_for_json_value(BasicJsonType& v)
|
|
||||||
{
|
|
||||||
if (m_lexer_ref)
|
|
||||||
{
|
|
||||||
// Lexer has read past the current field value, so set the end position to the current position.
|
|
||||||
// The start position will be set below based on the length of the string representation
|
|
||||||
// of the value.
|
|
||||||
v.end_position = m_lexer_ref->get_position();
|
|
||||||
|
|
||||||
switch (v.type())
|
|
||||||
{
|
|
||||||
case value_t::boolean:
|
|
||||||
{
|
|
||||||
// 4 and 5 are the string length of "true" and "false"
|
|
||||||
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::null:
|
|
||||||
{
|
|
||||||
// 4 is the string length of "null"
|
|
||||||
v.start_position = v.end_position - 4;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::string:
|
|
||||||
{
|
|
||||||
// escape sequences make the token longer than the value it
|
|
||||||
// parses to, so the start position cannot be derived from
|
|
||||||
// the value; use the offset the lexer recorded instead
|
|
||||||
v.start_position = m_lexer_ref->get_token_start_position();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::discarded:
|
|
||||||
{
|
|
||||||
v.end_position = std::string::npos;
|
|
||||||
v.start_position = v.end_position;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::binary:
|
|
||||||
case value_t::number_integer:
|
|
||||||
case value_t::number_unsigned:
|
|
||||||
case value_t::number_float:
|
|
||||||
{
|
|
||||||
v.start_position = v.end_position - m_lexer_ref->get_string().size();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::object:
|
|
||||||
case value_t::array:
|
|
||||||
{
|
|
||||||
// object and array are handled in start_object() and start_array() handlers
|
|
||||||
// skip setting the values here.
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
default: // LCOV_EXCL_LINE
|
|
||||||
// Handle all possible types discretely, default handler should never be reached.
|
|
||||||
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert,-warnings-as-errors) LCOV_EXCL_LINE
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
#endif
|
|
||||||
|
|
||||||
/// if there is a pending duplicate-key stash entry for this exact slot,
|
/// if there is a pending duplicate-key stash entry for this exact slot,
|
||||||
/// remove it from the stash; if restore_value is true, the stashed
|
/// remove it from the stash; if restore_value is true, the stashed
|
||||||
/// previous value is moved back into the slot first (use this when the
|
/// previous value is moved back into the slot first (use this when the
|
||||||
@@ -1030,7 +977,7 @@ class json_sax_dom_callback_parser
|
|||||||
auto value = BasicJsonType(std::forward<Value>(v));
|
auto value = BasicJsonType(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(value);
|
diagnostic_positions::set_from_lexer(value, m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
// check callback
|
// check callback
|
||||||
|
|||||||
@@ -10,6 +10,7 @@
|
|||||||
|
|
||||||
#include <array> // array
|
#include <array> // array
|
||||||
#include <cstddef> // size_t
|
#include <cstddef> // size_t
|
||||||
|
#include <cstdint> // uint32_t
|
||||||
#include <cstdio> // snprintf
|
#include <cstdio> // snprintf
|
||||||
#include <initializer_list> // initializer_list
|
#include <initializer_list> // initializer_list
|
||||||
#include <string> // char_traits, string
|
#include <string> // char_traits, string
|
||||||
@@ -22,6 +23,7 @@
|
|||||||
#include <nlohmann/detail/input/string_scan.hpp>
|
#include <nlohmann/detail/input/string_scan.hpp>
|
||||||
#include <nlohmann/detail/macro_scope.hpp>
|
#include <nlohmann/detail/macro_scope.hpp>
|
||||||
#include <nlohmann/detail/meta/type_traits.hpp>
|
#include <nlohmann/detail/meta/type_traits.hpp>
|
||||||
|
#include <nlohmann/detail/string_utils.hpp>
|
||||||
|
|
||||||
NLOHMANN_JSON_NAMESPACE_BEGIN
|
NLOHMANN_JSON_NAMESPACE_BEGIN
|
||||||
namespace detail
|
namespace detail
|
||||||
@@ -486,32 +488,10 @@ class lexer : public lexer_base<BasicJsonType>
|
|||||||
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
|
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
|
||||||
|
|
||||||
// translate codepoint into bytes
|
// translate codepoint into bytes
|
||||||
if (codepoint < 0x80)
|
encode_utf8(static_cast<std::uint32_t>(codepoint), [this](std::uint32_t byte)
|
||||||
{
|
{
|
||||||
// 1-byte characters: 0xxxxxxx (ASCII)
|
add(static_cast<char_int_type>(byte));
|
||||||
add(static_cast<char_int_type>(codepoint));
|
});
|
||||||
}
|
|
||||||
else if (codepoint <= 0x7FF)
|
|
||||||
{
|
|
||||||
// 2-byte characters: 110xxxxx 10xxxxxx
|
|
||||||
add(static_cast<char_int_type>(0xC0u | (static_cast<unsigned int>(codepoint) >> 6u)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
|
|
||||||
}
|
|
||||||
else if (codepoint <= 0xFFFF)
|
|
||||||
{
|
|
||||||
// 3-byte characters: 1110xxxx 10xxxxxx 10xxxxxx
|
|
||||||
add(static_cast<char_int_type>(0xE0u | (static_cast<unsigned int>(codepoint) >> 12u)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 6u) & 0x3Fu)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
// 4-byte characters: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
|
|
||||||
add(static_cast<char_int_type>(0xF0u | (static_cast<unsigned int>(codepoint) >> 18u)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 12u) & 0x3Fu)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 6u) & 0x3Fu)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
|
|
||||||
}
|
|
||||||
|
|
||||||
break;
|
break;
|
||||||
}
|
}
|
||||||
@@ -1409,45 +1389,30 @@ scan_number_done:
|
|||||||
*/
|
*/
|
||||||
token_type convert_number(token_type number_type, std::size_t mantissa_end)
|
token_type convert_number(token_type number_type, std::size_t mantissa_end)
|
||||||
{
|
{
|
||||||
// If the caller does not need the converted value (only whether the
|
// accept() only needs to know whether the input is valid, so it sets
|
||||||
// input is syntactically valid; see json_sax_acceptor/accept()), an
|
// discard_number_values (see json.hpp), and an integer token whose
|
||||||
// unsigned/integer token can be reported without calling
|
// digit count shows that it fits is reported without calling
|
||||||
// strtoull()/strtoll() at all, *provided* we can already tell from
|
// convert_integer(). A number with up to 18 digits always fits into
|
||||||
// the digit count alone that the conversion cannot overflow 64 bits.
|
// both std::uint64_t and std::int64_t (18 nines is about 1e18, below
|
||||||
// Such tokens are always finite and are accepted unconditionally by
|
// INT64_MAX, which is about 9.2e18). Longer tokens take the exact path
|
||||||
// the parser regardless of their actual value (parser::sax_parse_internal()
|
// below, including the fallback to floating point when the value does
|
||||||
// never checks finiteness for value_unsigned/value_integer), so the
|
// not fit.
|
||||||
// classification below is all that is needed.
|
|
||||||
//
|
//
|
||||||
// A decimal number with up to 18 digits is always representable in
|
// With a narrower number_unsigned_t/number_integer_t (e.g.
|
||||||
// both std::uint64_t and std::int64_t (18 nines is ~1e18, well below
|
// std::uint32_t), the exact path would reclassify some of these tokens
|
||||||
// both UINT64_MAX ~1.8e19 and INT64_MAX ~9.2e18), so strtoull()/strtoll()
|
// as (finite) floats, while this check reports integers. That does not
|
||||||
// could not have set errno to ERANGE for it. Numbers with more digits
|
// change the result of accept(): it always parses through
|
||||||
// (rare in practice) fall through to the exact code below, unchanged,
|
// json_sax_acceptor, whose number callbacks discard their argument and
|
||||||
// so their handling -- including reclassification to value_float when
|
// return true, and the parser rejects neither integers nor finite
|
||||||
// the value overflows 64 bits, and rejection when it is not even
|
// floats. value_unsigned/value_integer are left unset here, so a caller
|
||||||
// finite as a double -- is bit-for-bit identical to before this
|
// that reads the converted value must not set discard_number_values.
|
||||||
// optimization.
|
|
||||||
//
|
//
|
||||||
// Note this reasons about std::uint64_t/std::int64_t, not about
|
// On contiguous input, scan_number_bulk_contiguous() converts integer
|
||||||
// number_unsigned_t/number_integer_t (BasicJsonType's own, possibly
|
// tokens itself and does not pass them to this function, unless
|
||||||
// narrower, template parameters -- e.g. std::uint32_t). That is fine
|
// JSON_DIAGNOSTIC_POSITIONS is enabled. This check is therefore only
|
||||||
// *only* because discard_number_values is exclusively set by
|
// reached for input without bulk access (e.g. streams), with
|
||||||
// accept() (see json.hpp), and accept() always parses through the
|
// JSON_DIAGNOSTIC_POSITIONS, or when scan_number_bulk_contiguous()
|
||||||
// library's own json_sax_acceptor -- never a user-supplied SAX
|
// falls back to scan_number().
|
||||||
// consumer -- whose number_unsigned()/number_integer()/number_float()
|
|
||||||
// callbacks unconditionally discard their argument and return true.
|
|
||||||
// So for every caller that can reach this branch, neither the token
|
|
||||||
// classification below nor the eventual (possibly narrowed, and on
|
|
||||||
// this fast path left stale/unset) value_unsigned/value_integer is
|
|
||||||
// ever consulted -- an unsigned/integer token is accepted outright,
|
|
||||||
// and even a >18-digit token that this fast path deliberately falls
|
|
||||||
// through for is, once reclassified to value_float, still finite
|
|
||||||
// (and thus accepted) for any digit count that fits in number_unsigned_t
|
|
||||||
// or number_integer_t regardless of that type's width. If this
|
|
||||||
// function is ever taught to run with discard_number_values true for
|
|
||||||
// a caller that *does* read the converted value, this reasoning (and
|
|
||||||
// the fast path below) would need to be revisited.
|
|
||||||
if (discard_number_values)
|
if (discard_number_values)
|
||||||
{
|
{
|
||||||
constexpr std::size_t safe_digit_count = 18;
|
constexpr std::size_t safe_digit_count = 18;
|
||||||
@@ -1876,7 +1841,7 @@ scan_number_done:
|
|||||||
return value_float;
|
return value_float;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// return current string value (implicitly resets the token; useful only once)
|
/// return current string value
|
||||||
string_t& get_string()
|
string_t& get_string()
|
||||||
{
|
{
|
||||||
// a number token holds '.' regardless of the locale (#4084)
|
// a number token holds '.' regardless of the locale (#4084)
|
||||||
@@ -2178,11 +2143,11 @@ scan_number_done:
|
|||||||
/// the position of the decimal point in token_buffer
|
/// the position of the decimal point in token_buffer
|
||||||
std::size_t decimal_point_position = std::string::npos;
|
std::size_t decimal_point_position = std::string::npos;
|
||||||
|
|
||||||
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
|
/// whether the caller only needs the token types and never looks at the
|
||||||
/// token classification and never looks at the converted numeric value;
|
/// converted numeric values; set only by accept(), which parses through
|
||||||
/// when set, scan_number() may skip strtoull()/strtoll() for
|
/// json_sax_acceptor. When set, convert_number() skips converting integer
|
||||||
/// value_unsigned/value_integer tokens whose digit count guarantees they
|
/// tokens whose digit count guarantees that they fit into 64 bits (see
|
||||||
/// fit into 64 bits (see scan_number())
|
/// there)
|
||||||
const bool discard_number_values = false;
|
const bool discard_number_values = false;
|
||||||
};
|
};
|
||||||
|
|
||||||
|
|||||||
@@ -54,7 +54,8 @@ using parser_callback_t =
|
|||||||
/*!
|
/*!
|
||||||
@brief syntax analysis
|
@brief syntax analysis
|
||||||
|
|
||||||
This class implements a recursive descent parser.
|
This class implements an iterative parser that keeps the open containers on
|
||||||
|
an explicit stack and reports what it reads as SAX events.
|
||||||
*/
|
*/
|
||||||
template<typename BasicJsonType, typename InputAdapterType>
|
template<typename BasicJsonType, typename InputAdapterType>
|
||||||
class parser
|
class parser
|
||||||
@@ -98,28 +99,9 @@ class parser
|
|||||||
if (callback)
|
if (callback)
|
||||||
{
|
{
|
||||||
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
|
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
|
||||||
sax_parse_internal(&sdp);
|
|
||||||
|
|
||||||
if (strict)
|
|
||||||
{
|
|
||||||
// in strict mode, input must be completely read
|
|
||||||
if (get_token() != token_type::end_of_input)
|
|
||||||
{
|
|
||||||
sdp.parse_error(m_lexer.get_position(),
|
|
||||||
m_lexer.get_token_string(),
|
|
||||||
parse_error::create(101, m_lexer.get_position(),
|
|
||||||
exception_message(token_type::end_of_input, "value"), nullptr));
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
// the caller keeps using the input: position it right after
|
|
||||||
// the value by leaving the character that terminated it
|
|
||||||
m_lexer.release_lookahead();
|
|
||||||
}
|
|
||||||
|
|
||||||
// in case of an error, return a discarded value
|
// in case of an error, return a discarded value
|
||||||
if (sdp.is_errored())
|
if (!parse_dom(sdp, strict))
|
||||||
{
|
{
|
||||||
result = value_t::discarded;
|
result = value_t::discarded;
|
||||||
return;
|
return;
|
||||||
@@ -135,26 +117,9 @@ class parser
|
|||||||
else
|
else
|
||||||
{
|
{
|
||||||
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
|
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
|
||||||
sax_parse_internal(&sdp);
|
|
||||||
|
|
||||||
if (strict)
|
|
||||||
{
|
|
||||||
// in strict mode, input must be completely read
|
|
||||||
if (get_token() != token_type::end_of_input)
|
|
||||||
{
|
|
||||||
sdp.parse_error(m_lexer.get_position(),
|
|
||||||
m_lexer.get_token_string(),
|
|
||||||
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr));
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
// see above
|
|
||||||
m_lexer.release_lookahead();
|
|
||||||
}
|
|
||||||
|
|
||||||
// in case of an error, return a discarded value
|
// in case of an error, return a discarded value
|
||||||
if (sdp.is_errored())
|
if (!parse_dom(sdp, strict))
|
||||||
{
|
{
|
||||||
result = value_t::discarded;
|
result = value_t::discarded;
|
||||||
return;
|
return;
|
||||||
@@ -207,6 +172,46 @@ class parser
|
|||||||
}
|
}
|
||||||
|
|
||||||
private:
|
private:
|
||||||
|
/*!
|
||||||
|
@brief run a DOM SAX parser to completion and position the lexer
|
||||||
|
|
||||||
|
Shared by both branches of @ref parse(): builds no SAX parser itself,
|
||||||
|
but drives an already-constructed @a json_sax_dom_parser or
|
||||||
|
@ref json_sax_dom_callback_parser through @ref sax_parse_internal(),
|
||||||
|
then applies the strict-EOF check (reporting parse_error.101 through
|
||||||
|
@a sdp on failure) or, in non-strict mode, releases the lookahead so
|
||||||
|
the caller can keep reading the input right after the parsed value.
|
||||||
|
|
||||||
|
@param[in,out] sdp the DOM SAX parser to run
|
||||||
|
@param[in] strict whether to expect the last token to be EOF
|
||||||
|
@return whether @a sdp did not report an error
|
||||||
|
*/
|
||||||
|
template<typename DomSax>
|
||||||
|
bool parse_dom(DomSax& sdp, const bool strict)
|
||||||
|
{
|
||||||
|
sax_parse_internal(&sdp);
|
||||||
|
|
||||||
|
if (strict)
|
||||||
|
{
|
||||||
|
// in strict mode, input must be completely read
|
||||||
|
if (get_token() != token_type::end_of_input)
|
||||||
|
{
|
||||||
|
sdp.parse_error(m_lexer.get_position(),
|
||||||
|
m_lexer.get_token_string(),
|
||||||
|
parse_error::create(101, m_lexer.get_position(),
|
||||||
|
exception_message(token_type::end_of_input, "value"), nullptr));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
// the caller keeps using the input: position it right after
|
||||||
|
// the value by leaving the character that terminated it
|
||||||
|
m_lexer.release_lookahead();
|
||||||
|
}
|
||||||
|
|
||||||
|
return !sdp.is_errored();
|
||||||
|
}
|
||||||
|
|
||||||
template<typename SAX>
|
template<typename SAX>
|
||||||
JSON_HEDLEY_NON_NULL(2)
|
JSON_HEDLEY_NON_NULL(2)
|
||||||
bool sax_parse_internal(SAX* sax)
|
bool sax_parse_internal(SAX* sax)
|
||||||
@@ -439,8 +444,9 @@ class parser
|
|||||||
|
|
||||||
// We are done with this array. Before we can parse a
|
// We are done with this array. Before we can parse a
|
||||||
// new value, we need to evaluate the new state first.
|
// new value, we need to evaluate the new state first.
|
||||||
// By setting skip_to_state_evaluation to false, we
|
// By setting skip_to_state_evaluation to true, the next
|
||||||
// are effectively jumping to the beginning of this if.
|
// iteration skips parsing a value and evaluates the
|
||||||
|
// enclosing state directly.
|
||||||
JSON_ASSERT(!states.empty());
|
JSON_ASSERT(!states.empty());
|
||||||
states.pop_back();
|
states.pop_back();
|
||||||
skip_to_state_evaluation = true;
|
skip_to_state_evaluation = true;
|
||||||
@@ -500,8 +506,9 @@ class parser
|
|||||||
|
|
||||||
// We are done with this object. Before we can parse a
|
// We are done with this object. Before we can parse a
|
||||||
// new value, we need to evaluate the new state first.
|
// new value, we need to evaluate the new state first.
|
||||||
// By setting skip_to_state_evaluation to false, we
|
// By setting skip_to_state_evaluation to true, the next
|
||||||
// are effectively jumping to the beginning of this if.
|
// iteration skips parsing a value and evaluates the
|
||||||
|
// enclosing state directly.
|
||||||
JSON_ASSERT(!states.empty());
|
JSON_ASSERT(!states.empty());
|
||||||
states.pop_back();
|
states.pop_back();
|
||||||
skip_to_state_evaluation = true;
|
skip_to_state_evaluation = true;
|
||||||
|
|||||||
@@ -47,6 +47,61 @@ inline std::string hex_byte(const std::uint8_t byte)
|
|||||||
return result;
|
return result;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
///////////////////
|
||||||
|
// UTF-8 encoding //
|
||||||
|
///////////////////
|
||||||
|
|
||||||
|
/*!
|
||||||
|
@brief encode a Unicode code point as UTF-8
|
||||||
|
|
||||||
|
Used to turn a decoded code point back into bytes: by the wide-string input
|
||||||
|
adapters in input_adapters.hpp (one code point per UTF-32 unit, per UTF-16
|
||||||
|
unit outside the surrogate range, and per valid UTF-16 surrogate pair), and
|
||||||
|
by the lexer's `\uXXXX`/`\uXXXX\uYYYY` handling in lexer.hpp. Passing a
|
||||||
|
code point above U+10FFFF, or one in the surrogate range U+D800..U+DFFF, is
|
||||||
|
undefined behavior; callers are expected to have rejected those already
|
||||||
|
(the wide-string adapters pass malformed units through unencoded instead of
|
||||||
|
calling this function, and the lexer rejects unpaired surrogates before
|
||||||
|
reaching it).
|
||||||
|
|
||||||
|
@tparam Out a callable invoked with one byte (as std::uint32_t, 0x00..0xFF)
|
||||||
|
at a time, most significant byte first
|
||||||
|
@param[in] cp the code point to encode (at most U+10FFFF)
|
||||||
|
@param[in] out called once for each byte of the UTF-8 encoding of @a cp
|
||||||
|
*/
|
||||||
|
template<typename Out>
|
||||||
|
void encode_utf8(std::uint32_t cp, Out&& out)
|
||||||
|
{
|
||||||
|
JSON_ASSERT(cp <= 0x10FFFF);
|
||||||
|
|
||||||
|
if (cp < 0x80)
|
||||||
|
{
|
||||||
|
// 1-byte characters: 0xxxxxxx (ASCII)
|
||||||
|
out(cp);
|
||||||
|
}
|
||||||
|
else if (cp <= 0x7FF)
|
||||||
|
{
|
||||||
|
// 2-byte characters: 110xxxxx 10xxxxxx
|
||||||
|
out(0xC0u | (cp >> 6u));
|
||||||
|
out(0x80u | (cp & 0x3Fu));
|
||||||
|
}
|
||||||
|
else if (cp <= 0xFFFF)
|
||||||
|
{
|
||||||
|
// 3-byte characters: 1110xxxx 10xxxxxx 10xxxxxx
|
||||||
|
out(0xE0u | (cp >> 12u));
|
||||||
|
out(0x80u | ((cp >> 6u) & 0x3Fu));
|
||||||
|
out(0x80u | (cp & 0x3Fu));
|
||||||
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
// 4-byte characters: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
|
||||||
|
out(0xF0u | (cp >> 18u));
|
||||||
|
out(0x80u | ((cp >> 12u) & 0x3Fu));
|
||||||
|
out(0x80u | ((cp >> 6u) & 0x3Fu));
|
||||||
|
out(0x80u | (cp & 0x3Fu));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
///////////////////
|
///////////////////
|
||||||
// UTF-8 decoding //
|
// UTF-8 decoding //
|
||||||
///////////////////
|
///////////////////
|
||||||
@@ -62,11 +117,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
|
|||||||
written by Björn Hoehrmann. See
|
written by Björn Hoehrmann. See
|
||||||
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
||||||
|
|
||||||
This decoder is the single source of truth for UTF-8 validation in this
|
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
|
||||||
library: it is used both by the serializer (to escape and, in strict mode,
|
places, which differ in speed, diagnostics, and how they read the input:
|
||||||
reject ill-formed UTF-8 when dumping a string) and by the binary readers
|
|
||||||
(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at
|
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
|
||||||
decode time; see @ref is_valid_utf8 below).
|
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
|
||||||
|
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
|
||||||
|
text strings at decode time).
|
||||||
|
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
|
||||||
|
diagnostic for each kind of error.
|
||||||
|
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
|
||||||
|
bulk string scan, the bulk path of the BON8 reader, and the BON8 writer.
|
||||||
|
They must accept exactly what the lexer's switch accepts.
|
||||||
|
- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk
|
||||||
|
access, and the bytes the bulk path leaves to it.
|
||||||
|
|
||||||
|
All four must accept the same set of sequences, so a change to one needs a
|
||||||
|
matching change to the others.
|
||||||
|
|
||||||
@param[in,out] state the current decoder state
|
@param[in,out] state the current decoder state
|
||||||
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
||||||
|
|||||||
@@ -164,6 +164,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
|
|||||||
friend class ::nlohmann::detail::json_sax_dom_parser;
|
friend class ::nlohmann::detail::json_sax_dom_parser;
|
||||||
template<typename BasicJsonType, typename InputAdapterType>
|
template<typename BasicJsonType, typename InputAdapterType>
|
||||||
friend class ::nlohmann::detail::json_sax_dom_callback_parser;
|
friend class ::nlohmann::detail::json_sax_dom_callback_parser;
|
||||||
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
|
friend struct ::nlohmann::detail::diagnostic_positions;
|
||||||
|
#endif
|
||||||
friend class ::nlohmann::detail::exception;
|
friend class ::nlohmann::detail::exception;
|
||||||
|
|
||||||
/// workaround type for MSVC
|
/// workaround type for MSVC
|
||||||
|
|||||||
+285
-319
@@ -6269,6 +6269,61 @@ inline std::string hex_byte(const std::uint8_t byte)
|
|||||||
return result;
|
return result;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
///////////////////
|
||||||
|
// UTF-8 encoding //
|
||||||
|
///////////////////
|
||||||
|
|
||||||
|
/*!
|
||||||
|
@brief encode a Unicode code point as UTF-8
|
||||||
|
|
||||||
|
Used to turn a decoded code point back into bytes: by the wide-string input
|
||||||
|
adapters in input_adapters.hpp (one code point per UTF-32 unit, per UTF-16
|
||||||
|
unit outside the surrogate range, and per valid UTF-16 surrogate pair), and
|
||||||
|
by the lexer's `\uXXXX`/`\uXXXX\uYYYY` handling in lexer.hpp. Passing a
|
||||||
|
code point above U+10FFFF, or one in the surrogate range U+D800..U+DFFF, is
|
||||||
|
undefined behavior; callers are expected to have rejected those already
|
||||||
|
(the wide-string adapters pass malformed units through unencoded instead of
|
||||||
|
calling this function, and the lexer rejects unpaired surrogates before
|
||||||
|
reaching it).
|
||||||
|
|
||||||
|
@tparam Out a callable invoked with one byte (as std::uint32_t, 0x00..0xFF)
|
||||||
|
at a time, most significant byte first
|
||||||
|
@param[in] cp the code point to encode (at most U+10FFFF)
|
||||||
|
@param[in] out called once for each byte of the UTF-8 encoding of @a cp
|
||||||
|
*/
|
||||||
|
template<typename Out>
|
||||||
|
void encode_utf8(std::uint32_t cp, Out&& out)
|
||||||
|
{
|
||||||
|
JSON_ASSERT(cp <= 0x10FFFF);
|
||||||
|
|
||||||
|
if (cp < 0x80)
|
||||||
|
{
|
||||||
|
// 1-byte characters: 0xxxxxxx (ASCII)
|
||||||
|
out(cp);
|
||||||
|
}
|
||||||
|
else if (cp <= 0x7FF)
|
||||||
|
{
|
||||||
|
// 2-byte characters: 110xxxxx 10xxxxxx
|
||||||
|
out(0xC0u | (cp >> 6u));
|
||||||
|
out(0x80u | (cp & 0x3Fu));
|
||||||
|
}
|
||||||
|
else if (cp <= 0xFFFF)
|
||||||
|
{
|
||||||
|
// 3-byte characters: 1110xxxx 10xxxxxx 10xxxxxx
|
||||||
|
out(0xE0u | (cp >> 12u));
|
||||||
|
out(0x80u | ((cp >> 6u) & 0x3Fu));
|
||||||
|
out(0x80u | (cp & 0x3Fu));
|
||||||
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
// 4-byte characters: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
|
||||||
|
out(0xF0u | (cp >> 18u));
|
||||||
|
out(0x80u | ((cp >> 12u) & 0x3Fu));
|
||||||
|
out(0x80u | ((cp >> 6u) & 0x3Fu));
|
||||||
|
out(0x80u | (cp & 0x3Fu));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
///////////////////
|
///////////////////
|
||||||
// UTF-8 decoding //
|
// UTF-8 decoding //
|
||||||
///////////////////
|
///////////////////
|
||||||
@@ -6284,11 +6339,23 @@ This is a single-byte step of a "shift-based" UTF-8 decoder originally
|
|||||||
written by Björn Hoehrmann. See
|
written by Björn Hoehrmann. See
|
||||||
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
http://bjoern.hoehrmann.de/utf-8/decoder/dfa/ for details.
|
||||||
|
|
||||||
This decoder is the single source of truth for UTF-8 validation in this
|
The library checks UTF-8 well-formedness (RFC 3629, section 4) in four
|
||||||
library: it is used both by the serializer (to escape and, in strict mode,
|
places, which differ in speed, diagnostics, and how they read the input:
|
||||||
reject ill-formed UTF-8 when dumping a string) and by the binary readers
|
|
||||||
(to reject ill-formed UTF-8 in CBOR/MessagePack/BSON/UBJSON text strings at
|
- decode() and @ref is_valid_utf8 below: the serializer (to escape and, in
|
||||||
decode time; see @ref is_valid_utf8 below).
|
strict mode, reject ill-formed UTF-8 when dumping a string) and the CBOR,
|
||||||
|
MessagePack, BSON, UBJSON and BJData readers (to reject ill-formed UTF-8 in
|
||||||
|
text strings at decode time).
|
||||||
|
- the per-lead-byte switch in lexer::scan_string(): JSON text, with a
|
||||||
|
diagnostic for each kind of error.
|
||||||
|
- validate_one_utf8() and valid_utf8_prefix() in string_scan.hpp: the lexer's
|
||||||
|
bulk string scan, the bulk path of the BON8 reader, and the BON8 writer.
|
||||||
|
They must accept exactly what the lexer's switch accepts.
|
||||||
|
- the byte path of binary_reader::get_bon8_string(): BON8 input without bulk
|
||||||
|
access, and the bytes the bulk path leaves to it.
|
||||||
|
|
||||||
|
All four must accept the same set of sequences, so a change to one needs a
|
||||||
|
matching change to the others.
|
||||||
|
|
||||||
@param[in,out] state the current decoder state
|
@param[in,out] state the current decoder state
|
||||||
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
@param[in,out] codep codepoint (valid only if resulting state is UTF8_ACCEPT)
|
||||||
@@ -7591,12 +7658,12 @@ NLOHMANN_JSON_NAMESPACE_END
|
|||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
#include <algorithm> // min
|
||||||
#include <array> // array
|
#include <array> // array
|
||||||
#include <cstddef> // size_t
|
#include <cstddef> // size_t
|
||||||
|
#include <cstdint> // uint32_t
|
||||||
#include <cstring> // strlen
|
#include <cstring> // strlen
|
||||||
#include <iterator> // begin, end, iterator_traits, random_access_iterator_tag, distance, next
|
#include <iterator> // begin, end, iterator_traits, random_access_iterator_tag, distance, next
|
||||||
#include <memory> // shared_ptr, make_shared, addressof
|
|
||||||
#include <numeric> // accumulate
|
|
||||||
#include <streambuf> // streambuf
|
#include <streambuf> // streambuf
|
||||||
#include <string> // string, char_traits
|
#include <string> // string, char_traits
|
||||||
#include <type_traits> // enable_if, is_base_of, is_pointer, is_integral, remove_pointer
|
#include <type_traits> // enable_if, is_base_of, is_pointer, is_integral, remove_pointer
|
||||||
@@ -7615,6 +7682,8 @@ NLOHMANN_JSON_NAMESPACE_END
|
|||||||
|
|
||||||
// #include <nlohmann/detail/meta/type_traits.hpp>
|
// #include <nlohmann/detail/meta/type_traits.hpp>
|
||||||
|
|
||||||
|
// #include <nlohmann/detail/string_utils.hpp>
|
||||||
|
|
||||||
|
|
||||||
NLOHMANN_JSON_NAMESPACE_BEGIN
|
NLOHMANN_JSON_NAMESPACE_BEGIN
|
||||||
namespace detail
|
namespace detail
|
||||||
@@ -7669,8 +7738,9 @@ class file_input_adapter
|
|||||||
};
|
};
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
Input adapter for a (caching) istream. Ignores a UFT Byte Order Mark at
|
Input adapter for a (caching) istream. Does not skip a UTF Byte Order Mark
|
||||||
beginning of input. Does not support changing the underlying std::streambuf
|
itself; that is done by the lexer's skip_bom(). Does not support changing
|
||||||
|
the underlying std::streambuf
|
||||||
in mid-input. Maintains underlying std::istream and std::streambuf to support
|
in mid-input. Maintains underlying std::istream and std::streambuf to support
|
||||||
subsequent use of standard std::istream operations to process any input
|
subsequent use of standard std::istream operations to process any input
|
||||||
characters following those used in parsing the JSON input. Clears the
|
characters following those used in parsing the JSON input. Clears the
|
||||||
@@ -8041,32 +8111,14 @@ struct wide_string_input_helper<BaseInputAdapter, 4>
|
|||||||
// get the current character
|
// get the current character
|
||||||
const auto wc = input.get_character();
|
const auto wc = input.get_character();
|
||||||
|
|
||||||
// UTF-32 to UTF-8 encoding
|
if (wc <= 0x10FFFF)
|
||||||
if (wc < 0x80)
|
|
||||||
{
|
{
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
|
// UTF-32 to UTF-8 encoding
|
||||||
utf8_bytes_filled = 1;
|
utf8_bytes_filled = 0;
|
||||||
}
|
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
||||||
else if (wc <= 0x7FF)
|
{
|
||||||
{
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xC0u | ((static_cast<unsigned int>(wc) >> 6u) & 0x1Fu));
|
});
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 2;
|
|
||||||
}
|
|
||||||
else if (wc <= 0xFFFF)
|
|
||||||
{
|
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xE0u | ((static_cast<unsigned int>(wc) >> 12u) & 0x0Fu));
|
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
|
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 3;
|
|
||||||
}
|
|
||||||
else if (wc <= 0x10FFFF)
|
|
||||||
{
|
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | ((static_cast<unsigned int>(wc) >> 18u) & 0x07u));
|
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 12u) & 0x3Fu));
|
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
|
|
||||||
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 4;
|
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
@@ -8103,24 +8155,15 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
|
|||||||
// get the current character
|
// get the current character
|
||||||
const auto wc = input.get_character();
|
const auto wc = input.get_character();
|
||||||
|
|
||||||
// UTF-16 to UTF-8 encoding
|
if (0xD800 > wc || wc >= 0xE000)
|
||||||
if (wc < 0x80)
|
|
||||||
{
|
{
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
|
// a UTF-16 code unit outside the surrogate range is a valid
|
||||||
utf8_bytes_filled = 1;
|
// code point (at most U+FFFF) on its own
|
||||||
}
|
utf8_bytes_filled = 0;
|
||||||
else if (wc <= 0x7FF)
|
encode_utf8(static_cast<std::uint32_t>(wc), [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
||||||
{
|
{
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xC0u | ((static_cast<unsigned int>(wc) >> 6u)));
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
});
|
||||||
utf8_bytes_filled = 2;
|
|
||||||
}
|
|
||||||
else if (0xD800 > wc || wc >= 0xE000)
|
|
||||||
{
|
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xE0u | ((static_cast<unsigned int>(wc) >> 12u)));
|
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((static_cast<unsigned int>(wc) >> 6u) & 0x3Fu));
|
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | (static_cast<unsigned int>(wc) & 0x3Fu));
|
|
||||||
utf8_bytes_filled = 3;
|
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
@@ -8138,11 +8181,11 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
|
|||||||
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
|
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
|
||||||
{
|
{
|
||||||
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
|
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
|
||||||
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
|
utf8_bytes_filled = 0;
|
||||||
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
|
encode_utf8(charcode, [&utf8_bytes, &utf8_bytes_filled](std::uint32_t byte)
|
||||||
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
|
{
|
||||||
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
|
utf8_bytes[utf8_bytes_filled++] = static_cast<std::char_traits<char>::int_type>(byte);
|
||||||
utf8_bytes_filled = 4;
|
});
|
||||||
valid_pair = true;
|
valid_pair = true;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -8471,9 +8514,9 @@ auto input_adapter(T (&array)[N]) -> decltype(input_adapter(array, array + N)) /
|
|||||||
return input_adapter(array, array + N);
|
return input_adapter(array, array + N);
|
||||||
}
|
}
|
||||||
|
|
||||||
// This class only handles inputs of input_buffer_adapter type.
|
// This class only handles inputs that construct a contiguous_bytes_input_adapter
|
||||||
// It's required so that expressions like {ptr, len} can be implicitly cast
|
// (e.g. span_input_adapter). It's required so that expressions like {ptr, len}
|
||||||
// to the correct adapter.
|
// can be implicitly cast to the correct adapter.
|
||||||
class span_input_adapter
|
class span_input_adapter
|
||||||
{
|
{
|
||||||
public:
|
public:
|
||||||
@@ -8518,6 +8561,7 @@ NLOHMANN_JSON_NAMESPACE_END
|
|||||||
|
|
||||||
#include <algorithm> // find_if, min
|
#include <algorithm> // find_if, min
|
||||||
#include <cstddef>
|
#include <cstddef>
|
||||||
|
#include <limits> // numeric_limits
|
||||||
#include <string> // string
|
#include <string> // string
|
||||||
#include <type_traits> // enable_if_t
|
#include <type_traits> // enable_if_t
|
||||||
#include <utility> // move, pair
|
#include <utility> // move, pair
|
||||||
@@ -8538,6 +8582,7 @@ NLOHMANN_JSON_NAMESPACE_END
|
|||||||
|
|
||||||
#include <array> // array
|
#include <array> // array
|
||||||
#include <cstddef> // size_t
|
#include <cstddef> // size_t
|
||||||
|
#include <cstdint> // uint32_t
|
||||||
#include <cstdio> // snprintf
|
#include <cstdio> // snprintf
|
||||||
#include <initializer_list> // initializer_list
|
#include <initializer_list> // initializer_list
|
||||||
#include <string> // char_traits, string
|
#include <string> // char_traits, string
|
||||||
@@ -10053,6 +10098,8 @@ NLOHMANN_JSON_NAMESPACE_END
|
|||||||
|
|
||||||
// #include <nlohmann/detail/meta/type_traits.hpp>
|
// #include <nlohmann/detail/meta/type_traits.hpp>
|
||||||
|
|
||||||
|
// #include <nlohmann/detail/string_utils.hpp>
|
||||||
|
|
||||||
|
|
||||||
NLOHMANN_JSON_NAMESPACE_BEGIN
|
NLOHMANN_JSON_NAMESPACE_BEGIN
|
||||||
namespace detail
|
namespace detail
|
||||||
@@ -10517,32 +10564,10 @@ class lexer : public lexer_base<BasicJsonType>
|
|||||||
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
|
JSON_ASSERT(0x00 <= codepoint && codepoint <= 0x10FFFF);
|
||||||
|
|
||||||
// translate codepoint into bytes
|
// translate codepoint into bytes
|
||||||
if (codepoint < 0x80)
|
encode_utf8(static_cast<std::uint32_t>(codepoint), [this](std::uint32_t byte)
|
||||||
{
|
{
|
||||||
// 1-byte characters: 0xxxxxxx (ASCII)
|
add(static_cast<char_int_type>(byte));
|
||||||
add(static_cast<char_int_type>(codepoint));
|
});
|
||||||
}
|
|
||||||
else if (codepoint <= 0x7FF)
|
|
||||||
{
|
|
||||||
// 2-byte characters: 110xxxxx 10xxxxxx
|
|
||||||
add(static_cast<char_int_type>(0xC0u | (static_cast<unsigned int>(codepoint) >> 6u)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
|
|
||||||
}
|
|
||||||
else if (codepoint <= 0xFFFF)
|
|
||||||
{
|
|
||||||
// 3-byte characters: 1110xxxx 10xxxxxx 10xxxxxx
|
|
||||||
add(static_cast<char_int_type>(0xE0u | (static_cast<unsigned int>(codepoint) >> 12u)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 6u) & 0x3Fu)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
// 4-byte characters: 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
|
|
||||||
add(static_cast<char_int_type>(0xF0u | (static_cast<unsigned int>(codepoint) >> 18u)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 12u) & 0x3Fu)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | ((static_cast<unsigned int>(codepoint) >> 6u) & 0x3Fu)));
|
|
||||||
add(static_cast<char_int_type>(0x80u | (static_cast<unsigned int>(codepoint) & 0x3Fu)));
|
|
||||||
}
|
|
||||||
|
|
||||||
break;
|
break;
|
||||||
}
|
}
|
||||||
@@ -11440,45 +11465,30 @@ scan_number_done:
|
|||||||
*/
|
*/
|
||||||
token_type convert_number(token_type number_type, std::size_t mantissa_end)
|
token_type convert_number(token_type number_type, std::size_t mantissa_end)
|
||||||
{
|
{
|
||||||
// If the caller does not need the converted value (only whether the
|
// accept() only needs to know whether the input is valid, so it sets
|
||||||
// input is syntactically valid; see json_sax_acceptor/accept()), an
|
// discard_number_values (see json.hpp), and an integer token whose
|
||||||
// unsigned/integer token can be reported without calling
|
// digit count shows that it fits is reported without calling
|
||||||
// strtoull()/strtoll() at all, *provided* we can already tell from
|
// convert_integer(). A number with up to 18 digits always fits into
|
||||||
// the digit count alone that the conversion cannot overflow 64 bits.
|
// both std::uint64_t and std::int64_t (18 nines is about 1e18, below
|
||||||
// Such tokens are always finite and are accepted unconditionally by
|
// INT64_MAX, which is about 9.2e18). Longer tokens take the exact path
|
||||||
// the parser regardless of their actual value (parser::sax_parse_internal()
|
// below, including the fallback to floating point when the value does
|
||||||
// never checks finiteness for value_unsigned/value_integer), so the
|
// not fit.
|
||||||
// classification below is all that is needed.
|
|
||||||
//
|
//
|
||||||
// A decimal number with up to 18 digits is always representable in
|
// With a narrower number_unsigned_t/number_integer_t (e.g.
|
||||||
// both std::uint64_t and std::int64_t (18 nines is ~1e18, well below
|
// std::uint32_t), the exact path would reclassify some of these tokens
|
||||||
// both UINT64_MAX ~1.8e19 and INT64_MAX ~9.2e18), so strtoull()/strtoll()
|
// as (finite) floats, while this check reports integers. That does not
|
||||||
// could not have set errno to ERANGE for it. Numbers with more digits
|
// change the result of accept(): it always parses through
|
||||||
// (rare in practice) fall through to the exact code below, unchanged,
|
// json_sax_acceptor, whose number callbacks discard their argument and
|
||||||
// so their handling -- including reclassification to value_float when
|
// return true, and the parser rejects neither integers nor finite
|
||||||
// the value overflows 64 bits, and rejection when it is not even
|
// floats. value_unsigned/value_integer are left unset here, so a caller
|
||||||
// finite as a double -- is bit-for-bit identical to before this
|
// that reads the converted value must not set discard_number_values.
|
||||||
// optimization.
|
|
||||||
//
|
//
|
||||||
// Note this reasons about std::uint64_t/std::int64_t, not about
|
// On contiguous input, scan_number_bulk_contiguous() converts integer
|
||||||
// number_unsigned_t/number_integer_t (BasicJsonType's own, possibly
|
// tokens itself and does not pass them to this function, unless
|
||||||
// narrower, template parameters -- e.g. std::uint32_t). That is fine
|
// JSON_DIAGNOSTIC_POSITIONS is enabled. This check is therefore only
|
||||||
// *only* because discard_number_values is exclusively set by
|
// reached for input without bulk access (e.g. streams), with
|
||||||
// accept() (see json.hpp), and accept() always parses through the
|
// JSON_DIAGNOSTIC_POSITIONS, or when scan_number_bulk_contiguous()
|
||||||
// library's own json_sax_acceptor -- never a user-supplied SAX
|
// falls back to scan_number().
|
||||||
// consumer -- whose number_unsigned()/number_integer()/number_float()
|
|
||||||
// callbacks unconditionally discard their argument and return true.
|
|
||||||
// So for every caller that can reach this branch, neither the token
|
|
||||||
// classification below nor the eventual (possibly narrowed, and on
|
|
||||||
// this fast path left stale/unset) value_unsigned/value_integer is
|
|
||||||
// ever consulted -- an unsigned/integer token is accepted outright,
|
|
||||||
// and even a >18-digit token that this fast path deliberately falls
|
|
||||||
// through for is, once reclassified to value_float, still finite
|
|
||||||
// (and thus accepted) for any digit count that fits in number_unsigned_t
|
|
||||||
// or number_integer_t regardless of that type's width. If this
|
|
||||||
// function is ever taught to run with discard_number_values true for
|
|
||||||
// a caller that *does* read the converted value, this reasoning (and
|
|
||||||
// the fast path below) would need to be revisited.
|
|
||||||
if (discard_number_values)
|
if (discard_number_values)
|
||||||
{
|
{
|
||||||
constexpr std::size_t safe_digit_count = 18;
|
constexpr std::size_t safe_digit_count = 18;
|
||||||
@@ -11907,7 +11917,7 @@ scan_number_done:
|
|||||||
return value_float;
|
return value_float;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// return current string value (implicitly resets the token; useful only once)
|
/// return current string value
|
||||||
string_t& get_string()
|
string_t& get_string()
|
||||||
{
|
{
|
||||||
// a number token holds '.' regardless of the locale (#4084)
|
// a number token holds '.' regardless of the locale (#4084)
|
||||||
@@ -12209,11 +12219,11 @@ scan_number_done:
|
|||||||
/// the position of the decimal point in token_buffer
|
/// the position of the decimal point in token_buffer
|
||||||
std::size_t decimal_point_position = std::string::npos;
|
std::size_t decimal_point_position = std::string::npos;
|
||||||
|
|
||||||
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
|
/// whether the caller only needs the token types and never looks at the
|
||||||
/// token classification and never looks at the converted numeric value;
|
/// converted numeric values; set only by accept(), which parses through
|
||||||
/// when set, scan_number() may skip strtoull()/strtoll() for
|
/// json_sax_acceptor. When set, convert_number() skips converting integer
|
||||||
/// value_unsigned/value_integer tokens whose digit count guarantees they
|
/// tokens whose digit count guarantees that they fit into 64 bits (see
|
||||||
/// fit into 64 bits (see scan_number())
|
/// there)
|
||||||
const bool discard_number_values = false;
|
const bool discard_number_values = false;
|
||||||
};
|
};
|
||||||
|
|
||||||
@@ -12381,6 +12391,88 @@ template<typename ArrayType>
|
|||||||
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
|
inline void reserve_array(ArrayType& /*arr*/, std::size_t /*len*/, priority_tag<0> /*unused*/)
|
||||||
{}
|
{}
|
||||||
|
|
||||||
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
|
/*!
|
||||||
|
@brief set the diagnostic positions of a value the DOM SAX parsers just stored
|
||||||
|
|
||||||
|
Shared by json_sax_dom_parser and json_sax_dom_callback_parser. basic_json
|
||||||
|
befriends this struct, as the position members are private.
|
||||||
|
*/
|
||||||
|
struct diagnostic_positions
|
||||||
|
{
|
||||||
|
/*!
|
||||||
|
@param[in,out] v the value that was just parsed
|
||||||
|
@param[in] lexer the lexer that read it, or nullptr to leave @a v alone
|
||||||
|
*/
|
||||||
|
template<typename BasicJsonType, typename LexerType>
|
||||||
|
static void set_from_lexer(BasicJsonType& v, LexerType* lexer)
|
||||||
|
{
|
||||||
|
if (lexer)
|
||||||
|
{
|
||||||
|
// Lexer has read past the current field value, so set the end position to the current position.
|
||||||
|
// The start position will be set below based on the length of the string representation
|
||||||
|
// of the value.
|
||||||
|
v.end_position = lexer->get_position();
|
||||||
|
|
||||||
|
switch (v.type())
|
||||||
|
{
|
||||||
|
case value_t::boolean:
|
||||||
|
{
|
||||||
|
// 4 and 5 are the string length of "true" and "false"
|
||||||
|
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::null:
|
||||||
|
{
|
||||||
|
// 4 is the string length of "null"
|
||||||
|
v.start_position = v.end_position - 4;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::string:
|
||||||
|
{
|
||||||
|
// escape sequences make the token longer than the value it
|
||||||
|
// parses to, so the start position cannot be derived from
|
||||||
|
// the value; use the offset the lexer recorded instead
|
||||||
|
v.start_position = lexer->get_token_start_position();
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::discarded:
|
||||||
|
{
|
||||||
|
// an object or array the callback of
|
||||||
|
// json_sax_dom_callback_parser rejected has no position
|
||||||
|
v.end_position = std::string::npos;
|
||||||
|
v.start_position = v.end_position;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::binary:
|
||||||
|
case value_t::number_integer:
|
||||||
|
case value_t::number_unsigned:
|
||||||
|
case value_t::number_float:
|
||||||
|
{
|
||||||
|
v.start_position = v.end_position - lexer->get_string().size();
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
case value_t::object:
|
||||||
|
case value_t::array:
|
||||||
|
{
|
||||||
|
// object and array are handled in start_object() and start_array() handlers
|
||||||
|
// skip setting the values here.
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
default: // LCOV_EXCL_LINE
|
||||||
|
// Handle all possible types discretely, default handler should never be reached.
|
||||||
|
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
};
|
||||||
|
#endif
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
@brief SAX implementation to create a JSON value from SAX events
|
@brief SAX implementation to create a JSON value from SAX events
|
||||||
|
|
||||||
@@ -12582,76 +12674,6 @@ class json_sax_dom_parser
|
|||||||
|
|
||||||
private:
|
private:
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
|
||||||
void handle_diagnostic_positions_for_json_value(BasicJsonType& v)
|
|
||||||
{
|
|
||||||
if (m_lexer_ref)
|
|
||||||
{
|
|
||||||
// Lexer has read past the current field value, so set the end position to the current position.
|
|
||||||
// The start position will be set below based on the length of the string representation
|
|
||||||
// of the value.
|
|
||||||
v.end_position = m_lexer_ref->get_position();
|
|
||||||
|
|
||||||
switch (v.type())
|
|
||||||
{
|
|
||||||
case value_t::boolean:
|
|
||||||
{
|
|
||||||
// 4 and 5 are the string length of "true" and "false"
|
|
||||||
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::null:
|
|
||||||
{
|
|
||||||
// 4 is the string length of "null"
|
|
||||||
v.start_position = v.end_position - 4;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::string:
|
|
||||||
{
|
|
||||||
// escape sequences make the token longer than the value it
|
|
||||||
// parses to, so the start position cannot be derived from
|
|
||||||
// the value; use the offset the lexer recorded instead
|
|
||||||
v.start_position = m_lexer_ref->get_token_start_position();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
// As we handle the start and end positions for values created during parsing,
|
|
||||||
// we do not expect the following value type to be called. Regardless, set the positions
|
|
||||||
// in case this is created manually or through a different constructor. Exclude from lcov
|
|
||||||
// since the exact condition of this switch is esoteric.
|
|
||||||
// LCOV_EXCL_START
|
|
||||||
case value_t::discarded:
|
|
||||||
{
|
|
||||||
v.end_position = std::string::npos;
|
|
||||||
v.start_position = v.end_position;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
// LCOV_EXCL_STOP
|
|
||||||
case value_t::binary:
|
|
||||||
case value_t::number_integer:
|
|
||||||
case value_t::number_unsigned:
|
|
||||||
case value_t::number_float:
|
|
||||||
{
|
|
||||||
v.start_position = v.end_position - m_lexer_ref->get_string().size();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
case value_t::object:
|
|
||||||
case value_t::array:
|
|
||||||
{
|
|
||||||
// object and array are handled in start_object() and start_array() handlers
|
|
||||||
// skip setting the values here.
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
default: // LCOV_EXCL_LINE
|
|
||||||
// Handle all possible types discretely, default handler should never be reached.
|
|
||||||
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert,-warnings-as-errors) LCOV_EXCL_LINE
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
#endif
|
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
@invariant If the ref stack is empty, then the passed value will be the new
|
@invariant If the ref stack is empty, then the passed value will be the new
|
||||||
root.
|
root.
|
||||||
@@ -12667,7 +12689,7 @@ class json_sax_dom_parser
|
|||||||
root = BasicJsonType(std::forward<Value>(v));
|
root = BasicJsonType(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(root);
|
diagnostic_positions::set_from_lexer(root, m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
return &root;
|
return &root;
|
||||||
@@ -12680,7 +12702,7 @@ class json_sax_dom_parser
|
|||||||
ref_stack.back()->m_data.m_value.array->emplace_back(std::forward<Value>(v));
|
ref_stack.back()->m_data.m_value.array->emplace_back(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(ref_stack.back()->m_data.m_value.array->back());
|
diagnostic_positions::set_from_lexer(ref_stack.back()->m_data.m_value.array->back(), m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
return &(ref_stack.back()->m_data.m_value.array->back());
|
return &(ref_stack.back()->m_data.m_value.array->back());
|
||||||
@@ -12691,7 +12713,7 @@ class json_sax_dom_parser
|
|||||||
*object_element = BasicJsonType(std::forward<Value>(v));
|
*object_element = BasicJsonType(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(*object_element);
|
diagnostic_positions::set_from_lexer(*object_element, m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
return object_element;
|
return object_element;
|
||||||
@@ -12880,7 +12902,7 @@ class json_sax_dom_callback_parser
|
|||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
// Set start/end positions for discarded object.
|
// Set start/end positions for discarded object.
|
||||||
handle_diagnostic_positions_for_json_value(*ref_stack.back());
|
diagnostic_positions::set_from_lexer(*ref_stack.back(), m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -12996,7 +13018,7 @@ class json_sax_dom_callback_parser
|
|||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
// Set start/end positions for discarded array.
|
// Set start/end positions for discarded array.
|
||||||
handle_diagnostic_positions_for_json_value(*ref_stack.back());
|
diagnostic_positions::set_from_lexer(*ref_stack.back(), m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -13049,72 +13071,6 @@ class json_sax_dom_callback_parser
|
|||||||
|
|
||||||
private:
|
private:
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
|
||||||
void handle_diagnostic_positions_for_json_value(BasicJsonType& v)
|
|
||||||
{
|
|
||||||
if (m_lexer_ref)
|
|
||||||
{
|
|
||||||
// Lexer has read past the current field value, so set the end position to the current position.
|
|
||||||
// The start position will be set below based on the length of the string representation
|
|
||||||
// of the value.
|
|
||||||
v.end_position = m_lexer_ref->get_position();
|
|
||||||
|
|
||||||
switch (v.type())
|
|
||||||
{
|
|
||||||
case value_t::boolean:
|
|
||||||
{
|
|
||||||
// 4 and 5 are the string length of "true" and "false"
|
|
||||||
v.start_position = v.end_position - (v.m_data.m_value.boolean ? 4 : 5);
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::null:
|
|
||||||
{
|
|
||||||
// 4 is the string length of "null"
|
|
||||||
v.start_position = v.end_position - 4;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::string:
|
|
||||||
{
|
|
||||||
// escape sequences make the token longer than the value it
|
|
||||||
// parses to, so the start position cannot be derived from
|
|
||||||
// the value; use the offset the lexer recorded instead
|
|
||||||
v.start_position = m_lexer_ref->get_token_start_position();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::discarded:
|
|
||||||
{
|
|
||||||
v.end_position = std::string::npos;
|
|
||||||
v.start_position = v.end_position;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::binary:
|
|
||||||
case value_t::number_integer:
|
|
||||||
case value_t::number_unsigned:
|
|
||||||
case value_t::number_float:
|
|
||||||
{
|
|
||||||
v.start_position = v.end_position - m_lexer_ref->get_string().size();
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
case value_t::object:
|
|
||||||
case value_t::array:
|
|
||||||
{
|
|
||||||
// object and array are handled in start_object() and start_array() handlers
|
|
||||||
// skip setting the values here.
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
default: // LCOV_EXCL_LINE
|
|
||||||
// Handle all possible types discretely, default handler should never be reached.
|
|
||||||
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert,-warnings-as-errors) LCOV_EXCL_LINE
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
#endif
|
|
||||||
|
|
||||||
/// if there is a pending duplicate-key stash entry for this exact slot,
|
/// if there is a pending duplicate-key stash entry for this exact slot,
|
||||||
/// remove it from the stash; if restore_value is true, the stashed
|
/// remove it from the stash; if restore_value is true, the stashed
|
||||||
/// previous value is moved back into the slot first (use this when the
|
/// previous value is moved back into the slot first (use this when the
|
||||||
@@ -13236,7 +13192,7 @@ class json_sax_dom_callback_parser
|
|||||||
auto value = BasicJsonType(std::forward<Value>(v));
|
auto value = BasicJsonType(std::forward<Value>(v));
|
||||||
|
|
||||||
#if JSON_DIAGNOSTIC_POSITIONS
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
handle_diagnostic_positions_for_json_value(value);
|
diagnostic_positions::set_from_lexer(value, m_lexer_ref);
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
// check callback
|
// check callback
|
||||||
@@ -16784,8 +16740,8 @@ class binary_reader
|
|||||||
return enter_object(detail::unknown_size());
|
return enter_object(detail::unknown_size());
|
||||||
}
|
}
|
||||||
|
|
||||||
// Note, no reader for UBJSON binary types is implemented because they do
|
// Note, UBJSON has no binary type of its own; BJData, which shares this
|
||||||
// not exist
|
// reader, decodes optimized 'B' arrays as binary in get_ubjson_array().
|
||||||
|
|
||||||
bool get_ubjson_high_precision_number()
|
bool get_ubjson_high_precision_number()
|
||||||
{
|
{
|
||||||
@@ -17491,7 +17447,7 @@ class binary_reader
|
|||||||
#endif
|
#endif
|
||||||
}
|
}
|
||||||
|
|
||||||
/*
|
/*!
|
||||||
@brief read a number from the input
|
@brief read a number from the input
|
||||||
|
|
||||||
@tparam NumberType the type of the number
|
@tparam NumberType the type of the number
|
||||||
@@ -17501,10 +17457,10 @@ class binary_reader
|
|||||||
@return whether conversion completed
|
@return whether conversion completed
|
||||||
|
|
||||||
@note This function needs to respect the system's endianness, because
|
@note This function needs to respect the system's endianness, because
|
||||||
bytes in CBOR, MessagePack, and UBJSON are stored in network order
|
bytes in CBOR, MessagePack, UBJSON, and BON8 are stored in network
|
||||||
(big endian) and therefore need reordering on little endian systems.
|
order (big endian) and therefore need reordering on little endian
|
||||||
On the other hand, BSON and BJData use little endian and should reorder
|
systems. On the other hand, BSON and BJData use little endian and
|
||||||
on big endian systems.
|
should reorder on big endian systems.
|
||||||
*/
|
*/
|
||||||
template<typename NumberType, bool InputIsLittleEndian = false>
|
template<typename NumberType, bool InputIsLittleEndian = false>
|
||||||
bool get_number(const input_format_t format, NumberType& result)
|
bool get_number(const input_format_t format, NumberType& result)
|
||||||
@@ -17945,7 +17901,8 @@ using parser_callback_t =
|
|||||||
/*!
|
/*!
|
||||||
@brief syntax analysis
|
@brief syntax analysis
|
||||||
|
|
||||||
This class implements a recursive descent parser.
|
This class implements an iterative parser that keeps the open containers on
|
||||||
|
an explicit stack and reports what it reads as SAX events.
|
||||||
*/
|
*/
|
||||||
template<typename BasicJsonType, typename InputAdapterType>
|
template<typename BasicJsonType, typename InputAdapterType>
|
||||||
class parser
|
class parser
|
||||||
@@ -17989,28 +17946,9 @@ class parser
|
|||||||
if (callback)
|
if (callback)
|
||||||
{
|
{
|
||||||
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
|
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
|
||||||
sax_parse_internal(&sdp);
|
|
||||||
|
|
||||||
if (strict)
|
|
||||||
{
|
|
||||||
// in strict mode, input must be completely read
|
|
||||||
if (get_token() != token_type::end_of_input)
|
|
||||||
{
|
|
||||||
sdp.parse_error(m_lexer.get_position(),
|
|
||||||
m_lexer.get_token_string(),
|
|
||||||
parse_error::create(101, m_lexer.get_position(),
|
|
||||||
exception_message(token_type::end_of_input, "value"), nullptr));
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
// the caller keeps using the input: position it right after
|
|
||||||
// the value by leaving the character that terminated it
|
|
||||||
m_lexer.release_lookahead();
|
|
||||||
}
|
|
||||||
|
|
||||||
// in case of an error, return a discarded value
|
// in case of an error, return a discarded value
|
||||||
if (sdp.is_errored())
|
if (!parse_dom(sdp, strict))
|
||||||
{
|
{
|
||||||
result = value_t::discarded;
|
result = value_t::discarded;
|
||||||
return;
|
return;
|
||||||
@@ -18026,26 +17964,9 @@ class parser
|
|||||||
else
|
else
|
||||||
{
|
{
|
||||||
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
|
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
|
||||||
sax_parse_internal(&sdp);
|
|
||||||
|
|
||||||
if (strict)
|
|
||||||
{
|
|
||||||
// in strict mode, input must be completely read
|
|
||||||
if (get_token() != token_type::end_of_input)
|
|
||||||
{
|
|
||||||
sdp.parse_error(m_lexer.get_position(),
|
|
||||||
m_lexer.get_token_string(),
|
|
||||||
parse_error::create(101, m_lexer.get_position(), exception_message(token_type::end_of_input, "value"), nullptr));
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
// see above
|
|
||||||
m_lexer.release_lookahead();
|
|
||||||
}
|
|
||||||
|
|
||||||
// in case of an error, return a discarded value
|
// in case of an error, return a discarded value
|
||||||
if (sdp.is_errored())
|
if (!parse_dom(sdp, strict))
|
||||||
{
|
{
|
||||||
result = value_t::discarded;
|
result = value_t::discarded;
|
||||||
return;
|
return;
|
||||||
@@ -18098,6 +18019,46 @@ class parser
|
|||||||
}
|
}
|
||||||
|
|
||||||
private:
|
private:
|
||||||
|
/*!
|
||||||
|
@brief run a DOM SAX parser to completion and position the lexer
|
||||||
|
|
||||||
|
Shared by both branches of @ref parse(): builds no SAX parser itself,
|
||||||
|
but drives an already-constructed @a json_sax_dom_parser or
|
||||||
|
@ref json_sax_dom_callback_parser through @ref sax_parse_internal(),
|
||||||
|
then applies the strict-EOF check (reporting parse_error.101 through
|
||||||
|
@a sdp on failure) or, in non-strict mode, releases the lookahead so
|
||||||
|
the caller can keep reading the input right after the parsed value.
|
||||||
|
|
||||||
|
@param[in,out] sdp the DOM SAX parser to run
|
||||||
|
@param[in] strict whether to expect the last token to be EOF
|
||||||
|
@return whether @a sdp did not report an error
|
||||||
|
*/
|
||||||
|
template<typename DomSax>
|
||||||
|
bool parse_dom(DomSax& sdp, const bool strict)
|
||||||
|
{
|
||||||
|
sax_parse_internal(&sdp);
|
||||||
|
|
||||||
|
if (strict)
|
||||||
|
{
|
||||||
|
// in strict mode, input must be completely read
|
||||||
|
if (get_token() != token_type::end_of_input)
|
||||||
|
{
|
||||||
|
sdp.parse_error(m_lexer.get_position(),
|
||||||
|
m_lexer.get_token_string(),
|
||||||
|
parse_error::create(101, m_lexer.get_position(),
|
||||||
|
exception_message(token_type::end_of_input, "value"), nullptr));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
// the caller keeps using the input: position it right after
|
||||||
|
// the value by leaving the character that terminated it
|
||||||
|
m_lexer.release_lookahead();
|
||||||
|
}
|
||||||
|
|
||||||
|
return !sdp.is_errored();
|
||||||
|
}
|
||||||
|
|
||||||
template<typename SAX>
|
template<typename SAX>
|
||||||
JSON_HEDLEY_NON_NULL(2)
|
JSON_HEDLEY_NON_NULL(2)
|
||||||
bool sax_parse_internal(SAX* sax)
|
bool sax_parse_internal(SAX* sax)
|
||||||
@@ -18330,8 +18291,9 @@ class parser
|
|||||||
|
|
||||||
// We are done with this array. Before we can parse a
|
// We are done with this array. Before we can parse a
|
||||||
// new value, we need to evaluate the new state first.
|
// new value, we need to evaluate the new state first.
|
||||||
// By setting skip_to_state_evaluation to false, we
|
// By setting skip_to_state_evaluation to true, the next
|
||||||
// are effectively jumping to the beginning of this if.
|
// iteration skips parsing a value and evaluates the
|
||||||
|
// enclosing state directly.
|
||||||
JSON_ASSERT(!states.empty());
|
JSON_ASSERT(!states.empty());
|
||||||
states.pop_back();
|
states.pop_back();
|
||||||
skip_to_state_evaluation = true;
|
skip_to_state_evaluation = true;
|
||||||
@@ -18391,8 +18353,9 @@ class parser
|
|||||||
|
|
||||||
// We are done with this object. Before we can parse a
|
// We are done with this object. Before we can parse a
|
||||||
// new value, we need to evaluate the new state first.
|
// new value, we need to evaluate the new state first.
|
||||||
// By setting skip_to_state_evaluation to false, we
|
// By setting skip_to_state_evaluation to true, the next
|
||||||
// are effectively jumping to the beginning of this if.
|
// iteration skips parsing a value and evaluates the
|
||||||
|
// enclosing state directly.
|
||||||
JSON_ASSERT(!states.empty());
|
JSON_ASSERT(!states.empty());
|
||||||
states.pop_back();
|
states.pop_back();
|
||||||
skip_to_state_evaluation = true;
|
skip_to_state_evaluation = true;
|
||||||
@@ -26806,6 +26769,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
|
|||||||
friend class ::nlohmann::detail::json_sax_dom_parser;
|
friend class ::nlohmann::detail::json_sax_dom_parser;
|
||||||
template<typename BasicJsonType, typename InputAdapterType>
|
template<typename BasicJsonType, typename InputAdapterType>
|
||||||
friend class ::nlohmann::detail::json_sax_dom_callback_parser;
|
friend class ::nlohmann::detail::json_sax_dom_callback_parser;
|
||||||
|
#if JSON_DIAGNOSTIC_POSITIONS
|
||||||
|
friend struct ::nlohmann::detail::diagnostic_positions;
|
||||||
|
#endif
|
||||||
friend class ::nlohmann::detail::exception;
|
friend class ::nlohmann::detail::exception;
|
||||||
|
|
||||||
/// workaround type for MSVC
|
/// workaround type for MSVC
|
||||||
|
|||||||
Reference in New Issue
Block a user