mirror of
https://github.com/nlohmann/json.git
synced 2026-08-08 18:23:18 +00:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a13902a33f | ||
|
|
c021a09b08 | ||
|
|
e4aaf46d38 | ||
|
|
5bc24e876b | ||
|
|
da7b9bdb3d | ||
|
|
634f49bc5b |
@@ -69,7 +69,8 @@ The SAX event lister must follow the interface of [`json_sax`](../json_sax/index
|
|||||||
[`input_format_t`](input_format_t.md) for more information
|
[`input_format_t`](input_format_t.md) for more information
|
||||||
|
|
||||||
`strict` (in)
|
`strict` (in)
|
||||||
: whether the input has to be consumed completely (optional, `#!cpp true` by default)
|
: whether the input has to be consumed completely (optional, `#!cpp true` by default); when `#!cpp false` and the
|
||||||
|
input is a `#!cpp std::istream`, the stream is left positioned right after the parsed value
|
||||||
|
|
||||||
`ignore_comments` (in)
|
`ignore_comments` (in)
|
||||||
: whether comments should be ignored and treated like whitespace (`#!cpp true`) or yield a parse error
|
: whether comments should be ignored and treated like whitespace (`#!cpp true`) or yield a parse error
|
||||||
@@ -136,6 +137,8 @@ A UTF-8 byte order mark is silently ignored.
|
|||||||
- Added `ignore_trailing_commas` in version 3.13.0.
|
- Added `ignore_trailing_commas` in version 3.13.0.
|
||||||
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
|
- Extended container support (1) to include types with lvalue-only ADL `begin`/`end` (matching `std::begin`/`std::end` semantics) in version 3.13.0.
|
||||||
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
|
- Extended overload (2) to accept heterogeneous iterator+sentinel pairs (C++20 ranges support) in version 3.13.0.
|
||||||
|
- Changed in version 4.0.0 to leave a `#!cpp std::istream` positioned right after the parsed value when `strict` is
|
||||||
|
`#!cpp false`; see [`operator>>`](../operator_gtgt.md#notes).
|
||||||
|
|
||||||
!!! warning "Deprecation"
|
!!! warning "Deprecation"
|
||||||
|
|
||||||
|
|||||||
@@ -24,7 +24,6 @@ header. See also the [macro overview page](../../features/macros.md).
|
|||||||
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
|
- [**JSON_NO_IO**](json_no_io.md) - switch off functions relying on certain C++ I/O headers
|
||||||
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
|
- [**JSON_SKIP_UNSUPPORTED_COMPILER_CHECK**](json_skip_unsupported_compiler_check.md) - do not warn about unsupported compilers
|
||||||
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
|
- [**JSON_USE_GLOBAL_UDLS**](json_use_global_udls.md) - place user-defined string literals (UDLs) into the global namespace
|
||||||
- [**JSON_USE_SIMDUTF**](json_use_simdutf.md) - use the simdutf library to accelerate UTF-8 validation
|
|
||||||
|
|
||||||
## Library version
|
## Library version
|
||||||
|
|
||||||
|
|||||||
@@ -1,52 +0,0 @@
|
|||||||
# JSON_USE_SIMDUTF
|
|
||||||
|
|
||||||
```cpp
|
|
||||||
#define JSON_USE_SIMDUTF
|
|
||||||
```
|
|
||||||
|
|
||||||
When defined, the parser validates the UTF-8 content of JSON strings that come from a **contiguous byte input**
|
|
||||||
(`std::string`, `std::vector<char>`/`<std::uint8_t>`, string literals, `const char*` ranges, …) using the
|
|
||||||
[simdutf](https://github.com/simdutf/simdutf) library instead of the built-in scalar validator. On text with many
|
|
||||||
non-ASCII characters (e.g. CJK or emoji) this can validate several times faster.
|
|
||||||
|
|
||||||
This is an **opt-in external dependency**. The library itself remains header-only and its behavior is unchanged: the
|
|
||||||
same input is accepted or rejected either way, and every parse error is reported at the same position with the same
|
|
||||||
message (simdutf is only used to fast-path *valid* runs; anything it flags falls back to the scalar path so the exact
|
|
||||||
diagnostic is preserved). Streaming inputs (files, `std::istream`, wide strings, user-defined adapters) always use the
|
|
||||||
scalar path.
|
|
||||||
|
|
||||||
When `JSON_USE_SIMDUTF` is defined you must make the `simdutf.h` header available on the include path and link the
|
|
||||||
simdutf library. When it is not defined, no simdutf header is included and there is no dependency.
|
|
||||||
|
|
||||||
## Default definition
|
|
||||||
|
|
||||||
By default, `#!cpp JSON_USE_SIMDUTF` is not defined and the portable C++11 scalar validator is used.
|
|
||||||
|
|
||||||
```cpp
|
|
||||||
#undef JSON_USE_SIMDUTF
|
|
||||||
```
|
|
||||||
|
|
||||||
## Examples
|
|
||||||
|
|
||||||
??? example
|
|
||||||
|
|
||||||
The code below enables the simdutf backend for UTF-8 validation.
|
|
||||||
|
|
||||||
```cpp
|
|
||||||
#define JSON_USE_SIMDUTF 1
|
|
||||||
#include <simdutf.h>
|
|
||||||
#include <nlohmann/json.hpp>
|
|
||||||
|
|
||||||
...
|
|
||||||
```
|
|
||||||
|
|
||||||
The project must also link against simdutf, e.g. with CMake:
|
|
||||||
|
|
||||||
```cmake
|
|
||||||
target_compile_definitions(your_target PRIVATE JSON_USE_SIMDUTF)
|
|
||||||
target_link_libraries(your_target PRIVATE simdutf::simdutf)
|
|
||||||
```
|
|
||||||
|
|
||||||
## Version history
|
|
||||||
|
|
||||||
- Added in version 3.12.1.
|
|
||||||
@@ -33,41 +33,26 @@ A UTF-8 byte order mark is silently ignored.
|
|||||||
Invalid Unicode escapes and unpaired surrogates in the input are reported as
|
Invalid Unicode escapes and unpaired surrogates in the input are reported as
|
||||||
[`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101) with a detailed message.
|
[`parse_error.101`](../home/exceptions.md#jsonexceptionparse_error101) with a detailed message.
|
||||||
|
|
||||||
`operator>>` parses exactly one JSON value, so it can be called repeatedly to read a sequence of concatenated JSON
|
`operator>>` parses exactly one JSON value and leaves the stream positioned right after it, so it can be called
|
||||||
values from the same stream:
|
repeatedly to read a sequence of concatenated JSON values from the same stream:
|
||||||
|
|
||||||
```cpp
|
```cpp
|
||||||
json j1, j2;
|
std::istringstream input("1true[2]");
|
||||||
input >> j1; // parses the first value
|
json j1, j2, j3;
|
||||||
input >> j2; // parses the next value
|
input >> j1; // j1 == 1, stream now positioned right after it
|
||||||
|
input >> j2; // j2 == true
|
||||||
|
input >> j3; // j3 == [2]
|
||||||
```
|
```
|
||||||
|
|
||||||
!!! warning "A number must be followed by whitespace"
|
!!! note "Changed behavior for numbers"
|
||||||
|
|
||||||
A number is only terminated by the character that follows it. That character is read from the stream to detect the
|
A number is the only value whose end can be detected solely by reading the character that follows it. Up to
|
||||||
end of the number, and it is **not** put back. When a value that is a number is immediately followed by the next
|
version 3.13.0 that character was consumed and not put back, so the stream was left one byte too far whenever a
|
||||||
value, the first character of that next value is lost:
|
number was immediately followed by another value: reading `1true` yielded `1` and left the stream at `rue`.
|
||||||
|
Values had to be separated by whitespace to work around this.
|
||||||
|
|
||||||
```cpp
|
The terminating character is now only looked at and left in the stream, so no separator is required. Code that
|
||||||
std::istringstream input("1true");
|
relied on the extra byte being swallowed will observe it again.
|
||||||
json j1, j2;
|
|
||||||
input >> j1; // j1 == 1
|
|
||||||
input >> j2; // throws parse_error.101: the stream now starts at "rue"
|
|
||||||
```
|
|
||||||
|
|
||||||
Separating the values with whitespace avoids this, because the character that is eaten is then the separator:
|
|
||||||
|
|
||||||
```cpp
|
|
||||||
std::istringstream input("1 true");
|
|
||||||
json j1, j2;
|
|
||||||
input >> j1; // j1 == 1
|
|
||||||
input >> j2; // j2 == true
|
|
||||||
```
|
|
||||||
|
|
||||||
Only numbers are affected. Values ending in a self-delimiting character do not read past themselves, so
|
|
||||||
`truefalse`, `[1][2]`, `{"a":1}{"b":2}`, and `"a""b"` can be read back to back without a separator.
|
|
||||||
|
|
||||||
This is tracked in [#5340](https://github.com/nlohmann/json/issues/5340).
|
|
||||||
|
|
||||||
Note that reading concatenated values does **not** work for [JSON Lines](../features/parsing/json_lines.md)
|
Note that reading concatenated values does **not** work for [JSON Lines](../features/parsing/json_lines.md)
|
||||||
(newline-delimited JSON) input -- see that page for why and for the recommended alternative.
|
(newline-delimited JSON) input -- see that page for why and for the recommended alternative.
|
||||||
@@ -102,3 +87,5 @@ Note that reading concatenated values does **not** work for [JSON Lines](../feat
|
|||||||
## Version history
|
## Version history
|
||||||
|
|
||||||
- Added in version 1.0.0.
|
- Added in version 1.0.0.
|
||||||
|
- Changed in version 4.0.0 to leave the character that terminates a number in the stream, so that the stream is
|
||||||
|
positioned right after the parsed value for every value type.
|
||||||
|
|||||||
@@ -296,7 +296,6 @@ nav:
|
|||||||
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
|
- 'JSON_USE_GLOBAL_UDLS': api/macros/json_use_global_udls.md
|
||||||
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
|
- 'JSON_USE_IMPLICIT_CONVERSIONS': api/macros/json_use_implicit_conversions.md
|
||||||
- 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md
|
- 'JSON_USE_LEGACY_DISCARDED_VALUE_COMPARISON': api/macros/json_use_legacy_discarded_value_comparison.md
|
||||||
- 'JSON_USE_SIMDUTF': api/macros/json_use_simdutf.md
|
|
||||||
- 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md
|
- 'NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_INTRUSIVE_ONLY_SERIALIZE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_DERIVED_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_derived_type.md
|
||||||
- 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md
|
- 'NLOHMANN_DEFINE_TYPE_INTRUSIVE, NLOHMANN_DEFINE_TYPE_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_intrusive.md
|
||||||
- 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md
|
- 'NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_ONLY_SERIALIZE': api/macros/nlohmann_define_type_non_intrusive.md
|
||||||
|
|||||||
@@ -101,6 +101,9 @@ class input_stream_adapter
|
|||||||
// maintain ifstream flags, except eof
|
// maintain ifstream flags, except eof
|
||||||
if (is != nullptr)
|
if (is != nullptr)
|
||||||
{
|
{
|
||||||
|
// consume the character last returned by get_character() unless it
|
||||||
|
// was given back with release_lookahead()
|
||||||
|
commit_lookahead();
|
||||||
is->clear(is->rdstate() & std::ios::eofbit);
|
is->clear(is->rdstate() & std::ios::eofbit);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -115,29 +118,60 @@ class input_stream_adapter
|
|||||||
input_stream_adapter& operator=(input_stream_adapter&&) = delete;
|
input_stream_adapter& operator=(input_stream_adapter&&) = delete;
|
||||||
|
|
||||||
input_stream_adapter(input_stream_adapter&& rhs) noexcept
|
input_stream_adapter(input_stream_adapter&& rhs) noexcept
|
||||||
: is(rhs.is), sb(rhs.sb)
|
: is(rhs.is), sb(rhs.sb), lookahead(rhs.lookahead)
|
||||||
{
|
{
|
||||||
rhs.is = nullptr;
|
rhs.is = nullptr;
|
||||||
rhs.sb = nullptr;
|
rhs.sb = nullptr;
|
||||||
|
rhs.lookahead = false;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Whether the character last returned by get_character() can be given back
|
||||||
|
// to the input with release_lookahead().
|
||||||
|
static constexpr bool supports_lookahead = true;
|
||||||
|
|
||||||
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
|
// std::istream/std::streambuf use std::char_traits<char>::to_int_type, to
|
||||||
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
|
// ensure that std::char_traits<char>::eof() and the character 0xFF do not
|
||||||
// end up as the same value, e.g., 0xFFFFFFFF.
|
// end up as the same value, e.g., 0xFFFFFFFF.
|
||||||
|
//
|
||||||
|
// The character is peeked rather than consumed: it is only stepped over
|
||||||
|
// once the next character is requested, or when the adapter is destroyed.
|
||||||
|
// Until then, release_lookahead() can leave it in the input.
|
||||||
std::char_traits<char>::int_type get_character()
|
std::char_traits<char>::int_type get_character()
|
||||||
{
|
{
|
||||||
auto res = sb->sbumpc();
|
if (lookahead)
|
||||||
|
{
|
||||||
|
// step over the character returned by the previous call
|
||||||
|
sb->sbumpc();
|
||||||
|
}
|
||||||
|
|
||||||
|
auto res = sb->sgetc();
|
||||||
// set eof manually, as we don't use the istream interface.
|
// set eof manually, as we don't use the istream interface.
|
||||||
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
|
if (JSON_HEDLEY_UNLIKELY(res == std::char_traits<char>::eof()))
|
||||||
{
|
{
|
||||||
|
// there is nothing to step over next time
|
||||||
|
lookahead = false;
|
||||||
is->clear(is->rdstate() | std::ios::eofbit);
|
is->clear(is->rdstate() | std::ios::eofbit);
|
||||||
}
|
}
|
||||||
|
else
|
||||||
|
{
|
||||||
|
lookahead = true;
|
||||||
|
}
|
||||||
return res;
|
return res;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Leave the character last returned by get_character() in the input, so
|
||||||
|
// that the next read from the stream - by this adapter or by the caller
|
||||||
|
// once parsing is done - sees it again. Unlike putting a consumed
|
||||||
|
// character back, this cannot fail.
|
||||||
|
void release_lookahead() noexcept
|
||||||
|
{
|
||||||
|
lookahead = false;
|
||||||
|
}
|
||||||
|
|
||||||
template<class T>
|
template<class T>
|
||||||
std::size_t get_elements(T* dest, std::size_t count = 1)
|
std::size_t get_elements(T* dest, std::size_t count = 1)
|
||||||
{
|
{
|
||||||
|
commit_lookahead();
|
||||||
auto res = static_cast<std::size_t>(sb->sgetn(reinterpret_cast<char*>(dest), static_cast<std::streamsize>(count * sizeof(T))));
|
auto res = static_cast<std::size_t>(sb->sgetn(reinterpret_cast<char*>(dest), static_cast<std::streamsize>(count * sizeof(T))));
|
||||||
if (JSON_HEDLEY_UNLIKELY(res < count * sizeof(T)))
|
if (JSON_HEDLEY_UNLIKELY(res < count * sizeof(T)))
|
||||||
{
|
{
|
||||||
@@ -147,9 +181,23 @@ class input_stream_adapter
|
|||||||
}
|
}
|
||||||
|
|
||||||
private:
|
private:
|
||||||
|
// Step over the character last returned by get_character(). The character
|
||||||
|
// has already been peeked successfully, so for every streambuf with a get
|
||||||
|
// area this is a pointer increment that cannot fail.
|
||||||
|
void commit_lookahead()
|
||||||
|
{
|
||||||
|
if (lookahead)
|
||||||
|
{
|
||||||
|
lookahead = false;
|
||||||
|
sb->sbumpc();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// the associated input stream
|
/// the associated input stream
|
||||||
std::istream* is = nullptr;
|
std::istream* is = nullptr;
|
||||||
std::streambuf* sb = nullptr;
|
std::streambuf* sb = nullptr;
|
||||||
|
/// whether get_character() peeked a character that is not consumed yet
|
||||||
|
bool lookahead = false;
|
||||||
};
|
};
|
||||||
#endif // JSON_NO_IO
|
#endif // JSON_NO_IO
|
||||||
|
|
||||||
@@ -231,33 +279,6 @@ class iterator_input_adapter
|
|||||||
std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
|
std::is_same<IteratorType, SentinelType>::value && std::is_pointer<IteratorType>::value;
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
public:
|
|
||||||
// Whether the remaining input is a single contiguous block of 1-byte
|
|
||||||
// elements that the lexer can inspect directly (used for the SWAR string
|
|
||||||
// fast path). Restricted to same-type iterator/sentinel pairs so that plain
|
|
||||||
// std::distance/std::advance are well-defined in all standards.
|
|
||||||
static constexpr bool supports_bulk_scan =
|
|
||||||
iterator_is_contiguous && std::is_same<IteratorType, SentinelType>::value && sizeof(char_type) == 1;
|
|
||||||
|
|
||||||
// Pointer to the next unread element; only valid when bulk_remaining() > 0.
|
|
||||||
const char_type* bulk_data() const
|
|
||||||
{
|
|
||||||
return &*current;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Number of unread elements available as one contiguous block.
|
|
||||||
std::size_t bulk_remaining() const
|
|
||||||
{
|
|
||||||
return static_cast<std::size_t>(std::distance(current, end));
|
|
||||||
}
|
|
||||||
|
|
||||||
// Consume @a n elements previously inspected via bulk_data().
|
|
||||||
void bulk_skip(std::size_t n)
|
|
||||||
{
|
|
||||||
std::advance(current, static_cast<typename std::iterator_traits<IteratorType>::difference_type>(n));
|
|
||||||
}
|
|
||||||
|
|
||||||
private:
|
|
||||||
// contiguous fast path: bulk copy the remaining range with std::memcpy
|
// contiguous fast path: bulk copy the remaining range with std::memcpy
|
||||||
template<class T>
|
template<class T>
|
||||||
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
|
std::size_t get_elements_impl(T* dest, std::size_t count, std::true_type /*contiguous*/)
|
||||||
@@ -593,24 +614,6 @@ typename iterator_input_adapter_factory<IteratorType, SentinelType>::adapter_typ
|
|||||||
return factory_type::create(first, last);
|
return factory_type::create(first, last);
|
||||||
}
|
}
|
||||||
|
|
||||||
// Detect a container that stores its elements contiguously as single bytes
|
|
||||||
// (std::string, std::vector<char/unsigned char>, std::array<char, N>,
|
|
||||||
// std::string_view, ...). Such inputs are wrapped in a pointer-based adapter so
|
|
||||||
// they benefit from the contiguous fast paths (bulk string scanning, memcpy for
|
|
||||||
// binary formats) in every C++ standard - not only in C++20, where the standard
|
|
||||||
// library iterators model std::contiguous_iterator and are detected directly.
|
|
||||||
template<typename ContainerType, typename = void>
|
|
||||||
struct is_contiguous_byte_container : std::false_type {};
|
|
||||||
|
|
||||||
template<typename ContainerType>
|
|
||||||
struct is_contiguous_byte_container < ContainerType, void_t <
|
|
||||||
decltype(std::declval<const ContainerType&>().data()),
|
|
||||||
decltype(std::declval<const ContainerType&>().size()) >>
|
|
||||||
: std::integral_constant < bool,
|
|
||||||
std::is_pointer<decltype(std::declval<const ContainerType&>().data())>::value&&
|
|
||||||
std::is_integral<typename std::remove_pointer<decltype(std::declval<const ContainerType&>().data())>::type>::value&&
|
|
||||||
sizeof(typename std::remove_pointer<decltype(std::declval<const ContainerType&>().data())>::type) == 1 > {};
|
|
||||||
|
|
||||||
// Convenience shorthand from container to iterator
|
// Convenience shorthand from container to iterator
|
||||||
// Enables ADL on begin(container) and end(container)
|
// Enables ADL on begin(container) and end(container)
|
||||||
// Encloses the using declarations in namespace for not to leak them to outside scope
|
// Encloses the using declarations in namespace for not to leak them to outside scope
|
||||||
@@ -638,32 +641,12 @@ struct container_input_adapter_factory< ContainerType,
|
|||||||
|
|
||||||
} // namespace container_input_adapter_factory_impl
|
} // namespace container_input_adapter_factory_impl
|
||||||
|
|
||||||
// General container path (iterator-based). Contiguous single-byte containers
|
template<typename ContainerType>
|
||||||
// are excluded here and routed through the pointer-based overload below.
|
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType&& container)
|
||||||
template < typename ContainerType,
|
|
||||||
enable_if_t < !is_contiguous_byte_container<ContainerType>::value, int > = 0 >
|
|
||||||
typename container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::adapter_type input_adapter(ContainerType && container)
|
|
||||||
{
|
{
|
||||||
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
|
return container_input_adapter_factory_impl::container_input_adapter_factory<ContainerType>::create(std::forward<ContainerType>(container));
|
||||||
}
|
}
|
||||||
|
|
||||||
// Contiguous single-byte containers (std::string, std::vector<char>, ...) are
|
|
||||||
// wrapped in a pointer-based adapter so the contiguous fast paths apply in every
|
|
||||||
// standard. The pointer keeps the container's own element type (const char* for
|
|
||||||
// std::string, const std::uint8_t* for std::vector<std::uint8_t>, ...), so the
|
|
||||||
// resulting char_type - and therefore the parsing behavior - is byte-for-byte
|
|
||||||
// identical to the iterator-based path; only the raw pointer additionally
|
|
||||||
// enables the bulk fast paths. The container outlives the adapter for the whole
|
|
||||||
// parse (temporaries live until the end of the full expression), exactly as the
|
|
||||||
// iterators it replaces did.
|
|
||||||
template < typename ContainerType,
|
|
||||||
enable_if_t < is_contiguous_byte_container<ContainerType>::value, int > = 0 >
|
|
||||||
auto input_adapter(const ContainerType& container)
|
|
||||||
-> decltype(input_adapter(container.data(), container.data() + container.size()))
|
|
||||||
{
|
|
||||||
return input_adapter(container.data(), container.data() + container.size());
|
|
||||||
}
|
|
||||||
|
|
||||||
// specialization for std::string
|
// specialization for std::string
|
||||||
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
|
using string_input_adapter_type = decltype(input_adapter(std::declval<std::string>()));
|
||||||
|
|
||||||
|
|||||||
@@ -19,9 +19,7 @@
|
|||||||
#include <vector> // vector
|
#include <vector> // vector
|
||||||
|
|
||||||
#include <nlohmann/detail/input/input_adapters.hpp>
|
#include <nlohmann/detail/input/input_adapters.hpp>
|
||||||
#include <nlohmann/detail/input/number_parse.hpp>
|
|
||||||
#include <nlohmann/detail/input/position_t.hpp>
|
#include <nlohmann/detail/input/position_t.hpp>
|
||||||
#include <nlohmann/detail/input/string_scan.hpp>
|
|
||||||
#include <nlohmann/detail/macro_scope.hpp>
|
#include <nlohmann/detail/macro_scope.hpp>
|
||||||
#include <nlohmann/detail/meta/type_traits.hpp>
|
#include <nlohmann/detail/meta/type_traits.hpp>
|
||||||
|
|
||||||
@@ -127,21 +125,20 @@ constexpr bool input_adapter_supports_seek(std::false_type /*detected*/)
|
|||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Detect whether an input adapter exposes a contiguous byte block that the
|
// Detect whether an input adapter reads with one character of lookahead that
|
||||||
// lexer can scan directly (see iterator_input_adapter::supports_bulk_scan).
|
// can be left in the input (see input_stream_adapter::supports_lookahead),
|
||||||
// Adapters without the flag - file, stream, wide-string, user-defined - fall
|
// detected like supports_seek above.
|
||||||
// back to the character-at-a-time string scanner.
|
|
||||||
template<typename InputAdapterType>
|
template<typename InputAdapterType>
|
||||||
using detect_supports_bulk_scan = decltype(InputAdapterType::supports_bulk_scan);
|
using detect_supports_lookahead = decltype(InputAdapterType::supports_lookahead);
|
||||||
|
|
||||||
template<typename InputAdapterType>
|
template<typename InputAdapterType>
|
||||||
constexpr bool input_adapter_supports_bulk_scan(std::true_type /*detected*/)
|
constexpr bool input_adapter_supports_lookahead(std::true_type /*detected*/)
|
||||||
{
|
{
|
||||||
return InputAdapterType::supports_bulk_scan;
|
return InputAdapterType::supports_lookahead;
|
||||||
}
|
}
|
||||||
|
|
||||||
template<typename InputAdapterType>
|
template<typename InputAdapterType>
|
||||||
constexpr bool input_adapter_supports_bulk_scan(std::false_type /*detected*/)
|
constexpr bool input_adapter_supports_lookahead(std::false_type /*detected*/)
|
||||||
{
|
{
|
||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
@@ -167,13 +164,11 @@ class lexer : public lexer_base<BasicJsonType>
|
|||||||
static constexpr bool lazy_token_string =
|
static constexpr bool lazy_token_string =
|
||||||
input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {});
|
input_adapter_supports_seek<InputAdapterType>(is_detected<detect_supports_seek, InputAdapterType> {});
|
||||||
|
|
||||||
/// whether string scanning may bulk-consume runs of ordinary characters
|
/// whether a simulated unget can be passed on to the input adapter, which
|
||||||
/// directly from a contiguous input buffer (SWAR fast path). This requires
|
/// then leaves the character in the input; see
|
||||||
/// the token to be reconstructible lazily (lazy_token_string), so bypassing
|
/// input_adapter_supports_lookahead
|
||||||
/// the per-character capture in get() cannot lose error diagnostics.
|
static constexpr bool can_release_lookahead =
|
||||||
static constexpr bool bulk_scan =
|
input_adapter_supports_lookahead<InputAdapterType>(is_detected<detect_supports_lookahead, InputAdapterType> {});
|
||||||
lazy_token_string
|
|
||||||
&& input_adapter_supports_bulk_scan<InputAdapterType>(is_detected<detect_supports_bulk_scan, InputAdapterType> {});
|
|
||||||
|
|
||||||
public:
|
public:
|
||||||
using token_type = typename lexer_base<BasicJsonType>::token_type;
|
using token_type = typename lexer_base<BasicJsonType>::token_type;
|
||||||
@@ -294,40 +289,6 @@ class lexer : public lexer_base<BasicJsonType>
|
|||||||
return true;
|
return true;
|
||||||
}
|
}
|
||||||
|
|
||||||
/// contiguous input: bulk-append the run of ordinary characters and complete
|
|
||||||
/// well-formed UTF-8 sequences starting at the current read position, leaving
|
|
||||||
/// the first byte that needs individual handling (the closing quote, an
|
|
||||||
/// escape, a control character, or an ill-formed UTF-8 byte) for get()
|
|
||||||
void scan_string_bulk(std::true_type /*bulk*/)
|
|
||||||
{
|
|
||||||
// a pending unget must be consumed through the normal path first
|
|
||||||
if (next_unget)
|
|
||||||
{
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
const std::size_t remaining = ia.bulk_remaining();
|
|
||||||
if (remaining == 0)
|
|
||||||
{
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
const auto* const data = reinterpret_cast<const unsigned char*>(ia.bulk_data());
|
|
||||||
|
|
||||||
const std::size_t pos = string_bulk_run(data, remaining);
|
|
||||||
if (pos == 0)
|
|
||||||
{
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), pos);
|
|
||||||
ia.bulk_skip(pos);
|
|
||||||
// the run contains no newline (all bytes < 0x20 are treated as special),
|
|
||||||
// so only the flat character counters advance
|
|
||||||
position.chars_read_total += pos;
|
|
||||||
position.chars_read_current_line += pos;
|
|
||||||
}
|
|
||||||
|
|
||||||
/// streaming input: no bulk fast path
|
|
||||||
void scan_string_bulk(std::false_type /*bulk*/) const noexcept {}
|
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
@brief scan a string literal
|
@brief scan a string literal
|
||||||
|
|
||||||
@@ -353,10 +314,6 @@ class lexer : public lexer_base<BasicJsonType>
|
|||||||
|
|
||||||
while (true)
|
while (true)
|
||||||
{
|
{
|
||||||
// bulk-consume ordinary characters from contiguous input, then
|
|
||||||
// handle the next special byte through the switch below
|
|
||||||
scan_string_bulk(std::integral_constant<bool, bulk_scan> {});
|
|
||||||
|
|
||||||
// get the next character
|
// get the next character
|
||||||
switch (get())
|
switch (get())
|
||||||
{
|
{
|
||||||
@@ -1346,56 +1303,45 @@ scan_number_done:
|
|||||||
// we are done scanning a number)
|
// we are done scanning a number)
|
||||||
unget();
|
unget();
|
||||||
|
|
||||||
return convert_number(number_type);
|
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
|
||||||
}
|
errno = 0;
|
||||||
|
|
||||||
/*!
|
// try to parse integers first and fall back to floats
|
||||||
@brief convert the number text in token_buffer to its value and token type
|
|
||||||
|
|
||||||
The digit sequence in token_buffer has already been validated (by the
|
|
||||||
scan_number() state machine or by the contiguous fast path) and holds the
|
|
||||||
locale decimal point in place of '.'. Integers are parsed first and fall
|
|
||||||
back to floating point on overflow. This is shared so both scanners produce
|
|
||||||
identical results.
|
|
||||||
*/
|
|
||||||
token_type convert_number(token_type number_type)
|
|
||||||
{
|
|
||||||
const char* const num_begin = token_buffer.data();
|
|
||||||
const char* const num_end = num_begin + token_buffer.size();
|
|
||||||
|
|
||||||
// try to parse integers first and fall back to floats; the digit
|
|
||||||
// sequence has already been validated, so a dedicated parser can avoid
|
|
||||||
// the locale/errno overhead of strtoull
|
|
||||||
if (number_type == token_type::value_unsigned)
|
if (number_type == token_type::value_unsigned)
|
||||||
{
|
{
|
||||||
if (parse_integer_unsigned(num_begin, num_end, value_unsigned))
|
const auto x = std::strtoull(token_buffer.data(), &endptr, 10);
|
||||||
|
|
||||||
|
// we checked the number format before
|
||||||
|
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
|
||||||
|
|
||||||
|
if (errno != ERANGE)
|
||||||
{
|
{
|
||||||
return token_type::value_unsigned;
|
value_unsigned = static_cast<number_unsigned_t>(x);
|
||||||
|
if (value_unsigned == x)
|
||||||
|
{
|
||||||
|
return token_type::value_unsigned;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
else if (number_type == token_type::value_integer)
|
else if (number_type == token_type::value_integer)
|
||||||
{
|
{
|
||||||
if (parse_integer_signed(num_begin, num_end, value_integer))
|
const auto x = std::strtoll(token_buffer.data(), &endptr, 10);
|
||||||
|
|
||||||
|
// we checked the number format before
|
||||||
|
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
|
||||||
|
|
||||||
|
if (errno != ERANGE)
|
||||||
{
|
{
|
||||||
return token_type::value_integer;
|
value_integer = static_cast<number_integer_t>(x);
|
||||||
|
if (value_integer == x)
|
||||||
|
{
|
||||||
|
return token_type::value_integer;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// this code is reached if we parse a floating-point number or if an
|
// this code is reached if we parse a floating-point number or if an
|
||||||
// integer conversion above overflowed. Prefer std::from_chars
|
// integer conversion above failed
|
||||||
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
|
|
||||||
// otherwise the exact Clinger fast path (double only); otherwise the
|
|
||||||
// locale-aware strtof/strtod.
|
|
||||||
if (parse_float_from_chars(num_begin, num_end, value_float))
|
|
||||||
{
|
|
||||||
return token_type::value_float;
|
|
||||||
}
|
|
||||||
if (parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
|
|
||||||
{
|
|
||||||
return token_type::value_float;
|
|
||||||
}
|
|
||||||
|
|
||||||
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
|
|
||||||
strtof(value_float, token_buffer.data(), &endptr);
|
strtof(value_float, token_buffer.data(), &endptr);
|
||||||
|
|
||||||
// we checked the number format before
|
// we checked the number format before
|
||||||
@@ -1404,130 +1350,6 @@ scan_number_done:
|
|||||||
return token_type::value_float;
|
return token_type::value_float;
|
||||||
}
|
}
|
||||||
|
|
||||||
/*!
|
|
||||||
@brief contiguous fast path for scanning a number
|
|
||||||
|
|
||||||
Parses the whole number token straight from the input buffer, avoiding the
|
|
||||||
per-character get()/add() of scan_number(). On success it fills token_buffer
|
|
||||||
(with the locale decimal point substituted, as scan_number() does) and
|
|
||||||
returns the token type. On anything it does not fully recognize as a
|
|
||||||
well-formed number it makes no state change and returns
|
|
||||||
token_type::uninitialized, so the caller falls back to scan_number(), which
|
|
||||||
then produces the exact diagnostic. @a current is the first digit or the
|
|
||||||
leading minus (already read); the remaining bytes are taken from the adapter.
|
|
||||||
*/
|
|
||||||
token_type scan_number_bulk_contiguous()
|
|
||||||
{
|
|
||||||
// a pending unget offsets the buffer position from current; fall back
|
|
||||||
if (next_unget)
|
|
||||||
{
|
|
||||||
return token_type::uninitialized;
|
|
||||||
}
|
|
||||||
const std::size_t rem = ia.bulk_remaining();
|
|
||||||
if (rem == 0)
|
|
||||||
{
|
|
||||||
// the first digit is the last input byte; let scan_number() finish
|
|
||||||
return token_type::uninitialized;
|
|
||||||
}
|
|
||||||
// the byte before the next unread one is current (contiguous input)
|
|
||||||
const char* const data = reinterpret_cast<const char*>(ia.bulk_data()) - 1;
|
|
||||||
const std::size_t avail = rem + 1;
|
|
||||||
|
|
||||||
// validate + classify the number extent (mirrors scan_number()'s grammar)
|
|
||||||
std::size_t i = 0;
|
|
||||||
std::size_t dot_index = std::string::npos;
|
|
||||||
token_type number_type = token_type::value_unsigned;
|
|
||||||
if (data[0] == '-')
|
|
||||||
{
|
|
||||||
number_type = token_type::value_integer;
|
|
||||||
i = 1;
|
|
||||||
if (i >= avail)
|
|
||||||
{
|
|
||||||
return token_type::uninitialized;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if (data[i] == '0')
|
|
||||||
{
|
|
||||||
++i;
|
|
||||||
}
|
|
||||||
else if (data[i] >= '1' && data[i] <= '9')
|
|
||||||
{
|
|
||||||
++i;
|
|
||||||
while (i < avail && data[i] >= '0' && data[i] <= '9')
|
|
||||||
{
|
|
||||||
++i;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
return token_type::uninitialized;
|
|
||||||
}
|
|
||||||
if (i < avail && data[i] == '.')
|
|
||||||
{
|
|
||||||
number_type = token_type::value_float;
|
|
||||||
dot_index = i;
|
|
||||||
++i;
|
|
||||||
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
|
|
||||||
{
|
|
||||||
return token_type::uninitialized;
|
|
||||||
}
|
|
||||||
while (i < avail && data[i] >= '0' && data[i] <= '9')
|
|
||||||
{
|
|
||||||
++i;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if (i < avail && (data[i] == 'e' || data[i] == 'E'))
|
|
||||||
{
|
|
||||||
number_type = token_type::value_float;
|
|
||||||
++i;
|
|
||||||
if (i < avail && (data[i] == '+' || data[i] == '-'))
|
|
||||||
{
|
|
||||||
++i;
|
|
||||||
}
|
|
||||||
if (i >= avail || !(data[i] >= '0' && data[i] <= '9'))
|
|
||||||
{
|
|
||||||
return token_type::uninitialized;
|
|
||||||
}
|
|
||||||
while (i < avail && data[i] >= '0' && data[i] <= '9')
|
|
||||||
{
|
|
||||||
++i;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
const std::size_t len = i;
|
|
||||||
|
|
||||||
// materialize the token exactly as scan_number() would, substituting the
|
|
||||||
// locale decimal point so convert_number()'s strtof fallback stays valid.
|
|
||||||
// reset() already cleared token_buffer, so append() fills it (assign() is
|
|
||||||
// avoided because custom string_t types need not provide it)
|
|
||||||
reset();
|
|
||||||
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
|
|
||||||
if (dot_index != std::string::npos)
|
|
||||||
{
|
|
||||||
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
|
|
||||||
decimal_point_position = dot_index;
|
|
||||||
}
|
|
||||||
|
|
||||||
// consume the remaining bytes of the number (current was already read)
|
|
||||||
ia.bulk_skip(len - 1);
|
|
||||||
position.chars_read_total += (len - 1);
|
|
||||||
position.chars_read_current_line += (len - 1);
|
|
||||||
|
|
||||||
return convert_number(number_type);
|
|
||||||
}
|
|
||||||
|
|
||||||
/// contiguous input: try the number fast path, else the byte-path scanner
|
|
||||||
token_type scan_number_dispatch(std::true_type /*bulk*/)
|
|
||||||
{
|
|
||||||
const token_type t = scan_number_bulk_contiguous();
|
|
||||||
return (t != token_type::uninitialized) ? t : scan_number();
|
|
||||||
}
|
|
||||||
|
|
||||||
/// streaming input: always use the byte-path scanner
|
|
||||||
token_type scan_number_dispatch(std::false_type /*bulk*/)
|
|
||||||
{
|
|
||||||
return scan_number();
|
|
||||||
}
|
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
@param[in] literal_text the literal text to expect
|
@param[in] literal_text the literal text to expect
|
||||||
@param[in] length the length of the passed literal text
|
@param[in] length the length of the passed literal text
|
||||||
@@ -1658,6 +1480,21 @@ scan_number_done:
|
|||||||
uncapture_char(std::integral_constant<bool, lazy_token_string> {});
|
uncapture_char(std::integral_constant<bool, lazy_token_string> {});
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/// adapter without lookahead: nothing to do (see release_lookahead)
|
||||||
|
void release_lookahead_impl(std::false_type /*can_release*/) const noexcept {}
|
||||||
|
|
||||||
|
/// adapter with lookahead: leave the character in the input instead
|
||||||
|
void release_lookahead_impl(std::true_type /*can_release*/)
|
||||||
|
{
|
||||||
|
if (next_unget)
|
||||||
|
{
|
||||||
|
// the character is read from the input again rather than replayed
|
||||||
|
// from current, so the adapter must not step over it
|
||||||
|
next_unget = false;
|
||||||
|
ia.release_lookahead();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
/// seekable adapter: nothing was captured, so nothing to undo
|
/// seekable adapter: nothing was captured, so nothing to undo
|
||||||
void uncapture_char(std::true_type /*lazy*/) const noexcept {}
|
void uncapture_char(std::true_type /*lazy*/) const noexcept {}
|
||||||
|
|
||||||
@@ -1721,6 +1558,29 @@ scan_number_done:
|
|||||||
return position;
|
return position;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/*!
|
||||||
|
@brief pass a pending simulated unget on to the input
|
||||||
|
|
||||||
|
unget() only rewinds the lexer's own bookkeeping, so the character that
|
||||||
|
terminated the last token (e.g. the character after a number) would still
|
||||||
|
be stepped over when the input adapter is done. Callers that hand the
|
||||||
|
input back to the user afterwards - operator>> and non-strict sax_parse -
|
||||||
|
call this once when scanning is done, so that the input is positioned
|
||||||
|
right after the value.
|
||||||
|
|
||||||
|
Adapters without lookahead (see input_adapter_supports_lookahead) are not
|
||||||
|
handed back to the user, so this is a no-op for them.
|
||||||
|
|
||||||
|
Scanning may continue after this call: @a next_unget is cleared, and the
|
||||||
|
character is read from the input again instead of being replayed from
|
||||||
|
@a current. A pending unget of EOF needs no special case, because reaching
|
||||||
|
EOF leaves no lookahead to release.
|
||||||
|
*/
|
||||||
|
void release_lookahead()
|
||||||
|
{
|
||||||
|
release_lookahead_impl(std::integral_constant<bool, can_release_lookahead> {});
|
||||||
|
}
|
||||||
|
|
||||||
/// seekable adapter: rebuild the last read token from the input on demand
|
/// seekable adapter: rebuild the last read token from the input on demand
|
||||||
const std::vector<char_type>& collect_token_chars(std::vector<char_type>& out, std::true_type /*lazy*/) const
|
const std::vector<char_type>& collect_token_chars(std::vector<char_type>& out, std::true_type /*lazy*/) const
|
||||||
{
|
{
|
||||||
@@ -1882,7 +1742,7 @@ scan_number_done:
|
|||||||
case '7':
|
case '7':
|
||||||
case '8':
|
case '8':
|
||||||
case '9':
|
case '9':
|
||||||
return scan_number_dispatch(std::integral_constant<bool, bulk_scan> {});
|
return scan_number();
|
||||||
|
|
||||||
// end of input (the null byte is needed when parsing from
|
// end of input (the null byte is needed when parsing from
|
||||||
// string literals)
|
// string literals)
|
||||||
|
|||||||
@@ -1,302 +0,0 @@
|
|||||||
// __ _____ _____ _____
|
|
||||||
// __| | __| | | | JSON for Modern C++
|
|
||||||
// | | |__ | | | | | | version 3.12.0
|
|
||||||
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
||||||
//
|
|
||||||
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
||||||
// SPDX-License-Identifier: MIT
|
|
||||||
|
|
||||||
#pragma once
|
|
||||||
|
|
||||||
#include <array> // array
|
|
||||||
#include <cfloat> // FLT_EVAL_METHOD
|
|
||||||
#include <cstddef> // size_t
|
|
||||||
#include <cstdint> // int64_t, uint64_t
|
|
||||||
#include <limits> // numeric_limits
|
|
||||||
|
|
||||||
#include <nlohmann/detail/macro_scope.hpp>
|
|
||||||
|
|
||||||
// std::from_chars lives in <charconv>, but being in C++17 mode does not
|
|
||||||
// guarantee the header exists: GCC 7 sets __cplusplus to C++17 yet ships no
|
|
||||||
// <charconv> (added in GCC 8; floating-point support in GCC 11). Guard the
|
|
||||||
// include with __has_include so such toolchains fall back to the scalar path.
|
|
||||||
#if defined(JSON_HAS_CPP_17) && defined(__has_include)
|
|
||||||
#if __has_include(<charconv>)
|
|
||||||
#include <charconv> // from_chars (only used when __cpp_lib_to_chars is defined)
|
|
||||||
#include <system_error> // errc
|
|
||||||
#endif
|
|
||||||
#endif
|
|
||||||
|
|
||||||
// This file contains the value-conversion helpers used by the lexer to turn an
|
|
||||||
// already-validated number token into a value, without the locale/errno
|
|
||||||
// overhead of std::strtoull/std::strtod. They are free functions so the lexer
|
|
||||||
// stays focused on scanning; see lexer::convert_number().
|
|
||||||
|
|
||||||
NLOHMANN_JSON_NAMESPACE_BEGIN
|
|
||||||
namespace detail
|
|
||||||
{
|
|
||||||
|
|
||||||
/*!
|
|
||||||
@brief fast integer parser for an already-validated unsigned integer
|
|
||||||
|
|
||||||
The number scanner has already checked that [first, last) is a valid JSON
|
|
||||||
integer, so this only needs to accumulate the digits and detect overflow. This
|
|
||||||
avoids the locale/errno machinery of std::strtoull, which dominates
|
|
||||||
integer-heavy inputs.
|
|
||||||
|
|
||||||
@param[in] first pointer to the first character (a digit)
|
|
||||||
@param[in] last pointer past the last character
|
|
||||||
@param[out] value the parsed value on success
|
|
||||||
@return true if the value fit into @a NumberUnsignedType; false on overflow, in
|
|
||||||
which case the caller falls back to floating-point parsing (matching the
|
|
||||||
previous std::strtoull behavior)
|
|
||||||
*/
|
|
||||||
template<typename NumberUnsignedType>
|
|
||||||
bool parse_integer_unsigned(const char* first, const char* last, NumberUnsignedType& value) noexcept
|
|
||||||
{
|
|
||||||
// accumulate in the widest unsigned type used by the previous strtoull
|
|
||||||
// path so the overflow behavior is unchanged for custom number types
|
|
||||||
std::uint64_t x = 0;
|
|
||||||
constexpr std::uint64_t cutoff = (std::numeric_limits<std::uint64_t>::max)() / 10u;
|
|
||||||
constexpr std::uint64_t cutlim = (std::numeric_limits<std::uint64_t>::max)() % 10u;
|
|
||||||
for (const char* p = first; p != last; ++p)
|
|
||||||
{
|
|
||||||
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(x > cutoff || (x == cutoff && digit > cutlim)))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
x = (x * 10u) + digit;
|
|
||||||
}
|
|
||||||
value = static_cast<NumberUnsignedType>(x);
|
|
||||||
// reject values that do not round-trip into a narrower NumberUnsignedType
|
|
||||||
return static_cast<std::uint64_t>(value) == x;
|
|
||||||
}
|
|
||||||
|
|
||||||
/*!
|
|
||||||
@brief fast integer parser for an already-validated negative integer
|
|
||||||
|
|
||||||
@param[in] first pointer to the leading '-'
|
|
||||||
@param[in] last pointer past the last character
|
|
||||||
@param[out] value the parsed (negative) value on success
|
|
||||||
@return true on success; false on overflow (caller falls back to float)
|
|
||||||
*/
|
|
||||||
template<typename NumberIntegerType>
|
|
||||||
bool parse_integer_signed(const char* first, const char* last, NumberIntegerType& value) noexcept
|
|
||||||
{
|
|
||||||
// the state machine only reaches the signed path via a leading '-'
|
|
||||||
JSON_ASSERT(first != last && *first == '-');
|
|
||||||
std::uint64_t magnitude = 0;
|
|
||||||
// |INT64_MIN| == INT64_MAX + 1; this is the largest admissible magnitude
|
|
||||||
constexpr std::uint64_t limit = static_cast<std::uint64_t>((std::numeric_limits<std::int64_t>::max)()) + 1u;
|
|
||||||
for (const char* p = first + 1; p != last; ++p)
|
|
||||||
{
|
|
||||||
const auto digit = static_cast<std::uint64_t>(static_cast<unsigned char>(*p) - static_cast<unsigned char>('0'));
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(magnitude > (limit - digit) / 10u))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
magnitude = (magnitude * 10u) + digit;
|
|
||||||
}
|
|
||||||
const std::int64_t x = (magnitude == limit)
|
|
||||||
? (std::numeric_limits<std::int64_t>::min)()
|
|
||||||
: -static_cast<std::int64_t>(magnitude);
|
|
||||||
value = static_cast<NumberIntegerType>(x);
|
|
||||||
// reject values that do not round-trip into a narrower NumberIntegerType
|
|
||||||
return static_cast<std::int64_t>(value) == x;
|
|
||||||
}
|
|
||||||
|
|
||||||
/*!
|
|
||||||
@brief exact fast path for parsing a `double` (Clinger's algorithm)
|
|
||||||
|
|
||||||
For the common case - at most 19 significant digits, a decimal exponent in
|
|
||||||
[-22, 22], and a significand below 2^53 - the value equals significand *
|
|
||||||
10^exp computed in IEEE-754 double arithmetic, which is exact under
|
|
||||||
round-to-nearest because both operands are exactly representable. This is the
|
|
||||||
same fast path used by fast_float/simdjson; the general cases are left to
|
|
||||||
std::strtod. The parser only activates for number_float_t == double; float and
|
|
||||||
long double keep the std::strtof/std::strtold paths (see the templated overload
|
|
||||||
below).
|
|
||||||
|
|
||||||
@param[in] first pointer to the first character of the number
|
|
||||||
@param[in] last pointer past the last character
|
|
||||||
@param[in] decimal_point the (locale-dependent) decimal point character
|
|
||||||
@param[out] out the parsed value on success
|
|
||||||
@return true if the value was parsed exactly; false to fall back to strtod
|
|
||||||
*/
|
|
||||||
template<typename DecimalPointType>
|
|
||||||
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
|
|
||||||
{
|
|
||||||
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
|
|
||||||
// Clinger's fast path is only exact when double operations are evaluated in
|
|
||||||
// true double precision. On platforms that keep intermediates in extended
|
|
||||||
// precision (e.g. the x87 FPU on 32-bit x86, where FLT_EVAL_METHOD == 2) the
|
|
||||||
// single significand * 10^scale step is double-rounded and can be 1 ULP off,
|
|
||||||
// so decline and let the caller fall back to the correctly-rounded
|
|
||||||
// std::from_chars / std::strtod path.
|
|
||||||
static_cast<void>(first);
|
|
||||||
static_cast<void>(last);
|
|
||||||
static_cast<void>(decimal_point);
|
|
||||||
static_cast<void>(out);
|
|
||||||
return false;
|
|
||||||
#else
|
|
||||||
static const std::array<double, 23> powers_of_ten =
|
|
||||||
{
|
|
||||||
{
|
|
||||||
1e0, 1e1, 1e2, 1e3, 1e4, 1e5, 1e6, 1e7, 1e8, 1e9, 1e10, 1e11,
|
|
||||||
1e12, 1e13, 1e14, 1e15, 1e16, 1e17, 1e18, 1e19, 1e20, 1e21, 1e22
|
|
||||||
}
|
|
||||||
};
|
|
||||||
|
|
||||||
const char* p = first;
|
|
||||||
bool negative = false;
|
|
||||||
if (p != last && (*p == '-' || *p == '+'))
|
|
||||||
{
|
|
||||||
negative = (*p == '-');
|
|
||||||
++p;
|
|
||||||
}
|
|
||||||
|
|
||||||
std::uint64_t significand = 0;
|
|
||||||
int num_digits = 0;
|
|
||||||
int fractional_digits = 0;
|
|
||||||
bool seen_dot = false;
|
|
||||||
bool any_digit = false;
|
|
||||||
for (; p != last; ++p)
|
|
||||||
{
|
|
||||||
const char c = *p;
|
|
||||||
if (c >= '0' && c <= '9')
|
|
||||||
{
|
|
||||||
any_digit = true;
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(num_digits >= 19))
|
|
||||||
{
|
|
||||||
return false; // significand may not fit into uint64_t
|
|
||||||
}
|
|
||||||
significand = (significand * 10u) + static_cast<std::uint64_t>(c - '0');
|
|
||||||
++num_digits;
|
|
||||||
fractional_digits += static_cast<int>(seen_dot);
|
|
||||||
}
|
|
||||||
else if (static_cast<DecimalPointType>(c) == decimal_point)
|
|
||||||
{
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(seen_dot))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
seen_dot = true;
|
|
||||||
}
|
|
||||||
else if (c == 'e' || c == 'E')
|
|
||||||
{
|
|
||||||
++p;
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(!any_digit))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
|
|
||||||
int exponent = 0;
|
|
||||||
if (p != last) // an exponent part remains
|
|
||||||
{
|
|
||||||
bool exp_negative = false;
|
|
||||||
if (p != last && (*p == '-' || *p == '+'))
|
|
||||||
{
|
|
||||||
exp_negative = (*p == '-');
|
|
||||||
++p;
|
|
||||||
}
|
|
||||||
bool any_exp_digit = false;
|
|
||||||
for (; p != last; ++p)
|
|
||||||
{
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(*p < '0' || *p > '9'))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
exponent = (exponent * 10) + (*p - '0');
|
|
||||||
any_exp_digit = true;
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(exponent > 9999))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(!any_exp_digit))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
if (exp_negative)
|
|
||||||
{
|
|
||||||
exponent = -exponent;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
const int scale = exponent - fractional_digits;
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(significand >= (static_cast<std::uint64_t>(1) << 53)))
|
|
||||||
{
|
|
||||||
return false; // significand not exactly representable as double
|
|
||||||
}
|
|
||||||
|
|
||||||
auto result = static_cast<double>(significand);
|
|
||||||
if (scale >= 0)
|
|
||||||
{
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(scale > 22))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
result *= powers_of_ten[static_cast<std::size_t>(scale)];
|
|
||||||
}
|
|
||||||
else
|
|
||||||
{
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(-scale > 22))
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
result /= powers_of_ten[static_cast<std::size_t>(-scale)];
|
|
||||||
}
|
|
||||||
out = negative ? -result : result;
|
|
||||||
return true;
|
|
||||||
#endif
|
|
||||||
}
|
|
||||||
|
|
||||||
/// fast float path is only exact for `double`; decline for float/long double
|
|
||||||
template<typename DecimalPointType, typename FloatType>
|
|
||||||
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
|
|
||||||
{
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
|
|
||||||
/*!
|
|
||||||
@brief parse a float with std::from_chars (Eisel-Lemire) when available
|
|
||||||
|
|
||||||
std::from_chars is locale-independent, correctly rounded, and - via the
|
|
||||||
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
|
|
||||||
over the whole value range (not just the Clinger subset). It is used only when
|
|
||||||
__cpp_lib_to_chars indicates full floating-point support and only when it
|
|
||||||
consumes the entire token ([first, last)); a partial parse means the buffer
|
|
||||||
uses a non-'.' locale decimal point, in which case the caller falls back to the
|
|
||||||
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
|
|
||||||
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
|
|
||||||
expects (side-stepping the P4168 divergence between implementations).
|
|
||||||
|
|
||||||
@return true if the value was parsed exactly and fully; false to fall back
|
|
||||||
*/
|
|
||||||
template<typename FloatType>
|
|
||||||
bool parse_float_from_chars(const char* first, const char* last, FloatType& out) noexcept
|
|
||||||
{
|
|
||||||
// JSON_HAS_CPP_17 must gate the use as well as the <charconv> include above:
|
|
||||||
// some standard libraries (e.g. libstdc++ 15) define __cpp_lib_to_chars even
|
|
||||||
// in C++14 mode, where <charconv> is not included.
|
|
||||||
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
|
|
||||||
const auto result = std::from_chars(first, last, out);
|
|
||||||
return result.ec == std::errc() && result.ptr == last;
|
|
||||||
#else
|
|
||||||
static_cast<void>(first);
|
|
||||||
static_cast<void>(last);
|
|
||||||
static_cast<void>(out);
|
|
||||||
return false;
|
|
||||||
#endif
|
|
||||||
}
|
|
||||||
|
|
||||||
} // namespace detail
|
|
||||||
NLOHMANN_JSON_NAMESPACE_END
|
|
||||||
@@ -99,8 +99,14 @@ class parser
|
|||||||
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
|
json_sax_dom_callback_parser<BasicJsonType, InputAdapterType> sdp(result, callback, allow_exceptions, &m_lexer);
|
||||||
sax_parse_internal(&sdp);
|
sax_parse_internal(&sdp);
|
||||||
|
|
||||||
|
if (!strict)
|
||||||
|
{
|
||||||
|
// the caller keeps using the input: position it right after
|
||||||
|
// the value by leaving the character that terminated it
|
||||||
|
m_lexer.release_lookahead();
|
||||||
|
}
|
||||||
// in strict mode, input must be completely read
|
// in strict mode, input must be completely read
|
||||||
if (strict && (get_token() != token_type::end_of_input))
|
else if (get_token() != token_type::end_of_input)
|
||||||
{
|
{
|
||||||
sdp.parse_error(m_lexer.get_position(),
|
sdp.parse_error(m_lexer.get_position(),
|
||||||
m_lexer.get_token_string(),
|
m_lexer.get_token_string(),
|
||||||
@@ -127,8 +133,13 @@ class parser
|
|||||||
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
|
json_sax_dom_parser<BasicJsonType, InputAdapterType> sdp(result, allow_exceptions, &m_lexer);
|
||||||
sax_parse_internal(&sdp);
|
sax_parse_internal(&sdp);
|
||||||
|
|
||||||
|
if (!strict)
|
||||||
|
{
|
||||||
|
// see above
|
||||||
|
m_lexer.release_lookahead();
|
||||||
|
}
|
||||||
// in strict mode, input must be completely read
|
// in strict mode, input must be completely read
|
||||||
if (strict && (get_token() != token_type::end_of_input))
|
else if (get_token() != token_type::end_of_input)
|
||||||
{
|
{
|
||||||
sdp.parse_error(m_lexer.get_position(),
|
sdp.parse_error(m_lexer.get_position(),
|
||||||
m_lexer.get_token_string(),
|
m_lexer.get_token_string(),
|
||||||
@@ -165,8 +176,14 @@ class parser
|
|||||||
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
|
(void)detail::is_sax_static_asserts<SAX, BasicJsonType> {};
|
||||||
const bool result = sax_parse_internal(sax);
|
const bool result = sax_parse_internal(sax);
|
||||||
|
|
||||||
|
if (result && !strict)
|
||||||
|
{
|
||||||
|
// the caller keeps using the input: position it right after the
|
||||||
|
// value by leaving the character that terminated it
|
||||||
|
m_lexer.release_lookahead();
|
||||||
|
}
|
||||||
// strict mode: next byte must be EOF
|
// strict mode: next byte must be EOF
|
||||||
if (result && strict && (get_token() != token_type::end_of_input))
|
else if (result && strict && (get_token() != token_type::end_of_input))
|
||||||
{
|
{
|
||||||
return sax->parse_error(m_lexer.get_position(),
|
return sax->parse_error(m_lexer.get_position(),
|
||||||
m_lexer.get_token_string(),
|
m_lexer.get_token_string(),
|
||||||
|
|||||||
@@ -1,289 +0,0 @@
|
|||||||
// __ _____ _____ _____
|
|
||||||
// __| | __| | | | JSON for Modern C++
|
|
||||||
// | | |__ | | | | | | version 3.12.0
|
|
||||||
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
|
|
||||||
//
|
|
||||||
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
|
|
||||||
// SPDX-License-Identifier: MIT
|
|
||||||
|
|
||||||
#pragma once
|
|
||||||
|
|
||||||
#include <cstddef> // size_t
|
|
||||||
#include <cstdint> // uint64_t
|
|
||||||
#include <cstring> // memcpy
|
|
||||||
|
|
||||||
#if defined(JSON_USE_SIMDUTF)
|
|
||||||
// Optional SIMD backend for bulk UTF-8 validation. This is an opt-in
|
|
||||||
// external dependency: nlohmann/json itself stays header-only and the C++11
|
|
||||||
// scalar validator below is always available; defining JSON_USE_SIMDUTF
|
|
||||||
// additionally requires the simdutf headers on the include path and linking
|
|
||||||
// the simdutf library. See string_bulk_run().
|
|
||||||
#include <simdutf.h>
|
|
||||||
#endif
|
|
||||||
|
|
||||||
#include <nlohmann/detail/macro_scope.hpp>
|
|
||||||
|
|
||||||
// This file contains the byte-level string-scanning helpers used by the lexer's
|
|
||||||
// contiguous fast path. They operate purely on raw bytes (no dependency on the
|
|
||||||
// lexer's template parameters) so they are free functions, keeping the lexer
|
|
||||||
// itself focused on the state machine; see lexer::scan_string_bulk().
|
|
||||||
|
|
||||||
NLOHMANN_JSON_NAMESPACE_BEGIN
|
|
||||||
namespace detail
|
|
||||||
{
|
|
||||||
|
|
||||||
// classify a single byte as needing individual string handling: the closing
|
|
||||||
// quote, an escape, a control character, or a non-ASCII (UTF-8)
|
|
||||||
// lead/continuation byte. Ordinary bytes (0x20..0x7F except '"' and '\\') are
|
|
||||||
// copied verbatim, which the bulk scanner does 8 bytes at a time.
|
|
||||||
inline bool is_string_special(unsigned char c) noexcept
|
|
||||||
{
|
|
||||||
return c == '\"' || c == '\\' || c < 0x20u || c >= 0x80u;
|
|
||||||
}
|
|
||||||
|
|
||||||
// SWAR helper: return a word whose high bit is set in every byte of @a v that
|
|
||||||
// is_string_special(); zero if the 8 bytes are all ordinary.
|
|
||||||
inline std::uint64_t swar_string_special(std::uint64_t v) noexcept
|
|
||||||
{
|
|
||||||
constexpr std::uint64_t ones = 0x0101010101010101ull;
|
|
||||||
constexpr std::uint64_t high = 0x8080808080808080ull;
|
|
||||||
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
|
|
||||||
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
|
|
||||||
const std::uint64_t has_quote = (q - ones) & ~q & high;
|
|
||||||
const std::uint64_t has_backslash = (b - ones) & ~b & high;
|
|
||||||
const std::uint64_t has_control = (v - 0x2020202020202020ull) & ~v & high; // < 0x20
|
|
||||||
const std::uint64_t has_non_ascii = v & high; // >= 0x80
|
|
||||||
return has_quote | has_backslash | has_control | has_non_ascii;
|
|
||||||
}
|
|
||||||
|
|
||||||
// return the index of the first is_string_special() byte in [data, data+n), or
|
|
||||||
// n if every byte is ordinary; scans 8 bytes at a time
|
|
||||||
inline std::size_t find_string_special(const unsigned char* data, std::size_t n) noexcept
|
|
||||||
{
|
|
||||||
std::size_t i = 0;
|
|
||||||
for (; i + 8 <= n; i += 8)
|
|
||||||
{
|
|
||||||
std::uint64_t word = 0;
|
|
||||||
std::memcpy(&word, data + i, sizeof(word));
|
|
||||||
if (swar_string_special(word) != 0)
|
|
||||||
{
|
|
||||||
// a special byte is in this word; locate it (endian-agnostic)
|
|
||||||
for (std::size_t j = 0; j < 8; ++j)
|
|
||||||
{
|
|
||||||
if (is_string_special(data[i + j]))
|
|
||||||
{
|
|
||||||
return i + j;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
for (; i < n; ++i)
|
|
||||||
{
|
|
||||||
if (is_string_special(data[i]))
|
|
||||||
{
|
|
||||||
return i;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return n;
|
|
||||||
}
|
|
||||||
|
|
||||||
// classify a byte as one the serializer must NOT copy verbatim when
|
|
||||||
// ensure_ascii is requested: the closing quote, an escape, a control character
|
|
||||||
// (< 0x20), DEL (0x7F), or any non-ASCII byte (>= 0x80). Everything else -
|
|
||||||
// printable ASCII except '"' and '\\' - is emitted unchanged. Note this differs
|
|
||||||
// from is_string_special() only in that 0x7F is also a stop (it is escaped as
|
|
||||||
// \u007f under ensure_ascii).
|
|
||||||
inline bool is_ascii_copyable(unsigned char c) noexcept
|
|
||||||
{
|
|
||||||
return c >= 0x20u && c < 0x7Fu && c != '\"' && c != '\\';
|
|
||||||
}
|
|
||||||
|
|
||||||
// return the index of the first byte in [data, data+n) that is NOT
|
|
||||||
// is_ascii_copyable(), or n if every byte can be copied verbatim; scans 8 bytes
|
|
||||||
// at a time. Used by the serializer's ensure_ascii fast path.
|
|
||||||
inline std::size_t find_ascii_copyable_run(const unsigned char* data, std::size_t n) noexcept
|
|
||||||
{
|
|
||||||
constexpr std::uint64_t ones = 0x0101010101010101ull;
|
|
||||||
constexpr std::uint64_t high = 0x8080808080808080ull;
|
|
||||||
std::size_t i = 0;
|
|
||||||
for (; i + 8 <= n; i += 8)
|
|
||||||
{
|
|
||||||
std::uint64_t v = 0;
|
|
||||||
std::memcpy(&v, data + i, sizeof(v));
|
|
||||||
const std::uint64_t q = v ^ 0x2222222222222222ull; // '"' (0x22)
|
|
||||||
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull; // '\\' (0x5C)
|
|
||||||
const std::uint64_t d = v ^ 0x7F7F7F7F7F7F7F7Full; // DEL (0x7F)
|
|
||||||
const std::uint64_t stop = ((q - ones) & ~q & high) // == '"'
|
|
||||||
| ((b - ones) & ~b & high) // == '\\'
|
|
||||||
| ((d - ones) & ~d & high) // == 0x7F
|
|
||||||
| ((v - 0x2020202020202020ull) & ~v & high) // < 0x20
|
|
||||||
| (v & high); // >= 0x80
|
|
||||||
if (stop != 0)
|
|
||||||
{
|
|
||||||
for (std::size_t j = 0; j < 8; ++j)
|
|
||||||
{
|
|
||||||
if (!is_ascii_copyable(data[i + j]))
|
|
||||||
{
|
|
||||||
return i + j;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
for (; i < n; ++i)
|
|
||||||
{
|
|
||||||
if (!is_ascii_copyable(data[i]))
|
|
||||||
{
|
|
||||||
return i;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return n;
|
|
||||||
}
|
|
||||||
|
|
||||||
// Validate one UTF-8 sequence at the front of [data, data+avail). Returns its
|
|
||||||
// length (2..4) only when the bytes form a *well-formed* sequence using exactly
|
|
||||||
// the same ranges as scan_string()'s per-byte switch, so the bulk path accepts
|
|
||||||
// precisely what the byte path accepts. Returns 0 for anything that is invalid,
|
|
||||||
// incomplete, or that the byte path must diagnose (the caller then defers to
|
|
||||||
// that path, keeping error messages unchanged). Lead bytes < 0x80 are handled
|
|
||||||
// by the caller and never passed here.
|
|
||||||
inline std::size_t validate_one_utf8(const unsigned char* data, std::size_t avail) noexcept
|
|
||||||
{
|
|
||||||
const unsigned char c0 = data[0];
|
|
||||||
if (c0 >= 0xC2 && c0 <= 0xDF) // U+0080..U+07FF
|
|
||||||
{
|
|
||||||
if (avail >= 2 && data[1] >= 0x80 && data[1] <= 0xBF)
|
|
||||||
{
|
|
||||||
return 2;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else if (c0 == 0xE0) // U+0800..U+0FFF
|
|
||||||
{
|
|
||||||
if (avail >= 3 && data[1] >= 0xA0 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
|
|
||||||
{
|
|
||||||
return 3;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else if ((c0 >= 0xE1 && c0 <= 0xEC) || c0 == 0xEE || c0 == 0xEF) // U+1000..U+CFFF, U+E000..U+FFFF
|
|
||||||
{
|
|
||||||
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF)
|
|
||||||
{
|
|
||||||
return 3;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else if (c0 == 0xED) // U+D000..U+D7FF (excludes surrogates)
|
|
||||||
{
|
|
||||||
if (avail >= 3 && data[1] >= 0x80 && data[1] <= 0x9F && data[2] >= 0x80 && data[2] <= 0xBF)
|
|
||||||
{
|
|
||||||
return 3;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else if (c0 == 0xF0) // U+10000..U+3FFFF
|
|
||||||
{
|
|
||||||
if (avail >= 4 && data[1] >= 0x90 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
|
|
||||||
{
|
|
||||||
return 4;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else if (c0 >= 0xF1 && c0 <= 0xF3) // U+40000..U+FFFFF
|
|
||||||
{
|
|
||||||
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0xBF && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
|
|
||||||
{
|
|
||||||
return 4;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
else if (c0 == 0xF4) // U+100000..U+10FFFF
|
|
||||||
{
|
|
||||||
if (avail >= 4 && data[1] >= 0x80 && data[1] <= 0x8F && data[2] >= 0x80 && data[2] <= 0xBF && data[3] >= 0x80 && data[3] <= 0xBF)
|
|
||||||
{
|
|
||||||
return 4;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return 0; // invalid, incomplete, or must be diagnosed by the byte path
|
|
||||||
}
|
|
||||||
|
|
||||||
// Scalar (C++11) computation of the bulk run length: the number of leading
|
|
||||||
// bytes in [data, data+n) that are ordinary ASCII or complete well-formed UTF-8
|
|
||||||
// sequences, stopping before the first byte that needs individual handling (the
|
|
||||||
// closing quote, an escape, a control character, or an ill-formed/truncated
|
|
||||||
// sequence). ASCII is skipped 8 bytes at a time.
|
|
||||||
inline std::size_t scalar_string_bulk_run(const unsigned char* data, std::size_t n) noexcept
|
|
||||||
{
|
|
||||||
std::size_t pos = 0;
|
|
||||||
while (pos < n)
|
|
||||||
{
|
|
||||||
pos += find_string_special(data + pos, n - pos);
|
|
||||||
if (pos >= n || data[pos] < 0x80u)
|
|
||||||
{
|
|
||||||
break; // end of buffer, or a quote/escape/control byte
|
|
||||||
}
|
|
||||||
const std::size_t seq = validate_one_utf8(data + pos, n - pos);
|
|
||||||
if (seq == 0)
|
|
||||||
{
|
|
||||||
break; // ill-formed or truncated: let the byte path diagnose it
|
|
||||||
}
|
|
||||||
pos += seq;
|
|
||||||
}
|
|
||||||
return pos;
|
|
||||||
}
|
|
||||||
|
|
||||||
#if defined(JSON_USE_SIMDUTF)
|
|
||||||
// Index of the first quote/escape/control byte in [data, data+n) (non-ASCII
|
|
||||||
// bytes are *not* stops here - the whole run is handed to simdutf), or n.
|
|
||||||
inline std::size_t find_string_delimiter(const unsigned char* data, std::size_t n) noexcept
|
|
||||||
{
|
|
||||||
constexpr std::uint64_t ones = 0x0101010101010101ull;
|
|
||||||
constexpr std::uint64_t high = 0x8080808080808080ull;
|
|
||||||
std::size_t i = 0;
|
|
||||||
for (; i + 8 <= n; i += 8)
|
|
||||||
{
|
|
||||||
std::uint64_t v = 0;
|
|
||||||
std::memcpy(&v, data + i, sizeof(v));
|
|
||||||
const std::uint64_t q = v ^ 0x2222222222222222ull;
|
|
||||||
const std::uint64_t b = v ^ 0x5C5C5C5C5C5C5C5Cull;
|
|
||||||
const std::uint64_t hit = ((q - ones) & ~q & high)
|
|
||||||
| ((b - ones) & ~b & high)
|
|
||||||
| ((v - 0x2020202020202020ull) & ~v & high);
|
|
||||||
if (hit != 0)
|
|
||||||
{
|
|
||||||
for (std::size_t j = 0; j < 8; ++j)
|
|
||||||
{
|
|
||||||
const unsigned char c = data[i + j];
|
|
||||||
if (c == '\"' || c == '\\' || c < 0x20u)
|
|
||||||
{
|
|
||||||
return i + j;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
for (; i < n; ++i)
|
|
||||||
{
|
|
||||||
const unsigned char c = data[i];
|
|
||||||
if (c == '\"' || c == '\\' || c < 0x20u)
|
|
||||||
{
|
|
||||||
return i;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return n;
|
|
||||||
}
|
|
||||||
#endif
|
|
||||||
|
|
||||||
// Backend-dispatched bulk run length. With JSON_USE_SIMDUTF the run up to the
|
|
||||||
// next delimiter is validated in one shot by simdutf; on the rare failure the
|
|
||||||
// scalar helper recomputes the exact valid prefix so the byte path still
|
|
||||||
// produces the precise diagnostic. Without it, the pure scalar path is used.
|
|
||||||
inline std::size_t string_bulk_run(const unsigned char* data, std::size_t n) noexcept
|
|
||||||
{
|
|
||||||
#if defined(JSON_USE_SIMDUTF)
|
|
||||||
const std::size_t run = find_string_delimiter(data, n);
|
|
||||||
if (run != 0 && simdutf::validate_utf8(reinterpret_cast<const char*>(data), run))
|
|
||||||
{
|
|
||||||
return run;
|
|
||||||
}
|
|
||||||
return scalar_string_bulk_run(data, n);
|
|
||||||
#else
|
|
||||||
return scalar_string_bulk_run(data, n);
|
|
||||||
#endif
|
|
||||||
}
|
|
||||||
|
|
||||||
} // namespace detail
|
|
||||||
NLOHMANN_JSON_NAMESPACE_END
|
|
||||||
@@ -16,7 +16,6 @@
|
|||||||
#include <cstddef> // size_t, ptrdiff_t
|
#include <cstddef> // size_t, ptrdiff_t
|
||||||
#include <cstdint> // uint8_t
|
#include <cstdint> // uint8_t
|
||||||
#include <cstdio> // snprintf
|
#include <cstdio> // snprintf
|
||||||
#include <cstring> // memcpy
|
|
||||||
#include <limits> // numeric_limits
|
#include <limits> // numeric_limits
|
||||||
#include <string> // string, char_traits
|
#include <string> // string, char_traits
|
||||||
#include <iomanip> // setfill, setw
|
#include <iomanip> // setfill, setw
|
||||||
@@ -25,7 +24,6 @@
|
|||||||
|
|
||||||
#include <nlohmann/detail/conversions/to_chars.hpp>
|
#include <nlohmann/detail/conversions/to_chars.hpp>
|
||||||
#include <nlohmann/detail/exceptions.hpp>
|
#include <nlohmann/detail/exceptions.hpp>
|
||||||
#include <nlohmann/detail/input/string_scan.hpp>
|
|
||||||
#include <nlohmann/detail/macro_scope.hpp>
|
#include <nlohmann/detail/macro_scope.hpp>
|
||||||
#include <nlohmann/detail/meta/cpp_future.hpp>
|
#include <nlohmann/detail/meta/cpp_future.hpp>
|
||||||
#include <nlohmann/detail/output/binary_writer.hpp>
|
#include <nlohmann/detail/output/binary_writer.hpp>
|
||||||
@@ -111,25 +109,6 @@ class serializer
|
|||||||
const bool ensure_ascii,
|
const bool ensure_ascii,
|
||||||
const unsigned int indent_step,
|
const unsigned int indent_step,
|
||||||
const unsigned int current_indent = 0)
|
const unsigned int current_indent = 0)
|
||||||
{
|
|
||||||
dump_internal(val, pretty_print, ensure_ascii, indent_step, current_indent);
|
|
||||||
flush();
|
|
||||||
}
|
|
||||||
|
|
||||||
JSON_PRIVATE_UNLESS_TESTED:
|
|
||||||
/*!
|
|
||||||
@brief recursive worker for @ref dump
|
|
||||||
|
|
||||||
Identical in behavior to the historical @ref dump, but writes into the
|
|
||||||
serializer's internal @ref write_buffer instead of issuing a virtual call
|
|
||||||
per token. The public @ref dump wraps this and flushes the buffer once the
|
|
||||||
top-level value has been serialized.
|
|
||||||
*/
|
|
||||||
void dump_internal(const BasicJsonType& val,
|
|
||||||
const bool pretty_print,
|
|
||||||
const bool ensure_ascii,
|
|
||||||
const unsigned int indent_step,
|
|
||||||
const unsigned int current_indent = 0)
|
|
||||||
{
|
{
|
||||||
switch (val.m_data.m_type)
|
switch (val.m_data.m_type)
|
||||||
{
|
{
|
||||||
@@ -137,13 +116,13 @@ class serializer
|
|||||||
{
|
{
|
||||||
if (val.m_data.m_value.object->empty())
|
if (val.m_data.m_value.object->empty())
|
||||||
{
|
{
|
||||||
put_chars("{}", 2);
|
o->write_characters("{}", 2);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
if (pretty_print)
|
if (pretty_print)
|
||||||
{
|
{
|
||||||
put_chars("{\n", 2);
|
o->write_characters("{\n", 2);
|
||||||
|
|
||||||
// variable to hold indentation for recursive calls
|
// variable to hold indentation for recursive calls
|
||||||
const auto new_indent = current_indent + indent_step;
|
const auto new_indent = current_indent + indent_step;
|
||||||
@@ -156,51 +135,51 @@ class serializer
|
|||||||
auto i = val.m_data.m_value.object->cbegin();
|
auto i = val.m_data.m_value.object->cbegin();
|
||||||
for (std::size_t cnt = 0; cnt < val.m_data.m_value.object->size() - 1; ++cnt, ++i)
|
for (std::size_t cnt = 0; cnt < val.m_data.m_value.object->size() - 1; ++cnt, ++i)
|
||||||
{
|
{
|
||||||
put_chars(indent_string.c_str(), new_indent);
|
o->write_characters(indent_string.c_str(), new_indent);
|
||||||
put_char('\"');
|
o->write_character('\"');
|
||||||
dump_escaped(i->first, ensure_ascii);
|
dump_escaped(i->first, ensure_ascii);
|
||||||
put_chars("\": ", 3);
|
o->write_characters("\": ", 3);
|
||||||
dump_internal(i->second, true, ensure_ascii, indent_step, new_indent);
|
dump(i->second, true, ensure_ascii, indent_step, new_indent);
|
||||||
put_chars(",\n", 2);
|
o->write_characters(",\n", 2);
|
||||||
}
|
}
|
||||||
|
|
||||||
// last element
|
// last element
|
||||||
JSON_ASSERT(i != val.m_data.m_value.object->cend());
|
JSON_ASSERT(i != val.m_data.m_value.object->cend());
|
||||||
JSON_ASSERT(std::next(i) == val.m_data.m_value.object->cend());
|
JSON_ASSERT(std::next(i) == val.m_data.m_value.object->cend());
|
||||||
put_chars(indent_string.c_str(), new_indent);
|
o->write_characters(indent_string.c_str(), new_indent);
|
||||||
put_char('\"');
|
o->write_character('\"');
|
||||||
dump_escaped(i->first, ensure_ascii);
|
dump_escaped(i->first, ensure_ascii);
|
||||||
put_chars("\": ", 3);
|
o->write_characters("\": ", 3);
|
||||||
dump_internal(i->second, true, ensure_ascii, indent_step, new_indent);
|
dump(i->second, true, ensure_ascii, indent_step, new_indent);
|
||||||
|
|
||||||
put_char('\n');
|
o->write_character('\n');
|
||||||
put_chars(indent_string.c_str(), current_indent);
|
o->write_characters(indent_string.c_str(), current_indent);
|
||||||
put_char('}');
|
o->write_character('}');
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
put_char('{');
|
o->write_character('{');
|
||||||
|
|
||||||
// first n-1 elements
|
// first n-1 elements
|
||||||
auto i = val.m_data.m_value.object->cbegin();
|
auto i = val.m_data.m_value.object->cbegin();
|
||||||
for (std::size_t cnt = 0; cnt < val.m_data.m_value.object->size() - 1; ++cnt, ++i)
|
for (std::size_t cnt = 0; cnt < val.m_data.m_value.object->size() - 1; ++cnt, ++i)
|
||||||
{
|
{
|
||||||
put_char('\"');
|
o->write_character('\"');
|
||||||
dump_escaped(i->first, ensure_ascii);
|
dump_escaped(i->first, ensure_ascii);
|
||||||
put_chars("\":", 2);
|
o->write_characters("\":", 2);
|
||||||
dump_internal(i->second, false, ensure_ascii, indent_step, current_indent);
|
dump(i->second, false, ensure_ascii, indent_step, current_indent);
|
||||||
put_char(',');
|
o->write_character(',');
|
||||||
}
|
}
|
||||||
|
|
||||||
// last element
|
// last element
|
||||||
JSON_ASSERT(i != val.m_data.m_value.object->cend());
|
JSON_ASSERT(i != val.m_data.m_value.object->cend());
|
||||||
JSON_ASSERT(std::next(i) == val.m_data.m_value.object->cend());
|
JSON_ASSERT(std::next(i) == val.m_data.m_value.object->cend());
|
||||||
put_char('\"');
|
o->write_character('\"');
|
||||||
dump_escaped(i->first, ensure_ascii);
|
dump_escaped(i->first, ensure_ascii);
|
||||||
put_chars("\":", 2);
|
o->write_characters("\":", 2);
|
||||||
dump_internal(i->second, false, ensure_ascii, indent_step, current_indent);
|
dump(i->second, false, ensure_ascii, indent_step, current_indent);
|
||||||
|
|
||||||
put_char('}');
|
o->write_character('}');
|
||||||
}
|
}
|
||||||
|
|
||||||
return;
|
return;
|
||||||
@@ -210,13 +189,13 @@ class serializer
|
|||||||
{
|
{
|
||||||
if (val.m_data.m_value.array->empty())
|
if (val.m_data.m_value.array->empty())
|
||||||
{
|
{
|
||||||
put_chars("[]", 2);
|
o->write_characters("[]", 2);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
if (pretty_print)
|
if (pretty_print)
|
||||||
{
|
{
|
||||||
put_chars("[\n", 2);
|
o->write_characters("[\n", 2);
|
||||||
|
|
||||||
// variable to hold indentation for recursive calls
|
// variable to hold indentation for recursive calls
|
||||||
const auto new_indent = current_indent + indent_step;
|
const auto new_indent = current_indent + indent_step;
|
||||||
@@ -229,37 +208,37 @@ class serializer
|
|||||||
for (auto i = val.m_data.m_value.array->cbegin();
|
for (auto i = val.m_data.m_value.array->cbegin();
|
||||||
i != val.m_data.m_value.array->cend() - 1; ++i)
|
i != val.m_data.m_value.array->cend() - 1; ++i)
|
||||||
{
|
{
|
||||||
put_chars(indent_string.c_str(), new_indent);
|
o->write_characters(indent_string.c_str(), new_indent);
|
||||||
dump_internal(*i, true, ensure_ascii, indent_step, new_indent);
|
dump(*i, true, ensure_ascii, indent_step, new_indent);
|
||||||
put_chars(",\n", 2);
|
o->write_characters(",\n", 2);
|
||||||
}
|
}
|
||||||
|
|
||||||
// last element
|
// last element
|
||||||
JSON_ASSERT(!val.m_data.m_value.array->empty());
|
JSON_ASSERT(!val.m_data.m_value.array->empty());
|
||||||
put_chars(indent_string.c_str(), new_indent);
|
o->write_characters(indent_string.c_str(), new_indent);
|
||||||
dump_internal(val.m_data.m_value.array->back(), true, ensure_ascii, indent_step, new_indent);
|
dump(val.m_data.m_value.array->back(), true, ensure_ascii, indent_step, new_indent);
|
||||||
|
|
||||||
put_char('\n');
|
o->write_character('\n');
|
||||||
put_chars(indent_string.c_str(), current_indent);
|
o->write_characters(indent_string.c_str(), current_indent);
|
||||||
put_char(']');
|
o->write_character(']');
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
put_char('[');
|
o->write_character('[');
|
||||||
|
|
||||||
// first n-1 elements
|
// first n-1 elements
|
||||||
for (auto i = val.m_data.m_value.array->cbegin();
|
for (auto i = val.m_data.m_value.array->cbegin();
|
||||||
i != val.m_data.m_value.array->cend() - 1; ++i)
|
i != val.m_data.m_value.array->cend() - 1; ++i)
|
||||||
{
|
{
|
||||||
dump_internal(*i, false, ensure_ascii, indent_step, current_indent);
|
dump(*i, false, ensure_ascii, indent_step, current_indent);
|
||||||
put_char(',');
|
o->write_character(',');
|
||||||
}
|
}
|
||||||
|
|
||||||
// last element
|
// last element
|
||||||
JSON_ASSERT(!val.m_data.m_value.array->empty());
|
JSON_ASSERT(!val.m_data.m_value.array->empty());
|
||||||
dump_internal(val.m_data.m_value.array->back(), false, ensure_ascii, indent_step, current_indent);
|
dump(val.m_data.m_value.array->back(), false, ensure_ascii, indent_step, current_indent);
|
||||||
|
|
||||||
put_char(']');
|
o->write_character(']');
|
||||||
}
|
}
|
||||||
|
|
||||||
return;
|
return;
|
||||||
@@ -267,9 +246,9 @@ class serializer
|
|||||||
|
|
||||||
case value_t::string:
|
case value_t::string:
|
||||||
{
|
{
|
||||||
put_char('\"');
|
o->write_character('\"');
|
||||||
dump_escaped(*val.m_data.m_value.string, ensure_ascii);
|
dump_escaped(*val.m_data.m_value.string, ensure_ascii);
|
||||||
put_char('\"');
|
o->write_character('\"');
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -277,7 +256,7 @@ class serializer
|
|||||||
{
|
{
|
||||||
if (pretty_print)
|
if (pretty_print)
|
||||||
{
|
{
|
||||||
put_chars("{\n", 2);
|
o->write_characters("{\n", 2);
|
||||||
|
|
||||||
// variable to hold indentation for recursive calls
|
// variable to hold indentation for recursive calls
|
||||||
const auto new_indent = current_indent + indent_step;
|
const auto new_indent = current_indent + indent_step;
|
||||||
@@ -286,9 +265,9 @@ class serializer
|
|||||||
indent_string.resize(indent_string.size() * 2, ' ');
|
indent_string.resize(indent_string.size() * 2, ' ');
|
||||||
}
|
}
|
||||||
|
|
||||||
put_chars(indent_string.c_str(), new_indent);
|
o->write_characters(indent_string.c_str(), new_indent);
|
||||||
|
|
||||||
put_chars("\"bytes\": [", 10);
|
o->write_characters("\"bytes\": [", 10);
|
||||||
|
|
||||||
if (!val.m_data.m_value.binary->empty())
|
if (!val.m_data.m_value.binary->empty())
|
||||||
{
|
{
|
||||||
@@ -296,30 +275,30 @@ class serializer
|
|||||||
i != val.m_data.m_value.binary->cend() - 1; ++i)
|
i != val.m_data.m_value.binary->cend() - 1; ++i)
|
||||||
{
|
{
|
||||||
dump_integer(*i);
|
dump_integer(*i);
|
||||||
put_chars(", ", 2);
|
o->write_characters(", ", 2);
|
||||||
}
|
}
|
||||||
dump_integer(val.m_data.m_value.binary->back());
|
dump_integer(val.m_data.m_value.binary->back());
|
||||||
}
|
}
|
||||||
|
|
||||||
put_chars("],\n", 3);
|
o->write_characters("],\n", 3);
|
||||||
put_chars(indent_string.c_str(), new_indent);
|
o->write_characters(indent_string.c_str(), new_indent);
|
||||||
|
|
||||||
put_chars("\"subtype\": ", 11);
|
o->write_characters("\"subtype\": ", 11);
|
||||||
if (val.m_data.m_value.binary->has_subtype())
|
if (val.m_data.m_value.binary->has_subtype())
|
||||||
{
|
{
|
||||||
dump_integer(val.m_data.m_value.binary->subtype());
|
dump_integer(val.m_data.m_value.binary->subtype());
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
put_chars("null", 4);
|
o->write_characters("null", 4);
|
||||||
}
|
}
|
||||||
put_char('\n');
|
o->write_character('\n');
|
||||||
put_chars(indent_string.c_str(), current_indent);
|
o->write_characters(indent_string.c_str(), current_indent);
|
||||||
put_char('}');
|
o->write_character('}');
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
put_chars("{\"bytes\":[", 10);
|
o->write_characters("{\"bytes\":[", 10);
|
||||||
|
|
||||||
if (!val.m_data.m_value.binary->empty())
|
if (!val.m_data.m_value.binary->empty())
|
||||||
{
|
{
|
||||||
@@ -327,20 +306,20 @@ class serializer
|
|||||||
i != val.m_data.m_value.binary->cend() - 1; ++i)
|
i != val.m_data.m_value.binary->cend() - 1; ++i)
|
||||||
{
|
{
|
||||||
dump_integer(*i);
|
dump_integer(*i);
|
||||||
put_char(',');
|
o->write_character(',');
|
||||||
}
|
}
|
||||||
dump_integer(val.m_data.m_value.binary->back());
|
dump_integer(val.m_data.m_value.binary->back());
|
||||||
}
|
}
|
||||||
|
|
||||||
put_chars("],\"subtype\":", 12);
|
o->write_characters("],\"subtype\":", 12);
|
||||||
if (val.m_data.m_value.binary->has_subtype())
|
if (val.m_data.m_value.binary->has_subtype())
|
||||||
{
|
{
|
||||||
dump_integer(val.m_data.m_value.binary->subtype());
|
dump_integer(val.m_data.m_value.binary->subtype());
|
||||||
put_char('}');
|
o->write_character('}');
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
put_chars("null}", 5);
|
o->write_characters("null}", 5);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return;
|
return;
|
||||||
@@ -350,11 +329,11 @@ class serializer
|
|||||||
{
|
{
|
||||||
if (val.m_data.m_value.boolean)
|
if (val.m_data.m_value.boolean)
|
||||||
{
|
{
|
||||||
put_chars("true", 4);
|
o->write_characters("true", 4);
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
put_chars("false", 5);
|
o->write_characters("false", 5);
|
||||||
}
|
}
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
@@ -379,13 +358,13 @@ class serializer
|
|||||||
|
|
||||||
case value_t::discarded:
|
case value_t::discarded:
|
||||||
{
|
{
|
||||||
put_chars("<discarded>", 11);
|
o->write_characters("<discarded>", 11);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
case value_t::null:
|
case value_t::null:
|
||||||
{
|
{
|
||||||
put_chars("null", 4);
|
o->write_characters("null", 4);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -421,45 +400,6 @@ class serializer
|
|||||||
|
|
||||||
for (std::size_t i = 0; i < s.size(); ++i)
|
for (std::size_t i = 0; i < s.size(); ++i)
|
||||||
{
|
{
|
||||||
// Fast path: at a character boundary (state == UTF8_ACCEPT),
|
|
||||||
// bulk-copy the longest run of bytes that need no escaping using a
|
|
||||||
// SWAR scanner shared with the lexer's contiguous path. The scanner
|
|
||||||
// stops exactly at the first byte dump_escaped would handle
|
|
||||||
// individually, so that byte is left to the byte-at-a-time path
|
|
||||||
// below, keeping escaping output and error diagnostics unchanged.
|
|
||||||
//
|
|
||||||
// - ensure_ascii == false: string_bulk_run() copies ordinary bytes
|
|
||||||
// and complete well-formed UTF-8, stopping at a quote, backslash,
|
|
||||||
// control character (< 0x20), or ill-formed/truncated sequence.
|
|
||||||
// - ensure_ascii == true: only printable ASCII may be copied
|
|
||||||
// verbatim; find_ascii_copyable_run() additionally stops at 0x7F
|
|
||||||
// and every non-ASCII byte (>= 0x80), which must be \u-escaped.
|
|
||||||
if (state == UTF8_ACCEPT)
|
|
||||||
{
|
|
||||||
const auto* const data = reinterpret_cast<const unsigned char*>(s.data());
|
|
||||||
const std::size_t run = ensure_ascii
|
|
||||||
? find_ascii_copyable_run(data + i, s.size() - i)
|
|
||||||
: string_bulk_run(data + i, s.size() - i);
|
|
||||||
if (run != 0)
|
|
||||||
{
|
|
||||||
// emit any bytes still pending in string_buffer first to
|
|
||||||
// preserve output order, then write the run directly
|
|
||||||
if (bytes != 0)
|
|
||||||
{
|
|
||||||
put_chars(string_buffer.data(), bytes);
|
|
||||||
bytes = 0;
|
|
||||||
}
|
|
||||||
put_chars(s.data() + i, run);
|
|
||||||
bytes_after_last_accept = 0;
|
|
||||||
undumped_chars = 0;
|
|
||||||
i += run;
|
|
||||||
if (i >= s.size())
|
|
||||||
{
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
const auto byte = static_cast<std::uint8_t>(s[i]);
|
const auto byte = static_cast<std::uint8_t>(s[i]);
|
||||||
|
|
||||||
switch (decode(state, codepoint, byte))
|
switch (decode(state, codepoint, byte))
|
||||||
@@ -548,7 +488,7 @@ class serializer
|
|||||||
// written ("\uxxxx\uxxxx\0") for one code point
|
// written ("\uxxxx\uxxxx\0") for one code point
|
||||||
if (string_buffer.size() - bytes < 13)
|
if (string_buffer.size() - bytes < 13)
|
||||||
{
|
{
|
||||||
put_chars(string_buffer.data(), bytes);
|
o->write_characters(string_buffer.data(), bytes);
|
||||||
bytes = 0;
|
bytes = 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -607,7 +547,7 @@ class serializer
|
|||||||
// written ("\uxxxx\uxxxx\0") for one code point
|
// written ("\uxxxx\uxxxx\0") for one code point
|
||||||
if (string_buffer.size() - bytes < 13)
|
if (string_buffer.size() - bytes < 13)
|
||||||
{
|
{
|
||||||
put_chars(string_buffer.data(), bytes);
|
o->write_characters(string_buffer.data(), bytes);
|
||||||
bytes = 0;
|
bytes = 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -646,7 +586,7 @@ class serializer
|
|||||||
// write buffer
|
// write buffer
|
||||||
if (bytes > 0)
|
if (bytes > 0)
|
||||||
{
|
{
|
||||||
put_chars(string_buffer.data(), bytes);
|
o->write_characters(string_buffer.data(), bytes);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
@@ -662,22 +602,22 @@ class serializer
|
|||||||
case error_handler_t::ignore:
|
case error_handler_t::ignore:
|
||||||
{
|
{
|
||||||
// write all accepted bytes
|
// write all accepted bytes
|
||||||
put_chars(string_buffer.data(), bytes_after_last_accept);
|
o->write_characters(string_buffer.data(), bytes_after_last_accept);
|
||||||
break;
|
break;
|
||||||
}
|
}
|
||||||
|
|
||||||
case error_handler_t::replace:
|
case error_handler_t::replace:
|
||||||
{
|
{
|
||||||
// write all accepted bytes
|
// write all accepted bytes
|
||||||
put_chars(string_buffer.data(), bytes_after_last_accept);
|
o->write_characters(string_buffer.data(), bytes_after_last_accept);
|
||||||
// add a replacement character
|
// add a replacement character
|
||||||
if (ensure_ascii)
|
if (ensure_ascii)
|
||||||
{
|
{
|
||||||
put_chars("\\ufffd", 6);
|
o->write_characters("\\ufffd", 6);
|
||||||
}
|
}
|
||||||
else
|
else
|
||||||
{
|
{
|
||||||
put_chars("\xEF\xBF\xBD", 3);
|
o->write_characters("\xEF\xBF\xBD", 3);
|
||||||
}
|
}
|
||||||
break;
|
break;
|
||||||
}
|
}
|
||||||
@@ -688,66 +628,6 @@ class serializer
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
private:
|
|
||||||
/*!
|
|
||||||
@brief append a single character to the write buffer
|
|
||||||
|
|
||||||
Structural characters ('{', '"', ',', ...) previously went straight to the
|
|
||||||
output adapter, one virtual call each. Buffering them and flushing in bulk
|
|
||||||
turns those many indirect calls into a single memcpy plus an occasional
|
|
||||||
flush, which dominates the cost of serializing object/array-heavy values.
|
|
||||||
*/
|
|
||||||
void put_char(char c)
|
|
||||||
{
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(write_buffer_pos == write_buffer.size()))
|
|
||||||
{
|
|
||||||
flush();
|
|
||||||
}
|
|
||||||
write_buffer[write_buffer_pos++] = c;
|
|
||||||
}
|
|
||||||
|
|
||||||
/*!
|
|
||||||
@brief append @a length characters to the write buffer
|
|
||||||
|
|
||||||
Runs that do not fit the buffer are written straight through the output
|
|
||||||
adapter (after flushing what is pending), so large string/number payloads
|
|
||||||
are not copied an extra time.
|
|
||||||
*/
|
|
||||||
JSON_HEDLEY_NON_NULL(2)
|
|
||||||
void put_chars(const char* s, std::size_t length)
|
|
||||||
{
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(length >= write_buffer.size()))
|
|
||||||
{
|
|
||||||
flush();
|
|
||||||
o->write_characters(s, length);
|
|
||||||
return;
|
|
||||||
}
|
|
||||||
if (JSON_HEDLEY_UNLIKELY(write_buffer_pos + length > write_buffer.size()))
|
|
||||||
{
|
|
||||||
flush();
|
|
||||||
}
|
|
||||||
std::memcpy(write_buffer.data() + write_buffer_pos, s, length);
|
|
||||||
write_buffer_pos += length;
|
|
||||||
}
|
|
||||||
|
|
||||||
JSON_PRIVATE_UNLESS_TESTED:
|
|
||||||
/*!
|
|
||||||
@brief flush the write buffer to the output adapter
|
|
||||||
|
|
||||||
Writing zero characters is a well-defined no-op for every output adapter, so
|
|
||||||
the buffered length is passed through unconditionally (no empty-guard branch
|
|
||||||
to leave uncovered).
|
|
||||||
|
|
||||||
@note dump_escaped() and dump_integer()/dump_float() write into the internal
|
|
||||||
write buffer; callers that invoke them directly (rather than through the
|
|
||||||
public dump()) must call flush() before inspecting the output.
|
|
||||||
*/
|
|
||||||
void flush()
|
|
||||||
{
|
|
||||||
o->write_characters(write_buffer.data(), write_buffer_pos);
|
|
||||||
write_buffer_pos = 0;
|
|
||||||
}
|
|
||||||
|
|
||||||
private:
|
private:
|
||||||
/*!
|
/*!
|
||||||
@brief count digits
|
@brief count digits
|
||||||
@@ -872,7 +752,7 @@ class serializer
|
|||||||
// special case for "0"
|
// special case for "0"
|
||||||
if (x == 0)
|
if (x == 0)
|
||||||
{
|
{
|
||||||
put_char('0');
|
o->write_character('0');
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -925,7 +805,7 @@ class serializer
|
|||||||
*(--buffer_ptr) = static_cast<char>('0' + abs_value);
|
*(--buffer_ptr) = static_cast<char>('0' + abs_value);
|
||||||
}
|
}
|
||||||
|
|
||||||
put_chars(number_buffer.data(), n_chars);
|
o->write_characters(number_buffer.data(), n_chars);
|
||||||
}
|
}
|
||||||
|
|
||||||
/*!
|
/*!
|
||||||
@@ -941,7 +821,7 @@ class serializer
|
|||||||
// NaN / inf
|
// NaN / inf
|
||||||
if (!std::isfinite(x))
|
if (!std::isfinite(x))
|
||||||
{
|
{
|
||||||
put_chars("null", 4);
|
o->write_characters("null", 4);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -962,7 +842,7 @@ class serializer
|
|||||||
auto* begin = number_buffer.data();
|
auto* begin = number_buffer.data();
|
||||||
auto* end = ::nlohmann::detail::to_chars(begin, begin + number_buffer.size(), x);
|
auto* end = ::nlohmann::detail::to_chars(begin, begin + number_buffer.size(), x);
|
||||||
|
|
||||||
put_chars(begin, static_cast<size_t>(end - begin));
|
o->write_characters(begin, static_cast<size_t>(end - begin));
|
||||||
}
|
}
|
||||||
|
|
||||||
JSON_HEDLEY_NON_NULL(1)
|
JSON_HEDLEY_NON_NULL(1)
|
||||||
@@ -1013,7 +893,7 @@ class serializer
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
put_chars(number_buffer.data(), static_cast<std::size_t>(len));
|
o->write_characters(number_buffer.data(), static_cast<std::size_t>(len));
|
||||||
|
|
||||||
// determine if we need to append ".0"
|
// determine if we need to append ".0"
|
||||||
const bool value_is_int_like =
|
const bool value_is_int_like =
|
||||||
@@ -1025,7 +905,7 @@ class serializer
|
|||||||
|
|
||||||
if (value_is_int_like)
|
if (value_is_int_like)
|
||||||
{
|
{
|
||||||
put_chars(".0", 2);
|
o->write_characters(".0", 2);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1135,12 +1015,6 @@ class serializer
|
|||||||
|
|
||||||
/// error_handler how to react on decoding errors
|
/// error_handler how to react on decoding errors
|
||||||
const error_handler_t error_handler;
|
const error_handler_t error_handler;
|
||||||
|
|
||||||
/// buffer collecting output before it is flushed to the output adapter, so
|
|
||||||
/// that the many small structural writes become few bulk writes
|
|
||||||
std::array<char, 1024> write_buffer{{}};
|
|
||||||
/// number of valid bytes currently held in @ref write_buffer
|
|
||||||
std::size_t write_buffer_pos = 0;
|
|
||||||
};
|
};
|
||||||
|
|
||||||
} // namespace detail
|
} // namespace detail
|
||||||
|
|||||||
+222
-1084
File diff suppressed because it is too large
Load Diff
@@ -12,10 +12,6 @@
|
|||||||
#include <nlohmann/json.hpp>
|
#include <nlohmann/json.hpp>
|
||||||
using nlohmann::json;
|
using nlohmann::json;
|
||||||
|
|
||||||
#include <sstream> // stringstream
|
|
||||||
#include <string> // string
|
|
||||||
#include <vector> // vector
|
|
||||||
|
|
||||||
namespace
|
namespace
|
||||||
{
|
{
|
||||||
// shortcut to scan a string literal
|
// shortcut to scan a string literal
|
||||||
@@ -228,75 +224,3 @@ TEST_CASE("lexer class")
|
|||||||
CHECK((scan_string("/**//**//**/", true) == json::lexer::token_type::end_of_input));
|
CHECK((scan_string("/**//**//**/", true) == json::lexer::token_type::end_of_input));
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
TEST_CASE("lexer number fast path")
|
|
||||||
{
|
|
||||||
// The contiguous fast path (used for pointer/string input) must agree with
|
|
||||||
// the streaming byte path (used for std::istream) on token type, numeric
|
|
||||||
// value, and round-trip text for every well-formed number, and reject the
|
|
||||||
// same malformed numbers with the same message.
|
|
||||||
SECTION("contiguous vs streaming parity")
|
|
||||||
{
|
|
||||||
const std::vector<std::string> numbers =
|
|
||||||
{
|
|
||||||
"0", "-0", "1", "-1", "42", "-42", "10", "100", "1234567890",
|
|
||||||
"0.0", "-0.0", "3.14", "-3.14", "0.5", "-0.001", "123.456789",
|
|
||||||
"1e0", "1E0", "1e10", "1e-10", "1e+10", "1.5e3", "-2.5E-4",
|
|
||||||
"9223372036854775807", // INT64_MAX -> unsigned
|
|
||||||
"9223372036854775808", // INT64_MAX + 1 -> unsigned
|
|
||||||
"18446744073709551615", // UINT64_MAX -> unsigned
|
|
||||||
"18446744073709551616", // UINT64_MAX + 1 -> float
|
|
||||||
"-9223372036854775808", // INT64_MIN -> integer
|
|
||||||
"-9223372036854775809", // INT64_MIN - 1 -> float
|
|
||||||
"123456789012345678901234567890", // huge -> float
|
|
||||||
"0.30000000000000004", "2.2250738585072014e-308", "1e308",
|
|
||||||
// high-precision / wide-exponent values that exercise the
|
|
||||||
// std::from_chars (Eisel-Lemire) path beyond the Clinger subset
|
|
||||||
"1.7976931348623157e308", "1.2345678901234567e-250",
|
|
||||||
"9007199254740993", "5e-324", "1e-320"
|
|
||||||
};
|
|
||||||
|
|
||||||
for (const auto& n : numbers)
|
|
||||||
{
|
|
||||||
const std::string doc = "[" + n + "]";
|
|
||||||
|
|
||||||
// contiguous fast path
|
|
||||||
const json a = json::parse(doc);
|
|
||||||
// streaming byte path
|
|
||||||
std::stringstream ss(doc);
|
|
||||||
const json b = json::parse(ss);
|
|
||||||
|
|
||||||
CAPTURE(n);
|
|
||||||
CHECK(a == b);
|
|
||||||
CHECK(a.dump() == b.dump());
|
|
||||||
CHECK(a[0].type() == b[0].type());
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
SECTION("token type classification")
|
|
||||||
{
|
|
||||||
CHECK((scan_string("0") == json::lexer::token_type::value_unsigned));
|
|
||||||
CHECK((scan_string("-1") == json::lexer::token_type::value_integer));
|
|
||||||
CHECK((scan_string("1.5") == json::lexer::token_type::value_float));
|
|
||||||
CHECK((scan_string("1e5") == json::lexer::token_type::value_float));
|
|
||||||
CHECK((scan_string("18446744073709551615") == json::lexer::token_type::value_unsigned));
|
|
||||||
CHECK((scan_string("18446744073709551616") == json::lexer::token_type::value_float));
|
|
||||||
CHECK((scan_string("-9223372036854775808") == json::lexer::token_type::value_integer));
|
|
||||||
CHECK((scan_string("-9223372036854775809") == json::lexer::token_type::value_float));
|
|
||||||
}
|
|
||||||
|
|
||||||
SECTION("malformed numbers are rejected identically")
|
|
||||||
{
|
|
||||||
for (const char* bad :
|
|
||||||
{"-", "1.", "1e", "1e+", "1.2e", "01", "-01", "1..2", "1.2.3"
|
|
||||||
})
|
|
||||||
{
|
|
||||||
CAPTURE(bad);
|
|
||||||
// the contiguous fast path must decline and let the byte path report
|
|
||||||
const std::string doc = std::string("[") + bad + "]";
|
|
||||||
CHECK_FALSE(json::accept(doc));
|
|
||||||
std::stringstream ss(doc);
|
|
||||||
CHECK_FALSE(json::accept(ss));
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|||||||
@@ -100,7 +100,6 @@ void check_escaped(const char* original, const char* escaped, const bool ensure_
|
|||||||
std::stringstream ss;
|
std::stringstream ss;
|
||||||
json::serializer s(nlohmann::detail::output_adapter<char>(ss), ' ');
|
json::serializer s(nlohmann::detail::output_adapter<char>(ss), ' ');
|
||||||
s.dump_escaped(original, ensure_ascii);
|
s.dump_escaped(original, ensure_ascii);
|
||||||
s.flush(); // dump_escaped writes into the serializer's internal buffer
|
|
||||||
CHECK(ss.str() == escaped);
|
CHECK(ss.str() == escaped);
|
||||||
}
|
}
|
||||||
} // namespace
|
} // namespace
|
||||||
|
|||||||
@@ -14,10 +14,15 @@ using nlohmann::json;
|
|||||||
using namespace nlohmann::literals; // NOLINT(google-build-using-namespace)
|
using namespace nlohmann::literals; // NOLINT(google-build-using-namespace)
|
||||||
#endif
|
#endif
|
||||||
|
|
||||||
|
#include <cstddef>
|
||||||
#include <iostream>
|
#include <iostream>
|
||||||
#include <iterator>
|
#include <iterator>
|
||||||
#include <sstream>
|
#include <sstream>
|
||||||
|
#include <streambuf>
|
||||||
|
#include <string>
|
||||||
|
#include <utility>
|
||||||
#include <valarray>
|
#include <valarray>
|
||||||
|
#include <vector>
|
||||||
|
|
||||||
#if defined(_WIN32)
|
#if defined(_WIN32)
|
||||||
#define NOMINMAX
|
#define NOMINMAX
|
||||||
@@ -219,6 +224,58 @@ class proxy_iterator
|
|||||||
iterator* m_it = nullptr;
|
iterator* m_it = nullptr;
|
||||||
};
|
};
|
||||||
|
|
||||||
|
// A streambuf that keeps no get area at all and refuses every putback: with an
|
||||||
|
// empty get area, sungetc() always ends up in pbackfail(). Used to check that
|
||||||
|
// the character terminating a number is left in the input without relying on
|
||||||
|
// the streambuf being able to put a consumed character back.
|
||||||
|
class no_putback_streambuf : public std::streambuf
|
||||||
|
{
|
||||||
|
public:
|
||||||
|
explicit no_putback_streambuf(std::string s) : m_data(std::move(s)) {}
|
||||||
|
|
||||||
|
protected:
|
||||||
|
// peek at the next character without consuming it
|
||||||
|
int_type underflow() override
|
||||||
|
{
|
||||||
|
if (m_pos >= m_data.size())
|
||||||
|
{
|
||||||
|
return traits_type::eof();
|
||||||
|
}
|
||||||
|
return traits_type::to_int_type(m_data[m_pos]);
|
||||||
|
}
|
||||||
|
|
||||||
|
// consume the next character
|
||||||
|
int_type uflow() override
|
||||||
|
{
|
||||||
|
if (m_pos >= m_data.size())
|
||||||
|
{
|
||||||
|
return traits_type::eof();
|
||||||
|
}
|
||||||
|
return traits_type::to_int_type(m_data[m_pos++]);
|
||||||
|
}
|
||||||
|
|
||||||
|
int_type pbackfail(int_type /*c*/) override
|
||||||
|
{
|
||||||
|
return traits_type::eof();
|
||||||
|
}
|
||||||
|
|
||||||
|
private:
|
||||||
|
std::string m_data;
|
||||||
|
std::size_t m_pos = 0;
|
||||||
|
};
|
||||||
|
|
||||||
|
// read the characters that are left in a stream
|
||||||
|
std::string remaining(std::istream& is)
|
||||||
|
{
|
||||||
|
std::string result;
|
||||||
|
char c = 0;
|
||||||
|
while (is.get(c))
|
||||||
|
{
|
||||||
|
result += c;
|
||||||
|
}
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
// JSON_HAS_CPP_20
|
// JSON_HAS_CPP_20
|
||||||
#if defined(__cpp_char8_t)
|
#if defined(__cpp_char8_t)
|
||||||
bool check_utf8()
|
bool check_utf8()
|
||||||
@@ -1181,6 +1238,122 @@ TEST_CASE("deserialization")
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
SECTION("stream position after extraction (#5340)")
|
||||||
|
{
|
||||||
|
SECTION("a number does not consume the character that terminates it")
|
||||||
|
{
|
||||||
|
// a number is only terminated by the character following it; that
|
||||||
|
// character must be given back so the stream is positioned right
|
||||||
|
// after the value
|
||||||
|
const std::vector<std::pair<std::string, std::string>> tests =
|
||||||
|
{
|
||||||
|
{"1true", "true"},
|
||||||
|
{"1[2]", "[2]"},
|
||||||
|
{"1{}", "{}"},
|
||||||
|
{R"(1"a")", R"("a")"},
|
||||||
|
{"1 true", " true"},
|
||||||
|
{"12,", ","},
|
||||||
|
{"-0.5e3x", "x"},
|
||||||
|
{"1null", "null"}
|
||||||
|
};
|
||||||
|
|
||||||
|
for (const auto& test : tests)
|
||||||
|
{
|
||||||
|
CAPTURE(test.first);
|
||||||
|
std::istringstream ss(test.first);
|
||||||
|
json j;
|
||||||
|
ss >> j;
|
||||||
|
CHECK(j == json::parse(test.first.substr(0, test.first.size() - test.second.size())));
|
||||||
|
CHECK(remaining(ss) == test.second);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
SECTION("values that are self-delimiting are unaffected")
|
||||||
|
{
|
||||||
|
const std::vector<std::pair<std::string, std::string>> tests =
|
||||||
|
{
|
||||||
|
{"truefalse", "false"},
|
||||||
|
{"[1][2]", "[2]"},
|
||||||
|
{R"({"a":1}{"b":2})", R"({"b":2})"},
|
||||||
|
{R"("a""b")", R"("b")"},
|
||||||
|
{"null null", " null"}
|
||||||
|
};
|
||||||
|
|
||||||
|
for (const auto& test : tests)
|
||||||
|
{
|
||||||
|
CAPTURE(test.first);
|
||||||
|
std::istringstream ss(test.first);
|
||||||
|
json j;
|
||||||
|
ss >> j;
|
||||||
|
CHECK(remaining(ss) == test.second);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
SECTION("a number at the end of the input leaves nothing behind")
|
||||||
|
{
|
||||||
|
for (const std::string s :
|
||||||
|
{"1", "12", "-3.5e2", " 7 "
|
||||||
|
})
|
||||||
|
{
|
||||||
|
CAPTURE(s);
|
||||||
|
std::istringstream ss(s);
|
||||||
|
json j;
|
||||||
|
ss >> j;
|
||||||
|
CHECK(remaining(ss).find_first_not_of(" \t\n\r") == std::string::npos);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
SECTION("repeated extraction of concatenated values")
|
||||||
|
{
|
||||||
|
std::istringstream ss(R"(1true[2]3"x"{"a":4}5)");
|
||||||
|
const std::vector<json> expected =
|
||||||
|
{
|
||||||
|
json(1), json(true), json::parse("[2]"), json(3),
|
||||||
|
json("x"), json::parse(R"({"a":4})"), json(5)
|
||||||
|
};
|
||||||
|
|
||||||
|
for (const auto& e : expected)
|
||||||
|
{
|
||||||
|
json j;
|
||||||
|
ss >> j;
|
||||||
|
CHECK(j == e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
SECTION("sax_parse with strict == false")
|
||||||
|
{
|
||||||
|
std::istringstream ss("1true");
|
||||||
|
SaxEventLogger l;
|
||||||
|
CHECK(json::sax_parse(ss, &l, nlohmann::detail::input_format_t::json, false));
|
||||||
|
CHECK(l.events.size() == 1);
|
||||||
|
CHECK(l.events[0] == "number_unsigned(1)");
|
||||||
|
CHECK(remaining(ss) == "true");
|
||||||
|
}
|
||||||
|
|
||||||
|
SECTION("strict parsing still rejects trailing data")
|
||||||
|
{
|
||||||
|
std::istringstream ss("1true");
|
||||||
|
json _;
|
||||||
|
CHECK_THROWS_WITH_AS(_ = json::parse(ss),
|
||||||
|
"[json.exception.parse_error.101] parse error at line 1, column 5: syntax error while parsing value - unexpected true literal; expected end of input", json::parse_error&);
|
||||||
|
|
||||||
|
std::istringstream ss2("1true");
|
||||||
|
CHECK_FALSE(json::accept(ss2));
|
||||||
|
}
|
||||||
|
|
||||||
|
SECTION("a streambuf that cannot put back is not needed")
|
||||||
|
{
|
||||||
|
// the terminating character is never consumed, so no putback
|
||||||
|
// position is required
|
||||||
|
no_putback_streambuf buf("1true");
|
||||||
|
std::istream is(&buf);
|
||||||
|
json j;
|
||||||
|
is >> j;
|
||||||
|
CHECK(j == json(1));
|
||||||
|
CHECK(remaining(is) == "true");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// build with C++20
|
// build with C++20
|
||||||
// JSON_HAS_CPP_20
|
// JSON_HAS_CPP_20
|
||||||
#if defined(__cpp_char8_t)
|
#if defined(__cpp_char8_t)
|
||||||
|
|||||||
@@ -382,91 +382,3 @@ TEST_CASE("dump for basic_json with long double number_float_t")
|
|||||||
check_same(100.0L, 100.0);
|
check_same(100.0L, 100.0);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
TEST_CASE("serialization of strings (bulk fast path)")
|
|
||||||
{
|
|
||||||
// These cases exercise the SWAR bulk-copy fast path in dump_escaped and the
|
|
||||||
// internal write buffer: long runs, escapes interrupting runs, 0x7F/DEL,
|
|
||||||
// multibyte UTF-8 under both ensure_ascii settings, and payloads larger than
|
|
||||||
// the write buffer.
|
|
||||||
|
|
||||||
SECTION("long unescaped ASCII exceeds the write buffer")
|
|
||||||
{
|
|
||||||
const std::string big(3000, 'a');
|
|
||||||
const json j = big;
|
|
||||||
CHECK(j.dump() == '"' + big + '"');
|
|
||||||
CHECK(j.dump(-1, ' ', true) == '"' + big + '"');
|
|
||||||
// round-trips
|
|
||||||
CHECK(json::parse(j.dump()) == j);
|
|
||||||
}
|
|
||||||
|
|
||||||
SECTION("runs interrupted by escapes")
|
|
||||||
{
|
|
||||||
const json j = std::string(500, 'x') + "\n\"\\" + std::string(500, 'y');
|
|
||||||
const std::string out = j.dump();
|
|
||||||
CHECK(out == '"' + std::string(500, 'x') + "\\n\\\"\\\\" + std::string(500, 'y') + '"');
|
|
||||||
CHECK(json::parse(out) == j);
|
|
||||||
}
|
|
||||||
|
|
||||||
SECTION("DEL (0x7F) depends on ensure_ascii")
|
|
||||||
{
|
|
||||||
const json j = std::string("a\x7f" "b");
|
|
||||||
CHECK(j.dump(-1, ' ', false) == "\"a\x7f" "b\""); // copied verbatim
|
|
||||||
CHECK(j.dump(-1, ' ', true) == "\"a\\u007fb\""); // escaped
|
|
||||||
}
|
|
||||||
|
|
||||||
SECTION("multibyte UTF-8 under both ensure_ascii settings")
|
|
||||||
{
|
|
||||||
const json j = std::string("A\xc3\xa9\xe4\xbd\xa0\xf0\x9f\x98\x80Z"); // A é 你 😀 Z
|
|
||||||
// not escaping non-ASCII: bytes are copied through the bulk validator
|
|
||||||
CHECK(j.dump(-1, ' ', false) == "\"A\xc3\xa9\xe4\xbd\xa0\xf0\x9f\x98\x80Z\"");
|
|
||||||
// ensure_ascii: escaped (with a surrogate pair for the emoji)
|
|
||||||
CHECK(j.dump(-1, ' ', true) == "\"A\\u00e9\\u4f60\\ud83d\\ude00Z\"");
|
|
||||||
CHECK(json::parse(j.dump(-1, ' ', true)) == j);
|
|
||||||
}
|
|
||||||
|
|
||||||
SECTION("many small structural writes exceed the write buffer")
|
|
||||||
{
|
|
||||||
json arr = json::array();
|
|
||||||
for (int i = 0; i < 2000; ++i)
|
|
||||||
{
|
|
||||||
arr.push_back(i);
|
|
||||||
}
|
|
||||||
const std::string out = arr.dump();
|
|
||||||
CHECK(out.front() == '[');
|
|
||||||
CHECK(out.back() == ']');
|
|
||||||
CHECK(json::parse(out) == arr);
|
|
||||||
|
|
||||||
json obj = json::object();
|
|
||||||
for (int i = 0; i < 500; ++i)
|
|
||||||
{
|
|
||||||
obj["key" + std::to_string(i)] = i;
|
|
||||||
}
|
|
||||||
CHECK(json::parse(obj.dump()) == obj);
|
|
||||||
CHECK(json::parse(obj.dump(2)) == obj);
|
|
||||||
|
|
||||||
// an array of many empty strings emits a long run of single-character
|
|
||||||
// writes ('"', '"', ',') at shallow nesting depth, so the write buffer
|
|
||||||
// fills and flushes mid-run without the deep recursion that would
|
|
||||||
// overflow the stack on some debug builds
|
|
||||||
json many_empty = json::array();
|
|
||||||
for (int i = 0; i < 500; ++i)
|
|
||||||
{
|
|
||||||
many_empty.push_back("");
|
|
||||||
}
|
|
||||||
const std::string out2 = many_empty.dump();
|
|
||||||
CHECK(out2.size() > 1024); // spans multiple write-buffer flushes
|
|
||||||
CHECK(out2.front() == '[');
|
|
||||||
CHECK(out2.back() == ']');
|
|
||||||
CHECK(json::parse(out2) == many_empty);
|
|
||||||
}
|
|
||||||
|
|
||||||
SECTION("invalid UTF-8 handling is unaffected by the fast path")
|
|
||||||
{
|
|
||||||
const json j = std::string("valid\xff" "more");
|
|
||||||
CHECK_THROWS_WITH_AS(j.dump(), "[json.exception.type_error.316] invalid UTF-8 byte at index 5: 0xFF", json::type_error&);
|
|
||||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::replace) == "\"valid\xef\xbf\xbd" "more\"");
|
|
||||||
CHECK(j.dump(-1, ' ', true, json::error_handler_t::replace) == "\"valid\\ufffdmore\"");
|
|
||||||
CHECK(j.dump(-1, ' ', false, json::error_handler_t::ignore) == "\"validmore\"");
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|||||||
Reference in New Issue
Block a user