Compare commits

..
Author SHA1 Message Date
Niels Lohmann 93fe62fc0b Assert on missing array indices in const operator[] and document the JSON pointer case
The const operator[] overloads are unchecked by design, and a missing key
or index is undefined behavior. The key overload guards this with a
runtime assertion, but the index overload did not, although the element
access documentation says an assertion fires in both cases. The const
JSON pointer overload inherits both through json_pointer::get_unchecked(),
so a pointer to a missing array index read out of bounds even in debug
builds, and its documentation promised out_of_range.404 for any pointer
that cannot be resolved.

Add JSON_ASSERT(idx < size()) to const operator[](size_type), which also
covers the index leg of the const JSON pointer overload. Document the
undefined behavior for the const JSON pointer overload in operator[].md
and in the runtime assertions page. Release builds are unchanged.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-27 23:17:42 +02:00
22 changed files with 276 additions and 1610 deletions
+2 -2
View File
@@ -80,8 +80,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from CBOR were
used in the given input or if the input is not valid CBOR
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key,
but not found
## Complexity
@@ -73,8 +73,8 @@ Strong guarantee: if an exception is thrown, there are no changes in the JSON va
the end of the file was not reached when `strict` was set to true
- Throws [parse_error.112](../../home/exceptions.md#jsonexceptionparse_error112) if unsupported features from
MessagePack were used in the given input or if the input is not valid MessagePack
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a map key is not a string (keys of other
types are not supported, as JSON object keys are always strings) or a string is malformed
- Throws [parse_error.113](../../home/exceptions.md#jsonexceptionparse_error113) if a string was expected as a map key,
but not found
## Complexity
+11 -3
View File
@@ -86,6 +86,9 @@ Strong exception safety: if an exception occurs, the original value stays intact
- Throws [`out_of_range.410`](../../home/exceptions.md#jsonexceptionout_of_range410) if an array index in the passed
JSON pointer `ptr` exceeds the range of `size_type` (e.g., on 32-bit platforms).
For the **const** version, an object key or array index in `ptr` that does not exist is not reported by an
exception, but is undefined behavior (see the notes below). Use [`at`](at.md) for checked access.
## Complexity
1. Constant if `idx` is in the range of the array. Otherwise, linear in `idx - size()`.
@@ -100,9 +103,12 @@ Strong exception safety: if an exception occurs, the original value stays intact
The following cases apply to the **const** overloads; the non-const overloads instead insert the missing element
(see the notes below).
1. If the element at index `idx` does not exist, the behavior is undefined.
1. If the element at index `idx` does not exist, the behavior is undefined and is **guarded by a
[runtime assertion](../../features/assertions.md)**!
2. If the element with key `key` does not exist, the behavior is undefined and is **guarded by a
[runtime assertion](../../features/assertions.md)**!
3. If the JSON pointer `ptr` refers to an object key or an array index that does not exist, the behavior is
undefined and is **guarded by a [runtime assertion](../../features/assertions.md)**!
1. The non-const version may add values: If `idx` is beyond the range of the array (i.e., `idx >= size()`), then the
array is silently filled up with `#!json null` values to make `idx` a valid reference to the last stored element. In
@@ -257,9 +263,11 @@ Strong exception safety: if an exception occurs, the original value stays intact
## Version history
1. Added in version 1.0.0.
1. Added in version 1.0.0. A missing index in the const version is guarded by a runtime assertion since version
3.13.0.
2. Added in version 1.0.0. Added overloads for `T* key` in version 1.1.0. Removed overloads for `T* key` (replaced by 3)
in version 3.11.0.
3. Added in version 3.11.0. Fixed in version 3.13.0 to consistently accept `std::string_view`-convertible keys, as
already supported by [`at`](at.md), [`value`](value.md), [`find`](find.md), and other lookup functions.
4. Added in version 2.0.0.
4. Added in version 2.0.0. A missing array index in the const version is guarded by a runtime assertion since
version 3.13.0.
+31 -7
View File
@@ -16,14 +16,15 @@ before including the `json.hpp` header.
## Function with runtime assertions
### Unchecked object access to a const value
### Unchecked access to a const value
Function [`operator[]`](../api/basic_json/operator%5B%5D.md) implements unchecked access for objects. Whereas a missing
key is added in the case of non-const objects, accessing a const object with a missing key is undefined behavior (think
of a dereferenced null pointer) and yields a runtime assertion.
Function [`operator[]`](../api/basic_json/operator%5B%5D.md) implements unchecked access for arrays and objects. Whereas
a missing element is added in the case of non-const values, accessing a const value with a missing object key or an
invalid array index is undefined behavior (think of a dereferenced null pointer) and yields a runtime assertion. This
also applies to a [JSON pointer](json_pointer.md) that refers to a missing key or an invalid index.
If you are not sure whether an element in an object exists, use checked access with the
[`at` function](../api/basic_json/at.md) or call the [`contains` function](../api/basic_json/contains.md) before.
If you are not sure whether an element exists, use checked access with the [`at` function](../api/basic_json/at.md)
or call the [`contains` function](../api/basic_json/contains.md) before.
See also the documentation on [element access](element_access/index.md).
@@ -46,7 +47,30 @@ See also the documentation on [element access](element_access/index.md).
Output:
```
Assertion failed: (m_value.object->find(key) != m_value.object->end()), function operator[], file json.hpp, line 2144.
Assertion failed: (it != m_data.m_value.object->end()), function operator[], file json.hpp, line 28795.
```
??? example "Example 2: Invalid array index in a JSON pointer"
The following code will trigger an assertion at runtime:
```cpp
#include <nlohmann/json.hpp>
using json = nlohmann::json;
using namespace nlohmann::literals;
int main()
{
const json j = {{"array", {1, 2, 3}}};
auto v = j["/array/5"_json_pointer];
}
```
Output:
```
Assertion failed: (idx < m_data.m_value.array->size()), function operator[], file json.hpp, line 28758.
```
### Constructing from an uninitialized iterator range
@@ -174,20 +174,7 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Object keys"
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps
with keys other than text strings (major type 3) are rejected with a
[`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with `allow_exceptions` set
to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. This is a deliberate restriction of the library's JSON value model, not an oversight: formats built on CBOR
maps with integer keys, such as COSE ([RFC 9052](https://www.rfc-editor.org/rfc/rfc9052.html)) or CWT
([RFC 8392](https://www.rfc-editor.org/rfc/rfc8392.html)), cannot be read with this library and need a
general-purpose CBOR library instead.
CBOR allows map keys of any type, whereas JSON only allows strings as keys in object values. Therefore, CBOR maps with keys other than UTF-8 strings are rejected.
!!! warning "UTF-8 validation of text strings"
@@ -138,21 +138,6 @@ The library maps MessagePack types to JSON value types as follows:
Any MessagePack output created by `to_msgpack` can be successfully parsed by `from_msgpack`.
!!! warning "Object keys"
MessagePack allows map keys of any type, whereas JSON only allows strings as keys in object values. Like the
JSON-compatible [profile](https://github.com/msgpack/msgpack/blob/master/spec.md#profile) sketched in the
MessagePack specification, this library restricts map keys to `str` values. Maps with keys of any other type are
rejected with a [`parse_error.113`](../../home/exceptions.md#jsonexceptionparse_error113) exception (or, with
`allow_exceptions` set to `false`, a discarded value) naming the type of the key that was found, for instance:
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
This applies to the [SAX interface](../parsing/sax_interface.md) as well, as the key is read before it is passed
on. Such input needs a general-purpose MessagePack library instead.
!!! warning "UTF-8 validation of string values"
The MessagePack specification requires `str` values (`fixstr`, `str 8`, `str 16`, `str 32`) to be valid UTF-8.
+2 -9
View File
@@ -343,20 +343,13 @@ A string could not be read from a [binary format](../features/binary_formats/ind
string was read where one was required (for instance as a map key), the string's length specification is invalid, or
the string's bytes are not valid UTF-8.
CBOR and MessagePack allow map keys of any type, but JSON object keys are always strings. Maps with keys of any other
type (for instance integers or `null`) are therefore not supported; see the notes on
[CBOR](../features/binary_formats/cbor.md) and [MessagePack](../features/binary_formats/messagepack.md).
!!! failure "Example messages"
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an unsigned integer; last byte: 0x01
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found nil; last byte: 0xC0
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF
```
```
[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON char: byte after 'C' must be in range 0x00..0x7F; last byte: 0x82
+2 -168
View File
@@ -1324,80 +1324,6 @@ class binary_reader
}
}
/*!
@brief reads a CBOR object key
RFC 8949 allows any data item as a map key, but only strings have a
counterpart in JSON. A key of any other type is rejected with a message
naming that type, rather than the one @ref get_cbor_string gives for a
malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_cbor_object_key(string_t& result)
{
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
}
const char* found = nullptr;
switch (static_cast<unsigned int>(current) >> 5u)
{
case 0:
found = "an unsigned integer";
break;
case 1:
found = "a negative integer";
break;
case 2:
found = "a byte string";
break;
case 4:
found = "an array";
break;
case 5:
found = "a map";
break;
case 6:
found = "a tag";
break;
default: // major type 7
switch (current)
{
case 0xF4:
case 0xF5:
found = "a boolean";
break;
case 0xF6:
found = "null";
break;
case 0xF7:
found = "undefined";
break;
case 0xF9:
case 0xFA:
case 0xFB:
found = "a floating-point number";
break;
case 0xFF:
found = "a break stop code";
break;
default:
found = "a simple value";
break;
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::cbor, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a definite-length CBOR byte array
@@ -1642,7 +1568,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_cbor_object_key(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string(key) || !sax->key(key)))
{
return false;
}
@@ -2143,98 +2069,6 @@ class binary_reader
}
}
/*!
@brief reads a MessagePack object key
The MessagePack specification allows any type as a map key, but only
strings have a counterpart in JSON. A key of any other type is rejected
with a message naming that type, rather than the one @ref
get_msgpack_string gives for a malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_msgpack_object_key(string_t& result)
{
const char* found = nullptr;
switch (current)
{
case 0xC0:
found = "nil";
break;
case 0xC2:
case 0xC3:
found = "a boolean";
break;
case 0xCA:
case 0xCB:
found = "a float";
break;
case 0xC4:
case 0xC5:
case 0xC6:
found = "a bin";
break;
case 0xC7:
case 0xC8:
case 0xC9:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
found = "an ext";
break;
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
found = "an integer";
break;
case 0xDC:
case 0xDD:
found = "an array";
break;
case 0xDE:
case 0xDF:
found = "a map";
break;
default:
// fixint, fixmap, and fixarray; strings, EOF, and the unused
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
}
if (current <= 0x7F || current >= 0xE0)
{
found = "an integer";
}
else if (current <= 0x8F)
{
found = "a map";
}
else if (current <= 0x9F)
{
found = "an array";
}
else
{
return get_msgpack_string(result);
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::msgpack, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a MessagePack byte array
@@ -2397,7 +2231,7 @@ class binary_reader
{
get();
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_object_key(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
{
return false;
}
+39 -76
View File
@@ -206,6 +206,7 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter))
, ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_)
{}
@@ -221,7 +222,8 @@ class lexer : public lexer_base<BasicJsonType>
// locales
/////////////////////
/// return the decimal point of the current locale
/// return the locale-dependent decimal point
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept
{
const auto* loc = localeconv();
@@ -1090,10 +1092,9 @@ class lexer : public lexer_base<BasicJsonType>
token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise
@note The scanner is independent of the current locale: token_buffer
always holds `.`. Only the std::strtod fallback of convert_number()
depends on the locale, and it looks up the decimal point right
before converting (see convert_float_locale_aware()).
@note The scanner is independent of the current locale. Internally, the
locale's decimal point is used instead of `.` to work with the
locale-dependent converters.
*/
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{
@@ -1182,7 +1183,7 @@ scan_number_zero:
{
case '.':
{
add(current);
add(decimal_point_char);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -1219,7 +1220,7 @@ scan_number_any1:
case '.':
{
add(current);
add(decimal_point_char);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -1461,9 +1462,9 @@ scan_number_done:
// Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. The
// fraction is located through decimal_point_position rather than by
// searching '.'.
// answer - so this loop is skipped for all but a few tokens. Note
// token_buffer holds the locale's decimal point, so the fraction is
// located through decimal_point_position rather than by searching '.'.
if (lead_zero != 0)
{
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
@@ -1481,8 +1482,8 @@ scan_number_done:
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds '.'
as decimal point, independent of the locale. Integers are parsed first and fall
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
@@ -1562,7 +1563,7 @@ scan_number_done:
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod/strtold.
// locale-aware strtof/strtod.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
@@ -1571,75 +1572,26 @@ scan_number_done:
// extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, value_float))
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
{
return token_type::value_float;
}
convert_float_locale_aware();
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
return token_type::value_float;
}
/*!
@brief convert the float in token_buffer with strtof/strtod/strtold
These functions expect the decimal point of the *current* locale, so it is
looked up right before the conversion instead of once when the lexer is
constructed: a locale change in between (by a parser callback, a SAX
handler, or another thread) must not truncate the value (#5198). The
token has been validated before, so if the conversion stops early and the
decimal point changed in the meantime, the locale changed between the
lookup and the call, and the conversion is repeated with the new decimal
point. If the decimal point did not change, a retry cannot succeed: the
locale's decimal point is not a single character (e.g., the two-byte
U+066B of ar_EG.UTF-8 or fa_IR.UTF-8) and cannot be substituted in place.
The value strtod parsed up to that point is kept, as before this change.
Note that changing the locale in another thread *while* strtod runs is
undefined behavior of the C library, which this function cannot prevent.
*/
void convert_float_locale_aware()
{
const bool has_dot = decimal_point_position != std::string::npos;
char decimal_point = get_decimal_point();
for (;;)
{
const bool substitute = has_dot && decimal_point != '.';
if (substitute)
{
token_buffer[decimal_point_position] = static_cast<typename string_t::value_type>(decimal_point);
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
if (substitute)
{
// get_string() hands the token to the SAX interface with '.'
token_buffer[decimal_point_position] = '.';
}
if (JSON_HEDLEY_LIKELY(endptr == token_buffer.data() + token_buffer.size()))
{
return;
}
// retry only if the locale changed; otherwise, this would loop forever
const char current_decimal_point = get_decimal_point();
if (current_decimal_point == decimal_point)
{
return;
}
decimal_point = current_decimal_point;
}
}
/*!
@brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(as scan_number() does) and
(with the locale decimal point substituted, as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
@@ -1755,11 +1707,16 @@ scan_number_done:
}
#endif
// materialize the token exactly as scan_number() would. reset() already
// cleared token_buffer, so append() fills it (assign() is avoided
// because custom string_t types need not provide it)
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
decimal_point_position = dot_index;
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
@@ -2026,7 +1983,11 @@ scan_number_done:
/// return current string value (implicitly resets the token; useful only once)
string_t& get_string()
{
// a number token holds '.' regardless of the locale (#4084)
// translate decimal points from locale back to '.' (#4084)
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
{
token_buffer[decimal_point_position] = '.';
}
return token_buffer;
}
@@ -2322,7 +2283,9 @@ scan_number_done:
number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0;
/// the position of the decimal point in token_buffer
/// the decimal point
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
+13 -8
View File
@@ -118,12 +118,14 @@ std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[out] out the parsed value on success
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
@@ -134,6 +136,7 @@ inline bool parse_float_fast(const char* first, const char* last, double& out) n
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
@@ -172,7 +175,7 @@ inline bool parse_float_fast(const char* first, const char* last, double& out) n
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (c == '.')
else if (static_cast<DecimalPointType>(c) == decimal_point)
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
@@ -257,8 +260,8 @@ inline bool parse_float_fast(const char* first, const char* last, double& out) n
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
{
return false;
}
@@ -270,7 +273,9 @@ std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
+8 -2
View File
@@ -535,6 +535,10 @@ class json_pointer
@return const reference to the JSON value pointed to by the JSON
pointer
@pre Every object key and array index the pointer refers to exists.
Like the const operator[] for keys and indices, a missing one is
undefined behavior, guarded by a runtime assertion.
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if an array index was not a number
@throw out_of_range.402 if the array index '-' is used
@@ -549,7 +553,8 @@ class json_pointer
{
case detail::value_t::object:
{
// use unchecked object access
// use unchecked object access; the const operator[]
// asserts that the key exists
ptr = &ptr->operator[](reference_token);
break;
}
@@ -562,7 +567,8 @@ class json_pointer
JSON_THROW(detail::out_of_range::create(402, detail::concat("array index '-' (", std::to_string(ptr->m_data.m_value.array->size()), ") is out of range"), ptr));
}
// use unchecked array access
// use unchecked array access; the const operator[]
// asserts that the index exists
ptr = &ptr->operator[](array_index<BasicJsonType>(reference_token));
break;
}
+48 -226
View File
@@ -908,11 +908,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/*!
@brief how many levels the operation going on in this thread has descended into
Copying a value, converting one from another specialization, and comparing
two values share this count. The library never nests one of them inside
another - none of them does either of the other two on the way - and where
user code nests them anyway, sharing the count only ends a descent sooner
than it had to, which costs a little speed and is never wrong.
Copying a value and comparing two values share this count. The library never
nests one inside the other - copying a value does not compare one, and
comparing two values does not copy them - and where user code nests them
anyway, sharing the count only ends a descent sooner than it had to, which
costs a little speed and is never wrong.
A byte is enough: the count never exceeds the limit by more than the single
level that notices the limit has been reached.
@@ -1268,221 +1268,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
copy_iteratively(src);
}
/*!
@brief convert the value @a val of another specialization into this null
value; @a val must be neither an object nor an array
Converting such a value never descends, so both ways of converting an
object or an array (@ref convert_structured) leave their elements of this
kind to the converting constructor, which leaves them to this.
*/
template<typename BasicJsonType>
void convert_leaf(const BasicJsonType& val)
{
using other_boolean_t = typename BasicJsonType::boolean_t;
using other_number_float_t = typename BasicJsonType::number_float_t;
using other_number_integer_t = typename BasicJsonType::number_integer_t;
using other_number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using other_string_t = typename BasicJsonType::string_t;
using other_binary_t = typename BasicJsonType::binary_t;
switch (val.type())
{
case value_t::boolean:
JSONSerializer<other_boolean_t>::to_json(*this, val.template get<other_boolean_t>());
break;
case value_t::number_float:
JSONSerializer<other_number_float_t>::to_json(*this, val.template get<other_number_float_t>());
break;
case value_t::number_integer:
JSONSerializer<other_number_integer_t>::to_json(*this, val.template get<other_number_integer_t>());
break;
case value_t::number_unsigned:
JSONSerializer<other_number_unsigned_t>::to_json(*this, val.template get<other_number_unsigned_t>());
break;
case value_t::string:
JSONSerializer<other_string_t>::to_json(*this, val.template get_ref<const other_string_t&>());
break;
case value_t::binary:
JSONSerializer<other_binary_t>::to_json(*this, val.template get_ref<const other_binary_t&>());
break;
case value_t::null:
// this value is null already; assigning null to it would also
// reset its positions (JSON_DIAGNOSTIC_POSITIONS)
break;
case value_t::discarded:
m_data.m_type = value_t::discarded;
break;
case value_t::object: // LCOV_EXCL_LINE
case value_t::array: // LCOV_EXCL_LINE
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
}
/// scratch space for the converted elements of the arrays that
/// @ref convert_iteratively has yet to create
using convert_scratch_t = std::vector<basic_json, AllocatorType<basic_json>>;
/*!
@brief create the object or array @a val converted into this null value
Its converted elements are the last `val.size()` entries of @a elements (an
array) or of @a members (an object); they are moved into the container in
one go and then removed.
*/
template<typename BasicJsonType>
void convert_level(const BasicJsonType& val, convert_scratch_t& elements, copy_scratch_t& members)
{
if (val.is_object())
{
const auto first = members.end() - static_cast<typename copy_scratch_t::difference_type>(val.size());
m_data.m_value.object = create<object_t>(std::make_move_iterator(first),
std::make_move_iterator(members.end()));
// only now that the object exists may this stop being a null value
m_data.m_type = value_t::object;
members.erase(first, members.end());
}
else
{
const auto first = elements.end() - static_cast<typename convert_scratch_t::difference_type>(val.size());
m_data.m_value.array = create<array_t>(std::make_move_iterator(first),
std::make_move_iterator(elements.end()));
// only now that the array exists may this stop being a null value
m_data.m_type = value_t::array;
elements.erase(first, elements.end());
}
set_parents();
}
/*!
@brief convert the object or array @a val of another specialization into
this null value without recursing
The containers whose conversion has begun are kept on an explicit stack
rather than on the call stack. Unlike @ref copy_iteratively, this builds
every container from the bottom up: all its elements are converted first,
and the container is then created from them in one go, the way the range
constructor that converts the levels above the bound does. The two object
types need not enumerate their members in the same order, so the members
could not be paired up by position anyway, and building from a range keeps
what the range constructor does with keys that become equal on conversion.
Every value is complete before it is handed on, and a container gets its
type only once it exists, so whatever throws, every value left behind can
be destroyed.
*/
template<typename BasicJsonType>
void convert_iteratively(const BasicJsonType& val)
{
using other_const_iterator = typename BasicJsonType::const_iterator;
// the containers whose conversion has begun, innermost last, each with
// its element to convert next
std::vector<std::pair<const BasicJsonType*, other_const_iterator>> pending;
// the converted elements of the pending arrays and the converted
// members of the pending objects, those of the innermost one last
convert_scratch_t elements;
copy_scratch_t members;
pending.emplace_back(&val, val.cbegin());
for (;;)
{
const BasicJsonType& container = *pending.back().first;
other_const_iterator& next = pending.back().second;
if (next != container.cend())
{
if (next->is_structured())
{
// convert its elements first; next stays where it is until
// the converted container is handed back to this one
pending.emplace_back(&*next, next->cbegin());
continue;
}
// the converting constructor does not descend into this value
if (container.is_object())
{
members.emplace_back(next.key(), *next);
}
else
{
elements.emplace_back(*next);
}
++next;
continue;
}
// all elements of the container are converted: create it
pending.pop_back();
if (pending.empty())
{
convert_level(container, elements, members);
return;
}
basic_json converted;
converted.convert_level(container, elements, members);
#if JSON_DIAGNOSTIC_POSITIONS
converted.start_position = container.start_pos();
converted.end_position = container.end_pos();
#endif
// hand it to the container it is an element of
if (pending.back().first->is_object())
{
members.emplace_back(pending.back().second.key(), std::move(converted));
}
else
{
elements.push_back(std::move(converted));
}
++pending.back().second;
}
}
/*!
@brief convert the object or array @a val of another specialization into
this null value
Converting a container converts its elements, so a value nested deeply
enough used to exhaust the call stack. The descent is bounded here as in
@ref copy_structured: the first @ref nesting_depth_limit levels are
converted by the containers' range constructors, just as they always were,
and anything below that is converted without the call stack by
@ref convert_iteratively.
@sa https://github.com/nlohmann/json/issues/5650
*/
template<typename BasicJsonType>
void convert_structured(const BasicJsonType& val)
{
const nesting_depth_guard guard;
if (JSON_HEDLEY_LIKELY(guard.okay()))
{
// every element comes back to the converting constructor
if (val.is_object())
{
using other_object_t = typename BasicJsonType::object_t;
JSONSerializer<other_object_t>::to_json(*this, val.template get_ref<const other_object_t&>());
}
else
{
using other_array_t = typename BasicJsonType::array_t;
JSONSerializer<other_array_t>::to_json(*this, val.template get_ref<const other_array_t&>());
}
return;
}
convert_iteratively(val);
}
/// the result of comparing two values, including values that cannot be
/// ordered at all, such as a discarded value or a NaN
@@ -1831,13 +1616,49 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
end_position(val.end_pos())
#endif
{
if (val.is_structured())
using other_boolean_t = typename BasicJsonType::boolean_t;
using other_number_float_t = typename BasicJsonType::number_float_t;
using other_number_integer_t = typename BasicJsonType::number_integer_t;
using other_number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using other_string_t = typename BasicJsonType::string_t;
using other_object_t = typename BasicJsonType::object_t;
using other_array_t = typename BasicJsonType::array_t;
using other_binary_t = typename BasicJsonType::binary_t;
switch (val.type())
{
convert_structured(val);
}
else
{
convert_leaf(val);
case value_t::boolean:
JSONSerializer<other_boolean_t>::to_json(*this, val.template get<other_boolean_t>());
break;
case value_t::number_float:
JSONSerializer<other_number_float_t>::to_json(*this, val.template get<other_number_float_t>());
break;
case value_t::number_integer:
JSONSerializer<other_number_integer_t>::to_json(*this, val.template get<other_number_integer_t>());
break;
case value_t::number_unsigned:
JSONSerializer<other_number_unsigned_t>::to_json(*this, val.template get<other_number_unsigned_t>());
break;
case value_t::string:
JSONSerializer<other_string_t>::to_json(*this, val.template get_ref<const other_string_t&>());
break;
case value_t::object:
JSONSerializer<other_object_t>::to_json(*this, val.template get_ref<const other_object_t&>());
break;
case value_t::array:
JSONSerializer<other_array_t>::to_json(*this, val.template get_ref<const other_array_t&>());
break;
case value_t::binary:
JSONSerializer<other_binary_t>::to_json(*this, val.template get_ref<const other_binary_t&>());
break;
case value_t::null:
*this = nullptr;
break;
case value_t::discarded:
m_data.m_type = value_t::discarded;
break;
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
JSON_ASSERT(m_data.m_type == val.type());
@@ -3045,6 +2866,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// const operator[] only works for arrays
if (JSON_HEDLEY_LIKELY(is_array()))
{
JSON_ASSERT(idx < m_data.m_value.array->size());
return m_data.m_value.array->operator[](idx);
}
+110 -480
View File
@@ -8605,12 +8605,14 @@ std::strtod. The parser only activates for number_float_t == double; float and
long double keep the std::strtof/std::strtold paths (see the templated overload
below).
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[out] out the parsed value on success
@param[in] first pointer to the first character of the number
@param[in] last pointer past the last character
@param[in] decimal_point the (locale-dependent) decimal point character
@param[out] out the parsed value on success
@return true if the value was parsed exactly; false to fall back to strtod
*/
inline bool parse_float_fast(const char* first, const char* last, double& out) noexcept
template<typename DecimalPointType>
bool parse_float_fast(const char* first, const char* last, DecimalPointType decimal_point, double& out) noexcept
{
#if defined(FLT_EVAL_METHOD) && FLT_EVAL_METHOD != 0
// Clinger's fast path is only exact when double operations are evaluated in
@@ -8621,6 +8623,7 @@ inline bool parse_float_fast(const char* first, const char* last, double& out) n
// std::from_chars / std::strtod path.
static_cast<void>(first);
static_cast<void>(last);
static_cast<void>(decimal_point);
static_cast<void>(out);
return false;
#else
@@ -8659,7 +8662,7 @@ inline bool parse_float_fast(const char* first, const char* last, double& out) n
++num_digits;
fractional_digits += static_cast<int>(seen_dot);
}
else if (c == '.')
else if (static_cast<DecimalPointType>(c) == decimal_point)
{
if (JSON_HEDLEY_UNLIKELY(seen_dot))
{
@@ -8744,8 +8747,8 @@ inline bool parse_float_fast(const char* first, const char* last, double& out) n
}
/// fast float path is only exact for `double`; decline for float/long double
template<typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, FloatType& /*out*/) noexcept
template<typename DecimalPointType, typename FloatType>
bool parse_float_fast(const char* /*first*/, const char* /*last*/, DecimalPointType /*decimal_point*/, FloatType& /*out*/) noexcept
{
return false;
}
@@ -8757,7 +8760,9 @@ std::from_chars is locale-independent, correctly rounded, and - via the
Eisel-Lemire algorithm in modern standard libraries - much faster than strtod
over the whole value range (not just the Clinger subset). It is used only when
__cpp_lib_to_chars indicates full floating-point support and only when it
consumes the entire token ([first, last)). An under-/overflow (result_out_of_range) also declines, so
consumes the entire token ([first, last)); a partial parse means the buffer
uses a non-'.' locale decimal point, in which case the caller falls back to the
locale-aware path. An under-/overflow (result_out_of_range) also declines, so
the caller's strtod fallback supplies the well-defined ±inf/0 result the parser
expects (side-stepping the P4168 divergence between implementations).
@@ -9298,6 +9303,7 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false, bool discard_number_values_ = false) noexcept
: ia(std::move(adapter))
, ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
, discard_number_values(discard_number_values_)
{}
@@ -9313,7 +9319,8 @@ class lexer : public lexer_base<BasicJsonType>
// locales
/////////////////////
/// return the decimal point of the current locale
/// return the locale-dependent decimal point
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept
{
const auto* loc = localeconv();
@@ -10182,10 +10189,9 @@ class lexer : public lexer_base<BasicJsonType>
token_type::value_float if number could be successfully scanned,
token_type::parse_error otherwise
@note The scanner is independent of the current locale: token_buffer
always holds `.`. Only the std::strtod fallback of convert_number()
depends on the locale, and it looks up the decimal point right
before converting (see convert_float_locale_aware()).
@note The scanner is independent of the current locale. Internally, the
locale's decimal point is used instead of `.` to work with the
locale-dependent converters.
*/
token_type scan_number() // lgtm [cpp/use-of-goto] `goto` is used in this function to implement the number-parsing state machine described above. By design, any finite input will eventually reach the "done" state or return token_type::parse_error. In each intermediate state, 1 byte of the input is appended to the token_buffer vector, and only the already initialized variables token_buffer, number_type, and error_message are manipulated.
{
@@ -10274,7 +10280,7 @@ scan_number_zero:
{
case '.':
{
add(current);
add(decimal_point_char);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -10311,7 +10317,7 @@ scan_number_any1:
case '.':
{
add(current);
add(decimal_point_char);
decimal_point_position = token_buffer.size() - 1;
goto scan_number_decimal1;
}
@@ -10553,9 +10559,9 @@ scan_number_done:
// Only a number below 1 can carry further insignificant zeros, and only
// while the count stays at the limit does removing them change the
// answer - so this loop is skipped for all but a few tokens. The
// fraction is located through decimal_point_position rather than by
// searching '.'.
// answer - so this loop is skipped for all but a few tokens. Note
// token_buffer holds the locale's decimal point, so the fraction is
// located through decimal_point_position rather than by searching '.'.
if (lead_zero != 0)
{
JSON_ASSERT(has_dot != 0); // an integer "0" cannot reach the limit
@@ -10573,8 +10579,8 @@ scan_number_done:
@brief convert the number text in token_buffer to its value and token type
The digit sequence in token_buffer has already been validated (by the
scan_number() state machine or by the contiguous fast path) and holds '.'
as decimal point, independent of the locale. Integers are parsed first and fall
scan_number() state machine or by the contiguous fast path) and holds the
locale decimal point in place of '.'. Integers are parsed first and fall
back to floating point on overflow. This is shared so both scanners produce
identical results.
@@ -10654,7 +10660,7 @@ scan_number_done:
// integer conversion above overflowed. Prefer std::from_chars
// (Eisel-Lemire, locale-independent, correctly rounded) when available;
// otherwise the exact Clinger fast path (double only); otherwise the
// locale-aware strtof/strtod/strtold.
// locale-aware strtof/strtod.
if (parse_float_from_chars(num_begin, num_end, value_float))
{
return token_type::value_float;
@@ -10663,75 +10669,26 @@ scan_number_done:
// extra pass over the token's bytes, which otherwise shows up on
// high-precision inputs such as canada.json
if (mantissa_fits_clinger(mantissa_end)
&& parse_float_fast(num_begin, num_end, value_float))
&& parse_float_fast(num_begin, num_end, decimal_point_char, value_float))
{
return token_type::value_float;
}
convert_float_locale_aware();
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
return token_type::value_float;
}
/*!
@brief convert the float in token_buffer with strtof/strtod/strtold
These functions expect the decimal point of the *current* locale, so it is
looked up right before the conversion instead of once when the lexer is
constructed: a locale change in between (by a parser callback, a SAX
handler, or another thread) must not truncate the value (#5198). The
token has been validated before, so if the conversion stops early and the
decimal point changed in the meantime, the locale changed between the
lookup and the call, and the conversion is repeated with the new decimal
point. If the decimal point did not change, a retry cannot succeed: the
locale's decimal point is not a single character (e.g., the two-byte
U+066B of ar_EG.UTF-8 or fa_IR.UTF-8) and cannot be substituted in place.
The value strtod parsed up to that point is kept, as before this change.
Note that changing the locale in another thread *while* strtod runs is
undefined behavior of the C library, which this function cannot prevent.
*/
void convert_float_locale_aware()
{
const bool has_dot = decimal_point_position != std::string::npos;
char decimal_point = get_decimal_point();
for (;;)
{
const bool substitute = has_dot && decimal_point != '.';
if (substitute)
{
token_buffer[decimal_point_position] = static_cast<typename string_t::value_type>(decimal_point);
}
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
strtof(value_float, token_buffer.data(), &endptr);
if (substitute)
{
// get_string() hands the token to the SAX interface with '.'
token_buffer[decimal_point_position] = '.';
}
if (JSON_HEDLEY_LIKELY(endptr == token_buffer.data() + token_buffer.size()))
{
return;
}
// retry only if the locale changed; otherwise, this would loop forever
const char current_decimal_point = get_decimal_point();
if (current_decimal_point == decimal_point)
{
return;
}
decimal_point = current_decimal_point;
}
}
/*!
@brief contiguous fast path for scanning a number
Parses the whole number token straight from the input buffer, avoiding the
per-character get()/add() of scan_number(). On success it fills token_buffer
(as scan_number() does) and
(with the locale decimal point substituted, as scan_number() does) and
returns the token type. On anything it does not fully recognize as a
well-formed number it makes no state change and returns
token_type::uninitialized, so the caller falls back to scan_number(), which
@@ -10847,11 +10804,16 @@ scan_number_done:
}
#endif
// materialize the token exactly as scan_number() would. reset() already
// cleared token_buffer, so append() fills it (assign() is avoided
// because custom string_t types need not provide it)
// materialize the token exactly as scan_number() would, substituting the
// locale decimal point so convert_number()'s strtof fallback stays valid.
// reset() already cleared token_buffer, so append() fills it (assign() is
// avoided because custom string_t types need not provide it)
token_buffer.append(reinterpret_cast<const typename string_t::value_type*>(data), len);
decimal_point_position = dot_index;
if (dot_index != std::string::npos)
{
token_buffer[dot_index] = static_cast<typename string_t::value_type>(decimal_point_char);
decimal_point_position = dot_index;
}
ia.bulk_skip(len - 1);
position.chars_read_total += (len - 1);
@@ -11118,7 +11080,11 @@ scan_number_done:
/// return current string value (implicitly resets the token; useful only once)
string_t& get_string()
{
// a number token holds '.' regardless of the locale (#4084)
// translate decimal points from locale back to '.' (#4084)
if (decimal_point_char != '.' && decimal_point_position != std::string::npos)
{
token_buffer[decimal_point_position] = '.';
}
return token_buffer;
}
@@ -11414,7 +11380,9 @@ scan_number_done:
number_unsigned_t value_unsigned = 0;
number_float_t value_float = 0;
/// the position of the decimal point in token_buffer
/// the decimal point
const char_int_type decimal_point_char = '.';
/// the position of the decimal point in the input
std::size_t decimal_point_position = std::string::npos;
/// whether the caller (e.g. accept()/json_sax_acceptor) only needs the
@@ -14091,80 +14059,6 @@ class binary_reader
}
}
/*!
@brief reads a CBOR object key
RFC 8949 allows any data item as a map key, but only strings have a
counterpart in JSON. A key of any other type is rejected with a message
naming that type, rather than the one @ref get_cbor_string gives for a
malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_cbor_object_key(string_t& result)
{
// EOF and major type 3 (text string) are left to get_cbor_string
if (current == char_traits<char_type>::eof() || (static_cast<unsigned int>(current) & 0xE0u) == 0x60u)
{
return get_cbor_string(result);
}
const char* found = nullptr;
switch (static_cast<unsigned int>(current) >> 5u)
{
case 0:
found = "an unsigned integer";
break;
case 1:
found = "a negative integer";
break;
case 2:
found = "a byte string";
break;
case 4:
found = "an array";
break;
case 5:
found = "a map";
break;
case 6:
found = "a tag";
break;
default: // major type 7
switch (current)
{
case 0xF4:
case 0xF5:
found = "a boolean";
break;
case 0xF6:
found = "null";
break;
case 0xF7:
found = "undefined";
break;
case 0xF9:
case 0xFA:
case 0xFB:
found = "a floating-point number";
break;
case 0xFF:
found = "a break stop code";
break;
default:
found = "a simple value";
break;
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::cbor, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a definite-length CBOR byte array
@@ -14409,7 +14303,7 @@ class binary_reader
if (top.is_object)
{
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_cbor_object_key(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_cbor_string(key) || !sax->key(key)))
{
return false;
}
@@ -14910,98 +14804,6 @@ class binary_reader
}
}
/*!
@brief reads a MessagePack object key
The MessagePack specification allows any type as a map key, but only
strings have a counterpart in JSON. A key of any other type is rejected
with a message naming that type, rather than the one @ref
get_msgpack_string gives for a malformed string.
@param[out] result created key
@return whether key creation completed
*/
bool get_msgpack_object_key(string_t& result)
{
const char* found = nullptr;
switch (current)
{
case 0xC0:
found = "nil";
break;
case 0xC2:
case 0xC3:
found = "a boolean";
break;
case 0xCA:
case 0xCB:
found = "a float";
break;
case 0xC4:
case 0xC5:
case 0xC6:
found = "a bin";
break;
case 0xC7:
case 0xC8:
case 0xC9:
case 0xD4:
case 0xD5:
case 0xD6:
case 0xD7:
case 0xD8:
found = "an ext";
break;
case 0xCC:
case 0xCD:
case 0xCE:
case 0xCF:
case 0xD0:
case 0xD1:
case 0xD2:
case 0xD3:
found = "an integer";
break;
case 0xDC:
case 0xDD:
found = "an array";
break;
case 0xDE:
case 0xDF:
found = "a map";
break;
default:
// fixint, fixmap, and fixarray; strings, EOF, and the unused
// byte 0xC1 are left to get_msgpack_string
if (current == char_traits<char_type>::eof())
{
return get_msgpack_string(result);
}
if (current <= 0x7F || current >= 0xE0)
{
found = "an integer";
}
else if (current <= 0x8F)
{
found = "a map";
}
else if (current <= 0x9F)
{
found = "an array";
}
else
{
return get_msgpack_string(result);
}
break;
}
auto last_token = get_token_string();
return sax->parse_error(chars_read, last_token, parse_error::create(113, chars_read,
exception_message(input_format_t::msgpack, concat("only string keys are supported, but found ", found, "; last byte: 0x", last_token), "object key"), nullptr));
}
/*!
@brief reads a MessagePack byte array
@@ -15164,7 +14966,7 @@ class binary_reader
{
get();
key.clear();
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_object_key(key) || !sax->key(key)))
if (JSON_HEDLEY_UNLIKELY(!get_msgpack_string(key) || !sax->key(key)))
{
return false;
}
@@ -19377,6 +19179,10 @@ class json_pointer
@return const reference to the JSON value pointed to by the JSON
pointer
@pre Every object key and array index the pointer refers to exists.
Like the const operator[] for keys and indices, a missing one is
undefined behavior, guarded by a runtime assertion.
@throw parse_error.106 if an array index begins with '0'
@throw parse_error.109 if an array index was not a number
@throw out_of_range.402 if the array index '-' is used
@@ -19391,7 +19197,8 @@ class json_pointer
{
case detail::value_t::object:
{
// use unchecked object access
// use unchecked object access; the const operator[]
// asserts that the key exists
ptr = &ptr->operator[](reference_token);
break;
}
@@ -19404,7 +19211,8 @@ class json_pointer
JSON_THROW(detail::out_of_range::create(402, detail::concat("array index '-' (", std::to_string(ptr->m_data.m_value.array->size()), ") is out of range"), ptr));
}
// use unchecked array access
// use unchecked array access; the const operator[]
// asserts that the index exists
ptr = &ptr->operator[](array_index<BasicJsonType>(reference_token));
break;
}
@@ -26989,11 +26797,11 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
/*!
@brief how many levels the operation going on in this thread has descended into
Copying a value, converting one from another specialization, and comparing
two values share this count. The library never nests one of them inside
another - none of them does either of the other two on the way - and where
user code nests them anyway, sharing the count only ends a descent sooner
than it had to, which costs a little speed and is never wrong.
Copying a value and comparing two values share this count. The library never
nests one inside the other - copying a value does not compare one, and
comparing two values does not copy them - and where user code nests them
anyway, sharing the count only ends a descent sooner than it had to, which
costs a little speed and is never wrong.
A byte is enough: the count never exceeds the limit by more than the single
level that notices the limit has been reached.
@@ -27349,221 +27157,6 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
copy_iteratively(src);
}
/*!
@brief convert the value @a val of another specialization into this null
value; @a val must be neither an object nor an array
Converting such a value never descends, so both ways of converting an
object or an array (@ref convert_structured) leave their elements of this
kind to the converting constructor, which leaves them to this.
*/
template<typename BasicJsonType>
void convert_leaf(const BasicJsonType& val)
{
using other_boolean_t = typename BasicJsonType::boolean_t;
using other_number_float_t = typename BasicJsonType::number_float_t;
using other_number_integer_t = typename BasicJsonType::number_integer_t;
using other_number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using other_string_t = typename BasicJsonType::string_t;
using other_binary_t = typename BasicJsonType::binary_t;
switch (val.type())
{
case value_t::boolean:
JSONSerializer<other_boolean_t>::to_json(*this, val.template get<other_boolean_t>());
break;
case value_t::number_float:
JSONSerializer<other_number_float_t>::to_json(*this, val.template get<other_number_float_t>());
break;
case value_t::number_integer:
JSONSerializer<other_number_integer_t>::to_json(*this, val.template get<other_number_integer_t>());
break;
case value_t::number_unsigned:
JSONSerializer<other_number_unsigned_t>::to_json(*this, val.template get<other_number_unsigned_t>());
break;
case value_t::string:
JSONSerializer<other_string_t>::to_json(*this, val.template get_ref<const other_string_t&>());
break;
case value_t::binary:
JSONSerializer<other_binary_t>::to_json(*this, val.template get_ref<const other_binary_t&>());
break;
case value_t::null:
// this value is null already; assigning null to it would also
// reset its positions (JSON_DIAGNOSTIC_POSITIONS)
break;
case value_t::discarded:
m_data.m_type = value_t::discarded;
break;
case value_t::object: // LCOV_EXCL_LINE
case value_t::array: // LCOV_EXCL_LINE
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
}
/// scratch space for the converted elements of the arrays that
/// @ref convert_iteratively has yet to create
using convert_scratch_t = std::vector<basic_json, AllocatorType<basic_json>>;
/*!
@brief create the object or array @a val converted into this null value
Its converted elements are the last `val.size()` entries of @a elements (an
array) or of @a members (an object); they are moved into the container in
one go and then removed.
*/
template<typename BasicJsonType>
void convert_level(const BasicJsonType& val, convert_scratch_t& elements, copy_scratch_t& members)
{
if (val.is_object())
{
const auto first = members.end() - static_cast<typename copy_scratch_t::difference_type>(val.size());
m_data.m_value.object = create<object_t>(std::make_move_iterator(first),
std::make_move_iterator(members.end()));
// only now that the object exists may this stop being a null value
m_data.m_type = value_t::object;
members.erase(first, members.end());
}
else
{
const auto first = elements.end() - static_cast<typename convert_scratch_t::difference_type>(val.size());
m_data.m_value.array = create<array_t>(std::make_move_iterator(first),
std::make_move_iterator(elements.end()));
// only now that the array exists may this stop being a null value
m_data.m_type = value_t::array;
elements.erase(first, elements.end());
}
set_parents();
}
/*!
@brief convert the object or array @a val of another specialization into
this null value without recursing
The containers whose conversion has begun are kept on an explicit stack
rather than on the call stack. Unlike @ref copy_iteratively, this builds
every container from the bottom up: all its elements are converted first,
and the container is then created from them in one go, the way the range
constructor that converts the levels above the bound does. The two object
types need not enumerate their members in the same order, so the members
could not be paired up by position anyway, and building from a range keeps
what the range constructor does with keys that become equal on conversion.
Every value is complete before it is handed on, and a container gets its
type only once it exists, so whatever throws, every value left behind can
be destroyed.
*/
template<typename BasicJsonType>
void convert_iteratively(const BasicJsonType& val)
{
using other_const_iterator = typename BasicJsonType::const_iterator;
// the containers whose conversion has begun, innermost last, each with
// its element to convert next
std::vector<std::pair<const BasicJsonType*, other_const_iterator>> pending;
// the converted elements of the pending arrays and the converted
// members of the pending objects, those of the innermost one last
convert_scratch_t elements;
copy_scratch_t members;
pending.emplace_back(&val, val.cbegin());
for (;;)
{
const BasicJsonType& container = *pending.back().first;
other_const_iterator& next = pending.back().second;
if (next != container.cend())
{
if (next->is_structured())
{
// convert its elements first; next stays where it is until
// the converted container is handed back to this one
pending.emplace_back(&*next, next->cbegin());
continue;
}
// the converting constructor does not descend into this value
if (container.is_object())
{
members.emplace_back(next.key(), *next);
}
else
{
elements.emplace_back(*next);
}
++next;
continue;
}
// all elements of the container are converted: create it
pending.pop_back();
if (pending.empty())
{
convert_level(container, elements, members);
return;
}
basic_json converted;
converted.convert_level(container, elements, members);
#if JSON_DIAGNOSTIC_POSITIONS
converted.start_position = container.start_pos();
converted.end_position = container.end_pos();
#endif
// hand it to the container it is an element of
if (pending.back().first->is_object())
{
members.emplace_back(pending.back().second.key(), std::move(converted));
}
else
{
elements.push_back(std::move(converted));
}
++pending.back().second;
}
}
/*!
@brief convert the object or array @a val of another specialization into
this null value
Converting a container converts its elements, so a value nested deeply
enough used to exhaust the call stack. The descent is bounded here as in
@ref copy_structured: the first @ref nesting_depth_limit levels are
converted by the containers' range constructors, just as they always were,
and anything below that is converted without the call stack by
@ref convert_iteratively.
@sa https://github.com/nlohmann/json/issues/5650
*/
template<typename BasicJsonType>
void convert_structured(const BasicJsonType& val)
{
const nesting_depth_guard guard;
if (JSON_HEDLEY_LIKELY(guard.okay()))
{
// every element comes back to the converting constructor
if (val.is_object())
{
using other_object_t = typename BasicJsonType::object_t;
JSONSerializer<other_object_t>::to_json(*this, val.template get_ref<const other_object_t&>());
}
else
{
using other_array_t = typename BasicJsonType::array_t;
JSONSerializer<other_array_t>::to_json(*this, val.template get_ref<const other_array_t&>());
}
return;
}
convert_iteratively(val);
}
/// the result of comparing two values, including values that cannot be
/// ordered at all, such as a discarded value or a NaN
@@ -27912,13 +27505,49 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
end_position(val.end_pos())
#endif
{
if (val.is_structured())
using other_boolean_t = typename BasicJsonType::boolean_t;
using other_number_float_t = typename BasicJsonType::number_float_t;
using other_number_integer_t = typename BasicJsonType::number_integer_t;
using other_number_unsigned_t = typename BasicJsonType::number_unsigned_t;
using other_string_t = typename BasicJsonType::string_t;
using other_object_t = typename BasicJsonType::object_t;
using other_array_t = typename BasicJsonType::array_t;
using other_binary_t = typename BasicJsonType::binary_t;
switch (val.type())
{
convert_structured(val);
}
else
{
convert_leaf(val);
case value_t::boolean:
JSONSerializer<other_boolean_t>::to_json(*this, val.template get<other_boolean_t>());
break;
case value_t::number_float:
JSONSerializer<other_number_float_t>::to_json(*this, val.template get<other_number_float_t>());
break;
case value_t::number_integer:
JSONSerializer<other_number_integer_t>::to_json(*this, val.template get<other_number_integer_t>());
break;
case value_t::number_unsigned:
JSONSerializer<other_number_unsigned_t>::to_json(*this, val.template get<other_number_unsigned_t>());
break;
case value_t::string:
JSONSerializer<other_string_t>::to_json(*this, val.template get_ref<const other_string_t&>());
break;
case value_t::object:
JSONSerializer<other_object_t>::to_json(*this, val.template get_ref<const other_object_t&>());
break;
case value_t::array:
JSONSerializer<other_array_t>::to_json(*this, val.template get_ref<const other_array_t&>());
break;
case value_t::binary:
JSONSerializer<other_binary_t>::to_json(*this, val.template get_ref<const other_binary_t&>());
break;
case value_t::null:
*this = nullptr;
break;
case value_t::discarded:
m_data.m_type = value_t::discarded;
break;
default: // LCOV_EXCL_LINE
JSON_ASSERT(false); // NOLINT(cert-dcl03-c,hicpp-static-assert,misc-static-assert) LCOV_EXCL_LINE
}
JSON_ASSERT(m_data.m_type == val.type());
@@ -29126,6 +28755,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
// const operator[] only works for arrays
if (JSON_HEDLEY_LIKELY(is_array()))
{
JSON_ASSERT(idx < m_data.m_value.array->size());
return m_data.m_value.array->operator[](idx);
}
-71
View File
@@ -352,77 +352,6 @@ TEST_CASE("deep copy uses the provided allocator")
CHECK(copy == j);
}
namespace
{
// the number of constructions countdown_allocator lets happen, including the
// one that fails; 0 means none ever fails
std::size_t constructions_until_failure = 0;
template<class T>
struct countdown_allocator : std::allocator<T>
{
using std::allocator<T>::allocator;
template<class U, class... Args>
void construct(U* p, Args&& ... args)
{
if (constructions_until_failure != 0 && --constructions_until_failure == 0)
{
throw std::bad_alloc();
}
::new (static_cast<void*>(p)) U(std::forward<Args>(args)...);
}
template <class U>
struct rebind
{
using other = countdown_allocator<U>;
};
};
} // namespace
TEST_CASE("converting a deeply nested value from another specialization fails cleanly (#5650)")
{
using countdown_json = nlohmann::basic_json<std::map,
std::vector,
std::string,
bool,
std::int64_t,
std::uint64_t,
double,
countdown_allocator>;
// deeper than the 128 levels the converting constructor descends into, so
// that failures land on both sides of the bound - or, built with
// JSON_NO_THREAD_LOCAL, all in the iterative conversion
json j = {1, "two", {{"three", 3}}};
for (std::size_t i = 0; i < 150; ++i)
{
j = json{{"a", json::array({j, "sibling"})}};
}
// Fail every construction in turn. Each failure has to reach the caller,
// and everything built until then has to be destroyed cleanly.
std::size_t failures = 0;
for (std::size_t n = 1;; ++n)
{
constructions_until_failure = n;
try
{
const countdown_json converted = j;
constructions_until_failure = 0;
CHECK(converted.dump() == j.dump());
break;
}
catch (const std::bad_alloc&)
{
++failures;
}
}
CHECK(failures > 0);
}
namespace
{
template<class T>
+2 -43
View File
@@ -1830,51 +1830,10 @@ TEST_CASE("CBOR")
SECTION("invalid string in map")
{
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&);
CHECK(json::from_cbor(std::vector<uint8_t>({0xa1, 0xff, 0x01}), true, false).is_discarded());
}
SECTION("non-string key (see #2766 and #3381)")
{
// only text strings map to JSON object keys; any other key is
// rejected with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0xA1, 0x01, 0x01}, "an unsigned integer; last byte: 0x01"},
{{0xA1, 0x20, 0x01}, "a negative integer; last byte: 0x20"},
{{0xA1, 0x41, 0x61, 0x01}, "a byte string; last byte: 0x41"},
{{0xA1, 0x80, 0x01}, "an array; last byte: 0x80"},
{{0xA1, 0xA0, 0x01}, "a map; last byte: 0xA0"},
{{0xA1, 0xC0, 0x61, 0x61, 0x01}, "a tag; last byte: 0xC0"},
{{0xA1, 0xF4, 0x01}, "a boolean; last byte: 0xF4"},
{{0xA1, 0xF5, 0x01}, "a boolean; last byte: 0xF5"},
{{0xA1, 0xF6, 0x01}, "null; last byte: 0xF6"},
{{0xA1, 0xF7, 0x01}, "undefined; last byte: 0xF7"},
{{0xA1, 0xF9, 0x3C, 0x00, 0x01}, "a floating-point number; last byte: 0xF9"},
{{0xA1, 0xFA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFA"},
{{0xA1, 0xFB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a floating-point number; last byte: 0xFB"},
{{0xA1, 0xE0, 0x01}, "a simple value; last byte: 0xE0"},
{{0xA1, 0xF8, 0x20, 0x01}, "a simple value; last byte: 0xF8"},
// indefinite-length map
{{0xBF, 0x01, 0x01, 0xFF}, "an unsigned integer; last byte: 0x01"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_cbor(c.first, true, false).is_discarded());
}
// a key of major type 3 with a reserved length is still reported as
// a malformed string, and a missing key as the end of input
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing CBOR string: unexpected end of input", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0x7C, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x7C", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
{
// a two-character text string (major type 3) whose bytes are not
@@ -2325,7 +2284,7 @@ TEST_CASE("CBOR indefinite-length strings do not recurse per chunk")
SECTION("a break marker outside an indefinite-length string is not a string")
{
// 0xFF only closes a string that was opened; on its own it is not one
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found a break stop code; last byte: 0xFF", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(std::vector<uint8_t>({0xA1, 0xFF, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xFF", json::parse_error&);
}
}
+1 -1
View File
@@ -666,7 +666,7 @@ TEST_CASE("parse_float_fast declines what it cannot convert exactly")
// always safe: the caller then falls back to a slower, exact conversion.
const auto fast = [](const std::string & s, double & out)
{
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), out);
return nlohmann::detail::parse_float_fast(s.data(), s.data() + s.size(), '.', out);
};
double out = 0;
-52
View File
@@ -141,58 +141,6 @@ TEST_CASE("Better diagnostics with positions")
check_objects(300);
}
SECTION("converting keeps the positions of nested values (#5650)")
{
// Values nested deeper than the converting constructor's descent bound
// are converted without the call stack, on a path that has to carry the
// positions of every value over itself. Objects and arrays take turns,
// and the innermost value is null, which used to lose its positions.
const auto check_conversion = [](std::size_t depth)
{
CAPTURE(depth)
std::string text;
std::string closing;
for (std::size_t i = 0; i < depth; ++i)
{
text += (i % 2 == 0) ? "[12, " : R"({"b":1, "a":)";
closing += (i % 2 == 0) ? ']' : '}';
}
text += "null";
text.append(closing.rbegin(), closing.rend());
const json original = json::parse(text);
const nlohmann::ordered_json converted = original;
const json* o = &original;
const nlohmann::ordered_json* c = &converted;
for (std::size_t level = 0; level <= depth; ++level)
{
CAPTURE(level)
REQUIRE(c->start_pos() == o->start_pos());
REQUIRE(c->end_pos() == o->end_pos());
if (level < depth)
{
// the number beside the value nested next
const json& o_number = o->is_object() ? o->at("b") : o->at(0);
const nlohmann::ordered_json& c_number = c->is_object() ? c->at("b") : c->at(0);
REQUIRE(c_number.start_pos() == o_number.start_pos());
REQUIRE(c_number.end_pos() == o_number.end_pos());
o = o->is_object() ? &o->at("a") : &o->at(1);
c = c->is_object() ? &c->at("a") : &c->at(1);
}
}
};
check_conversion(1);
check_conversion(127);
check_conversion(128);
check_conversion(129);
check_conversion(300);
}
SECTION("JSON patch add to primitive parent (#4292)")
{
// the JSON Patch "add" target /foo/bar/baz has a string parent
-30
View File
@@ -331,36 +331,6 @@ TEST_CASE("Regression tests for extended diagnostics")
}
}
SECTION("Regression test for issue #5650 - converting keeps the parents of nested values")
{
// A value nested deeper than the converting constructor's descent bound
// is converted without the call stack. Every container that path creates
// has to have the parents of its children set, or the JSON Pointer in the
// diagnostic is cut short. Objects and arrays take turns.
const std::size_t pairs = 150;
json j = "not a number";
std::string pointer;
for (std::size_t i = 0; i < pairs; ++i)
{
j = json{{"a", json::array({j})}};
pointer += "/a/0";
}
const nlohmann::ordered_json converted = j;
const nlohmann::ordered_json* inner = &converted;
for (std::size_t i = 0; i < pairs; ++i)
{
inner = &inner->at("a").at(0);
}
std::string const expected = "[json.exception.type_error.302] (" + pointer + ") type must be number, but is string";
int i = 0;
CHECK_THROWS_WITH_AS(i = inner->get<int>(), expected.c_str(), nlohmann::ordered_json::type_error);
CHECK(i == 0);
}
SECTION("Regression test - swap(array_t&)/swap(object_t&) must update JSON_DIAGNOSTICS parent pointers")
{
// swap(array_t&)
-128
View File
@@ -13,7 +13,6 @@ using nlohmann::json;
#include <algorithm>
#include <string>
#include <vector>
TEST_CASE("tests on very large JSONs")
{
@@ -54,24 +53,6 @@ const json* innermost_value(const json& j, std::size_t& depth)
return current;
}
// The text of a value nested depth levels deep around the number 0. Level i is
// an array if pattern[i % pattern.size()] is '[', and otherwise an object with
// the single member "a", which every object type enumerates in the same order.
std::string nested_text(std::size_t depth, const std::string& pattern)
{
std::string text;
std::string closing;
for (std::size_t i = 0; i < depth; ++i)
{
const bool array = pattern[i % pattern.size()] == '[';
text += array ? "[" : "{\"a\":";
closing += array ? ']' : '}';
}
text += '0';
text.append(closing.rbegin(), closing.rend());
return text;
}
} // namespace
TEST_CASE("tests on deeply nested JSONs")
@@ -243,114 +224,5 @@ TEST_CASE("tests on deeply nested JSONs")
CHECK(*innermost_value(j, unused) == 0);
}
}
SECTION("issue #5650 - stack overflow converting between specializations")
{
const std::vector<std::string> patterns = {"[", "{", "[{"};
SECTION("json to ordered_json")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const json j = json::parse(text);
const nlohmann::ordered_json converted = j;
CHECK(converted.dump() == text);
}
}
SECTION("ordered_json to json")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const nlohmann::ordered_json o = nlohmann::ordered_json::parse(text);
const json converted = o;
CHECK(converted.dump() == text);
}
}
SECTION("get<ordered_json>()")
{
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(depth, pattern);
const json j = json::parse(text);
CHECK(j.get<nlohmann::ordered_json>().dump() == text);
}
}
SECTION("depths around the bound of the recursive descent")
{
for (std::size_t d = 1; d <= 300; ++d)
{
CAPTURE(d);
for (const auto& pattern : patterns)
{
CAPTURE(pattern);
const std::string text = nested_text(d, pattern);
const json j = json::parse(text);
const nlohmann::ordered_json converted = j;
CHECK(converted.dump() == text);
const json back = converted;
CHECK(back.dump() == text);
}
}
}
SECTION("values below the bound are converted as values above it")
{
// Bury a value below the bound, where it is converted without the
// call stack, and compare it with the same value converted on its
// own by the containers' range constructors. Its objects have
// members that the two object types enumerate in different orders.
const auto bury = [](nlohmann::ordered_json value)
{
for (std::size_t i = 0; i < 200; ++i)
{
value = nlohmann::ordered_json::array({std::move(value)});
}
return value;
};
const auto dig = [](const json & value)
{
const json* current = &value;
for (std::size_t i = 0; i < 200; ++i)
{
current = &current->at(0);
}
return current;
};
nlohmann::ordered_json value = nlohmann::ordered_json::object();
value["z"] = {1, -2, 3U, 4.5, true, nullptr, "six", nlohmann::ordered_json::binary({7, 8}, 9),
nlohmann::ordered_json::binary({10}), nlohmann::ordered_json::array(), nlohmann::ordered_json::object()
};
value["y"] = {{"x", {{"w", 1}, {"v", 2}}}, {"u", {3, {{"t", 4}, {"s", 5}}}}};
value["r"] = nlohmann::ordered_json::array({nlohmann::ordered_json(nlohmann::ordered_json::value_t::discarded)});
const json converted_above = value;
const json buried = bury(value);
const json& converted_below = *dig(buried);
CHECK(converted_below.dump() == converted_above.dump());
CHECK(converted_below.at("z").at(7).get_binary().subtype() == 9);
CHECK_FALSE(converted_below.at("z").at(8).get_binary().has_subtype());
CHECK(converted_below.at("r").at(0).is_discarded());
// a discarded value is never equal to anything, so compare the rest
value.erase("r");
const json without_discarded_above = value;
const json without_discarded_buried = bury(value);
CHECK(*dig(without_discarded_buried) == without_discarded_above);
}
}
}
-210
View File
@@ -12,12 +12,7 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <array>
#include <clocale>
#include <map>
#include <string>
#include <utility>
#include <vector>
struct ParserImpl final: public nlohmann::json_sax<json>
{
@@ -180,208 +175,3 @@ TEST_CASE("locale-dependent test (LC_NUMERIC=de_DE)")
MESSAGE("locale de_DE is not usable");
}
}
namespace
{
// records the numbers of a flat array and switches LC_NUMERIC to the given
// locale once the array opens - after the lexer was constructed, but before
// any number in the array is lexed
struct LocaleSwitchingSax final: public nlohmann::json_sax<json>
{
explicit LocaleSwitchingSax(const char* switch_to)
: locale_after_open(switch_to)
{}
bool null() override
{
return true;
}
bool boolean(bool /*val*/) override
{
return true;
}
bool number_integer(json::number_integer_t /*val*/) override
{
return true;
}
bool number_unsigned(json::number_unsigned_t /*val*/) override
{
return true;
}
bool number_float(json::number_float_t val, const json::string_t& s) override
{
values.push_back(val);
strings.push_back(s);
return true;
}
bool string(json::string_t& /*val*/) override
{
return true;
}
bool binary(json::binary_t& /*val*/) override
{
return true;
}
bool start_object(std::size_t /*val*/) override
{
return true;
}
bool key(json::string_t& /*val*/) override
{
return true;
}
bool end_object() override
{
return true;
}
bool start_array(std::size_t /*val*/) override
{
switched = std::setlocale(LC_NUMERIC, locale_after_open.c_str()) != nullptr;
return true;
}
bool end_array() override
{
return true;
}
bool parse_error(std::size_t /*val*/, const std::string& /*val*/, const nlohmann::detail::exception& /*val*/) override
{
return false;
}
std::string locale_after_open;
bool switched = false;
std::vector<json::number_float_t> values {}; // NOLINT(readability-redundant-member-init)
std::vector<json::string_t> strings {}; // NOLINT(readability-redundant-member-init)
};
} // namespace
TEST_CASE("locale changes between lexer construction and number conversion (#5198)")
{
// The numbers are chosen so that the conversion also takes the strtod
// fallback, which honors the locale that is current at conversion time:
// too many significant digits for Clinger's fast path, an underflow that
// std::from_chars rejects, and a plain value.
const std::vector<std::string> numbers = {"3.14159265358979323846", "1.5e-400", "12.34", "-0.000123456789012345678"};
std::string text = "[";
for (const auto& n : numbers)
{
text += (text.size() == 1 ? "" : ",") + n;
}
text += "]";
using long_double_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int64_t, std::uint64_t, long double>;
// reference values, parsed without a locale switch
REQUIRE(std::setlocale(LC_NUMERIC, "C") != nullptr);
const json expected = json::parse(text);
const long_double_json expected_ld = long_double_json::parse(text);
const std::array<std::pair<const char*, const char*>, 2> transitions =
{
{
{"C", "de_DE"},
{"de_DE", "C"}
}
};
for (const auto& transition : transitions)
{
CAPTURE(transition.first);
CAPTURE(transition.second);
if (std::setlocale(LC_NUMERIC, transition.first) == nullptr)
{
MESSAGE("locale is not usable");
continue;
}
// SAX parsing
{
LocaleSwitchingSax sax(transition.second);
CHECK(json::sax_parse(text, &sax));
if (sax.switched)
{
CHECK(sax.values == expected.get<std::vector<json::number_float_t>>());
CHECK(sax.strings == numbers);
}
}
// DOM parsing with a callback
{
bool switched = false;
const auto cb = [&](int /*depth*/, json::parse_event_t event, json& /*parsed*/) noexcept
{
if (event == json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const json j = json::parse(text, cb);
if (switched)
{
CHECK(j == expected);
}
}
// a long double goes through std::strtold unless std::from_chars supports it
{
bool switched = false;
const auto cb = [&](int /*depth*/, long_double_json::parse_event_t event, long_double_json& /*parsed*/) noexcept
{
if (event == long_double_json::parse_event_t::array_start)
{
switched = std::setlocale(LC_NUMERIC, transition.second) != nullptr;
}
return true;
};
const long_double_json j = long_double_json::parse(text, cb);
if (switched)
{
CHECK(j == expected_ld);
}
}
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
TEST_CASE("locale with a multi-byte decimal point")
{
// Some locales use a decimal point that is not a single character, e.g.
// U+066B ARABIC DECIMAL SEPARATOR (two bytes in UTF-8). It cannot be
// substituted in place for '.', so the strtod fallback stops early. The
// conversion must still terminate rather than retry forever.
const std::array<const char*, 6> names = {{"ar_EG.UTF-8", "ar_SA.UTF-8", "fa_IR.UTF-8", "ps_AF.UTF-8", "ar_EG", "fa_IR"}};
bool tested = false;
for (const char* name : names)
{
if (std::setlocale(LC_NUMERIC, name) == nullptr)
{
continue;
}
const std::string decimal_point = std::localeconv()->decimal_point;
if (decimal_point.size() < 2)
{
continue;
}
CAPTURE(name);
tested = true;
// too many significant digits for Clinger's fast path, and an underflow
// that std::from_chars rejects: both reach the strtod fallback
json j;
CHECK_NOTHROW(j = json::parse("[3.14159265358979323846, 1.5e-400, -0.000123456789012345678]"));
CHECK(j.is_array());
CHECK(json::accept("3.14159265358979323846"));
// a value the locale-independent paths convert is not affected
CHECK(json::parse("12.5") == 12.5);
}
if (!tested)
{
MESSAGE("no locale with a multi-byte decimal point is usable");
}
CHECK(std::setlocale(LC_NUMERIC, "C") != nullptr);
}
+1 -60
View File
@@ -1551,69 +1551,10 @@ TEST_CASE("MessagePack")
SECTION("invalid string in map")
{
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found an integer; last byte: 0xFF", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xFF", json::parse_error&);
CHECK(json::from_msgpack(std::vector<uint8_t>({0x81, 0xff, 0x01}), true, false).is_discarded());
}
SECTION("non-string key (see #3381)")
{
// only strings map to JSON object keys; any other key is rejected
// with a message naming its type
const std::vector<std::pair<std::vector<std::uint8_t>, std::string>> cases =
{
{{0x81, 0xC0, 0x01}, "nil; last byte: 0xC0"},
{{0x81, 0xC2, 0x01}, "a boolean; last byte: 0xC2"},
{{0x81, 0xC3, 0x01}, "a boolean; last byte: 0xC3"},
{{0x81, 0xCA, 0x3F, 0x80, 0x00, 0x00, 0x01}, "a float; last byte: 0xCA"},
{{0x81, 0xCB, 0x3F, 0xF0, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "a float; last byte: 0xCB"},
{{0x81, 0xC4, 0x00, 0x01}, "a bin; last byte: 0xC4"},
{{0x81, 0xC5, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC5"},
{{0x81, 0xC6, 0x00, 0x00, 0x00, 0x00, 0x01}, "a bin; last byte: 0xC6"},
{{0x81, 0xC7, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC7"},
{{0x81, 0xC8, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC8"},
{{0x81, 0xC9, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an ext; last byte: 0xC9"},
{{0x81, 0xD4, 0x01, 0x00, 0x01}, "an ext; last byte: 0xD4"},
{{0x81, 0xD5, 0x01, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD5"},
{{0x81, 0xD6, 0x01, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD6"},
{{0x81, 0xD7, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD7"},
{{0x81, 0xD8, 0x01, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01}, "an ext; last byte: 0xD8"},
{{0x81, 0xCC, 0x01, 0x01}, "an integer; last byte: 0xCC"},
{{0x81, 0xCD, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCD"},
{{0x81, 0xCE, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCE"},
{{0x81, 0xCF, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xCF"},
{{0x81, 0xD0, 0x01, 0x01}, "an integer; last byte: 0xD0"},
{{0x81, 0xD1, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD1"},
{{0x81, 0xD2, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD2"},
{{0x81, 0xD3, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x01, 0x01}, "an integer; last byte: 0xD3"},
{{0x81, 0x00, 0x01}, "an integer; last byte: 0x00"},
{{0x81, 0x7F, 0x01}, "an integer; last byte: 0x7F"},
{{0x81, 0xE0, 0x01}, "an integer; last byte: 0xE0"},
{{0x81, 0x80, 0x01}, "a map; last byte: 0x80"},
{{0x81, 0x8F, 0x01}, "a map; last byte: 0x8F"},
{{0x81, 0xDE, 0x00, 0x00, 0x01}, "a map; last byte: 0xDE"},
{{0x81, 0xDF, 0x00, 0x00, 0x00, 0x00, 0x01}, "a map; last byte: 0xDF"},
{{0x81, 0x90, 0x01}, "an array; last byte: 0x90"},
{{0x81, 0x9F, 0x01}, "an array; last byte: 0x9F"},
{{0x81, 0xDC, 0x00, 0x00, 0x01}, "an array; last byte: 0xDC"},
{{0x81, 0xDD, 0x00, 0x00, 0x00, 0x00, 0x01}, "an array; last byte: 0xDD"},
};
for (const auto& c : cases)
{
CAPTURE(c.first)
const std::string expected = "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack object key: only string keys are supported, but found " + c.second;
json _;
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(c.first), expected.c_str(), json::parse_error&);
CHECK(json::from_msgpack(c.first, true, false).is_discarded());
}
json _;
// the unused byte 0xC1 is still reported as a malformed string
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81, 0xC1, 0x01})), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing MessagePack string: expected length specification (0xA0-0xBF, 0xD9-0xDB); last byte: 0xC1", json::parse_error&);
// a missing key is still reported as the end of input
CHECK_THROWS_WITH_AS(_ = json::from_msgpack(std::vector<uint8_t>({0x81})), "[json.exception.parse_error.110] parse error at byte 2: syntax error while parsing MessagePack string: unexpected end of input", json::parse_error&);
}
SECTION("invalid UTF-8 in string (see #5529)")
{
// a fixstr of length 2 (0xA0 | 2) whose bytes are not valid UTF-8
+3 -3
View File
@@ -1018,7 +1018,7 @@ TEST_CASE("regression tests 1")
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR object key: only string keys are supported, but found an array; last byte: 0x98", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0x98", json::parse_error&);
// related test case: nonempty UTF-8 string (indefinite length)
std::vector<uint8_t> const vec1 {0x7f, 0x61, 0x61};
@@ -1065,7 +1065,7 @@ TEST_CASE("regression tests 1")
};
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec1), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&);
// related test case: double-precision
std::vector<uint8_t> const vec2
@@ -1077,7 +1077,7 @@ TEST_CASE("regression tests 1")
0x96, 0x96, 0xb4, 0xb4, 0xfa, 0x94, 0x94, 0x61,
0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0x61, 0xfb
};
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR object key: only string keys are supported, but found a map; last byte: 0xB4", json::parse_error&);
CHECK_THROWS_WITH_AS(_ = json::from_cbor(vec2), "[json.exception.parse_error.113] parse error at byte 13: syntax error while parsing CBOR string: expected length specification (0x60-0x7B) or indefinite string type (0x7F); last byte: 0xB4", json::parse_error&);
}
SECTION("issue #452 - Heap-buffer-overflow (OSS-Fuzz issue 585)")