Compare commits

..
Author SHA1 Message Date
Niels Lohmann 998456a218 Merge branch 'develop' into claude/ordered-map-growth-moves-values
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 22:22:19 +02:00
Niels Lohmann 4b2d2244c3 Move instead of deep-copy ordered_json values when an object grows
ordered_map keeps its elements in a std::vector<std::pair<const Key, T>>.
With a std::string key, that pair is not nothrow move constructible (the
const key has to be copied), so std::vector copies every element when it
reallocates. For ordered_json, this deep-copies every member value an
object already holds, including whole nested subtrees, on each growth
step.

Grow the storage in ordered_map instead, copying the keys and moving the
values. This happens in two phases, so the strong exception guarantee is
kept without try/catch. The first phase may throw, but only touches a
temporary buffer: it copies the keys, value-initializes the values, and
constructs the new element. The second phase moves the values (noexcept)
and swaps the buffers. Because the new element is constructed before any
value is moved, arguments that refer to elements of the container stay
valid, as with std::vector. Types that cannot take this path keep the
std::vector behavior.

Parsing into ordered_json (ParseStringOrdered, Apple M1 Max, clang -O3):
twitter 3.20 -> 1.70 ms, citm_catalog 7.73 -> 3.67 ms, jeopardy 219 ->
177 ms, canada unchanged. The number of allocations for twitter and
citm_catalog drops by two thirds.

Also add ParseStringOrdered rows to the benchmarks, and document the
growth behavior and the exception safety of ordered_map.

Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-09-28 18:48:05 +02:00
14 changed files with 656 additions and 453 deletions
@@ -55,10 +55,6 @@ This implementation does exactly follow this approach, as it uses double precisi
smaller than `-1.79769313486232e+308` and values greater than `1.79769313486232e+308` will be stored as NaN internally
and be serialized to `null`.
During deserialization (from JSON text or any of the binary formats), a finite number that does not fit into
`number_float_t` is rejected with [`out_of_range.406`](../../home/exceptions.md#jsonexceptionout_of_range406), for
example a double-precision number in a binary format when `number_float_t` is `#!cpp float`.
#### Storage
Floating-point number values are stored directly inside a `basic_json` type.
@@ -47,9 +47,8 @@ With the default values for `NumberIntegerType` (`std::int64_t`), the default va
When the default type is used, the maximal integer number that can be stored is `9223372036854775807` (INT64_MAX) and
the minimal integer number that can be stored is `-9223372036854775808` (INT64_MIN). Integer numbers that are out of
range will yield over/underflow when used in a constructor. During deserialization (from JSON text or any of the binary
formats), too large or small integer numbers will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md)
or [`number_float_t`](number_float_t.md).
range will yield over/underflow when used in a constructor. During deserialization, too large or small integer numbers
will automatically be stored as [`number_unsigned_t`](number_unsigned_t.md) or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
@@ -48,9 +48,8 @@ With the default values for `NumberUnsignedType` (`std::uint64_t`), the default
When the default type is used, the maximal integer number that can be stored is `18446744073709551615` (UINT64_MAX) and
the minimal integer number that can be stored is `0`. Integer numbers that are out of range will yield over/underflow
when used in a constructor. During deserialization (from JSON text or any of the binary formats), too large or small
integer numbers will automatically be stored as [`number_integer_t`](number_integer_t.md) or
[`number_float_t`](number_float_t.md).
when used in a constructor. During deserialization, too large or small integer numbers will automatically be stored
as [`number_integer_t`](number_integer_t.md) or [`number_float_t`](number_float_t.md).
[RFC 8259](https://tools.ietf.org/html/rfc8259) further states:
> Note that when such software is used, numbers that are integers and are in the range $[-2^{53}+1, 2^{53}-1]$ are
+11
View File
@@ -28,6 +28,11 @@ A minimal map-like container that preserves insertion order for use within [`nlo
The type uses a `std::vector` to store object elements. Therefore, adding elements can yield a reallocation in which
case all iterators (including the `end()` iterator) and all references to the elements are invalidated.
When the storage grows, the keys are copied and the mapped values are moved to the new storage. A plain `std::vector`
would copy the whole elements instead, because their `#!cpp const` keys make them not nothrow move constructible; for
[`ordered_json`](ordered_json.md), this would be a deep copy of every nested value. The values are only copied if
`T` is not default constructible or not nothrow move assignable.
## Member types
- **key_type** - key type (`Key`)
@@ -56,6 +61,11 @@ std::equal_to<> // since C++14
- **find**
- **insert**
## Exception safety
**emplace**, **operator\[\]**, and **insert(value)** have the strong exception guarantee: if an exception is thrown (for
instance, because copying a key or allocating memory fails), the contents of the container are unchanged.
## Complexity
Because the elements are stored in a `std::vector` in insertion order, there is no index to look a key up by. Every
@@ -122,3 +132,4 @@ This differs from `#!cpp std::map`, where the same operations are O(log n).
- Added in version 3.9.0 to implement [`nlohmann::ordered_json`](ordered_json.md).
- Added **key_compare** member in version 3.11.0.
- Changed in version 3.13.0: growing the storage moves the mapped values instead of copying them.
@@ -168,9 +168,9 @@ The library maps CBOR types to JSON value types as follows:
!!! warning "Negative integer overflow"
CBOR negative integers (major type 1) are decoded as `-1 - n`. If the encoded magnitude `n` is too large for the
result to fit into `number_integer_t` (`std::int64_t` by default), the result is stored as `number_float_t`, like
a too small integer in JSON text. For example, `-18446744073709551616` (`0x3B` followed by eight `0xFF` bytes) is
stored as `-1.8446744073709552e+19`.
result to fit into `number_integer_t` (`std::int64_t` by default), parsing fails with a
[`parse_error.112`](../../home/exceptions.md#jsonexceptionparse_error112) exception rather than overflowing
silently.
!!! warning "Object keys"
+5 -7
View File
@@ -331,6 +331,9 @@ An unexpected byte was read in a [binary format](../features/binary_formats/inde
[json.exception.parse_error.112] parse error at byte 15: syntax error while parsing BSON binary: byte array length cannot be negative, is -1
```
```
[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow
```
```
[json.exception.parse_error.112] parse error at byte 5: syntax error while parsing BSON document: document size 6 does not match the number of bytes read (5)
```
@@ -851,18 +854,13 @@ The JSON Patch operations 'remove' and 'add' cannot be applied to the root eleme
### json.exception.out_of_range.406
A parsed number could not be stored without changing it to NaN or INF. For the binary formats, this happens when a
finite floating-point number does not fit into [`number_float_t`](../api/basic_json/number_float_t.md), for example a
double-precision number when `number_float_t` is `#!cpp float`.
A parsed number could not be stored as without changing it to NaN or INF.
!!! failure "Example messages"
!!! failure "Example message"
```
number overflow parsing '10E1000'
```
```
[json.exception.out_of_range.406] syntax error while parsing CBOR value: number overflow
```
### json.exception.out_of_range.407
+45 -133
View File
@@ -559,7 +559,7 @@ class binary_reader
case 0x01: // double
{
double number{};
return get_number<double, true>(input_format_t::bson, number) && emit_float(input_format_t::bson, number);
return get_number<double, true>(input_format_t::bson, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0x02: // string
@@ -600,19 +600,19 @@ class binary_reader
case 0x10: // int32
{
std::int32_t value{};
return get_number<std::int32_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value);
return get_number<std::int32_t, true>(input_format_t::bson, value) && sax->number_integer(value);
}
case 0x12: // int64
{
std::int64_t value{};
return get_number<std::int64_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value);
return get_number<std::int64_t, true>(input_format_t::bson, value) && sax->number_integer(value);
}
case 0x11: // uint64
{
std::uint64_t value{};
return get_number<std::uint64_t, true>(input_format_t::bson, value) && emit_unsigned(input_format_t::bson, value);
return get_number<std::uint64_t, true>(input_format_t::bson, value) && sax->number_unsigned(value);
}
default: // anything else is not supported (yet)
@@ -638,19 +638,14 @@ class binary_reader
{
return false;
}
// the value is -1 - number, which fits into number_integer_t
// whenever number does
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
const auto max_val = static_cast<NumberType>((std::numeric_limits<number_integer_t>::max)());
if (number > max_val)
{
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(112, chars_read,
exception_message(input_format_t::cbor, "negative integer overflow", "value"), nullptr));
}
// like the lexer does for JSON text, store a value too small for
// number_integer_t as number_float_t; compute it as long double so
// that emit_float sees a finite value and can detect an overflow of
// number_float_t
return emit_float(input_format_t::cbor, static_cast<long double>(-1) - static_cast<long double>(number));
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
}
/*!
@@ -707,25 +702,25 @@ class binary_reader
case 0x18: // Unsigned integer (one-byte uint8_t follows)
{
std::uint8_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
case 0x19: // Unsigned integer (two-byte uint16_t follows)
{
std::uint16_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
case 0x1A: // Unsigned integer (four-byte uint32_t follows)
{
std::uint32_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
case 0x1B: // Unsigned integer (eight-byte uint64_t follows)
{
std::uint64_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
// Negative integer -1-0x00..-1-0x17 (-1..-24)
@@ -1170,13 +1165,13 @@ class binary_reader
case 0xFA: // Single-Precision Float (four-byte IEEE 754)
{
float number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xFB: // Double-Precision Float (eight-byte IEEE 754)
{
double number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
default: // anything else (0xFF is handled inside the other types)
@@ -1940,61 +1935,61 @@ class binary_reader
case 0xCA: // float 32
{
float number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xCB: // float 64
{
double number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xCC: // uint 8
{
std::uint8_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xCD: // uint 16
{
std::uint16_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xCE: // uint 32
{
std::uint32_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xCF: // uint 64
{
std::uint64_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xD0: // int 8
{
std::int8_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xD1: // int 16
{
std::int16_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xD2: // int 32
{
std::int32_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xD3: // int 64
{
std::int64_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xDC: // array 16
@@ -2927,7 +2922,7 @@ class binary_reader
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!emit_unsigned(input_format, i)))
if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(static_cast<number_unsigned_t>(i))))
{
return false;
}
@@ -3059,37 +3054,37 @@ class binary_reader
break;
}
std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'U':
{
std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'i':
{
std::int8_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'I':
{
std::int16_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'l':
{
std::int32_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'L':
{
std::int64_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'u':
@@ -3099,7 +3094,7 @@ class binary_reader
break;
}
std::uint16_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'm':
@@ -3109,7 +3104,7 @@ class binary_reader
break;
}
std::uint32_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'M':
@@ -3119,7 +3114,7 @@ class binary_reader
break;
}
std::uint64_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'h':
@@ -3177,13 +3172,13 @@ class binary_reader
case 'd':
{
float number{};
return get_number(input_format, number) && emit_float(input_format, number);
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 'D':
{
double number{};
return get_number(input_format, number) && emit_float(input_format, number);
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 'H':
@@ -3650,13 +3645,13 @@ class binary_reader
case 0x8E: // binary32
{
float number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0x8F: // binary64
{
double number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xF8:
@@ -3722,9 +3717,7 @@ class binary_reader
@brief pass an integer to the SAX parser
Non-negative integers are passed as unsigned, negative integers as signed
numbers, like the other binary formats do. A value that does not fit the
number type is passed as described for @ref emit_unsigned and
@ref emit_signed.
numbers, like the other binary formats do.
@param[in] number the integer
@return whether the SAX parser accepted the value
@@ -3733,9 +3726,9 @@ class binary_reader
{
if (number >= 0)
{
return emit_unsigned(input_format_t::bon8, static_cast<std::uint64_t>(number));
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_signed(input_format_t::bon8, number);
return sax->number_integer(static_cast<number_integer_t>(number));
}
/*!
@@ -3792,7 +3785,8 @@ class binary_reader
value = (value << 8) | static_cast<std::int64_t>(current);
}
return emit_bon8_integer(negative ? -(value + offset) : value + offset);
return negative ? sax->number_integer(static_cast<number_integer_t>(-(value + offset)))
: sax->number_unsigned(static_cast<number_unsigned_t>(value + offset));
}
/*!
@@ -4089,88 +4083,6 @@ class binary_reader
return true;
}
/*!
@brief pass a signed integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_integer_t is passed as number_unsigned_t if it is non-negative and
fits there, and as number_float_t otherwise. With the default number
types, every integer the binary formats can encode fits, so this only
matters for narrower custom number types.
@tparam NumberType a signed integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_signed(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->number_integer(static_cast<number_integer_t>(number));
}
if (value_in_range_of<number_unsigned_t>(number))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass an unsigned integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_unsigned_t is passed as number_float_t.
@tparam NumberType an unsigned integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_unsigned(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_unsigned_t>(number)))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass a floating-point number read from the input to the SAX parser
Like the lexer does for JSON text, a finite value that overflows
number_float_t is rejected instead of silently becoming infinity. Infinity
and NaN in the input are passed on unchanged. Integers only overflow if
number_float_t cannot represent 2^64, e.g., a half-precision type.
@tparam NumberType a floating-point or integer type
@param[in] format the current format (for diagnostics)
@param[in] number the number
@return whether the SAX parser accepted the value
@throw out_of_range.406 if a finite @a number overflows number_float_t
*/
template<typename NumberType>
bool emit_float(const input_format_t format, const NumberType number)
{
const auto result = static_cast<number_float_t>(number);
if (JSON_HEDLEY_UNLIKELY(std::isfinite(number) && !std::isfinite(result)))
{
return sax->parse_error(chars_read, get_token_string(),
out_of_range::create(406, exception_message(format, "number overflow", "value"), nullptr));
}
return sax->number_float(result, "");
}
/*!
@brief create a string by reading characters from the input
+65 -5
View File
@@ -8,13 +8,15 @@
#pragma once
#include <algorithm> // max, min
#include <functional> // equal_to, less
#include <initializer_list> // initializer_list
#include <iterator> // input_iterator_tag, iterator_traits
#include <memory> // allocator
#include <stdexcept> // for out_of_range
#include <type_traits> // enable_if, is_convertible
#include <utility> // pair
#include <tuple> // forward_as_tuple
#include <type_traits> // enable_if, integral_constant, is_convertible, is_nothrow_move_constructible
#include <utility> // forward, move, pair, piecewise_construct
#include <vector> // vector
#include <nlohmann/detail/macro_scope.hpp>
@@ -79,7 +81,7 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
return {it, false};
}
}
Container::emplace_back(key, std::forward<T>(t));
append(key, std::forward<T>(t));
return {std::prev(this->end()), true};
}
@@ -94,7 +96,7 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
return {it, false};
}
}
Container::emplace_back(std::forward<KeyType>(key), std::forward<T>(t));
append(std::forward<KeyType>(key), std::forward<T>(t));
return {std::prev(this->end()), true};
}
@@ -368,7 +370,7 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
return {it, false};
}
}
Container::push_back(value);
append(value);
return {--this->end(), true};
}
@@ -386,6 +388,64 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
}
private:
/*!
@brief add an element whose key is not yet contained at the end
A std::vector copies all elements when it grows, because their const keys
make them not nothrow move constructible. For ordered_json, this is a deep
copy of every value. Where the strong exception guarantee can be kept, grow
the storage here instead, copying only the keys and moving the values.
*/
template<typename... Args>
void append(Args&& ... args)
{
// evaluated here rather than at class scope, because T is still
// incomplete when basic_json instantiates its object_t
using move_values = std::integral_constant<bool, detail::conjunction<
detail::negation<std::is_nothrow_move_constructible<value_type>>,
std::is_copy_constructible<key_type>,
detail::is_default_constructible<mapped_type>,
std::is_nothrow_move_assignable<mapped_type>>::value>;
append_impl(move_values{}, std::forward<Args>(args)...);
}
template<typename... Args>
void append_impl(std::true_type /*unused*/, Args&& ... args)
{
if (this->size() < this->capacity())
{
Container::emplace_back(std::forward<Args>(args)...);
return;
}
// 1. May throw, but only changes tmp: copy the keys, value-initialize
// the values, and add the new element. The arguments may refer to
// elements of this container, so they are used before any value is
// moved out of it.
Container tmp(this->get_allocator()); // equal allocators, so swap() is valid
tmp.reserve((std::min)(this->max_size(), (std::max)(size_type{1}, 2 * this->size())));
for (const auto& element : *this)
{
tmp.emplace_back(std::piecewise_construct, std::forward_as_tuple(element.first), std::forward_as_tuple());
}
tmp.emplace_back(std::forward<Args>(args)...);
// 2. Cannot throw: move the values over and adopt the new storage.
auto it = tmp.begin();
for (auto& element : *this)
{
it->second = std::move(element.second);
++it;
}
Container::swap(tmp);
}
template<typename... Args>
void append_impl(std::false_type /*unused*/, Args&& ... args)
{
Container::emplace_back(std::forward<Args>(args)...);
}
JSON_NO_UNIQUE_ADDRESS key_compare m_compare = key_compare();
};
+110 -138
View File
@@ -13326,7 +13326,7 @@ class binary_reader
case 0x01: // double
{
double number{};
return get_number<double, true>(input_format_t::bson, number) && emit_float(input_format_t::bson, number);
return get_number<double, true>(input_format_t::bson, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0x02: // string
@@ -13367,19 +13367,19 @@ class binary_reader
case 0x10: // int32
{
std::int32_t value{};
return get_number<std::int32_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value);
return get_number<std::int32_t, true>(input_format_t::bson, value) && sax->number_integer(value);
}
case 0x12: // int64
{
std::int64_t value{};
return get_number<std::int64_t, true>(input_format_t::bson, value) && emit_signed(input_format_t::bson, value);
return get_number<std::int64_t, true>(input_format_t::bson, value) && sax->number_integer(value);
}
case 0x11: // uint64
{
std::uint64_t value{};
return get_number<std::uint64_t, true>(input_format_t::bson, value) && emit_unsigned(input_format_t::bson, value);
return get_number<std::uint64_t, true>(input_format_t::bson, value) && sax->number_unsigned(value);
}
default: // anything else is not supported (yet)
@@ -13405,19 +13405,14 @@ class binary_reader
{
return false;
}
// the value is -1 - number, which fits into number_integer_t
// whenever number does
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
const auto max_val = static_cast<NumberType>((std::numeric_limits<number_integer_t>::max)());
if (number > max_val)
{
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
return sax->parse_error(chars_read, get_token_string(),
parse_error::create(112, chars_read,
exception_message(input_format_t::cbor, "negative integer overflow", "value"), nullptr));
}
// like the lexer does for JSON text, store a value too small for
// number_integer_t as number_float_t; compute it as long double so
// that emit_float sees a finite value and can detect an overflow of
// number_float_t
return emit_float(input_format_t::cbor, static_cast<long double>(-1) - static_cast<long double>(number));
return sax->number_integer(static_cast<number_integer_t>(-1) - static_cast<number_integer_t>(number));
}
/*!
@@ -13474,25 +13469,25 @@ class binary_reader
case 0x18: // Unsigned integer (one-byte uint8_t follows)
{
std::uint8_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
case 0x19: // Unsigned integer (two-byte uint16_t follows)
{
std::uint16_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
case 0x1A: // Unsigned integer (four-byte uint32_t follows)
{
std::uint32_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
case 0x1B: // Unsigned integer (eight-byte uint64_t follows)
{
std::uint64_t number{};
return get_number(input_format_t::cbor, number) && emit_unsigned(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_unsigned(number);
}
// Negative integer -1-0x00..-1-0x17 (-1..-24)
@@ -13937,13 +13932,13 @@ class binary_reader
case 0xFA: // Single-Precision Float (four-byte IEEE 754)
{
float number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xFB: // Double-Precision Float (eight-byte IEEE 754)
{
double number{};
return get_number(input_format_t::cbor, number) && emit_float(input_format_t::cbor, number);
return get_number(input_format_t::cbor, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
default: // anything else (0xFF is handled inside the other types)
@@ -14707,61 +14702,61 @@ class binary_reader
case 0xCA: // float 32
{
float number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xCB: // float 64
{
double number{};
return get_number(input_format_t::msgpack, number) && emit_float(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xCC: // uint 8
{
std::uint8_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xCD: // uint 16
{
std::uint16_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xCE: // uint 32
{
std::uint32_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xCF: // uint 64
{
std::uint64_t number{};
return get_number(input_format_t::msgpack, number) && emit_unsigned(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_unsigned(number);
}
case 0xD0: // int 8
{
std::int8_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xD1: // int 16
{
std::int16_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xD2: // int 32
{
std::int32_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xD3: // int 64
{
std::int64_t number{};
return get_number(input_format_t::msgpack, number) && emit_signed(input_format_t::msgpack, number);
return get_number(input_format_t::msgpack, number) && sax->number_integer(number);
}
case 0xDC: // array 16
@@ -15694,7 +15689,7 @@ class binary_reader
{
return sax->parse_error(chars_read, get_token_string(), out_of_range::create(408, exception_message(input_format, "excessive ndarray size caused overflow", "size"), nullptr));
}
if (JSON_HEDLEY_UNLIKELY(!emit_unsigned(input_format, i)))
if (JSON_HEDLEY_UNLIKELY(!sax->number_unsigned(static_cast<number_unsigned_t>(i))))
{
return false;
}
@@ -15826,37 +15821,37 @@ class binary_reader
break;
}
std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'U':
{
std::uint8_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'i':
{
std::int8_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'I':
{
std::int16_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'l':
{
std::int32_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'L':
{
std::int64_t number{};
return get_number(input_format, number) && emit_signed(input_format, number);
return get_number(input_format, number) && sax->number_integer(number);
}
case 'u':
@@ -15866,7 +15861,7 @@ class binary_reader
break;
}
std::uint16_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'm':
@@ -15876,7 +15871,7 @@ class binary_reader
break;
}
std::uint32_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'M':
@@ -15886,7 +15881,7 @@ class binary_reader
break;
}
std::uint64_t number{};
return get_number(input_format, number) && emit_unsigned(input_format, number);
return get_number(input_format, number) && sax->number_unsigned(number);
}
case 'h':
@@ -15944,13 +15939,13 @@ class binary_reader
case 'd':
{
float number{};
return get_number(input_format, number) && emit_float(input_format, number);
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 'D':
{
double number{};
return get_number(input_format, number) && emit_float(input_format, number);
return get_number(input_format, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 'H':
@@ -16417,13 +16412,13 @@ class binary_reader
case 0x8E: // binary32
{
float number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0x8F: // binary64
{
double number{};
return get_number(input_format_t::bon8, number) && emit_float(input_format_t::bon8, number);
return get_number(input_format_t::bon8, number) && sax->number_float(static_cast<number_float_t>(number), "");
}
case 0xF8:
@@ -16489,9 +16484,7 @@ class binary_reader
@brief pass an integer to the SAX parser
Non-negative integers are passed as unsigned, negative integers as signed
numbers, like the other binary formats do. A value that does not fit the
number type is passed as described for @ref emit_unsigned and
@ref emit_signed.
numbers, like the other binary formats do.
@param[in] number the integer
@return whether the SAX parser accepted the value
@@ -16500,9 +16493,9 @@ class binary_reader
{
if (number >= 0)
{
return emit_unsigned(input_format_t::bon8, static_cast<std::uint64_t>(number));
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_signed(input_format_t::bon8, number);
return sax->number_integer(static_cast<number_integer_t>(number));
}
/*!
@@ -16559,7 +16552,8 @@ class binary_reader
value = (value << 8) | static_cast<std::int64_t>(current);
}
return emit_bon8_integer(negative ? -(value + offset) : value + offset);
return negative ? sax->number_integer(static_cast<number_integer_t>(-(value + offset)))
: sax->number_unsigned(static_cast<number_unsigned_t>(value + offset));
}
/*!
@@ -16856,88 +16850,6 @@ class binary_reader
return true;
}
/*!
@brief pass a signed integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_integer_t is passed as number_unsigned_t if it is non-negative and
fits there, and as number_float_t otherwise. With the default number
types, every integer the binary formats can encode fits, so this only
matters for narrower custom number types.
@tparam NumberType a signed integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_signed(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_integer_t>(number)))
{
return sax->number_integer(static_cast<number_integer_t>(number));
}
if (value_in_range_of<number_unsigned_t>(number))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass an unsigned integer read from the input to the SAX parser
Like the lexer does for JSON text, a value that does not fit into
number_unsigned_t is passed as number_float_t.
@tparam NumberType an unsigned integer type
@param[in] format the current format (for diagnostics)
@param[in] number the integer
@return whether the SAX parser accepted the value
@throw out_of_range.406 if @a number overflows number_float_t (see
@ref emit_float)
*/
template<typename NumberType>
bool emit_unsigned(const input_format_t format, const NumberType number)
{
if (JSON_HEDLEY_LIKELY(value_in_range_of<number_unsigned_t>(number)))
{
return sax->number_unsigned(static_cast<number_unsigned_t>(number));
}
return emit_float(format, number);
}
/*!
@brief pass a floating-point number read from the input to the SAX parser
Like the lexer does for JSON text, a finite value that overflows
number_float_t is rejected instead of silently becoming infinity. Infinity
and NaN in the input are passed on unchanged. Integers only overflow if
number_float_t cannot represent 2^64, e.g., a half-precision type.
@tparam NumberType a floating-point or integer type
@param[in] format the current format (for diagnostics)
@param[in] number the number
@return whether the SAX parser accepted the value
@throw out_of_range.406 if a finite @a number overflows number_float_t
*/
template<typename NumberType>
bool emit_float(const input_format_t format, const NumberType number)
{
const auto result = static_cast<number_float_t>(number);
if (JSON_HEDLEY_UNLIKELY(std::isfinite(number) && !std::isfinite(result)))
{
return sax->parse_error(chars_read, get_token_string(),
out_of_range::create(406, exception_message(format, "number overflow", "value"), nullptr));
}
return sax->number_float(result, "");
}
/*!
@brief create a string by reading characters from the input
@@ -25856,13 +25768,15 @@ NLOHMANN_JSON_NAMESPACE_END
#include <algorithm> // max, min
#include <functional> // equal_to, less
#include <initializer_list> // initializer_list
#include <iterator> // input_iterator_tag, iterator_traits
#include <memory> // allocator
#include <stdexcept> // for out_of_range
#include <type_traits> // enable_if, is_convertible
#include <utility> // pair
#include <tuple> // forward_as_tuple
#include <type_traits> // enable_if, integral_constant, is_convertible, is_nothrow_move_constructible
#include <utility> // forward, move, pair, piecewise_construct
#include <vector> // vector
// #include <nlohmann/detail/macro_scope.hpp>
@@ -25929,7 +25843,7 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
return {it, false};
}
}
Container::emplace_back(key, std::forward<T>(t));
append(key, std::forward<T>(t));
return {std::prev(this->end()), true};
}
@@ -25944,7 +25858,7 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
return {it, false};
}
}
Container::emplace_back(std::forward<KeyType>(key), std::forward<T>(t));
append(std::forward<KeyType>(key), std::forward<T>(t));
return {std::prev(this->end()), true};
}
@@ -26218,7 +26132,7 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
return {it, false};
}
}
Container::push_back(value);
append(value);
return {--this->end(), true};
}
@@ -26236,6 +26150,64 @@ template <class Key, class T, class IgnoredLess = std::less<Key>,
}
private:
/*!
@brief add an element whose key is not yet contained at the end
A std::vector copies all elements when it grows, because their const keys
make them not nothrow move constructible. For ordered_json, this is a deep
copy of every value. Where the strong exception guarantee can be kept, grow
the storage here instead, copying only the keys and moving the values.
*/
template<typename... Args>
void append(Args&& ... args)
{
// evaluated here rather than at class scope, because T is still
// incomplete when basic_json instantiates its object_t
using move_values = std::integral_constant<bool, detail::conjunction<
detail::negation<std::is_nothrow_move_constructible<value_type>>,
std::is_copy_constructible<key_type>,
detail::is_default_constructible<mapped_type>,
std::is_nothrow_move_assignable<mapped_type>>::value>;
append_impl(move_values{}, std::forward<Args>(args)...);
}
template<typename... Args>
void append_impl(std::true_type /*unused*/, Args&& ... args)
{
if (this->size() < this->capacity())
{
Container::emplace_back(std::forward<Args>(args)...);
return;
}
// 1. May throw, but only changes tmp: copy the keys, value-initialize
// the values, and add the new element. The arguments may refer to
// elements of this container, so they are used before any value is
// moved out of it.
Container tmp(this->get_allocator()); // equal allocators, so swap() is valid
tmp.reserve((std::min)(this->max_size(), (std::max)(size_type{1}, 2 * this->size())));
for (const auto& element : *this)
{
tmp.emplace_back(std::piecewise_construct, std::forward_as_tuple(element.first), std::forward_as_tuple());
}
tmp.emplace_back(std::forward<Args>(args)...);
// 2. Cannot throw: move the values over and adopt the new storage.
auto it = tmp.begin();
for (auto& element : *this)
{
it->second = std::move(element.second);
++it;
}
Container::swap(tmp);
}
template<typename... Args>
void append_impl(std::false_type /*unused*/, Args&& ... args)
{
Container::emplace_back(std::forward<Args>(args)...);
}
JSON_NO_UNIQUE_ADDRESS key_compare m_compare = key_compare();
};
+33
View File
@@ -119,6 +119,39 @@ BENCHMARK_CAPTURE(ParseIndented, canada / 4, TEST_DATA_DIRECTORY "/nativej
BENCHMARK_CAPTURE(ParseIndented, citm_catalog / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json", 4);
BENCHMARK_CAPTURE(ParseIndented, twitter / 4, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json", 4);
//////////////////////////////////////////////////////////////////////////////
// parse JSON from string into an ordered_json
//
// Same as ParseString above, but with nlohmann::ordered_json, whose objects
// keep their members in a vector: the pair of rows shows what preserving the
// insertion order costs.
//////////////////////////////////////////////////////////////////////////////
static void ParseStringOrdered(benchmark::State& state, const char* filename)
{
std::ifstream f(filename);
std::string str((std::istreambuf_iterator<char>(f)), std::istreambuf_iterator<char>());
while (state.KeepRunning())
{
state.PauseTiming();
auto* j = new nlohmann::ordered_json();
state.ResumeTiming();
*j = nlohmann::ordered_json::parse(str);
state.PauseTiming();
delete j;
state.ResumeTiming();
}
state.SetBytesProcessed(state.iterations() * str.size());
}
BENCHMARK_CAPTURE(ParseStringOrdered, jeopardy, TEST_DATA_DIRECTORY "/jeopardy/jeopardy.json");
BENCHMARK_CAPTURE(ParseStringOrdered, canada, TEST_DATA_DIRECTORY "/nativejson-benchmark/canada.json");
BENCHMARK_CAPTURE(ParseStringOrdered, citm_catalog, TEST_DATA_DIRECTORY "/nativejson-benchmark/citm_catalog.json");
BENCHMARK_CAPTURE(ParseStringOrdered, twitter, TEST_DATA_DIRECTORY "/nativejson-benchmark/twitter.json");
//////////////////////////////////////////////////////////////////////////////
// serialize JSON
//////////////////////////////////////////////////////////////////////////////
-141
View File
@@ -11,12 +11,7 @@
#include <nlohmann/json.hpp>
using nlohmann::json;
#include <cmath>
#include <fstream>
#include <limits>
#include <map>
#include <string>
#include <vector>
#include "make_test_data_available.hpp"
TEST_CASE("Binary Formats" * doctest::skip())
@@ -229,139 +224,3 @@ TEST_CASE("Binary Formats" * doctest::skip())
CHECK((100.0 * double(ubjson_3_size) / double(json_size)) == Approx(89.450));
}
}
namespace
{
// the binary formats as function pointers for "Binary formats with narrow number types";
// named functions rather than lambdas, because clang 3.5 cannot convert a lambda
// to a function pointer in the braced initializer of the format table
using narrow_json = nlohmann::basic_json<std::map, std::vector, std::string, bool, std::int32_t, std::uint32_t, float>;
using bytes = std::vector<std::uint8_t>;
bytes encode_cbor(const json& j)
{
return json::to_cbor(j);
}
narrow_json decode_cbor(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_cbor(v, true, allow_exceptions);
}
bytes encode_msgpack(const json& j)
{
return json::to_msgpack(j);
}
narrow_json decode_msgpack(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_msgpack(v, true, allow_exceptions);
}
bytes encode_ubjson(const json& j)
{
return json::to_ubjson(j);
}
narrow_json decode_ubjson(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_ubjson(v, true, allow_exceptions);
}
bytes encode_bjdata(const json& j)
{
return json::to_bjdata(j);
}
narrow_json decode_bjdata(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_bjdata(v, true, allow_exceptions);
}
// BSON can only store numbers as object members
bytes encode_bson(const json& j)
{
return json::to_bson(json{{"a", j}});
}
narrow_json decode_bson(const bytes& v, bool allow_exceptions)
{
const auto result = narrow_json::from_bson(v, true, allow_exceptions);
return result.is_discarded() ? result : result.at("a");
}
bytes encode_bon8(const json& j)
{
return json::to_bon8(j);
}
narrow_json decode_bon8(const bytes& v, bool allow_exceptions)
{
return narrow_json::from_bon8(v, true, allow_exceptions);
}
} // namespace
TEST_CASE("Binary formats with narrow number types")
{
// Numbers that do not fit the number types are handled like the lexer
// handles them in JSON text: an integer that fits neither integer type is
// stored as a floating-point number, and a finite floating-point number
// that overflows number_float_t is rejected with out_of_range.406.
struct binary_format
{
const char* name;
bytes (*encode)(const json&);
narrow_json (*decode)(const bytes&, bool);
};
const std::vector<binary_format> formats =
{
{"CBOR", encode_cbor, decode_cbor},
{"MessagePack", encode_msgpack, decode_msgpack},
{"UBJSON", encode_ubjson, decode_ubjson},
{"BJData", encode_bjdata, decode_bjdata},
{"BSON", encode_bson, decode_bson},
{"BON8", encode_bon8, decode_bon8},
};
for (const auto& format : formats)
{
const std::string name = format.name;
INFO("format := ", name);
const auto roundtrip = [&format](const json & j)
{
return format.decode(format.encode(j), true);
};
// integers that fit keep their type
CHECK(roundtrip(json(-5)).is_number_integer());
CHECK(roundtrip(json(-5)).get<std::int32_t>() == -5);
CHECK(roundtrip(json(3000000000u)).is_number_unsigned());
CHECK(roundtrip(json(3000000000u)).get<std::uint32_t>() == 3000000000u);
// integers that fit neither integer type are stored as float
CHECK(roundtrip(json(5000000000u)).is_number_float());
CHECK(roundtrip(json(5000000000u)).get<float>() == 5000000000.0f);
if (name != "BON8") // BON8 cannot encode integers above INT64_MAX
{
CHECK(roundtrip(json(10000000000000000000u)).is_number_float());
CHECK(roundtrip(json(10000000000000000000u)).get<float>() == 10000000000000000000.0f);
}
CHECK(roundtrip(json(-3000000000LL)).is_number_float());
CHECK(roundtrip(json(-3000000000LL)).get<float>() == -3000000000.0f);
CHECK(roundtrip(json(-5000000000LL)).is_number_float());
CHECK(roundtrip(json(-5000000000LL)).get<float>() == -5000000000.0f);
// floating-point numbers that fit
CHECK(roundtrip(json(1.5)).get<float>() == 1.5f);
const auto just_above_max = std::nextafter(static_cast<double>((std::numeric_limits<float>::max)()),
std::numeric_limits<double>::infinity());
CHECK(roundtrip(json(just_above_max)).get<float>() == (std::numeric_limits<float>::max)());
// infinity and NaN are passed on
CHECK(std::isinf(roundtrip(json(std::numeric_limits<double>::infinity())).get<float>()));
CHECK(std::isnan(roundtrip(json(std::numeric_limits<double>::quiet_NaN())).get<float>()));
// finite floating-point numbers that overflow number_float_t are rejected
const std::string message = "[json.exception.out_of_range.406] syntax error while parsing " + name
+ " value: number overflow";
CHECK_THROWS_WITH_AS(roundtrip(json(1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK_THROWS_WITH_AS(roundtrip(json(-1e300)), message.c_str(), narrow_json::out_of_range&);
CHECK(format.decode(format.encode(json(1e300)), false).is_discarded());
}
}
+14 -16
View File
@@ -3187,8 +3187,7 @@ TEST_CASE("Tagged values")
// CBOR encodes negative integers as: result = -1 - n
// For type 0x3B, n is an 8-byte uint64_t. Valid range for n with
// the default int64_t is [0, INT64_MAX], producing results in [INT64_MIN, -1].
// When n > INT64_MAX, the result exceeds int64_t range and is stored
// as a floating-point number, as the lexer does for JSON text.
// When n > INT64_MAX, the result exceeds int64_t range and is rejected.
SECTION("n = 0 is valid (result = -1)")
{
@@ -3209,34 +3208,33 @@ TEST_CASE("Tagged values")
CHECK(result.get<int64_t>() == (std::numeric_limits<int64_t>::min)());
}
SECTION("n = INT64_MAX + 1 is stored as float")
SECTION("n = INT64_MAX + 1 is rejected (overflow)")
{
// n = INT64_MAX + 1 (0x8000000000000000)
// result = -1 - n = -9223372036854775809, which exceeds int64_t range;
// the nearest double is -9223372036854775808.0
// result = -1 - n = -9223372036854775809, which exceeds int64_t range
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
const auto result = json::from_cbor(input);
CHECK(result.is_number_float());
CHECK(result.get<double>() == -9223372036854775808.0);
CHECK(result == json::parse("-9223372036854775809"));
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
json::parse_error);
}
SECTION("n = UINT64_MAX is stored as float")
SECTION("n = UINT64_MAX is rejected (overflow)")
{
// n = UINT64_MAX (0xFFFFFFFFFFFFFFFF)
// result = -1 - n = -18446744073709551616, which exceeds int64_t range
const std::vector<uint8_t> input = {0x3B, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
const auto result = json::from_cbor(input);
CHECK(result.is_number_float());
CHECK(result.get<double>() == -18446744073709551616.0);
CHECK(result == json::parse("-18446744073709551616"));
json _;
CHECK_THROWS_WITH_AS(_ = json::from_cbor(input),
"[json.exception.parse_error.112] parse error at byte 9: syntax error while parsing CBOR value: negative integer overflow",
json::parse_error);
}
SECTION("overflow with allow_exceptions=false is not an error")
SECTION("overflow with allow_exceptions=false returns discarded")
{
const std::vector<uint8_t> input = {0x3B, 0x80, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};
const auto result = json::from_cbor(input, true, false);
CHECK(result.is_number_float());
CHECK(result.is_discarded());
}
}
+18
View File
@@ -46,6 +46,24 @@ TEST_CASE("Tests with disabled exceptions")
CHECK(*sax_no_exception::error_string == "[json.exception.parse_error.101] parse error at line 1, column 1: syntax error while parsing value - invalid literal; last read: 'x'");
delete sax_no_exception::error_string; // NOLINT(cppcoreguidelines-owning-memory)
}
SECTION("growing an ordered_json object")
{
auto j = nlohmann::ordered_json::object();
for (int i = 0; i < 100; ++i)
{
j[std::to_string(i)] = {{"nested", i}};
}
CHECK(j.size() == 100);
int i = 0;
for (const auto& element : j.items())
{
CHECK(element.key() == std::to_string(i));
CHECK(element.value()["nested"] == i);
++i;
}
}
}
DOCTEST_GCC_SUPPRESS_WARNING_POP
+348
View File
@@ -11,6 +11,87 @@
#include <nlohmann/json.hpp>
using nlohmann::ordered_map;
#include <stdexcept>
#include <string>
#include <type_traits>
#include <utility>
#include <vector>
namespace
{
// number of copies made of counted values
int value_copies = 0;
// a mapped type that counts its copies; moving from it leaves -1 behind
struct counted // NOLINT(cppcoreguidelines-special-member-functions,hicpp-special-member-functions)
{
int payload = 0;
counted() = default;
explicit counted(int p) noexcept : payload(p) {}
counted(const counted& other) : payload(other.payload)
{
++value_copies;
}
counted(counted&& other) noexcept : payload(other.payload)
{
other.payload = -1;
}
counted& operator=(const counted&) = delete;
counted& operator=(counted&& other) noexcept
{
payload = other.payload;
other.payload = -1;
return *this;
}
};
#if !defined(JSON_NOEXCEPTION)
// number of throwing_key copies that still succeed; the next one throws
// (a negative value means that copies never throw)
int key_copies_until_throw = -1;
// a key type whose copy constructor can be made to throw
struct throwing_key // NOLINT(cppcoreguidelines-special-member-functions,hicpp-special-member-functions)
{
int id = 0;
explicit throwing_key(int i) noexcept : id(i) {}
throwing_key(const throwing_key& other) : id(other.id)
{
if (key_copies_until_throw == 0)
{
throw std::runtime_error("key copy failed");
}
if (key_copies_until_throw > 0)
{
--key_copies_until_throw;
}
}
throwing_key& operator=(const throwing_key&) = delete;
friend bool operator==(const throwing_key& lhs, const throwing_key& rhs) noexcept
{
return lhs.id == rhs.id;
}
};
#endif
// a mapped type that cannot be default-constructed
struct no_default
{
explicit no_default(int v) noexcept : value(v) {}
int value;
};
// ordered_json must keep moving its values when an object grows
using ordered_object_t = nlohmann::ordered_json::object_t;
static_assert(!std::is_nothrow_move_constructible<ordered_object_t::value_type>::value, "std::vector would move the elements itself");
static_assert(std::is_copy_constructible<ordered_object_t::key_type>::value, "keys must be copyable");
static_assert(std::is_default_constructible<ordered_object_t::mapped_type>::value, "values must be default-constructible");
static_assert(std::is_nothrow_move_assignable<ordered_object_t::mapped_type>::value, "values must be nothrow move-assignable");
} // namespace
TEST_CASE("ordered_map")
{
SECTION("constructor")
@@ -313,3 +394,270 @@ TEST_CASE("ordered_map")
}
}
}
TEST_CASE("ordered_map growth")
{
SECTION("values are moved, not copied, when the storage grows")
{
ordered_map<std::string, counted> om;
std::size_t growths = 0;
value_copies = 0;
// inserts 100 elements with the given function and counts the growths
const auto fill = [&om, &growths](void (*insert)(ordered_map<std::string, counted>&, int))
{
for (int i = 0; i < 100; ++i)
{
const auto old_capacity = om.capacity();
insert(om, i);
if (om.capacity() > old_capacity)
{
++growths;
}
}
};
// checks that the elements are in insertion order with their values
const auto check_contents = [&om]
{
CHECK(om.size() == 100);
int i = 0;
for (const auto& element : om)
{
CHECK(element.first == std::to_string(i));
CHECK(element.second.payload == i);
++i;
}
};
SECTION("emplace")
{
fill([](ordered_map<std::string, counted>& m, int i)
{
m.emplace(std::to_string(i), counted(i));
});
CHECK(growths >= 3);
CHECK(value_copies == 0);
check_contents();
}
SECTION("operator[]")
{
fill([](ordered_map<std::string, counted>& m, int i)
{
m[std::to_string(i)] = counted(i);
});
CHECK(growths >= 3);
CHECK(value_copies == 0);
check_contents();
}
SECTION("insert(value_type&&)")
{
fill([](ordered_map<std::string, counted>& m, int i)
{
m.insert({std::to_string(i), counted(i)});
});
CHECK(growths >= 3);
CHECK(value_copies == 0);
check_contents();
}
SECTION("insert(const value_type&)")
{
fill([](ordered_map<std::string, counted>& m, int i)
{
const std::pair<const std::string, counted> value(std::to_string(i), counted(i));
m.insert(value);
});
CHECK(growths >= 3);
// only the inserted values are copied
CHECK(value_copies == 100);
check_contents();
}
SECTION("insert(first, last)")
{
std::vector<std::pair<const std::string, counted>> values;
values.reserve(100);
for (int i = 0; i < 100; ++i)
{
values.emplace_back(std::to_string(i), counted(i));
}
value_copies = 0;
om.insert(values.cbegin(), values.cend());
// only the inserted values are copied
CHECK(value_copies == 100);
check_contents();
}
}
SECTION("elements keep their order and values over many growths")
{
ordered_map<std::string, counted> om;
for (int i = 0; i < 1000; ++i)
{
om.emplace(std::to_string(i), counted(i));
}
CHECK(om.size() == 1000);
int i = 0;
for (const auto& element : om)
{
CHECK(element.first == std::to_string(i));
CHECK(element.second.payload == i);
++i;
}
}
SECTION("arguments may refer to elements of the full container")
{
SECTION("moving a value out of the container")
{
ordered_map<std::string, counted> om;
om.reserve(4);
while (om.size() < om.capacity())
{
const auto i = static_cast<int>(om.size());
om.emplace(std::to_string(i), counted(i));
}
const auto size = om.size();
om.emplace("new", std::move(om.at("0")));
CHECK(om.size() == size + 1);
CHECK(om.at("new").payload == 0);
CHECK(om.at("0").payload == -1);
}
SECTION("using a value as key")
{
ordered_map<std::string, std::string> om;
om.reserve(4);
while (om.size() < om.capacity())
{
const auto i = std::to_string(om.size());
om.emplace("k" + i, "v" + i);
}
const auto size = om.size();
om.emplace(om.at("k0"), std::string("x"));
CHECK(om.size() == size + 1);
CHECK(om.at("k0") == "v0");
CHECK(om.at("v0") == "x");
}
SECTION("ordered_json")
{
auto j = nlohmann::ordered_json::object();
auto& object = j.get_ref<nlohmann::ordered_json::object_t&>();
object.reserve(4);
while (object.size() < object.capacity())
{
const auto i = std::to_string(object.size());
j[i] = "a value that is too long for the small string optimization " + i;
}
const auto size = j.size();
j.emplace("new", std::move(j["0"]));
CHECK(j.size() == size + 1);
CHECK(j["new"] == "a value that is too long for the small string optimization 0");
CHECK(j["0"].is_null());
}
}
#if !defined(JSON_NOEXCEPTION)
SECTION("the container is unchanged if growing it throws")
{
ordered_map<throwing_key, counted> om;
om.reserve(4);
while (om.size() < om.capacity())
{
const auto i = static_cast<int>(om.size());
om.emplace(throwing_key(i), counted(i));
}
const auto size = om.size();
const auto capacity = om.capacity();
// checks that the elements are unchanged
const auto check_unchanged = [&om, size, capacity]
{
CHECK(om.size() == size);
CHECK(om.capacity() == capacity);
int i = 0;
for (const auto& element : om)
{
CHECK(element.first.id == i);
CHECK(element.second.payload == i);
++i;
}
};
SECTION("emplace")
{
// growing copies the existing keys and then the new one; let each of these copies throw
for (std::size_t k = 0; k <= size; ++k)
{
counted value(100);
key_copies_until_throw = static_cast<int>(k);
CHECK_THROWS_AS(om.emplace(throwing_key(100), std::move(value)), std::runtime_error);
key_copies_until_throw = -1;
check_unchanged();
CHECK(value.payload == 100); // NOLINT(bugprone-use-after-move,hicpp-invalid-access-moved)
}
om.emplace(throwing_key(100), counted(100));
CHECK(om.size() == size + 1);
CHECK(om.capacity() > capacity);
CHECK(om.at(throwing_key(100)).payload == 100);
}
SECTION("insert(const value_type&)")
{
const std::pair<const throwing_key, counted> value(throwing_key(100), counted(100));
value_copies = 0;
key_copies_until_throw = static_cast<int>(size / 2);
CHECK_THROWS_AS(om.insert(value), std::runtime_error);
key_copies_until_throw = -1;
check_unchanged();
CHECK(value_copies == 0);
}
}
#endif
SECTION("elements that std::vector moves, or that cannot be moved back")
{
SECTION("nothrow move-constructible elements")
{
ordered_map<int, counted> om;
value_copies = 0;
for (int i = 0; i < 100; ++i)
{
om.emplace(i, counted(i));
}
CHECK(om.size() == 100);
CHECK(value_copies == 0);
}
SECTION("mapped type without default constructor")
{
ordered_map<std::string, no_default> om;
for (int i = 0; i < 100; ++i)
{
om.emplace(std::to_string(i), no_default(i));
}
CHECK(om.size() == 100);
int i = 0;
for (const auto& element : om)
{
CHECK(element.first == std::to_string(i));
CHECK(element.second.value == i);
++i;
}
}
}
}