Compare commits

..
Author SHA1 Message Date
Niels LohmannandClaude Opus 4.8 9ea2d28ef9 Reserve output capacity up front for binary serialization
The vector-returning to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson grew
the output buffer purely by geometric reallocation. Reserving an estimate
up front avoids the early reallocations, which is the dominant per-byte
cost for array/object-heavy output.

The estimate (binary_reserve_hint) is deliberately conservative and safe
against untrusted input: it consults only the top-level element count
(O(1), no walk of the DOM), guards the multiplication against overflow,
and clamps the result to a fixed 1 MiB ceiling, so a large or hostile DOM
can never force an oversized allocation here. The buffer still grows
geometrically past the hint, so an underestimate only costs a few later
reallocations; scalars/strings/binary are written in one shot and get no
hint. Reserving capacity does not change the bytes produced.

Throughput (g++/clang -O3, vs the previous commit):
  cbor int array     +10% / +13%
  cbor object array  +20% / +38%

Output is byte-for-byte identical to develop across the binary
differential corpus.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-22 11:01:03 +00:00
Niels LohmannandClaude Opus 4.8 ac6b505923 Encode big-endian numbers with a byte swap instead of std::reverse
write_number() reordered multi-byte numbers for the big-endian formats
(CBOR/MessagePack/UBJSON) with std::reverse over the byte array. GCC
lowered only some sizes to a bswap; clang kept a scalar byte shuffle
(0 bswap instructions in the CBOR number path). Replace the reverse with
size-dispatched __builtin_bswap16/32/64 helpers (portable shift fallback
for other compilers; std::reverse retained for exotic sizes such as a
long double number_float_t).

Codegen: the CBOR number path now emits bswap on both compilers
(gcc 2 -> 16, clang 0 -> 4). Output is byte-for-byte identical to the
previous implementation across the binary differential corpus.

Throughput (isolated vs the std::reverse version, best of 9):
  CBOR int64 array   gcc +7%   clang +10%
  CBOR uint16 array  gcc +27%  clang flat

Modest but consistent on number-dense encodings; negligible on
string/blob-heavy output, as expected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-22 10:20:33 +00:00
Niels LohmannandClaude Opus 4.8 7374730aed Fix CI failures from binary_writer output-sink change
Four CI jobs failed on the initial commit; all are addressed here without
changing any output (binary encodings remain byte-for-byte identical to
develop across the differential corpus):

1. ci_test_gcc / cuda (-Werror=duplicated-branches): for number_float_t ==
   float, static_cast<float>(n) is the identity, so write_compact_float's
   two branches are intentionally identical. Once the concrete vector sink
   is inlined, GCC constant-folds and diagnoses this (the type-erased path
   hid it behind a non-inlined virtual call). Silence -Wduplicated-branches
   for GCC (clang has no such warning) alongside the existing -Wfloat-equal
   pragma.

2. ci_static_analysis_clang (UBSan nonnull-attribute): binary_writer passes
   a null pointer with length 0 for empty strings/binary. output_vector_sink
   / output_adapter_sink declared write_characters JSON_HEDLEY_NON_NULL, so
   the sanitizer flagged the (harmless) zero-length call once the sink was
   called directly rather than through the attribute-free virtual base. Drop
   the attribute from both sinks, matching the pre-existing behavior.

3. ci_cpplint (build/include_what_you_use): output_adapter_sink uses
   std::move; add #include <utility>.

4. ci_cuda_example (nvcc 11.8): NVCC's front end rejects the default
   template argument on the binary_writer alias template. Revert the alias
   to its original single-parameter form (relying on binary_writer's own
   defaulted OutputSinkType) and spell out the full type in the vector-sink
   convenience functions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-22 08:47:24 +00:00
Niels LohmannandClaude Opus 4.8 ebb3abba41 Devirtualize binary_writer via a value-type output sink
to_cbor/to_msgpack/to_ubjson/to_bjdata/to_bson wrote every byte through
output_adapter_t, a shared_ptr<output_adapter_protocol> whose
write_character/write_characters are virtual. Unlike the lexer (templated
on a concrete InputAdapterType), the binary writer never got that
treatment, so binary output paid a vtable lookup per byte and a
make_shared per call.

Template binary_writer on an OutputSinkType and give it two concrete,
non-virtual sinks:

- output_vector_sink: appends straight into a std::vector (push_back /
  insert), used by the vector-returning to_* convenience functions. No
  vtable, no shared_ptr; the writes inline.
- output_adapter_sink: forwards to a type-erased output_adapter_t, so the
  existing to_*(j, output_adapter) overloads (streams, strings, custom
  adapters) keep working exactly as before -- one virtual call each,
  unchanged.

binary_writer keeps a convenience constructor taking output_adapter_t
(building the default output_adapter_sink), so the adapter overloads are
untouched; only the convenience functions switch to the vector sink. The
friend declaration and the basic_json binary_writer alias gain the new
(defaulted) template parameter.

Output is byte-for-byte identical: verified across ~3000 randomized
values plus curated edge cases (all scalar widths, strings with invalid
UTF-8, binary, nested arrays/objects) for CBOR, MessagePack, UBJSON (both
size/type settings), BJData, and BSON, plus the output_adapter path, in
C++11/17/20. Warning-clean under clang -Weverything and the gcc pedantic
set; clang-tidy clean on the changed headers; make check-amalgamation
clean.

Throughput (g++ -O3, vs develop): scalar-dense binary output such as
integer arrays ~1.4x; many small to_cbor calls ~1.04x (DOM traversal
bound); string/blob-heavy output unchanged (already bulk-bound). No
workload regressed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XAYM1qhSA2FDaDcGfPW3fG
Signed-off-by: Niels Lohmann <mail@nlohmann.me>
2026-07-20 23:27:51 +00:00
Angadi56andGitHub 3565f40229 reject negative UBJSON/BJData string length (#5284) 2026-07-20 20:32:38 +00:00
Angadi56andGitHub 1c5a953de5 reject unpaired UTF-16 surrogates in wide-string input (#5276) 2026-07-19 18:33:42 +02:00
dependabot[bot]andGitHub a03e65420c ⬆️ Bump the codeql-action group with 4 updates (#5275) 2026-07-18 13:59:20 +02:00
dependabot[bot]andGitHub d6ede37088 ⬆️ Bump actions/stale from 10.3.0 to 10.4.0 (#5277) 2026-07-18 13:58:45 +02:00
23 changed files with 970 additions and 2025 deletions
+3 -3
View File
@@ -38,14 +38,14 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/init@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
languages: c-cpp
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/autobuild@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/analyze@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
+1 -1
View File
@@ -43,6 +43,6 @@ jobs:
output: 'flawfinder_results.sarif'
- name: Upload analysis results to GitHub Security tab
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
sarif_file: ${{github.workspace}}/flawfinder_results.sarif
+1 -1
View File
@@ -76,6 +76,6 @@ jobs:
# Upload the results to GitHub's code scanning dashboard.
- name: "Upload to code-scanning"
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
sarif_file: results.sarif
+1 -1
View File
@@ -61,7 +61,7 @@ jobs:
# Upload SARIF file generated in previous step
- name: Upload SARIF file
uses: github/codeql-action/upload-sarif@54f647b7e1bb85c95cddabcd46b0c578ec92bc1a # v4.36.3
uses: github/codeql-action/upload-sarif@99df26d4f13ea111d4ec1a7dddef6063f76b97e9 # v4.37.0
with:
sarif_file: semgrep.sarif
if: always()
+1 -1
View File
@@ -20,7 +20,7 @@ jobs:
with:
egress-policy: audit
- uses: actions/stale@eb5cf3af3ac0a1aa4c9c45633dd1ae542a27a899 # v10.3.0
- uses: actions/stale@1e223db275d687790206a7acac4d1a11bd6fe629 # v10.4.0
with:
stale-issue-label: 'state: stale'
stale-pr-label: 'state: stale'
-1
View File
@@ -21,7 +21,6 @@ cc_library(
"include/nlohmann/adl_serializer.hpp",
"include/nlohmann/byte_container_with_subtype.hpp",
"include/nlohmann/detail/abi_macros.hpp",
"include/nlohmann/detail/conversions/from_chars.hpp",
"include/nlohmann/detail/conversions/from_json.hpp",
"include/nlohmann/detail/conversions/to_chars.hpp",
"include/nlohmann/detail/conversions/to_json.hpp",
@@ -1,636 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#pragma once
#include <nlohmann/detail/macro_scope.hpp>
#if JSON_HAS_FROM_CHARS
#include <algorithm> // find, find_if
#include <charconv> // from_chars, from_chars_result
#include <cmath> // isnan
#else
#include <cctype> // isspace
#include <cerrno> // ERANGE
#include <clocale> // localeconv, newlocale, freelocale
#include <cstdlib> // strto*
#include <memory> // unique_ptr
#include <type_traits> // enable_if, is_unsigned, remove_pointer
#include <utility> // declval
#ifdef __has_include
#if __has_include(<xlocale.h>)
#include <xlocale.h> // strto*_l, isspace_l
#endif
#endif
#endif
#include <limits> // numeric_limits<T>::[has_]infinity, quiet_NaN
#include <system_error> // errc
NLOHMANN_JSON_NAMESPACE_BEGIN
namespace detail
{
constexpr char get_locale_independent_decimal_point() noexcept
{
return '.';
}
#if JSON_HAS_FROM_CHARS
//////////////////////////////////////////////
// Delegate to std::from_chars if supported //
//////////////////////////////////////////////
/*!
@brief Number parsing implementation details
A traits class template that allows users of nlohmann::detail::from_chars to
query details of the used implementation.
@tparam T numeric type to parse, matching the `value` argument of from_chars
*/
template <typename T>
struct from_chars_traits
{
/// whether the from_chars in unaffected by the current global C locale
static constexpr bool is_locale_independent = true;
/// getter for the character used as decimal point by from_chars
static constexpr char(*get_decimal_point)() = &get_locale_independent_decimal_point;
};
using std::from_chars_result;
/*!
@brief Parse integer/floating-point numbers
Parses the string [first, last) into the numeric type T.
This implementation merely delegates to `std::from_chars` with one noteworthy
difference: Different C++ standard library implementations do not agree on
whether `value` should be set in the out-of-range case for floating-point
numbers. Presently, libstdc++ leaves `value` unchanged whereas libc++ and MS STL
set `value` to +/-0 or +/-inf which allows its users to distinguish between the
various cases of over- and underflow. The C++ proposal [P4168][1] details this
discrepancy and advocates to standardize the latter behavior.
Since we need to be able to distinguish absolute underflow (+/-0) from absolute
overflow (+/-inf), this implementation "corrects" the out-of-range behavior
accordingly by estimating the effective exponent through manual string parsing.
[1]: https://isocpp.org/files/papers/P4168R0.html
@tparam T numeric type to parse
@param[in] first points to the beginning of the parsed string (inclusive)
@param[in] end points to the end of the parsed string (exclusive)
@param[out] value references the variable to set the parsed number
@return structure consisting of `ptr` and `ec` such that:
- If the string starting at `first` matches a number, `ptr` points to
the address immediately after the last character that matches.
* If the matched number can be represented by `T`, `value` is set and
`ec` is value-initialized.
* If the matched number cannot be represented by `T`, `ec` is set to
`std::errc::result_out_of_range` and depending on the type of `T`:
+ `value` is left unchanged for integral `T`,
+ `value` is set to +/-0 or +/-inf for floating-point `T`.
- If the string starting at `first` doesn't match a number, `ptr`
points to `first`, `ec` is `std::errc::invalid_argument`, and `value`
is left unchanged.
*/
template <typename T>
JSON_HEDLEY_NON_NULL(1, 2)
inline from_chars_result from_chars(const char* first, const char* last, T& value)
{
// No implementation of std::from_chars will set its value argument to NaN,
// so by using NaN as the initial value, we can tell if it changed.
T inner_value = std::numeric_limits<T>::quiet_NaN();
auto res = std::from_chars(first, last, inner_value);
if constexpr (std::numeric_limits<T>::has_infinity)
{
if (res.ec == std::errc::result_out_of_range)
{
if (std::isnan(inner_value)) // inner_value was not set
{
// [first, res.ptr) matches a valid floating-point number that's
// out-of-range and our std::from_chars did not set `value`, so we
// need to distinguish the type of under-/overflow ourselves.
const char* mantissa_begin = *first == '-' ? first + 1 : first;
const char* mantissa_end = std::find_if(mantissa_begin, res.ptr, [](char c)
{
return c == 'e' || c == 'E';
});
const char* decimal_point = std::find(mantissa_begin, mantissa_end, '.');
const char* significant_digit = std::find_if(mantissa_begin, mantissa_end, [](char d)
{
return d >= '1' && d <= '9';
});
// Position of the significant digit relative to the decimal gives
// the order of magnitude of the mantissa.
auto effective_exponent = static_cast<int>(decimal_point - significant_digit);
// If the number includes an explicit exponent, add that to the the
// total exponent.
if (mantissa_end != res.ptr)
{
const char* exponent_begin = *(mantissa_end + 1) == '+'
? mantissa_end + 2 // skip the exponent's + sign
: mantissa_end + 1;
int exponent = 0;
auto exp_res = std::from_chars(exponent_begin, res.ptr, exponent);
JSON_ASSERT(exp_res.ptr == res.ptr);
if (exp_res.ec == std::errc::result_out_of_range)
{
// NOTE: Parentheses around function names mitigate min/max macro collision on Windows
effective_exponent = *exponent_begin == '-'
? (std::numeric_limits<int>::min)()
: (std::numeric_limits<int>::max)();
}
else
{
JSON_ASSERT(exp_res.ec == std::errc{});
effective_exponent += exponent;
}
}
// Set `value` to the sign-correct non-finite value based on the
// effective exponent.
value = (*first == '-' ? -1 : 1) * (effective_exponent < 0
? T{}
: std::numeric_limits<T>::infinity());
}
else
{
value = inner_value;
}
}
}
if (res.ec == std::errc{})
{
value = inner_value;
}
return res;
}
#else
//////////////////////////////////////////////
// Fallback implementation using strto*[_l] //
//////////////////////////////////////////////
#if defined(_WIN32)
/*!
@brief Extended locale functions on Windows
A trait object which provides wrappers for isspace and the strto* family of
functions, based on the platform's support for extended locale APIs.
This is the Windows implementation. The primary class template definition
falls back to the standard locale-dependent functions, which is used on
MinGW, whose C runtime only provides an incomplete subset of the
`_<funcname>_l` family (e.g., missing `_strtof_l`/`_strtold_l` depending on
the toolchain version). A specialization for genuine MSVC (including
clang-cl, which mimics the MSVC ABI and CRT) is used instead, as the full
`_<funcname>_l` family has reliably been available there since at least
Visual Studio 2005.
*/
template <typename = const char*, typename = float>
struct extended_locale_traits
{
static constexpr bool is_locale_independent = false;
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept
{
const auto* loc = localeconv();
JSON_ASSERT(loc != nullptr);
return (loc->decimal_point == nullptr) ? '.' : *(loc->decimal_point);
}
static int isspace(int c) noexcept
{
return std::isspace(c);
}
// NOLINTNEXTLINE(runtime/int)
static long long strtoll(const char* str, char** endptr) noexcept
{
return std::strtoll(str, endptr, 10);
}
// NOLINTNEXTLINE(runtime/int)
static unsigned long long strtoull(const char* str, char** endptr) noexcept
{
return std::strtoull(str, endptr, 10);
}
static constexpr float(&strtof)(const char*, char**) = std::strtof;
static constexpr double(&strtod)(const char*, char**) = std::strtod;
static constexpr long double(&strtold)(const char*, char**) = std::strtold;
};
#if defined(_MSC_VER)
/*!
@brief Get a pointer to a globally shared instance of the "C" locale on Windows
This function uses _create_locale/_free_locale to allocate a "C" locale instance
which can be passed to extended locale APIs. This instance is created on first
use and has static storage duration, such that it can be used for the duration
of the program.
*/
inline _locale_t get_c_locale_t() noexcept
{
struct LocaleTDeleter
{
void operator()(_locale_t loc) noexcept
{
::_free_locale(loc);
}
};
using LocaleUPtr = std::unique_ptr<std::remove_pointer<_locale_t>::type, LocaleTDeleter>;
static const LocaleUPtr c_locale = LocaleUPtr{::_create_locale(LC_ALL, "C"), LocaleTDeleter{}};
return c_locale.get();
}
template <typename T>
struct extended_locale_traits<T, float>
{
static constexpr bool is_locale_independent = true;
static constexpr char(&get_decimal_point)() = get_locale_independent_decimal_point;
static int isspace(int c) noexcept
{
return _isspace_l(c, get_c_locale_t());
}
// NOLINTNEXTLINE(runtime/int)
static long long strtoll(const char* str, char** endptr) noexcept
{
return _strtoll_l(str, endptr, 10, get_c_locale_t());
}
// NOLINTNEXTLINE(runtime/int)
static unsigned long long strtoull(const char* str, char** endptr) noexcept
{
return _strtoull_l(str, endptr, 10, get_c_locale_t());
}
static float strtof(T str, char** str_end)
{
return _strtof_l(str, str_end, get_c_locale_t());
}
static double strtod(T str, char** str_end)
{
return _strtod_l(str, str_end, get_c_locale_t());
}
static long double strtold(T str, char** str_end)
{
return _strtold_l(str, str_end, get_c_locale_t());
}
};
#endif
#else
/*!
@brief Get a pointer to a globally shared instance of the "C" locale on POSIX
This function uses newlocale/freelocale to allocate a "C" locale instance
which can be passed to extended locale APIs. This instance is created on first
use and has static storage duration, such that it can be used for the duration
of the program.
newlocale/freelocale themselves are in POSIX and are available independent of
the platform's support for extended locale APIs.
*/
inline locale_t get_c_locale_t() noexcept
{
struct LocaleTDeleter
{
void operator()(locale_t loc) noexcept
{
::freelocale(loc);
}
};
using LocaleUPtr = std::unique_ptr<std::remove_pointer<locale_t>::type, LocaleTDeleter>;
#if defined(__clang__)
#pragma clang diagnostic push
#pragma clang diagnostic ignored "-Wexit-time-destructors"
#endif
static const LocaleUPtr c_locale = LocaleUPtr {::newlocale(LC_ALL_MASK, "C", nullptr), LocaleTDeleter{}};
#if defined(__clang__)
#pragma clang diagnostic pop
#endif
return c_locale.get();
}
/*!
@brief Extended locale functions on POSIX
A trait object which provides wrappers for isspace and the strto* family of
functions, based on the platform's support for extended locale APIs.
This is the POSIX implementation where extended locale support might not be
available. The primary class template definition falls back to the standard
locale-dependent functions. A partial specialization uses SFINAE to detect the
availability of `strtof_l` (as a proxy for the whole `<funcname>_l` family) and
provides access to locale-independent functions.
*/
template <typename = const char*, typename = float>
struct extended_locale_traits
{
static constexpr bool is_locale_independent = false;
/*!
@brief Query the decimal point used by the current C locale
Note that calling this function while switching the global C locale from
another thread is undefined behavior.
*/
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept
{
const auto* loc = localeconv();
JSON_ASSERT(loc != nullptr);
return (loc->decimal_point == nullptr) ? '.' : *(loc->decimal_point);
}
static int isspace(int c) noexcept
{
return std::isspace(c);
}
// NOLINTNEXTLINE(runtime/int)
static long long strtoll(const char* str, char** endptr) noexcept
{
return std::strtoll(str, endptr, 10);
}
// NOLINTNEXTLINE(runtime/int)
static unsigned long long strtoull(const char* str, char** endptr) noexcept
{
return std::strtoull(str, endptr, 10);
}
static constexpr float(&strtof)(const char*, char**) = std::strtof;
static constexpr double(&strtod)(const char*, char**) = std::strtod;
static constexpr long double(&strtold)(const char*, char**) = std::strtold;
};
template <typename T>
struct extended_locale_traits<T, decltype(strtof_l(
std::declval<T>(),
std::declval<char**>(),
std::declval<locale_t>()))>
{
static constexpr bool is_locale_independent = true;
static constexpr char(&get_decimal_point)() = get_locale_independent_decimal_point;
static int isspace(int c) noexcept
{
return isspace_l(c, get_c_locale_t());
}
// NOLINTNEXTLINE(runtime/int)
static long long strtoll(const char* str, char** endptr) noexcept
{
return ::strtoll_l(str, endptr, 10, get_c_locale_t());
}
// NOLINTNEXTLINE(runtime/int)
static unsigned long long strtoull(const char* str, char** endptr) noexcept
{
return ::strtoull_l(str, endptr, 10, get_c_locale_t());
}
static float strtof(T str, char** str_end)
{
return ::strtof_l(str, str_end, get_c_locale_t());
}
static double strtod(T str, char** str_end)
{
return ::strtod_l(str, str_end, get_c_locale_t());
}
static long double strtold(T str, char** str_end)
{
return ::strtold_l(str, str_end, get_c_locale_t());
}
};
#endif
/*!
@brief Number parsing implementation details
A traits class template that allows users of nlohmann::detail::from_chars to
query details of the used implementation.
Each specialization provides the following static members:
- is_locale_independent: whether the from_chars in unaffected by the current
global C locale
- get_decimal_point: getter for the character used as decimal point by from_chars
- strto: dispatches to the C standard library strto* function for type T, or to
the corresponding extended locale function, if available.
@tparam T numeric type to parse, matching the `value` argument of from_chars
*/
template <typename T>
struct from_chars_traits;
template <>
struct from_chars_traits<long long> // NOLINT(runtime/int)
{
static constexpr bool is_locale_independent = extended_locale_traits<>::is_locale_independent;
static char get_decimal_point()
{
return extended_locale_traits<>::get_decimal_point();
}
// NOLINTNEXTLINE(runtime/int)
static constexpr long long(*strto)(const char*, char**) = &extended_locale_traits<>::strtoll;
};
template <>
struct from_chars_traits<unsigned long long> // NOLINT(runtime/int)
{
static constexpr bool is_locale_independent = extended_locale_traits<>::is_locale_independent;
static char get_decimal_point()
{
return extended_locale_traits<>::get_decimal_point();
}
// NOLINTNEXTLINE(runtime/int)
static constexpr unsigned long long(*strto)(const char*, char**) = &extended_locale_traits<>::strtoull;
};
template <>
struct from_chars_traits<float>
{
static constexpr bool is_locale_independent = extended_locale_traits<>::is_locale_independent;
static char get_decimal_point()
{
return extended_locale_traits<>::get_decimal_point();
}
static constexpr float(*strto)(const char*, char**) = &extended_locale_traits<>::strtof;
};
template <>
struct from_chars_traits<double>
{
static constexpr bool is_locale_independent = extended_locale_traits<>::is_locale_independent;
static char get_decimal_point()
{
return extended_locale_traits<>::get_decimal_point();
}
static constexpr double(*strto)(const char*, char**) = &extended_locale_traits<>::strtod;
};
template <>
struct from_chars_traits<long double>
{
static constexpr bool is_locale_independent = extended_locale_traits<>::is_locale_independent;
static char get_decimal_point()
{
return extended_locale_traits<>::get_decimal_point();
}
static constexpr long double(*strto)(const char*, char**) = &extended_locale_traits<>::strtold;
};
template <typename T>
constexpr typename std::enable_if<std::numeric_limits<T>::has_infinity, bool>::type is_out_of_range_value(T value) noexcept
{
#ifdef __GNUC__
#pragma GCC diagnostic push
#pragma GCC diagnostic ignored "-Wfloat-equal"
#endif
return value == T {} || value == std::numeric_limits<T>::infinity() || value == -std::numeric_limits<T>::infinity();
#ifdef __GNUC__
#pragma GCC diagnostic pop
#endif
}
template <typename T>
constexpr typename std::enable_if < !std::numeric_limits<T>::has_infinity, bool >::type is_out_of_range_value(T value) noexcept
{
// NOTE: Parentheses around function names mitigate min/max macro collision on Windows
return value == (std::numeric_limits<T>::max)() || value == (std::numeric_limits<T>::min)();
}
struct from_chars_result
{
const char* ptr;
std::errc ec;
};
/*!
@brief Parse integer/floating-point numbers
Parses the string [first, last) into the numeric type T.
This implementation uses the C standard library functions strto*, or their
extended locale counterparts strto*_l, to emulate the behavior of
`std::from_chars` on platforms where it is not (fully) supported.
Regarding the under-/overflow behavior, this implementation adopts the behavior
proposed by [P4168][1] and implemented by libc++ and MS STL of setting `value`
to the corresponding non-finite floating-point value in case of an out-of-range
result.
[1]: https://isocpp.org/files/papers/P4168R0.html
@pre Unlike std::from_chars, this implementation requires the string to be
null-terminated, such that `last` dereferences into a NUL byte.
@tparam T numeric type to parse
@param[in] first points to the beginning of the parsed string (inclusive)
@param[in] end points to the end of the parsed string (exclusive)
@param[out] value references the variable to set the parsed number
@return structure consisting of `ptr` and `ec` such that:
- If the string starting at `first` matches a number, `ptr` points to
the address immediately after the last character that matches.
* If the matched number can be represented by `T`, `value` is set and
`ec` is value-initialized.
* If the matched number cannot be represented by `T`, `ec` is set to
`std::errc::result_out_of_range` and depending on the type of `T`:
+ `value` is left unchanged for integral `T`,
+ `value` is set to +/-0 or +/-inf for floating-point `T`.
- If the string starting at `first` doesn't match a number, `ptr`
points to `first`, `ec` is `std::errc::invalid_argument`, and `value`
is left unchanged.
*/
template <typename T>
JSON_HEDLEY_NON_NULL(1, 2)
inline from_chars_result from_chars(const char* first, const char* last, T& value)
{
JSON_ASSERT(*last == '\0');
// Unlike strto*, from_chars does not accept leading whitespace or + signs
if (first == last || *first == '+'
|| (std::is_unsigned<T>::value && *first == '-')
|| extended_locale_traits<>::isspace(*first) != 0)
{
return {first, std::errc::invalid_argument};
}
errno = 0;
char* ptr = nullptr; // NOLINT(misc-const-correctness)
T result = from_chars_traits<T>::strto(first, &ptr);
if (ptr == first)
{
return {ptr, std::errc::invalid_argument};
}
// Upon under-/overflow, strto* returns a marginal value and sets errno.
// Note that it is NOT sufficient to just check errno: strto* only clears
// errno if the parsed string actually parses into 0/[U]LLONG_MIN/-_MAX
// and this result was returned without indicating an under-/overflow.
// Otherwise, errno may not be relied upon to indicate the *absence* of
// an out-of-range error.
if (is_out_of_range_value(result) && errno == ERANGE)
{
if (std::numeric_limits<T>::has_infinity)
{
// ONLY for floating-point types, set `value` to the same non-finite
// result returned by strto* to allow users to distinguish between
// different types of under-/overflow.
value = result;
}
return {ptr, std::errc::result_out_of_range};
}
value = result;
return {ptr, std::errc{}};
}
#endif
} // namespace detail
NLOHMANN_JSON_NAMESPACE_END
@@ -1846,6 +1846,29 @@ class binary_reader
return get_ubjson_value(get_char ? get_ignore_noop() : current);
}
/*!
@brief reject a negative UBJSON/BJData string length
String and key lengths are written with signed integer markers (i, I, l,
L). A negative value is malformed; without this check get_string() would
silently treat it as an empty string and leave the following bytes to be
misread as the next value. This mirrors the non-negative check the
optimized-container count path already performs in get_ubjson_size_value.
@param[in] len the string length read from the input
@return whether the length is valid (non-negative)
*/
template<typename NumberType>
bool check_ubjson_string_length(const NumberType len)
{
if (JSON_HEDLEY_UNLIKELY(len < 0))
{
return sax->parse_error(chars_read, get_token_string(), parse_error::create(113, chars_read,
exception_message(input_format, "string length must not be negative", "string"), nullptr));
}
return true;
}
/*!
@brief reads a UBJSON string
@@ -1883,25 +1906,25 @@ class binary_reader
case 'i':
{
std::int8_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'I':
{
std::int16_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'l':
{
std::int32_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'L':
{
std::int64_t len{};
return get_number(input_format, len) && get_string(input_format, len, result);
return get_number(input_format, len) && check_ubjson_string_length(len) && get_string(input_format, len, result);
}
case 'u':
@@ -395,17 +395,30 @@ struct wide_string_input_helper<BaseInputAdapter, 2>
}
else
{
if (JSON_HEDLEY_UNLIKELY(!input.empty()))
// A supplementary code point is a high surrogate (0xD800..0xDBFF)
// followed by a low surrogate (0xDC00..0xDFFF). A lone low
// surrogate, a high surrogate at the end of the input, or a high
// surrogate followed by any other unit is malformed UTF-16. In
// that case the offending unit is passed through unchanged so the
// UTF-8 decoder rejects it, matching how \uXXXX surrogate escapes
// are handled in the lexer.
bool valid_pair = false;
if (wc <= 0xDBFF && JSON_HEDLEY_UNLIKELY(!input.empty()))
{
const auto wc2 = static_cast<unsigned int>(input.get_character());
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
if (0xDC00 <= wc2 && wc2 <= 0xDFFF)
{
const auto charcode = 0x10000u + (((static_cast<unsigned int>(wc) & 0x3FFu) << 10u) | (wc2 & 0x3FFu));
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(0xF0u | (charcode >> 18u));
utf8_bytes[1] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 12u) & 0x3Fu));
utf8_bytes[2] = static_cast<std::char_traits<char>::int_type>(0x80u | ((charcode >> 6u) & 0x3Fu));
utf8_bytes[3] = static_cast<std::char_traits<char>::int_type>(0x80u | (charcode & 0x3Fu));
utf8_bytes_filled = 4;
valid_pair = true;
}
}
else
if (!valid_pair)
{
utf8_bytes[0] = static_cast<std::char_traits<char>::int_type>(wc);
utf8_bytes_filled = 1;
+45 -16
View File
@@ -9,15 +9,15 @@
#pragma once
#include <array> // array
#include <clocale> // localeconv
#include <cstddef> // size_t
#include <cstdio> // snprintf
#include <cstdlib> // strtof, strtod, strtold, strtoll, strtoull
#include <initializer_list> // initializer_list
#include <string> // char_traits, string
#include <system_error> // errc
#include <utility> // move
#include <vector> // vector
#include <nlohmann/detail/conversions/from_chars.hpp>
#include <nlohmann/detail/input/input_adapters.hpp>
#include <nlohmann/detail/input/position_t.hpp>
#include <nlohmann/detail/macro_scope.hpp>
@@ -152,7 +152,7 @@ class lexer : public lexer_base<BasicJsonType>
explicit lexer(InputAdapterType&& adapter, bool ignore_comments_ = false) noexcept
: ia(std::move(adapter))
, ignore_comments(ignore_comments_)
, decimal_point_char(static_cast<char_int_type>(from_chars_traits<number_float_t>::get_decimal_point()))
, decimal_point_char(static_cast<char_int_type>(get_decimal_point()))
{}
// deleted because of pointer members
@@ -163,6 +163,19 @@ class lexer : public lexer_base<BasicJsonType>
~lexer() = default;
private:
/////////////////////
// locales
/////////////////////
/// return the locale-dependent decimal point
JSON_HEDLEY_PURE
static char get_decimal_point() noexcept
{
const auto* loc = localeconv();
JSON_ASSERT(loc != nullptr);
return (loc->decimal_point == nullptr) ? '.' : *(loc->decimal_point);
}
/////////////////////
// scan functions
/////////////////////
@@ -928,6 +941,24 @@ class lexer : public lexer_base<BasicJsonType>
}
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(float& f, const char* str, char** endptr) noexcept
{
f = std::strtof(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(double& f, const char* str, char** endptr) noexcept
{
f = std::strtod(str, endptr);
}
JSON_HEDLEY_NON_NULL(2)
static void strtof(long double& f, const char* str, char** endptr) noexcept
{
f = std::strtold(str, endptr);
}
/*!
@brief scan a number literal
@@ -1248,18 +1279,19 @@ scan_number_done:
// we are done scanning a number)
unget();
char* endptr = nullptr; // NOLINT(misc-const-correctness,cppcoreguidelines-pro-type-vararg,hicpp-vararg)
errno = 0;
// try to parse integers first and fall back to floats
if (number_type == token_type::value_unsigned)
{
unsigned long long x{}; // NOLINT(runtime/int)
const auto res = ::nlohmann::detail::from_chars(token_buffer.data(), token_buffer.data() + token_buffer.size(), x);
const auto x = std::strtoull(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(res.ptr == token_buffer.data() + token_buffer.size());
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (res.ec != std::errc::result_out_of_range)
if (errno != ERANGE)
{
JSON_ASSERT(res.ec == std::errc{});
value_unsigned = static_cast<number_unsigned_t>(x);
if (value_unsigned == x)
{
@@ -1269,15 +1301,13 @@ scan_number_done:
}
else if (number_type == token_type::value_integer)
{
long long x{}; // NOLINT(runtime/int)
const auto res = ::nlohmann::detail::from_chars(token_buffer.data(), token_buffer.data() + token_buffer.size(), x);
const auto x = std::strtoll(token_buffer.data(), &endptr, 10);
// we checked the number format before
JSON_ASSERT(res.ptr == token_buffer.data() + token_buffer.size());
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
if (res.ec != std::errc::result_out_of_range)
if (errno != ERANGE)
{
JSON_ASSERT(res.ec == std::errc{});
value_integer = static_cast<number_integer_t>(x);
if (value_integer == x)
{
@@ -1288,11 +1318,10 @@ scan_number_done:
// this code is reached if we parse a floating-point number or if an
// integer conversion above failed
const auto res = ::nlohmann::detail::from_chars(token_buffer.data(), token_buffer.data() + token_buffer.size(), value_float);
strtof(value_float, token_buffer.data(), &endptr);
// we checked the number format before
JSON_ASSERT(res.ptr == token_buffer.data() + token_buffer.size());
JSON_ASSERT(res.ec != std::errc::invalid_argument);
JSON_ASSERT(endptr == token_buffer.data() + token_buffer.size());
return token_type::value_float;
}
-14
View File
@@ -124,20 +124,6 @@
#define JSON_HAS_FILESYSTEM 0
#endif
#ifndef JSON_HAS_FROM_CHARS
#if defined(JSON_HAS_CPP_17) && defined(__cpp_lib_to_chars)
// The std::from_chars<float> implementation in libstdc++ 11 gets some of the corner cases wrong;
// starting with libstdc++ 12, these are fixed by switching the implementation to fast_float.
#if defined(_GLIBCXX_RELEASE) && _GLIBCXX_RELEASE < 12
#define JSON_HAS_FROM_CHARS 0
#else
#define JSON_HAS_FROM_CHARS 1
#endif
#else
#define JSON_HAS_FROM_CHARS 0
#endif
#endif
#ifndef JSON_HAS_THREE_WAY_COMPARISON
#if defined(__cpp_impl_three_way_comparison) && __cpp_impl_three_way_comparison >= 201907L \
&& defined(__cpp_lib_three_way_comparison) && __cpp_lib_three_way_comparison >= 201907L
@@ -39,7 +39,6 @@
#undef JSON_HAS_CPP_26
#undef JSON_HAS_FILESYSTEM
#undef JSON_HAS_EXPERIMENTAL_FILESYSTEM
#undef JSON_HAS_FROM_CHARS
#undef JSON_HAS_THREE_WAY_COMPARISON
#undef JSON_HAS_RANGES
#undef JSON_HAS_STD_FORMAT
File diff suppressed because it is too large Load Diff
@@ -13,6 +13,7 @@
#include <iterator> // back_inserter
#include <memory> // shared_ptr, make_shared
#include <string> // basic_string
#include <utility> // move
#include <vector> // vector
#ifndef JSON_NO_IO
@@ -118,6 +119,72 @@ class output_string_adapter : public output_adapter_protocol<CharType>
StringType& str;
};
/// @brief non-virtual output sink writing into a std::vector
///
/// Unlike output_vector_adapter, this sink is not part of the virtual
/// output_adapter_protocol hierarchy: it is passed to binary_writer by value as
/// a template parameter, so write_character()/write_characters() are ordinary
/// (inlinable) calls with no vtable lookup and no shared_ptr. It is used for the
/// common `to_cbor`/`to_msgpack`/... into a std::vector.
template<typename CharType, typename AllocatorType = std::allocator<CharType>>
class output_vector_sink
{
public:
explicit output_vector_sink(std::vector<CharType, AllocatorType>& vec) noexcept
: v(vec)
{}
void write_character(CharType c)
{
v.push_back(c);
}
// no JSON_HEDLEY_NON_NULL here: binary_writer legitimately passes a null
// pointer with length 0 for empty strings/binary values. Appending an empty
// range is a no-op; the type-erased path tolerates this via the (unattributed)
// virtual base, and the concrete sink must do the same.
void write_characters(const CharType* s, std::size_t length)
{
v.insert(v.end(), s, s + length);
}
private:
std::vector<CharType, AllocatorType>& v;
};
/// @brief output sink forwarding to a type-erased output adapter
///
/// Wraps the polymorphic output_adapter_t so the same binary_writer template can
/// also target arbitrary adapters (output streams, strings, user-provided
/// adapters) via the `output_adapter`-based overloads. Each write still goes
/// through one virtual call, exactly as before; only the concrete sinks above
/// avoid it.
template<typename CharType>
class output_adapter_sink
{
public:
explicit output_adapter_sink(output_adapter_t<CharType> adapter)
: oa(std::move(adapter))
{
JSON_ASSERT(oa);
}
void write_character(CharType c)
{
oa->write_character(c);
}
// no JSON_HEDLEY_NON_NULL: forwards (null, 0) for empty payloads, exactly as
// the type-erased path already did before this sink existed
void write_characters(const CharType* s, std::size_t length)
{
oa->write_characters(s, length);
}
private:
output_adapter_t<CharType> oa = nullptr;
};
template<typename CharType, typename StringType = std::basic_string<CharType>>
class output_adapter
{
+16 -6
View File
@@ -140,7 +140,7 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
friend ::nlohmann::detail::serializer<basic_json>;
template<typename BasicJsonType>
friend class ::nlohmann::detail::iter_impl;
template<typename BasicJsonType, typename CharType>
template<typename BasicJsonType, typename CharType, typename OutputSinkType>
friend class ::nlohmann::detail::binary_writer;
template<typename BasicJsonType, typename InputType, typename SAX>
friend class ::nlohmann::detail::binary_reader;
@@ -4327,7 +4327,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static std::vector<std::uint8_t> to_cbor(const basic_json& j)
{
std::vector<std::uint8_t> result;
to_cbor(j, result);
result.reserve(detail::binary_reserve_hint(j));
detail::binary_writer<basic_json, std::uint8_t, detail::output_vector_sink<std::uint8_t>>(
detail::output_vector_sink<std::uint8_t>(result)).write_cbor(j);
return result;
}
@@ -4350,7 +4352,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static std::vector<std::uint8_t> to_msgpack(const basic_json& j)
{
std::vector<std::uint8_t> result;
to_msgpack(j, result);
result.reserve(detail::binary_reserve_hint(j));
detail::binary_writer<basic_json, std::uint8_t, detail::output_vector_sink<std::uint8_t>>(
detail::output_vector_sink<std::uint8_t>(result)).write_msgpack(j);
return result;
}
@@ -4375,7 +4379,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
const bool use_type = false)
{
std::vector<std::uint8_t> result;
to_ubjson(j, result, use_size, use_type);
result.reserve(detail::binary_reserve_hint(j));
detail::binary_writer<basic_json, std::uint8_t, detail::output_vector_sink<std::uint8_t>>(
detail::output_vector_sink<std::uint8_t>(result)).write_ubjson(j, use_size, use_type);
return result;
}
@@ -4403,7 +4409,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
const bjdata_version_t version = bjdata_version_t::draft2)
{
std::vector<std::uint8_t> result;
to_bjdata(j, result, use_size, use_type, version);
result.reserve(detail::binary_reserve_hint(j));
detail::binary_writer<basic_json, std::uint8_t, detail::output_vector_sink<std::uint8_t>>(
detail::output_vector_sink<std::uint8_t>(result)).write_ubjson(j, use_size, use_type, true, true, version);
return result;
}
@@ -4430,7 +4438,9 @@ class basic_json // NOLINT(cppcoreguidelines-special-member-functions,hicpp-spec
static std::vector<std::uint8_t> to_bson(const basic_json& j)
{
std::vector<std::uint8_t> result;
to_bson(j, result);
result.reserve(detail::binary_reserve_hint(j));
detail::binary_writer<basic_json, std::uint8_t, detail::output_vector_sink<std::uint8_t>>(
detail::output_vector_sink<std::uint8_t>(result)).write_bson(j);
return result;
}
File diff suppressed because it is too large Load Diff
-4
View File
@@ -121,10 +121,6 @@ json_test_set_test_options(test-disabled_exceptions
# raise timeout of expensive Unicode test
json_test_set_test_options(test-unicode4 TEST_PROPERTIES TIMEOUT 3000)
# link pthreads to tests that need it
find_package(Threads REQUIRED)
json_test_set_test_options(test-regression2 LINK_LIBRARIES Threads::Threads)
#############################################################################
# add unit tests
#############################################################################
+13
View File
@@ -2721,6 +2721,19 @@ TEST_CASE("BJData")
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing BJData string: expected length type specification (U, i, u, I, m, l, M, L); last byte: 0x31", json::parse_error&);
}
SECTION("negative length")
{
json _;
std::vector<uint8_t> const vi = {'S', 'i', 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vi), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing BJData string: string length must not be negative", json::parse_error&);
CHECK(json::from_bjdata(vi, true, false).is_discarded());
std::vector<uint8_t> const vl = {'S', 'l', 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_bjdata(vl), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing BJData string: string length must not be negative", json::parse_error&);
CHECK(json::from_bjdata(vl, true, false).is_discarded());
}
SECTION("parse bjdata markers in ubjson")
{
// create a single-character string for all number types
-9
View File
@@ -55,8 +55,6 @@ TEST_CASE("lexer class")
SECTION("numbers")
{
// Number parsing implementation uses std::from_chars where available,
// so run this test suite with JSON_HAS_CPP_17 as well.
CHECK((scan_string("0") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("1") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("2") == json::lexer::token_type::value_unsigned));
@@ -67,20 +65,13 @@ TEST_CASE("lexer class")
CHECK((scan_string("7") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("8") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("9") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("18446744073709551615") == json::lexer::token_type::value_unsigned));
CHECK((scan_string("-0") == json::lexer::token_type::value_integer));
CHECK((scan_string("-1") == json::lexer::token_type::value_integer));
CHECK((scan_string("-9223372036854775808") == json::lexer::token_type::value_integer));
CHECK((scan_string("1.1") == json::lexer::token_type::value_float));
CHECK((scan_string("-1.1") == json::lexer::token_type::value_float));
CHECK((scan_string("1E10") == json::lexer::token_type::value_float));
// out-of-range integers/floats are treated as value_float tokens
CHECK((scan_string("18446744073709551616") == json::lexer::token_type::value_float));
CHECK((scan_string("-9223372036854775809") == json::lexer::token_type::value_float));
CHECK((scan_string("1E400") == json::lexer::token_type::value_float));
}
SECTION("whitespace")
-278
View File
@@ -1,278 +0,0 @@
// __ _____ _____ _____
// __| | __| | | | JSON for Modern C++ (supporting code)
// | | |__ | | | | | | version 3.12.0
// |_____|_____|_____|_|___| https://github.com/nlohmann/json
//
// SPDX-FileCopyrightText: 2013-2026 Niels Lohmann <https://nlohmann.me>
// SPDX-License-Identifier: MIT
#include "doctest_compatibility.h"
#include <nlohmann/json.hpp>
#include <algorithm>
#include <cmath>
#include <cstddef>
#include <limits>
#include <string>
#include <system_error>
// nlohmann::details::from_chars is a wrapper around std::from_chars when it is
// available (__cpp_lib_to_chars is defined), otherwise our fallback
// implementation kicks in.
//
// Due to the incomplete implementation of this C++17 standard library feature
// (for example, lack of support for long double even in current libc++),
// JSON_HAS_CPP_17 does not guarantee that std::from_chars is available.
// However, the reverse holds: If the test is compiled in C++11 mode, the
// fallback implementation will be used for sure.
// By mentioning the JSON_HAS_CPP_17 macro in this here comment, the test will
// be compiled both using our fallback and, at least on some platforms,
// the C++17 standard library version, as a way of ensuring that the
// expectations formulated in this test align with `std::from_chars`.
namespace
{
struct init_val_t {};
constexpr init_val_t init_val{};
template <typename T>
void check_result(std::string& str, T& value, std::ptrdiff_t expected_ptr_offset, std::errc expected_ec)
{
// On platforms where neither `std::from_chars`, nor extended locale support (`strtof_l`) are
// available, the `strtof`-based implementation is in fact locale-dependent and might use a
// different decimal separator, so we need to adapt the test expectations accordingly.
std::replace(str.begin(), str.end(), '.', nlohmann::detail::from_chars_traits<T>::get_decimal_point());
auto res = nlohmann::detail::from_chars(str.data(), str.data() + str.size(), value);
CHECK_MESSAGE(res.ec == expected_ec, "Error code mismatch while parsing: \"", str,
"\": Actual: ", make_error_code(res.ec).message(),
"; expected: ", make_error_code(expected_ec).message());
CHECK_MESSAGE(res.ptr - str.data() == expected_ptr_offset, "Ptr offset mismatch while parsing: \"", str, "\"");
}
template <typename T>
typename std::enable_if<std::is_integral<T>::value, void>::type
check(std::string str, T expected_value, std::ptrdiff_t expected_ptr_offset, std::errc expected_ec)
{
T value = 42;
check_result(str, value, expected_ptr_offset, expected_ec);
CHECK_MESSAGE(value == expected_value, "while parsing: \"", str, "\"");
}
template <typename T>
typename std::enable_if<std::is_floating_point<T>::value, void>::type
check(std::string str, T expected_value, std::ptrdiff_t expected_ptr_offset, std::errc expected_ec)
{
{
T value = 42;
check_result(str, value, expected_ptr_offset, expected_ec);
CHECK_MESSAGE(value == expected_value, "while parsing: \"", str, "\"");
CHECK_MESSAGE(std::signbit(value) == std::signbit(expected_value), "while parsing: \"", str, "\"");
}
{
T value = std::numeric_limits<T>::quiet_NaN();
check_result(str, value, expected_ptr_offset, expected_ec);
CHECK_MESSAGE(value == expected_value, "while parsing: \"", str, "\"");
}
}
template <typename T>
typename std::enable_if<std::is_integral<T>::value, void>::type
check(std::string str, init_val_t /*unused*/, std::ptrdiff_t expected_ptr_offset, std::errc expected_ec)
{
T const init_value = 42;
T value = init_value;
check_result(str, value, expected_ptr_offset, expected_ec);
CHECK_MESSAGE(value == init_value, "while parsing: \"", str, "\"");
}
template <typename T>
typename std::enable_if<std::is_floating_point<T>::value, void>::type
check(std::string str, init_val_t /*unused*/, std::ptrdiff_t expected_ptr_offset, std::errc expected_ec)
{
{
T const init_value = 42;
T value = init_value;
check_result(str, value, expected_ptr_offset, expected_ec);
CHECK_MESSAGE(value == init_value, "while parsing: \"", str, "\"");
}
{
T value = std::numeric_limits<T>::quiet_NaN();
check_result(str, value, expected_ptr_offset, expected_ec);
CHECK_MESSAGE(doctest::IsNaN<T>(value), "while parsing: \"", str, "\"");
}
}
} // namespace
TEST_CASE("integral results consistent with std::from_chars")
{
SECTION("unsigned long long")
{
check<unsigned long long>("", init_val, 0, std::errc::invalid_argument);
check<unsigned long long>("0", 0ULL, 1, std::errc{});
check<unsigned long long>(" 123", init_val, 0, std::errc::invalid_argument);
check<unsigned long long>("123 ", 123ULL, 3, std::errc{});
check<unsigned long long>("+123", init_val, 0, std::errc::invalid_argument);
check<unsigned long long>("-123", init_val, 0, std::errc::invalid_argument);
check<unsigned long long>("123", 123ULL, 3, std::errc{});
check<unsigned long long>("123e10", 123ULL, 3, std::errc{});
check<unsigned long long>("18446744073709551615", 18446744073709551615ULL, 20, std::errc{});
check<unsigned long long>("18446744073709551616", init_val, 20, std::errc::result_out_of_range);
}
SECTION("long long")
{
check<long long>("", init_val, 0, std::errc::invalid_argument);
check<long long>("0", 0LL, 1, std::errc{});
check<long long>(" 123", init_val, 0, std::errc::invalid_argument);
check<long long>("123 ", 123LL, 3, std::errc{});
check<long long>("+123", init_val, 0, std::errc::invalid_argument);
check<long long>("-123", -123LL, 4, std::errc{});
check<long long>("123", 123LL, 3, std::errc{});
check<long long>("123e10", 123LL, 3, std::errc{});
check<long long>("9223372036854775807", 9223372036854775807LL, 19, std::errc{});
check<long long>("9223372036854775808", init_val, 19, std::errc::result_out_of_range);
check<long long>("-9223372036854775808", -9223372036854775807LL - 1, 20, std::errc{});
check<long long>("-9223372036854775809", init_val, 20, std::errc::result_out_of_range);
}
}
TEST_CASE("floating point results consistent with std::from_chars")
{
INFO("Using decimal point '",
std::string{1, nlohmann::detail::from_chars_traits<float>::get_decimal_point()},
"' to match our possibly-locale-dependent from_chars fallback implementation");
SECTION("single precision")
{
check<float>("", init_val, 0, std::errc::invalid_argument);
check<float>(" 123", init_val, 0, std::errc::invalid_argument);
check<float>("123 ", 123.0f, 3, std::errc{});
check<float>("+123", init_val, 0, std::errc::invalid_argument);
check<float>("-123", -123.0f, 4, std::errc{});
check<float>("123", 123.0f, 3, std::errc{});
check<float>("123e10", 123e10f, 6, std::errc{});
check<float>("123e+10", 123e10f, 7, std::errc{});
check<float>("123e-10", 123e-10f, 7, std::errc{});
check<float>("123.456", 123.456f, 7, std::errc{});
check<float>("123;456", 123.0f, 3, std::errc{});
check<float>("123.456 ", 123.456f, 7, std::errc{});
check<float>("123;456 ", 123.0f, 3, std::errc{});
check<float>("123456789.123456789", 123456789.123456789f, 19, std::errc{});
check<float>("1e40", std::numeric_limits<float>::infinity(), 4, std::errc::result_out_of_range);
check<float>("-1e40", -std::numeric_limits<float>::infinity(), 5, std::errc::result_out_of_range);
check<float>("2e308", std::numeric_limits<float>::infinity(), 5, std::errc::result_out_of_range);
check<float>("-2e308", -std::numeric_limits<float>::infinity(), 6, std::errc::result_out_of_range);
check<float>("123.456e-789", 0.0f, 12, std::errc::result_out_of_range);
check<float>("-123.456e-789", -0.0f, 13, std::errc::result_out_of_range);
check<float>("1e-45", 1e-45f, 5, std::errc{});
check<float>("1e-46", 0.0f, 5, std::errc::result_out_of_range);
check<float>("1E-45", 1e-45f, 5, std::errc{});
check<float>("1E-46", 0.0f, 5, std::errc::result_out_of_range);
check<float>("10e-46", 1e-45f, 6, std::errc{});
check<float>("0.1e-45", 0.0f, 7, std::errc::result_out_of_range);
check<float>("100000000000000000000000000000000000000", 1e38f, 39, std::errc{});
check<float>("100000000000000000000000000000000000000.0", 1e38f, 41, std::errc{});
check<float>("1000000000000000000000000000000000000000e-1", 1e38f, 43, std::errc{});
check<float>("0.0000000000000000000000000000000000000000000001e84", 1e38f, 51, std::errc{});
check<float>("0.000000000000000000000000000000000000000000001", 1e-45f, 47, std::errc{});
check<float>("00.000000000000000000000000000000000000000000001", 1e-45f, 48, std::errc{});
check<float>("0.0000000000000000000000000000000000000000000001e1", 1e-45f, 50, std::errc{});
check<float>("1000000000000000000000000000000000000000e-84", 1e-45f, 44, std::errc{});
check<float>("1000000000000000000000000000000000000000", std::numeric_limits<float>::infinity(), 40, std::errc::result_out_of_range);
check<float>("1000000000000000000000000000000000000000.0", std::numeric_limits<float>::infinity(), 42, std::errc::result_out_of_range);
check<float>("100000000000000000000000000000000000000e1", std::numeric_limits<float>::infinity(), 41, std::errc::result_out_of_range);
check<float>("0.0000000000000000000000000000000000000000000001e85", std::numeric_limits<float>::infinity(), 51, std::errc::result_out_of_range);
check<float>("1000000000000000000000000000000000000000e-85", 0.0f, 44, std::errc::result_out_of_range);
check<float>("0.0000000000000000000000000000000000000000000001", 0.0f, 48, std::errc::result_out_of_range);
check<float>("0.000000000000000000000000000000000000000000001e-1", 0.0f, 50, std::errc::result_out_of_range);
check<float>("1e-99999999999999999999", 0.0f, 23, std::errc::result_out_of_range);
check<float>("1e+99999999999999999999", std::numeric_limits<float>::infinity(), 23, std::errc::result_out_of_range);
check<float>("-1e-99999999999999999999", -0.0f, 24, std::errc::result_out_of_range);
check<float>("-1e+99999999999999999999", -std::numeric_limits<float>::infinity(), 24, std::errc::result_out_of_range);
}
SECTION("double precision")
{
check<double>("", init_val, 0, std::errc::invalid_argument);
check<double>(" 123", init_val, 0, std::errc::invalid_argument);
check<double>("123 ", 123.0, 3, std::errc{});
check<double>("+123", init_val, 0, std::errc::invalid_argument);
check<double>("-123", -123.0, 4, std::errc{});
check<double>("123", 123.0, 3, std::errc{});
check<double>("123e10", 123e10, 6, std::errc{});
check<double>("123e+10", 123e10, 7, std::errc{});
check<double>("123e-10", 123e-10, 7, std::errc{});
check<double>("123.456", 123.456, 7, std::errc{});
check<double>("123;456", 123.0, 3, std::errc{});
check<double>("123.456 ", 123.456, 7, std::errc{});
check<double>("123;456 ", 123.0, 3, std::errc{});
check<double>("123456789.123456789", 123456789.123456789, 19, std::errc{});
check<double>("1e40", 1e40, 4, std::errc{});
check<double>("-1e40", -1e40, 5, std::errc{});
check<double>("2e308", std::numeric_limits<double>::infinity(), 5, std::errc::result_out_of_range);
check<double>("-2e308", -std::numeric_limits<double>::infinity(), 6, std::errc::result_out_of_range);
check<double>("2e-324", 0.0, 6, std::errc::result_out_of_range);
check<double>("-2e-324", -0.0, 7, std::errc::result_out_of_range);
check<double>("123.456e-789", 0.0, 12, std::errc::result_out_of_range);
check<double>("-123.456e-789", -0.0, 13, std::errc::result_out_of_range);
check<double>("1e-99999999999999999999", 0.0, 23, std::errc::result_out_of_range);
check<double>("1e+99999999999999999999", std::numeric_limits<double>::infinity(), 23, std::errc::result_out_of_range);
check<double>("-1e-99999999999999999999", -0.0, 24, std::errc::result_out_of_range);
check<double>("-1e+99999999999999999999", -std::numeric_limits<double>::infinity(), 24, std::errc::result_out_of_range);
}
SECTION("long double precision")
{
check<long double>("", init_val, 0, std::errc::invalid_argument);
check<long double>(" 123", init_val, 0, std::errc::invalid_argument);
check<long double>("123 ", 123.0L, 3, std::errc{});
check<long double>("+123", init_val, 0, std::errc::invalid_argument);
check<long double>("-123", -123.0L, 4, std::errc{});
check<long double>("123", 123.0L, 3, std::errc{});
check<long double>("123e10", 123e10L, 6, std::errc{});
check<long double>("123e+10", 123e10L, 7, std::errc{});
check<long double>("123e-10", 123e-10L, 7, std::errc{});
check<long double>("123.456", 123.456L, 7, std::errc{});
check<long double>("123;456", 123.0L, 3, std::errc{});
check<long double>("123.456 ", 123.456L, 7, std::errc{});
check<long double>("123;456 ", 123.0L, 3, std::errc{});
check<long double>("123456789.123456789", 123456789.123456789L, 19, std::errc{});
check<long double>("1e40", 1e40L, 4, std::errc{});
check<long double>("-1e40", -1e40L, 5, std::errc{});
#if defined(_MSC_VER)
#pragma warning(push)
#pragma warning(disable: 4127)
#endif
if (sizeof(long double) > 8)
#if defined(_MSC_VER)
#pragma warning(pop)
#endif
{
// Right-hand-side is calculated to avoid warning about literal
// exceeding range on platforms where this branch is NOT taken.
check<long double>("2e308", 1e308L * 2, 5, std::errc{});
check<long double>("-2e308", -1e308L * 2, 6, std::errc{});
check<long double>("2e-324", 8e-324L / 4, 6, std::errc{});
check<long double>("-2e-324", -8e-324L / 4, 7, std::errc{});
}
else
{
check<long double>("2e308", std::numeric_limits<long double>::infinity(), 5, std::errc::result_out_of_range);
check<long double>("-2e308", -std::numeric_limits<long double>::infinity(), 6, std::errc::result_out_of_range);
check<long double>("2e-324", 0.0L, 6, std::errc::result_out_of_range);
check<long double>("-2e-324", -0.0L, 7, std::errc::result_out_of_range);
}
check<long double>("1e5000", std::numeric_limits<long double>::infinity(), 6, std::errc::result_out_of_range);
check<long double>("-1e5000", -std::numeric_limits<long double>::infinity(), 7, std::errc::result_out_of_range);
check<long double>("123.456e-7890", 0.0L, 13, std::errc::result_out_of_range);
check<long double>("-123.456e-7890", -0.0L, 14, std::errc::result_out_of_range);
check<long double>("1e-99999999999999999999", 0.0L, 23, std::errc::result_out_of_range);
check<long double>("1e+99999999999999999999", std::numeric_limits<long double>::infinity(), 23, std::errc::result_out_of_range);
check<long double>("-1e-99999999999999999999", -0.0L, 24, std::errc::result_out_of_range);
check<long double>("-1e+99999999999999999999", -std::numeric_limits<long double>::infinity(), 24, std::errc::result_out_of_range);
}
}
-41
View File
@@ -26,11 +26,8 @@ using ordered_json = nlohmann::ordered_json;
using namespace nlohmann::literals; // NOLINT(google-build-using-namespace)
#endif
#include <atomic>
#include <clocale>
#include <cstdio>
#include <list>
#include <thread>
#include <type_traits>
#include <utility>
@@ -1242,44 +1239,6 @@ TEST_CASE("regression tests 2")
CHECK(j == json({2, 4, 6}));
}
#endif
SECTION("issue #5198 - TOCTOU race between lexer construction and locale changes causes float truncation")
{
std::atomic<bool> stop_requested{};
std::thread switcher([&stop_requested]
{
bool german = true;
const std::string initial_locale = std::setlocale(LC_NUMERIC, nullptr);
while (!stop_requested)
{
const bool setlocale_res = std::setlocale(LC_NUMERIC, german ? "de_DE.UTF-8" : "C");
WARN_MESSAGE(setlocale_res,
"Setting locale failed, probably because de_DE.UTF-8 is not installed. "
"This test was skipped as it would be inconclusive.");
if (!setlocale_res)
{
stop_requested = true;
}
german = !german;
}
(void)std::setlocale(LC_NUMERIC, initial_locale.c_str()); // restore original locale
});
for (std::size_t i = 0; i < 10000 && !stop_requested; ++i)
{
const json j = json::parse("{\"val\": 99.123456789}");
const double parsed = j.value("val", 0.0);
if (!CHECK(parsed == 99.123456789))
{
break;
}
}
stop_requested = true;
switcher.join();
}
}
TEST_CASE_TEMPLATE("issue #4798 - nlohmann::json::to_msgpack() encode float NaN as double", T, double, float) // NOLINT(readability-math-missing-parentheses, bugprone-throwing-static-initialization)
+25
View File
@@ -1862,6 +1862,31 @@ TEST_CASE("UBJSON")
json _;
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(v), "[json.exception.parse_error.113] parse error at byte 2: syntax error while parsing UBJSON string: expected length type specification (U, i, I, l, L); last byte: 0x31", json::parse_error&);
}
SECTION("negative length")
{
json _;
std::vector<uint8_t> const vi = {'S', 'i', 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vi), "[json.exception.parse_error.113] parse error at byte 3: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vi, true, false).is_discarded());
std::vector<uint8_t> const vI = {'S', 'I', 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vI), "[json.exception.parse_error.113] parse error at byte 4: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vI, true, false).is_discarded());
std::vector<uint8_t> const vl = {'S', 'l', 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vl), "[json.exception.parse_error.113] parse error at byte 6: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vl, true, false).is_discarded());
std::vector<uint8_t> const vL = {'S', 'L', 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF, 0xFF};
CHECK_THROWS_WITH_AS(_ = json::from_ubjson(vL), "[json.exception.parse_error.113] parse error at byte 10: syntax error while parsing UBJSON string: string length must not be negative", json::parse_error&);
CHECK(json::from_ubjson(vL, true, false).is_discarded());
// a length of zero remains valid and yields an empty string
std::vector<uint8_t> const v0 = {'S', 'i', 0};
CHECK(json::from_ubjson(v0) == json(""));
}
}
SECTION("array")
+33 -1
View File
@@ -53,6 +53,27 @@ TEST_CASE("wide strings")
std::wstring const w = L"\"\xDBFF";
json _;
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// the exact message depends on the width of wchar_t: a 16-bit
// wchar_t passes the lone surrogate to the UTF-8 decoder unchanged
// (rejected as a single ill-formed byte at column 2), while a
// 32-bit wchar_t first encodes it as an ill-formed three-byte
// sequence (rejected one byte later, at column 3)
const char* const error_low_surrogate = sizeof(wchar_t) == 2
? "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'"
: "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xED\xB0'";
const char* const error_high_surrogate = sizeof(wchar_t) == 2
? "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'"
: "[json.exception.parse_error.101] parse error at line 1, column 3: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"\xED\xA0'";
// a lone low surrogate cannot start a pair
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xDC00), L'"'}), error_low_surrogate, json::parse_error&);
// a high surrogate followed by a non-low-surrogate unit is invalid
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xD800), L'a', L'"'}), error_high_surrogate, json::parse_error&);
// a lone low surrogate must not swallow the following unit: pairing
// it with any second unit would produce valid UTF-8, so the error
// has to report an ill-formed byte at the surrogate's own position
CHECK_THROWS_WITH_AS(_ = json::parse(std::wstring{L'"', static_cast<wchar_t>(0xDC00), L'a', L'"'}), error_low_surrogate, json::parse_error&);
}
}
@@ -68,11 +89,22 @@ TEST_CASE("wide strings")
SECTION("invalid std::u16string")
{
if (wstring_is_utf16())
if (u16string_is_utf16())
{
std::u16string const w = u"\"\xDBFF";
json _;
CHECK_THROWS_AS(_ = json::parse(w), json::parse_error&);
// a lone low surrogate cannot start a pair
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a high surrogate followed by a non-low-surrogate unit is invalid
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xD800, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a lone low surrogate must not swallow the following unit: pairing
// it with any second unit would produce valid UTF-8, so the error
// has to report an ill-formed byte at the surrogate's own position
CHECK_THROWS_WITH_AS(_ = json::parse(std::u16string{u'"', 0xDC00, u'a', u'"'}), "[json.exception.parse_error.101] parse error at line 1, column 2: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"<U+0000>'", json::parse_error&);
// a valid surrogate pair is still decoded (U+1F600)
CHECK(json::parse(std::u16string{u'"', 0xD83D, 0xDE00, u'"'}).get<std::string>() == "\xF0\x9F\x98\x80");
}
}